# `[2026-09-15]` A client timeout SOMETIMES cancels a vLLM generation and sometimes does not — the boundary is unknown ⚠⚠ **DO NOT carry "a client-side timeout is not a cancellation" as a rule. It is FALSE as stated, and it was disproved by the peer who coined it, on our own seat, within the hour.** `tts-dev` orphaned six unbounded generations on `vllm-erp-seat` (fv-ml1 GPU 1) by firing `char-rp-fast` probes with no `max_tokens` and letting clients time out at 110 s / 115 s / 600 s. They wrote the lesson up, then **controlled their own detector and the POSITIVE CONTROL FAILED** — chasing it produced this, measured against the live seat: t+1.6s running=1 kv=0.4% request reaches the engine client gave up (urlopen timeout=2) t+3.1s running=1 kv=0.8% still generating t+7.8s running=0 kv=0.0% CANCELLED, unprompted, ~6s after the client left **A clean client abandon DOES propagate.** Yet six requests genuinely orphaned — I observed that independently. **So some abandons propagate and some do not, and nobody has isolated the boundary.** Unseparated candidates: SIGTERM'd process vs clean client-side timeout; multi-minute unbounded generation vs short one; several stacked at once. ⭐ **That unknown is the argument FOR a detector and AGAINST a rule — a rule needs the boundary, a detector just looks.** tts-dev holds a standing request: if we ever isolate what makes an abandon stick, tell them; it is the input that would let them build a real positive control (theirs is SYNTHETIC and their file says so in place — detection logic proven, reproduction of the underlying bug not). ⭐⭐ **THE DISCRIMINATOR, and it is the durable artifact of the day: a serving engine's KV cache CYCLES; an orphaned one only CLIMBS.** Request count and throughput are **ambiguous** between a loaded seat and a wedged one — I read `vllm-erp-seat` twice off those signals and called it healthy both times, correctly on the evidence (39 completions/hour, 210–290 tok/s, `Running: 3 / Waiting: 3`, KV cycling 70→99→70%). The traffic was genuinely real; it then *ended*, and what remained were orphans. The tell was `prompt throughput 0.0` sustained, `Waiting: 0`, and KV **monotonic** 87.4 → 87.9 → 88.4 → 88.9 → 89.4. Now implemented in `tts-stack tools/engine_guard.py --watch` (`db9d847`). vLLM serves `/metrics` **unauthenticated** on the seat ports, so `num_requests_running`, `num_requests_waiting` and `kv_cache_usage_perc` are directly pollable — no gateway, no auth. ⚠ Its `settle` defaults to 20 s so normal cancellation lag is not reported as a leak: a guard that cries wolf gets disabled, and then you are back to a docstring. ⚠ **A `max_tokens` ceiling would NOT have prevented this.** tts-dev's worst offender ran with `max_tokens=16384` **explicitly set**, hit it exactly, and returned 24,594 characters of whitespace wrapping a correct three-field answer. **A ceiling bounds how long you wait for the failure, not whether it happens.** Escalated to the operator anyway as a two-layer choice (gateway-side LiteLLM default — one blast radius, misses direct-to-seat callers; vs per-seat limits — catches everything, nine seats to touch); gateway first and measure what it breaks is the right order. Related: [[feedback_detector_after_reflex_beats_reminder_before]]. **Remediation**: `docker restart vllm-erp-seat` 23:36:31 UTC, healthy in ~1 min, GPU 1 100% / 275 W (at the cap) / 74°C → 0% / 4.8 W / 42°C. The five other tenants on that card (`vllm-reward`, `vllm-rerank-a3`, `vllm-embed`, `vllm-coder`, `vllm-meromero-rp`) were untouched. Restarted rather than waiting — they DO self-terminate at the context limit and one dropped off mid-diagnosis (6→5, KV 89.4→86.8) — because KV at 89% and climbing starts costing the co-tenants through preemption. ⚠ **Noticed in passing, unresolved: `vllm-erp-seat` and `vllm-meromero-rp` advertise the SAME `--served-model-name`** (`G4-MeroMero-26B-A4B-it-uncensored-heretic-NVFP4A16`). Fine if it is deliberate replication for throughput; it is also the exact shape that makes gateway routing ambiguous and "which seat served this?" unanswerable after the fact. Surfaced to the operator, not yet answered.