4.2 KiB
[2026-09-15] A client timeout SOMETIMES cancels a vLLM generation and sometimes does not — the boundary is unknown
⚠⚠ DO NOT carry "a client-side timeout is not a cancellation" as a rule. It is FALSE as
stated, and it was disproved by the peer who coined it, on our own seat, within the hour.
tts-dev orphaned six unbounded generations on vllm-erp-seat (fv-ml1 GPU 1) by firing
char-rp-fast probes with no max_tokens and letting clients time out at 110 s / 115 s /
600 s. They wrote the lesson up, then controlled their own detector and the POSITIVE
CONTROL FAILED — chasing it produced this, measured against the live seat:
t+1.6s running=1 kv=0.4% request reaches the engine
client gave up (urlopen timeout=2)
t+3.1s running=1 kv=0.8% still generating
t+7.8s running=0 kv=0.0% CANCELLED, unprompted, ~6s after the client left
A clean client abandon DOES propagate. Yet six requests genuinely orphaned — I observed that independently. So some abandons propagate and some do not, and nobody has isolated the boundary. Unseparated candidates: SIGTERM'd process vs clean client-side timeout; multi-minute unbounded generation vs short one; several stacked at once. ⭐ That unknown is the argument FOR a detector and AGAINST a rule — a rule needs the boundary, a detector just looks. tts-dev holds a standing request: if we ever isolate what makes an abandon stick, tell them; it is the input that would let them build a real positive control (theirs is SYNTHETIC and their file says so in place — detection logic proven, reproduction of the underlying bug not).
⭐⭐ THE DISCRIMINATOR, and it is the durable artifact of the day: a serving engine's KV
cache CYCLES; an orphaned one only CLIMBS. Request count and throughput are ambiguous
between a loaded seat and a wedged one — I read vllm-erp-seat twice off those signals and
called it healthy both times, correctly on the evidence (39 completions/hour, 210–290 tok/s,
Running: 3 / Waiting: 3, KV cycling 70→99→70%). The traffic was genuinely real; it then
ended, and what remained were orphans. The tell was prompt throughput 0.0 sustained,
Waiting: 0, and KV monotonic 87.4 → 87.9 → 88.4 → 88.9 → 89.4. Now implemented in
tts-stack tools/engine_guard.py --watch (db9d847). vLLM serves /metrics
unauthenticated on the seat ports, so num_requests_running, num_requests_waiting and
kv_cache_usage_perc are directly pollable — no gateway, no auth. ⚠ Its settle defaults
to 20 s so normal cancellation lag is not reported as a leak: a guard that cries wolf gets
disabled, and then you are back to a docstring.
⚠ A max_tokens ceiling would NOT have prevented this. tts-dev's worst offender ran
with max_tokens=16384 explicitly set, hit it exactly, and returned 24,594 characters
of whitespace wrapping a correct three-field answer. A ceiling bounds how long you wait
for the failure, not whether it happens. Escalated to the operator anyway as a two-layer
choice (gateway-side LiteLLM default — one blast radius, misses direct-to-seat callers;
vs per-seat limits — catches everything, nine seats to touch); gateway first and measure
what it breaks is the right order. Related: feedback_detector_after_reflex_beats_reminder_before.
Remediation: docker restart vllm-erp-seat 23:36:31 UTC, healthy in ~1 min, GPU 1
100% / 275 W (at the cap) / 74°C → 0% / 4.8 W / 42°C. The five other tenants on that card
(vllm-reward, vllm-rerank-a3, vllm-embed, vllm-coder, vllm-meromero-rp) were
untouched. Restarted rather than waiting — they DO self-terminate at the context limit and
one dropped off mid-diagnosis (6→5, KV 89.4→86.8) — because KV at 89% and climbing starts
costing the co-tenants through preemption.
⚠ Noticed in passing, unresolved: vllm-erp-seat and vllm-meromero-rp advertise the
SAME --served-model-name (G4-MeroMero-26B-A4B-it-uncensored-heretic-NVFP4A16). Fine
if it is deliberate replication for throughput; it is also the exact shape that makes
gateway routing ambiguous and "which seat served this?" unanswerable after the fact.
Surfaced to the operator, not yet answered.