The gen-small EngineCore OOM (04:21 PT): our parked 3,582 MiB window cache left
no room for vLLM's runtime workspace. Seat-side fix, three controls:
- windowed path wraps every window in torch.cuda.empty_cache(), so the seat
returns to ~2,108 MiB rest after a 12-min file instead of parking at the
peak (measured: peak 3,028 MiB during, rest after, restarts=0);
- MEM_CAP_MIB=3840 hard set_per_process_memory_fraction: over-cap requests
answer 503 with the seat alive (proved at cap=2000), so the failure lands
on us, never on a neighbour;
- CUDA_GRAPHS=0: the graph decoder pins cache blocks that empty_cache must
free (illegal-memory-access wedge when both were on first try). Cost:
12-min file 3.0 s vs 1.2 s, short bins 35-62 ms vs 33-42 ms -- still 4-15x
under the sherpa seat.
Measured, not computed: gen-small moved ZERO from 36,116 MiB across three
realistic requests (1,351 in / ~180 out) -- its workspace lands at engine
init; the growth window is restart-relative, matching infra-ops's observation.
WINDOW_S is now a real compose tunable. README memory section rewritten.