8842ffe1fe
Operator: "yes, pin the kv and take it to 320k." The real finding, which took three crashes and two failed attempts to reach: --gpu-memory-utilization does not bound actual usage. It sizes the KV calculation, but peak activation is measured at profiling time and real long-context work exceeds the profile. vLLM's own budget line showed mog-sec running 0.9 GiB over its 47.48 GiB reservation -- 26.44 consumed plus 3.53 peak activation plus 0.89 CUDAGraph plus 17.52 KV equals 48.38 -- and gen was over by 0.33 on the same card. That overage came out of the shared card's slack, which is what kept OOMing after the utilization drop. The fix is the one vLLM printed itself: --kv-cache-memory=17697765376, its own recommended figure to fit inside the requested budget. Same discipline erp-seat already uses, and for the same stated reason -- an explicit figure is reproducible where a ratio silently yields a different cache depending on what else is resident at start time. The KV pin and the context length are coupled. 16.48 GiB yields about 383,730 tokens, so a 393,216 max_model_len falls under the 1.0x floor and vLLM refuses to start rather than crashing later; pinning the KV while keeping 384k was never an available combination. 327,680 leaves 1.15x, up from 1.03x. Verified: the engine now logs "reserved 16.48 GiB memory for KV Cache as specified by kv_cache_memory_bytes config and skipped memory profiling", KV 375,901 tokens, GPU0 down to 90,561 MiB from 91,313, RestartCount 0, and both sec and sec-reasoning return 200 through the gateway. Also records the BabyBronte eyeball A/B, whose result is the operator's own: the voice transferred and the sense did not. Curly quotes went 1 of 18 to 18 of 18 and worksheet collapse 3 of 18 to 0 of 18 between arms. That voice is separable from coherence at 0.6B is the premise the lightweight-adapter regime rests on, so this is the informative outcome rather than a disappointing one. A corpus-prep defect surfaced with it: the tuned output is hard-wrapped at about 70 characters because the Gutenberg source kept its line breaks and the adapter learned the typography too. Cost: 320k of context instead of 420k, on a seat whose crashes happened at 151k.