Pin mog-sec's KV in bytes at 16.48 GiB and take it to 320k context

Operator: "yes, pin the kv and take it to 320k."

The real finding, which took three crashes and two failed attempts to reach:
--gpu-memory-utilization does not bound actual usage. It sizes the KV
calculation, but peak activation is measured at profiling time and real
long-context work exceeds the profile. vLLM's own budget line showed mog-sec
running 0.9 GiB over its 47.48 GiB reservation -- 26.44 consumed plus 3.53 peak
activation plus 0.89 CUDAGraph plus 17.52 KV equals 48.38 -- and gen was over by
0.33 on the same card. That overage came out of the shared card's slack, which is
what kept OOMing after the utilization drop.

The fix is the one vLLM printed itself: --kv-cache-memory=17697765376, its own
recommended figure to fit inside the requested budget. Same discipline erp-seat
already uses, and for the same stated reason -- an explicit figure is
reproducible where a ratio silently yields a different cache depending on what
else is resident at start time.

The KV pin and the context length are coupled. 16.48 GiB yields about 383,730
tokens, so a 393,216 max_model_len falls under the 1.0x floor and vLLM refuses to
start rather than crashing later; pinning the KV while keeping 384k was never an
available combination. 327,680 leaves 1.15x, up from 1.03x.

Verified: the engine now logs "reserved 16.48 GiB memory for KV Cache as
specified by kv_cache_memory_bytes config and skipped memory profiling", KV
375,901 tokens, GPU0 down to 90,561 MiB from 91,313, RestartCount 0, and both sec
and sec-reasoning return 200 through the gateway.

Also records the BabyBronte eyeball A/B, whose result is the operator's own: the
voice transferred and the sense did not. Curly quotes went 1 of 18 to 18 of 18
and worksheet collapse 3 of 18 to 0 of 18 between arms. That voice is separable
from coherence at 0.6B is the premise the lightweight-adapter regime rests on, so
this is the informative outcome rather than a disappointing one. A corpus-prep
defect surfaced with it: the tuned output is hard-wrapped at about 70 characters
because the Gutenberg source kept its line breaks and the adapter learned the
typography too.

Cost: 320k of context instead of 420k, on a seat whose crashes happened at 151k.
This commit is contained in:
vh
2026-09-10 15:19:47 -07:00
parent 77224619ee
commit 8842ffe1fe
2 changed files with 22 additions and 1 deletions
+19
View File
@@ -99,7 +99,26 @@ services:
- ${MOG_QUANT:-compressed-tensors}
- --gpu-memory-utilization
- ${MOG_GPU_MEM_UTIL:-0.44}
# ⚠ KV PINNED IN BYTES, added 2026-09-10 after the crash pair above. The
# utilization ratio does NOT bound actual usage -- it sizes the KV calculation,
# but peak activation is measured at profiling time and real long-context work
# exceeds the profile. vLLM's own budget line proved this seat was running 0.9 GiB
# OVER its 47.48 GiB reservation at util 0.50 (26.44 consumed + 3.53 peak activation
# + 0.89 CUDAGraph + 17.52 KV = 48.38), and `gen` was over by 0.33 on the same card.
# That overage came out of the shared card's slack, which is what kept OOMing.
#
# 17,697,765,376 B = 16.48 GiB is vLLM's OWN recommended figure from that line
# ("Replace gpu_memory_utilization config with --kv-cache-memory=17697765376 to fit
# into requested memory"), not a value anyone here invented. Same discipline as
# stacks/erp-seat: an explicit figure is reproducible, a ratio silently yields a
# different cache depending on what else is resident at start time.
- --kv-cache-memory
- ${MOG_KV_CACHE_MEMORY:-17697765376}
- --max-model-len
# ⚠ 320k, NOT 384k, and the two settings are coupled -- 16.48 GiB of KV yields about
# 383,730 tokens, so a 393,216 max_model_len falls under the 1.0x floor and vLLM
# refuses to START rather than crashing later. Pinning the KV and keeping 384k was
# never an available combination. 327,680 leaves ~1.17x.
- ${MOG_MAX_MODEL_LEN:-262144}
- --max-num-seqs
- ${MOG_MAX_NUM_SEQS:-16}