Give GPU0 real headroom: mog-sec to util 0.50 and 384k context
sec/sec-reasoning crash-bounced twice in ten minutes, not once in thirteen days. My earlier read of "rare, not chronic" came off a RestartCount of 1 and was wrong; the operator pushed back and the second and third failures arrived while that recommendation was still on screen. The memory entry making that call is replaced rather than left standing. Cause is unchanged from the first diagnosis: mog-sec at 0.52 plus gen at 0.38 reserve 0.90 of the card, leaving about 4.6 GiB, and vLLM's utilization figure covers weights and the KV pool but not transient activation memory. A request about 151,700 tokens deep scheduling a further 15,700-token chunk asked for 1.04 GiB with roughly 600 MB free. Dropping utilization alone does not work, and fails in a worse way: a single 420,000-token sequence needs 17.88 GiB of KV, and at 0.50 the pool is 17.4 to 17.5 GiB, so vLLM refuses to start at all and the seat crash-loops during startup instead of during a request. The context length and the crash were directly coupled -- 420k was only reachable at the utilization that left no transient headroom. So both moved: 0.50 and 393,216. 384k rather than vLLM's suggested maximum, deliberately. It estimated 406,352 on one boot and 409,840 on the next, because the available-KV figure drifts about 0.1 GiB boot to boot; pinning the edge value fails to start on an unlucky boot. 393,216 sits 3% under the lower estimate and leaves roughly 0.7 GiB of the pool unspent, which is the transient headroom the change exists to buy. Verified after: KV 405,612 tokens, concurrency 1.03x at 393,216, and both sec and sec-reasoning return 200 through the gateway. num_speculative_tokens is documented as NOT the lever. The crash window logged 17.6% draft acceptance with positions 5 through 7 at 1.5 to 4.9 percent, which reads as an obvious cut from 7 to 3; across 180 samples the median acceptance length is 3.12 of 7 and median draft acceptance is 30.4%, so the crash window sat near the minimum and cutting would cap the workloads accepting nearly the full draft. Cost: 384k of context instead of 420k, an 8.5% reduction on a seat whose crashes were happening at 151k.
This commit is contained in:
@@ -23,6 +23,37 @@
|
||||
# ships a deployment kit for — neither is our vLLM serving surface. 262K is the honest
|
||||
# native ceiling here; a real 1M seat would be a separate SGLang project.
|
||||
#
|
||||
# ⚠⚠ GPU0 HEADROOM — the 2026-09-10 crash pair, and why the live .env now reads
|
||||
# MOG_GPU_MEM_UTIL=0.50 / MOG_MAX_MODEL_LEN=393216 rather than 0.52 / 420000.
|
||||
#
|
||||
# At util 0.52 this seat and `gen` (0.38) together reserve 0.90 of the card, leaving
|
||||
# ~4.6 GiB. vLLM's utilization figure covers weights and the KV pool but NOT all
|
||||
# transient activation memory, and a long-context prefill chunk with the 7-wide dflash
|
||||
# drafter lives in what is left. Twice in ten minutes (20:20:30Z and 20:30:02Z) a request
|
||||
# ~151,700 tokens deep scheduling a further ~15,700-token chunk asked for ~1.04 GiB with
|
||||
# ~600 MB free, EngineCore took a fatal error, and `restart: unless-stopped` bounced the
|
||||
# seat. Each bounce is a hard 500 to every in-flight caller.
|
||||
#
|
||||
# ⚠ 0.50 AND 420000 ARE MUTUALLY EXCLUSIVE — dropping util alone does NOT work and the
|
||||
# seat will crash-loop at startup instead of at runtime. A single 420,000-token sequence
|
||||
# needs 17.88 GiB of KV; at 0.50 the pool is 17.4-17.5 GiB, so vLLM refuses:
|
||||
# "To serve at least one request with the model's max seq len (420000), 17.88 GiB KV
|
||||
# cache is needed, which is larger than the available KV cache memory (17.41 GiB)."
|
||||
# The context length and the crash were directly coupled: 420k was only reachable at the
|
||||
# utilization that left no transient headroom.
|
||||
#
|
||||
# ⚠ DO NOT pin max_model_len to vLLM's suggested maximum. It estimated 406,352 on one
|
||||
# boot and 409,840 on the next -- the available-KV figure drifts ~0.1 GiB boot to boot, so
|
||||
# the edge value fails to start on an unlucky one. 393,216 (384k) sits 3% under the lower
|
||||
# estimate and leaves ~0.7 GiB of the pool unspent, which IS the transient headroom this
|
||||
# change exists to buy. Measured after: KV 405,612 tokens, concurrency 1.03x at 393,216.
|
||||
#
|
||||
# ⚠ NOT the lever: `num_speculative_tokens`. The crash window logged 17.6% draft
|
||||
# acceptance with positions 5-7 at 1.5-4.9%, which reads as an obvious cut from 7 to 3.
|
||||
# Across 180 samples of that counter the median acceptance LENGTH is 3.12 of 7 (range
|
||||
# 1.83-6.75) and median draft acceptance is 30.4% (range 11.9-82.1%). The crash window sat
|
||||
# near the minimum; cutting to 3 would cap the workloads accepting nearly the full draft.
|
||||
#
|
||||
# Two served-names (base + `-thinking`): LiteLLM keys deployments by (model, api_base),
|
||||
# so mog-sec and mog-sec-reasoning use distinct names to avoid the shared-config
|
||||
# enable_thinking clobber. Same pinned nightly as the gen seat (carries the #51113
|
||||
|
||||
Reference in New Issue
Block a user