memory: snapshot — run 3c gated, run 4 training, DAC revert, GPU rebalance, NAS exposure
This commit is contained in:
@@ -0,0 +1,41 @@
|
||||
# `[2026-09-04]` `gen` moved to ana-ml2 GPU0 — and I sized it against the wrong number, twice in one hour
|
||||
|
||||
`vllm-embed` had OOM-crashed **7 times** (`RestartCount 7`): GPU1 was at **216 MiB free** of
|
||||
97,887, and a ~3.4 GB tenant sharing a card with five other seats dies when it cannot get another
|
||||
96 MiB mid-inference. Root cause of the `qwen3-embed` 500s brokkr saw — **not** the LiteLLM
|
||||
gateway restart he attributed them to (different host, different component, 50 min earlier, and
|
||||
six of the seven crashes predate it).
|
||||
|
||||
**Operator ruling:** move `gen` to GPU0. *"we were keeping it free for training, but we don't have
|
||||
the power for sustained load."* GPU0 was never free — `mog-sec` has been there at util 0.52
|
||||
(55,126 MiB) since the August move.
|
||||
|
||||
## ⚠ THE SIZING ERROR, WHICH IS THE POINT OF THIS ENTRY
|
||||
|
||||
gen at util 0.43 wants 42,091 MiB GPU0 free 42,113 margin 22 MiB — would not start
|
||||
gen at util 0.41 wants 40,134 MiB margin 1,979 MiB — I called this safe. IT WAS NOT.
|
||||
result 1,548 MiB free, 56 FlashInfer autotuner OOM-fallbacks per 3 min
|
||||
gen at util 0.38 34,558 MiB actual 4,466-7,548 MiB free ZERO OOM events over 6 min
|
||||
|
||||
**`gpu-memory-utilization` governs vLLM's declared weights+KV budget. It does not cover what the
|
||||
process actually needs at runtime** — FlashInfer JIT autotuner workspace, CUDA graphs, expandable
|
||||
segments all allocate on top. My "1,979 MiB margin" was against the declared budget, so the seat
|
||||
came up and then silently fell back to slower kernels for want of 20 MB chunks.
|
||||
|
||||
⚠ **Same class of error as the DAC plateau the same day: checked a proxy, reported it as the
|
||||
thing.** The fix was to *measure* — count OOM-fallback events before and after — not to re-derive.
|
||||
|
||||
## Final state and what it cost
|
||||
|
||||
GPU0 mog-sec 55,126 + gen 34,558 = 89,701 / 97,887
|
||||
GPU1 50,507 / 97,887 — 46.7 GB free, was 216 MiB
|
||||
|
||||
`gen` KV cache is now **348,515 → 268,205 tokens, 1.02x concurrency at its 262,144 max-model-len**
|
||||
— it holds exactly one full-context request. Short/medium requests still batch; long-context
|
||||
throughput effectively serialises. The lever to reclaim it is `MOG_GPU_MEM_UTIL 0.52`, untouched.
|
||||
|
||||
.env.bak-preGPU0-20260904-164032 the GPU move
|
||||
.env.bak-preShrink-165133 the utilization
|
||||
|
||||
⚠ `/opt/docker/compose/*/.env` is root-owned — `docker compose restart` needs **sudo** on
|
||||
ana-docker/ana-ml2 or it fails with a bare `permission denied` on the .env.
|
||||
Reference in New Issue
Block a user