Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-04-ana-ml2-gpu-rebalance.md
T

2.4 KiB

[2026-09-04] gen moved to ana-ml2 GPU0 — and I sized it against the wrong number, twice in one hour

vllm-embed had OOM-crashed 7 times (RestartCount 7): GPU1 was at 216 MiB free of 97,887, and a ~3.4 GB tenant sharing a card with five other seats dies when it cannot get another 96 MiB mid-inference. Root cause of the qwen3-embed 500s brokkr saw — not the LiteLLM gateway restart he attributed them to (different host, different component, 50 min earlier, and six of the seven crashes predate it).

Operator ruling: move gen to GPU0. "we were keeping it free for training, but we don't have the power for sustained load." GPU0 was never free — mog-sec has been there at util 0.52 (55,126 MiB) since the August move.

⚠ THE SIZING ERROR, WHICH IS THE POINT OF THIS ENTRY

gen at util 0.43   wants 42,091 MiB    GPU0 free 42,113    margin 22 MiB   — would not start
gen at util 0.41   wants 40,134 MiB    margin 1,979 MiB    — I called this safe. IT WAS NOT.
result             1,548 MiB free, 56 FlashInfer autotuner OOM-fallbacks per 3 min
gen at util 0.38   34,558 MiB actual   4,466-7,548 MiB free   ZERO OOM events over 6 min

gpu-memory-utilization governs vLLM's declared weights+KV budget. It does not cover what the process actually needs at runtime — FlashInfer JIT autotuner workspace, CUDA graphs, expandable segments all allocate on top. My "1,979 MiB margin" was against the declared budget, so the seat came up and then silently fell back to slower kernels for want of 20 MB chunks.

Same class of error as the DAC plateau the same day: checked a proxy, reported it as the thing. The fix was to measure — count OOM-fallback events before and after — not to re-derive.

Final state and what it cost

GPU0   mog-sec 55,126 + gen 34,558 = 89,701 / 97,887
GPU1   50,507 / 97,887 — 46.7 GB free, was 216 MiB

gen KV cache is now 348,515 → 268,205 tokens, 1.02x concurrency at its 262,144 max-model-len — it holds exactly one full-context request. Short/medium requests still batch; long-context throughput effectively serialises. The lever to reclaim it is MOG_GPU_MEM_UTIL 0.52, untouched.

.env.bak-preGPU0-20260904-164032   the GPU move
.env.bak-preShrink-165133          the utilization

/opt/docker/compose/*/.env is root-owned — docker compose restart needs sudo on ana-docker/ana-ml2 or it fails with a bare permission denied on the .env.