Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-15-breeze-placement-sizing.md
T
vh 9d36c74572 memory: snapshot refresh — breeze settled, util does not predict residency, handoff defect
Incremental over 653f7fb, covering the three turns since.

breeze-tts stays on irv-ml1 and the TTS-stack move is parked at id 75. Adds the
full fv-ml1 per-seat residency table and the shuffle arithmetic behind the
recommendation: every arrangement that makes real room on GPU 1 spends the GPU 3
reserve anyway, so if breeze ever moves it should go straight to GPU 3 rather than
evacuating three working seats for the same cost plus a migration.

The finding worth keeping is that --gpu-memory-utilization does not predict resident
VRAM, and is wrong in both directions: cyberprev at util 0.40 holds 8 GB more than
its fraction, gen-small at 0.48 holds 10 GB less. An hour earlier in the session I
reasoned about placement from those fractions and would have been 8 GB out. Read
nvidia-smi --query-compute-apps instead.

Also records a characterized defect in this skill's own handoff generator. Across
two runs in one session it converted operator-deferred items into an imperative
next-steps list and twice invited the next session to commit files marked as
predating it. It fails in the blind spot of the documented fail-loud posture: the
output is structurally valid and exits 0 while inverting the intent. That matters
more than an ordinary bad summary because the handoff is the one artifact a fresh
context reads as instruction. Both runs were corrected in-session; not yet reported
to the skill's owner.
2026-09-15 10:08:57 -07:00

6.1 KiB

breeze-tts sizing and the fv-ml1 GPU 0 placement question (2026-09-15)

Operator asked how big breeze-tts is and whether fv-ml1 GPU 0 has room for it. Recommendation: do not move it. Numbers and reasoning below.

⚠ Measure it UNDER LOAD and measure it LATE — it does both badly at idle and early

when breeze VRAM
shortly after warm-up 9,218 MiB
after 53 min uptime 10,316 MiB
at idle on the wrong card invisible — shows nothing

Two traps stacked:

  1. ⚠ It reports nothing when you query the wrong GPU. BREEZE_GPU_DEVICES=0 and on irv-ml1 that is the RTX 3090, not the A6000. An idle query of GPU 1 shows breeze absent entirely — I read that as "released its memory" before checking the device map.
  2. ⚠ The footprint GROWS with use — +1.1 GiB over 53 minutes. Consistent with PyTorch's caching allocator not returning memory (already hypothesised in this repo for this host), so probably caching rather than a leak — but it is resident either way and it counts against any neighbour. n=2 points is a trend, not a characterised curve; whether it plateaus is unmeasured.

Other facts: image local/breeze-tts:v3 is 26.5 GB on disk; weights bind-mounted from /home/lkraven/breeze/weights; container pid 2048660, CUDA pid 2048873.

Why NOT to move it to fv-ml1 GPU 0

1. The margin is thinner than it first looked, and trending the wrong way.

fv-ml1 GPU 0 free : 11,982 MiB   (two vLLM seats at util 0.88, + parakeet ~800 MiB)
breeze at 9.2 GiB -> 2.8 GB margin   <- the number I first quoted
breeze at 10.3 GiB -> 1.7 GB margin  <- the number after an hour, still climbing

GPU 0 is the live chat serving path. Putting a growing tenant in shrinking headroom there is the wrong direction.

2. ⭐ The real objection is topology, not VRAM. tts-gateway runs on irv-ml1 and reaches breeze on the same box. Moving breeze alone to Fountain Valley puts an Irvine↔FV hop on every TTS call, against a budget tts-dev measured at 478 ms to first sample. Moving it properly means moving the gateway too — a project, not a placement tweak.

3. It is not constrained where it is. The 3090 still has 10,099 MiB free.

✅ RESOLVED — operator 2026-09-15: leave it, park the move

"leave it where it is, park moving breeze, bragi, and the tts-gateway to fv1 when we move the embedder, reranker, and reward models off gpu0"

Parked at move-the-tts-stack-breeze-tts-bragi-tts-gateway (park id 75).

All three move together, and only one is GPU-resident:

service footprint
breeze-tts ~10.3 GiB VRAM, growing
bragi CPU only — proxy, polls tts-gateway /health
tts-gateway CPU only — the ext-tts routing gateway

⭐ Co-location is the whole reason to move them as a set: tts-gateway reaches breeze same-box today, and moving breeze alone is what puts the cross-site hop on every call.

⚠ The trigger as stated has the card wrong, and it is worth confirming before executing. The operator said "off gpu0", but vllm-embed, vllm-rerank-a3 and vllm-reward are all pinned to GPU 1 (util 0.03 + 0.03 + 0.10 ≈ 0.16, ~15.7 GB). GPU 1 is the constrained card at 0.975 committed / 4,336 MiB free; GPU 0 has 11,982 MiB free and carries the live chat path. So evacuating those three relieves GPU 1, not GPU 0 — which changes where the TTS stack would land.

⚠ Migration gotcha: containers on irv-ml1 cannot resolve *.nh3.internal; bragi and tts-gateway reach each other by name only through extra_hosts pins. Those pins travel with them and need re-pointing at the new host.

fv-ml1 per-seat residency, measured 2026-09-15 (97,887 MiB per card)

GPU seats (resident MiB) resident free
0 cyberprev 47,124 · gen-small 36,942 · parakeet 1,178 85,244 11,982
1 meromero 44,056 · erp 27,748 · reward 9,192 · coder 6,292 · embed 3,472 · rerank 2,114 92,874 4,336
2 flash-next 95,344 95,344 1,767
3 — (deliberate reserve) 2 97,247

⚠⚠ --gpu-memory-utilization DOES NOT PREDICT RESIDENCY — measure, don't compute

Capacity planning off the configured util fractions is wrong in both directions:

vllm-cyberprev   util 0.40 -> expect ~39,155 MiB   ACTUAL 47,124   (+8.0 GB OVER)
vllm-gen-small   util 0.48 -> expect ~46,986 MiB   ACTUAL 36,942   (-10.0 GB UNDER)

An hour earlier in this session I reasoned about placement from the util fractions and would have been ~8 GB wrong on one card. ⭐ util is a ceiling the engine may not reach and evidently can exceed (weights, activations and non-KV allocations sit outside the KV budget the fraction governs). Always read nvidia-smi --query-compute-apps for the real number.

Does a shuffle make room for breeze? — worked through

  • GPU 0 today: YES, no shuffle. 11,982 free vs breeze 10,316 → 1.7 GB margin, against a tenant that grew +1.1 GB in an hour, on the live chat path. Fits; thin.
  • GPU 1 today: NO. 4,336 free.
  • Evacuate embed+rerank+reward (14,778 MiB) → GPU 1 has 19,114 free, breeze fits with ~8.8 GB headroom. But there is nowhere to put those three except GPU 3 — 14,778 does not fit GPU 0's 11,982, and GPU 2 has 1,767.
  • Partial shuffle fails: embed+rerank (5,586) to GPU 0 leaves GPU 1 at 9,922 free — ~400 MB short of breeze, and breeze is growing.

⭐ Conclusion: every arrangement that makes real room on GPU 1 spends the GPU 3 reserve anyway. So if breeze ever goes to FV, put it straight on GPU 3 rather than evacuating three working seats to achieve the same cost plus a migration.

If consolidation onto FV is the goal anyway

GPU 3 fits it comfortably (97,247 MiB free) and has none of the margin problem — but it spends the reserve, and the reserve exists because vLLM sizes KV against TOTAL VRAM and flash-next needs 93 of 96 GiB. Breeze at ~10 GiB leaves ~87 GiB, which no longer takes a full-card seat. That is the trade; it is the operator's call, and the gateway should move with it either way.