memory: breeze stays put; TTS-stack move to fv-ml1 parked at id 75
Operator ruling: leave breeze-tts on irv-ml1 and park moving it, bragi and tts-gateway to fv-ml1 until the embedder, reranker and reward seats are evacuated. Parked as move-the-tts-stack-breeze-tts-bragi-tts-gateway (id 75) with the trigger, the footprints and the migration gotchas, so it resurfaces with everything needed rather than as a bare line. Two things worth having recorded against the trigger. All three services move as a set because only breeze is GPU-resident at ~10.3 GiB and growing, while bragi and tts-gateway are CPU-only proxies - co-location with the gateway is the entire reason not to move breeze alone, since that is what puts a cross-site hop on every TTS call. And the trigger as stated names gpu0, but vllm-embed, vllm-rerank-a3 and vllm-reward are all pinned to GPU 1. GPU 1 is the constrained card at 0.975 committed with 4,336 MiB free, while GPU 0 has 11,982 MiB free and carries the live chat path, so evacuating those three relieves GPU 1 rather than GPU 0. Recorded as a confirm-before-executing rather than silently corrected, since it changes where the TTS stack would land. Also notes that bragi and tts-gateway reach each other by name only through extra_hosts pins, because containers on irv-ml1 cannot resolve nh3.internal - those pins travel with them and need re-pointing at the new host.
This commit is contained in:
@@ -44,6 +44,35 @@ placement tweak.
|
||||
|
||||
**3. It is not constrained where it is.** The 3090 still has **10,099 MiB free**.
|
||||
|
||||
## ✅ RESOLVED — operator 2026-09-15: leave it, park the move
|
||||
|
||||
> "leave it where it is, park moving breeze, bragi, and the tts-gateway to fv1 when we
|
||||
> move the embedder, reranker, and reward models off gpu0"
|
||||
|
||||
Parked at **`move-the-tts-stack-breeze-tts-bragi-tts-gateway` (park id 75)**.
|
||||
|
||||
**All three move together, and only one is GPU-resident:**
|
||||
|
||||
| service | footprint |
|
||||
|---|---|
|
||||
| `breeze-tts` | **~10.3 GiB VRAM**, growing |
|
||||
| `bragi` | **CPU only** — proxy, polls tts-gateway `/health` |
|
||||
| `tts-gateway` | **CPU only** — the `ext-tts` routing gateway |
|
||||
|
||||
⭐ Co-location is the whole reason to move them as a set: `tts-gateway` reaches breeze
|
||||
**same-box** today, and moving breeze alone is what puts the cross-site hop on every call.
|
||||
|
||||
⚠ **The trigger as stated has the card wrong, and it is worth confirming before
|
||||
executing.** The operator said "off gpu0", but `vllm-embed`, `vllm-rerank-a3` and
|
||||
`vllm-reward` are all pinned to **GPU 1** (util 0.03 + 0.03 + 0.10 ≈ 0.16, ~15.7 GB).
|
||||
GPU 1 is the constrained card at **0.975 committed / 4,336 MiB free**; GPU 0 has
|
||||
11,982 MiB free and carries the live chat path. So evacuating those three relieves
|
||||
**GPU 1**, not GPU 0 — which changes where the TTS stack would land.
|
||||
|
||||
⚠ **Migration gotcha:** containers on irv-ml1 cannot resolve `*.nh3.internal`; `bragi`
|
||||
and `tts-gateway` reach each other by name only through `extra_hosts` pins. Those pins
|
||||
travel with them and need re-pointing at the new host.
|
||||
|
||||
## If consolidation onto FV is the goal anyway
|
||||
|
||||
**GPU 3** fits it comfortably (97,247 MiB free) and has none of the margin problem — but
|
||||
|
||||
Reference in New Issue
Block a user