docs(gemma4-erp-tune): whole-card placement — gen moves to GPU1, sec stands down

Operator chose a third placement over the two the sizing offered: rather
than train beside gen on GPU0 or on GPU1 in mog-sec's slot, move gen to
GPU1 and empty GPU0 completely. The tune gets 95.60 GiB with no
co-tenant and gen never goes dark beyond its own restart.

Revised run parameters, since a whole card changes them:

- micro-batch 8 (71.8 GiB of 95.60) rather than 4, grad-accum 1, giving
  888 optimizer steps instead of 444. At one epoch the step count is
  worth having, and 8 x 8192 tokens puts ~4,096 rows through each expert
  per step against ~512 at micro-batch 1 — a far healthier GEMM on
  704-wide experts.
- Gradient checkpointing stays ON. Dropping it takes ~17% off wall-clock
  but pushes activations to ~24 GiB per sequence, which forces
  micro-batch 1 and costs 8x on MoE efficiency. Wide beats shallow.
- Scriberr stays on GPU1. The previous revision suggested moving it to
  GPU0, which was correct only while training was going to live on GPU1.

Records the ordering constraint in both directions, the elway identity
requirement, and that sec's aliases should be allowed to fail at the
gateway rather than be substituted with another model.
This commit is contained in:
vh
2026-08-24 18:40:42 -07:00
parent 9d70100867
commit c9943b1507
2 changed files with 67 additions and 28 deletions
+66 -27
View File
@@ -163,23 +163,33 @@ Card total: 97,887 MiB = **95.60 GiB** each.
| | GPU0 | GPU1 | | | GPU0 | GPU1 |
|---|---|---| |---|---|---|
| resident | `vllm-gen` 42,508 MiB (up 3 h) | `vllm-mog-sec` 56,624 MiB + `embed` 3,304 + `reward` 9,512 + `coder` 6,158 + `rerank-a3` 2,170; Scriberr pinned here, loads on demand | | resident before the window | `vllm-gen` 42,508 MiB (up 3 h) | `vllm-mog-sec` 56,624 MiB + `embed` 3,304 + `reward` 9,512 + `coder` 6,158 + `rerank-a3` 2,170; Scriberr pinned here, loads on demand |
| free | 54,741 MiB = **53.46 GiB** | 19,446 MiB = **18.99 GiB** | | free | 54,741 MiB = **53.46 GiB** | 19,446 MiB = **18.99 GiB** |
Three placements were on the table:
- **GPU0 beside `gen`: does not fit.** 53.46 GiB free against ~57.6 GiB needed — - **GPU0 beside `gen`: does not fit.** 53.46 GiB free against ~57.6 GiB needed —
short by ~4 GiB. And `gen` is only three hours old: measured footprint runs short by ~4 GiB. And `gen` is only three hours old: measured footprint runs
38.5 GiB fresh → 42.5 GiB at 3 h → 45.6 GiB at 3 days. Budgeting against the 38.5 GiB fresh → 42.5 GiB at 3 h → 45.6 GiB at 3 days. Budgeting against the
current number is budgeting against a moving one. current number is budgeting against a moving one.
- **GPU1 as-is: nowhere near.** 18.99 GiB. - **GPU1 with `mog-sec` stopped: 76,070 MiB free.** Fits, but shares a card with
- **GPU1 with `vllm-mog-sec` stopped: 76,070 MiB = 74.29 GiB free.** Fits four small seats and Scriberr.
micro-batch 4 (61.8 GiB) with ~12 GiB clear even after reserving ~6 GiB for - **GPU0 with `gen` MOVED OFF: the whole card.** ← what was chosen.
Scriberr's on-demand whisper load.
--- ---
## 4. Recommendation ## 4. The window, as executed
**Run on GPU1. Stand down `sec`, not `gen`.** **Operator call, 2026-08-24: move `gen` to GPU1 and stand `sec` down, so GPU0 is
emptied completely rather than shared.** This is strictly better than training
beside `gen`: the tune gets 95.60 GiB with no co-tenant, and the fleet's general
seat never goes dark beyond its own ~5-minute restart.
before: GPU0 [ gen 42.5 ] GPU1 [ sec 55.3 | small seats 20.7 ]
after: GPU0 [ ---- empty, 95.60 GiB ---- ] GPU1 [ gen ~41 | small seats 20.7 | ~33 free ]
`sec` is genuinely in use and this is not free — but it is the smaller blast
radius by a wide margin:
| | `gen` | `sec` | | | `gen` | `sec` |
|---|---|---| |---|---|---|
@@ -187,31 +197,60 @@ Card total: 97,887 MiB = **95.60 GiB** each.
| standing role | the fleet's general seat; a documented always-available dependency in global `CLAUDE.md` | M.O.G.-SEC, niche | | standing role | the fleet's general seat; a documented always-available dependency in global `CLAUDE.md` | M.O.G.-SEC, niche |
| measured traffic | 765 busy-engine log lines in 24 h — continuously in use | bursty; peak 8 concurrent, **last request ~5 h ago** | | measured traffic | 765 busy-engine log lines in 24 h — continuously in use | bursty; peak 8 concurrent, **last request ~5 h ago** |
`sec` is genuinely in use and this is not free — but it is the smaller blast Traffic to `sec` arrives from `10.250.50.70` (the LiteLLM gateway), so the
radius by a wide margin, and it is the difference between a 4–10 hour window aliases will fail at the gateway for the duration. Per the standing rule, let
that nobody outside the security work notices and one that takes the fleet's them fail — **do not route `sec` to another model as a stand-in.**
default model offline for a working day.
Also fold in, cheaply and reversibly: Both directions are playbooks, and **the order in each is load-bearing**:
- **Move Scriberr to GPU0 for the window** (`device_ids: ["0"]`, one compose scripts/elway infra-ops@10.250.50.54 --playbook playbooks/ana-ml2-training-window-open.yaml
edit + `up -d`). GPU0 will be sitting on ~53 GiB free with `gen` up, and it scripts/elway infra-ops@10.250.50.54 --playbook playbooks/ana-ml2-training-window-close.yaml
removes contention from the training card entirely.
- **Micro-batch 4, grad-accum 4** (effective 16, 444 steps). Micro-batch 6 is ⚠ `gen` runs at `--gpu-memory-utilization 0.43`, which vLLM reads as a fraction
the measured ceiling; the gap is deliberate margin. A seat crash-looped of **total** card memory: 42,091 MiB must be *free at startup* or the engine
earlier the same day on 0.6 GiB of assumed headroom — this document does not refuses to boot. GPU1 has 19,446 MiB free while `mog-sec` is up. **Recreating
repeat that. `gen` onto GPU1 before stopping `mog-sec` takes the fleet's main seat down and
leaves it down.** The open playbook stops `mog-sec` first and hard-gates on the
freed memory; the close playbook mirrors it, because `mog-sec` needs 50,901 MiB
of its own and cannot start until `gen` has vacated GPU1.
⚠ Invoke elway as `infra-ops@10.250.50.54`, not the `ana-ml2` ssh-target — that
resolves to `lkraven`, which has no NOPASSWD sudo, and elway aborts at its sudo
probe.
### Revised run parameters, now that it is a whole card
| micro-batch | GiB of 95.60 | |
|---:|---:|---|
| 4 | 61.8 | (the shared-card recommendation, now superseded) |
| **8** | **71.8** | **recommended — ~24 GiB clear** |
| 12 | 81.8 | ~14 GiB clear |
| 16 | 91.8 | too tight |
**Micro-batch 8, grad-accum 1 → 888 optimizer steps** (up from 444 on the shared
card). Two independent wins: the step count doubles, which matters at only one
epoch, and 8 × 8192 tokens gives ~4,096 rows per expert per step against ~512 at
micro-batch 1 — a far healthier GEMM on 704-wide experts, so MFU should land at
the top of the 10–25% band rather than the bottom.
⚠ **Keep gradient checkpointing ON.** With a whole card it is tempting to drop
it (~17% off wall-clock by removing the recompute forward), but activations then
run ~24 GiB per sequence, which forces micro-batch 1 — and micro-batch 1 craters
MoE GEMM efficiency by 8×. Going wide beats going shallow here. Do not
"optimize" this later without re-reading this paragraph.
Also fold in:
- **Scriberr STAYS on GPU1.** (An earlier draft of this document suggested moving
it to GPU0; that was written when training was going to live on GPU1, and it is
now exactly backwards. GPU0 is the training card and wants no co-tenant.)
- **Package as a `uv` venv on `/tank`, not a Docker image.** Root is at **91% - **Package as a `uv` venv on `/tank`, not a Docker image.** Root is at **91%
(36 GB free)** and `/var/lib/docker` lives on it; a PyTorch training image (36 GB free)** and `/var/lib/docker` lives on it; a PyTorch training image
would come close to filling it. `/tank` has 4.0 TB. would come close to filling it. `/tank` has 4.0 TB.
- **The run is still resumable-by-design** (INV-T7 + wall-clock checkpointing).
### If `sec` may not be stood down Nothing about a dedicated card removes that requirement — a 4–10 hour window
is long enough that an unresumable run is a bad bet regardless of who owns the
Fallback is GPU0 with `gen` stopped — the run fits with ~38 GiB to spare, which GPU.
buys micro-batch 8 and a shorter wall-clock. Make it **resumable and split** in
that case: INV-T7 plus wall-clock checkpointing already allows the window to be
broken into two or three shorter stretches with `gen` restored between them,
rather than one long outage.
--- ---
+1 -1
View File
@@ -110,7 +110,7 @@ no longer deployed sidecars here. See Recent decisions.)
_As of 2026-08-24 (late) — a very long ops session. The homepage arc and the char-rp arc both closed. **The live thread is the ERP/RP trainee: QLoRA sizing is the next conversation and brokkr-smithy-dev is waiting on it.**_ _As of 2026-08-24 (late) — a very long ops session. The homepage arc and the char-rp arc both closed. **The live thread is the ERP/RP trainee: QLoRA sizing is the next conversation and brokkr-smithy-dev is waiting on it.**_
- **🔴 SIZING DONE — AND IT KILLED THE QLoRA PREMISE. AWAITING THE SEAT CALL.** → `docs/pfi/gemma4-erp-tune-sizing.md` (measured, not estimated). **QLoRA is structurally unavailable on this architecture:** the checkpoint stores each layer's 128 experts as two fused 3-D `nn.Parameter` tensors (`layers.N.experts.gate_up_proj` `[128,1408,2816]`, `.down_proj` `[128,2816,704]` — note the missing `.weight` suffix), and `bitsandbytes` 4-bit replacement walks `nn.Linear` modules only. **88.5% of the model (22.84B params / 42.54 GiB) is untouchable; `load_in_4bit` saves ~3.1 GiB of 48.07 and does NOT error.** ⚠ **Eitri must not hard-code `BitsAndBytesConfig`** — it loads, reports 4-bit, and silently leaves 42.5 GiB BF16. Verdict: **plain LoRA on BF16**, ~57.6 GiB at micro-batch 1, +2.5 GiB per extra 8192-seq. ⚠ **The loss head is the real driver and was NOT in the brief:** vocab 262,144 × 8192 = 2.147B logits, plus `final_logit_softcapping 30.0` → **~28-30 GiB transient at BATCH 1** through naive HF CE (= ~85.6 GiB total, an OOM-on-first-long-sample). **Fused/chunked linear CE is mandatory and must be smoke-proven before the window is booked** (Liger may lack a Gemma-4 MoE patch). ⚠ **`v_proj` DOES NOT EXIST on layers 5/11/17/23/29** (`attention_k_eq_v` on the full-attention layers) — a `v_proj` target silently no-ops there and `k_proj` adapts K and V at once; 45.96M trainable at r64. ⚠ 7,104 packed seqs → **only 444 optimizer steps** at effective batch 16, which is why wall-clock checkpointing matters concretely. **RECOMMENDATION: run on GPU1, stand down `sec` (2 aliases, last request ~5h ago), NOT `gen` (7 aliases, 765 busy-engine lines/24h).** Stopping mog-sec frees 74.29 GiB; micro-batch 4 = 61.8 GiB. Est. **4-10h, likely ~6.** Also: move Scriberr to GPU0 for the window, and package as a `uv` venv on `/tank` — **root is 91% full (36 GB)** and `/var/lib/docker` is on it. - **🔴 SIZING DONE — AND IT KILLED THE QLoRA PREMISE. AWAITING THE SEAT CALL.** → `docs/pfi/gemma4-erp-tune-sizing.md` (measured, not estimated). **QLoRA is structurally unavailable on this architecture:** the checkpoint stores each layer's 128 experts as two fused 3-D `nn.Parameter` tensors (`layers.N.experts.gate_up_proj` `[128,1408,2816]`, `.down_proj` `[128,2816,704]` — note the missing `.weight` suffix), and `bitsandbytes` 4-bit replacement walks `nn.Linear` modules only. **88.5% of the model (22.84B params / 42.54 GiB) is untouchable; `load_in_4bit` saves ~3.1 GiB of 48.07 and does NOT error.** ⚠ **Eitri must not hard-code `BitsAndBytesConfig`** — it loads, reports 4-bit, and silently leaves 42.5 GiB BF16. Verdict: **plain LoRA on BF16**, ~57.6 GiB at micro-batch 1, +2.5 GiB per extra 8192-seq. ⚠ **The loss head is the real driver and was NOT in the brief:** vocab 262,144 × 8192 = 2.147B logits, plus `final_logit_softcapping 30.0` → **~28-30 GiB transient at BATCH 1** through naive HF CE (= ~85.6 GiB total, an OOM-on-first-long-sample). **Fused/chunked linear CE is mandatory and must be smoke-proven before the window is booked** (Liger may lack a Gemma-4 MoE patch). ⚠ **`v_proj` DOES NOT EXIST on layers 5/11/17/23/29** (`attention_k_eq_v` on the full-attention layers) — a `v_proj` target silently no-ops there and `k_proj` adapts K and V at once; 45.96M trainable at r64. ⚠ 7,104 packed seqs → **only 444 optimizer steps** at effective batch 16, which is why wall-clock checkpointing matters concretely. **OPERATOR CALL, EXECUTED 2026-08-24 EVENING: `gen` MOVED to GPU1, `sec` STOOD DOWN, GPU0 emptied entirely (95.60 GiB) for the tune** — strictly better than training beside gen. Both directions are playbooks: `playbooks/ana-ml2-training-window-{open,close}.yaml`. ⚠⚠ **ORDER IS LOAD-BEARING IN BOTH:** gen's `--gpu-memory-utilization 0.43` is a fraction of TOTAL card memory, so **42,091 MiB must be FREE AT STARTUP or the engine refuses to boot** — recreating gen onto GPU1 before stopping mog-sec takes the fleet's main seat down and LEAVES it down. Close mirrors it (mog-sec needs 50,901 MiB, so gen must vacate GPU1 first) and gates on GPU0 being idle so it cannot evict a running tune (`--var allow_busy_gpu0=true` to override). ⚠ **Invoke elway as `infra-ops@10.250.50.54`** — the `ana-ml2` ssh-target resolves to `lkraven`, no NOPASSWD sudo, elway aborts at its sudo probe. Revised run params on a whole card: **micro-batch 8, grad-accum 1 → 888 steps** (double the shared-card count, and ~4,096 rows/expert/step vs ~512 at mb=1, so much better MoE GEMM). ⚠ **KEEP gradient checkpointing ON** — dropping it saves ~17% wall-clock but forces mb=1, which craters MoE efficiency 8×; wide beats shallow. Scriberr STAYS on GPU1 (an earlier draft said move it to GPU0 — backwards now that GPU0 is the training card). Est. **4-10h, likely ~6.** Package as a `uv` venv on `/tank` — **root is 91% full (36 GB)** and `/var/lib/docker` is on it. ⚠ `sec` traffic arrives from the LiteLLM gateway (10.250.50.70); let the aliases FAIL, do not substitute another model.
- **⚠ TELL EITRI BEFORE HE HARD-CODES: the trainee base changed.** Contract still names the stock BF16. It is now `/tank/aimodels/gemma4-26b-a4b-it-heretic-bf16` (llmfan46). **Base path AND chat-template path must be config keys, not constants** — and the template must point at upstream's (`gemma4-26b-a4b-it-bf16/chat_template.jinja`), never the base's own, or training renders a different prompt than production serves. - **⚠ TELL EITRI BEFORE HE HARD-CODES: the trainee base changed.** Contract still names the stock BF16. It is now `/tank/aimodels/gemma4-26b-a4b-it-heretic-bf16` (llmfan46). **Base path AND chat-template path must be config keys, not constants** — and the template must point at upstream's (`gemma4-26b-a4b-it-bf16/chat_template.jinja`), never the base's own, or training renders a different prompt than production serves.
- **🟢 char-rp seat = Gemma-4 26B-A4B MoE NVFP4** on `:8016`, both aliases on ONE backend. **Currently DOWN by operator instruction** to hold GPU0 headroom for the tune. `gen` is UP and verified. MeroMero-v2 retained stopped in `created` state for rollback (stop-then-start; both bind :8016). → `persistent-memory.d/2026-08-24-charrp-gemma4-moe-swap-and-trainee.md` - **🟢 char-rp seat = Gemma-4 26B-A4B MoE NVFP4** on `:8016`, both aliases on ONE backend. **Currently DOWN by operator instruction** to hold GPU0 headroom for the tune. `gen` is UP and verified. MeroMero-v2 retained stopped in `created` state for rollback (stop-then-start; both bind :8016). → `persistent-memory.d/2026-08-24-charrp-gemma4-moe-swap-and-trainee.md`
- **🟢 THREE trainee-relevant model dirs on `/tank/aimodels/`, NOT interchangeable:** `gemma4-26b-a4b-it-bf16` (stock, 49 GB — its chat_template is the canonical upstream one), `gemma4-26b-a4b-it-heretic-bf16` (llmfan46 abliterated, the trainee), `gemma4-26b-a4b-it-abliterated-bf16` (TrevorJS, KL 0.09, alternate). Plus `-nvfp4` (served) and `-nvfp4a16` (activation control). ⚠ **BF16 cannot coexist with `gen`** — 48.07 GiB of weights on a 94.97 GiB card. Every BF16 window means gen stops. - **🟢 THREE trainee-relevant model dirs on `/tank/aimodels/`, NOT interchangeable:** `gemma4-26b-a4b-it-bf16` (stock, 49 GB — its chat_template is the canonical upstream one), `gemma4-26b-a4b-it-heretic-bf16` (llmfan46 abliterated, the trainee), `gemma4-26b-a4b-it-abliterated-bf16` (TrevorJS, KL 0.09, alternate). Plus `-nvfp4` (served) and `-nvfp4a16` (activation control). ⚠ **BF16 cannot coexist with `gen`** — 48.07 GiB of weights on a 94.97 GiB card. Every BF16 window means gen stops.