diff --git a/docs/pfi/gemma4-erp-tune-sizing.md b/docs/pfi/gemma4-erp-tune-sizing.md new file mode 100644 index 0000000..ba82d0e --- /dev/null +++ b/docs/pfi/gemma4-erp-tune-sizing.md @@ -0,0 +1,235 @@ +# Gemma-4 26B-A4B ERP/RP tune — GPU sizing adjudication + +_Measured 2026-08-24 on `ana-ml2` against +`/tank/aimodels/gemma4-26b-a4b-it-heretic-bf16` (llmfan46 abliterated trainee)._ + +Division of labour for this run: **Eitri writes the harness, brokkr-smithy-dev +audits, infra-ops owns the GPU window and executes.** This document is the +sizing infra-ops owes; it is arithmetic against the real checkpoint and the +real card, not an estimate. + +--- + +## 1. ⚠ QLoRA IS NOT AVAILABLE ON THIS ARCHITECTURE + +**The proposed shape was QLoRA r64. It cannot be run as specified**, and the +reason is structural rather than a tuning preference. + +The checkpoint stores each layer's 128 experts as **two fused 3-D +`nn.Parameter` tensors**, not as 128 `nn.Linear` modules: + + model.language_model.layers.N.experts.gate_up_proj BF16 [128, 1408, 2816] + model.language_model.layers.N.experts.down_proj BF16 [128, 2816, 704] + +Note the absence of a `.weight` suffix — compare `mlp.down_proj.weight` +(an `nn.Linear`) against `experts.down_proj` (a bare parameter). That is the +tell, and it is decisive: **`bitsandbytes` 4-bit replacement walks `nn.Linear` +modules.** A fused 3-D parameter is not one, so it is skipped and stays BF16. + +What `load_in_4bit=True` would actually buy on this model: + +| block | params | BF16 | after bnb NF4 | saved | +|---|---:|---:|---:|---:| +| **MoE experts** (fused 3-D — **NOT quantized**) | 22.84 B | 42.54 GiB | **42.54 GiB** | **0** | +| lm attention (`nn.Linear`) | 1.11 B | 2.07 GiB | 0.52 GiB | 1.55 | +| dense shared MLP (`nn.Linear`) | 0.54 B | 1.00 GiB | 0.25 GiB | 0.75 | +| vision tower (`nn.Linear`) | 0.57 B | 1.06 GiB | 0.27 GiB | 0.79 | +| embed (tied, normally kept BF16) | 0.74 B | 1.38 GiB | 1.38 GiB | 0 | +| router + norms | 0.01 B | 0.02 GiB | 0.02 GiB | 0 | +| **total** | **25.81 B** | **48.07 GiB** | **~44.98 GiB** | **~3.1 GiB** | + +**88.5% of the model is in tensors bitsandbytes cannot touch.** "QLoRA" here +means paying the NF4 dequant tax on 6% of the weights to save 6% of the +footprint. The premise does not survive contact with the checkpoint. + +> **Eitri: do not hard-code a `BitsAndBytesConfig` / `load_in_4bit` path.** +> It will not error loudly — it will load, report a 4-bit model, and quietly +> leave 42.5 GiB in BF16. Same silent-failure shape as the stale chat template. + +**The one thing that could overturn this** is a third-party fork shipping +custom grouped-GEMM 4-bit MoE kernels for this specific architecture (Unsloth +is the candidate). **Not chased, deliberately** — see §4, where the run fits in +BF16 without displacing anything the fleet depends on, which collapses QLoRA's +value to zero. If it is ever revisited, it must be *before* the harness +hard-codes a quantization path, not after. + +**Verdict: plain LoRA on BF16 weights.** + +--- + +## 2. What the run actually costs + +Adapter targeting `q_proj,k_proj,v_proj,o_proj` at r64, computed from the real +tensor shapes: + +| | layers | per layer | total | +|---|---:|---:|---:| +| sliding-attention (q 4096, kv 2048, o 4096) | 25 | 1,507,328 | 37,683,200 | +| full-attention (q 8192, kv 1024, o 8192) | 5 | 1,654,784 | 8,273,920 | +| **trainable** | | | **45,957,120** (0.178% of base) | + +⚠ **`v_proj` DOES NOT EXIST ON LAYERS 5, 11, 17, 23, 29.** Those are the +`full_attention` layers, and `attention_k_eq_v: true` means one projection +serves both K and V. Consequences the harness must respect: + +- PEFT matches by name suffix, so a `v_proj` target **silently produces no + adapter** on those five layers. Do not assert a fixed adapter count. +- Adapting `k_proj` on a global layer **adapts K and V simultaneously** — a + different intervention than on the sliding layers. If that asymmetry matters + to the recipe, say so explicitly rather than discovering it in the loss curve. + +### Memory budget, batch 1, `max_seq_len` 8192 + +| item | GiB | note | +|---|---:|---| +| base weights BF16 | 48.07 | measured: 25,805,936,206 params × 2 B | +| adapters + grads + AdamW fp32 m/v | 0.75 | 45.96 M trainable — rounding error | +| checkpointed layer inputs | 1.29 | 30 × 8192 × 2816 × 2 B | +| recompute peak, one layer | ~2.5 | 8192 tok × top-8 of 128, `moe_intermediate 704` | +| loss head, **fused/chunked CE** | ~2.0 | see the warning below | +| CUDA context + cuBLAS + fragmentation | ~3.0 | the item `--gpu-memory-utilization` never covered | +| **total** | **~57.6** | | + +Marginal cost per extra sequence in the micro-batch: **~2.5 GiB.** + +| micro-batch | GiB | +|---:|---:| +| 1 | 54.3 | +| 2 | 56.8 | +| **4** | **61.8** | +| 6 | 66.8 | +| 8 | 71.8 | + +### ⚠ The loss head is the whole ballgame, and it is not in the brief + +`vocab_size` is **262,144** and `final_logit_softcapping` is **30.0**. One +8192-token sequence produces **2.147 billion logits**. Through a naive HF +`ForCausalLM` loss that is: + + BF16 logits 4.0 GiB + fp32 upcast 8.0 GiB + softcap tanh saved 8.0 GiB (autograd keeps the pre-cap tensor) + softmax + grad 8.0 GiB + ------------------------------ + ~28-30 GiB transient, at BATCH 1 + +Naive CE at batch 1 lands the run at **~85.6 GiB on a 95.6 GiB card** — it will +appear to work and then OOM on the first long sample. At micro-batch 4 it is +~120 GiB and never starts. **Fused/chunked linear cross-entropy is mandatory, +not an optimization.** + +⚠ Honest uncertainty: Liger ships per-architecture patches and Gemma-4 MoE with +softcapping may not have one. Three ways out, in order of preference — +(a) generic `LigerFusedLinearCrossEntropyLoss` wired against the lm_head with +softcapping applied inside the chunk; (b) `cut-cross-entropy`; (c) hand-rolled +sequence-chunked CE. **This must be proven on a 10-step smoke run before the +window is booked**, because everything else in this document assumes it works. + +### Step count + +58.2 M tokens / 20,576 samples = **2,829 tokens/sample average** — well under +8192, so packing matters. + +- Packed to 8192: **7,104 sequences.** At micro-batch 4 × grad-accum 4 + (effective 16) → **444 optimizer steps for the whole epoch.** +- ⚠ That is a *small* step count. A "checkpoint every 100 steps" default gives + four checkpoints across a multi-hour run. This is exactly why the amendment + asked for **wall-clock-interval checkpointing, not step-count** — the case is + now concrete, not hypothetical. +- ⚠ **Packing must use `position_ids` + varlen/block-diagonal attention.** Naive + concatenation bleeds samples into each other. `sliding_window` is 1024 on 25 + of 30 layers so the damage is bounded there — but the 5 `full_attention` + layers see the entire packed sequence. + +**Open question for brokkr/Eitri:** what fraction of the 20,576 samples exceed +8192 tokens? Below ~2%, 8192 is right. A long tail means truncation is cutting +the ends off RP scenes, which is where the signal lives. + +### Runtime + +Active parameters per token ≈ **3.67 B** (2.93 B routed + attention, plus the +0.74 B tied lm_head matmul). Forward + backward + gradient-checkpoint recompute +≈ 6 × active × tokens = **1.28e18 FLOPs** for the epoch. + +At 10–25% MFU on a 300 W-capped Max-Q card — HF MoE paths with 704-wide experts +are not efficient — **4 to 10 hours, most likely ~6.** Treat as a band, not a +number; it will be measured on the smoke run. + +--- + +## 3. Where it fits (measured 2026-08-24, 18:20 PDT) + +Card total: 97,887 MiB = **95.60 GiB** each. + +| | GPU0 | GPU1 | +|---|---|---| +| resident | `vllm-gen` 42,508 MiB (up 3 h) | `vllm-mog-sec` 56,624 MiB + `embed` 3,304 + `reward` 9,512 + `coder` 6,158 + `rerank-a3` 2,170; Scriberr pinned here, loads on demand | +| free | 54,741 MiB = **53.46 GiB** | 19,446 MiB = **18.99 GiB** | + +- **GPU0 beside `gen`: does not fit.** 53.46 GiB free against ~57.6 GiB needed — + short by ~4 GiB. And `gen` is only three hours old: measured footprint runs + 38.5 GiB fresh → 42.5 GiB at 3 h → 45.6 GiB at 3 days. Budgeting against the + current number is budgeting against a moving one. +- **GPU1 as-is: nowhere near.** 18.99 GiB. +- **GPU1 with `vllm-mog-sec` stopped: 76,070 MiB = 74.29 GiB free.** Fits + micro-batch 4 (61.8 GiB) with ~12 GiB clear even after reserving ~6 GiB for + Scriberr's on-demand whisper load. + +--- + +## 4. Recommendation + +**Run on GPU1. Stand down `sec`, not `gen`.** + +| | `gen` | `sec` | +|---|---|---| +| aliases | 7 (`gen`, `gen-reasoning`, `chat-judge`, `image-judge`, summarizer/classifier family) | 2 (`sec`, `sec-reasoning`) | +| standing role | the fleet's general seat; a documented always-available dependency in global `CLAUDE.md` | M.O.G.-SEC, niche | +| measured traffic | 765 busy-engine log lines in 24 h — continuously in use | bursty; peak 8 concurrent, **last request ~5 h ago** | + +`sec` is genuinely in use and this is not free — but it is the smaller blast +radius by a wide margin, and it is the difference between a 4–10 hour window +that nobody outside the security work notices and one that takes the fleet's +default model offline for a working day. + +Also fold in, cheaply and reversibly: + +- **Move Scriberr to GPU0 for the window** (`device_ids: ["0"]`, one compose + edit + `up -d`). GPU0 will be sitting on ~53 GiB free with `gen` up, and it + removes contention from the training card entirely. +- **Micro-batch 4, grad-accum 4** (effective 16, 444 steps). Micro-batch 6 is + the measured ceiling; the gap is deliberate margin. A seat crash-looped + earlier the same day on 0.6 GiB of assumed headroom — this document does not + repeat that. +- **Package as a `uv` venv on `/tank`, not a Docker image.** Root is at **91% + (36 GB free)** and `/var/lib/docker` lives on it; a PyTorch training image + would come close to filling it. `/tank` has 4.0 TB. + +### If `sec` may not be stood down + +Fallback is GPU0 with `gen` stopped — the run fits with ~38 GiB to spare, which +buys micro-batch 8 and a shorter wall-clock. Make it **resumable and split** in +that case: INV-T7 plus wall-clock checkpointing already allows the window to be +broken into two or three shorter stretches with `gen` restored between them, +rather than one long outage. + +--- + +## 5. Standing warnings that apply to this run + +- **Never render training examples through the base's own + `chat_template.jinja`.** Every third-party Gemma-4 derivative ships a stale + one; the trainee's is 365 lines against upstream's 390. Use + `/tank/aimodels/gemma4-26b-a4b-it-bf16/chat_template.jinja`. Training through + the wrong template is train/serve skew with no error — it presents as a + tuning failure. +- **Base path and chat-template path are config keys, not constants.** The + trainee base already moved once (stock BF16 → `-heretic-bf16`). +- **`--gpu-memory-utilization` sizes the KV cache only.** It does not cover CUDA + context, graphs, or non-torch overhead — the same misreading that OOM'd the + char-rp seat. +- **Serving the result is not settled.** LoRA-on-NVFP4 hot-swap was a silent + no-op on vLLM 0.24.0 (#47639, proven quant-agnostic). Retest on the tagged + `vllm/vllm-openai:v0.27.1` already on disk. **If it still no-ops, the harness + must emit merged weights** — and Eitri needs that requirement while he is + early, not after the run. diff --git a/persistent-memory.md b/persistent-memory.md index 1341cc7..8ac61a2 100644 --- a/persistent-memory.md +++ b/persistent-memory.md @@ -110,7 +110,7 @@ no longer deployed sidecars here. See Recent decisions.) _As of 2026-08-24 (late) — a very long ops session. The homepage arc and the char-rp arc both closed. **The live thread is the ERP/RP trainee: QLoRA sizing is the next conversation and brokkr-smithy-dev is waiting on it.**_ -- **🔴 NEXT UP — QLoRA SIZING, AND I OWN THE RUN.** Operator set the division of labour: **Eitri writes the training harness, brokkr audits, infra-ops owns the GPU window AND executes the run.** Harness contract requires headless-from-a-config (no notebook, no interactive steps) and INV-T7 resumable-from-checkpoint. I asked for one amendment: **checkpoint on a wall-clock interval, not only step count**, since step time under contention is not knowable in advance. Rough shape: QLoRA r64, attention-only adapters, max_seq_len 8192, 1 epoch, ~58.2M tokens / 20,576 samples. **I have NOT sized it and must not guess** — GPU0 carries `gen` (~38.5 GiB fresh, ~45.6 GiB after days up) and GPU1 carries `sec` + Scriberr with ~19.4 GiB free. The window will require standing something down; which seat is the arithmetic I owe. +- **🔴 SIZING DONE — AND IT KILLED THE QLoRA PREMISE. AWAITING THE SEAT CALL.** → `docs/pfi/gemma4-erp-tune-sizing.md` (measured, not estimated). **QLoRA is structurally unavailable on this architecture:** the checkpoint stores each layer's 128 experts as two fused 3-D `nn.Parameter` tensors (`layers.N.experts.gate_up_proj` `[128,1408,2816]`, `.down_proj` `[128,2816,704]` — note the missing `.weight` suffix), and `bitsandbytes` 4-bit replacement walks `nn.Linear` modules only. **88.5% of the model (22.84B params / 42.54 GiB) is untouchable; `load_in_4bit` saves ~3.1 GiB of 48.07 and does NOT error.** ⚠ **Eitri must not hard-code `BitsAndBytesConfig`** — it loads, reports 4-bit, and silently leaves 42.5 GiB BF16. Verdict: **plain LoRA on BF16**, ~57.6 GiB at micro-batch 1, +2.5 GiB per extra 8192-seq. ⚠ **The loss head is the real driver and was NOT in the brief:** vocab 262,144 × 8192 = 2.147B logits, plus `final_logit_softcapping 30.0` → **~28-30 GiB transient at BATCH 1** through naive HF CE (= ~85.6 GiB total, an OOM-on-first-long-sample). **Fused/chunked linear CE is mandatory and must be smoke-proven before the window is booked** (Liger may lack a Gemma-4 MoE patch). ⚠ **`v_proj` DOES NOT EXIST on layers 5/11/17/23/29** (`attention_k_eq_v` on the full-attention layers) — a `v_proj` target silently no-ops there and `k_proj` adapts K and V at once; 45.96M trainable at r64. ⚠ 7,104 packed seqs → **only 444 optimizer steps** at effective batch 16, which is why wall-clock checkpointing matters concretely. **RECOMMENDATION: run on GPU1, stand down `sec` (2 aliases, last request ~5h ago), NOT `gen` (7 aliases, 765 busy-engine lines/24h).** Stopping mog-sec frees 74.29 GiB; micro-batch 4 = 61.8 GiB. Est. **4-10h, likely ~6.** Also: move Scriberr to GPU0 for the window, and package as a `uv` venv on `/tank` — **root is 91% full (36 GB)** and `/var/lib/docker` is on it. - **⚠ TELL EITRI BEFORE HE HARD-CODES: the trainee base changed.** Contract still names the stock BF16. It is now `/tank/aimodels/gemma4-26b-a4b-it-heretic-bf16` (llmfan46). **Base path AND chat-template path must be config keys, not constants** — and the template must point at upstream's (`gemma4-26b-a4b-it-bf16/chat_template.jinja`), never the base's own, or training renders a different prompt than production serves. - **🟢 char-rp seat = Gemma-4 26B-A4B MoE NVFP4** on `:8016`, both aliases on ONE backend. **Currently DOWN by operator instruction** to hold GPU0 headroom for the tune. `gen` is UP and verified. MeroMero-v2 retained stopped in `created` state for rollback (stop-then-start; both bind :8016). → `persistent-memory.d/2026-08-24-charrp-gemma4-moe-swap-and-trainee.md` - **🟢 THREE trainee-relevant model dirs on `/tank/aimodels/`, NOT interchangeable:** `gemma4-26b-a4b-it-bf16` (stock, 49 GB — its chat_template is the canonical upstream one), `gemma4-26b-a4b-it-heretic-bf16` (llmfan46 abliterated, the trainee), `gemma4-26b-a4b-it-abliterated-bf16` (TrevorJS, KL 0.09, alternate). Plus `-nvfp4` (served) and `-nvfp4a16` (activation control). ⚠ **BF16 cannot coexist with `gen`** — 48.07 GiB of weights on a 94.97 GiB card. Every BF16 window means gen stops.