feat(erp-seat): erp-tune-v6-nvfp4a16 quantized (data-free W4A16, ~90 s) and serving on ana-ml2 GPU1 :8021; playbook §3.16 (data-free NVFP4A16 still bakes the tokenizer cap); memory: gate state, brokkr after-window asks, ana-ml2 non-persistent mesh routes

This commit is contained in:
vh
2026-09-08 22:14:09 -07:00
parent 911ff20356
commit 8512dd4d31
2 changed files with 34 additions and 0 deletions
+14
View File
@@ -457,6 +457,20 @@ fallback, the incumbent-vs-candidate A/Bs (47.2% acceptance, PPL 6.910, and the
2026-08-20 Heretic-300 build) are apples-to-apples. This is unrealised upside, not a 2026-08-20 Heretic-300 build) are apples-to-apples. This is unrealised upside, not a
correction to past numbers. correction to past numbers.
### 3.16 Weight-only NVFP4A16 with a minmax observer is DATA-FREE — your calibration corpus is ignored, but its tokenizer side-effect is not
Measured 2026-09-08 (Gemma-4 26B-A4B MoE, ERP run 6, llm-compressor 0.13): with
`scheme="NVFP4A16"` (default `memoryless_minmax` weights, no activation quant) llm-compressor
logs `Inferred DataFreePipeline for QuantizationModifier` and never touches the dataset — the
whole 26B quant ran in ~90 s on one Blackwell. Two consequences: (1) do not budget calibration
time or believe a corpus "shaped" the result — only `imatrix_mse`/activation observers consume
data; (2) building the calibration set still calls the fast tokenizer with
`truncation=True, max_length=N`, so §3.14's baked cap (`max_length: 8192` here) lands in the
saved `tokenizer.json` **even though no calibration happened**. The §4.3 post-step caught it.
Reference: `services/erp-seat-quant/quant_nvfp4a16_gemma4_moe.py` (linearize_moe + assert
11,520 expert Linears + post-steps; the published `prithivMLmods/gemma-4-26B-A4B-it-NVFP4A16`
recipe replicated, 222→252 ignore entries with audio/norm/router regexes added).
### 3.14 ⭐⭐ Calibration BAKES a truncation cap into the shipped tokenizer ### 3.14 ⭐⭐ Calibration BAKES a truncation cap into the shipped tokenizer
**Symptom (on a newer transformers, at startup, on a vision model):** **Symptom (on a newer transformers, at startup, on a vision model):**
+20
View File
@@ -158,6 +158,26 @@ COMPLETE and gated RESCUED (02:13 PDT). Live open items:_
installed over it (`*.repo` kept). ⚠ `hf download` ignores `--include` with several patterns after one flag. installed over it (`*.repo` kept). ⚠ `hf download` ignores `--include` with several patterns after one flag.
**LiteLLM `trial` alias DARK until the gate ends** (then repoint 5→6 only on the operator's word). **LiteLLM `trial` alias DARK until the gate ends** (then repoint 5→6 only on the operator's word).
Miranda informed 16:07 PT. → `docs/runbooks/gx10-run-06.md`. Miranda informed 16:07 PT. → `docs/runbooks/gx10-run-06.md`.
- **🔥 GATE: erp-tune-v6 (bf16) SERVING on gx10:8098 since 21:46 PT** (`vllm-run06.pid`); brokkr's tuned window
(~1.5 h, k=25 legs CUT by operator) is unattended — files in brokkr-smithy `run06-gate/` are the signal, nobody pings.
**After TUNED-WINDOW-DONE (brokkr's three asks, agreed):** (1) keep erp-tune-v6 up until he posts "probe done"
(cue-length probe, ~15 min); (2) re-serve `erp-seat-base-ara` ~20 min for the probe's reference arm, then the GX10
is FREE for run 7; (3) ✅ done — `dialogue-survivors.jsonl` (610, sha b2b6a0e9) copied to
`/mnt/smithy/datasets/derived/_recipes/erp-seat-sft-r3/`. Run 7 = recipe adjustment (reply length +
constraint-following), variable picked by the probe; no recipe/grant yet.
- **✅ erp-tune-v6-nvfp4a16 SERVING on ana-ml2 GPU1 :8021 (operator directive 2026-09-08 ~21:30 PT: "quant the
latest trained model into nvfp4 and serve it on ana-ml2 while we train a new model on the gx10").** Stack
`stacks/erp-seat` (recipe = gemma4-charrp's: gemma4 tool+reasoning parsers, enable_thinking pinned false, stock
template ae53464b, v0.26.0, util 0.35, 32K ctx) — TRUE name only; **`trial` alias NOT repointed (operator's call)**.
Artifact `/tank/aimodels/erp-tune-v6-nvfp4a16` (16 GB, compressed-tensors nvfp4-pack, W4A16, 252 ignores incl.
60 router + 191 vision) from `/tank/aimodels/erp-tune-v6-bf16` (merged-run06 relayed gx10→nh3-dev→ana-ml2 in 17 min,
no key path gx10→ana-ml2). Quant = `services/erp-seat-quant/` — DATA-FREE (~90 s; playbook §3.16), tokenizer cap
reset, encode-equal to source. Smoke: prose in `content`, ~200 tok/s single-stream (n=3, spread <1%).
⚠ **NOT gate-parity: the gated artifact is the bf16 arm on the GX10; the NVFP4 seat has only a smoke test** — a
brokkr battery subset on :8021 is the honest acceptance gate (offered, operator hasn't ruled).
⚠ **ana-ml2 mesh return routes are NON-PERSISTENT** (`ip route replace 10.100/16, 10.0/16, 10.6.110/24, 100.64/10
via 10.250.50.45` on `enp97s0f0np0.50`, 2026-09-08) — before that nh3-dev→ana-ml2 timed out (two DHCP defaults,
reply left the wrong NIC). Lost on reboot; make durable (netplan/networkd) or expect the timeout to return.
- **📮 althing reachability on THIS bg seat = the cc-channel route, NOT the waiter.** `althing-listen` - **📮 althing reachability on THIS bg seat = the cc-channel route, NOT the waiter.** `althing-listen`
reaps within ~seconds-to-minutes on an idle background seat (harness kills detached bg tasks; hit it reaps within ~seconds-to-minutes on an idle background seat (harness kills detached bg tasks; hit it
repeatedly this session). Fix (operator-directed 2026-09-08): `althing-route declare --handle infra-ops repeatedly this session). Fix (operator-directed 2026-09-08): `althing-route declare --handle infra-ops