From 8512dd4d31ce93384440bd3098d6b2d0a20a1bd1 Mon Sep 17 00:00:00 2001 From: Vuong Hoang Date: Tue, 8 Sep 2026 22:14:09 -0700 Subject: [PATCH] =?UTF-8?q?feat(erp-seat):=20erp-tune-v6-nvfp4a16=20quanti?= =?UTF-8?q?zed=20(data-free=20W4A16,=20~90=20s)=20and=20serving=20on=20ana?= =?UTF-8?q?-ml2=20GPU1=20:8021;=20playbook=20=C2=A73.16=20(data-free=20NVF?= =?UTF-8?q?P4A16=20still=20bakes=20the=20tokenizer=20cap);=20memory:=20gat?= =?UTF-8?q?e=20state,=20brokkr=20after-window=20asks,=20ana-ml2=20non-pers?= =?UTF-8?q?istent=20mesh=20routes?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- docs/pfi/model-quantization-playbook.md | 14 ++++++++++++++ persistent-memory.md | 20 ++++++++++++++++++++ 2 files changed, 34 insertions(+) diff --git a/docs/pfi/model-quantization-playbook.md b/docs/pfi/model-quantization-playbook.md index 4f0c905..86b7b6d 100644 --- a/docs/pfi/model-quantization-playbook.md +++ b/docs/pfi/model-quantization-playbook.md @@ -457,6 +457,20 @@ fallback, the incumbent-vs-candidate A/Bs (47.2% acceptance, PPL 6.910, and the 2026-08-20 Heretic-300 build) are apples-to-apples. This is unrealised upside, not a correction to past numbers. +### 3.16 Weight-only NVFP4A16 with a minmax observer is DATA-FREE — your calibration corpus is ignored, but its tokenizer side-effect is not + +Measured 2026-09-08 (Gemma-4 26B-A4B MoE, ERP run 6, llm-compressor 0.13): with +`scheme="NVFP4A16"` (default `memoryless_minmax` weights, no activation quant) llm-compressor +logs `Inferred DataFreePipeline for QuantizationModifier` and never touches the dataset — the +whole 26B quant ran in ~90 s on one Blackwell. Two consequences: (1) do not budget calibration +time or believe a corpus "shaped" the result — only `imatrix_mse`/activation observers consume +data; (2) building the calibration set still calls the fast tokenizer with +`truncation=True, max_length=N`, so §3.14's baked cap (`max_length: 8192` here) lands in the +saved `tokenizer.json` **even though no calibration happened**. The §4.3 post-step caught it. +Reference: `services/erp-seat-quant/quant_nvfp4a16_gemma4_moe.py` (linearize_moe + assert +11,520 expert Linears + post-steps; the published `prithivMLmods/gemma-4-26B-A4B-it-NVFP4A16` +recipe replicated, 222→252 ignore entries with audio/norm/router regexes added). + ### 3.14 ⭐⭐ Calibration BAKES a truncation cap into the shipped tokenizer **Symptom (on a newer transformers, at startup, on a vision model):** diff --git a/persistent-memory.md b/persistent-memory.md index 872b09c..7d4b734 100644 --- a/persistent-memory.md +++ b/persistent-memory.md @@ -158,6 +158,26 @@ COMPLETE and gated RESCUED (02:13 PDT). Live open items:_ installed over it (`*.repo` kept). ⚠ `hf download` ignores `--include` with several patterns after one flag. **LiteLLM `trial` alias DARK until the gate ends** (then repoint 5→6 only on the operator's word). Miranda informed 16:07 PT. → `docs/runbooks/gx10-run-06.md`. +- **🔥 GATE: erp-tune-v6 (bf16) SERVING on gx10:8098 since 21:46 PT** (`vllm-run06.pid`); brokkr's tuned window + (~1.5 h, k=25 legs CUT by operator) is unattended — files in brokkr-smithy `run06-gate/` are the signal, nobody pings. + **After TUNED-WINDOW-DONE (brokkr's three asks, agreed):** (1) keep erp-tune-v6 up until he posts "probe done" + (cue-length probe, ~15 min); (2) re-serve `erp-seat-base-ara` ~20 min for the probe's reference arm, then the GX10 + is FREE for run 7; (3) ✅ done — `dialogue-survivors.jsonl` (610, sha b2b6a0e9) copied to + `/mnt/smithy/datasets/derived/_recipes/erp-seat-sft-r3/`. Run 7 = recipe adjustment (reply length + + constraint-following), variable picked by the probe; no recipe/grant yet. +- **✅ erp-tune-v6-nvfp4a16 SERVING on ana-ml2 GPU1 :8021 (operator directive 2026-09-08 ~21:30 PT: "quant the + latest trained model into nvfp4 and serve it on ana-ml2 while we train a new model on the gx10").** Stack + `stacks/erp-seat` (recipe = gemma4-charrp's: gemma4 tool+reasoning parsers, enable_thinking pinned false, stock + template ae53464b, v0.26.0, util 0.35, 32K ctx) — TRUE name only; **`trial` alias NOT repointed (operator's call)**. + Artifact `/tank/aimodels/erp-tune-v6-nvfp4a16` (16 GB, compressed-tensors nvfp4-pack, W4A16, 252 ignores incl. + 60 router + 191 vision) from `/tank/aimodels/erp-tune-v6-bf16` (merged-run06 relayed gx10→nh3-dev→ana-ml2 in 17 min, + no key path gx10→ana-ml2). Quant = `services/erp-seat-quant/` — DATA-FREE (~90 s; playbook §3.16), tokenizer cap + reset, encode-equal to source. Smoke: prose in `content`, ~200 tok/s single-stream (n=3, spread <1%). + ⚠ **NOT gate-parity: the gated artifact is the bf16 arm on the GX10; the NVFP4 seat has only a smoke test** — a + brokkr battery subset on :8021 is the honest acceptance gate (offered, operator hasn't ruled). + ⚠ **ana-ml2 mesh return routes are NON-PERSISTENT** (`ip route replace 10.100/16, 10.0/16, 10.6.110/24, 100.64/10 + via 10.250.50.45` on `enp97s0f0np0.50`, 2026-09-08) — before that nh3-dev→ana-ml2 timed out (two DHCP defaults, + reply left the wrong NIC). Lost on reboot; make durable (netplan/networkd) or expect the timeout to return. - **📮 althing reachability on THIS bg seat = the cc-channel route, NOT the waiter.** `althing-listen` reaps within ~seconds-to-minutes on an idle background seat (harness kills detached bg tasks; hit it repeatedly this session). Fix (operator-directed 2026-09-08): `althing-route declare --handle infra-ops