2656196f47
Run-01 completed in 7:21:52 (47% faster than the 13.85h round-1 projection),
lora_B gate 205/205 non-zero at median norm 1.708, and the acceptance gate says
it did the thing it was built for: diversity +0.178 against a 0.008 floor (22x),
attractor hit rate -11.3pt against a 2.0pt floor, memorisation 0.0000 on both
arms — which closes the R20 licensed-prose exposure on measurement rather than
argument.
Five new detail files carry the substance:
erp-tune-run2-complete the run, the gate, the noise-floor near-miss
(brokkr was one step from reporting a 13-point
T6 regression sitting inside twice his
instrument's own variance)
mfu-root-caused-attention 8.6% MFU was an accounting artifact; real
utilisation 17-20%, cost was attention on
AMPERE kernels. Two independent methods agreed
to 2.6 points.
nvfp4-serving-pipeline merged weights are MANDATORY — vLLM cannot
serve a LoRA on ANY Gemma-4 — plus the recipe
that silently misses all 11,520 expert tensors
refusal-retention-probe measured base 0/100 -> tuned 29/100, then had
to accept it was the wrong axis
worldtree-b188-b189-and-selene three arcs closed, and a #411 diagnosis I got
wrong twice before a directory probe settled it
Current state rewritten end to end — the previous snapshot had the run in
flight at ~17h with MFU unexplained. Both are now closed.
The open operator decision is run 2's base, deliberately unstaged and flagged
against being filed as a config knob: it is a reversal of the trainee-selection
decision, and the pretrained-base option removes the last non-lexical floor on
the CSAM axis given stage-2-detector-inert and contamination-scan-absent are
both already overridden.
Tried-and-abandoned gains four measured-dead throughput levers, the packing
correction (bucketing wins under sdpa and the conclusion flips under flex — do
not carry it past the backend decision), and the merge-back-undoes-abliteration
trap brokkr caught in his own advice.
Index stays at 291 lines, under the soft cap. No archival this run.
85 lines
4.5 KiB
Markdown
85 lines
4.5 KiB
Markdown
# NVFP4A16 serving pipeline — built, validated, and the MoE landmine it found
|
|
|
|
`[2026-08-25]`
|
|
|
|
Pipeline at `scripts/erp-tune-serve/` (commits `6a85829`, `ab980e9`).
|
|
Validated end-to-end against checkpoint-100 before the real adapter existed.
|
|
|
|
## ⚠⚠ THE LANDMINE: a `targets=["Linear"]` recipe misses EVERY MoE expert
|
|
|
|
before linearize_moe: 427 Linears, 205 targeted, experts 0
|
|
after linearize_moe: 11,947 Linears, 11,725 targeted, experts 11,520
|
|
(30 layers x 128 experts x 3 projections)
|
|
|
|
Gemma-4 stores each layer's 128 experts as two fused 3-D `nn.Parameter` tensors
|
|
(`gate_up_proj` [128,1408,2816], `down_proj` [128,2816,704]) — note the absent
|
|
`.weight` suffix. A Linear-targeting recipe resolves 205 of 427 modules and
|
|
**zero experts**, leaving 22.84 B params (88.5% of the model) in BF16 with no
|
|
warning.
|
|
|
|
**This is the same defect that killed QLoRA here via bitsandbytes.** The blind
|
|
spot is in the *checkpoint layout*, not the tool. Fix:
|
|
`llmcompressor.modeling.moe.linearize.linearize_moe` — no registration needed,
|
|
Gemma-4 satisfies `FusedExpertsProtocol` structurally. Playbook §3.15.
|
|
|
|
## Scheme: NVFP4A16, deviating from the playbook default, on measured grounds
|
|
|
|
brokkr benched the W4A4 quant of this checkpoint at **12% on contradiction
|
|
detection with CoT off against gen's 81%** — the signature of 4-bit input
|
|
activations on a reasoning-dense task. Plus W4A4 KLD is 2-4x worse past ~10k ctx
|
|
on sm_120. This is a 16,384-ctx RP seat. Marlin's prefill cost accepted.
|
|
|
|
⚠ Several HF repos named `…-NVFP4A16` declare `input_activations num_bits 4` —
|
|
W4A4 wearing an A16 label. The script refuses if the emitted config says 4.
|
|
|
|
## Four silent defects the dry run found
|
|
|
|
1. **transformers 5.15 MIGRATES the config schema on save** — drops
|
|
`global_head_dim`/`num_global_key_value_heads`, writes `per_layer_config`.
|
|
transformers 5.10 (the llmcompressor venv) then reads `num_key_value_heads`
|
|
as None and dies with `TypeError: unsupported operand type(s) for //`.
|
|
Every working artifact on the box uses the OLD schema. Merge now downgrades it.
|
|
2. **llmcompressor cannot auto-init a processor for a multimodal checkpoint** —
|
|
pass the tokenizer explicitly as `processor`.
|
|
3. **`save_pretrained` does not carry `processor_config.json`** — vLLM then fails
|
|
with "Can't load feature extractor", which reads as a vision bug.
|
|
4. **The quant needs more than GPU1's free 32 GiB.** `quant_with_gen_down.sh`
|
|
stops `vllm-gen` and restores it from a trap on EVERY exit path, using
|
|
`docker start` not `compose up` so the container returns with its exact config.
|
|
|
|
## Verified on the emitted artifact
|
|
|
|
49 GB -> 17 GB, format nvfp4-pack-quantized, a=null (genuine A16)
|
|
weight_packed 11,725 of which expert 11,520
|
|
tokenizer truncation: clean (§3.14 trap avoided by calibrating on the
|
|
encode cache, so the tokenizer is never called
|
|
with truncation=True at all)
|
|
served: Marlin NVFP4 kernel + Marlin MoE backend, coherent generation
|
|
|
|
⚠ The reference `nvfp4a16` artifact triggers a vLLM warning that q/k/v carry
|
|
*different* weight global scales ("likely reduced accuracy"). **Ours does not** —
|
|
llmcompressor 0.12 links weight observers across fused groups automatically. The
|
|
in-house quant is better than the downloaded one on that axis.
|
|
|
|
## ⚠ MERGED WEIGHTS ARE MANDATORY — and not for the reason we assumed
|
|
|
|
The open question was whether LoRA-on-NVFP4 hot-swap still silently no-ops.
|
|
Retested on `vllm/vllm-openai:latest`: **it refuses to start.**
|
|
|
|
AttributeError: To support LoRA for MoE model,
|
|
'get_expert_mapping' must be implemented
|
|
|
|
The check is in `vllm/lora/utils.py::process_packed_modules_mapping` and branches
|
|
on `is_moe_model()` — **quantization is not in the condition.** `gemma4.py`,
|
|
`gemma4_mm.py`, `gemma4_mtp.py`, `gemma4_unified.py` all have ZERO occurrences;
|
|
`deepseek_v2`, `mixtral`, `glm4_moe`, `ernie45_moe` implement it.
|
|
|
|
**vLLM cannot serve a LoRA on ANY Gemma-4, bf16 or quantized.** Merging is the
|
|
only path for this architecture, and it would have bitten identically on the
|
|
unquantized base. A loud refusal is strictly better than the 0.24.0 silent no-op,
|
|
which shipped a base model wearing the tune's name.
|
|
|
|
⚠ Base-viability pre-flight is now playbook §3.11 — three greps before picking a
|
|
base. **Grep the CLASS, not the file**: `mistral.py` greps as `SupportsLoRA=0`
|
|
and is fully LoRA-capable via inheritance from `LlamaForCausalLM`.
|