1a5bc2ddf1
The v2 dense quant had failed four times. Attempt 5 lands it at 19 G. The blocker was not what it looked like. `AmbiguousGlobalPerLayerAttributeError` on `head_dim` read as a malformed upload -- DogOnKeyboard's config carries a `per_layer_config` key zerofata's canonical one lacks -- and the standing fix was to force `allow_global_per_layer_attribute_access=True`. Both halves were wrong. `pip install llmcompressor==0.13.0` downgrades transformers 5.16.1 -> 5.14.1. The config was serialized by 5.16.1, which materializes `per_layer_config` from `global_head_dim` + `layer_types`; 5.14.1 carries the heterogeneity guard but not the gemma4 resolver. Under the image's own transformers the same config loads fine. `:latest` was also re-pulled during attempt 4 and no earlier run, so the toolchain moved mid-diagnosis. Two things separated "malformed upload" from "moved toolchain": reproducing the real failing call (a bare AutoConfig load does not reproduce it; the trigger is reached through AutoTokenizer) and keeping zerofata's canonical tree, quantized cleanly on 2026-08-21, as a positive control. The fix drops `per_layer_config` rather than forcing global access. It is exactly redundant -- keys are precisely the ten full_attention layer indices, sole value (512, 4), verbatim the global fields -- and forcing instead would make `config.head_dim` answer 256 to the callers building the 512-wide layers. patch_perlayer.py re-proves that redundancy at apply time and refuses if it ever stops holding. Verified on the tensor table rather than the exit code: the output is identical family-for-family and count-for-count to the August canonical quant, with 356 BF16 vision-tower tensors preserved and input_activations=None. A GPU-free load leaves 0 tensors on meta and generates coherent prose. The section 4.4 serve test has NOT run -- GPU1 has 19.9 GB free against 19.5 GB of weights, so it needs a live seat displaced. Also fixes the A4B output, which had a truncation cap baked into its tokenizer (max_length 8192) from being quantized with the calibration corpus. Playbook gains section 3.17 for the pinned-transformers class and sharpens 3.16 to say drop the dataset outright for any A16 scheme.
139 lines
8.1 KiB
Markdown
139 lines
8.1 KiB
Markdown
# MeroMero NVFP4A16 quants (Gemma-4) — instruments and the failure chain
|
||
|
||
Two in-house W4A16 quants of the abliterated MeroMero models, 2026-09-10. The
|
||
operator's brief was three clauses long: *"w4a16 vllm servable, vision towers
|
||
intact, mtp if applicable."* Every published quant of these models is W4A4 (our own
|
||
measured long-context collapse) or, for v2, does not exist at all — 0 of 27 v2 repos.
|
||
|
||
| output | source | result |
|
||
|---|---|---|
|
||
| `G4-MeroMero-26B-A4B-it-uncensored-heretic-NVFP4A16` | `DogOnKeyboard` A4B ablit | **16 G**, first try, 2m08s |
|
||
| `G4-MeroMero-v2-31B-heretic-NVFP4A16` | `DogOnKeyboard` v2-31B ablit | **19 G**, attempt **5** |
|
||
|
||
Both live on ana-ml2 under `/tank/aimodels/`. Neither is serving yet — see *Owed*.
|
||
|
||
## What the v2 dense cost, and why each layer mattered
|
||
|
||
Five attempts, five different causes. The order matters because each one masked the
|
||
next.
|
||
|
||
1. **`num_key_value_heads` is None** at `Gemma4TextAttention.__init__`. The uploader
|
||
set `attention_k_eq_v: true` but omitted `num_global_key_value_heads` and
|
||
`global_head_dim` — a malformed upload, not a toolchain problem. Patched from
|
||
zerofata's canonical values (4 / 512) after shape-verifying the checkpoint
|
||
(`shape_verify.py`): full-attn `k_proj [2048,5376]` = 4×512, sliding
|
||
`[4096,5376]` = 16×256, identical to canonical.
|
||
⚠ **A 2-layer truncation test PASSED and hid this.** The failing branch is chosen
|
||
per layer type and only `full_attention` layers take it. Testing each layer type
|
||
individually found it in seconds.
|
||
2. **`initialize model processor ... required when a dataset is provided`.** This
|
||
upload ships no `processor_config.json`. Rather than supply one, the dataset was
|
||
dropped entirely — NVFP4A16 is weight-only and runs a `DataFreePipeline`, so the
|
||
corpus was never read anyway (playbook §3.16), and passing one also bakes a
|
||
truncation cap into the shipped tokenizer (§3.14). Removing it kills both for
|
||
zero loss. `quant_a16_datafree.py`.
|
||
3. **Our own bug**: the reference module argparses at import with `required=True`,
|
||
so blanking `sys.argv` still exited 2. Placeholder argv, real one restored after.
|
||
4. **`AmbiguousGlobalPerLayerAttributeError: 'head_dim' is a per-layer attribute`.**
|
||
See below — this one was not what it looked like.
|
||
5. Clean. `rc=0`, 19 G, 3m07s.
|
||
|
||
## The attempt-4 trap: the transformers you measured is not the one that ran
|
||
|
||
The obvious reading was "DogOnKeyboard's config carries a `per_layer_config` key
|
||
zerofata's lacks, so that key is the defect." Two measurements said otherwise.
|
||
|
||
`tok_repro.py` reproduces the **actual failing call** rather than a paraphrase of it
|
||
— a bare `AutoConfig.from_pretrained` does not reproduce it, and testing that
|
||
instead would have sent us patching a file that was never the problem — with
|
||
zerofata's canonical tree, which quantized cleanly on 2026-08-21, as the positive
|
||
control. Against the container's shipped transformers **5.16.1**, every variant
|
||
passes, the unmodified heretic config included. Run the same script *after*
|
||
`pip install llmcompressor==0.13.0` and transformers is **5.14.1**: canonical
|
||
passes, heretic fails. **llmcompressor pins transformers and silently downgrades
|
||
it**, so the version printed before the install is not the version that runs.
|
||
|
||
Compounding it, `vllm/vllm-openai:latest` was re-pulled during attempt 4 and not
|
||
before it — the pull line appears in that run's log block and in no earlier one — so
|
||
the toolchain moved mid-campaign (the previous session recorded 5.12.1 in-container
|
||
while diagnosing attempt 1). That is why attempt 4's error read as a *new config
|
||
problem* and was not one. **`run_v2_quant.sh` now pins the image by digest.**
|
||
|
||
`per_layer_config` was in fact a 5.16.1 serialization artifact, and an exactly
|
||
redundant one: its keys are precisely the ten `full_attention` layer indices
|
||
[5,11,…,59] and its sole distinct value is `(head_dim 512, num_key_value_heads 4)` —
|
||
verbatim what `global_head_dim: 512` / `num_global_key_value_heads: 4` already say.
|
||
`patch_perlayer.py` re-proves that redundancy at patch time and refuses to drop the
|
||
key if it ever stops holding.
|
||
|
||
**Why drop it rather than set `allow_global_per_layer_attribute_access=True`.** The
|
||
forcing flag leaves the config heterogeneous and makes `config.head_dim` answer 256
|
||
to every caller — including the ones building the 512-wide full-attention layers.
|
||
`perlayer_test.py` builds all four variants on the meta device and reads the k_proj
|
||
widths back: geometry survived the flag, so it was not wrong, but llmcompressor's
|
||
own onloading is a caller nobody here has audited and the flag's warning names
|
||
exactly that hazard. The lossless option verified identically, so there was no
|
||
reason to take the risk.
|
||
|
||
## Verification — the tensor table, not the exit code
|
||
|
||
`rc=0` and a plausible file size prove neither of the operator's two checkable
|
||
requirements. `verify_quant.py` parses the safetensors headers directly (no torch,
|
||
no GPU, no 20 GB load) and reports dtypes by module family. The new quant is
|
||
**identical, family for family and count for count, to the 2026-08-21 known-good
|
||
canonical quant** — 410 U8 packed + 410 F8_E4M3 scales + 410 F32 global scales on the
|
||
LM Linears, **356 BF16 vision-tower tensors preserved**, `input_activations=None`
|
||
(genuinely A16, not A4). Shard sizes match that tree byte for byte. Full transcript
|
||
in `raw/verification.txt`.
|
||
|
||
MTP is N/A and was checked on the sources, not assumed: Gemma-4 ships no MTP head at
|
||
all — 0 mtp tensors and no mtp config keys in either bf16 source or in any published
|
||
quant. The "mtp if applicable" clause is a no-op for this family.
|
||
|
||
`cpu_smoke.py` then loads the finished tree with no GPU at all, confirms **0 tensors
|
||
left on the meta device** (a hole `from_pretrained` will not always raise on),
|
||
decompresses, and generates:
|
||
|
||
> *"A lighthouse is a tower with a bright light used to guide ships at sea and warn
|
||
> them of dangerous coastlines."*
|
||
|
||
24 greedy tokens, 3.4 s/tok on CPU. That is an "is it wired up" check and is offered
|
||
as nothing more — n=1 says nothing about quality, and it says nothing about whether
|
||
vLLM's sm_120 NVFP4 kernels serve the thing.
|
||
|
||
## Owed
|
||
|
||
- **§4.4 serve test on a temp port.** Not run. GPU1 has 19.9 GB free against 19.5 GB
|
||
of weights, so it cannot happen without displacing a live seat, which is the
|
||
operator's call. Until it runs, *"vllm servable"* is unverified for this tree.
|
||
- The A4B output **had the §3.14 truncation cap baked in** (`max_length: 8192`) — it
|
||
was quantized *with* the calibration corpus, before the data-free path existed.
|
||
Fixed 2026-09-10 by `post_quant_gemma4.py`; backup at
|
||
`tokenizer.json.bak-pre-truncfix`. Both trees now pass `--check` clean, as does the
|
||
August tree the checker is calibrated against.
|
||
|
||
## Files
|
||
|
||
| file | what it does |
|
||
|---|---|
|
||
| `run_v2_quant.sh` | attempt-5 runner, image pinned by digest |
|
||
| `run_quant_batch.sh` | the original two-model batch (A4B succeeded here, v2 did not) |
|
||
| `quant_a16_datafree.py` | NVFP4A16 driver with no dataset, wrapping the reference recipe |
|
||
| `patch_perlayer.py` | drops `per_layer_config`, re-proving its redundancy first |
|
||
| `shape_verify.py` | do the checkpoint's k/v shapes agree with the patched config? |
|
||
| `perlayer_test.py` | meta-device geometry across all four config variants |
|
||
| `tok_repro.py` | reproduces the real failing call; canonical tree as positive control |
|
||
| `post_quant_gemma4.py` | playbook §4.3 post-steps, idempotent, `--check` mode |
|
||
| `verify_quant.py` | dtypes by module family, straight from safetensors headers |
|
||
| `cpu_smoke.py` | GPU-free load + generate |
|
||
| `raw/` | run logs and the verification transcript, so the claims can be re-derived |
|
||
|
||
`raw/` holds `batch-quant-run-2026-09-10.txt` (attempt 1, and the A4B success),
|
||
`v2-quant-run-2026-09-10.txt` (attempts 2-5), `cpu-smoke-2026-09-10.txt`, and
|
||
`verification.txt`. Progress-bar redraws are collapsed to one line per bar, final
|
||
state; nothing else is edited. (`.txt` rather than `.log` because the repo ignores
|
||
`*.log` — same convention as `scripts/training-probes/`.)
|
||
|
||
General lessons live in `docs/pfi/model-quantization-playbook.md` (§3.4, §3.14,
|
||
§3.16, §3.17, §4.3) — read that first, and where it disagrees with this file, it wins.
|