# MeroMero NVFP4A16 quants (Gemma-4) — instruments and the failure chain Two in-house W4A16 quants of the abliterated MeroMero models, 2026-09-10. The operator's brief was three clauses long: *"w4a16 vllm servable, vision towers intact, mtp if applicable."* Every published quant of these models is W4A4 (our own measured long-context collapse) or, for v2, does not exist at all — 0 of 27 v2 repos. | output | source | result | |---|---|---| | `G4-MeroMero-26B-A4B-it-uncensored-heretic-NVFP4A16` | `DogOnKeyboard` A4B ablit | **16 G**, first try, 2m08s | | `G4-MeroMero-v2-31B-heretic-NVFP4A16` | `DogOnKeyboard` v2-31B ablit | **19 G**, attempt **5** | Both live on ana-ml2 under `/tank/aimodels/`. Neither is serving yet — see *Owed*. ## What the v2 dense cost, and why each layer mattered Five attempts, five different causes. The order matters because each one masked the next. 1. **`num_key_value_heads` is None** at `Gemma4TextAttention.__init__`. The uploader set `attention_k_eq_v: true` but omitted `num_global_key_value_heads` and `global_head_dim` — a malformed upload, not a toolchain problem. Patched from zerofata's canonical values (4 / 512) after shape-verifying the checkpoint (`shape_verify.py`): full-attn `k_proj [2048,5376]` = 4×512, sliding `[4096,5376]` = 16×256, identical to canonical. ⚠ **A 2-layer truncation test PASSED and hid this.** The failing branch is chosen per layer type and only `full_attention` layers take it. Testing each layer type individually found it in seconds. 2. **`initialize model processor ... required when a dataset is provided`.** This upload ships no `processor_config.json`. Rather than supply one, the dataset was dropped entirely — NVFP4A16 is weight-only and runs a `DataFreePipeline`, so the corpus was never read anyway (playbook §3.16), and passing one also bakes a truncation cap into the shipped tokenizer (§3.14). Removing it kills both for zero loss. `quant_a16_datafree.py`. 3. **Our own bug**: the reference module argparses at import with `required=True`, so blanking `sys.argv` still exited 2. Placeholder argv, real one restored after. 4. **`AmbiguousGlobalPerLayerAttributeError: 'head_dim' is a per-layer attribute`.** See below — this one was not what it looked like. 5. Clean. `rc=0`, 19 G, 3m07s. ## The attempt-4 trap: the transformers you measured is not the one that ran The obvious reading was "DogOnKeyboard's config carries a `per_layer_config` key zerofata's lacks, so that key is the defect." Two measurements said otherwise. `tok_repro.py` reproduces the **actual failing call** rather than a paraphrase of it — a bare `AutoConfig.from_pretrained` does not reproduce it, and testing that instead would have sent us patching a file that was never the problem — with zerofata's canonical tree, which quantized cleanly on 2026-08-21, as the positive control. Against the container's shipped transformers **5.16.1**, every variant passes, the unmodified heretic config included. Run the same script *after* `pip install llmcompressor==0.13.0` and transformers is **5.14.1**: canonical passes, heretic fails. **llmcompressor pins transformers and silently downgrades it**, so the version printed before the install is not the version that runs. Compounding it, `vllm/vllm-openai:latest` was re-pulled during attempt 4 and not before it — the pull line appears in that run's log block and in no earlier one — so the toolchain moved mid-campaign (the previous session recorded 5.12.1 in-container while diagnosing attempt 1). That is why attempt 4's error read as a *new config problem* and was not one. **`run_v2_quant.sh` now pins the image by digest.** `per_layer_config` was in fact a 5.16.1 serialization artifact, and an exactly redundant one: its keys are precisely the ten `full_attention` layer indices [5,11,…,59] and its sole distinct value is `(head_dim 512, num_key_value_heads 4)` — verbatim what `global_head_dim: 512` / `num_global_key_value_heads: 4` already say. `patch_perlayer.py` re-proves that redundancy at patch time and refuses to drop the key if it ever stops holding. **Why drop it rather than set `allow_global_per_layer_attribute_access=True`.** The forcing flag leaves the config heterogeneous and makes `config.head_dim` answer 256 to every caller — including the ones building the 512-wide full-attention layers. `perlayer_test.py` builds all four variants on the meta device and reads the k_proj widths back: geometry survived the flag, so it was not wrong, but llmcompressor's own onloading is a caller nobody here has audited and the flag's warning names exactly that hazard. The lossless option verified identically, so there was no reason to take the risk. ## Verification — the tensor table, not the exit code `rc=0` and a plausible file size prove neither of the operator's two checkable requirements. `verify_quant.py` parses the safetensors headers directly (no torch, no GPU, no 20 GB load) and reports dtypes by module family. The new quant is **identical, family for family and count for count, to the 2026-08-21 known-good canonical quant** — 410 U8 packed + 410 F8_E4M3 scales + 410 F32 global scales on the LM Linears, **356 BF16 vision-tower tensors preserved**, `input_activations=None` (genuinely A16, not A4). Shard sizes match that tree byte for byte. Full transcript in `raw/verification.txt`. MTP is N/A and was checked on the sources, not assumed: Gemma-4 ships no MTP head at all — 0 mtp tensors and no mtp config keys in either bf16 source or in any published quant. The "mtp if applicable" clause is a no-op for this family. `cpu_smoke.py` then loads the finished tree with no GPU at all, confirms **0 tensors left on the meta device** (a hole `from_pretrained` will not always raise on), decompresses, and generates: > *"A lighthouse is a tower with a bright light used to guide ships at sea and warn > them of dangerous coastlines."* 24 greedy tokens, 3.4 s/tok on CPU. That is an "is it wired up" check and is offered as nothing more — n=1 says nothing about quality, and it says nothing about whether vLLM's sm_120 NVFP4 kernels serve the thing. ## Owed - **§4.4 serve test on a temp port.** Not run. GPU1 has 19.9 GB free against 19.5 GB of weights, so it cannot happen without displacing a live seat, which is the operator's call. Until it runs, *"vllm servable"* is unverified for this tree. - The A4B output **had the §3.14 truncation cap baked in** (`max_length: 8192`) — it was quantized *with* the calibration corpus, before the data-free path existed. Fixed 2026-09-10 by `post_quant_gemma4.py`; backup at `tokenizer.json.bak-pre-truncfix`. Both trees now pass `--check` clean, as does the August tree the checker is calibrated against. ## Files | file | what it does | |---|---| | `run_v2_quant.sh` | attempt-5 runner, image pinned by digest | | `run_quant_batch.sh` | the original two-model batch (A4B succeeded here, v2 did not) | | `quant_a16_datafree.py` | NVFP4A16 driver with no dataset, wrapping the reference recipe | | `patch_perlayer.py` | drops `per_layer_config`, re-proving its redundancy first | | `shape_verify.py` | do the checkpoint's k/v shapes agree with the patched config? | | `perlayer_test.py` | meta-device geometry across all four config variants | | `tok_repro.py` | reproduces the real failing call; canonical tree as positive control | | `post_quant_gemma4.py` | playbook §4.3 post-steps, idempotent, `--check` mode | | `verify_quant.py` | dtypes by module family, straight from safetensors headers | | `cpu_smoke.py` | GPU-free load + generate | | `raw/` | run logs and the verification transcript, so the claims can be re-derived | `raw/` holds `batch-quant-run-2026-09-10.txt` (attempt 1, and the A4B success), `v2-quant-run-2026-09-10.txt` (attempts 2-5), `cpu-smoke-2026-09-10.txt`, and `verification.txt`. Progress-bar redraws are collapsed to one line per bar, final state; nothing else is edited. (`.txt` rather than `.log` because the repo ignores `*.log` — same convention as `scripts/training-probes/`.) General lessons live in `docs/pfi/model-quantization-playbook.md` (§3.4, §3.14, §3.16, §3.17, §4.3) — read that first, and where it disagrees with this file, it wins.