Files
esh-pfi-infrastructure/services/meromero-quant/run_quant_batch.sh
T
vh 1a5bc2ddf1 Land the MeroMero v2-31B NVFP4A16 quant and record the pinned-transformers trap
The v2 dense quant had failed four times. Attempt 5 lands it at 19 G.

The blocker was not what it looked like. `AmbiguousGlobalPerLayerAttributeError`
on `head_dim` read as a malformed upload -- DogOnKeyboard's config carries a
`per_layer_config` key zerofata's canonical one lacks -- and the standing fix was
to force `allow_global_per_layer_attribute_access=True`. Both halves were wrong.

`pip install llmcompressor==0.13.0` downgrades transformers 5.16.1 -> 5.14.1. The
config was serialized by 5.16.1, which materializes `per_layer_config` from
`global_head_dim` + `layer_types`; 5.14.1 carries the heterogeneity guard but not
the gemma4 resolver. Under the image's own transformers the same config loads
fine. `:latest` was also re-pulled during attempt 4 and no earlier run, so the
toolchain moved mid-diagnosis. Two things separated "malformed upload" from
"moved toolchain": reproducing the real failing call (a bare AutoConfig load does
not reproduce it; the trigger is reached through AutoTokenizer) and keeping
zerofata's canonical tree, quantized cleanly on 2026-08-21, as a positive control.

The fix drops `per_layer_config` rather than forcing global access. It is exactly
redundant -- keys are precisely the ten full_attention layer indices, sole value
(512, 4), verbatim the global fields -- and forcing instead would make
`config.head_dim` answer 256 to the callers building the 512-wide layers.
patch_perlayer.py re-proves that redundancy at apply time and refuses if it ever
stops holding.

Verified on the tensor table rather than the exit code: the output is identical
family-for-family and count-for-count to the August canonical quant, with 356
BF16 vision-tower tensors preserved and input_activations=None. A GPU-free load
leaves 0 tensors on meta and generates coherent prose. The section 4.4 serve test
has NOT run -- GPU1 has 19.9 GB free against 19.5 GB of weights, so it needs a
live seat displaced.

Also fixes the A4B output, which had a truncation cap baked into its tokenizer
(max_length 8192) from being quantized with the calibration corpus.

Playbook gains section 3.17 for the pinned-transformers class and sharpens 3.16
to say drop the dataset outright for any A16 scheme.
2026-09-10 10:56:16 -07:00

41 lines
1.8 KiB
Bash
Executable File

#!/usr/bin/env bash
# Two NVFP4A16 quants, 2026-09-10. Operator: "run our own quant. w4a16 vllm
# servable, vision towers intact, mtp if applicable."
# W4A16 -> --scheme NVFP4A16 (weight-only, NOT plain NVFP4/W4A4)
# vision intact -> recipe ignore-list keeps vision/audio towers BF16
# MTP -> N/A: verified 0 mtp tensors in BOTH bf16 sources
# GPU1 not GPU0: the script onloads one layer at a time (GPU-light) but its own
# docstring warns of OOM when the card is not fairly free. GPU0 has 4.6 GiB spare
# (gen + mog-sec resident); GPU1 has ~19.3 GiB.
set -uo pipefail
WORK=/tank/aimodels/meromero-v2-nvfp4-work
CALIB=/tank/aimodels/heretic2-nvfp4-work/production_calib_512.jsonl
exec >> /home/infra-ops/quant/batch.log 2>&1
run_one () {
local name="$1" src="$2" out="$3"
echo "=== $(date -Is) START $name -> $out"
docker rm -f meromero-quant >/dev/null 2>&1
docker run --rm --name meromero-quant --gpus '"device=1"' --ipc host \
-e PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
-v /tank/aimodels:/tank/aimodels \
--entrypoint bash vllm/vllm-openai:latest -c "
set -e
pip install -q llmcompressor==0.13.0 tiktoken sentencepiece 2>&1 | tail -1
python3 $WORK/quant_nvfp4_gemma.py \
--model '$src' --calib '$CALIB' \
--num-samples 512 --seqlen 8192 --scheme NVFP4A16 \
--out '$out'
"
local rc=$?
echo "=== $(date -Is) END $name rc=$rc size=$(du -sh "$out" 2>/dev/null | cut -f1)"
}
run_one A4B-heretic \
/tank/aimodels/G4-MeroMero-26B-A4B-it-uncensored-heretic-bf16 \
/tank/aimodels/G4-MeroMero-26B-A4B-it-uncensored-heretic-NVFP4A16
run_one v2-31B-heretic \
/tank/aimodels/G4-MeroMero-v2-31B-heretic-bf16 \
/tank/aimodels/G4-MeroMero-v2-31B-heretic-NVFP4A16
echo "=== $(date -Is) BATCH DONE"