fix(heretic2-nvfp4): quant as ConditionalGeneration (namespace fix) + modelopt recipe for working MTP
Root-caused the NVFP4 gibberish to a quant-namespace bug: quant_nvfp4.py loaded via AutoModelForCausalLM -> text-only Qwen3_5ForCausalLM -> flat model.layers.* keys, but vLLM 0.24 serves only Qwen3_5ForConditionalGeneration (whose weight mapper needs model.language_model.*). Fixed by loading as AutoModelForImageTextToText; NVFP4 now serves coherent (validated greedy on ana-ml2 GPU0). Base NVFP4 (compressed-tensors) measured ~53 tok/s (~= GGUF at batch-1, no single-stream win) and its MTP is 0% acceptance (vLLM's Qwen3_5MTP drafter loads the bf16 mtp head only off a modelopt main-model checkpoint). Added quant_modelopt.py (nvidia-modelopt PTQ, matches AEON's NVFP4 W4A4 g16 + lm_head/linear_attn/visual exclusions) as the path to working native MTP; graft + splice + serve otherwise unchanged.
This commit is contained in:
@@ -8,6 +8,24 @@ inside soong's latency window). R36 fast-seat spike, 2026-07-14.
|
||||
Fleet-first **local** NVFP4 quant — every other fleet NVFP4 model is *pulled*
|
||||
pre-quantized; Heretic2 has none published, so we quantize it.
|
||||
|
||||
## 2026-07-14 status — gibberish FIXED, format PIVOTED to modelopt for MTP
|
||||
|
||||
- **Root cause of the `!!!!` was the quant NAMESPACE**, not calib/scheme: `quant_nvfp4.py`
|
||||
loaded via `AutoModelForCausalLM` → text-only `Qwen3_5ForCausalLM` → flat `model.layers.*`
|
||||
keys, but vLLM 0.24 serves only `Qwen3_5ForConditionalGeneration`, whose weight mapper needs
|
||||
`model.language_model.*`. **Fixed** by loading as `AutoModelForImageTextToText`
|
||||
(= `Qwen3_5ForConditionalGeneration`) → keys born `model.language_model.*` + `model.visual.*`.
|
||||
NVFP4 now serves **coherent**.
|
||||
- **But this (llm-compressor / compressed-tensors) format can't deliver the speed goal:** base
|
||||
NVFP4 ≈ 53 tok/s ≈ the GGUF seat's ~59.5 at batch-1 (no single-stream win), and **MTP =
|
||||
0% acceptance** (vLLM's `Qwen3_5MTP` drafter won't load the bf16 mtp head off a compressed-tensors
|
||||
main model). The mtp tensors are identical to AEON's; the blocker is purely the main-model format.
|
||||
- **→ Working native MTP requires the MODELOPT format** (what AEON uses, ~3.3/3 accept). Use
|
||||
**`quant_modelopt.py`** (nvidia-modelopt PTQ; same graft + same `splice_mtp.py` + serve
|
||||
`--quantization modelopt`). AEON `/tank/aimodels/qwen36-27b-aeon-nvfp4` (`vllm-aeon-rp`) is the
|
||||
exact reference. `quant_nvfp4.py` (below) is kept for the coherent-but-MTP-inert compressed-tensors
|
||||
artifact and as the namespace-fix record.
|
||||
|
||||
## Fire sequence
|
||||
1. **`graft_mtp.py`** — graft the 15 base-Qwen3.6 MTP tensors into Heretic2 BF16.
|
||||
CPU-only, no GPU window. (Heretic2's finetune dropped the head; config declares
|
||||
|
||||
Reference in New Issue
Block a user