fix(heretic2-nvfp4): quant as ConditionalGeneration (namespace fix) + modelopt recipe for working MTP
Root-caused the NVFP4 gibberish to a quant-namespace bug: quant_nvfp4.py loaded via AutoModelForCausalLM -> text-only Qwen3_5ForCausalLM -> flat model.layers.* keys, but vLLM 0.24 serves only Qwen3_5ForConditionalGeneration (whose weight mapper needs model.language_model.*). Fixed by loading as AutoModelForImageTextToText; NVFP4 now serves coherent (validated greedy on ana-ml2 GPU0). Base NVFP4 (compressed-tensors) measured ~53 tok/s (~= GGUF at batch-1, no single-stream win) and its MTP is 0% acceptance (vLLM's Qwen3_5MTP drafter loads the bf16 mtp head only off a modelopt main-model checkpoint). Added quant_modelopt.py (nvidia-modelopt PTQ, matches AEON's NVFP4 W4A4 g16 + lm_head/linear_attn/visual exclusions) as the path to working native MTP; graft + splice + serve otherwise unchanged.
This commit is contained in:
@@ -82,12 +82,24 @@ def main() -> int:
|
||||
ap.add_argument("--seqlen", type=int, default=8192)
|
||||
args = ap.parse_args()
|
||||
|
||||
from transformers import AutoModelForCausalLM, AutoTokenizer
|
||||
from transformers import AutoModelForImageTextToText, AutoTokenizer
|
||||
from llmcompressor import oneshot
|
||||
from llmcompressor.modifiers.quantization import QuantizationModifier
|
||||
|
||||
print(f"loading grafted model: {args.model}", flush=True)
|
||||
model = AutoModelForCausalLM.from_pretrained(
|
||||
# Load as the FULL multimodal Qwen3_5ForConditionalGeneration (NOT AutoModelForCausalLM).
|
||||
# AutoModelForCausalLM resolves qwen3_5 -> Qwen3_5ForCausalLM (text-only), whose weight
|
||||
# keys are flat `model.layers.*` with no vision tower. But vLLM 0.24 only registers
|
||||
# Qwen3_5ForConditionalGeneration, and its hf_to_vllm_mapper expects the checkpoint keyed
|
||||
# `model.language_model.layers.*` (+ `model.visual.*`) — a bare `model.layers.` prefix has
|
||||
# NO mapping rule, so every transformer-layer weight fails to load -> uninitialized weights
|
||||
# -> degenerate `!!!!` output. AutoModelForImageTextToText resolves qwen3_5 ->
|
||||
# Qwen3_5ForConditionalGeneration, so keys are born `model.language_model.*` / `model.visual.*`
|
||||
# matching the working pantheon-27b-mtp-nvfp4 reference. The vision tower loads in BF16 and is
|
||||
# ignored by the quant (re:.*visual.*); calibration is text-only (no pixel_values needed).
|
||||
# (R36 fast-seat namespace fix, 2026-07-14 — the config-merge in the prior recipe was a doomed
|
||||
# patch over a checkpoint quantized in the wrong namespace.)
|
||||
model = AutoModelForImageTextToText.from_pretrained(
|
||||
args.model, torch_dtype="auto", device_map="auto", trust_remote_code=True,
|
||||
)
|
||||
tok = AutoTokenizer.from_pretrained(args.model, trust_remote_code=True)
|
||||
|
||||
Reference in New Issue
Block a user