Files
esh-pfi-infrastructure/services/heretic2-nvfp4-quant/finalize_modelopt_mtp.py
T
vh 982c319d9f feat(heretic2-nvfp4): WORKING modelopt NVFP4+MTP seat + full recipe runbook
The fast char-rp-reasoning seat works: ~77 tok/s (vs GGUF ~59.5, base NVFP4 ~53),
MTP draft-acceptance 32-40%, mean acceptance length 2.19. Same Heretic2/NEO-CODE
model, NVFP4 + native qwen3_5_mtp spec-decode.

Full end-to-end recipe + the four landmines in docs/runbooks/heretic2-nvfp4-mtp-seat.md:
(1) load as AutoModelForImageTextToText not AutoModelForCausalLM (namespace/gibberish);
(2) modelopt format not compressed-tensors (compressed-tensors MTP = 0% accept);
(3) modelopt 0.45 <-> transformers 5.12.1 FusedMoE crash (guarded in quant_modelopt.py);
(4) vLLM 0.24.0 does NOT propagate modelopt exclude_modules to the spec-decode draft
model -> BF16 mtp head gets quantized -> shape crash; no checkpoint config fixes it
(is_layer_skipped is exact-membership not glob) -> fix is a mounted sitecustomize that
force-skips mtp.* in is_layer_skipped (upstream vLLM bug to report).

Scripts: quant_modelopt.py (FusedMoE guard + single-shard export + multimodal load),
finalize_modelopt_mtp.py (splice bf16 mtp), serve_modelopt_mtp.sh, run_quant_modelopt.sh,
sitecustomize-mtp-workaround.py.
2026-07-14 14:41:48 -07:00

56 lines
2.3 KiB
Python

#!/usr/bin/env python3
"""Splice the 15 BF16 mtp.* tensors into the modelopt NVFP4 quant output → the servable seat.
Runs AFTER quant_modelopt.py. Copies heretic2-modelopt-nvfp4 → heretic2-modelopt-nvfp4-mtp,
splices the grafted BF16 mtp head into the single shard (transformers never builds an mtp module
at load, so mtp is always post-quant-spliced — same as AEON/pantheon), and adds the mtp module
names to config.json exclude_modules for tidiness. NOTE: the actual thing that keeps the mtp head
BF16 at serve time is the sitecustomize MTP workaround (see runbook landmine #4); the config
exclude here is belt-and-suspenders and does NOT by itself prevent the drafter-quant crash.
Run in a vLLM container (root; /tank/aimodels files are root-owned):
docker run --rm -v /tank/aimodels:/tank/aimodels -v /home/lkraven:/lk \
--entrypoint python3 vllm/vllm-openai:v0.24.0 /lk/finalize_modelopt_mtp.py
"""
import json
import os
import shutil
from safetensors import safe_open
from safetensors.torch import save_file
WORK = "/tank/aimodels/heretic2-nvfp4-work"
SRC = f"{WORK}/heretic2-modelopt-nvfp4"
DST = f"{WORK}/heretic2-modelopt-nvfp4-mtp"
GRAFT = f"{WORK}/heretic2-mtp-bf16"
if os.path.exists(DST):
shutil.rmtree(DST)
print(f"copying {SRC} -> {DST}", flush=True)
shutil.copytree(SRC, DST)
out_st = f"{DST}/model.safetensors" # single shard (quant_modelopt.py forces max_shard_size huge)
mtp_st = f"{GRAFT}/model-mtp.safetensors"
tensors = {}
with safe_open(out_st, framework="pt") as f:
for k in f.keys():
tensors[k] = f.get_tensor(k)
n_main = len(tensors)
with safe_open(mtp_st, framework="pt") as f:
mtp_keys = list(f.keys())
for k in mtp_keys:
tensors[k] = f.get_tensor(k)
assert not any("mtp" in k.lower() for k in list(tensors)[:n_main]), "output already had mtp?"
save_file(tensors, out_st, metadata={"format": "pt"})
print(f"spliced {len(mtp_keys)} bf16 mtp tensors -> {len(tensors)} total", flush=True)
cfgp = f"{DST}/config.json"
cfg = json.load(open(cfgp))
qc = cfg.setdefault("quantization_config", {})
exc = qc.setdefault("exclude_modules", [])
mtp_mods = sorted({k.rsplit(".", 1)[0] for k in mtp_keys})
exc.extend(m for m in mtp_mods if m not in exc)
json.dump(cfg, open(cfgp, "w"), indent=2)
print(f"config exclude_modules += {len(mtp_mods)} mtp modules; total {len(exc)}", flush=True)
print(f"DONE: {DST}", flush=True)