The char-rp-reasoning seat on ana-ml2 GPU0 — NEO-CODE Heretic2 27B at modelopt NVFP4 with a grafted BF16 MTP head, ~77 tok/s via qwen3_5_mtp spec-decode, replacing the retired GGUF seat. It had been running untracked. Includes conf/mtp-workaround/sitecustomize.py, which is not optional: vLLM 0.24.0 does not propagate modelopt exclude_modules to the spec-decode DRAFT model, so the BF16 MTP head gets quantized and the engine dies at load. The shim force-skips mtp.* in is_layer_skipped. Both the mount and PYTHONPATH are load-bearing. Adds the two files house convention expects and the directory lacked: a .env.example naming every knob (all values are the compose defaults; the live host overrides only the three VRAM ones) and a README that points at docs/runbooks/heretic2-nvfp4-mtp-seat.md rather than duplicating it. No secrets: API_KEY is empty by default and the real .env stays on the host.
heretic2-charrp-reasoning — NVFP4 + native MTP reasoning seat (ana-ml2)
The char-rp-reasoning seat: NEO-CODE Heretic2 27B quantized to modelopt NVFP4
with a grafted BF16 MTP head, served by vLLM on ana-ml2 GPU0, port 8018.
Replaces the retired GGUF seat (llama-charrp-reasoning) at roughly 77 tok/s
(~1.3x) via qwen3_5_mtp speculative decode.
- Served model name:
char-rp-reasoning— the gateway alias consumers use - Endpoint:
http://10.250.50.54:8018(/docsfor the card link)
⚠️ It does not boot without the MTP workaround
vLLM 0.24.0 does not propagate modelopt's exclude_modules to the draft
model in a spec-decode config, so the BF16 MTP head gets quantized along with
everything else and the engine dies at load on a shape mismatch.
conf/mtp-workaround/sitecustomize.py is mounted at PYTHONPATH and patches
is_layer_skipped to force-skip mtp.*, keeping the head BF16. The mount and
the PYTHONPATH env are both load-bearing — remove either and the seat
crash-loops at startup with an error that looks like a bad quant rather than a
missing shim.
Canonical copy of the shim and the reasoning behind it live in
services/heretic2-nvfp4-quant/; the full build-and-serve recipe, including the
other landmines hit on the way, is in
docs/runbooks/heretic2-nvfp4-mtp-seat.md.
Read the runbook before changing anything here — this README is the pointer, not
the spec.
VRAM is the constraint, and it is shared
GPU0 also hosts vllm-aeon-gen (gen) and llama-charrp (char-rp). NVFP4
27B weights are ~26 GB plus KV, which is why --gpu-memory-utilization sits at
0.30 and --max-model-len at 32768 — the old GGUF seat managed 256K on
much lighter Q5 weights. Raising either without rebalancing GPU0 first will
OOM the neighbours, not just this container. Tunables are in .env; see
.env.example.
Deploy
scripts/deploy-stack.sh ana-ml2 heretic2-charrp-reasoning
ssh infra-ops@10.250.50.54 \
'cd /opt/docker/compose/heretic2-charrp-reasoning && sudo docker compose up -d'
First start is slow — the healthcheck allows a 600s start_period because
loading NVFP4 weights plus the draft head takes minutes.