Files
vh dc3e47b3a2 feat(heretic2-charrp-reasoning): track the NVFP4+MTP reasoning seat
The char-rp-reasoning seat on ana-ml2 GPU0 — NEO-CODE Heretic2 27B at modelopt
NVFP4 with a grafted BF16 MTP head, ~77 tok/s via qwen3_5_mtp spec-decode,
replacing the retired GGUF seat. It had been running untracked.

Includes conf/mtp-workaround/sitecustomize.py, which is not optional: vLLM
0.24.0 does not propagate modelopt exclude_modules to the spec-decode DRAFT
model, so the BF16 MTP head gets quantized and the engine dies at load. The
shim force-skips mtp.* in is_layer_skipped. Both the mount and PYTHONPATH are
load-bearing.

Adds the two files house convention expects and the directory lacked: a
.env.example naming every knob (all values are the compose defaults; the live
host overrides only the three VRAM ones) and a README that points at
docs/runbooks/heretic2-nvfp4-mtp-seat.md rather than duplicating it.

No secrets: API_KEY is empty by default and the real .env stays on the host.
2026-08-18 23:53:59 -07:00
..

heretic2-charrp-reasoning — NVFP4 + native MTP reasoning seat (ana-ml2)

The char-rp-reasoning seat: NEO-CODE Heretic2 27B quantized to modelopt NVFP4 with a grafted BF16 MTP head, served by vLLM on ana-ml2 GPU0, port 8018. Replaces the retired GGUF seat (llama-charrp-reasoning) at roughly 77 tok/s (~1.3x) via qwen3_5_mtp speculative decode.

  • Served model name: char-rp-reasoning — the gateway alias consumers use
  • Endpoint: http://10.250.50.54:8018 (/docs for the card link)

⚠️ It does not boot without the MTP workaround

vLLM 0.24.0 does not propagate modelopt's exclude_modules to the draft model in a spec-decode config, so the BF16 MTP head gets quantized along with everything else and the engine dies at load on a shape mismatch.

conf/mtp-workaround/sitecustomize.py is mounted at PYTHONPATH and patches is_layer_skipped to force-skip mtp.*, keeping the head BF16. The mount and the PYTHONPATH env are both load-bearing — remove either and the seat crash-loops at startup with an error that looks like a bad quant rather than a missing shim.

Canonical copy of the shim and the reasoning behind it live in services/heretic2-nvfp4-quant/; the full build-and-serve recipe, including the other landmines hit on the way, is in docs/runbooks/heretic2-nvfp4-mtp-seat.md. Read the runbook before changing anything here — this README is the pointer, not the spec.

VRAM is the constraint, and it is shared

GPU0 also hosts vllm-aeon-gen (gen) and llama-charrp (char-rp). NVFP4 27B weights are ~26 GB plus KV, which is why --gpu-memory-utilization sits at 0.30 and --max-model-len at 32768 — the old GGUF seat managed 256K on much lighter Q5 weights. Raising either without rebalancing GPU0 first will OOM the neighbours, not just this container. Tunables are in .env; see .env.example.

Deploy

scripts/deploy-stack.sh ana-ml2 heretic2-charrp-reasoning
ssh infra-ops@10.250.50.54 \
  'cd /opt/docker/compose/heretic2-charrp-reasoning && sudo docker compose up -d'

First start is slow — the healthcheck allows a 600s start_period because loading NVFP4 weights plus the draft head takes minutes.