# heretic2-charrp-reasoning — NVFP4 + native MTP reasoning seat (fv-ml1) The `char-rp-reasoning` seat: NEO-CODE Heretic2 27B quantized to modelopt NVFP4 with a grafted BF16 MTP head, served by vLLM on **fv-ml1 GPU0**, port **8018**. Replaces the retired GGUF seat (`llama-charrp-reasoning`) at roughly **77 tok/s (~1.3x)** via `qwen3_5_mtp` speculative decode. - **Served model name:** `char-rp-reasoning` — the gateway alias consumers use - **Endpoint:** `http://10.251.50.54:8018` (`/docs` for the card link) ## ⚠️ It does not boot without the MTP workaround vLLM 0.24.0 does not propagate modelopt's `exclude_modules` to the **draft** model in a spec-decode config, so the BF16 MTP head gets quantized along with everything else and the engine dies at load on a shape mismatch. `conf/mtp-workaround/sitecustomize.py` is mounted at `PYTHONPATH` and patches `is_layer_skipped` to force-skip `mtp.*`, keeping the head BF16. **The mount and the `PYTHONPATH` env are both load-bearing** — remove either and the seat crash-loops at startup with an error that looks like a bad quant rather than a missing shim. Canonical copy of the shim and the reasoning behind it live in `services/heretic2-nvfp4-quant/`; the full build-and-serve recipe, including the other landmines hit on the way, is in [`docs/runbooks/heretic2-nvfp4-mtp-seat.md`](../../docs/runbooks/heretic2-nvfp4-mtp-seat.md). Read the runbook before changing anything here — this README is the pointer, not the spec. ## VRAM is the constraint, and it is shared GPU0 also hosts `vllm-aeon-gen` (`gen`) and `llama-charrp` (`char-rp`). NVFP4 27B weights are ~26 GB plus KV, which is why `--gpu-memory-utilization` sits at **0.30** and `--max-model-len` at **32768** — the old GGUF seat managed 256K on much lighter Q5 weights. **Raising either without rebalancing GPU0 first will OOM the neighbours, not just this container.** Tunables are in `.env`; see `.env.example`. ## Deploy ```bash scripts/deploy-stack.sh fv-ml1 heretic2-charrp-reasoning ssh infra-ops@10.251.50.54 \ 'cd /opt/docker/compose/heretic2-charrp-reasoning && sudo docker compose up -d' ``` First start is slow — the healthcheck allows a 600s `start_period` because loading NVFP4 weights plus the draft head takes minutes.