dc3e47b3a2
The char-rp-reasoning seat on ana-ml2 GPU0 — NEO-CODE Heretic2 27B at modelopt NVFP4 with a grafted BF16 MTP head, ~77 tok/s via qwen3_5_mtp spec-decode, replacing the retired GGUF seat. It had been running untracked. Includes conf/mtp-workaround/sitecustomize.py, which is not optional: vLLM 0.24.0 does not propagate modelopt exclude_modules to the spec-decode DRAFT model, so the BF16 MTP head gets quantized and the engine dies at load. The shim force-skips mtp.* in is_layer_skipped. Both the mount and PYTHONPATH are load-bearing. Adds the two files house convention expects and the directory lacked: a .env.example naming every knob (all values are the compose defaults; the live host overrides only the three VRAM ones) and a README that points at docs/runbooks/heretic2-nvfp4-mtp-seat.md rather than duplicating it. No secrets: API_KEY is empty by default and the real .env stays on the host.
49 lines
2.2 KiB
Markdown
49 lines
2.2 KiB
Markdown
# heretic2-charrp-reasoning — NVFP4 + native MTP reasoning seat (ana-ml2)
|
|
|
|
The `char-rp-reasoning` seat: NEO-CODE Heretic2 27B quantized to modelopt NVFP4
|
|
with a grafted BF16 MTP head, served by vLLM on **ana-ml2 GPU0**, port **8018**.
|
|
Replaces the retired GGUF seat (`llama-charrp-reasoning`) at roughly **77 tok/s
|
|
(~1.3x)** via `qwen3_5_mtp` speculative decode.
|
|
|
|
- **Served model name:** `char-rp-reasoning` — the gateway alias consumers use
|
|
- **Endpoint:** `http://10.250.50.54:8018` (`/docs` for the card link)
|
|
|
|
## ⚠️ It does not boot without the MTP workaround
|
|
|
|
vLLM 0.24.0 does not propagate modelopt's `exclude_modules` to the **draft**
|
|
model in a spec-decode config, so the BF16 MTP head gets quantized along with
|
|
everything else and the engine dies at load on a shape mismatch.
|
|
|
|
`conf/mtp-workaround/sitecustomize.py` is mounted at `PYTHONPATH` and patches
|
|
`is_layer_skipped` to force-skip `mtp.*`, keeping the head BF16. **The mount and
|
|
the `PYTHONPATH` env are both load-bearing** — remove either and the seat
|
|
crash-loops at startup with an error that looks like a bad quant rather than a
|
|
missing shim.
|
|
|
|
Canonical copy of the shim and the reasoning behind it live in
|
|
`services/heretic2-nvfp4-quant/`; the full build-and-serve recipe, including the
|
|
other landmines hit on the way, is in
|
|
[`docs/runbooks/heretic2-nvfp4-mtp-seat.md`](../../docs/runbooks/heretic2-nvfp4-mtp-seat.md).
|
|
Read the runbook before changing anything here — this README is the pointer, not
|
|
the spec.
|
|
|
|
## VRAM is the constraint, and it is shared
|
|
|
|
GPU0 also hosts `vllm-aeon-gen` (`gen`) and `llama-charrp` (`char-rp`). NVFP4
|
|
27B weights are ~26 GB plus KV, which is why `--gpu-memory-utilization` sits at
|
|
**0.30** and `--max-model-len` at **32768** — the old GGUF seat managed 256K on
|
|
much lighter Q5 weights. **Raising either without rebalancing GPU0 first will
|
|
OOM the neighbours, not just this container.** Tunables are in `.env`; see
|
|
`.env.example`.
|
|
|
|
## Deploy
|
|
|
|
```bash
|
|
scripts/deploy-stack.sh ana-ml2 heretic2-charrp-reasoning
|
|
ssh infra-ops@10.250.50.54 \
|
|
'cd /opt/docker/compose/heretic2-charrp-reasoning && sudo docker compose up -d'
|
|
```
|
|
|
|
First start is slow — the healthcheck allows a 600s `start_period` because
|
|
loading NVFP4 weights plus the draft head takes minutes.
|