8.0 KiB
gen-seat mixed NVFP4+FP8 requant + char-rp tool-parser fix (2026-08-15, overnight)
Autonomous overnight session. Two operator-queued items, both closed.
1. char-rp / MeroMero tool-call parser (parked since the prior session)
Symptom: every tools-bearing request to char-rp (:8016) returned
400 "auto" tool choice requires --enable-auto-tool-choice and --tool-call-parser.
The seat had no tool parser configured at all — the migration from the
Magidonia GGUF seat dropped it.
Fix. MeroMero-v2 is Gemma-4 and emits its own native
<|tool_call>call:name{...}<tool_call|> syntax, not the qwen3_coder XML the
Qwen-family seats use. vLLM 0.24 ships a gemma4 tool parser whose
TOOL_CALL_START/END + CHANNEL_START/END + escape constants match this
tokenizer's etc_token/eoc_token/escape_token exactly (verified before
deploying, not assumed).
Four flags, and they are a set:
--tool-call-parser gemma4
--enable-auto-tool-choice
--reasoning-parser gemma4
--default-chat-template-kwargs '{"enable_thinking": false}'
- Without the reasoning parser, the post-tool-response turn leaks a literal
<|channel>thought\n<channel|>prefix intocontent(upstream vllm #45834 — the chat template leaves the prompt inside an open channel block). - The
enable_thinking: falseis mandatory, not cosmetic. The parser reads it fromchat_template_kwargsand defaults it toTrue(vllm/parser/gemma4.py:439). True →is_reasoning_end()returns False at a new turn → engine pre-initialises to REASONING → all plain RP prose lands inreasoning_contentandcontentcomes back null, breaking every char-rp consumer. Caught by reading the parser before deploying it. - Zero behavioural risk, proven not asserted:
chat_template.jinja:350already defaultsenable_thinkingto false, so passing it explicitly renders a byte-identical prompt — diffed across plain / with-tools / post-tool-response / system-prompt shapes before the flag went anywhere near the live seat.
Verified green: tool call (streaming + non-streaming), tool-result round-trip
(leak gone), plain prose in content with reasoning null, vision. Commit
b8f0f4c.
2. gen seat requant — the "W4A8" framing was wrong
The queued task was not servable as specified. vLLM 0.24's compressed-tensors
dispatcher (compressed_tensors.py:704-713) accepts NVFP4 weights with exactly
two activation options — None (W4A16, which forces the Marlin kernel,
kernels/linear/__init__.py:881-883) or NVFP4 (W4A4). Anything else, FP8
included, raises ValueError: For NVFP4 weights, input quantization must also be NVFP4 format. CompressedTensorsW4A8Fp8 exists but is INT4 weights
(W4A8_SUPPORTED_TYPES_MAP = {4: int4}) gated on _check_scheme_supported(90, match_exact=True) — Hopper only, so on Blackwell it is closed twice over.
The ~20% intuition was right; the scheme name was wrong. FP8 has to enter per-layer-group, not as activations on NVFP4 weights.
Two baseline corrections. The handoff's "~68 tok/s, ~42% acceptance" did not reproduce. Cache-busted (unique prompt per run — with a fixed prompt, prefix caching returns byte-identical timings and you measure nothing), the incumbent W4A16 build already did 80.12 tok/s at 47.8% acceptance — i.e. essentially at the handoff's stated W4A8 target of ~82. Had that not been re-measured the whole chase would have been declared a success for doing nothing.
The shortcut that saved hours. unsloth/Qwen3.8-27B-NVFP4 was already on-box
(pulled the previous day) — same architecture, same size, a published
mixed-precision scheme. Serving it as a probe measured +19.1% at identical MTP
acceptance — proving the gain was real and kernel-level before committing to
a requant. Its config was then read out as the reference recipe.
The recipe (byte-for-byte unsloth's, applied to the abliterated weights):
| group | scheme | targets |
|---|---|---|
group_0 |
FP8 W8A8, channel weights + per-token dynamic acts | self_attn.{q,k,v,o}_proj, linear_attn.{in_proj_qkv,in_proj_z,out_proj}, lm_head, layers 56-63 MLPs |
group_1 |
NVFP4 W4A4, tensor_group gsize16, fp8 scales, imatrix_mse weights, dynamic:"local" acts |
layers 0-55 MLP {gate,up,down}_proj |
| kv | FP8 static tensor | — |
| ignore | vision tower, linear_attn.{norm,in_proj_a,in_proj_b}, re:^mtp.* |
— |
Holding the last 8 layers' MLPs at FP8 is the accuracy trick. Targets were
made explicitly non-overlapping (group_1 enumerates 0-55) rather than trusting
group precedence, and validate_targets.py proved coverage against real module
names — 0 overlap, MLP union = layers 0-63 — before any GPU time was spent.
Results (cache-busted, bs=1):
| metric | W4A16 | mixed | delta |
|---|---|---|---|
| decode tok/s | 80.12 | 94.53 | +18.0% |
| prefill tok/s (~6.7k prompt) | 3,206 | 6,334 | +98% |
| prefill tok/s (~27k prompt) | 2,862 | 5,085 | +78% |
| TTFT on a ~27k doc | 9.43 s | 5.31 s | −44% |
| MTP acceptance | 47.8% | 47.7% | unchanged |
| perplexity (6 passages) | 6.941 | 7.059 | +1.7% worse |
| abliteration compliance | 4/4 | 4/4 | preserved |
| weights on disk | 27.7 GB | 22.5 GB | −19% |
Surface test 6/6 on the live seat (plain chat, vision, tool calling, thinking
split, 36K-token needle retrieval, streaming); all 7 LiteLLM aliases verified
routing. Commit 74f596b.
Foot-guns banked
llm-compressorPRUNESignoreentries that matched no module at quant time. The wrapper class never loads the MTP head, sore:^mtp.*matched nothing and was silently dropped from the saved config — the exact bug that cost two prior rounds (vLLM then loads the grafted BF16 MTP as quantized → uninitialised → 0% acceptance).post_quant.pynow re-injects it after the graft and re-verifies. That check fired on this run — it was not hypothetical.- Prefix caching silently fakes prefill numbers too. The prefill harness originally used a
seeded nonce, so run 2 regenerated run 1's prompts verbatim and read ~41k tok/s of
cache-hit instead of ~5k of real prefill. Same class of error as the decode bench. Use
SystemRandom; never seed a cache-busting nonce. - vLLM's
prompt_logprobsare garbage while speculative decoding is on — ~uniform over the vocab (median rank ~10⁵, logprob ≈ log(1/vocab); " Paris" after "The capital of France is" ranked 69698). Perplexity must be measured on a seat served without--speculative-config. The harness now raises rather than reporting the garbage. gen-seat/.envis mode 0600 / lkraven-owned → everydocker composecall needssudo. Without it compose failspermission deniedreading.env, leaves the old container running, and the change silently does not take — which produced one round of "benchmark results" that were just the unchanged baseline. Hard-verify againstdocker inspectargv after any such change.- GPU0 co-residency is a zero-sum budget. The smaller mixed weights meant gen
at the old util 0.45 absorbed the slack as KV (17.0 GiB / 477K tokens) and left
meromero 0.18 GiB short of its 0.52 → crash-loop. Fixed at
GEN_GPU_MEM_UTIL=0.43(15.1 GiB / 422K tokens, still 1.6× the 262K context). Both seats now 94.4/97.9 GB.
Measured negatives — do not re-chase
GEN_SPEC_TOKENSis already optimal at 3. Swept on the live seat: n=2 → 77.1, n=3 → 80.1, n=4 → 78.7, n=5 → 75.9 tok/s. Higher n trades acceptance for draft width and loses.- W4A4-everywhere was never attempted and should not be — the accuracy-safe shape is precisely the mixed one (FP8 on attention + late MLPs).
Artifacts
- Pipeline + acceptance harness + raw JSON:
services/gen-seat-mixed-quant/ - Stack docs:
stacks/gen-seat/README.md,stacks/meromero-charrp/README.md - Rollback:
sudo cp /opt/docker/compose/gen-seat/.env.bak-w4a16-20260815 …/.envthensudo docker compose up -d vllm-gen; old build untouched at/tank/aimodels/qwen38-27b-uncensored-nvfp4.