Files
esh-pfi-infrastructure/persistent-memory.d/2026-08-15-gen-seat-mixed-requant.md
T

148 lines
8.0 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# gen-seat mixed NVFP4+FP8 requant + char-rp tool-parser fix (2026-08-15, overnight)
Autonomous overnight session. Two operator-queued items, both closed.
## 1. char-rp / MeroMero tool-call parser (parked since the prior session)
**Symptom:** every tools-bearing request to `char-rp` (:8016) returned
`400 "auto" tool choice requires --enable-auto-tool-choice and --tool-call-parser`.
The seat had **no tool parser configured at all** — the migration from the
Magidonia GGUF seat dropped it.
**Fix.** MeroMero-v2 is Gemma-4 and emits its own native
`<|tool_call>call:name{...}<tool_call|>` syntax, not the qwen3_coder XML the
Qwen-family seats use. vLLM 0.24 ships a `gemma4` tool parser whose
TOOL_CALL_START/END + CHANNEL_START/END + escape constants match this
tokenizer's `etc_token`/`eoc_token`/`escape_token` exactly (verified before
deploying, not assumed).
Four flags, and they are a **set**:
```
--tool-call-parser gemma4
--enable-auto-tool-choice
--reasoning-parser gemma4
--default-chat-template-kwargs '{"enable_thinking": false}'
```
- Without the **reasoning parser**, the post-tool-response turn leaks a literal
`<|channel>thought\n<channel|>` prefix into `content` (upstream vllm #45834
the chat template leaves the prompt inside an open channel block).
- The **`enable_thinking: false`** is mandatory, not cosmetic. The parser reads
it from `chat_template_kwargs` and **defaults it to `True`**
(`vllm/parser/gemma4.py:439`). True → `is_reasoning_end()` returns False at a
new turn → engine pre-initialises to REASONING → **all plain RP prose lands in
`reasoning_content` and `content` comes back null**, breaking every char-rp
consumer. Caught by reading the parser before deploying it.
- **Zero behavioural risk, proven not asserted:** `chat_template.jinja:350`
already defaults `enable_thinking` to false, so passing it explicitly renders a
**byte-identical prompt** — diffed across plain / with-tools / post-tool-response
/ system-prompt shapes before the flag went anywhere near the live seat.
Verified green: tool call (streaming + non-streaming), tool-result round-trip
(leak gone), plain prose in `content` with `reasoning` null, vision. Commit
`b8f0f4c`.
## 2. gen seat requant — the "W4A8" framing was wrong
**The queued task was not servable as specified.** vLLM 0.24's compressed-tensors
dispatcher (`compressed_tensors.py:704-713`) accepts NVFP4 weights with exactly
two activation options — `None` (W4A16, which **forces the Marlin kernel**,
`kernels/linear/__init__.py:881-883`) or NVFP4 (W4A4). Anything else, FP8
included, raises `ValueError: For NVFP4 weights, input quantization must also be
NVFP4 format`. `CompressedTensorsW4A8Fp8` exists but is **INT4** weights
(`W4A8_SUPPORTED_TYPES_MAP = {4: int4}`) gated on `_check_scheme_supported(90,
match_exact=True)` — Hopper only, so on Blackwell it is closed twice over.
The ~20% intuition was right; the *scheme name* was wrong. FP8 has to enter
**per-layer-group**, not as activations on NVFP4 weights.
**Two baseline corrections.** The handoff's "~68 tok/s, ~42% acceptance" did not
reproduce. Cache-busted (unique prompt per run — with a fixed prompt, prefix
caching returns byte-identical timings and you measure nothing), the incumbent
W4A16 build already did **80.12 tok/s at 47.8% acceptance** — i.e. essentially
*at* the handoff's stated W4A8 target of ~82. Had that not been re-measured the
whole chase would have been declared a success for doing nothing.
**The shortcut that saved hours.** `unsloth/Qwen3.8-27B-NVFP4` was already on-box
(pulled the previous day) — same architecture, same size, a published
mixed-precision scheme. Serving it as a probe measured **+19.1% at identical MTP
acceptance** — proving the gain was real and kernel-level *before* committing to
a requant. Its config was then read out as the reference recipe.
**The recipe** (byte-for-byte unsloth's, applied to the abliterated weights):
| group | scheme | targets |
|---|---|---|
| `group_0` | FP8 W8A8, channel weights + per-token dynamic acts | `self_attn.{q,k,v,o}_proj`, `linear_attn.{in_proj_qkv,in_proj_z,out_proj}`, `lm_head`, **layers 56-63** MLPs |
| `group_1` | NVFP4 W4A4, tensor_group gsize16, fp8 scales, `imatrix_mse` weights, `dynamic:"local"` acts | **layers 0-55** MLP `{gate,up,down}_proj` |
| kv | FP8 static tensor | — |
| ignore | vision tower, `linear_attn.{norm,in_proj_a,in_proj_b}`, `re:^mtp.*` | — |
Holding the **last 8 layers' MLPs at FP8** is the accuracy trick. Targets were
made explicitly non-overlapping (group_1 enumerates 0-55) rather than trusting
group precedence, and `validate_targets.py` proved coverage against real module
names — 0 overlap, MLP union = layers 0-63 — before any GPU time was spent.
**Results (cache-busted, bs=1):**
| metric | W4A16 | mixed | delta |
|---|---|---|---|
| decode tok/s | 80.12 | **94.53** | **+18.0%** |
| prefill tok/s (~6.7k prompt) | 3,206 | **6,334** | **+98%** |
| prefill tok/s (~27k prompt) | 2,862 | **5,085** | **+78%** |
| TTFT on a ~27k doc | 9.43 s | **5.31 s** | 44% |
| MTP acceptance | 47.8% | 47.7% | unchanged |
| perplexity (6 passages) | 6.941 | 7.059 | +1.7% worse |
| abliteration compliance | 4/4 | 4/4 | preserved |
| weights on disk | 27.7 GB | 22.5 GB | 19% |
Surface test 6/6 on the live seat (plain chat, vision, tool calling, thinking
split, 36K-token needle retrieval, streaming); all 7 LiteLLM aliases verified
routing. Commit `74f596b`.
## Foot-guns banked
- **`llm-compressor` PRUNES `ignore` entries that matched no module at quant
time.** The wrapper class never loads the MTP head, so `re:^mtp.*` matched
nothing and was silently dropped from the saved config — the exact bug that
cost two prior rounds (vLLM then loads the grafted BF16 MTP as quantized →
uninitialised → 0% acceptance). `post_quant.py` now **re-injects it after the
graft and re-verifies**. That check *fired on this run* — it was not
hypothetical.
- **Prefix caching silently fakes prefill numbers too.** The prefill harness originally used a
*seeded* nonce, so run 2 regenerated run 1's prompts verbatim and read **~41k tok/s of
cache-hit** instead of ~5k of real prefill. Same class of error as the decode bench. Use
`SystemRandom`; never seed a cache-busting nonce.
- **vLLM's `prompt_logprobs` are garbage while speculative decoding is on** —
~uniform over the vocab (median rank ~10⁵, logprob ≈ log(1/vocab); " Paris"
after "The capital of France is" ranked 69698). Perplexity must be measured on
a seat served **without** `--speculative-config`. The harness now raises rather
than reporting the garbage.
- **`gen-seat/.env` is mode 0600 / lkraven-owned** → *every* `docker compose`
call needs `sudo`. Without it compose fails `permission denied` reading `.env`,
**leaves the old container running**, and the change silently does not take —
which produced one round of "benchmark results" that were just the unchanged
baseline. Hard-verify against `docker inspect` argv after any such change.
- **GPU0 co-residency is a zero-sum budget.** The smaller mixed weights meant gen
at the old util 0.45 absorbed the slack as KV (17.0 GiB / 477K tokens) and left
meromero **0.18 GiB** short of its 0.52 → crash-loop. Fixed at
`GEN_GPU_MEM_UTIL=0.43` (15.1 GiB / 422K tokens, still 1.6× the 262K context).
Both seats now 94.4/97.9 GB.
## Measured negatives — do not re-chase
- **`GEN_SPEC_TOKENS` is already optimal at 3.** Swept on the live seat:
n=2 → 77.1, **n=3 → 80.1**, n=4 → 78.7, n=5 → 75.9 tok/s. Higher n trades
acceptance for draft width and loses.
- **W4A4-everywhere was never attempted** and should not be — the accuracy-safe
shape is precisely the mixed one (FP8 on attention + late MLPs).
## Artifacts
- Pipeline + acceptance harness + raw JSON: `services/gen-seat-mixed-quant/`
- Stack docs: `stacks/gen-seat/README.md`, `stacks/meromero-charrp/README.md`
- Rollback: `sudo cp /opt/docker/compose/gen-seat/.env.bak-w4a16-20260815 …/.env`
then `sudo docker compose up -d vllm-gen`; old build untouched at
`/tank/aimodels/qwen38-27b-uncensored-nvfp4`.