Files
esh-pfi-infrastructure/services/erp-seat-quant/raw/char-rp-fast-seat-verification-2026-09-10.txt
T
vh 9a916a759f Swap the MeroMero A4B onto the erp-seat seat as char-rp-fast, retire the Pfish-6 alias
Operator: "replace that a4b moe over pfish-6 -- remove the pfish-6 alias and
create an alias for char-rp-fast."

G4-MeroMero-26B-A4B-it-uncensored-heretic-NVFP4A16 is live on ana-ml2 :8021 under
its own served name, behind the new gateway alias char-rp-fast. Pfish-6 is gone
from the gateway and now returns an explicit 400 rather than a substitution; 0 of
17 LiteLLM keys scoped it, so nothing was orphaned. The compose project name stays
erp-seat because asset-engine derives seat liveness from it.

The first quant of that A4B served NaN and passed its healthcheck doing it. It was
built with the dense v2-31B recipe, whose ignore list has no router regex, so all
30 MoE routers were quantized to 4 bits -- and a 4-bit router changes which experts
run rather than degrading them. Quant rc=0, healthcheck green, correct KV pool,
correct served name, and every completion returned finish_reason=length with the
full token count and content: null. The model was emitting a full budget of tokens
that decoded to the empty string. Raw /v1/completions was empty too, ruling out the
chat template and the reasoning parser. The signal that named it was logprobs:
vLLM refused to serialize the response, "Out of range float values are not JSON
compliant: nan".

The lesson is about the control rather than the router. That tree had already been
structurally diffed and passed -- against a verified-good DENSE quant of the same
Gemma-4 family. A dense model has no routers, so the single thing that was wrong
was the single thing the control could not distinguish. Diffing instead against
Pfish-6, a known-good quant of the same 26B-A4B MoE, gave it in one line: 222
ignore entries against 252, the 30 missing being layers.N.router.proj. A positive
control is only worth what it can distinguish, and "same family" is not "same
architecture class".

Re-quantized with the MoE recipe, whose dry-run asserts 11,520 expert Linears and
refuses a router in the quantize set before any GPU time. The live seat then passed
prose with no channel-prefix leak, a solid-colour image read correctly, an auto
tool call parsed, finite logprobs, and KV 534,649 tokens / 2.04x carried over from
Pfish-6 unchanged. The broken tree is parked on ana-ml2 as
...-NVFP4A16.BROKEN-routers-quantized-20260910.

Section 4.4's temp port was not reachable: 15.9 GiB of weights plus KV plus
multimodal encoder-cache profiling does not fit in the ~19 GiB free beside GPU1's
six other tenants -- 0.20 utilization refused admission, 0.185 OOM'd in encoder
profiling. The substitute was reversibility and ordering: named .env backup, prove
the seat on its real port while no alias points at it, move the alias last. That is
why a NaN-serving seat never reached a consumer. The seat was down about 16 minutes
across two attempts; no consumer saw a broken alias.

Playbook gains the router-quant failure signature and the control-class rule in
3.15, and a logprobs check in 4.4. seat_verify.py carries that check as check 6.

Quality is NOT established: no RP eval, no long-context check, no A/B against
Pfish-6 or char-rp. Samplers are the author's card values, untuned here.
2026-09-10 11:34:05 -07:00

40 lines
2.2 KiB
Plaintext

### char-rp-fast seat verification — ana-ml2 :8021, 2026-09-10
### model: G4-MeroMero-26B-A4B-it-uncensored-heretic-NVFP4A16 (MoE-recipe re-quant)
$ docker logs vllm-erp-seat | grep 'GPU KV cache size'
(EngineCore pid=663) INFO 09-10 18:30:39 [kv_cache_utils.py:1869] GPU KV cache size: 534,649 tokens, Maximum concurrency for 262,144 tokens per request: 2.04x
$ python3 seat_verify.py http://127.0.0.1:8021/v1 <served-name>
== 1. served name + context
served: ['G4-MeroMero-26B-A4B-it-uncensored-heretic-NVFP4A16']
max_model_len: {'G4-MeroMero-26B-A4B-it-uncensored-heretic-NVFP4A16': 262144}
OK 'G4-MeroMero-26B-A4B-it-uncensored-heretic-NVFP4A16' is served
== 3. prose, non-thinking (the <|channel>thought leak)
content (277 chars): 'Oil-slicked puddles mirror the fractured glow of a flickering neon sign, casting distorted crimson light across the uneven cobblestones. The sharp, metallic tang of wet iron clings to the air as water cascades rhythmical'
reasoning_content: None
OK clean prose in content, no reasoning, no channel prefix
== 4. vision (towers preserved, tested not inferred)
answer: 'Blue' (image was solid RGB(30,60,200) = blue)
OK image was decoded and read correctly
== 5. tool call (auto)
tool_calls: [{"id": "chatcmpl-tool-ba6a1874968381f1", "type": "function", "function": {"name": "get_weather", "arguments": "{\"city\": \"Anaheim\"}"}}]
content: ''
OK parsed a get_weather call, arguments='{"city": "Anaheim"}'
== 6. logprobs (NaN logits, the router-quant tell)
text: '</b></b></b></b></b></b></b><b>'
token_logprobs: [-1.3935617208480835, -0.1289057433605194, -0.006735478527843952, -0.006430173758417368, -0.005962086841464043]
OK finite logprobs, non-empty raw text
============================================================
ALL CHECKS PASSED
### ignore-list diff vs Pfish-6 (the known-good MoE quant of the SAME architecture class)
Pfish-6 (known good) : 252 ignore entries
A4B re-quant (live) : 252 ignore entries identical to Pfish-6: True
A4B FIRST quant (bad) : 222 ignore entries missing vs good: 30
the missing ones : ['model.language_model.layers.0.router.proj', 'model.language_model.layers.1.router.proj', 'model.language_model.layers.10.router.proj'] ... (all 30 are layers.N.router.proj)