Swap the MeroMero A4B onto the erp-seat seat as char-rp-fast, retire the Pfish-6 alias

Operator: "replace that a4b moe over pfish-6 -- remove the pfish-6 alias and
create an alias for char-rp-fast."

G4-MeroMero-26B-A4B-it-uncensored-heretic-NVFP4A16 is live on ana-ml2 :8021 under
its own served name, behind the new gateway alias char-rp-fast. Pfish-6 is gone
from the gateway and now returns an explicit 400 rather than a substitution; 0 of
17 LiteLLM keys scoped it, so nothing was orphaned. The compose project name stays
erp-seat because asset-engine derives seat liveness from it.

The first quant of that A4B served NaN and passed its healthcheck doing it. It was
built with the dense v2-31B recipe, whose ignore list has no router regex, so all
30 MoE routers were quantized to 4 bits -- and a 4-bit router changes which experts
run rather than degrading them. Quant rc=0, healthcheck green, correct KV pool,
correct served name, and every completion returned finish_reason=length with the
full token count and content: null. The model was emitting a full budget of tokens
that decoded to the empty string. Raw /v1/completions was empty too, ruling out the
chat template and the reasoning parser. The signal that named it was logprobs:
vLLM refused to serialize the response, "Out of range float values are not JSON
compliant: nan".

The lesson is about the control rather than the router. That tree had already been
structurally diffed and passed -- against a verified-good DENSE quant of the same
Gemma-4 family. A dense model has no routers, so the single thing that was wrong
was the single thing the control could not distinguish. Diffing instead against
Pfish-6, a known-good quant of the same 26B-A4B MoE, gave it in one line: 222
ignore entries against 252, the 30 missing being layers.N.router.proj. A positive
control is only worth what it can distinguish, and "same family" is not "same
architecture class".

Re-quantized with the MoE recipe, whose dry-run asserts 11,520 expert Linears and
refuses a router in the quantize set before any GPU time. The live seat then passed
prose with no channel-prefix leak, a solid-colour image read correctly, an auto
tool call parsed, finite logprobs, and KV 534,649 tokens / 2.04x carried over from
Pfish-6 unchanged. The broken tree is parked on ana-ml2 as
...-NVFP4A16.BROKEN-routers-quantized-20260910.

Section 4.4's temp port was not reachable: 15.9 GiB of weights plus KV plus
multimodal encoder-cache profiling does not fit in the ~19 GiB free beside GPU1's
six other tenants -- 0.20 utilization refused admission, 0.185 OOM'd in encoder
profiling. The substitute was reversibility and ordering: named .env backup, prove
the seat on its real port while no alias points at it, move the alias last. That is
why a NaN-serving seat never reached a consumer. The seat was down about 16 minutes
across two attempts; no consumer saw a broken alias.

Playbook gains the router-quant failure signature and the control-class rule in
3.15, and a logprobs check in 4.4. seat_verify.py carries that check as check 6.

Quality is NOT established: no RP eval, no long-context check, no A/B against
Pfish-6 or char-rp. Samplers are the author's card values, untuned here.
This commit is contained in:
vh
2026-09-10 11:34:05 -07:00
parent 1a5bc2ddf1
commit 9a916a759f
9 changed files with 588 additions and 41 deletions
+43 -13
View File
@@ -872,9 +872,10 @@ model_list:
# base profile ~0% on 30/35 axes per brokkr) -- this seat is for the operator's ear; treat
# it as unrated on every safety axis.
#
# ROLLBACK: /tank/aimodels/erp-tune-v6-nvfp4a16 is still on disk, and the previous host env
# is at /tmp/erp-seat-env.v6.bak on ana-ml2 -- flip ERP_MODEL/ERP_SERVED_NAME/ERP_CHAT_TEMPLATE
# in /opt/docker/compose/erp-seat/.env back to v6 and `docker compose up -d`.
# ROLLBACK (updated 2026-09-10, the seat now serves MeroMero A4B): Pfish-6 itself is
# the rollback target. /tank/aimodels/erp-tune-v6-nvfp4a16 is still on disk and the
# pre-swap host env is at /opt/docker/compose/erp-seat/.env.pfish6.bak-20260910 --
# `cp .env.pfish6.bak-20260910 .env && docker compose up -d` restores Pfish-6 in ~4 min.
#
# THIS GATEWAY IS THE SHARED-KEY SURFACE: `all-agents-local` reaches every model here,
# in every session and project. Removing this alias does not remove the operator's
@@ -882,21 +883,50 @@ model_list:
#
# Same-site: seat and gateway are both at Anaheim (local hop, no mesh crossing).
# Pfish-6 — the STANDING seat as of 2026-09-09. NVFP4A16 quant of the run-6
# LoRA merge on the jenerallee78 ARA abliteration, on ana-ml2 GPU1 (:8021).
# Operator ruling: "declare run 6 as Pfish-6 ... we're gonna stay on 6 for now."
# char-rp-fast — the MeroMero A4B MoE, on ana-ml2 GPU1 (:8021). Operator, 2026-09-10:
# "replace that a4b moe over pfish-6 -- remove the pfish-6 alias and create an alias
# for char-rp-fast." REPLACES the `Pfish-6` alias, which is removed with this change.
#
# REPLACES the `trial` alias, which is retired with run 7. Run 7's CSAM gate
# failure turned out to be a DETECTOR BUG (the adjective "minor" matching a
# HARD rule — fixed cc42d76 in brokkr-smithy), but run 7 was independently a
# poor run and is not coming back.
# The seat itself is unchanged in every dimension that matters to a caller: the A4B is
# 30 layers / kv 8 / sliding_window 1024 / 128 experts top-8 — field for field the same
# geometry as Pfish-6 — so the 9.114 GB KV pinning transfers exactly and the seat still
# reports 534,649 tokens and 2.04x concurrency at 262,144. That was verified from the
# engine log, not assumed, because KV-per-token is normally NOT transferable.
#
# Served under its TRUE name. This is now a named seat, not a trial.
- model_name: Pfish-6
# ⚠ Pfish-6 IS GONE from this gateway and its seat no longer serves that name. A caller
# still asking for `Pfish-6` gets a clean 404 rather than a silent substitution, which is
# the intended behaviour. The artifact is still on disk (see the ROLLBACK note above).
#
# SAMPLERS, and they are the author's, not inherited: the model card states Temp 0.8-1.0
# and MinP 0.05. Temperature/top_p/top_k already come from the tree's own
# generation_config (1.0 / 0.95 / 64) which vLLM applies server-side, and 1.0 sits at the
# top of the card's stated range, so the only value that needs stating here is min_p.
# Deliberately NOT copying char-rp's temp 1.1 / min_p 0.10 — those were A/B-tuned against
# a retired Mistral seat and are inherited, not canonical, as that block's own note says.
#
# enable_thinking:false is ALSO pinned process-level on the seat
# (--default-chat-template-kwargs). Stated here as well so the intent is visible at the
# routing layer: without it the gemma4 reasoning parser pre-initialises to REASONING and
# plain prose comes back with a null `content`.
#
# ⚠ ITS FIRST QUANT WAS BROKEN AND SERVED NaN. The 2026-09-10 08:15 build was made with
# the DENSE recipe, whose ignore list has no router regex, so all 30 MoE routers were
# quantized to 4 bits — a 4-bit router picks different experts (playbook §3.15). The seat
# came up healthy, answered every request with 120 tokens, and decoded to the empty string;
# logprobs were NaN. Re-quantized with services/erp-seat-quant/quant_nvfp4a16_gemma4_moe.py,
# whose target guard refuses that exact mistake. The broken tree is parked on ana-ml2 as
# ...-NVFP4A16.BROKEN-routers-quantized-20260910. Do not serve it.
- model_name: char-rp-fast
litellm_params:
model: hosted_vllm/Pfish-6
model: hosted_vllm/G4-MeroMero-26B-A4B-it-uncensored-heretic-NVFP4A16
api_base: http://10.250.50.54:8021/v1
api_key: os.environ/VLLM_API_KEY
extra_body:
min_p: 0.05
chat_template_kwargs:
enable_thinking: false
model_info:
mode: chat
general_settings:
master_key: os.environ/LITELLM_MASTER_KEY