Swap the MeroMero A4B onto the erp-seat seat as char-rp-fast, retire the Pfish-6 alias
Operator: "replace that a4b moe over pfish-6 -- remove the pfish-6 alias and create an alias for char-rp-fast." G4-MeroMero-26B-A4B-it-uncensored-heretic-NVFP4A16 is live on ana-ml2 :8021 under its own served name, behind the new gateway alias char-rp-fast. Pfish-6 is gone from the gateway and now returns an explicit 400 rather than a substitution; 0 of 17 LiteLLM keys scoped it, so nothing was orphaned. The compose project name stays erp-seat because asset-engine derives seat liveness from it. The first quant of that A4B served NaN and passed its healthcheck doing it. It was built with the dense v2-31B recipe, whose ignore list has no router regex, so all 30 MoE routers were quantized to 4 bits -- and a 4-bit router changes which experts run rather than degrading them. Quant rc=0, healthcheck green, correct KV pool, correct served name, and every completion returned finish_reason=length with the full token count and content: null. The model was emitting a full budget of tokens that decoded to the empty string. Raw /v1/completions was empty too, ruling out the chat template and the reasoning parser. The signal that named it was logprobs: vLLM refused to serialize the response, "Out of range float values are not JSON compliant: nan". The lesson is about the control rather than the router. That tree had already been structurally diffed and passed -- against a verified-good DENSE quant of the same Gemma-4 family. A dense model has no routers, so the single thing that was wrong was the single thing the control could not distinguish. Diffing instead against Pfish-6, a known-good quant of the same 26B-A4B MoE, gave it in one line: 222 ignore entries against 252, the 30 missing being layers.N.router.proj. A positive control is only worth what it can distinguish, and "same family" is not "same architecture class". Re-quantized with the MoE recipe, whose dry-run asserts 11,520 expert Linears and refuses a router in the quantize set before any GPU time. The live seat then passed prose with no channel-prefix leak, a solid-colour image read correctly, an auto tool call parsed, finite logprobs, and KV 534,649 tokens / 2.04x carried over from Pfish-6 unchanged. The broken tree is parked on ana-ml2 as ...-NVFP4A16.BROKEN-routers-quantized-20260910. Section 4.4's temp port was not reachable: 15.9 GiB of weights plus KV plus multimodal encoder-cache profiling does not fit in the ~19 GiB free beside GPU1's six other tenants -- 0.20 utilization refused admission, 0.185 OOM'd in encoder profiling. The substitute was reversibility and ordering: named .env backup, prove the seat on its real port while no alias points at it, move the alias last. That is why a NaN-serving seat never reached a consumer. The seat was down about 16 minutes across two attempts; no consumer saw a broken alias. Playbook gains the router-quant failure signature and the control-class rule in 3.15, and a logprobs check in 4.4. seat_verify.py carries that check as check 6. Quality is NOT established: no RP eval, no long-context check, no A/B against Pfish-6 or char-rp. Samplers are the author's card values, untuned here.
This commit is contained in:
@@ -1,12 +1,15 @@
|
||||
# erp-seat — ana-ml2 GPU1. Real .env lives on the host at /opt/docker/compose/erp-seat/.env.
|
||||
ERP_IMAGE=vllm/vllm-openai:nightly-311b3513af33bc29b4acb2fde2e9313e5e9966a0
|
||||
ERP_MODEL=/tank/aimodels/erp-tune-v6-nvfp4a16
|
||||
ERP_SERVED_NAME=erp-tune-v6-nvfp4a16
|
||||
ERP_CHAT_TEMPLATE=/tank/aimodels/erp-tune-v6-nvfp4a16/chat_template.jinja
|
||||
ERP_MODEL=/tank/aimodels/G4-MeroMero-26B-A4B-it-uncensored-heretic-NVFP4A16
|
||||
ERP_SERVED_NAME=G4-MeroMero-26B-A4B-it-uncensored-heretic-NVFP4A16
|
||||
ERP_CHAT_TEMPLATE=/tank/aimodels/G4-MeroMero-26B-A4B-it-uncensored-heretic-NVFP4A16/chat_template.jinja
|
||||
ERP_PORT=8021
|
||||
ERP_GPU_ID=1
|
||||
# 0.35 x 97.9 GiB = 34 GiB. GPU1 had ~47 GiB free on 2026-09-08 (scriberr/embed/rerank/coder/reward resident).
|
||||
ERP_GPU_MEM_UTIL=0.35
|
||||
ERP_MAX_MODEL_LEN=32768
|
||||
ERP_MAX_NUM_SEQS=8
|
||||
ERP_GPU_MEM_UTIL=0.30
|
||||
ERP_MAX_MODEL_LEN=262144
|
||||
ERP_MAX_NUM_SEQS=32
|
||||
API_KEY=
|
||||
# 8.49 GiB -> 534,649 KV tokens -> 2.04x a 262,144 context (operator's KV = 2x rule).
|
||||
ERP_KV_CACHE_MEMORY=9114000000
|
||||
ERP_MOE_BACKEND=auto
|
||||
|
||||
@@ -1,17 +1,38 @@
|
||||
# erp-seat — the ERP-tune seat on ana-ml2 GPU1: NVFP4A16 quant of **Pfish-6**, the run-6 LoRA
|
||||
# merge on the jenerallee78 ARA abliteration, served under that name.
|
||||
# erp-seat — the RP seat on ana-ml2 GPU1. Serves the **MeroMero A4B MoE** NVFP4A16 quant
|
||||
# (G4-MeroMero-26B-A4B-it-uncensored-heretic) behind the gateway alias `char-rp-fast`.
|
||||
#
|
||||
# ⚠ RUN 7 IS RETIRED (operator ruling 2026-09-09): "we're gonna stay on 6 for now". Run 7's
|
||||
# gate failure turned out to be a DETECTOR BUG (the adjective "minor" in a HARD rule, fixed
|
||||
# cc42d76 in brokkr-smithy) — but run 7 was independently a poor run (primary FLAT +2, both
|
||||
# diversity families reduced, long-context coherence 1.0 -> 0.875). Run 6 is the standing seat.
|
||||
# Routing aliases (e.g. LiteLLM `trial`) are the operator's call and live in the gateway, not here.
|
||||
# ⚠ THE STACK NAME IS HISTORICAL. It served Pfish-6 (the run-6 ERP-tune LoRA merge) until
|
||||
# 2026-09-10, when the operator swapped the occupant: "replace that a4b moe over pfish-6 --
|
||||
# remove the pfish-6 alias and create an alias for char-rp-fast." The compose PROJECT name is
|
||||
# deliberately NOT renamed: asset-engine derives seat liveness from it, so a rename reads as
|
||||
# OFFLINE. Pfish-6 remains on disk at /tank/aimodels/erp-tune-v6-nvfp4a16 and the pre-swap host
|
||||
# env is at /opt/docker/compose/erp-seat/.env.pfish6.bak-20260910 -- one cp plus `up -d` back.
|
||||
#
|
||||
# ⚠ THE FIRST A4B QUANT SERVED NaN AND LOOKED HEALTHY DOING IT. It was built with the DENSE
|
||||
# recipe, whose ignore list carries no router regex, so all 30 MoE routers were quantized to
|
||||
# 4 bits and expert selection was destroyed (playbook §3.15). The seat passed its healthcheck,
|
||||
# returned finish_reason=length with the full token count, and every response decoded to the
|
||||
# empty string; the give-away was NaN logprobs. Re-quantized with
|
||||
# services/erp-seat-quant/quant_nvfp4a16_gemma4_moe.py, whose target guard refuses exactly
|
||||
# that. Use the MoE recipe for anything in this family; the dense one is for the v2-31B.
|
||||
#
|
||||
# PFISH-6 PROVENANCE, kept because it is still the rollback target. Run 7 was retired by
|
||||
# operator ruling 2026-09-09 ("we're gonna stay on 6 for now"): its gate failure turned out to
|
||||
# be a DETECTOR BUG (the adjective "minor" in a HARD rule, fixed cc42d76 in brokkr-smithy), but
|
||||
# run 7 was independently a poor run (primary FLAT +2, both diversity families reduced,
|
||||
# long-context coherence 1.0 -> 0.875). Run 6 was the standing seat here until the 2026-09-10
|
||||
# swap above. Routing aliases live in the gateway, not here.
|
||||
#
|
||||
# Serve recipe copied from stacks/gemma4-charrp (same architecture + quant format, proven on this
|
||||
# box): gemma4 tool + reasoning parsers, enable_thinking pinned false, model's own stock template.
|
||||
# GPU1 is SHARED (charrp-MoE moved? no — scriberr, embed, rerank, coder, reward live there):
|
||||
# ~47 GiB was free on 2026-09-08; 0.35 x 97.9 GiB = 34 GiB keeps ~13 GiB of real margin.
|
||||
# Quant pipeline: services/erp-seat-quant/. Tunables in .env.
|
||||
# It carries over to MeroMero A4B unchanged -- verified 2026-09-10 end to end: clean prose with
|
||||
# no channel-prefix leak, a solid-colour image read correctly (vision towers intact), and an
|
||||
# auto tool_choice call parsed.
|
||||
# GPU1 is SHARED (scriberr, embed, rerank, coder, reward, charrp live there): the live .env runs
|
||||
# ERP_GPU_MEM_UTIL=0.30, and the KV pool is pinned in bytes below regardless, so the ratio only
|
||||
# has to clear admission.
|
||||
# Quant pipeline: services/erp-seat-quant/ (MoE recipe -- NOT services/meromero-quant/, which is
|
||||
# the dense one). Tunables in .env.
|
||||
|
||||
name: erp-seat
|
||||
|
||||
@@ -28,11 +49,11 @@ services:
|
||||
environment:
|
||||
- VLLM_API_KEY=${API_KEY:-}
|
||||
command:
|
||||
- ${ERP_MODEL:-/tank/aimodels/erp-tune-v6-nvfp4a16}
|
||||
- ${ERP_MODEL:-/tank/aimodels/G4-MeroMero-26B-A4B-it-uncensored-heretic-NVFP4A16}
|
||||
- --quantization
|
||||
- compressed-tensors
|
||||
- --served-model-name
|
||||
- ${ERP_SERVED_NAME:-Pfish-6}
|
||||
- ${ERP_SERVED_NAME:-G4-MeroMero-26B-A4B-it-uncensored-heretic-NVFP4A16}
|
||||
- --tool-call-parser
|
||||
- gemma4
|
||||
- --enable-auto-tool-choice
|
||||
@@ -52,7 +73,7 @@ services:
|
||||
# an empty turn. The flag drops the tools from the prompt so the model answers in prose.
|
||||
- --exclude-tools-when-tool-choice-none
|
||||
- --chat-template
|
||||
- ${ERP_CHAT_TEMPLATE:-/tank/aimodels/erp-tune-v6-nvfp4a16/chat_template.jinja}
|
||||
- ${ERP_CHAT_TEMPLATE:-/tank/aimodels/G4-MeroMero-26B-A4B-it-uncensored-heretic-NVFP4A16/chat_template.jinja}
|
||||
- --max-model-len
|
||||
- "${ERP_MAX_MODEL_LEN:-262144}"
|
||||
# KV pool pinned in BYTES, not inferred from the utilization ratio. GPU1 is
|
||||
@@ -71,6 +92,12 @@ services:
|
||||
#
|
||||
# 8.49 GiB -> 534,649 tokens -> 2.04x a full 262,144-token context, which is
|
||||
# the operator's sizing rule (KV = 2x max context, 2026-09-09).
|
||||
#
|
||||
# This figure SURVIVED the 2026-09-10 Pfish-6 -> MeroMero-A4B swap unchanged, and that
|
||||
# is not luck: the two are the same architecture field for field (30 layers, kv 8,
|
||||
# head_dim 256, sliding_window 1024, 25 sliding / 5 full, 128 experts top-8), so the
|
||||
# KV-per-token is the same number. Confirmed by reading 534,649 tokens / 2.04x back out
|
||||
# of the new engine's log rather than assuming the pinning carried.
|
||||
- --kv-cache-memory
|
||||
- "${ERP_KV_CACHE_MEMORY:-9114000000}"
|
||||
- --max-num-seqs
|
||||
@@ -122,9 +149,9 @@ services:
|
||||
- tnet
|
||||
labels:
|
||||
- homepage.group=AI - Inference
|
||||
- homepage.name=Pfish-6 (Gemma-4 26B-A4B ARA, NVFP4A16)
|
||||
- homepage.name=char-rp-fast (MeroMero 26B-A4B, NVFP4A16 MoE)
|
||||
- homepage.icon=mdi-fire
|
||||
- homepage.description=Pfish-6 — the run-6 LoRA merge on the jenerallee78 abliteration, NVFP4A16 MoE (ana-ml2 GPU1)
|
||||
- homepage.description=MeroMero A4B abliterated RP seat, NVFP4A16 weight-only, vision intact (ana-ml2 GPU1)
|
||||
- homepage.href=http://10.250.50.54:${ERP_PORT:-8021}/docs
|
||||
|
||||
networks:
|
||||
|
||||
@@ -872,9 +872,10 @@ model_list:
|
||||
# base profile ~0% on 30/35 axes per brokkr) -- this seat is for the operator's ear; treat
|
||||
# it as unrated on every safety axis.
|
||||
#
|
||||
# ROLLBACK: /tank/aimodels/erp-tune-v6-nvfp4a16 is still on disk, and the previous host env
|
||||
# is at /tmp/erp-seat-env.v6.bak on ana-ml2 -- flip ERP_MODEL/ERP_SERVED_NAME/ERP_CHAT_TEMPLATE
|
||||
# in /opt/docker/compose/erp-seat/.env back to v6 and `docker compose up -d`.
|
||||
# ROLLBACK (updated 2026-09-10, the seat now serves MeroMero A4B): Pfish-6 itself is
|
||||
# the rollback target. /tank/aimodels/erp-tune-v6-nvfp4a16 is still on disk and the
|
||||
# pre-swap host env is at /opt/docker/compose/erp-seat/.env.pfish6.bak-20260910 --
|
||||
# `cp .env.pfish6.bak-20260910 .env && docker compose up -d` restores Pfish-6 in ~4 min.
|
||||
#
|
||||
# THIS GATEWAY IS THE SHARED-KEY SURFACE: `all-agents-local` reaches every model here,
|
||||
# in every session and project. Removing this alias does not remove the operator's
|
||||
@@ -882,21 +883,50 @@ model_list:
|
||||
#
|
||||
# Same-site: seat and gateway are both at Anaheim (local hop, no mesh crossing).
|
||||
|
||||
# Pfish-6 — the STANDING seat as of 2026-09-09. NVFP4A16 quant of the run-6
|
||||
# LoRA merge on the jenerallee78 ARA abliteration, on ana-ml2 GPU1 (:8021).
|
||||
# Operator ruling: "declare run 6 as Pfish-6 ... we're gonna stay on 6 for now."
|
||||
# char-rp-fast — the MeroMero A4B MoE, on ana-ml2 GPU1 (:8021). Operator, 2026-09-10:
|
||||
# "replace that a4b moe over pfish-6 -- remove the pfish-6 alias and create an alias
|
||||
# for char-rp-fast." REPLACES the `Pfish-6` alias, which is removed with this change.
|
||||
#
|
||||
# REPLACES the `trial` alias, which is retired with run 7. Run 7's CSAM gate
|
||||
# failure turned out to be a DETECTOR BUG (the adjective "minor" matching a
|
||||
# HARD rule — fixed cc42d76 in brokkr-smithy), but run 7 was independently a
|
||||
# poor run and is not coming back.
|
||||
# The seat itself is unchanged in every dimension that matters to a caller: the A4B is
|
||||
# 30 layers / kv 8 / sliding_window 1024 / 128 experts top-8 — field for field the same
|
||||
# geometry as Pfish-6 — so the 9.114 GB KV pinning transfers exactly and the seat still
|
||||
# reports 534,649 tokens and 2.04x concurrency at 262,144. That was verified from the
|
||||
# engine log, not assumed, because KV-per-token is normally NOT transferable.
|
||||
#
|
||||
# Served under its TRUE name. This is now a named seat, not a trial.
|
||||
- model_name: Pfish-6
|
||||
# ⚠ Pfish-6 IS GONE from this gateway and its seat no longer serves that name. A caller
|
||||
# still asking for `Pfish-6` gets a clean 404 rather than a silent substitution, which is
|
||||
# the intended behaviour. The artifact is still on disk (see the ROLLBACK note above).
|
||||
#
|
||||
# SAMPLERS, and they are the author's, not inherited: the model card states Temp 0.8-1.0
|
||||
# and MinP 0.05. Temperature/top_p/top_k already come from the tree's own
|
||||
# generation_config (1.0 / 0.95 / 64) which vLLM applies server-side, and 1.0 sits at the
|
||||
# top of the card's stated range, so the only value that needs stating here is min_p.
|
||||
# Deliberately NOT copying char-rp's temp 1.1 / min_p 0.10 — those were A/B-tuned against
|
||||
# a retired Mistral seat and are inherited, not canonical, as that block's own note says.
|
||||
#
|
||||
# enable_thinking:false is ALSO pinned process-level on the seat
|
||||
# (--default-chat-template-kwargs). Stated here as well so the intent is visible at the
|
||||
# routing layer: without it the gemma4 reasoning parser pre-initialises to REASONING and
|
||||
# plain prose comes back with a null `content`.
|
||||
#
|
||||
# ⚠ ITS FIRST QUANT WAS BROKEN AND SERVED NaN. The 2026-09-10 08:15 build was made with
|
||||
# the DENSE recipe, whose ignore list has no router regex, so all 30 MoE routers were
|
||||
# quantized to 4 bits — a 4-bit router picks different experts (playbook §3.15). The seat
|
||||
# came up healthy, answered every request with 120 tokens, and decoded to the empty string;
|
||||
# logprobs were NaN. Re-quantized with services/erp-seat-quant/quant_nvfp4a16_gemma4_moe.py,
|
||||
# whose target guard refuses that exact mistake. The broken tree is parked on ana-ml2 as
|
||||
# ...-NVFP4A16.BROKEN-routers-quantized-20260910. Do not serve it.
|
||||
- model_name: char-rp-fast
|
||||
litellm_params:
|
||||
model: hosted_vllm/Pfish-6
|
||||
model: hosted_vllm/G4-MeroMero-26B-A4B-it-uncensored-heretic-NVFP4A16
|
||||
api_base: http://10.250.50.54:8021/v1
|
||||
api_key: os.environ/VLLM_API_KEY
|
||||
extra_body:
|
||||
min_p: 0.05
|
||||
chat_template_kwargs:
|
||||
enable_thinking: false
|
||||
model_info:
|
||||
mode: chat
|
||||
|
||||
general_settings:
|
||||
master_key: os.environ/LITELLM_MASTER_KEY
|
||||
|
||||
Reference in New Issue
Block a user