feat(char-rp): swap the seat to the Gemma-4 26B-A4B MoE, NVFP4, same port

Straight-across replacement of the dense G4-MeroMero-v2-31B-NVFP4A16 seat with
google/gemma-4-26B-A4B-it on ana-ml2 GPU0. Port, served-model-names and every
gateway route are unchanged, so no consumer sees a difference in addressing:
`char-rp` -> hosted_vllm/char-rp and `char-rp-reasoning` ->
hosted_vllm/char-rp-thinking, both still :8016. The seat's requirements now
include chain-of-thought, which makes throughput more critical rather than less
— the user waits through the whole reasoning block before the first visible
token, and the MoE measures ~114 tok/s @32K against the dense 31B's ~40.7.

Both artifacts are on disk and they are NOT interchangeable. The BF16 weights
(/tank/aimodels/gemma4-26b-a4b-it-bf16, 49 GB) are the QLoRA tuning base, since
QLoRA does its own quantization. They CANNOT be served here: 48.10 GiB of
weights against ~49 GiB of free GPU0 leaves nothing for KV cache, and the
engine would die at allocation exactly the way the predecessor did this
afternoon. The serving copy is RedHatAI/gemma-4-26B-A4B-it-NVFP4 (16 GB),
chosen over the other -it quants because it is compressed-tensors
(nvfp4-pack-quantized) — the same loader path the outgoing seat used — from the
llm-compressor team at 357k downloads. The nvidia/ repo is the base rather than
-it, and the thinking channel lives in the instruction-tuned weights.

Smaller weights at the same 0.47 memory budget buy a much larger KV pool:
27.37 GiB and 1,724,110 tokens, against the predecessor's 371,023 at the same
budget. That is 6.5 full-length 262K sequences concurrent rather than 1.4.

The gemma4 tool-call parser, reasoning parser and the enable_thinking:false
default all carry over unchanged — they are architecture-level, not
checkpoint-level. The --chat-template override does NOT carry over: MeroMero
pointed at a jinja hand-patched against that checkpoint, and this model ships
its own. Verified that dropping it did not reintroduce the failure that flag
existed to prevent — non-thinking prose lands in content with reasoning_content
empty, and the thinking alias populates reasoning_content with content
carrying the answer.

⚠ Scheme differs from the incumbent and the bench should say so: this quant
declares 4-bit input activations (W4A4) where the outgoing seat was NVFP4A16.
Faster, and not like-for-like on the activation axis.

meromero-charrp is retained stopped in `created` state and relabelled to
AI - Dormant, per the house rollback pattern. Both stacks want :8016, so
rolling back means stopping the gemma4 seat first.
This commit is contained in:
vh
2026-08-24 12:02:59 -07:00
parent 850e0c3351
commit 27155c0f3b
3 changed files with 158 additions and 3 deletions
+31
View File
@@ -0,0 +1,31 @@
# gemma4-charrp — char-rp seat on ana-ml2 GPU0. Real .env lives on the host.
# ⚠ GPU0 IS SHARED WITH `vllm-gen`. gen runs at --gpu-memory-utilization 0.43
# but actually holds ~45.6 GiB of the 94.97 GiB card — that flag sizes the KV
# cache and does NOT cover CUDA context, graphs and non-torch overhead. The
# predecessor seat sat at 0.51, the pair summed to 0.94, and on 2026-08-24 it
# stopped fitting and crash-looped 13 times with
# `torch.OutOfMemoryError: ... 195.19 MiB is free`.
#
# 0.47 keeps ~4.8 GiB of real margin. This model's weights are only ~15.3 GiB
# (NVFP4) against the predecessor's ~19.5 GiB, so the same budget buys MORE KV
# cache than before, not less. Raising this means lowering gen's in the same
# change — and check the real numbers, not the flags:
# nvidia-smi --query-compute-apps=pid,used_memory --format=csv
GEMMA4_GPU_MEM_UTIL=0.47
# Native context. config.json declares max_position_embeddings 262144, same as
# the outgoing seat, so this is a straight-across swap on context too.
GEMMA4_MAX_MODEL_LEN=262144
GEMMA4_MAX_NUM_SEQS=32
# ⚠ THE NVFP4 QUANT, NOT THE BF16. /tank/aimodels/gemma4-26b-a4b-it-bf16 is the
# QLoRA tuning base and is 48.10 GiB of weights — it does not fit beside gen.
GEMMA4_MODEL=/tank/aimodels/gemma4-26b-a4b-it-nvfp4
GEMMA4_PORT=8016
GEMMA4_GPU_ID=0
GEMMA4_CONTAINER=vllm-gemma4-charrp
# Same value as every other vLLM seat on this host — the gateway presents it.
API_KEY=