Straight-across replacement of the dense G4-MeroMero-v2-31B-NVFP4A16 seat with google/gemma-4-26B-A4B-it on ana-ml2 GPU0. Port, served-model-names and every gateway route are unchanged, so no consumer sees a difference in addressing: `char-rp` -> hosted_vllm/char-rp and `char-rp-reasoning` -> hosted_vllm/char-rp-thinking, both still :8016. The seat's requirements now include chain-of-thought, which makes throughput more critical rather than less — the user waits through the whole reasoning block before the first visible token, and the MoE measures ~114 tok/s @32K against the dense 31B's ~40.7. Both artifacts are on disk and they are NOT interchangeable. The BF16 weights (/tank/aimodels/gemma4-26b-a4b-it-bf16, 49 GB) are the QLoRA tuning base, since QLoRA does its own quantization. They CANNOT be served here: 48.10 GiB of weights against ~49 GiB of free GPU0 leaves nothing for KV cache, and the engine would die at allocation exactly the way the predecessor did this afternoon. The serving copy is RedHatAI/gemma-4-26B-A4B-it-NVFP4 (16 GB), chosen over the other -it quants because it is compressed-tensors (nvfp4-pack-quantized) — the same loader path the outgoing seat used — from the llm-compressor team at 357k downloads. The nvidia/ repo is the base rather than -it, and the thinking channel lives in the instruction-tuned weights. Smaller weights at the same 0.47 memory budget buy a much larger KV pool: 27.37 GiB and 1,724,110 tokens, against the predecessor's 371,023 at the same budget. That is 6.5 full-length 262K sequences concurrent rather than 1.4. The gemma4 tool-call parser, reasoning parser and the enable_thinking:false default all carry over unchanged — they are architecture-level, not checkpoint-level. The --chat-template override does NOT carry over: MeroMero pointed at a jinja hand-patched against that checkpoint, and this model ships its own. Verified that dropping it did not reintroduce the failure that flag existed to prevent — non-thinking prose lands in content with reasoning_content empty, and the thinking alias populates reasoning_content with content carrying the answer. ⚠ Scheme differs from the incumbent and the bench should say so: this quant declares 4-bit input activations (W4A4) where the outgoing seat was NVFP4A16. Faster, and not like-for-like on the activation axis. meromero-charrp is retained stopped in `created` state and relabelled to AI - Dormant, per the house rollback pattern. Both stacks want :8016, so rolling back means stopping the gemma4 seat first.
meromero-charrp — MeroMero-v2 char-rp prose seat (ana-ml2 GPU0)
The non-thinking, multimodal RP prose seat. Serves the LiteLLM char-rp alias.
- Model:
G4-MeroMero-v2-31B-NVFP4A16(Gemma-4-31B, home-quantized weight-only NVFP4A16). - Host/GPU: ana-ml2, GPU0 (co-located with
gen/ vllm-aeon-gen). - Port: :8016 → LiteLLM
char-rp. - Context: 256K (
--max-model-len 262144). Gemma-4 uses sliding-window attention → KV-efficient, ~2× concurrency at full context. - Vision: enabled (image + text).
preprocessor_config.jsonwas materialized from the model's ownprocessor_config.json(Gemma4ImageProcessor); audio is config-declared but weightless. - Tool-calling: enabled via the
gemma4parser (notqwen3_coder— that's the Qwen-family XML the other seats use). Gemma-4 emits its own native<|tool_call>call:name{...}<tool_call|>syntax.
Tool-calling — the three flags are a set, don't split them
- --tool-call-parser gemma4 # native <|tool_call> syntax; without it ANY tools request 400s
- --enable-auto-tool-choice
- --reasoning-parser gemma4 # absorbs the <|channel>…<channel|> thought markers
- --default-chat-template-kwargs '{"enable_thinking": false}' # MANDATORY, see below
Why the last one is mandatory: the gemma4 parser reads enable_thinking out of
chat_template_kwargs and defaults it to True (vllm/parser/gemma4.py:439). With True,
is_reasoning_end() returns False at a new turn, which pre-initialises the parser engine to
REASONING — so all plain RP prose lands in reasoning_content and content comes back
null, breaking every char-rp consumer. This model's chat_template.jinja:350 already
defaults enable_thinking to false, so passing it explicitly renders a byte-identical
prompt (verified across plain / tools / post-tool-response / system-prompt shapes) — it changes
nothing about generation, it only corrects the parser's state machine.
Without --reasoning-parser gemma4, the post-tool-response turn leaks a literal
<|channel>thought\n<channel|> prefix into content (upstream vllm #45834 — the chat template
leaves the prompt sitting inside an open channel block).
Verified green after the fix: tool call (streaming + non-streaming), tool-result round-trip,
plain prose in content, vision.
- Tuning:
.env—MEROMERO_GPU_MEM_UTIL=0.52(leaves ~4.6 GB GPU0 headroom),MEROMERO_MAX_MODEL_LEN=262144,MEROMERO_GPU_ID=0.
Replaces the retired char-rp-gguf (Magidonia-24B GGUF / llama.cpp) seat. The quant pipeline
lives in ana-ml2:/tank/aimodels/meromero-v2-nvfp4-work/.
Deploy
scripts/deploy-stack.sh ana-ml2 meromero-charrp # diffs vs live, prompts y/N
# on host: cp .env.example .env; docker compose up -d