Files
esh-pfi-infrastructure/stacks/meromero-charrp
vh 27155c0f3b feat(char-rp): swap the seat to the Gemma-4 26B-A4B MoE, NVFP4, same port
Straight-across replacement of the dense G4-MeroMero-v2-31B-NVFP4A16 seat with
google/gemma-4-26B-A4B-it on ana-ml2 GPU0. Port, served-model-names and every
gateway route are unchanged, so no consumer sees a difference in addressing:
`char-rp` -> hosted_vllm/char-rp and `char-rp-reasoning` ->
hosted_vllm/char-rp-thinking, both still :8016. The seat's requirements now
include chain-of-thought, which makes throughput more critical rather than less
— the user waits through the whole reasoning block before the first visible
token, and the MoE measures ~114 tok/s @32K against the dense 31B's ~40.7.

Both artifacts are on disk and they are NOT interchangeable. The BF16 weights
(/tank/aimodels/gemma4-26b-a4b-it-bf16, 49 GB) are the QLoRA tuning base, since
QLoRA does its own quantization. They CANNOT be served here: 48.10 GiB of
weights against ~49 GiB of free GPU0 leaves nothing for KV cache, and the
engine would die at allocation exactly the way the predecessor did this
afternoon. The serving copy is RedHatAI/gemma-4-26B-A4B-it-NVFP4 (16 GB),
chosen over the other -it quants because it is compressed-tensors
(nvfp4-pack-quantized) — the same loader path the outgoing seat used — from the
llm-compressor team at 357k downloads. The nvidia/ repo is the base rather than
-it, and the thinking channel lives in the instruction-tuned weights.

Smaller weights at the same 0.47 memory budget buy a much larger KV pool:
27.37 GiB and 1,724,110 tokens, against the predecessor's 371,023 at the same
budget. That is 6.5 full-length 262K sequences concurrent rather than 1.4.

The gemma4 tool-call parser, reasoning parser and the enable_thinking:false
default all carry over unchanged — they are architecture-level, not
checkpoint-level. The --chat-template override does NOT carry over: MeroMero
pointed at a jinja hand-patched against that checkpoint, and this model ships
its own. Verified that dropping it did not reintroduce the failure that flag
existed to prevent — non-thinking prose lands in content with reasoning_content
empty, and the thinking alias populates reasoning_content with content
carrying the answer.

⚠ Scheme differs from the incumbent and the bench should say so: this quant
declares 4-bit input activations (W4A4) where the outgoing seat was NVFP4A16.
Faster, and not like-for-like on the activation axis.

meromero-charrp is retained stopped in `created` state and relabelled to
AI - Dormant, per the house rollback pattern. Both stacks want :8016, so
rolling back means stopping the gemma4 seat first.
2026-08-24 12:02:59 -07:00
..

meromero-charrp — MeroMero-v2 char-rp prose seat (ana-ml2 GPU0)

The non-thinking, multimodal RP prose seat. Serves the LiteLLM char-rp alias.

  • Model: G4-MeroMero-v2-31B-NVFP4A16 (Gemma-4-31B, home-quantized weight-only NVFP4A16).
  • Host/GPU: ana-ml2, GPU0 (co-located with gen / vllm-aeon-gen).
  • Port: :8016 → LiteLLM char-rp.
  • Context: 256K (--max-model-len 262144). Gemma-4 uses sliding-window attention → KV-efficient, ~2× concurrency at full context.
  • Vision: enabled (image + text). preprocessor_config.json was materialized from the model's own processor_config.json (Gemma4ImageProcessor); audio is config-declared but weightless.
  • Tool-calling: enabled via the gemma4 parser (not qwen3_coder — that's the Qwen-family XML the other seats use). Gemma-4 emits its own native <|tool_call>call:name{...}<tool_call|> syntax.

Tool-calling — the three flags are a set, don't split them

- --tool-call-parser gemma4        # native <|tool_call> syntax; without it ANY tools request 400s
- --enable-auto-tool-choice
- --reasoning-parser gemma4        # absorbs the <|channel>…<channel|> thought markers
- --default-chat-template-kwargs '{"enable_thinking": false}'   # MANDATORY, see below

Why the last one is mandatory: the gemma4 parser reads enable_thinking out of chat_template_kwargs and defaults it to True (vllm/parser/gemma4.py:439). With True, is_reasoning_end() returns False at a new turn, which pre-initialises the parser engine to REASONING — so all plain RP prose lands in reasoning_content and content comes back null, breaking every char-rp consumer. This model's chat_template.jinja:350 already defaults enable_thinking to false, so passing it explicitly renders a byte-identical prompt (verified across plain / tools / post-tool-response / system-prompt shapes) — it changes nothing about generation, it only corrects the parser's state machine.

Without --reasoning-parser gemma4, the post-tool-response turn leaks a literal <|channel>thought\n<channel|> prefix into content (upstream vllm #45834 — the chat template leaves the prompt sitting inside an open channel block).

Verified green after the fix: tool call (streaming + non-streaming), tool-result round-trip, plain prose in content, vision.

  • Tuning: .env — MEROMERO_GPU_MEM_UTIL=0.52 (leaves ~4.6 GB GPU0 headroom), MEROMERO_MAX_MODEL_LEN=262144, MEROMERO_GPU_ID=0.

Replaces the retired char-rp-gguf (Magidonia-24B GGUF / llama.cpp) seat. The quant pipeline lives in ana-ml2:/tank/aimodels/meromero-v2-nvfp4-work/.

Deploy

scripts/deploy-stack.sh ana-ml2 meromero-charrp     # diffs vs live, prompts y/N
# on host: cp .env.example .env; docker compose up -d