Adds an ephemeral stack for serving the BF16 trainee base on :8016 under the char-rp aliases, so the abliterated base can be measured on the same battery and the same gateway routes as the served seat with no harness edit. It is a separate stack rather than another variable on gemma4-charrp because that compose hardcodes `--quantization compressed-tensors` for the NVFP4 build. Pointing it at unquantized BF16 weights crash-loops immediately — `TypeError: CompressedTensorsConfig.__init__() missing 3 required positional arguments: 'target_scheme_map', 'ignore', 'quant_format'` — vLLM trying to read a quantization config out of a checkpoint that has none. 35 restarts before it was caught. `restart: "no"` here so a bench seat cannot resurrect itself and block gen's restore, and no homepage labels so it leaves no permanently-offline dashboard card. It cannot coexist with gen and says so: 48.07 GiB of BF16 weights plus gen's footprint exceeds the 94.97 GiB card before any KV cache. Running it means gen is stopped. THE MORE USEFUL FINDING is in the meromero env note: gen's memory footprint GROWS WITH UPTIME. Measured today at 46,726 MiB (45.6 GiB) after ~3 days up, and 39,424 MiB (38.5 GiB) immediately after a restart — same container, same --gpu-memory-utilization 0.43, ~7 GiB apart. That is the missing half of this afternoon's crash-loop: the char-rp seat "fit on the 21st and stopped fitting on the 24th" because nothing about char-rp changed and gen crept up underneath it. Headroom arithmetic done against a long-running gen is measuring a moving number, so the note now says to measure against a freshly-restarted one. Operator's requested end state reached and verified through the gateway: gen and summarizer both 200, char-rp down deliberately to hold GPU0 headroom for the upcoming trainee run, bench seat stopped.
meromero-charrp — MeroMero-v2 char-rp prose seat (ana-ml2 GPU0)
The non-thinking, multimodal RP prose seat. Serves the LiteLLM char-rp alias.
- Model:
G4-MeroMero-v2-31B-NVFP4A16(Gemma-4-31B, home-quantized weight-only NVFP4A16). - Host/GPU: ana-ml2, GPU0 (co-located with
gen/ vllm-aeon-gen). - Port: :8016 → LiteLLM
char-rp. - Context: 256K (
--max-model-len 262144). Gemma-4 uses sliding-window attention → KV-efficient, ~2× concurrency at full context. - Vision: enabled (image + text).
preprocessor_config.jsonwas materialized from the model's ownprocessor_config.json(Gemma4ImageProcessor); audio is config-declared but weightless. - Tool-calling: enabled via the
gemma4parser (notqwen3_coder— that's the Qwen-family XML the other seats use). Gemma-4 emits its own native<|tool_call>call:name{...}<tool_call|>syntax.
Tool-calling — the three flags are a set, don't split them
- --tool-call-parser gemma4 # native <|tool_call> syntax; without it ANY tools request 400s
- --enable-auto-tool-choice
- --reasoning-parser gemma4 # absorbs the <|channel>…<channel|> thought markers
- --default-chat-template-kwargs '{"enable_thinking": false}' # MANDATORY, see below
Why the last one is mandatory: the gemma4 parser reads enable_thinking out of
chat_template_kwargs and defaults it to True (vllm/parser/gemma4.py:439). With True,
is_reasoning_end() returns False at a new turn, which pre-initialises the parser engine to
REASONING — so all plain RP prose lands in reasoning_content and content comes back
null, breaking every char-rp consumer. This model's chat_template.jinja:350 already
defaults enable_thinking to false, so passing it explicitly renders a byte-identical
prompt (verified across plain / tools / post-tool-response / system-prompt shapes) — it changes
nothing about generation, it only corrects the parser's state machine.
Without --reasoning-parser gemma4, the post-tool-response turn leaks a literal
<|channel>thought\n<channel|> prefix into content (upstream vllm #45834 — the chat template
leaves the prompt sitting inside an open channel block).
Verified green after the fix: tool call (streaming + non-streaming), tool-result round-trip,
plain prose in content, vision.
- Tuning:
.env—MEROMERO_GPU_MEM_UTIL=0.52(leaves ~4.6 GB GPU0 headroom),MEROMERO_MAX_MODEL_LEN=262144,MEROMERO_GPU_ID=0.
Replaces the retired char-rp-gguf (Magidonia-24B GGUF / llama.cpp) seat. The quant pipeline
lives in ana-ml2:/tank/aimodels/meromero-v2-nvfp4-work/.
Deploy
scripts/deploy-stack.sh ana-ml2 meromero-charrp # diffs vs live, prompts y/N
# on host: cp .env.example .env; docker compose up -d