# meromero-charrp — MeroMero-v2 char-rp prose seat (ana-ml2 GPU0) The **non-thinking, multimodal** RP prose seat. Serves the LiteLLM `char-rp` alias. - **Model:** `G4-MeroMero-v2-31B-NVFP4A16` (Gemma-4-31B, home-quantized weight-only NVFP4A16). - **Host/GPU:** ana-ml2, GPU0 (co-located with `gen` / vllm-aeon-gen). - **Port:** :8016 → LiteLLM `char-rp`. - **Context:** 256K (`--max-model-len 262144`). Gemma-4 uses sliding-window attention → KV-efficient, ~2× concurrency at full context. - **Vision:** enabled (image + text). `preprocessor_config.json` was materialized from the model's own `processor_config.json` (`Gemma4ImageProcessor`); audio is config-declared but weightless. - **Tool-calling:** enabled via the `gemma4` parser (**not** `qwen3_coder` — that's the Qwen-family XML the other seats use). Gemma-4 emits its own native `<|tool_call>call:name{...}` syntax. ## Tool-calling — the three flags are a set, don't split them ```yaml - --tool-call-parser gemma4 # native <|tool_call> syntax; without it ANY tools request 400s - --enable-auto-tool-choice - --reasoning-parser gemma4 # absorbs the <|channel>… thought markers - --default-chat-template-kwargs '{"enable_thinking": false}' # MANDATORY, see below ``` Why the last one is mandatory: the gemma4 parser reads `enable_thinking` out of `chat_template_kwargs` and **defaults it to `True`** (`vllm/parser/gemma4.py:439`). With `True`, `is_reasoning_end()` returns `False` at a new turn, which pre-initialises the parser engine to `REASONING` — so **all plain RP prose lands in `reasoning_content` and `content` comes back `null`**, breaking every `char-rp` consumer. This model's `chat_template.jinja:350` already defaults `enable_thinking` to `false`, so passing it explicitly renders a **byte-identical prompt** (verified across plain / tools / post-tool-response / system-prompt shapes) — it changes nothing about generation, it only corrects the parser's state machine. Without `--reasoning-parser gemma4`, the post-tool-response turn leaks a literal `<|channel>thought\n` prefix into `content` (upstream vllm #45834 — the chat template leaves the prompt sitting inside an open channel block). Verified green after the fix: tool call (streaming + non-streaming), tool-result round-trip, plain prose in `content`, vision. - **Tuning:** `.env` — `MEROMERO_GPU_MEM_UTIL=0.52` (leaves ~4.6 GB GPU0 headroom), `MEROMERO_MAX_MODEL_LEN=262144`, `MEROMERO_GPU_ID=0`. Replaces the retired **char-rp-gguf** (Magidonia-24B GGUF / llama.cpp) seat. The quant pipeline lives in `ana-ml2:/tank/aimodels/meromero-v2-nvfp4-work/`. ## Deploy ```bash scripts/deploy-stack.sh ana-ml2 meromero-charrp # diffs vs live, prompts y/N # on host: cp .env.example .env; docker compose up -d ```