Files
esh-pfi-infrastructure/stacks/meromero-charrp/README.md
T
vh b8f0f4c568 fix(char-rp): enable Gemma-4 tool-calling on the MeroMero seat
The char-rp seat shipped with no tool-call parser at all, so every
tools-bearing request was rejected outright:

  400 "auto" tool choice requires --enable-auto-tool-choice and
      --tool-call-parser to be set

MeroMero-v2 is Gemma-4, which emits its own native
`<|tool_call>call:name{...}<tool_call|>` syntax rather than the
qwen3_coder XML the Qwen-family seats use. vLLM 0.24 ships a matching
`gemma4` parser whose TOOL_CALL_START/END, CHANNEL_START/END and escape
token constants line up with this tokenizer's etc/eoc/escape tokens
exactly.

Three flags, and they are a set:

- --tool-call-parser gemma4 + --enable-auto-tool-choice: the fix proper.
- --reasoning-parser gemma4: without it the post-tool-response turn
  leaks a literal `<|channel>thought\n<channel|>` prefix into content
  (upstream vllm #45834 — the chat template leaves the prompt inside an
  open channel block).
- --default-chat-template-kwargs '{"enable_thinking": false}': MANDATORY
  companion to the reasoning parser. The parser reads enable_thinking
  from chat_template_kwargs and defaults it to True
  (vllm/parser/gemma4.py:439); True makes is_reasoning_end() return
  False at a new turn, pre-initialising the engine to REASONING, which
  routes ALL plain RP prose into reasoning_content and returns a null
  content — breaking every char-rp consumer. This template already
  defaults enable_thinking to false (chat_template.jinja:350), so
  passing it explicitly renders a byte-identical prompt (verified across
  plain / tools / post-tool-response / system-prompt shapes). It changes
  generation not at all; it only corrects the parser state machine.

Verified green on the live seat after deploy: tool call streaming and
non-streaming, tool-result round-trip (leak gone), plain prose in
content with reasoning null, vision unchanged.
2026-08-15 00:51:29 -07:00

47 lines
2.8 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# meromero-charrp — MeroMero-v2 char-rp prose seat (ana-ml2 GPU0)
The **non-thinking, multimodal** RP prose seat. Serves the LiteLLM `char-rp` alias.
- **Model:** `G4-MeroMero-v2-31B-NVFP4A16` (Gemma-4-31B, home-quantized weight-only NVFP4A16).
- **Host/GPU:** ana-ml2, GPU0 (co-located with `gen` / vllm-aeon-gen).
- **Port:** :8016 → LiteLLM `char-rp`.
- **Context:** 256K (`--max-model-len 262144`). Gemma-4 uses sliding-window attention → KV-efficient, ~2× concurrency at full context.
- **Vision:** enabled (image + text). `preprocessor_config.json` was materialized from the model's own `processor_config.json` (`Gemma4ImageProcessor`); audio is config-declared but weightless.
- **Tool-calling:** enabled via the `gemma4` parser (**not** `qwen3_coder` — that's the Qwen-family XML the other seats use). Gemma-4 emits its own native `<|tool_call>call:name{...}<tool_call|>` syntax.
## Tool-calling — the three flags are a set, don't split them
```yaml
- --tool-call-parser gemma4 # native <|tool_call> syntax; without it ANY tools request 400s
- --enable-auto-tool-choice
- --reasoning-parser gemma4 # absorbs the <|channel>…<channel|> thought markers
- --default-chat-template-kwargs '{"enable_thinking": false}' # MANDATORY, see below
```
Why the last one is mandatory: the gemma4 parser reads `enable_thinking` out of
`chat_template_kwargs` and **defaults it to `True`** (`vllm/parser/gemma4.py:439`). With `True`,
`is_reasoning_end()` returns `False` at a new turn, which pre-initialises the parser engine to
`REASONING` — so **all plain RP prose lands in `reasoning_content` and `content` comes back
`null`**, breaking every `char-rp` consumer. This model's `chat_template.jinja:350` already
defaults `enable_thinking` to `false`, so passing it explicitly renders a **byte-identical
prompt** (verified across plain / tools / post-tool-response / system-prompt shapes) — it changes
nothing about generation, it only corrects the parser's state machine.
Without `--reasoning-parser gemma4`, the post-tool-response turn leaks a literal
`<|channel>thought\n<channel|>` prefix into `content` (upstream vllm #45834 — the chat template
leaves the prompt sitting inside an open channel block).
Verified green after the fix: tool call (streaming + non-streaming), tool-result round-trip,
plain prose in `content`, vision.
- **Tuning:** `.env``MEROMERO_GPU_MEM_UTIL=0.52` (leaves ~4.6 GB GPU0 headroom), `MEROMERO_MAX_MODEL_LEN=262144`, `MEROMERO_GPU_ID=0`.
Replaces the retired **char-rp-gguf** (Magidonia-24B GGUF / llama.cpp) seat. The quant pipeline
lives in `ana-ml2:/tank/aimodels/meromero-v2-nvfp4-work/`.
## Deploy
```bash
scripts/deploy-stack.sh ana-ml2 meromero-charrp # diffs vs live, prompts y/N
# on host: cp .env.example .env; docker compose up -d
```