feat(meromero): MeroMero-v2 dual-mode (prose + streaming CoT) live on one seat — no re-quant
The multi-turn Gemma-4 CoT problem is solved. One MeroMero-v2 seat, one weight set,
two aliases: char-rp (prose) + char-rp-reasoning (streaming chain-of-thought).
The winning stack, traced from vLLM source by the four-arm brokkr/dwarf panel:
- vllm/vllm-openai:v0.26.0 — ships transformers 5.14.1 natively, below the
head_dim guard, so Gemma-4-31B loads with no pin and no custom image. It also
carries the #48217 streaming pre-arm fix.
- A patched chat template whose enable_thinking:true branch force-opens a BARE
<|channel> (not <|channel>thought\n -- full-open defeats _preprocess_feed's
injection). --chat-template override, no re-quant.
- Two served-names char-rp / char-rp-thinking; --reasoning-parser gemma4;
default enable_thinking:false. LiteLLM char-rp -> prose, char-rp-reasoning ->
the thinking served-name with enable_thinking:true.
Verified: streaming CoT split 6/6 direct on :8016 and 3/3 through the gateway;
char-rp prose clean on both transports with no trailing-token leak.
Two hard-won facts recorded in persistent-memory:
- STREAMING ONLY. Non-streaming can't split -- extract_reasoning never receives
prompt_token_ids so the pre-arm can't fire (a vLLM one-shot bug unchanged
across v0.24-0.27). Fine here: Lobe/OWUI stream. Upstream PR #49797 fixes
non-streaming too, landing ~v0.28.0 -- then it's a clean image bump.
- KEY-NAME TRAP: vLLM streams reasoning in delta.reasoning; LiteLLM normalizes
to delta.reasoning_content. I lost two false-negative test rounds to this.
Canonical: stacks/meromero-charrp/ (compose + patched_chat_template.jinja) and
stacks/litellm/conf/config.yaml. Rollback is the .env image line + dropping
--chat-template.
This commit is contained in:
@@ -30,6 +30,7 @@ services:
|
||||
- compressed-tensors
|
||||
- --served-model-name
|
||||
- char-rp
|
||||
- char-rp-thinking
|
||||
# Tool-calling: Gemma-4 emits its OWN native syntax
|
||||
# (<|tool_call>call:name{...}<tool_call|>), NOT the qwen3_coder XML the
|
||||
# other seats use. vLLM 0.24 ships a matching `gemma4` parser whose token
|
||||
@@ -44,6 +45,9 @@ services:
|
||||
# the prompt inside an open channel block).
|
||||
- --reasoning-parser
|
||||
- gemma4
|
||||
# Dwarf-panel CoT patch 2026-08-21: force-open <|channel> on enable_thinking:true
|
||||
- --chat-template
|
||||
- /tank/aimodels/meromero-v2-nvfp4-work/patched_chat_template.jinja
|
||||
# MANDATORY companion to the reasoning parser on this seat. The parser
|
||||
# reads enable_thinking from chat_template_kwargs and DEFAULTS IT TO TRUE
|
||||
# (vllm/parser/gemma4.py:439). True makes is_reasoning_end() return False
|
||||
|
||||
Reference in New Issue
Block a user