feat(meromero): MeroMero-v2 dual-mode (prose + streaming CoT) live on one seat — no re-quant

The multi-turn Gemma-4 CoT problem is solved. One MeroMero-v2 seat, one weight set,
two aliases: char-rp (prose) + char-rp-reasoning (streaming chain-of-thought).

The winning stack, traced from vLLM source by the four-arm brokkr/dwarf panel:

  - vllm/vllm-openai:v0.26.0 — ships transformers 5.14.1 natively, below the
    head_dim guard, so Gemma-4-31B loads with no pin and no custom image. It also
    carries the #48217 streaming pre-arm fix.
  - A patched chat template whose enable_thinking:true branch force-opens a BARE
    <|channel> (not <|channel>thought\n -- full-open defeats _preprocess_feed's
    injection). --chat-template override, no re-quant.
  - Two served-names char-rp / char-rp-thinking; --reasoning-parser gemma4;
    default enable_thinking:false. LiteLLM char-rp -> prose, char-rp-reasoning ->
    the thinking served-name with enable_thinking:true.

Verified: streaming CoT split 6/6 direct on :8016 and 3/3 through the gateway;
char-rp prose clean on both transports with no trailing-token leak.

Two hard-won facts recorded in persistent-memory:
  - STREAMING ONLY. Non-streaming can't split -- extract_reasoning never receives
    prompt_token_ids so the pre-arm can't fire (a vLLM one-shot bug unchanged
    across v0.24-0.27). Fine here: Lobe/OWUI stream. Upstream PR #49797 fixes
    non-streaming too, landing ~v0.28.0 -- then it's a clean image bump.
  - KEY-NAME TRAP: vLLM streams reasoning in delta.reasoning; LiteLLM normalizes
    to delta.reasoning_content. I lost two false-negative test rounds to this.

Canonical: stacks/meromero-charrp/ (compose + patched_chat_template.jinja) and
stacks/litellm/conf/config.yaml. Rollback is the .env image line + dropping
--chat-template.
This commit is contained in:
vh
2026-08-21 13:03:53 -07:00
parent 76834777a4
commit 5e47a59b32
4 changed files with 392 additions and 4 deletions
+4
View File
@@ -30,6 +30,7 @@ services:
- compressed-tensors
- --served-model-name
- char-rp
- char-rp-thinking
# Tool-calling: Gemma-4 emits its OWN native syntax
# (<|tool_call>call:name{...}<tool_call|>), NOT the qwen3_coder XML the
# other seats use. vLLM 0.24 ships a matching `gemma4` parser whose token
@@ -44,6 +45,9 @@ services:
# the prompt inside an open channel block).
- --reasoning-parser
- gemma4
# Dwarf-panel CoT patch 2026-08-21: force-open <|channel> on enable_thinking:true
- --chat-template
- /tank/aimodels/meromero-v2-nvfp4-work/patched_chat_template.jinja
# MANDATORY companion to the reasoning parser on this seat. The parser
# reads enable_thinking from chat_template_kwargs and DEFAULTS IT TO TRUE
# (vllm/parser/gemma4.py:439). True makes is_reasoning_end() return False