feat(meromero): MeroMero-v2 dual-mode (prose + streaming CoT) live on one seat — no re-quant

The multi-turn Gemma-4 CoT problem is solved. One MeroMero-v2 seat, one weight set,
two aliases: char-rp (prose) + char-rp-reasoning (streaming chain-of-thought).

The winning stack, traced from vLLM source by the four-arm brokkr/dwarf panel:

  - vllm/vllm-openai:v0.26.0 — ships transformers 5.14.1 natively, below the
    head_dim guard, so Gemma-4-31B loads with no pin and no custom image. It also
    carries the #48217 streaming pre-arm fix.
  - A patched chat template whose enable_thinking:true branch force-opens a BARE
    <|channel> (not <|channel>thought\n -- full-open defeats _preprocess_feed's
    injection). --chat-template override, no re-quant.
  - Two served-names char-rp / char-rp-thinking; --reasoning-parser gemma4;
    default enable_thinking:false. LiteLLM char-rp -> prose, char-rp-reasoning ->
    the thinking served-name with enable_thinking:true.

Verified: streaming CoT split 6/6 direct on :8016 and 3/3 through the gateway;
char-rp prose clean on both transports with no trailing-token leak.

Two hard-won facts recorded in persistent-memory:
  - STREAMING ONLY. Non-streaming can't split -- extract_reasoning never receives
    prompt_token_ids so the pre-arm can't fire (a vLLM one-shot bug unchanged
    across v0.24-0.27). Fine here: Lobe/OWUI stream. Upstream PR #49797 fixes
    non-streaming too, landing ~v0.28.0 -- then it's a clean image bump.
  - KEY-NAME TRAP: vLLM streams reasoning in delta.reasoning; LiteLLM normalizes
    to delta.reasoning_content. I lost two false-negative test rounds to this.

Canonical: stacks/meromero-charrp/ (compose + patched_chat_template.jinja) and
stacks/litellm/conf/config.yaml. Rollback is the .env image line + dropping
--chat-template.
This commit is contained in:
vh
2026-08-21 13:03:53 -07:00
parent 76834777a4
commit 5e47a59b32
4 changed files with 392 additions and 4 deletions
+30 -4
View File
@@ -8,8 +8,7 @@
# Model-name → upstream mapping:
# phi4-mini → vLLM :8004 (generative chat)
# qwen3-embedding → vLLM :8001 (/v1/embeddings)
# reranker → vLLM :8013 (/rerank; bge-v2-m3. The old qwen3-reranker
# alias on :8002 was retired 2026-08-20.)
# qwen3-reranker → vLLM :8002 (/rerank)
# * (wildcard) → llama-swap :9292 (the swappable generative zoo)
#
# The wildcard fronts llama-swap so its whole model zoo logs through the
@@ -223,6 +222,34 @@ model_list:
extra_body:
min_p: 0.10
top_k: 0
# EXPLICIT since 2026-08-21: the meromero seat no longer forces
# enable_thinking:false at the process level (it now also serves the
# char-rp-thinking variant for char-rp-reasoning). This false keeps the
# gemma4 parser out of the reasoning state so prose lands in content.
chat_template_kwargs:
enable_thinking: false
model_info:
mode: chat
# char-rp-reasoning -> MeroMero-v2 WITH CoT, 2026-08-21. Same physical seat as
# char-rp (:8016) but a DISTINCT served-name (char-rp-thinking) so LiteLLM keys
# it as its own deployment (no shared-param mutation with char-rp), and
# enable_thinking:true so the gemma4 parser splits the <|channel>thought block
# into reasoning_content while content stays clean prose. MeroMero-v2 is
# GRPO-trained with thinking (its own card: "Stage 3 RP logic GRPO, think
# enabled"). Same creative RP samplers as char-rp, thinking on.
- model_name: char-rp-reasoning
litellm_params:
model: hosted_vllm/char-rp-thinking
api_base: http://10.250.50.54:8016/v1
api_key: os.environ/VLLM_API_KEY
temperature: 1.1
top_p: 0.95
extra_body:
min_p: 0.10
top_k: 0
chat_template_kwargs:
enable_thinking: true
model_info:
mode: chat
# char-rp-reasoning -> GGUF managed-REASONING seat (:8018, llama.cpp, char-rp-gguf stack).
@@ -357,8 +384,7 @@ model_list:
top_p: 0.9
model_info:
mode: chat
# reranker → generic capability name for rerank. THE reranker alias — the only
# one left as of 2026-08-20. Every consumer pins this name, never a model name.
# reranker → generic capability name for rerank (currently qwen3-reranker).
- model_name: reranker
litellm_params:
model: hosted_vllm/BAAI/bge-reranker-v2-m3