feat(meromero): MeroMero-v2 dual-mode (prose + streaming CoT) live on one seat — no re-quant
The multi-turn Gemma-4 CoT problem is solved. One MeroMero-v2 seat, one weight set,
two aliases: char-rp (prose) + char-rp-reasoning (streaming chain-of-thought).
The winning stack, traced from vLLM source by the four-arm brokkr/dwarf panel:
- vllm/vllm-openai:v0.26.0 — ships transformers 5.14.1 natively, below the
head_dim guard, so Gemma-4-31B loads with no pin and no custom image. It also
carries the #48217 streaming pre-arm fix.
- A patched chat template whose enable_thinking:true branch force-opens a BARE
<|channel> (not <|channel>thought\n -- full-open defeats _preprocess_feed's
injection). --chat-template override, no re-quant.
- Two served-names char-rp / char-rp-thinking; --reasoning-parser gemma4;
default enable_thinking:false. LiteLLM char-rp -> prose, char-rp-reasoning ->
the thinking served-name with enable_thinking:true.
Verified: streaming CoT split 6/6 direct on :8016 and 3/3 through the gateway;
char-rp prose clean on both transports with no trailing-token leak.
Two hard-won facts recorded in persistent-memory:
- STREAMING ONLY. Non-streaming can't split -- extract_reasoning never receives
prompt_token_ids so the pre-arm can't fire (a vLLM one-shot bug unchanged
across v0.24-0.27). Fine here: Lobe/OWUI stream. Upstream PR #49797 fixes
non-streaming too, landing ~v0.28.0 -- then it's a clean image bump.
- KEY-NAME TRAP: vLLM streams reasoning in delta.reasoning; LiteLLM normalizes
to delta.reasoning_content. I lost two false-negative test rounds to this.
Canonical: stacks/meromero-charrp/ (compose + patched_chat_template.jinja) and
stacks/litellm/conf/config.yaml. Rollback is the .env image line + dropping
--chat-template.
This commit is contained in:
@@ -8,8 +8,7 @@
|
||||
# Model-name → upstream mapping:
|
||||
# phi4-mini → vLLM :8004 (generative chat)
|
||||
# qwen3-embedding → vLLM :8001 (/v1/embeddings)
|
||||
# reranker → vLLM :8013 (/rerank; bge-v2-m3. The old qwen3-reranker
|
||||
# alias on :8002 was retired 2026-08-20.)
|
||||
# qwen3-reranker → vLLM :8002 (/rerank)
|
||||
# * (wildcard) → llama-swap :9292 (the swappable generative zoo)
|
||||
#
|
||||
# The wildcard fronts llama-swap so its whole model zoo logs through the
|
||||
@@ -223,6 +222,34 @@ model_list:
|
||||
extra_body:
|
||||
min_p: 0.10
|
||||
top_k: 0
|
||||
# EXPLICIT since 2026-08-21: the meromero seat no longer forces
|
||||
# enable_thinking:false at the process level (it now also serves the
|
||||
# char-rp-thinking variant for char-rp-reasoning). This false keeps the
|
||||
# gemma4 parser out of the reasoning state so prose lands in content.
|
||||
chat_template_kwargs:
|
||||
enable_thinking: false
|
||||
model_info:
|
||||
mode: chat
|
||||
|
||||
# char-rp-reasoning -> MeroMero-v2 WITH CoT, 2026-08-21. Same physical seat as
|
||||
# char-rp (:8016) but a DISTINCT served-name (char-rp-thinking) so LiteLLM keys
|
||||
# it as its own deployment (no shared-param mutation with char-rp), and
|
||||
# enable_thinking:true so the gemma4 parser splits the <|channel>thought block
|
||||
# into reasoning_content while content stays clean prose. MeroMero-v2 is
|
||||
# GRPO-trained with thinking (its own card: "Stage 3 RP logic GRPO, think
|
||||
# enabled"). Same creative RP samplers as char-rp, thinking on.
|
||||
- model_name: char-rp-reasoning
|
||||
litellm_params:
|
||||
model: hosted_vllm/char-rp-thinking
|
||||
api_base: http://10.250.50.54:8016/v1
|
||||
api_key: os.environ/VLLM_API_KEY
|
||||
temperature: 1.1
|
||||
top_p: 0.95
|
||||
extra_body:
|
||||
min_p: 0.10
|
||||
top_k: 0
|
||||
chat_template_kwargs:
|
||||
enable_thinking: true
|
||||
model_info:
|
||||
mode: chat
|
||||
# char-rp-reasoning -> GGUF managed-REASONING seat (:8018, llama.cpp, char-rp-gguf stack).
|
||||
@@ -357,8 +384,7 @@ model_list:
|
||||
top_p: 0.9
|
||||
model_info:
|
||||
mode: chat
|
||||
# reranker → generic capability name for rerank. THE reranker alias — the only
|
||||
# one left as of 2026-08-20. Every consumer pins this name, never a model name.
|
||||
# reranker → generic capability name for rerank (currently qwen3-reranker).
|
||||
- model_name: reranker
|
||||
litellm_params:
|
||||
model: hosted_vllm/BAAI/bge-reranker-v2-m3
|
||||
|
||||
Reference in New Issue
Block a user