feat(meromero): MeroMero-v2 dual-mode (prose + streaming CoT) live on one seat — no re-quant

The multi-turn Gemma-4 CoT problem is solved. One MeroMero-v2 seat, one weight set,
two aliases: char-rp (prose) + char-rp-reasoning (streaming chain-of-thought).

The winning stack, traced from vLLM source by the four-arm brokkr/dwarf panel:

  - vllm/vllm-openai:v0.26.0 — ships transformers 5.14.1 natively, below the
    head_dim guard, so Gemma-4-31B loads with no pin and no custom image. It also
    carries the #48217 streaming pre-arm fix.
  - A patched chat template whose enable_thinking:true branch force-opens a BARE
    <|channel> (not <|channel>thought\n -- full-open defeats _preprocess_feed's
    injection). --chat-template override, no re-quant.
  - Two served-names char-rp / char-rp-thinking; --reasoning-parser gemma4;
    default enable_thinking:false. LiteLLM char-rp -> prose, char-rp-reasoning ->
    the thinking served-name with enable_thinking:true.

Verified: streaming CoT split 6/6 direct on :8016 and 3/3 through the gateway;
char-rp prose clean on both transports with no trailing-token leak.

Two hard-won facts recorded in persistent-memory:
  - STREAMING ONLY. Non-streaming can't split -- extract_reasoning never receives
    prompt_token_ids so the pre-arm can't fire (a vLLM one-shot bug unchanged
    across v0.24-0.27). Fine here: Lobe/OWUI stream. Upstream PR #49797 fixes
    non-streaming too, landing ~v0.28.0 -- then it's a clean image bump.
  - KEY-NAME TRAP: vLLM streams reasoning in delta.reasoning; LiteLLM normalizes
    to delta.reasoning_content. I lost two false-negative test rounds to this.

Canonical: stacks/meromero-charrp/ (compose + patched_chat_template.jinja) and
stacks/litellm/conf/config.yaml. Rollback is the .env image line + dropping
--chat-template.
This commit is contained in:
vh
2026-08-21 13:03:53 -07:00
parent 76834777a4
commit 5e47a59b32
4 changed files with 392 additions and 4 deletions
+2
View File
@@ -124,6 +124,8 @@ _As of 2026-08-21 00:35 — **the Heretic-300 session, and its reversal** (see t
- **🟢 OPEN WEBUI — deployed as a Lobe bake-off, esh-docker-vm:3211 (2026-08-21).** Operator-approved candidate replacement for `lobe-chat` (:3210), stood up **parallel** — Lobe untouched. `stacks/open-webui/` (v0.11.0, `ENABLE_PERSISTENT_CONFIG=False` = deploy is the config source of truth). Gates (verified on the box): **G1** declarative-config PASS both directions (env change takes on bounce, UI change reverts on restart — no persistent-config bug bit it); **G2** picker auto-tracks the 31 live gateway models 1:1, no pins (also shows non-chat seats — the flip side of no-hand-listing); **G3** `POST /api/v1/models/sync` genuinely reconciles (create+delete), `export` round-trips; **G5** task model pinned `summarizer`; **G4** (TTS, direct at `:8198`) handed to tts-dev. Admin = **lkraven** (temp pw, signup then locked off). Fresh **capped** key `open-webui-esh` (`all-proxy-models` + **$50/1mo** cap — NOT inherited from uncapped `lobe-chat-esh`). Secrets vaulted `esh-docker-vm/open-webui-{litellm-key,secret-key,admin}`. Folded in a `docker image prune -af` → **73.6 GB reclaimed**. ⚠ **LESSON:** in Open WebUI a `.env` var only reaches the container if `compose.yaml` names it in `environment:` (Compose uses `.env` for `${VAR}` substitution, not as an `env_file`); and the API-key toggle env var is **`ENABLE_API_KEYS`** (plural) — singular is inert. Detail lives in `stacks/open-webui/README.md`. **Operator's open call:** whether Lobe retires once G4 passes.
- **✅✅ SOLVED 2026-08-21 — MeroMero-v2 DUAL-MODE (prose + streaming CoT) IS LIVE on ONE seat, ONE weight set, TWO aliases. No re-quant.** The multi-turn saga below is resolved. **Config:** `meromero-charrp` seat on **`vllm/vllm-openai:v0.26.0`** (ships transformers **5.14.1** natively — below the head_dim guard, so Gemma-4-31B loads with NO pin/custom image) + a **patched chat template** (`stacks/meromero-charrp/patched_chat_template.jinja`, `--chat-template` override) whose Think branch force-opens a **bare `<|channel>`** (NOT `<|channel>thought\n` — full-open defeats the parser) + **two served-names** `char-rp`/`char-rp-thinking` + `--reasoning-parser gemma4` + default `enable_thinking:false`. LiteLLM: `char-rp` (enable_thinking:false → prose) + `char-rp-reasoning` (→ char-rp-thinking served-name, enable_thinking:true → CoT). **★ STREAMING ONLY** — verified 6/6 direct + 3/3 via gateway; **non-streaming does NOT split** (structural: `extract_reasoning` never gets prompt_token_ids so the pre-arm can't fire — vLLM one-shot bug, unchanged across v0.24-0.27; fine because Lobe/OWUI stream). **★ KEY-NAME TRAP that cost me two false negatives:** vLLM streams reasoning in delta.**`reasoning`**; LiteLLM normalizes it to delta.**`reasoning_content`**. Test the RIGHT key per path or you'll wrongly conclude failure. **Credit: the four-arm brokkr/dwarf panel** (thread `01M0JKW44Y…`) traced it from vLLM source — the fix is the force-open template + v0.26.0's #48217 streaming pre-arm. Upstream PR #49797 (full fix, non-streaming too) lands ~v0.28.0 → then it's a clean image bump. char-rp prose verified clean on v0.26.0 (no #49955 trailing-token leak observed). Canonical: `stacks/meromero-charrp/` (compose + patched template), `stacks/litellm/conf/config.yaml`. ROLLBACK: `.env` MEROMERO_IMAGE→latest + drop --chat-template.
- **⛔ RESULT 2026-08-21 — the gemma4 CoT test on a STABLE (v0.27.1) is BLOCKED by a config incompatibility, NOT the parser.** Tried serving the MeroMero NVFP4A16 quant on `vllm/vllm-openai:v0.27.1`. Two-stage failure: (1) v0.27.1's stricter transformers raised `AmbiguousGlobalPerLayerAttributeError: 'head_dim' is per-layer` on the Gemma-4 config; setting `allow_global_per_layer_attribute_access:true` on `text_config` downgraded it to a warning BUT (2) then `gemma4.py load_weights` crashed with **`AssertionError: load weight (512) into parameter (256)`** — **Gemma-4-31B is genuinely HETEROGENEOUS (some layers head_dim 512, not a uniform 256)**, so forcing the global value built wrong-shaped params. The transformers guard was CORRECT; there is no safe override. **The MeroMero quant's config was authored for v0.24.0's Gemma4 loader and cannot load on v0.27.x without a config migration (proper per_layer_config) or a re-quant against the newer transformers.** ⚠ **This also means the eventual gen-seat move to v0.27.2 stable must re-verify any Gemma-4 seat's config-compat** — the transformers heterogeneity change affects all Gemma-4 quants of this vintage. **FULLY REVERTED:** config.json restored (flags removed), compose + image back to `latest` (v0.24.0), gateway char-rp-reasoning removed, char-rp prose verified on v0.24.0. Net: char-rp stays pinned to v0.24.0; MeroMero CoT remains undelivered. **The per-request-kwargs hypothesis was never even reachable** — couldn't load the model to test it. **For RP-with-CoT: gen-reasoning (works now) or a re-quant of MeroMero against v0.27.x transformers (real work, unproven payoff).**
- **⚠️ CORRECTED 2026-08-21 — MeroMero-v2 CoT: NOT a hard wall, and NOT MeroMero-specific. My first conclusion ("gemma4 parser is process-wide") was WRONG.** Read the actual code, not the stale compose comment. **The real mechanism (gemma4-GENERAL, applies to any gemma4 finetune on this template family):** thinking is a **per-request** template toggle — `chat_template.jinja:347-352` emits the generation prompt `<|turn>model\n`, and **only when `enable_thinking` is false** does it prefill an empty `<|channel>thought\n<channel|>` to SUPPRESS thinking; `enable_thinking:true` omits the prefill so the model is free to open a real `<|channel>thought…<channel|>` block. The vLLM parser (`vllm/reasoning/gemma4_utils.py:parse_thinking_output`) **splits on `<|channel>`/`<channel|>` tag PRESENCE — "works with or without enable_thinking," NOT a process-wide flag.** The stale compose comment I trusted cited an OLD parser API (`vllm/parser/gemma4.py:439`) that this container does not run. **So there is no architectural blocker; the two-served-name gen pattern SHOULD work.** **What actually failed my test:** meromero runs `vllm/vllm-openai:latest` (v0.24.0); per-request `chat_template_kwargs.enable_thinking:true` produced no thinking on it, whereas the **gen seat's pinned nightly demonstrably applies per-request `chat_template_kwargs`** (gen-reasoning works). So the practical block is a **vLLM-version / per-request-plumbing issue on v0.24.0**, not the model and not the architecture — and it would hit ANY gemma4 finetune served on that image the same way. **UNVERIFIED FIX (needs a GPU window): re-serve meromero on the nightly image + no process default + per-request enable_thinking; likely yields clean split CoT.** Currently REVERTED to known-good (char-rp prose, process default false, single served-name). ⚠ Kept `MEROMERO_GPU_MEM_UTIL` 0.52→0.51 (0.52 no longer boots next to the bigger orcarouter gen; free 49.02 < 49.38 GiB; 0.51 = KV 2.00× @ 262K).