diff --git a/stacks/gemma4-charrp/README.md b/stacks/gemma4-charrp/README.md new file mode 100644 index 0000000..5093be5 --- /dev/null +++ b/stacks/gemma4-charrp/README.md @@ -0,0 +1,131 @@ +# gemma4-charrp — the char-rp seat (ana-ml2 GPU0) + +`google/gemma-4-26B-A4B-it`, NVFP4, serving both halves of the char-rp pair on +`:8016`. Replaced the dense `G4-MeroMero-v2-31B-NVFP4A16` seat on 2026-08-24. + +``` +char-rp non-thinking -> http://10.250.50.54:8016/v1 +char-rp-reasoning thinking -> http://10.250.50.54:8016/v1 +``` + +Two LiteLLM aliases, **one backend**. They are not two seats — this trips people +up, and it cost a peer a mis-attributed benchmark before it was noticed. + +## Three model directories, and they are not interchangeable + +| path | size | what it is | +|---|---|---| +| `gemma4-26b-a4b-it-bf16` | 49 GB | **QLoRA tuning base.** BF16, unquantized. Cannot be served here — 48.10 GiB of weights against ~49 GiB of free GPU0 leaves nothing for KV cache. | +| `gemma4-26b-a4b-it-nvfp4` | 16 GB | **What is served.** RedHatAI, compressed-tensors, W4**A4**. | +| `gemma4-26b-a4b-it-nvfp4a16` | 17 GB | **Activation-axis control**, for benching only. prithivMLmods, compressed-tensors, W4**A16**. | + +All under `/tank/aimodels/`, pulled by the `gemma4-26b-*-dl.py` scripts beside +them with revisions pinned. + +## GPU0 is shared and the budgets must sum under ~0.92 + +`vllm-gen` runs at `--gpu-memory-utilization 0.43` but actually holds ~45.6 GiB +of the 94.97 GiB card — that flag sizes the KV cache and does **not** cover CUDA +context, graphs and non-torch overhead. The predecessor seat sat at 0.51, the +pair summed to 0.94, and on 2026-08-24 it stopped fitting and crash-looped 13 +times with `torch.OutOfMemoryError: ... 195.19 MiB is free`. + +This seat runs 0.47. Raising it means lowering gen's in the same change, and +**check the real numbers, not the flags**: + +```bash +nvidia-smi --query-compute-apps=pid,used_memory --format=csv +``` + +There is no room for a second concurrent seat on this card. That is why the A16 +control below is a swap rather than a parallel deployment. + +## Running the A16 activation-axis control + +Why: brokkr-smithy-dev's 96-item battery scored **12% on contradiction detection +with CoT off**, against gen's 81%, while state tracking, constraint following, +long-context state and the confabulation control all sat at 100%. Not general +degradation — the shape 4-bit input activations produce on the most +reasoning-dense task. Recorded as CONFOUNDED until this run separates the two. + +⚠ **The A16 build ships a STALE CHAT TEMPLATE and the run MUST override it.** +Verified by hash 2026-08-24: upstream `google/gemma-4-26B-A4B-it` is 390 lines, +the A4 build's is 389 and byte-identical to upstream once trailing newlines are +normalised, the A16 build's is 266 and is not. Upstream and A4 open the thinking +path with `{%- set enable_thinking = enable_thinking | default(false) -%}`; the +A16 template has no such set. Its `tokenizer_config.json` `response_schema` also +lacks the `thinking` property. Served with its own template the two arms render +**different prompts**, and a score delta could be the template rather than the +activations. + +Overriding is safe: the tokenizers agree — vocab identical at 262,144 entries, +`added_tokens` identical — so the same template over the same vocab renders the +same token ids. brokkr independently diffed every non-quantization config field +of both builds against the upstream BF16 and found only `transformers_version` +differing, which is a save-time library version rather than a model property. +Residual risk stated honestly on both sides: config identity is not weight +identity, and nobody has done a dequantization pass. + +The procedure, ~2 minutes each way: + +1. Add the template flag to `compose.yaml` (a **no-op for production** — the A4 + build ships this exact file, so pinning it explicitly changes nothing and + guards against precisely the staleness above): + + ```yaml + - --chat-template + - ${GEMMA4_CHAT_TEMPLATE:-/tank/aimodels/gemma4-26b-a4b-it-nvfp4/chat_template.jinja} + ``` + + ⚠ If `GEMMA4_MODEL` is ever pointed at a *different checkpoint*, this default + must move with it. A pinned template is only correct for the checkpoint it + came from — that is the mistake the outgoing MeroMero seat's hand-patched + template was warning about, inverted. + +2. `GEMMA4_MODEL=/tank/aimodels/gemma4-26b-a4b-it-nvfp4a16` in the host `.env`, + then `docker compose up -d --force-recreate vllm-gemma4-charrp`. +3. Tell brokkr; they fire one CoT-off arm, 96 items, under a minute of wall + clock, and report back. +4. Revert `GEMMA4_MODEL` to the `nvfp4` path and recreate. + +Only the **CoT-off** arm is worth running. With thinking on the model already +scores 100% on every task it completes, and a ceiling cannot move. + +⚠ Displacing this seat is an operator decision, not a routine one. It is live. + +## Rollback to MeroMero-v2 + +`stacks/meromero-charrp/` is retained stopped in `created` state, labelled +`AI - Dormant`. Both stacks bind `:8016`, so rollback is **stop-then-start**: + +```bash +cd /opt/docker/compose/gemma4-charrp && docker compose stop vllm-gemma4-charrp +cd /opt/docker/compose/meromero-charrp && docker compose up -d vllm-meromero-rp +``` + +## Gemma-4 flags that are load-bearing + +Carried over from the MeroMero seat because they are architecture-level, not +checkpoint-level: + +- `--tool-call-parser gemma4` + `--enable-auto-tool-choice` — Gemma-4 emits its + own native tool syntax, not the qwen3_coder XML the other seats use. Without + these, any tools-bearing request 400s outright. +- `--reasoning-parser gemma4` — absorbs the `<|channel>` thought markers; without + it they leak into `content` verbatim on the post-tool turn (vllm #45834). +- `--default-chat-template-kwargs '{"enable_thinking": false}'` — **mandatory + companion to the reasoning parser.** The parser defaults `enable_thinking` to + TRUE, which pre-initialises the engine to REASONING, so all plain prose lands + in `reasoning_content` with a null `content` and every char-rp consumer breaks. + Production runs thinking off; the thinking route reaches it per-request. + +## Known defect, not ours + +With thinking on, the model does not reliably terminate on constraint-following. +brokkr measured **32 of 96 calls truncating at a 12,000-token cap, with all 16 +constraint items among them**, and inspected the traces: the reasoning is sound +right up to the point it fails to stop. A non-termination defect, not a +wrong-answer one. `VLLM_USE_V2_MODEL_RUNNER=0` enables +`thinking_token_budget` and is the first lever to try if a usable thinking seat +is ever wanted — deliberately **not** applied here, since it costs the faster +model runner for a mode production does not serve.