Captures an evening that ran from an OOM crash-loop to a measured trainee base. The durable lessons, none of which CLAUDE.md can carry: --gpu-memory-utilization sizes the KV cache and does not cover CUDA context or graphs, which is half of why a seat that fit on the 21st stopped fitting on the 24th; the other half is that gen's footprint GROWS WITH UPTIME (38.5 GiB fresh against 45.6 GiB after three days, same container, same flag), so headroom arithmetic against a long-running gen measures a moving number. The stale-chat-template trap turned out to be endemic across third-party Gemma-4 derivatives rather than a one-off, and it is silent in both directions — wrong prompt when serving, train/serve skew when tuning. And a benchmark finding was retracted because 12% on a five-option task is below the 20% chance floor: a below-chance score indicts the instrument before the model, and a preflight can be thorough while aimed in the wrong direction. Records the serving decision for the tuned model with its history intact: LoRA-on-NVFP4 is preferred if it works, merged weights the expected fallback, but the archived root-cause says the objection was never NVFP4-specific — vLLM 0.24.0's LoRA application was a silent no-op proven quant-agnostic, and ana-ml2 now runs 0.26.0. Retest before designing around merge; the answer changes what Eitri's harness must emit, and he is still early. Auto-archival moved 5 entries (Recent decisions) to archival-memory.md; the guards held back the rest of the 78 age-eligible candidates because their bodies carry open deferred-work language, per the keep-when-unsure rule. Index sits at 286 lines, above the ~250 target and reported rather than forced.
8.5 KiB
[2026-08-24] char-rp seat: OOM root-cause, Gemma-4 MoE swap, and the abliterated trainee base
One evening, one thread with brokkr-smithy-dev, five commits: 850e0c3,
27155c0, f509668+24e8826+1bd90ea+3446367+8d6a939, 14ff4a3,
019ccff, 5415fd4.
1. The seat was crash-looping, and the cause was NOT its config
vllm-meromero-rp reported up-but-unreachable, RestartCount climbing (13 by the
time it was examined, not the 4 first reported). Startup logs looked clean all
the way through weights, torch.compile and CUDA-graph capture, then:
torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 336.00 MiB.
GPU 0 has a total capacity of 94.97 GiB of which 195.19 MiB is free.
⚠ --gpu-memory-utilization SIZES THE KV CACHE AND DOES NOT COVER CUDA
CONTEXT, GRAPHS OR NON-TORCH OVERHEAD. gen is configured at 0.43 and actually
held 45.6 GiB. char-rp was at 0.51. The pair was committed to 0.94 of the card
with ~0.6 GiB of real headroom — it fit on the 21st and stopped fitting on the
24th.
Dropped char-rp to 0.47: ~4.8 GiB margin, KV 27.36 → 23.56 GiB, 430,825 → 371,023 tokens against a max-model-len of 262,144. Cost nothing usable — the pool still holds 1.4x a full-length sequence; what is lost is concurrent long requests, not context.
⚠⚠ THE MISSING HALF, found later that evening: gen's footprint GROWS WITH UPTIME. Same container, same 0.43: 45.6 GiB after ~3 days up, 38.5 GiB freshly restarted — ~7 GiB apart. Nothing about char-rp changed between the 21st and the 24th; gen crept up underneath it. Headroom arithmetic done against a long-running gen is measuring a moving number. Measure against a freshly-restarted one.
2. char-rp and char-rp-reasoning are ONE seat, not two
Both LiteLLM routes point at 10.250.50.54:8016/v1 — hosted_vllm/char-rp and
hosted_vllm/char-rp-thinking. brokkr had reported 30/80 and 80/80 failure
rates as two failing services; it was one outage sampled twice. This also
improved a result of theirs: their CoT on/off battery had assumed both aliases
were the same weights under two chat templates, and the routing detail turned an
assumption into a verified fact.
(vllm-charrp-reasoning-nvfp4, the Heretic2 NVFP4+MTP container, has been
stopped for 12+ days and is unrelated — it is not what that alias resolves to.)
3. The seat swapped to the Gemma-4 26B-A4B MoE
Operator-directed straight-across replacement: same port, same served-model-names, so no gateway route or consumer config moved. Rationale is throughput under CoT — the user waits through the whole reasoning block before the first visible token, and the MoE measures ~114 tok/s @32K against the dense 31B's ~40.7.
Serving copy is RedHatAI/gemma-4-26B-A4B-it-NVFP4 (16 GB), chosen over the
other -it quants because it is compressed-tensors (nvfp4-pack-quantized) —
the same loader path the outgoing seat used. Smaller weights at the same 0.47
budget bought 1,724,110 KV tokens against the predecessor's 371,023.
meromero-charrp is retained stopped in created state, labelled
AI - Dormant. Both stacks bind :8016, so rollback is stop-then-start.
4. ⚠ THE STALE-CHAT-TEMPLATE TRAP IS ENDEMIC, NOT A ONE-OFF
Verified by hash across every third-party Gemma-4 derivative pulled:
| build | lines | sha256 (normalised) |
|---|---|---|
upstream google/gemma-4-26B-A4B-it |
390 | 6a1015c47ccfcfa6 |
| RedHatAI NVFP4 (served) | 389 | 6a1015c47ccfcfa6 — the only match |
| llmfan46 heretic | 365 | 0a52be69cda5ab8a |
| TrevorJS abliterated | 266 | 58c66fdee4afa297 |
| jenerallee78 abliterated | 266 | 58c66fdee4afa297 |
| prithivMLmods NVFP4A16 | 266 | 58c66fdee4afa297 |
Three independent repos carrying the identical stale file means it propagated
through the ecosystem. Consequences differ by use and both are silent:
serving a mismatched template renders a different prompt; training through
base/chat_template.jinja means training on a different prompt format than
production serves — train/serve skew, no error, presents as a tuning failure.
The production compose now pins the template explicitly. It is a no-op for the
served weights (the A4 build ships that exact file) and permanently closes the
class. ⚠ If GEMMA4_MODEL ever points at a different checkpoint, the template
default must move with it.
5. A benchmark result was RETRACTED — below chance indicts the instrument
A battery appeared to show Gemma at 12% contradiction detection with CoT off against gen's 81%. An A16 activation-precision control was staged to test whether the quant scheme owned it. Then the operator asked to see the individual items, and the item was ill-posed: it presented two mutually contradicting statements and asked for "the contradicting statement", but contradiction is symmetric. The model consistently named the absolute claim — a defensible reading the labelling scored wrong every time.
⚠ THE TELL WAS IN PLAIN SIGHT: 12% ON A FIVE-OPTION TASK IS BELOW THE 20% CHANCE FLOOR. A below-chance score indicts the instrument before it indicts the model, and neither side reacted to it. I spent the afternoon verifying repo names, config fields, template hashes and tokenizer vocabs — every layer of plumbing — and never asked whether the number itself was possible. A preflight can be thorough and still be aimed in the wrong direction.
Retracted: "the model owns the contradiction deficit"; "domain tuning costs 43
points of contradiction detection" (on a sound instrument it reverses); all
pre-fix T2 numbers. Recorded as a dated superseded-claims table in
stacks/gemma4-charrp/README.md rather than a silent edit.
What survived: the A16 control result — activation precision is close to free on this battery, every other task identical across W4A4 and W4A16 builds.
6. The abliterated trainee base — measured, not assumed
Operator directed a low-damage abliterated instruct build. "Low damage" was treated as a measurable claim; the field spreads from KL 0.09 to 0.4118:
| build | method | KL | refusals |
|---|---|---|---|
| llmfan46 (operator's pick) | Heretic v1.2.0 ARA | 0.1237 | 3/100 |
| TrevorJS | ARA-family | 0.09 | 1/100 effective, 5/686 cross-dataset |
| jenerallee78 | ARA 2-pass | 0.1299 | 7.7% StrongREJECT |
| huihui-ai | remove-refusals | none published | none published |
Fleet anchor: our own work found Heretic at KL 0.12 preserved the MTP head at 83.7% acceptance, so both staged builds sit inside an already-measured band. huihui-ai rejected — no metrics, its card calls the method "a crude, proof-of-concept implementation", it abliterates both thinking and non-thinking modes, and its parameter count runs ~738M over upstream. Operator's independent read matched ("huihui produces garbage").
Abliteration isolated properly (stock BF16 vs llmfan46 BF16, same precision, same pinned template, same 192 items):
T2 contradiction 75% → 59% (−5 items)
T6 spatial 75% → 88% (+4 items)
core 90.0% → 89.4% (−0.6 pts)
It MOVED capability rather than removing it — five lost on contradiction, four gained on spatial, nearly cancelling. Nobody predicted a gain. llmfan46 stands; no case for re-staging on TrevorJS over 0.6 points.
⚠ Read as ~5 and ~4 items at n=32, not as −15.6/+12.5 percent. ⚠ Says nothing about quantization — the stock-NVFP4 T2 figure came from n=16 against n=32, different item sets, n-confounded.
7. ⚠ The production compose hardcodes --quantization compressed-tensors
Pointing the char-rp stack at unquantized BF16 weights crash-loops immediately:
TypeError: CompressedTensorsConfig.__init__() missing 3 required
positional arguments: 'target_scheme_map', 'ignore', 'quant_format'
vLLM trying to read a quantization config out of a checkpoint that has none. 35
restarts before it was caught. Hence stacks/gemma4-trainee-bench/ — a separate
ephemeral stack with no quantization flag, restart: "no" so a bench seat cannot
resurrect itself and block gen's restore, and no homepage labels so it leaves no
permanently-offline card. That detour is why a base swap is now ~5 minutes
instead of ~15.
8. BF16 cannot coexist with gen
48.07 GiB of BF16 weights plus gen's footprint exceeds the 94.97 GiB card before a byte of KV cache. Every BF16 bench window means gen is stopped. Two such windows were run and gen restored both times; the restore was triggered by observing the seat's own throughput logs (a large prefill burst then zero running/zero waiting) rather than waiting on a courtesy message.
Cross-links: 2026-08-24-homepage-uniform-grid