From b5bbc29b9140665972d875cb7fc1eee7b4b65d5e Mon Sep 17 00:00:00 2001 From: Vuong Hoang Date: Wed, 26 Aug 2026 02:25:35 -0700 Subject: [PATCH] =?UTF-8?q?memory:=20gate=20verdict=20FAIL=20=E2=80=94=20a?= =?UTF-8?q?nd=20the=20T6/T3=20trade=20is=20what=20the=20pair=20of=20runs?= =?UTF-8?q?=20bought?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Records the verdict as a FAIL without rounding it off, and the three findings worth more than the verdict: - T6 spatial +15.0 where run 1 failed the same axis at -3.5, with the base swap as the only intended variable. Neither run ships; together they price what the abliteration was costing, which neither could answer alone. - an output-stability regression visible ONLY on long-form (truncated 0->38, degenerate 0->19 per 384) that the reasoning battery could not see across four passes because its answers are short - PIPPA's 123-word product clip sitting in the length signal at 70.3% of bot TURNS against 37.5% of bot WORDS, with the counter-evidence recorded too (the tune landed near the median, not the cap) Also records why keeping the tune out of the LiteLLM gateway now reads as clearly right rather than merely cautious: a FAILED tune must not be one alias resolution away from a consumer who has not read the thread. --- persistent-memory.md | 8 +++++++- 1 file changed, 7 insertions(+), 1 deletion(-) diff --git a/persistent-memory.md b/persistent-memory.md index f651a1f..d40fe9d 100644 --- a/persistent-memory.md +++ b/persistent-memory.md @@ -112,7 +112,11 @@ _As of 2026-08-26 ~02:10 PDT โ€” **run 2 is trained, merged, coherence-gated and - **๐ŸŸข RUN 2 COMPLETE AND SERVING as `erp-tune-v2`.** 1312/1312 in **7:22:44**, `train_loss` **2.839**, on the **stock `google/gemma-4-26B-A4B-it`** base (operator's call, 08-25). Endpoint `http://10.250.50.54:8098/v1`, container `erp-eval-v2` on GPU0 (84,272 MiB, `restart unless-stopped`), weights `/tank/erp-tune/serve/merged-run02` (bf16 merged, 51.6 GB), adapter `/tank/erp-tune/run-02/adapter/`. Gates all passed: `lora_B` **205/205** (min 0.6826), vision_tower 0, merge verified, upstream template shipped, **coherence 5/5**. โ†’ `persistent-memory.d/2026-08-26-erp-run2-complete-and-served.md` - **โš  THE MASK IS PROVEN BY THE LOSS-TOKEN DELTA, NOT BY THE p50 MATCH.** โˆ’221,712 loss tokens against **byte-identical context tokens**, unchanged records, same nine drops; cross-checks to brokkr's independent corpus-side figure within 4%. The p50 match (19.79 vs run 1's 19.75) is evidence about *timing stability across a base swap* and **cannot** evidence masking โ€” step time is insensitive to which positions carry loss. Do not re-assert it; I did, and brokkr was right. -- **๐Ÿ”„ BATTERY IS RUNNING โ€” TUNED ARM LIVE, BASE SET COMPLETE.** Base arm ran 08:54:52Zโ†’09:03:43Z, **zero errors/truncations/empties/degenerates across 1,728 generations**. Swapped to `erp-tune-v2` on `:8098`; **image digest `sha256:4091d5593f77โ€ฆ` VERIFIED IDENTICAL across both arms** (brokkr's requirement โ€” a differing digest voids the run). brokkr is now running the tuned set with `--freeze-from` the base's marker list. **The floors the tuned deltas must clear:** reasoning core **0.5 pt**, diversity overall **0.0125** (rp 0.0153, story 0.0096), story attractor **0.0000**, memorisation signal 0.00%. โš  **The rp family froze ZERO markers**, so its attractor metric is structurally 0.0 on both arms and cannot discriminate โ€” rp is measured on the distance axis only; do not quote an rp attractor delta. โš  Base's slop attractor: **"Elias" in 92/96 bare-prompt stories** (published figure on the same family is 102/144 โ€” ours is an independent replication at a higher rate, not a novel finding). +- **๐Ÿ”ด GATE VERDICT: FAIL โ€” recorded as FAIL, and the T6/T3 trade is the actual result.** brokkr's full write-up at `4973991`, `research/R47-premium-corpus-gate/run02-gate/RESULT-run02-gate.md`. 3,456 generations, zero API errors, digests verified identical so the run is not void. **Gate 1 T6-must-not-regress: 73.5 โ†’ 88.5, +15.0 โœ…** (run 1 FAILED this same axis at โˆ’3.5). **Gate 2 no-task-regresses->1-item: T3 โˆ’12.0, T4 โˆ’5.5 โŒ.** Gate 3 T5 control 100% โœ…. Gate 4 latency 0.11s โœ…. T3 reads 88 on both tuned passes โ€” not variance. **Neither run ships; together they price what the abliteration was costing, which is a question neither could answer alone.** Should-help axis worked: diversity **+0.196 (~15x floor)**, confound-tested against attrition (matched-k=8 keeps +0.19, attrition โ‰ˆ7% of the effect) and against length (story got LONGER 669โ†’727 while rp got SHORTER 137โ†’88, same gain both โ€” a length artifact would have opposite signs). Memorisation none. โ†’ `persistent-memory.d/2026-08-26-erp-run2-complete-and-served.md` +- **โš  THE FINDING WITH LEGS BEYOND THIS TUNE: an output-stability regression VISIBLE ONLY ON LONG-FORM.** Truncated **0/384 โ†’ 38/384**, degenerate **0/384 โ†’ 19/384** โ€” and the reasoning battery saw **zero of it on either arm across four passes**, because its answers are short. **A gate composed only of short-answer tasks would have passed this cleanly.** Playbook ยง4.6.3. brokkr's block-1 rp VOID (degeneracy 12/95 past the preregistered 10% budget) was the guard catching it first, recorded as VOID not reinterpreted. +- **โš  PIPPA'S 123-WORD PRODUCT CLIP IS IN THE MIX AND THE DPO PAIRS WOULD INHERIT IT.** Measured 08-26: pippa bot turns max **123 exactly**, 100% at-or-under, **0.00% in the 124-130 band** โ€” a wall, not a preference (every other root crosses its p99 smoothly). The mechanism is the exposure asymmetry: **pippa is 70.3% of bot TURNS but 37.5% of bot WORDS, and length is learned per turn.** Generalises โ€” a clipped root is over-represented in the length signal by exactly the ratio its clipping creates. โš  Counter-evidence: the tune landed near pippa's **median 67**, not its **cap 123**, so "learned central tendency" is better supported than "learned the limit". Closing test is the tuned arm's own rp length distribution vs pippa's and bluemoon's. Scripts `/tank/erp-tune/pippa_clip.py`, `clip_share.py`. โ†’ `docs/pfi/erp-dpo-stage-prep.md` +- **๐ŸŸข SEATS: `erp-tune-v2` UP on `:8098`, base arm torn down, GPU0 has ~13 GB spare.** The tune stays served (operator asked; it is coherent โ€” the gate verdict is a research result, not a serving fault) and stays **OUT of the LiteLLM gateway**, which now reads as clearly right rather than merely cautious: **a FAILED tune must not be one alias resolution away from a consumer who has not read the thread.** +- **โณ SUPERSEDED โ€” the mid-swap state.** Base arm ran 08:54:52Zโ†’09:03:43Z, **zero errors/truncations/empties/degenerates across 1,728 generations**. Swapped to `erp-tune-v2` on `:8098`; **image digest `sha256:4091d5593f77โ€ฆ` VERIFIED IDENTICAL across both arms** (brokkr's requirement โ€” a differing digest voids the run). brokkr is now running the tuned set with `--freeze-from` the base's marker list. **The floors the tuned deltas must clear:** reasoning core **0.5 pt**, diversity overall **0.0125** (rp 0.0153, story 0.0096), story attractor **0.0000**, memorisation signal 0.00%. โš  **The rp family froze ZERO markers**, so its attractor metric is structurally 0.0 on both arms and cannot discriminate โ€” rp is measured on the distance axis only; do not quote an rp attractor delta. โš  Base's slop attractor: **"Elias" in 92/96 bare-prompt stories** (published figure on the same family is 102/144 โ€” ours is an independent replication at a higher rate, not a novel finding). - **โณ SUPERSEDED โ€” the mid-swap state.** brokkr **withdrew** his both-arms-concurrent requirement (his diversity battery emits the frozen marker list to a *file*, so the arms were never a live dependency; the real requirement is all of ONE arm's passes on ONE instance before the swap). So: **sequential, no fleet seats displaced, operator not woken.** Live now: `erp-eval-base` on `:8099` (`erp-base-stock`, stock bf16 via symlinks into `/tank/aimodels`). **`erp-eval-v2` is DOWN and its container REMOVED**; weights are on disk so it is a cheap restart. Swap back with `/tank/erp-tune/serve-arm.sh tuned` when he reports the base set complete. โš  **Both arms go through `serve-arm.sh`**, which pins the vLLM image **by digest** (`sha256:251eba5cc7c1โ€ฆ`) โ€” a version change between arms is a base swap nobody would see โ€” and mounts `/tank/aimodels` for BOTH arms even though only base needs it, because a mount that differs between arms is a difference between arms. - **โณ SUPERSEDED โ€” the GPU arithmetic that forced sequential.** He needs **both arms reachable in one window** (the diversity battery mines its frozen marker list from base). **Two bf16 26B arms = 98 GB of weights on a 97.9 GB card**; GPU1 has ~30 GB free under six shared seats. Reduced `gpu-memory-utilization` does not help โ€” weights are the floor. Options put to him: sequential-with-every-nuisance-variable-pinned, or displace GPU1 seats (**operator call, affects other consumers**). Thread `01M0WQ8W5574KMEVCHCEKEXNS5`. - **๐Ÿ”ด `erp-tune-v1` IS STILL IN THE LITELLM GATEWAY AND RETURNS HTTP 500.** I stopped its container to free GPU0 and left the route dangling. It is **config-file-defined (`db_model: false`)**, so removal needs a config edit **plus a gateway reload**, and a reload briefly interrupts all fleet LLM traffic. Judged a 2am restart the worse trade at 1 request/24h. **Batch the removal with the v2-registration decision.** @@ -125,6 +129,8 @@ _As of 2026-08-26 ~02:10 PDT โ€” **run 2 is trained, merged, coherence-gated and ## Recent decisions +- `[2026-08-26]` **Run 2's gate FAILED and is recorded as a FAIL** โ€” T3 constraint โˆ’12.0 against a ~1 pt floor. But gate 1 is the result: **T6 spatial +15.0 where run 1 failed the same axis at โˆ’3.5**, base swap the only intended variable. Neither run ships; the pair prices what the abliteration cost. Plus the long-form-only stability regression a short-answer gate would have passed, and PIPPA's 123-word clip in the length signal. โ†’ `persistent-memory.d/2026-08-26-erp-run2-complete-and-served.md` + - `[2026-08-26]` **Run 2 complete, merged, coherence-gated and serving as `erp-tune-v2`** โ€” stock-instruct base, and the mask proven by a โˆ’221,712 loss-token delta against byte-identical context. Also the p50 claim I asserted and had to withdraw. โ†’ `persistent-memory.d/2026-08-26-erp-run2-complete-and-served.md` - `[2026-08-26]` **Playbook ยง4 written: "when the artifact lies about itself"** โ€” seven landmines plus a pre-launch checklist, from a night in which *three separate fixes each shipped a check that could not fail*. The unifying line is brokkr's: when you change what an artifact means, every derived artifact keyed on the old meaning is now a liar. Commits `dae6ede` โ†’ `d54f256`; the doc is `docs/pfi/training-throughput-playbook.md` (filename kept for inbound links; scope is now wider than the name). - `[2026-08-26]` **Served under a NEW name on a NEW port (`erp-tune-v2` / :8098), never re-pointing `erp-tune-v1`.** Run 1's artifact still exists and is still what that name refers to; re-pointing would be the silent substitution the standing no-false-aliases rule forbids. brokkr independently asked for the same and additionally wants the concrete backing model + date in provenance, not just the alias โ€” an alias has silently changed meaning under recorded results before.