diff --git a/persistent-memory.md b/persistent-memory.md index f651a1f..d40fe9d 100644 --- a/persistent-memory.md +++ b/persistent-memory.md @@ -112,7 +112,11 @@ _As of 2026-08-26 ~02:10 PDT โ€” **run 2 is trained, merged, coherence-gated and - **๐ŸŸข RUN 2 COMPLETE AND SERVING as `erp-tune-v2`.** 1312/1312 in **7:22:44**, `train_loss` **2.839**, on the **stock `google/gemma-4-26B-A4B-it`** base (operator's call, 08-25). Endpoint `http://10.250.50.54:8098/v1`, container `erp-eval-v2` on GPU0 (84,272 MiB, `restart unless-stopped`), weights `/tank/erp-tune/serve/merged-run02` (bf16 merged, 51.6 GB), adapter `/tank/erp-tune/run-02/adapter/`. Gates all passed: `lora_B` **205/205** (min 0.6826), vision_tower 0, merge verified, upstream template shipped, **coherence 5/5**. โ†’ `persistent-memory.d/2026-08-26-erp-run2-complete-and-served.md` - **โš  THE MASK IS PROVEN BY THE LOSS-TOKEN DELTA, NOT BY THE p50 MATCH.** โˆ’221,712 loss tokens against **byte-identical context tokens**, unchanged records, same nine drops; cross-checks to brokkr's independent corpus-side figure within 4%. The p50 match (19.79 vs run 1's 19.75) is evidence about *timing stability across a base swap* and **cannot** evidence masking โ€” step time is insensitive to which positions carry loss. Do not re-assert it; I did, and brokkr was right. -- **๐Ÿ”„ BATTERY IS RUNNING โ€” TUNED ARM LIVE, BASE SET COMPLETE.** Base arm ran 08:54:52Zโ†’09:03:43Z, **zero errors/truncations/empties/degenerates across 1,728 generations**. Swapped to `erp-tune-v2` on `:8098`; **image digest `sha256:4091d5593f77โ€ฆ` VERIFIED IDENTICAL across both arms** (brokkr's requirement โ€” a differing digest voids the run). brokkr is now running the tuned set with `--freeze-from` the base's marker list. **The floors the tuned deltas must clear:** reasoning core **0.5 pt**, diversity overall **0.0125** (rp 0.0153, story 0.0096), story attractor **0.0000**, memorisation signal 0.00%. โš  **The rp family froze ZERO markers**, so its attractor metric is structurally 0.0 on both arms and cannot discriminate โ€” rp is measured on the distance axis only; do not quote an rp attractor delta. โš  Base's slop attractor: **"Elias" in 92/96 bare-prompt stories** (published figure on the same family is 102/144 โ€” ours is an independent replication at a higher rate, not a novel finding). +- **๐Ÿ”ด GATE VERDICT: FAIL โ€” recorded as FAIL, and the T6/T3 trade is the actual result.** brokkr's full write-up at `4973991`, `research/R47-premium-corpus-gate/run02-gate/RESULT-run02-gate.md`. 3,456 generations, zero API errors, digests verified identical so the run is not void. **Gate 1 T6-must-not-regress: 73.5 โ†’ 88.5, +15.0 โœ…** (run 1 FAILED this same axis at โˆ’3.5). **Gate 2 no-task-regresses->1-item: T3 โˆ’12.0, T4 โˆ’5.5 โŒ.** Gate 3 T5 control 100% โœ…. Gate 4 latency 0.11s โœ…. T3 reads 88 on both tuned passes โ€” not variance. **Neither run ships; together they price what the abliteration was costing, which is a question neither could answer alone.** Should-help axis worked: diversity **+0.196 (~15x floor)**, confound-tested against attrition (matched-k=8 keeps +0.19, attrition โ‰ˆ7% of the effect) and against length (story got LONGER 669โ†’727 while rp got SHORTER 137โ†’88, same gain both โ€” a length artifact would have opposite signs). Memorisation none. โ†’ `persistent-memory.d/2026-08-26-erp-run2-complete-and-served.md` +- **โš  THE FINDING WITH LEGS BEYOND THIS TUNE: an output-stability regression VISIBLE ONLY ON LONG-FORM.** Truncated **0/384 โ†’ 38/384**, degenerate **0/384 โ†’ 19/384** โ€” and the reasoning battery saw **zero of it on either arm across four passes**, because its answers are short. **A gate composed only of short-answer tasks would have passed this cleanly.** Playbook ยง4.6.3. brokkr's block-1 rp VOID (degeneracy 12/95 past the preregistered 10% budget) was the guard catching it first, recorded as VOID not reinterpreted. +- **โš  PIPPA'S 123-WORD PRODUCT CLIP IS IN THE MIX AND THE DPO PAIRS WOULD INHERIT IT.** Measured 08-26: pippa bot turns max **123 exactly**, 100% at-or-under, **0.00% in the 124-130 band** โ€” a wall, not a preference (every other root crosses its p99 smoothly). The mechanism is the exposure asymmetry: **pippa is 70.3% of bot TURNS but 37.5% of bot WORDS, and length is learned per turn.** Generalises โ€” a clipped root is over-represented in the length signal by exactly the ratio its clipping creates. โš  Counter-evidence: the tune landed near pippa's **median 67**, not its **cap 123**, so "learned central tendency" is better supported than "learned the limit". Closing test is the tuned arm's own rp length distribution vs pippa's and bluemoon's. Scripts `/tank/erp-tune/pippa_clip.py`, `clip_share.py`. โ†’ `docs/pfi/erp-dpo-stage-prep.md` +- **๐ŸŸข SEATS: `erp-tune-v2` UP on `:8098`, base arm torn down, GPU0 has ~13 GB spare.** The tune stays served (operator asked; it is coherent โ€” the gate verdict is a research result, not a serving fault) and stays **OUT of the LiteLLM gateway**, which now reads as clearly right rather than merely cautious: **a FAILED tune must not be one alias resolution away from a consumer who has not read the thread.** +- **โณ SUPERSEDED โ€” the mid-swap state.** Base arm ran 08:54:52Zโ†’09:03:43Z, **zero errors/truncations/empties/degenerates across 1,728 generations**. Swapped to `erp-tune-v2` on `:8098`; **image digest `sha256:4091d5593f77โ€ฆ` VERIFIED IDENTICAL across both arms** (brokkr's requirement โ€” a differing digest voids the run). brokkr is now running the tuned set with `--freeze-from` the base's marker list. **The floors the tuned deltas must clear:** reasoning core **0.5 pt**, diversity overall **0.0125** (rp 0.0153, story 0.0096), story attractor **0.0000**, memorisation signal 0.00%. โš  **The rp family froze ZERO markers**, so its attractor metric is structurally 0.0 on both arms and cannot discriminate โ€” rp is measured on the distance axis only; do not quote an rp attractor delta. โš  Base's slop attractor: **"Elias" in 92/96 bare-prompt stories** (published figure on the same family is 102/144 โ€” ours is an independent replication at a higher rate, not a novel finding). - **โณ SUPERSEDED โ€” the mid-swap state.** brokkr **withdrew** his both-arms-concurrent requirement (his diversity battery emits the frozen marker list to a *file*, so the arms were never a live dependency; the real requirement is all of ONE arm's passes on ONE instance before the swap). So: **sequential, no fleet seats displaced, operator not woken.** Live now: `erp-eval-base` on `:8099` (`erp-base-stock`, stock bf16 via symlinks into `/tank/aimodels`). **`erp-eval-v2` is DOWN and its container REMOVED**; weights are on disk so it is a cheap restart. Swap back with `/tank/erp-tune/serve-arm.sh tuned` when he reports the base set complete. โš  **Both arms go through `serve-arm.sh`**, which pins the vLLM image **by digest** (`sha256:251eba5cc7c1โ€ฆ`) โ€” a version change between arms is a base swap nobody would see โ€” and mounts `/tank/aimodels` for BOTH arms even though only base needs it, because a mount that differs between arms is a difference between arms. - **โณ SUPERSEDED โ€” the GPU arithmetic that forced sequential.** He needs **both arms reachable in one window** (the diversity battery mines its frozen marker list from base). **Two bf16 26B arms = 98 GB of weights on a 97.9 GB card**; GPU1 has ~30 GB free under six shared seats. Reduced `gpu-memory-utilization` does not help โ€” weights are the floor. Options put to him: sequential-with-every-nuisance-variable-pinned, or displace GPU1 seats (**operator call, affects other consumers**). Thread `01M0WQ8W5574KMEVCHCEKEXNS5`. - **๐Ÿ”ด `erp-tune-v1` IS STILL IN THE LITELLM GATEWAY AND RETURNS HTTP 500.** I stopped its container to free GPU0 and left the route dangling. It is **config-file-defined (`db_model: false`)**, so removal needs a config edit **plus a gateway reload**, and a reload briefly interrupts all fleet LLM traffic. Judged a 2am restart the worse trade at 1 request/24h. **Batch the removal with the v2-registration decision.** @@ -125,6 +129,8 @@ _As of 2026-08-26 ~02:10 PDT โ€” **run 2 is trained, merged, coherence-gated and ## Recent decisions +- `[2026-08-26]` **Run 2's gate FAILED and is recorded as a FAIL** โ€” T3 constraint โˆ’12.0 against a ~1 pt floor. But gate 1 is the result: **T6 spatial +15.0 where run 1 failed the same axis at โˆ’3.5**, base swap the only intended variable. Neither run ships; the pair prices what the abliteration cost. Plus the long-form-only stability regression a short-answer gate would have passed, and PIPPA's 123-word clip in the length signal. โ†’ `persistent-memory.d/2026-08-26-erp-run2-complete-and-served.md` + - `[2026-08-26]` **Run 2 complete, merged, coherence-gated and serving as `erp-tune-v2`** โ€” stock-instruct base, and the mask proven by a โˆ’221,712 loss-token delta against byte-identical context. Also the p50 claim I asserted and had to withdraw. โ†’ `persistent-memory.d/2026-08-26-erp-run2-complete-and-served.md` - `[2026-08-26]` **Playbook ยง4 written: "when the artifact lies about itself"** โ€” seven landmines plus a pre-launch checklist, from a night in which *three separate fixes each shipped a check that could not fail*. The unifying line is brokkr's: when you change what an artifact means, every derived artifact keyed on the old meaning is now a liar. Commits `dae6ede` โ†’ `d54f256`; the doc is `docs/pfi/training-throughput-playbook.md` (filename kept for inbound links; scope is now wider than the name). - `[2026-08-26]` **Served under a NEW name on a NEW port (`erp-tune-v2` / :8098), never re-pointing `erp-tune-v1`.** Run 1's artifact still exists and is still what that name refers to; re-pointing would be the silent substitution the standing no-false-aliases rule forbids. brokkr independently asked for the same and additionally wants the concrete backing model + date in provenance, not just the alias โ€” an alias has silently changed meaning under recorded results before.