docs(lv-mccarthy): record the gate result -- voice passes, memorisation is the cleanest in the line, NOT shipped
720 generations, 3 arms x 60 held-out beats x 4 seeds, against the design frozen in GATE-PREREG.md before any arm existed. AXIS A VOICE -- PASS, both candidates, both reads. Span 0.661 -> 0.370 = 0.291 achievable; ckpt900 closed 59.1% (+0.172, but only 1.2x its floor), ckpt450 52.2% (+0.152 at 2.9x). The normalised secondary read HOLDS at +0.124 / +0.114, so about three quarters of the gain survives stripping every punctuation mark -- the voice is not the cheap win the register made available. AXIS B NOT COPIED -- ckpt450 is the cleanest result in the line. 0.12 hit-rate against the author's own held-out 0.12, and its longest match (11 words) is SHORTER than the author's coincidental longest (12). All 96 matched runs were read: stock grammar in the commonest words, the name-shaped hits are the RENAMED inventions, nothing protectable. The amendment is why this reads as clean -- the defective base control would have shown 0.12 vs 0.00 as a 12x red flag. Separately measured: the "his register makes collisions inevitable" story that was FALSE for Hemingway (0.01) is TRUE for McCarthy (0.12). Neither transfers. AXIS C NO DAMAGE -- FAIL, both, and it survives reading. 20% (ckpt450) / 28% (ckpt900) of generations overshoot the 90-140 band against base's 1%; p90 171/190 words, max 297/279. The worst case is degenerate looping, not a long McCarthy sentence. Base is GOOD on this axis here (0.89 in-band vs Hemingway's 0.05), so the adapter measurably makes instruction-following worse. NOT SHIPPED. Section 7 rule 3 makes axis C disqualifying outright. Recorded honestly: my own prereg's axis C transcribed score_beats.py's v1 criteria, including "in-band up on base", which the operator RETIRED on 2026-09-15 for exactly the reason it fails here -- base maxes it, so it is unsatisfiable on this corpus regardless of adapter quality. Under the operator's v2 (ran-on only) ckpt450 passes by 0.01 against a 0.200 floor. That reading was found AFTER the numbers and was NOT used; lv-bronte's floor defect is the in-line precedent for finding one and declining to exploit it. The prereg gets fixed prospectively for the next author, not re-read for this one. And the finding worth more than the adapter: the two-epoch recipe is now 0 for 3, and this time the loss curve was CONFIDENTLY wrong. On Bronte and Hemingway the epoch-1/epoch-2 checkpoints were tied, so preferring the earlier one cost nothing. Here the curve resolved epoch 2 as better at 4.9x the median neighbour jitter -- and epoch 2 lost every axis that resolves: 4.0x wider seed spread, 1.8x the author's memorisation rate against 1.0x, more ran-on, worse on-beat. Its only win is a 0.019 voice point estimate, inside the floor, and its spread is one outlier seed -- the third occurrence of that shape in the later checkpoint after lv-bronte's ckpt925 and lv-hemingway's ckpt1750. Raw artifacts committed at scripts/mccarthy-corpus/gate-results/ so the claims can be re-read without gx10.
This commit is contained in:
+34
-11
@@ -1,6 +1,6 @@
|
||||
# Persistent memory — eshpfi-management
|
||||
|
||||
_Last updated: 2026-09-21 ~14:30 PT (⭐ **the ops log is BUILT** and then failed four ways in its first hours — every one recording something unfindable; the day's subject was instruments that report without looking. ⭐ Draupnir engine COMPLETE and acceptance-tested on irv-ml1. ⭐ Booth gained blur + a closed keep round trip after shipping two dead controls. ⭐ claude-bot is an org Owner; cicada+draupnir moved to `pfi`; vh-token use standing-authorized from the vault. ⭐ nh3-dev 84%→72%, and a LoRA adapter rescued from a 3-day-swept /tmp. ⚠ lv-mccarthy's run outcome STILL UNVERIFIED after two days.)_
|
||||
_Last updated: 2026-09-21 ~16:45 PT (⭐ **the ops log is BUILT** and then failed four ways in its first hours — every one recording something unfindable; the day's subject was instruments that report without looking. ⭐ Draupnir engine COMPLETE and acceptance-tested on irv-ml1. ⭐ Booth gained blur + a closed keep round trip after shipping two dead controls. ⭐ claude-bot is an org Owner; cicada+draupnir moved to `pfi`; vh-token use standing-authorized from the vault. ⭐ nh3-dev 84%→72%, and a LoRA adapter rescued from a 3-day-swept /tmp. ⭐⭐ lv-mccarthy GATED: voice passes and its memorisation is the cleanest in the line, but 20-28% of generations blow the length band — NOT shipped, and the two-epoch recipe is now 0 for 3 with the loss curve confidently wrong.)_
|
||||
|
||||
> **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its
|
||||
> `Written:` stamp is under **8 hours** old, read it (it carries the in-flight
|
||||
@@ -115,18 +115,36 @@ no longer deployed sidecars here. See Recent decisions.)
|
||||
|
||||
## Current state / in-flight
|
||||
|
||||
_As of 2026-09-21 ~14:30 PT._
|
||||
_As of 2026-09-21 ~16:40 PT._
|
||||
|
||||
### ⚠ FIRST — lv-mccarthy's run outcome is STILL UNVERIFIED
|
||||
### ⚠ FIRST — lv-mccarthy is GATED and NOT SHIPPED; the decision is the operator's
|
||||
|
||||
Carried unchanged from the 2026-09-19 handoff and untouched for two full days.
|
||||
Look at `~/r49-runs/mccarthy-4b-pairs-3ep/` before anything else. ⚠ `pfi-gx10`
|
||||
does not resolve from nh3-dev by that name. Assume nothing — it was never
|
||||
checked, not checked-and-found-good. Then: checkpoint selection off the loss
|
||||
curve → the v2 gate (`eval-*.sh`) → ship-or-park. ⚠ Before the gate, settle how
|
||||
the VOICE axis is read: D1 pre-registered a punctuation-normalised secondary
|
||||
read, and `--system-from` hands the register's tics to the BASE arm too, which
|
||||
changes what the primary number means.
|
||||
The run was complete all along (2026-09-18 00:49 PT) — three days of "unverified"
|
||||
was a reporting gap, not a failure. The v2 gate has now RUN, 720 generations, and
|
||||
both candidates are **disqualified by axis C**. Full record:
|
||||
`persistent-memory.d/2026-09-21-lv-mccarthy-gate.md`. ⚠ `pfi-gx10` does not
|
||||
resolve from nh3-dev by name — it is **10.100.50.60**.
|
||||
|
||||
**The shape of it:** voice PASSES on both candidates and both reads (ckpt450 at
|
||||
2.9× its floor, and ~3/4 of the gain survives having every punctuation mark
|
||||
stripped, so it is not the cheap win). Memorisation on **ckpt450 is the cleanest
|
||||
result in the line** — 0.12 against the author's own held-out 0.12, longest match
|
||||
11 words against the author's coincidental 12, all 96 matches read and every one
|
||||
stock grammar. **Damage is the blocker**: 20% (ckpt450) / 28% (ckpt900) of
|
||||
generations overshoot the 90–140 band against base's 1%, worst case a degenerate
|
||||
`he looked at X and he looked at Y` loop at 279 words.
|
||||
|
||||
**Two things a next session must not re-litigate:**
|
||||
1. **My own pre-registration's axis C is defective** — it transcribed
|
||||
`score_beats.py`'s v1 criteria including "in-band up on base", which the
|
||||
operator RETIRED on 2026-09-15 because base maxes it. Under the operator's v2
|
||||
(ran-on only) ckpt450 passes by **0.01 against a 0.200 floor**. That reading
|
||||
was found AFTER the numbers, so it was not used. **Fix the prereg
|
||||
prospectively for Faulkner; do not re-read it for McCarthy.**
|
||||
2. **ckpt450 is the candidate, not ckpt900** (the loss minimum). It wins every
|
||||
resolving axis. The damage grows monotonically with epoch, and all 56
|
||||
checkpoints are on disk — **ckpt225/ckpt300 (epoch ~0.5–0.65) are untested**
|
||||
and are the obvious next probe, one arm each.
|
||||
|
||||
### Draupnir engine — COMPLETE on irv-ml1, acceptance passing
|
||||
|
||||
@@ -160,6 +178,11 @@ the althing thread.
|
||||
|
||||
## Recent decisions
|
||||
|
||||
- `[2026-09-21]` ⭐⭐⭐ **lv-mccarthy GATED and NOT SHIPPED — voice passes, memorisation is the cleanest in the line, and the length discipline is gone.** 720 generations, 3 arms. Voice +0.152 at **2.9×** floor for ckpt450 (ckpt900 +0.172 at only 1.2×, its spread one outlier seed), and **~3/4 of the gain survives stripping every punctuation mark**, so it is not the cheap win. ⭐ Memorisation: **ckpt450 at 0.12 against the author's own held-out 0.12 — identical, longest match 11 words against the author's coincidental 12**, all 96 matches READ and every one stock grammar (`he looked at the wolf and he looked at him`); the name-shaped hits are the RENAMED inventions. Axis C is the blocker: **20% / 28% of generations overshoot the 90–140 band against base's 1%**, worst case a degenerate loop at 279 words. ⚠ My own prereg's axis C transcribed `score_beats.py`'s **v1** criteria including "in-band up on base", which the operator RETIRED 2026-09-15 because base maxes it — under the v2 (ran-on only) ckpt450 passes by **0.01 against a 0.200 floor**, but that reading was found AFTER the numbers and was not used. Fix the prereg prospectively. → `persistent-memory.d/2026-09-21-lv-mccarthy-gate.md`
|
||||
- `[2026-09-21]` ⭐⭐⭐ **The two-epoch recipe is now 0 for 3, and this time the loss curve was CONFIDENTLY wrong.** On Brontë and Hemingway the epoch-1/epoch-2 checkpoints were TIED on eval loss, so preferring the earlier one cost nothing. Here the curve RESOLVED epoch 2 as better — ckpt900 +4.9× the 0.00393 median neighbour jitter above ckpt450, nowhere near tied — and epoch 2 lost every axis that resolves: **4.0× wider seed spread** (0.148 vs 0.037), **1.8× the author's memorisation rate vs 1.0×**, more ran-on (0.28 vs 0.20), worse on-beat. Its only win is a 0.019 voice point estimate, inside the floor, and its spread is ONE outlier seed (0.605 vs 0.457/0.531/0.554) — the third occurrence of that shape in the later checkpoint after lv-bronte's ckpt925 and lv-hemingway's ckpt1750. **Durable: on this schedule the eval-loss minimum is not the ship candidate, and the curve's CONFIDENCE about it carries no information.** Default this for Faulkner/Morrison/Chandler rather than re-deriving it.
|
||||
- `[2026-09-21]` ⭐⭐ **The memorisation control that shipped broken is now committed, and it turned a 12× red flag into a clean pass.** `memorization_check.py` used **base-unadapted** as its negative control; base writes summary while the adapted arms write pastiche, so it cannot collide with a register it does not imitate and its zero is unearned. The lv-hemingway gate established this, computed the correct control (the author's own held-out text) **by hand**, and left no instrument — so the finding was not reproducible. Now `--train-only --heldout-reference`, refusing the unsafe combination, validated by reproducing Hemingway's hand-computed row **to the digit**. ⭐ On McCarthy it is the difference between reporting ckpt450 as memorising (0.12 vs base 0.00 = 12×) and clean (0.12 vs the author's 0.12 = 1.0×). ⭐ And the "the register makes collisions inevitable" story that was FALSE for Hemingway (0.01) is TRUE for McCarthy (0.12) — measured, not assumed; neither number transfers between authors. Also new: `show_memorisation_matches.py`, because rate and exposure are different questions and the reading was hand-done too. Commits `b4ba731` `7eadbd6`.
|
||||
- `[2026-09-21]` ⚠⚠ **My confound detector was counting apostrophes as quote marks, and I saw it FIRE before I saw the bug.** `voice_distance.py`'s quote class shipped with `'` and `’` in it — on the one corpus whose signature is `dont`/`aint`/`wont`. It reported the held-out reference at 121.1 "quote marks" per 10k for a corpus whose builder **ASSERTS 0.0**, and fired the pre-registered trigger on a base arm whose true density is 19.9. Fixing a detector to measure the quantity the frozen rule names is not moving the rule, but the fix un-fires the trigger, which is indistinguishable from shopping — **so the trigger was made MOOT instead of adjudicated**: the normalised read is load-bearing unconditionally, both columns reported, zero verdict effect. ⚠ The lesson: I controlled `strip_punct` (2500→0) and the byte-identity of the default path, and never asked the quote counter for a value whose answer I already knew. Commit `0d80e49`.
|
||||
|
||||
- `[2026-09-21]` ⭐⭐⭐ **The ops log shipped — and then failed FOUR ways in its first hours, every one recording something unfindable.** Claim dropped by a sub-tool, hook behind graphify's eight `exit 0`s, no handle in the env, ssh-target written as a hostname. The general shape is **configured ≠ effective**; twelve instruments reported confidently and wrongly across three days, five of them mine. → `persistent-memory.d/2026-09-21-ops-log-and-the-instruments-that-lied.md`
|
||||
|
||||
- `[2026-09-21]` ⭐⭐ **The Booth gained blur + a closed keep round trip, after shipping TWO controls that did nothing** — a reveal handler Jinja discarded for sitting after `{% endblock %}`, and a `×` a sibling form covered by 30×22 px. Both operator-found, both from reading templates instead of rendering them. `scripts/layout-probe.py` took four iterations to become trustworthy. ⚠ Blur is COSMETIC and a test asserts the 200 on purpose. → `persistent-memory.d/2026-09-21-booth-two-dead-controls.md`
|
||||
|
||||
Reference in New Issue
Block a user