Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-21-lv-mccarthy-gate.md
T
Vuong Hoang e8eb1594d9 docs(lv-mccarthy): five-arm ladder -- ckpt300 wins every axis, held at the gate by 0.02
1,200 generations across five arms. ckpt300 (epoch 0.652) is the best arm in the
run on every axis that resolves:

  VOICE          +0.177 at 3.2x its pairwise floor (primary), +0.128 at 2.8x with
                 every punctuation mark stripped. Best point estimate AND best
                 margin of any arm, spread 0.055/0.038 with no outlier seed.
  MEMORISATION   0.12 against real unseen McCarthy's own 0.12 -- identical -- with
                 a longest match of 10 words against the author's coincidental 12.
                 All 31 matches read: stock grammar, names are the renamed
                 inventions, nothing protectable.
  DAMAGE         ran-on +0.12. Clears the operator's ratified v2 floor of 0.200 by
                 40%. FAILS AMENDMENT 3's self-imposed 0.100 bar by 0.02.

NOT SHIPPED, and the reason is the bar rather than the adapter. AMENDMENT 3 fixed
ran-on <= 0.100 before either new arm existed, precisely so a marginal number could
not be talked into a ship, and shipping at 0.12 would make that pre-registration
theatre. But the bar's stated rationale was written against ckpt450's pass by 0.01
-- 5% of the threshold -- and ckpt300 clears by 40%. The number excludes a candidate
the reasoning does not. That is an operator call.

ckpt325/350/375 are on disk and one may sit under 0.100. They were deliberately NOT
gated: searching the checkpoint space until something clears is candidate-shopping,
the same family as threshold-shopping approached from the other side.

THREE CLAIMS FROM EARLIER THIS SESSION ARE REFUTED and are corrected in the record:

  1. "The damage is flat across epochs and only rotates direction" -- FALSE. ran-on
     is non-monotonic (0.38 -> 0.13 -> 0.20 -> 0.28 across epochs 0.49/0.65/0.98/
     1.96) with a real minimum near 0.65, and ckpt225 is 48% out-of-band against
     ckpt300's 35%.
  2. "ckpt300 runs far too short, ckpt225 will clear ran-on by being short" -- FALSE
     on both. ckpt225 runs LONG (median 127, 38% over-band) and is the worst arm in
     the run. I generalised from SIX generations of one arm, which is the exact n=1
     violation the measurement-discipline rule names, committed in the same breath
     as a note about being careful.
  3. The original "gate an earlier checkpoint, the overshoot may not have arrived
     yet" recommendation was RIGHT. Retracting it an hour later on a three-arm read
     was the error, not the recommendation.

What is true and unresolved by any checkpoint choice: 35% of ckpt300's generations
miss the 90-140 band against base's 11%, and in-band is 0.65 against 0.89. An
adapter that buys a voice and costs a third of the length compliance is a trade, not
a defect -- but it is the operator's trade to accept.

Raw artifacts for all five arms at scripts/mccarthy-corpus/gate-results/.
2026-09-21 17:53:43 -07:00

18 KiB
Raw Blame History

[2026-09-21] lv-mccarthy: gate run, NOT SHIPPED — the voice works, the memorisation is the cleanest in the line, and the length discipline is gone

Status: NOT SHIPPED. Both gated candidates are disqualified by axis C under the pre-registration as frozen. The operator had pre-authorised "ship it if the gate passes"; it did not pass, and their conditional preserved that branch.

Gate design pre-registered before any generation existed: scripts/mccarthy-corpus/GATE-PREREG.md, commit 9c8a4e9, amended b4ba731 and 0d80e49. Raw artifacts committed at scripts/mccarthy-corpus/gate-results/; arms and generations at gx10:~/r49-runs/mccarthy-eval/.

First: the run WAS complete, and had been for three days

persistent-memory.md carried "lv-mccarthy's run outcome STILL UNVERIFIED" as the top in-flight item for two days. It finished 2026-09-18 00:49 PT — 1,380/1,380 steps in 2h30m47s, train_loss 2.172. Three days of "unverified" was a reporting gap, not a failure. ⚠ ~/r49-runs/ does not exist on nh3-dev; the runs are on pfi-gx10 = 10.100.50.60, which does not resolve by name from nh3-dev.

The result — 3 arms × 60 held-out beats × 4 seeds = 720 generations

axis ckpt900 (ep 1.96, loss min) ckpt450 (ep 0.98) verdict
A. VOICE +0.172, 1.2× floor 0.148 +0.152, 2.9× floor 0.052 ✅ both PASS, both reads
B. NOT COPIED 0.22 = 1.8× the author 0.12 = 1.0× the author ✅ ckpt450 clean, ckpt900 elevated
C. NO DAMAGE fails all three criteria fails two of three ❌ both FAIL
same-author target (held-out McCarthy vs itself)   delta_cb 0.370   <- best achievable
ckpt900                                                    0.490
ckpt450                                                    0.509
base-unadapted                                             0.661

Span 0.661 → 0.370 = 0.291. ckpt900 closed 59.1%, ckpt450 52.2% — between lv-bronte (48%) and lv-hemingway (73.8%), on an axis deliberately made harder (§5a hands the register's punctuation tics to the control).

⭐⭐ AXIS B: the corrected control turned a 12× red flag into a clean pass

The pre-registration as first written inherited memorization_check.py's base-unadapted negative control, which the lv-hemingway record had already established is defective. Caught while the base arm was still generating and amended append-only before any McCarthy number was read (AMENDMENT 1). The correct innocent sample is the author himself.

sample                          n  hit-rate  mean-longest   max
HELD-OUT McCARTHY (untrained) 364      0.12           1.0    12   <- the innocent rate
base-unadapted                240      0.00           0.0     8
ckpt450                       240      0.12           1.1    11
ckpt900                       240      0.22           1.9    11
positive control (train vs train)                             160  <- not blind

⭐ ckpt450 is INDISTINGUISHABLE from real unseen McCarthy — 0.12 against 0.12, and its longest match (11 words) is SHORTER than the author's own coincidental longest (12). That is the best axis-B result in the line. Against the defective base control its 0.12-vs-0.00 would have read as a 12× red flag; against the correct one it is 1.0×. The amendment is the only reason this adapter is not on the record as memorising.

⭐ And the "register makes collisions inevitable" story — FALSE for Hemingway — is TRUE here, measured rather than assumed. Hemingway's held-out rate was 0.01; McCarthy's is 0.12. His idiom really does self-collide at 8 grams (and he looked at the), so the same argument that was a comfortable excuse there is a fact here. Authors are not interchangeable on this axis and neither number transfers.

All 96 matched runs were READ (gate-results/memorisation-matches.txt), not counted. Longest 11 words, every one stock grammar in the commonest words:

he looked at the wolf and he looked at him
leaned back in his chair and looked at the boy
He reached into his pocket and took out the
I dont know What are you goin to do

The name-shaped hits (Adam Caleb, Toribio) are the RENAMED invented names, not McCarthy's — the swift tristan precedent exactly. No plot, no imagery, nothing protectable. ⚠ McCarthy is in copyright with a living estate and this axis still comes out clean for ckpt450.

❌ AXIS C: the damage is real, and reading it confirms the criterion

arm        n  in-band  on-beat  ran-on  words   p90   max   over-140
base     240     0.89     0.71    0.01    110   126   149      1%
ckpt450  240     0.67     0.45    0.20    108   171   297     20%
ckpt900  240     0.67     0.42    0.28    114   190   279     28%
                                        floor 0.200

Not an artifact. Base is GOOD here (0.89 in-band against Hemingway's 0.05, because McCarthy's register prompt is far more prescriptive and Qwen3-4B-Instruct follows it), and the adapter measurably makes it worse — 22 points of in-band, 26–29 of on-beat, and 20–28× base's overshoot rate.

Read the worst case and it is degenerate looping, not a long McCarthy sentence:

…and then he looked at the wolf and he looked at the road and he looked at the sun and he looked at the road again. — ckpt900, b41, 279 words against a 90–140 ask

⚠ THE COMPLICATION, AND WHY IT DID NOT BECOME A SHIP

My pre-registration's §6 axis C transcribed score_beats.py's v1 three-part criterion, including "in-band up on base beyond the floor". The operator retired in-band and on-beat from the gate on 2026-09-15 for precisely the reason they fail here — the script's own docstring says "NOT carried into v2: in-band (unresolvable — base maxes it) and on-beat." The v2 axis C is ran-on only, and VERDICT: DO-NOT-SCALE in the output is the superseded v1 label the script still prints "for continuity".

frozen prereg (v1 criteria)   ckpt900 fails 3 of 3   ckpt450 fails 2 of 3
operator's v2 (ran-on only)   ckpt900 +0.27 FAILS    ckpt450 +0.19 passes by 0.01

So my own pre-registration contains a criterion unsatisfiable on this corpus regardless of adapter quality. That is a defect in the pre-registration, not in the adapter — and it is still not a reason to ship:

  1. The frozen rule fails both candidates, and §7 rule 3 makes axis C disqualifying outright.
  2. The reading that rescues ckpt450 was found after seeing the numbers. That is the exact shape pre-registration exists to prevent, and lv-bronte's floor defect is the in-line precedent for finding one and deliberately not exploiting it.
  3. Even under that reading the margin is 0.01 against a floor of 0.200 — noise-adjacent.
  4. The damage survives reading, not just the criterion. 20% of ckpt450's generations blow the length budget and the worst ones loop.

Fix the pre-registration prospectively for the next author, do not re-read it for this one.

⭐⭐⭐ THE TWO-EPOCH RECIPE IS NOW 0 FOR 3 — AND THIS TIME THE LOSS CURVE WAS CONFIDENTLY WRONG

On Brontë and Hemingway the epoch-1 and epoch-2 checkpoints were tied on eval loss, so preferring the earlier one cost nothing and could be dismissed as taste. Here the curve resolved epoch 2 as better — ckpt900 at 2.38706 against ckpt450's 2.4063, +4.9× the 0.00393 median neighbour jitter, nowhere near tied — and it was wrong on every axis that resolves:

ckpt450 ckpt900 ratio
seed spread (voice) 0.037 0.148 4.0× wider
memorisation vs the author's 0.12 0.12 (1.0×) 0.22 (1.8×)
ran-on 0.20 0.28
on-beat 0.45 0.42
epochs of overfit 0.98 1.96

ckpt900's only advantage is a 0.019 better voice point estimate, which sits inside the floor. And its spread is one outlier seed — 0.605 against 0.457 / 0.531 / 0.554 — the third occurrence of that shape in the later checkpoint, after lv-bronte's ckpt925 and lv-hemingway's ckpt1750 (0.491 / 0.449 / 0.468 / 0.562).

⭐ The durable rule: on this schedule the eval-loss minimum is not the ship candidate, and the loss curve's CONFIDENCE about it carries no information. Read the resolving axes. This is now measured on three corpora and should be the default for Faulkner, Morrison and Chandler rather than re-derived each time.

⚠ A bug in my own confound detector, and the order I found it in

§5c pre-registered a trigger: base quote density over 100 per 10k means the control did not take the punctuation win it was handed, and the normalised secondary read is promoted to load-bearing. The base arm finished first, so I evaluated it early, saw it FIRE at 224.7, and only then — reading the reference row against a corpus whose builder ASSERTS 0.0 — found that _QUOTE_RE contained ' and ’. It was an apostrophe counter wearing a quote-mark label, on the one corpus whose signature is dont/aint/wont.

                        as implemented   TRUE quotes   all apostrophes
held-out McCarthy ref          121.1            0.0             121.1
base-unadapted                 224.7           19.9             204.8
held-out Hemingway ref        1112.6          694.7             351.7

Fixing a detector to measure the quantity the frozen rule names is not moving the rule — but the fix un-fires the trigger, which is indistinguishable from shopping. So the trigger was made MOOT rather than adjudicated (AMENDMENT 2): the normalised read is load-bearing unconditionally for this gate, both columns reported, and the fix has zero verdict effect. There is a better reason anyway — base's true density is 19.9 against the reference's 0.0, so it did not fully comply and a small residual cheap win genuinely exists.

⭐ The normalised read HOLDS: ckpt900 +0.124 at 1.2× and ckpt450 +0.114 at 2.5×, against primaries of +0.172 and +0.152. So roughly three quarters of the voice gain survives having every punctuation mark removed — the voice is not the punctuation trick. Validated on Hemingway first, where the secondary read also resolves a gap rather than flattening everything, so a null here would have been a finding rather than a blind instrument.

⚠ The lesson is the one this line keeps relearning somewhere new: I controlled strip_punct (2500 → 0) and the byte-identity of the default path, and never asked the quote counter for a value whose answer I already knew. The corpus asserts 0.0. That check cost one line and was available before the gate launched.

Instruments built, each validated against a documented finding

instrument control
voice_distance.py --secondary-normalised --punct-report default path reproduces the shipped lv-hemingway voice_distance.txt byte for byte
memorization_check.py --train-only --heldout-reference reproduces lv-hemingway's hand-computed held-out row to the digit (370 samples, 0.01, 0.1, max 10, 101-word chunks)
show_memorisation_matches.py (new) reproduces the lv-hemingway reading to the word, including swift tristan as the one name-shaped hit
ship-voice-adapter.sh (new) the live lv-hemingway adapter verified byte-identical to gx10's checkpoint-850

⚠ The held-out control and the match reading were both done by hand for Hemingway and left no instrument, so the finding was not reproducible. Both are now committed.

⚠ The waiter I first armed was silently broken: pgrep -f "eval-mccarthy.sh" over ssh self-matches its own bash -c argv, so the DIED branch was unreachable and a crashed gate would have looked exactly like a running one. Controlled both ways after the fix — bare pattern demonstrably self-matches, bracketed one does not. → feedback_pkill_ssh_self_match

What to do next — ckpt450 is the candidate, and an earlier one may beat it

The damage grows monotonically with epoch: base 1% over-band → ckpt450 20% → ckpt900 28%. --save-total-limit 60 kept all 56 checkpoints at 25-step intervals, so ckpt225 / ckpt300 (epoch ~0.49 / 0.65) are on disk and untested. Voice at 0.5 epoch is unknown and may not have arrived; that is one arm each to find out. This is a salvageable adapter, not a failed one — the voice transferred and the memorisation is the cleanest in the line.

Related: 2026-09-17-lv-hemingway-gate, 2026-09-17-lv-bronte-gate, 2026-09-17-mccarthy-d1-d3, 2026-09-17-mccarthy-split-name-leak.


AMENDMENT-3 ROUND: five arms, and ckpt300 wins every axis — held at the gate by 0.02

Two more arms gated after the frozen axis C was proven unsatisfiable (a perfect adapter fails it by 0.09 — see GATE-PREREG.md AMENDMENT 3). Same fixture, same four seeds, same rule. 1,200 generations total.

The complete ladder

arm epoch VOICE primary VOICE normalised MEMORISATION (author = 0.12 / max 12) ran-on Δ in-band
ckpt300 0.652 +0.177, 3.2× +0.128, 2.8× 0.12 / 1.0 / max 10 +0.12 0.65
ckpt225 0.489 +0.176, 1.8× +0.122, 1.5× 0.12 / 1.0 / max 10 +0.37 0.52
ckpt450 0.980 +0.152, 2.9× +0.114, 2.5× 0.12 / 1.1 / max 11 +0.19 0.67
ckpt900 1.958 +0.172, 1.2× +0.124, 1.2× 0.22 / 1.9 / max 11 +0.27 0.67
base — — — 0.00 / 0.0 / max 8 — 0.89
length distribution, raw generations
arm        n  median  p10  p90  max  in-band  under90  over140  TOTAL out
base     240    110    89  126  149     0.89      10%       1%       11%
ckpt225  240    127    89  206  296     0.52      10%      38%       48%
ckpt300  240    101    83  151  216     0.65      22%      13%       35%
ckpt450  240    107    87  171  297     0.67      14%      20%       33%
ckpt900  240    114    92  190  279     0.67       5%      28%       33%

⭐⭐ ckpt300 is the best arm in the run on every axis that resolves. Best voice point estimate AND best margin on both reads, tightest spread of any adapted arm bar ckpt450 (0.055 / 0.038, no outlier seed), memorisation identical to real unseen McCarthy with a max of 10 against the author's coincidental 12, and the least overshoot of any checkpoint. Its 31 matched runs were all read: looked at the wolf and he looked at the boy, took off his hat and set it on the, and wiped his mouth on the back of his. Name-shaped hits are Adam Caleb / Colton — the renamed inventions. Nothing protectable.

⚠ THREE CLAIMS I MADE EARLIER THAT THE FIVE-ARM DATA REFUTES

  1. "The damage is flat across epochs and only rotates direction." FALSE. ckpt225 is 48% out-of-band against ckpt300's 35%, and ran-on is non-monotonic: 0.38 → 0.13 → 0.20 → 0.28 across epochs 0.49 → 0.65 → 0.98 → 1.96. There is a genuine minimum near epoch 0.65.
  2. "ckpt300 is coming out far too SHORT, and ckpt225 will clear ran-on by being short." FALSE on both. ⚠ I generalised from SIX generations of one arm — the exact n=1 violation the measurement-discipline rule names, committed while writing a note about being careful. ckpt225 runs long (median 127, 38% over-band) and is the worst arm in the run.
  3. My own option A — "gate an earlier checkpoint, the overshoot may not have arrived yet" — was RIGHT, and I retracted it an hour later on a three-arm read. ckpt300 has the best voice and the least damage. The retraction was the error, not the recommendation.

The gate's verdict on ckpt300, stated exactly

1. VOICE          +0.177 at 3.2x floor (primary), +0.128 at 2.8x (normalised)   PASS
2. MEMORISATION   0.12 vs the author's own 0.12, max 10 vs the author's 12      PASS
3. DAMAGE         ran-on delta +0.12
                    vs the operator's ratified v2 floor 0.200   -> PASS, 40% headroom
                    vs AMENDMENT 3's self-imposed bar   0.100   -> FAIL by 0.02

NOT SHIPPED, and the reason is the bar rather than the adapter. AMENDMENT 3 fixed ran-on ≤ 0.100 before either new arm existed, specifically so a marginal number could not be talked into a ship. ckpt300 is 0.12. Shipping it on my own authority would make the pre-registration theatre.

⚠ But the bar's stated RATIONALE does not describe this candidate. It was written against ckpt450's +0.19 vs 0.200 — a pass by 0.01, i.e. 5% of the threshold — and the sentence justifying it says exactly that. ckpt300 clears the operator's threshold by 40%. The bar's number excludes a candidate its reasoning does not. That is an operator call, not mine.

⚠⚠ AND I DID NOT GO LOOKING FOR A CHECKPOINT UNDER 0.100. ckpt325/350/375 are all on disk and one of them may well sit below the bar. Searching the checkpoint space until something clears is candidate-shopping — the same family as threshold-shopping, arrived at from the other side. The measuring stopped here deliberately.

What is actually true about this adapter

The voice transferred, ~3/4 of the gain survives stripping every punctuation mark, and the memorisation is indistinguishable from the author's own self-collision rate on a corpus in copyright with a living estate. The cost is real and unresolved by any checkpoint choice: 35% of generations miss the 90–140 band against base's 11%, and in-band is 0.65 against 0.89. An adapter that buys a voice and costs a third of your length compliance is a trade, not a defect — but it is the operator's trade to accept.

Recommendation: ship ckpt300. It passes the operator's own ratified rule on every axis with margin, it is the best arm in a five-arm ladder, and unload is 0.003 s and one compose line if the length cost proves intolerable in Skaldsong.