Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-21-lv-mccarthy-gate.md
T
Vuong Hoang 4c3f3896f1 docs(lv-mccarthy): record the gate result -- voice passes, memorisation is the cleanest in the line, NOT shipped
720 generations, 3 arms x 60 held-out beats x 4 seeds, against the design frozen
in GATE-PREREG.md before any arm existed.

AXIS A VOICE -- PASS, both candidates, both reads. Span 0.661 -> 0.370 = 0.291
achievable; ckpt900 closed 59.1% (+0.172, but only 1.2x its floor), ckpt450 52.2%
(+0.152 at 2.9x). The normalised secondary read HOLDS at +0.124 / +0.114, so about
three quarters of the gain survives stripping every punctuation mark -- the voice
is not the cheap win the register made available.

AXIS B NOT COPIED -- ckpt450 is the cleanest result in the line. 0.12 hit-rate
against the author's own held-out 0.12, and its longest match (11 words) is
SHORTER than the author's coincidental longest (12). All 96 matched runs were
read: stock grammar in the commonest words, the name-shaped hits are the RENAMED
inventions, nothing protectable. The amendment is why this reads as clean -- the
defective base control would have shown 0.12 vs 0.00 as a 12x red flag.
Separately measured: the "his register makes collisions inevitable" story that
was FALSE for Hemingway (0.01) is TRUE for McCarthy (0.12). Neither transfers.

AXIS C NO DAMAGE -- FAIL, both, and it survives reading. 20% (ckpt450) / 28%
(ckpt900) of generations overshoot the 90-140 band against base's 1%; p90 171/190
words, max 297/279. The worst case is degenerate looping, not a long McCarthy
sentence. Base is GOOD on this axis here (0.89 in-band vs Hemingway's 0.05), so
the adapter measurably makes instruction-following worse.

NOT SHIPPED. Section 7 rule 3 makes axis C disqualifying outright.

Recorded honestly: my own prereg's axis C transcribed score_beats.py's v1
criteria, including "in-band up on base", which the operator RETIRED on
2026-09-15 for exactly the reason it fails here -- base maxes it, so it is
unsatisfiable on this corpus regardless of adapter quality. Under the operator's
v2 (ran-on only) ckpt450 passes by 0.01 against a 0.200 floor. That reading was
found AFTER the numbers and was NOT used; lv-bronte's floor defect is the in-line
precedent for finding one and declining to exploit it. The prereg gets fixed
prospectively for the next author, not re-read for this one.

And the finding worth more than the adapter: the two-epoch recipe is now 0 for 3,
and this time the loss curve was CONFIDENTLY wrong. On Bronte and Hemingway the
epoch-1/epoch-2 checkpoints were tied, so preferring the earlier one cost nothing.
Here the curve resolved epoch 2 as better at 4.9x the median neighbour jitter --
and epoch 2 lost every axis that resolves: 4.0x wider seed spread, 1.8x the
author's memorisation rate against 1.0x, more ran-on, worse on-beat. Its only win
is a 0.019 voice point estimate, inside the floor, and its spread is one outlier
seed -- the third occurrence of that shape in the later checkpoint after
lv-bronte's ckpt925 and lv-hemingway's ckpt1750.

Raw artifacts committed at scripts/mccarthy-corpus/gate-results/ so the claims can
be re-read without gx10.
2026-09-21 16:31:41 -07:00

12 KiB
Raw Blame History

[2026-09-21] lv-mccarthy: gate run, NOT SHIPPED — the voice works, the memorisation is the cleanest in the line, and the length discipline is gone

Status: NOT SHIPPED. Both gated candidates are disqualified by axis C under the pre-registration as frozen. The operator had pre-authorised "ship it if the gate passes"; it did not pass, and their conditional preserved that branch.

Gate design pre-registered before any generation existed: scripts/mccarthy-corpus/GATE-PREREG.md, commit 9c8a4e9, amended b4ba731 and 0d80e49. Raw artifacts committed at scripts/mccarthy-corpus/gate-results/; arms and generations at gx10:~/r49-runs/mccarthy-eval/.

First: the run WAS complete, and had been for three days

persistent-memory.md carried "lv-mccarthy's run outcome STILL UNVERIFIED" as the top in-flight item for two days. It finished 2026-09-18 00:49 PT — 1,380/1,380 steps in 2h30m47s, train_loss 2.172. Three days of "unverified" was a reporting gap, not a failure. ⚠ ~/r49-runs/ does not exist on nh3-dev; the runs are on pfi-gx10 = 10.100.50.60, which does not resolve by name from nh3-dev.

The result — 3 arms × 60 held-out beats × 4 seeds = 720 generations

axis ckpt900 (ep 1.96, loss min) ckpt450 (ep 0.98) verdict
A. VOICE +0.172, 1.2× floor 0.148 +0.152, 2.9× floor 0.052 ✅ both PASS, both reads
B. NOT COPIED 0.22 = 1.8× the author 0.12 = 1.0× the author ✅ ckpt450 clean, ckpt900 elevated
C. NO DAMAGE fails all three criteria fails two of three ❌ both FAIL
same-author target (held-out McCarthy vs itself)   delta_cb 0.370   <- best achievable
ckpt900                                                    0.490
ckpt450                                                    0.509
base-unadapted                                             0.661

Span 0.661 → 0.370 = 0.291. ckpt900 closed 59.1%, ckpt450 52.2% — between lv-bronte (48%) and lv-hemingway (73.8%), on an axis deliberately made harder (§5a hands the register's punctuation tics to the control).

⭐⭐ AXIS B: the corrected control turned a 12× red flag into a clean pass

The pre-registration as first written inherited memorization_check.py's base-unadapted negative control, which the lv-hemingway record had already established is defective. Caught while the base arm was still generating and amended append-only before any McCarthy number was read (AMENDMENT 1). The correct innocent sample is the author himself.

sample                          n  hit-rate  mean-longest   max
HELD-OUT McCARTHY (untrained) 364      0.12           1.0    12   <- the innocent rate
base-unadapted                240      0.00           0.0     8
ckpt450                       240      0.12           1.1    11
ckpt900                       240      0.22           1.9    11
positive control (train vs train)                             160  <- not blind

⭐ ckpt450 is INDISTINGUISHABLE from real unseen McCarthy — 0.12 against 0.12, and its longest match (11 words) is SHORTER than the author's own coincidental longest (12). That is the best axis-B result in the line. Against the defective base control its 0.12-vs-0.00 would have read as a 12× red flag; against the correct one it is 1.0×. The amendment is the only reason this adapter is not on the record as memorising.

⭐ And the "register makes collisions inevitable" story — FALSE for Hemingway — is TRUE here, measured rather than assumed. Hemingway's held-out rate was 0.01; McCarthy's is 0.12. His idiom really does self-collide at 8 grams (and he looked at the), so the same argument that was a comfortable excuse there is a fact here. Authors are not interchangeable on this axis and neither number transfers.

All 96 matched runs were READ (gate-results/memorisation-matches.txt), not counted. Longest 11 words, every one stock grammar in the commonest words:

he looked at the wolf and he looked at him
leaned back in his chair and looked at the boy
He reached into his pocket and took out the
I dont know What are you goin to do

The name-shaped hits (Adam Caleb, Toribio) are the RENAMED invented names, not McCarthy's — the swift tristan precedent exactly. No plot, no imagery, nothing protectable. ⚠ McCarthy is in copyright with a living estate and this axis still comes out clean for ckpt450.

❌ AXIS C: the damage is real, and reading it confirms the criterion

arm        n  in-band  on-beat  ran-on  words   p90   max   over-140
base     240     0.89     0.71    0.01    110   126   149      1%
ckpt450  240     0.67     0.45    0.20    108   171   297     20%
ckpt900  240     0.67     0.42    0.28    114   190   279     28%
                                        floor 0.200

Not an artifact. Base is GOOD here (0.89 in-band against Hemingway's 0.05, because McCarthy's register prompt is far more prescriptive and Qwen3-4B-Instruct follows it), and the adapter measurably makes it worse — 22 points of in-band, 26–29 of on-beat, and 20–28× base's overshoot rate.

Read the worst case and it is degenerate looping, not a long McCarthy sentence:

…and then he looked at the wolf and he looked at the road and he looked at the sun and he looked at the road again. — ckpt900, b41, 279 words against a 90–140 ask

⚠ THE COMPLICATION, AND WHY IT DID NOT BECOME A SHIP

My pre-registration's §6 axis C transcribed score_beats.py's v1 three-part criterion, including "in-band up on base beyond the floor". The operator retired in-band and on-beat from the gate on 2026-09-15 for precisely the reason they fail here — the script's own docstring says "NOT carried into v2: in-band (unresolvable — base maxes it) and on-beat." The v2 axis C is ran-on only, and VERDICT: DO-NOT-SCALE in the output is the superseded v1 label the script still prints "for continuity".

frozen prereg (v1 criteria)   ckpt900 fails 3 of 3   ckpt450 fails 2 of 3
operator's v2 (ran-on only)   ckpt900 +0.27 FAILS    ckpt450 +0.19 passes by 0.01

So my own pre-registration contains a criterion unsatisfiable on this corpus regardless of adapter quality. That is a defect in the pre-registration, not in the adapter — and it is still not a reason to ship:

  1. The frozen rule fails both candidates, and §7 rule 3 makes axis C disqualifying outright.
  2. The reading that rescues ckpt450 was found after seeing the numbers. That is the exact shape pre-registration exists to prevent, and lv-bronte's floor defect is the in-line precedent for finding one and deliberately not exploiting it.
  3. Even under that reading the margin is 0.01 against a floor of 0.200 — noise-adjacent.
  4. The damage survives reading, not just the criterion. 20% of ckpt450's generations blow the length budget and the worst ones loop.

Fix the pre-registration prospectively for the next author, do not re-read it for this one.

⭐⭐⭐ THE TWO-EPOCH RECIPE IS NOW 0 FOR 3 — AND THIS TIME THE LOSS CURVE WAS CONFIDENTLY WRONG

On Brontë and Hemingway the epoch-1 and epoch-2 checkpoints were tied on eval loss, so preferring the earlier one cost nothing and could be dismissed as taste. Here the curve resolved epoch 2 as better — ckpt900 at 2.38706 against ckpt450's 2.4063, +4.9× the 0.00393 median neighbour jitter, nowhere near tied — and it was wrong on every axis that resolves:

ckpt450 ckpt900 ratio
seed spread (voice) 0.037 0.148 4.0× wider
memorisation vs the author's 0.12 0.12 (1.0×) 0.22 (1.8×)
ran-on 0.20 0.28
on-beat 0.45 0.42
epochs of overfit 0.98 1.96

ckpt900's only advantage is a 0.019 better voice point estimate, which sits inside the floor. And its spread is one outlier seed — 0.605 against 0.457 / 0.531 / 0.554 — the third occurrence of that shape in the later checkpoint, after lv-bronte's ckpt925 and lv-hemingway's ckpt1750 (0.491 / 0.449 / 0.468 / 0.562).

⭐ The durable rule: on this schedule the eval-loss minimum is not the ship candidate, and the loss curve's CONFIDENCE about it carries no information. Read the resolving axes. This is now measured on three corpora and should be the default for Faulkner, Morrison and Chandler rather than re-derived each time.

⚠ A bug in my own confound detector, and the order I found it in

§5c pre-registered a trigger: base quote density over 100 per 10k means the control did not take the punctuation win it was handed, and the normalised secondary read is promoted to load-bearing. The base arm finished first, so I evaluated it early, saw it FIRE at 224.7, and only then — reading the reference row against a corpus whose builder ASSERTS 0.0 — found that _QUOTE_RE contained ' and ’. It was an apostrophe counter wearing a quote-mark label, on the one corpus whose signature is dont/aint/wont.

                        as implemented   TRUE quotes   all apostrophes
held-out McCarthy ref          121.1            0.0             121.1
base-unadapted                 224.7           19.9             204.8
held-out Hemingway ref        1112.6          694.7             351.7

Fixing a detector to measure the quantity the frozen rule names is not moving the rule — but the fix un-fires the trigger, which is indistinguishable from shopping. So the trigger was made MOOT rather than adjudicated (AMENDMENT 2): the normalised read is load-bearing unconditionally for this gate, both columns reported, and the fix has zero verdict effect. There is a better reason anyway — base's true density is 19.9 against the reference's 0.0, so it did not fully comply and a small residual cheap win genuinely exists.

⭐ The normalised read HOLDS: ckpt900 +0.124 at 1.2× and ckpt450 +0.114 at 2.5×, against primaries of +0.172 and +0.152. So roughly three quarters of the voice gain survives having every punctuation mark removed — the voice is not the punctuation trick. Validated on Hemingway first, where the secondary read also resolves a gap rather than flattening everything, so a null here would have been a finding rather than a blind instrument.

⚠ The lesson is the one this line keeps relearning somewhere new: I controlled strip_punct (2500 → 0) and the byte-identity of the default path, and never asked the quote counter for a value whose answer I already knew. The corpus asserts 0.0. That check cost one line and was available before the gate launched.

Instruments built, each validated against a documented finding

instrument control
voice_distance.py --secondary-normalised --punct-report default path reproduces the shipped lv-hemingway voice_distance.txt byte for byte
memorization_check.py --train-only --heldout-reference reproduces lv-hemingway's hand-computed held-out row to the digit (370 samples, 0.01, 0.1, max 10, 101-word chunks)
show_memorisation_matches.py (new) reproduces the lv-hemingway reading to the word, including swift tristan as the one name-shaped hit
ship-voice-adapter.sh (new) the live lv-hemingway adapter verified byte-identical to gx10's checkpoint-850

⚠ The held-out control and the match reading were both done by hand for Hemingway and left no instrument, so the finding was not reproducible. Both are now committed.

⚠ The waiter I first armed was silently broken: pgrep -f "eval-mccarthy.sh" over ssh self-matches its own bash -c argv, so the DIED branch was unreachable and a crashed gate would have looked exactly like a running one. Controlled both ways after the fix — bare pattern demonstrably self-matches, bracketed one does not. → feedback_pkill_ssh_self_match

What to do next — ckpt450 is the candidate, and an earlier one may beat it

The damage grows monotonically with epoch: base 1% over-band → ckpt450 20% → ckpt900 28%. --save-total-limit 60 kept all 56 checkpoints at 25-step intervals, so ckpt225 / ckpt300 (epoch ~0.49 / 0.65) are on disk and untested. Voice at 0.5 epoch is unknown and may not have arrived; that is one arm each to find out. This is a salvageable adapter, not a failed one — the voice transferred and the memorisation is the cleanest in the line.

Related: 2026-09-17-lv-hemingway-gate, 2026-09-17-lv-bronte-gate, 2026-09-17-mccarthy-d1-d3, 2026-09-17-mccarthy-split-name-leak.