Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-21-lv-mccarthy-gate.md
T
Vuong Hoang e8eb1594d9 docs(lv-mccarthy): five-arm ladder -- ckpt300 wins every axis, held at the gate by 0.02
1,200 generations across five arms. ckpt300 (epoch 0.652) is the best arm in the
run on every axis that resolves:

  VOICE          +0.177 at 3.2x its pairwise floor (primary), +0.128 at 2.8x with
                 every punctuation mark stripped. Best point estimate AND best
                 margin of any arm, spread 0.055/0.038 with no outlier seed.
  MEMORISATION   0.12 against real unseen McCarthy's own 0.12 -- identical -- with
                 a longest match of 10 words against the author's coincidental 12.
                 All 31 matches read: stock grammar, names are the renamed
                 inventions, nothing protectable.
  DAMAGE         ran-on +0.12. Clears the operator's ratified v2 floor of 0.200 by
                 40%. FAILS AMENDMENT 3's self-imposed 0.100 bar by 0.02.

NOT SHIPPED, and the reason is the bar rather than the adapter. AMENDMENT 3 fixed
ran-on <= 0.100 before either new arm existed, precisely so a marginal number could
not be talked into a ship, and shipping at 0.12 would make that pre-registration
theatre. But the bar's stated rationale was written against ckpt450's pass by 0.01
-- 5% of the threshold -- and ckpt300 clears by 40%. The number excludes a candidate
the reasoning does not. That is an operator call.

ckpt325/350/375 are on disk and one may sit under 0.100. They were deliberately NOT
gated: searching the checkpoint space until something clears is candidate-shopping,
the same family as threshold-shopping approached from the other side.

THREE CLAIMS FROM EARLIER THIS SESSION ARE REFUTED and are corrected in the record:

  1. "The damage is flat across epochs and only rotates direction" -- FALSE. ran-on
     is non-monotonic (0.38 -> 0.13 -> 0.20 -> 0.28 across epochs 0.49/0.65/0.98/
     1.96) with a real minimum near 0.65, and ckpt225 is 48% out-of-band against
     ckpt300's 35%.
  2. "ckpt300 runs far too short, ckpt225 will clear ran-on by being short" -- FALSE
     on both. ckpt225 runs LONG (median 127, 38% over-band) and is the worst arm in
     the run. I generalised from SIX generations of one arm, which is the exact n=1
     violation the measurement-discipline rule names, committed in the same breath
     as a note about being careful.
  3. The original "gate an earlier checkpoint, the overshoot may not have arrived
     yet" recommendation was RIGHT. Retracting it an hour later on a three-arm read
     was the error, not the recommendation.

What is true and unresolved by any checkpoint choice: 35% of ckpt300's generations
miss the 90-140 band against base's 11%, and in-band is 0.65 against 0.89. An
adapter that buys a voice and costs a third of the length compliance is a trade, not
a defect -- but it is the operator's trade to accept.

Raw artifacts for all five arms at scripts/mccarthy-corpus/gate-results/.
2026-09-21 17:53:43 -07:00

305 lines
18 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# `[2026-09-21]` lv-mccarthy: gate run, NOT SHIPPED — the voice works, the memorisation is the cleanest in the line, and the length discipline is gone
**Status: NOT SHIPPED.** Both gated candidates are disqualified by axis C under the
pre-registration as frozen. The operator had pre-authorised "ship it if the gate passes";
it did not pass, and their conditional preserved that branch.
Gate design pre-registered before any generation existed:
`scripts/mccarthy-corpus/GATE-PREREG.md`, commit `9c8a4e9`, amended `b4ba731` and `0d80e49`.
Raw artifacts committed at `scripts/mccarthy-corpus/gate-results/`; arms and generations at
`gx10:~/r49-runs/mccarthy-eval/`.
## First: the run WAS complete, and had been for three days
`persistent-memory.md` carried "lv-mccarthy's run outcome STILL UNVERIFIED" as the top
in-flight item for two days. It finished **2026-09-18 00:49 PT** — 1,380/1,380 steps in
2h30m47s, `train_loss` 2.172. Three days of "unverified" was a reporting gap, not a failure.
⚠ `~/r49-runs/` does not exist on nh3-dev; the runs are on **pfi-gx10 = 10.100.50.60**, which
does not resolve by name from nh3-dev.
## The result — 3 arms × 60 held-out beats × 4 seeds = 720 generations
| axis | ckpt900 (ep 1.96, loss min) | ckpt450 (ep 0.98) | verdict |
|---|---|---|---|
| **A. VOICE** | +0.172, **1.2×** floor 0.148 | +0.152, **2.9×** floor 0.052 | ✅ both PASS, both reads |
| **B. NOT COPIED** | 0.22 = 1.8× the author | **0.12 = 1.0× the author** | ✅ ckpt450 clean, ckpt900 elevated |
| **C. NO DAMAGE** | fails all three criteria | fails two of three | ❌ **both FAIL** |
```
same-author target (held-out McCarthy vs itself) delta_cb 0.370 <- best achievable
ckpt900 0.490
ckpt450 0.509
base-unadapted 0.661
```
Span 0.661 → 0.370 = 0.291. ckpt900 closed **59.1%**, ckpt450 **52.2%** — between lv-bronte
(48%) and lv-hemingway (73.8%), on an axis deliberately made harder (§5a hands the register's
punctuation tics to the control).
## ⭐⭐ AXIS B: the corrected control turned a 12× red flag into a clean pass
The pre-registration as first written inherited `memorization_check.py`'s **base-unadapted**
negative control, which the lv-hemingway record had already established is defective. Caught
while the base arm was still generating and amended append-only before any McCarthy number
was read (AMENDMENT 1). The correct innocent sample is the author himself.
```
sample n hit-rate mean-longest max
HELD-OUT McCARTHY (untrained) 364 0.12 1.0 12 <- the innocent rate
base-unadapted 240 0.00 0.0 8
ckpt450 240 0.12 1.1 11
ckpt900 240 0.22 1.9 11
positive control (train vs train) 160 <- not blind
```
⭐ **ckpt450 is INDISTINGUISHABLE from real unseen McCarthy** — 0.12 against 0.12, and its
longest match (11 words) is **SHORTER than the author's own coincidental longest (12)**. That
is the best axis-B result in the line. Against the defective base control its 0.12-vs-0.00
would have read as a **12× red flag**; against the correct one it is **1.0×**. The amendment
is the only reason this adapter is not on the record as memorising.
⭐ **And the "register makes collisions inevitable" story — FALSE for Hemingway — is TRUE
here, measured rather than assumed.** Hemingway's held-out rate was 0.01; McCarthy's is 0.12.
His idiom really does self-collide at 8 grams (`and he looked at the`), so the same argument
that was a comfortable excuse there is a fact here. Authors are not interchangeable on this
axis and neither number transfers.
**All 96 matched runs were READ** (`gate-results/memorisation-matches.txt`), not counted.
Longest 11 words, every one stock grammar in the commonest words:
```
he looked at the wolf and he looked at him
leaned back in his chair and looked at the boy
He reached into his pocket and took out the
I dont know What are you goin to do
```
The name-shaped hits (`Adam Caleb`, `Toribio`) are the **RENAMED invented names**, not
McCarthy's — the `swift tristan` precedent exactly. No plot, no imagery, nothing protectable.
⚠ McCarthy is in copyright with a living estate and this axis still comes out clean for
ckpt450.
## ❌ AXIS C: the damage is real, and reading it confirms the criterion
```
arm n in-band on-beat ran-on words p90 max over-140
base 240 0.89 0.71 0.01 110 126 149 1%
ckpt450 240 0.67 0.45 0.20 108 171 297 20%
ckpt900 240 0.67 0.42 0.28 114 190 279 28%
floor 0.200
```
Not an artifact. **Base is GOOD here** (0.89 in-band against Hemingway's 0.05, because
McCarthy's register prompt is far more prescriptive and Qwen3-4B-Instruct follows it), and the
adapter measurably makes it worse — 22 points of in-band, 26–29 of on-beat, and 20–28× base's
overshoot rate.
Read the worst case and it is **degenerate looping**, not a long McCarthy sentence:
> *…and then he looked at the wolf and he looked at the road and he looked at the sun and he
> looked at the road again.* — ckpt900, b41, 279 words against a 90–140 ask
## ⚠ THE COMPLICATION, AND WHY IT DID NOT BECOME A SHIP
My pre-registration's §6 axis C transcribed `score_beats.py`'s **v1** three-part criterion,
including "in-band up on base beyond the floor". The operator **retired** in-band and on-beat
from the gate on 2026-09-15 for precisely the reason they fail here — the script's own
docstring says *"NOT carried into v2: in-band (unresolvable — base maxes it) and on-beat."*
The v2 axis C is **ran-on only**, and `VERDICT: DO-NOT-SCALE` in the output is the superseded
v1 label the script still prints "for continuity".
```
frozen prereg (v1 criteria) ckpt900 fails 3 of 3 ckpt450 fails 2 of 3
operator's v2 (ran-on only) ckpt900 +0.27 FAILS ckpt450 +0.19 passes by 0.01
```
So my own pre-registration contains a criterion **unsatisfiable on this corpus regardless of
adapter quality**. That is a defect in the pre-registration, not in the adapter — and it is
still not a reason to ship:
1. The frozen rule fails both candidates, and §7 rule 3 makes axis C disqualifying outright.
2. The reading that rescues ckpt450 was found **after** seeing the numbers. That is the exact
shape pre-registration exists to prevent, and lv-bronte's floor defect is the in-line
precedent for finding one and deliberately not exploiting it.
3. Even under that reading the margin is **0.01 against a floor of 0.200** — noise-adjacent.
4. The damage survives reading, not just the criterion. 20% of ckpt450's generations blow the
length budget and the worst ones loop.
**Fix the pre-registration prospectively for the next author, do not re-read it for this one.**
## ⭐⭐⭐ THE TWO-EPOCH RECIPE IS NOW 0 FOR 3 — AND THIS TIME THE LOSS CURVE WAS CONFIDENTLY WRONG
On Brontë and Hemingway the epoch-1 and epoch-2 checkpoints were **tied** on eval loss, so
preferring the earlier one cost nothing and could be dismissed as taste. Here the curve
**resolved** epoch 2 as better — ckpt900 at 2.38706 against ckpt450's 2.4063, **+4.9× the
0.00393 median neighbour jitter**, nowhere near tied — and it was wrong on every axis that
resolves:
| | ckpt450 | ckpt900 | ratio |
|---|---|---|---|
| seed spread (voice) | **0.037** | 0.148 | **4.0× wider** |
| memorisation vs the author's 0.12 | **0.12 (1.0×)** | 0.22 (1.8×) | |
| ran-on | **0.20** | 0.28 | |
| on-beat | **0.45** | 0.42 | |
| epochs of overfit | **0.98** | 1.96 | |
ckpt900's only advantage is a 0.019 better voice point estimate, which sits inside the floor.
And its spread is **one outlier seed** — 0.605 against 0.457 / 0.531 / 0.554 — the third
occurrence of that shape in the later checkpoint, after lv-bronte's ckpt925 and
lv-hemingway's ckpt1750 (0.491 / 0.449 / 0.468 / **0.562**).
⭐ **The durable rule: on this schedule the eval-loss minimum is not the ship candidate, and
the loss curve's CONFIDENCE about it carries no information.** Read the resolving axes. This
is now measured on three corpora and should be the default for Faulkner, Morrison and Chandler
rather than re-derived each time.
## ⚠ A bug in my own confound detector, and the order I found it in
§5c pre-registered a trigger: base quote density over **100 per 10k** means the control did
not take the punctuation win it was handed, and the normalised secondary read is promoted to
load-bearing. The base arm finished first, so I evaluated it early, **saw it FIRE at 224.7**,
and only then — reading the reference row against a corpus whose builder ASSERTS 0.0 — found
that `_QUOTE_RE` contained `'` and `’`. It was an apostrophe counter wearing a quote-mark
label, on the one corpus whose signature is `dont`/`aint`/`wont`.
```
as implemented TRUE quotes all apostrophes
held-out McCarthy ref 121.1 0.0 121.1
base-unadapted 224.7 19.9 204.8
held-out Hemingway ref 1112.6 694.7 351.7
```
Fixing a detector to measure the quantity the frozen rule names is not moving the rule — but
the fix un-fires the trigger, which is indistinguishable from shopping. **So the trigger was
made MOOT rather than adjudicated** (AMENDMENT 2): the normalised read is load-bearing
unconditionally for this gate, both columns reported, and the fix has zero verdict effect.
There is a better reason anyway — base's true density is 19.9 against the reference's 0.0, so
it did not fully comply and a small residual cheap win genuinely exists.
⭐ The normalised read **HOLDS**: ckpt900 +0.124 at 1.2× and ckpt450 +0.114 at 2.5×, against
primaries of +0.172 and +0.152. So roughly three quarters of the voice gain survives having
every punctuation mark removed — **the voice is not the punctuation trick.** Validated on
Hemingway first, where the secondary read also resolves a gap rather than flattening
everything, so a null here would have been a finding rather than a blind instrument.
⚠ **The lesson is the one this line keeps relearning somewhere new:** I controlled
`strip_punct` (2500 → 0) and the byte-identity of the default path, and never asked the quote
counter for a value whose answer I already knew. The corpus asserts 0.0. That check cost one
line and was available before the gate launched.
## Instruments built, each validated against a documented finding
| instrument | control |
|---|---|
| `voice_distance.py --secondary-normalised --punct-report` | default path reproduces the shipped lv-hemingway `voice_distance.txt` **byte for byte** |
| `memorization_check.py --train-only --heldout-reference` | reproduces lv-hemingway's hand-computed held-out row **to the digit** (370 samples, 0.01, 0.1, max 10, 101-word chunks) |
| `show_memorisation_matches.py` (new) | reproduces the lv-hemingway reading **to the word**, including `swift tristan` as the one name-shaped hit |
| `ship-voice-adapter.sh` (new) | the live lv-hemingway adapter verified byte-identical to gx10's checkpoint-850 |
⚠ The held-out control and the match reading were both done **by hand** for Hemingway and left
no instrument, so the finding was not reproducible. Both are now committed.
⚠ **The waiter I first armed was silently broken**: `pgrep -f "eval-mccarthy.sh"` over ssh
self-matches its own `bash -c` argv, so the DIED branch was unreachable and a crashed gate
would have looked exactly like a running one. Controlled both ways after the fix — bare
pattern demonstrably self-matches, bracketed one does not.
→ [[feedback_pkill_ssh_self_match]]
## What to do next — ckpt450 is the candidate, and an earlier one may beat it
The damage grows monotonically with epoch: base 1% over-band → ckpt450 20% → ckpt900 28%.
`--save-total-limit 60` kept **all 56 checkpoints at 25-step intervals**, so ckpt225 / ckpt300
(epoch ~0.49 / 0.65) are on disk and untested. Voice at 0.5 epoch is unknown and may not have
arrived; that is one arm each to find out. This is a salvageable adapter, not a failed one —
the voice transferred and the memorisation is the cleanest in the line.
Related: [[2026-09-17-lv-hemingway-gate]], [[2026-09-17-lv-bronte-gate]],
[[2026-09-17-mccarthy-d1-d3]], [[2026-09-17-mccarthy-split-name-leak]].
---
# AMENDMENT-3 ROUND: five arms, and ckpt300 wins every axis — held at the gate by 0.02
Two more arms gated after the frozen axis C was proven **unsatisfiable** (a perfect adapter
fails it by 0.09 — see `GATE-PREREG.md` AMENDMENT 3). Same fixture, same four seeds, same rule.
**1,200 generations total.**
## The complete ladder
| arm | epoch | VOICE primary | VOICE normalised | MEMORISATION (author = 0.12 / max 12) | ran-on Δ | in-band |
|---|---|---|---|---|---|---|
| **ckpt300** | 0.652 | **+0.177, 3.2×** | **+0.128, 2.8×** | **0.12 / 1.0 / max 10** | **+0.12** | 0.65 |
| ckpt225 | 0.489 | +0.176, 1.8× | +0.122, 1.5× | 0.12 / 1.0 / max 10 | +0.37 | 0.52 |
| ckpt450 | 0.980 | +0.152, 2.9× | +0.114, 2.5× | 0.12 / 1.1 / max 11 | +0.19 | 0.67 |
| ckpt900 | 1.958 | +0.172, 1.2× | +0.124, 1.2× | 0.22 / 1.9 / max 11 | +0.27 | 0.67 |
| base | — | — | — | 0.00 / 0.0 / max 8 | — | 0.89 |
```
length distribution, raw generations
arm n median p10 p90 max in-band under90 over140 TOTAL out
base 240 110 89 126 149 0.89 10% 1% 11%
ckpt225 240 127 89 206 296 0.52 10% 38% 48%
ckpt300 240 101 83 151 216 0.65 22% 13% 35%
ckpt450 240 107 87 171 297 0.67 14% 20% 33%
ckpt900 240 114 92 190 279 0.67 5% 28% 33%
```
⭐⭐ **ckpt300 is the best arm in the run on every axis that resolves.** Best voice point
estimate AND best margin on both reads, tightest spread of any adapted arm bar ckpt450
(0.055 / 0.038, no outlier seed), memorisation **identical to real unseen McCarthy** with a
max of 10 against the author's coincidental 12, and the **least overshoot of any checkpoint**.
Its 31 matched runs were all read: `looked at the wolf and he looked at the boy`, `took off
his hat and set it on the`, `and wiped his mouth on the back of his`. Name-shaped hits are
`Adam Caleb` / `Colton` — the renamed inventions. Nothing protectable.
## ⚠ THREE CLAIMS I MADE EARLIER THAT THE FIVE-ARM DATA REFUTES
1. **"The damage is flat across epochs and only rotates direction."** FALSE. ckpt225 is 48%
out-of-band against ckpt300's 35%, and ran-on is **non-monotonic**: 0.38 → 0.13 → 0.20 →
0.28 across epochs 0.49 → 0.65 → 0.98 → 1.96. There is a genuine minimum near epoch 0.65.
2. **"ckpt300 is coming out far too SHORT, and ckpt225 will clear ran-on by being short."**
FALSE on both. ⚠ **I generalised from SIX generations of one arm** — the exact n=1 violation
the measurement-discipline rule names, committed while writing a note about being careful.
ckpt225 runs **long** (median 127, 38% over-band) and is the worst arm in the run.
3. **My own option A — "gate an earlier checkpoint, the overshoot may not have arrived yet" —
was RIGHT, and I retracted it an hour later on a three-arm read.** ckpt300 has the best
voice *and* the least damage. The retraction was the error, not the recommendation.
## The gate's verdict on ckpt300, stated exactly
```
1. VOICE +0.177 at 3.2x floor (primary), +0.128 at 2.8x (normalised) PASS
2. MEMORISATION 0.12 vs the author's own 0.12, max 10 vs the author's 12 PASS
3. DAMAGE ran-on delta +0.12
vs the operator's ratified v2 floor 0.200 -> PASS, 40% headroom
vs AMENDMENT 3's self-imposed bar 0.100 -> FAIL by 0.02
```
**NOT SHIPPED, and the reason is the bar rather than the adapter.** AMENDMENT 3 fixed
`ran-on ≤ 0.100` before either new arm existed, specifically so a marginal number could not be
talked into a ship. ckpt300 is 0.12. Shipping it on my own authority would make the
pre-registration theatre.
⚠ **But the bar's stated RATIONALE does not describe this candidate.** It was written against
ckpt450's `+0.19 vs 0.200` — a pass by **0.01**, i.e. 5% of the threshold — and the sentence
justifying it says exactly that. ckpt300 clears the operator's threshold by **40%**. The bar's
*number* excludes a candidate its *reasoning* does not. That is an operator call, not mine.
⚠⚠ **AND I DID NOT GO LOOKING FOR A CHECKPOINT UNDER 0.100.** ckpt325/350/375 are all on disk
and one of them may well sit below the bar. Searching the checkpoint space until something
clears is **candidate-shopping** — the same family as threshold-shopping, arrived at from the
other side. The measuring stopped here deliberately.
## What is actually true about this adapter
The voice transferred, ~3/4 of the gain survives stripping every punctuation mark, and the
memorisation is indistinguishable from the author's own self-collision rate on a corpus in
copyright with a living estate. **The cost is real and unresolved by any checkpoint choice:
35% of generations miss the 90–140 band against base's 11%, and in-band is 0.65 against 0.89.**
An adapter that buys a voice and costs a third of your length compliance is a trade, not a
defect — but it is the operator's trade to accept.
**Recommendation: ship ckpt300.** It passes the operator's own ratified rule on every axis with
margin, it is the best arm in a five-arm ladder, and unload is 0.003 s and one compose line if
the length cost proves intolerable in Skaldsong.