# `[2026-09-21]` lv-mccarthy: gate run, NOT SHIPPED — the voice works, the memorisation is the cleanest in the line, and the length discipline is gone **Status: NOT SHIPPED.** Both gated candidates are disqualified by axis C under the pre-registration as frozen. The operator had pre-authorised "ship it if the gate passes"; it did not pass, and their conditional preserved that branch. Gate design pre-registered before any generation existed: `scripts/mccarthy-corpus/GATE-PREREG.md`, commit `9c8a4e9`, amended `b4ba731` and `0d80e49`. Raw artifacts committed at `scripts/mccarthy-corpus/gate-results/`; arms and generations at `gx10:~/r49-runs/mccarthy-eval/`. ## First: the run WAS complete, and had been for three days `persistent-memory.md` carried "lv-mccarthy's run outcome STILL UNVERIFIED" as the top in-flight item for two days. It finished **2026-09-18 00:49 PT** — 1,380/1,380 steps in 2h30m47s, `train_loss` 2.172. Three days of "unverified" was a reporting gap, not a failure. ⚠ `~/r49-runs/` does not exist on nh3-dev; the runs are on **pfi-gx10 = 10.100.50.60**, which does not resolve by name from nh3-dev. ## The result — 3 arms × 60 held-out beats × 4 seeds = 720 generations | axis | ckpt900 (ep 1.96, loss min) | ckpt450 (ep 0.98) | verdict | |---|---|---|---| | **A. VOICE** | +0.172, **1.2×** floor 0.148 | +0.152, **2.9×** floor 0.052 | ✅ both PASS, both reads | | **B. NOT COPIED** | 0.22 = 1.8× the author | **0.12 = 1.0× the author** | ✅ ckpt450 clean, ckpt900 elevated | | **C. NO DAMAGE** | fails all three criteria | fails two of three | ❌ **both FAIL** | ``` same-author target (held-out McCarthy vs itself) delta_cb 0.370 <- best achievable ckpt900 0.490 ckpt450 0.509 base-unadapted 0.661 ``` Span 0.661 → 0.370 = 0.291. ckpt900 closed **59.1%**, ckpt450 **52.2%** — between lv-bronte (48%) and lv-hemingway (73.8%), on an axis deliberately made harder (§5a hands the register's punctuation tics to the control). ## ⭐⭐ AXIS B: the corrected control turned a 12× red flag into a clean pass The pre-registration as first written inherited `memorization_check.py`'s **base-unadapted** negative control, which the lv-hemingway record had already established is defective. Caught while the base arm was still generating and amended append-only before any McCarthy number was read (AMENDMENT 1). The correct innocent sample is the author himself. ``` sample n hit-rate mean-longest max HELD-OUT McCARTHY (untrained) 364 0.12 1.0 12 <- the innocent rate base-unadapted 240 0.00 0.0 8 ckpt450 240 0.12 1.1 11 ckpt900 240 0.22 1.9 11 positive control (train vs train) 160 <- not blind ``` ⭐ **ckpt450 is INDISTINGUISHABLE from real unseen McCarthy** — 0.12 against 0.12, and its longest match (11 words) is **SHORTER than the author's own coincidental longest (12)**. That is the best axis-B result in the line. Against the defective base control its 0.12-vs-0.00 would have read as a **12× red flag**; against the correct one it is **1.0×**. The amendment is the only reason this adapter is not on the record as memorising. ⭐ **And the "register makes collisions inevitable" story — FALSE for Hemingway — is TRUE here, measured rather than assumed.** Hemingway's held-out rate was 0.01; McCarthy's is 0.12. His idiom really does self-collide at 8 grams (`and he looked at the`), so the same argument that was a comfortable excuse there is a fact here. Authors are not interchangeable on this axis and neither number transfers. **All 96 matched runs were READ** (`gate-results/memorisation-matches.txt`), not counted. Longest 11 words, every one stock grammar in the commonest words: ``` he looked at the wolf and he looked at him leaned back in his chair and looked at the boy He reached into his pocket and took out the I dont know What are you goin to do ``` The name-shaped hits (`Adam Caleb`, `Toribio`) are the **RENAMED invented names**, not McCarthy's — the `swift tristan` precedent exactly. No plot, no imagery, nothing protectable. ⚠ McCarthy is in copyright with a living estate and this axis still comes out clean for ckpt450. ## ❌ AXIS C: the damage is real, and reading it confirms the criterion ``` arm n in-band on-beat ran-on words p90 max over-140 base 240 0.89 0.71 0.01 110 126 149 1% ckpt450 240 0.67 0.45 0.20 108 171 297 20% ckpt900 240 0.67 0.42 0.28 114 190 279 28% floor 0.200 ``` Not an artifact. **Base is GOOD here** (0.89 in-band against Hemingway's 0.05, because McCarthy's register prompt is far more prescriptive and Qwen3-4B-Instruct follows it), and the adapter measurably makes it worse — 22 points of in-band, 26–29 of on-beat, and 20–28× base's overshoot rate. Read the worst case and it is **degenerate looping**, not a long McCarthy sentence: > *…and then he looked at the wolf and he looked at the road and he looked at the sun and he > looked at the road again.* — ckpt900, b41, 279 words against a 90–140 ask ## ⚠ THE COMPLICATION, AND WHY IT DID NOT BECOME A SHIP My pre-registration's §6 axis C transcribed `score_beats.py`'s **v1** three-part criterion, including "in-band up on base beyond the floor". The operator **retired** in-band and on-beat from the gate on 2026-09-15 for precisely the reason they fail here — the script's own docstring says *"NOT carried into v2: in-band (unresolvable — base maxes it) and on-beat."* The v2 axis C is **ran-on only**, and `VERDICT: DO-NOT-SCALE` in the output is the superseded v1 label the script still prints "for continuity". ``` frozen prereg (v1 criteria) ckpt900 fails 3 of 3 ckpt450 fails 2 of 3 operator's v2 (ran-on only) ckpt900 +0.27 FAILS ckpt450 +0.19 passes by 0.01 ``` So my own pre-registration contains a criterion **unsatisfiable on this corpus regardless of adapter quality**. That is a defect in the pre-registration, not in the adapter — and it is still not a reason to ship: 1. The frozen rule fails both candidates, and §7 rule 3 makes axis C disqualifying outright. 2. The reading that rescues ckpt450 was found **after** seeing the numbers. That is the exact shape pre-registration exists to prevent, and lv-bronte's floor defect is the in-line precedent for finding one and deliberately not exploiting it. 3. Even under that reading the margin is **0.01 against a floor of 0.200** — noise-adjacent. 4. The damage survives reading, not just the criterion. 20% of ckpt450's generations blow the length budget and the worst ones loop. **Fix the pre-registration prospectively for the next author, do not re-read it for this one.** ## ⭐⭐⭐ THE TWO-EPOCH RECIPE IS NOW 0 FOR 3 — AND THIS TIME THE LOSS CURVE WAS CONFIDENTLY WRONG On Brontë and Hemingway the epoch-1 and epoch-2 checkpoints were **tied** on eval loss, so preferring the earlier one cost nothing and could be dismissed as taste. Here the curve **resolved** epoch 2 as better — ckpt900 at 2.38706 against ckpt450's 2.4063, **+4.9× the 0.00393 median neighbour jitter**, nowhere near tied — and it was wrong on every axis that resolves: | | ckpt450 | ckpt900 | ratio | |---|---|---|---| | seed spread (voice) | **0.037** | 0.148 | **4.0× wider** | | memorisation vs the author's 0.12 | **0.12 (1.0×)** | 0.22 (1.8×) | | | ran-on | **0.20** | 0.28 | | | on-beat | **0.45** | 0.42 | | | epochs of overfit | **0.98** | 1.96 | | ckpt900's only advantage is a 0.019 better voice point estimate, which sits inside the floor. And its spread is **one outlier seed** — 0.605 against 0.457 / 0.531 / 0.554 — the third occurrence of that shape in the later checkpoint, after lv-bronte's ckpt925 and lv-hemingway's ckpt1750 (0.491 / 0.449 / 0.468 / **0.562**). ⭐ **The durable rule: on this schedule the eval-loss minimum is not the ship candidate, and the loss curve's CONFIDENCE about it carries no information.** Read the resolving axes. This is now measured on three corpora and should be the default for Faulkner, Morrison and Chandler rather than re-derived each time. ## ⚠ A bug in my own confound detector, and the order I found it in §5c pre-registered a trigger: base quote density over **100 per 10k** means the control did not take the punctuation win it was handed, and the normalised secondary read is promoted to load-bearing. The base arm finished first, so I evaluated it early, **saw it FIRE at 224.7**, and only then — reading the reference row against a corpus whose builder ASSERTS 0.0 — found that `_QUOTE_RE` contained `'` and `’`. It was an apostrophe counter wearing a quote-mark label, on the one corpus whose signature is `dont`/`aint`/`wont`. ``` as implemented TRUE quotes all apostrophes held-out McCarthy ref 121.1 0.0 121.1 base-unadapted 224.7 19.9 204.8 held-out Hemingway ref 1112.6 694.7 351.7 ``` Fixing a detector to measure the quantity the frozen rule names is not moving the rule — but the fix un-fires the trigger, which is indistinguishable from shopping. **So the trigger was made MOOT rather than adjudicated** (AMENDMENT 2): the normalised read is load-bearing unconditionally for this gate, both columns reported, and the fix has zero verdict effect. There is a better reason anyway — base's true density is 19.9 against the reference's 0.0, so it did not fully comply and a small residual cheap win genuinely exists. ⭐ The normalised read **HOLDS**: ckpt900 +0.124 at 1.2× and ckpt450 +0.114 at 2.5×, against primaries of +0.172 and +0.152. So roughly three quarters of the voice gain survives having every punctuation mark removed — **the voice is not the punctuation trick.** Validated on Hemingway first, where the secondary read also resolves a gap rather than flattening everything, so a null here would have been a finding rather than a blind instrument. ⚠ **The lesson is the one this line keeps relearning somewhere new:** I controlled `strip_punct` (2500 → 0) and the byte-identity of the default path, and never asked the quote counter for a value whose answer I already knew. The corpus asserts 0.0. That check cost one line and was available before the gate launched. ## Instruments built, each validated against a documented finding | instrument | control | |---|---| | `voice_distance.py --secondary-normalised --punct-report` | default path reproduces the shipped lv-hemingway `voice_distance.txt` **byte for byte** | | `memorization_check.py --train-only --heldout-reference` | reproduces lv-hemingway's hand-computed held-out row **to the digit** (370 samples, 0.01, 0.1, max 10, 101-word chunks) | | `show_memorisation_matches.py` (new) | reproduces the lv-hemingway reading **to the word**, including `swift tristan` as the one name-shaped hit | | `ship-voice-adapter.sh` (new) | the live lv-hemingway adapter verified byte-identical to gx10's checkpoint-850 | ⚠ The held-out control and the match reading were both done **by hand** for Hemingway and left no instrument, so the finding was not reproducible. Both are now committed. ⚠ **The waiter I first armed was silently broken**: `pgrep -f "eval-mccarthy.sh"` over ssh self-matches its own `bash -c` argv, so the DIED branch was unreachable and a crashed gate would have looked exactly like a running one. Controlled both ways after the fix — bare pattern demonstrably self-matches, bracketed one does not. → [[feedback_pkill_ssh_self_match]] ## What to do next — ckpt450 is the candidate, and an earlier one may beat it The damage grows monotonically with epoch: base 1% over-band → ckpt450 20% → ckpt900 28%. `--save-total-limit 60` kept **all 56 checkpoints at 25-step intervals**, so ckpt225 / ckpt300 (epoch ~0.49 / 0.65) are on disk and untested. Voice at 0.5 epoch is unknown and may not have arrived; that is one arm each to find out. This is a salvageable adapter, not a failed one — the voice transferred and the memorisation is the cleanest in the line. Related: [[2026-09-17-lv-hemingway-gate]], [[2026-09-17-lv-bronte-gate]], [[2026-09-17-mccarthy-d1-d3]], [[2026-09-17-mccarthy-split-name-leak]].