Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-21-lv-mccarthy-gate.md
T
Vuong Hoang 4c3f3896f1 docs(lv-mccarthy): record the gate result -- voice passes, memorisation is the cleanest in the line, NOT shipped
720 generations, 3 arms x 60 held-out beats x 4 seeds, against the design frozen
in GATE-PREREG.md before any arm existed.

AXIS A VOICE -- PASS, both candidates, both reads. Span 0.661 -> 0.370 = 0.291
achievable; ckpt900 closed 59.1% (+0.172, but only 1.2x its floor), ckpt450 52.2%
(+0.152 at 2.9x). The normalised secondary read HOLDS at +0.124 / +0.114, so about
three quarters of the gain survives stripping every punctuation mark -- the voice
is not the cheap win the register made available.

AXIS B NOT COPIED -- ckpt450 is the cleanest result in the line. 0.12 hit-rate
against the author's own held-out 0.12, and its longest match (11 words) is
SHORTER than the author's coincidental longest (12). All 96 matched runs were
read: stock grammar in the commonest words, the name-shaped hits are the RENAMED
inventions, nothing protectable. The amendment is why this reads as clean -- the
defective base control would have shown 0.12 vs 0.00 as a 12x red flag.
Separately measured: the "his register makes collisions inevitable" story that
was FALSE for Hemingway (0.01) is TRUE for McCarthy (0.12). Neither transfers.

AXIS C NO DAMAGE -- FAIL, both, and it survives reading. 20% (ckpt450) / 28%
(ckpt900) of generations overshoot the 90-140 band against base's 1%; p90 171/190
words, max 297/279. The worst case is degenerate looping, not a long McCarthy
sentence. Base is GOOD on this axis here (0.89 in-band vs Hemingway's 0.05), so
the adapter measurably makes instruction-following worse.

NOT SHIPPED. Section 7 rule 3 makes axis C disqualifying outright.

Recorded honestly: my own prereg's axis C transcribed score_beats.py's v1
criteria, including "in-band up on base", which the operator RETIRED on
2026-09-15 for exactly the reason it fails here -- base maxes it, so it is
unsatisfiable on this corpus regardless of adapter quality. Under the operator's
v2 (ran-on only) ckpt450 passes by 0.01 against a 0.200 floor. That reading was
found AFTER the numbers and was NOT used; lv-bronte's floor defect is the in-line
precedent for finding one and declining to exploit it. The prereg gets fixed
prospectively for the next author, not re-read for this one.

And the finding worth more than the adapter: the two-epoch recipe is now 0 for 3,
and this time the loss curve was CONFIDENTLY wrong. On Bronte and Hemingway the
epoch-1/epoch-2 checkpoints were tied, so preferring the earlier one cost nothing.
Here the curve resolved epoch 2 as better at 4.9x the median neighbour jitter --
and epoch 2 lost every axis that resolves: 4.0x wider seed spread, 1.8x the
author's memorisation rate against 1.0x, more ran-on, worse on-beat. Its only win
is a 0.019 voice point estimate, inside the floor, and its spread is one outlier
seed -- the third occurrence of that shape in the later checkpoint after
lv-bronte's ckpt925 and lv-hemingway's ckpt1750.

Raw artifacts committed at scripts/mccarthy-corpus/gate-results/ so the claims can
be re-read without gx10.
2026-09-21 16:31:41 -07:00

218 lines
12 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# `[2026-09-21]` lv-mccarthy: gate run, NOT SHIPPED — the voice works, the memorisation is the cleanest in the line, and the length discipline is gone
**Status: NOT SHIPPED.** Both gated candidates are disqualified by axis C under the
pre-registration as frozen. The operator had pre-authorised "ship it if the gate passes";
it did not pass, and their conditional preserved that branch.
Gate design pre-registered before any generation existed:
`scripts/mccarthy-corpus/GATE-PREREG.md`, commit `9c8a4e9`, amended `b4ba731` and `0d80e49`.
Raw artifacts committed at `scripts/mccarthy-corpus/gate-results/`; arms and generations at
`gx10:~/r49-runs/mccarthy-eval/`.
## First: the run WAS complete, and had been for three days
`persistent-memory.md` carried "lv-mccarthy's run outcome STILL UNVERIFIED" as the top
in-flight item for two days. It finished **2026-09-18 00:49 PT** — 1,380/1,380 steps in
2h30m47s, `train_loss` 2.172. Three days of "unverified" was a reporting gap, not a failure.
⚠ `~/r49-runs/` does not exist on nh3-dev; the runs are on **pfi-gx10 = 10.100.50.60**, which
does not resolve by name from nh3-dev.
## The result — 3 arms × 60 held-out beats × 4 seeds = 720 generations
| axis | ckpt900 (ep 1.96, loss min) | ckpt450 (ep 0.98) | verdict |
|---|---|---|---|
| **A. VOICE** | +0.172, **1.2×** floor 0.148 | +0.152, **2.9×** floor 0.052 | ✅ both PASS, both reads |
| **B. NOT COPIED** | 0.22 = 1.8× the author | **0.12 = 1.0× the author** | ✅ ckpt450 clean, ckpt900 elevated |
| **C. NO DAMAGE** | fails all three criteria | fails two of three | ❌ **both FAIL** |
```
same-author target (held-out McCarthy vs itself) delta_cb 0.370 <- best achievable
ckpt900 0.490
ckpt450 0.509
base-unadapted 0.661
```
Span 0.661 → 0.370 = 0.291. ckpt900 closed **59.1%**, ckpt450 **52.2%** — between lv-bronte
(48%) and lv-hemingway (73.8%), on an axis deliberately made harder (§5a hands the register's
punctuation tics to the control).
## ⭐⭐ AXIS B: the corrected control turned a 12× red flag into a clean pass
The pre-registration as first written inherited `memorization_check.py`'s **base-unadapted**
negative control, which the lv-hemingway record had already established is defective. Caught
while the base arm was still generating and amended append-only before any McCarthy number
was read (AMENDMENT 1). The correct innocent sample is the author himself.
```
sample n hit-rate mean-longest max
HELD-OUT McCARTHY (untrained) 364 0.12 1.0 12 <- the innocent rate
base-unadapted 240 0.00 0.0 8
ckpt450 240 0.12 1.1 11
ckpt900 240 0.22 1.9 11
positive control (train vs train) 160 <- not blind
```
⭐ **ckpt450 is INDISTINGUISHABLE from real unseen McCarthy** — 0.12 against 0.12, and its
longest match (11 words) is **SHORTER than the author's own coincidental longest (12)**. That
is the best axis-B result in the line. Against the defective base control its 0.12-vs-0.00
would have read as a **12× red flag**; against the correct one it is **1.0×**. The amendment
is the only reason this adapter is not on the record as memorising.
⭐ **And the "register makes collisions inevitable" story — FALSE for Hemingway — is TRUE
here, measured rather than assumed.** Hemingway's held-out rate was 0.01; McCarthy's is 0.12.
His idiom really does self-collide at 8 grams (`and he looked at the`), so the same argument
that was a comfortable excuse there is a fact here. Authors are not interchangeable on this
axis and neither number transfers.
**All 96 matched runs were READ** (`gate-results/memorisation-matches.txt`), not counted.
Longest 11 words, every one stock grammar in the commonest words:
```
he looked at the wolf and he looked at him
leaned back in his chair and looked at the boy
He reached into his pocket and took out the
I dont know What are you goin to do
```
The name-shaped hits (`Adam Caleb`, `Toribio`) are the **RENAMED invented names**, not
McCarthy's — the `swift tristan` precedent exactly. No plot, no imagery, nothing protectable.
⚠ McCarthy is in copyright with a living estate and this axis still comes out clean for
ckpt450.
## ❌ AXIS C: the damage is real, and reading it confirms the criterion
```
arm n in-band on-beat ran-on words p90 max over-140
base 240 0.89 0.71 0.01 110 126 149 1%
ckpt450 240 0.67 0.45 0.20 108 171 297 20%
ckpt900 240 0.67 0.42 0.28 114 190 279 28%
floor 0.200
```
Not an artifact. **Base is GOOD here** (0.89 in-band against Hemingway's 0.05, because
McCarthy's register prompt is far more prescriptive and Qwen3-4B-Instruct follows it), and the
adapter measurably makes it worse — 22 points of in-band, 26–29 of on-beat, and 20–28× base's
overshoot rate.
Read the worst case and it is **degenerate looping**, not a long McCarthy sentence:
> *…and then he looked at the wolf and he looked at the road and he looked at the sun and he
> looked at the road again.* — ckpt900, b41, 279 words against a 90–140 ask
## ⚠ THE COMPLICATION, AND WHY IT DID NOT BECOME A SHIP
My pre-registration's §6 axis C transcribed `score_beats.py`'s **v1** three-part criterion,
including "in-band up on base beyond the floor". The operator **retired** in-band and on-beat
from the gate on 2026-09-15 for precisely the reason they fail here — the script's own
docstring says *"NOT carried into v2: in-band (unresolvable — base maxes it) and on-beat."*
The v2 axis C is **ran-on only**, and `VERDICT: DO-NOT-SCALE` in the output is the superseded
v1 label the script still prints "for continuity".
```
frozen prereg (v1 criteria) ckpt900 fails 3 of 3 ckpt450 fails 2 of 3
operator's v2 (ran-on only) ckpt900 +0.27 FAILS ckpt450 +0.19 passes by 0.01
```
So my own pre-registration contains a criterion **unsatisfiable on this corpus regardless of
adapter quality**. That is a defect in the pre-registration, not in the adapter — and it is
still not a reason to ship:
1. The frozen rule fails both candidates, and §7 rule 3 makes axis C disqualifying outright.
2. The reading that rescues ckpt450 was found **after** seeing the numbers. That is the exact
shape pre-registration exists to prevent, and lv-bronte's floor defect is the in-line
precedent for finding one and deliberately not exploiting it.
3. Even under that reading the margin is **0.01 against a floor of 0.200** — noise-adjacent.
4. The damage survives reading, not just the criterion. 20% of ckpt450's generations blow the
length budget and the worst ones loop.
**Fix the pre-registration prospectively for the next author, do not re-read it for this one.**
## ⭐⭐⭐ THE TWO-EPOCH RECIPE IS NOW 0 FOR 3 — AND THIS TIME THE LOSS CURVE WAS CONFIDENTLY WRONG
On Brontë and Hemingway the epoch-1 and epoch-2 checkpoints were **tied** on eval loss, so
preferring the earlier one cost nothing and could be dismissed as taste. Here the curve
**resolved** epoch 2 as better — ckpt900 at 2.38706 against ckpt450's 2.4063, **+4.9× the
0.00393 median neighbour jitter**, nowhere near tied — and it was wrong on every axis that
resolves:
| | ckpt450 | ckpt900 | ratio |
|---|---|---|---|
| seed spread (voice) | **0.037** | 0.148 | **4.0× wider** |
| memorisation vs the author's 0.12 | **0.12 (1.0×)** | 0.22 (1.8×) | |
| ran-on | **0.20** | 0.28 | |
| on-beat | **0.45** | 0.42 | |
| epochs of overfit | **0.98** | 1.96 | |
ckpt900's only advantage is a 0.019 better voice point estimate, which sits inside the floor.
And its spread is **one outlier seed** — 0.605 against 0.457 / 0.531 / 0.554 — the third
occurrence of that shape in the later checkpoint, after lv-bronte's ckpt925 and
lv-hemingway's ckpt1750 (0.491 / 0.449 / 0.468 / **0.562**).
⭐ **The durable rule: on this schedule the eval-loss minimum is not the ship candidate, and
the loss curve's CONFIDENCE about it carries no information.** Read the resolving axes. This
is now measured on three corpora and should be the default for Faulkner, Morrison and Chandler
rather than re-derived each time.
## ⚠ A bug in my own confound detector, and the order I found it in
§5c pre-registered a trigger: base quote density over **100 per 10k** means the control did
not take the punctuation win it was handed, and the normalised secondary read is promoted to
load-bearing. The base arm finished first, so I evaluated it early, **saw it FIRE at 224.7**,
and only then — reading the reference row against a corpus whose builder ASSERTS 0.0 — found
that `_QUOTE_RE` contained `'` and `’`. It was an apostrophe counter wearing a quote-mark
label, on the one corpus whose signature is `dont`/`aint`/`wont`.
```
as implemented TRUE quotes all apostrophes
held-out McCarthy ref 121.1 0.0 121.1
base-unadapted 224.7 19.9 204.8
held-out Hemingway ref 1112.6 694.7 351.7
```
Fixing a detector to measure the quantity the frozen rule names is not moving the rule — but
the fix un-fires the trigger, which is indistinguishable from shopping. **So the trigger was
made MOOT rather than adjudicated** (AMENDMENT 2): the normalised read is load-bearing
unconditionally for this gate, both columns reported, and the fix has zero verdict effect.
There is a better reason anyway — base's true density is 19.9 against the reference's 0.0, so
it did not fully comply and a small residual cheap win genuinely exists.
⭐ The normalised read **HOLDS**: ckpt900 +0.124 at 1.2× and ckpt450 +0.114 at 2.5×, against
primaries of +0.172 and +0.152. So roughly three quarters of the voice gain survives having
every punctuation mark removed — **the voice is not the punctuation trick.** Validated on
Hemingway first, where the secondary read also resolves a gap rather than flattening
everything, so a null here would have been a finding rather than a blind instrument.
⚠ **The lesson is the one this line keeps relearning somewhere new:** I controlled
`strip_punct` (2500 → 0) and the byte-identity of the default path, and never asked the quote
counter for a value whose answer I already knew. The corpus asserts 0.0. That check cost one
line and was available before the gate launched.
## Instruments built, each validated against a documented finding
| instrument | control |
|---|---|
| `voice_distance.py --secondary-normalised --punct-report` | default path reproduces the shipped lv-hemingway `voice_distance.txt` **byte for byte** |
| `memorization_check.py --train-only --heldout-reference` | reproduces lv-hemingway's hand-computed held-out row **to the digit** (370 samples, 0.01, 0.1, max 10, 101-word chunks) |
| `show_memorisation_matches.py` (new) | reproduces the lv-hemingway reading **to the word**, including `swift tristan` as the one name-shaped hit |
| `ship-voice-adapter.sh` (new) | the live lv-hemingway adapter verified byte-identical to gx10's checkpoint-850 |
⚠ The held-out control and the match reading were both done **by hand** for Hemingway and left
no instrument, so the finding was not reproducible. Both are now committed.
⚠ **The waiter I first armed was silently broken**: `pgrep -f "eval-mccarthy.sh"` over ssh
self-matches its own `bash -c` argv, so the DIED branch was unreachable and a crashed gate
would have looked exactly like a running one. Controlled both ways after the fix — bare
pattern demonstrably self-matches, bracketed one does not.
→ [[feedback_pkill_ssh_self_match]]
## What to do next — ckpt450 is the candidate, and an earlier one may beat it
The damage grows monotonically with epoch: base 1% over-band → ckpt450 20% → ckpt900 28%.
`--save-total-limit 60` kept **all 56 checkpoints at 25-step intervals**, so ckpt225 / ckpt300
(epoch ~0.49 / 0.65) are on disk and untested. Voice at 0.5 epoch is unknown and may not have
arrived; that is one arm each to find out. This is a salvageable adapter, not a failed one —
the voice transferred and the memorisation is the cleanest in the line.
Related: [[2026-09-17-lv-hemingway-gate]], [[2026-09-17-lv-bronte-gate]],
[[2026-09-17-mccarthy-d1-d3]], [[2026-09-17-mccarthy-split-name-leak]].