720 generations, 3 arms x 60 held-out beats x 4 seeds, against the design frozen in GATE-PREREG.md before any arm existed. AXIS A VOICE -- PASS, both candidates, both reads. Span 0.661 -> 0.370 = 0.291 achievable; ckpt900 closed 59.1% (+0.172, but only 1.2x its floor), ckpt450 52.2% (+0.152 at 2.9x). The normalised secondary read HOLDS at +0.124 / +0.114, so about three quarters of the gain survives stripping every punctuation mark -- the voice is not the cheap win the register made available. AXIS B NOT COPIED -- ckpt450 is the cleanest result in the line. 0.12 hit-rate against the author's own held-out 0.12, and its longest match (11 words) is SHORTER than the author's coincidental longest (12). All 96 matched runs were read: stock grammar in the commonest words, the name-shaped hits are the RENAMED inventions, nothing protectable. The amendment is why this reads as clean -- the defective base control would have shown 0.12 vs 0.00 as a 12x red flag. Separately measured: the "his register makes collisions inevitable" story that was FALSE for Hemingway (0.01) is TRUE for McCarthy (0.12). Neither transfers. AXIS C NO DAMAGE -- FAIL, both, and it survives reading. 20% (ckpt450) / 28% (ckpt900) of generations overshoot the 90-140 band against base's 1%; p90 171/190 words, max 297/279. The worst case is degenerate looping, not a long McCarthy sentence. Base is GOOD on this axis here (0.89 in-band vs Hemingway's 0.05), so the adapter measurably makes instruction-following worse. NOT SHIPPED. Section 7 rule 3 makes axis C disqualifying outright. Recorded honestly: my own prereg's axis C transcribed score_beats.py's v1 criteria, including "in-band up on base", which the operator RETIRED on 2026-09-15 for exactly the reason it fails here -- base maxes it, so it is unsatisfiable on this corpus regardless of adapter quality. Under the operator's v2 (ran-on only) ckpt450 passes by 0.01 against a 0.200 floor. That reading was found AFTER the numbers and was NOT used; lv-bronte's floor defect is the in-line precedent for finding one and declining to exploit it. The prereg gets fixed prospectively for the next author, not re-read for this one. And the finding worth more than the adapter: the two-epoch recipe is now 0 for 3, and this time the loss curve was CONFIDENTLY wrong. On Bronte and Hemingway the epoch-1/epoch-2 checkpoints were tied, so preferring the earlier one cost nothing. Here the curve resolved epoch 2 as better at 4.9x the median neighbour jitter -- and epoch 2 lost every axis that resolves: 4.0x wider seed spread, 1.8x the author's memorisation rate against 1.0x, more ran-on, worse on-beat. Its only win is a 0.019 voice point estimate, inside the floor, and its spread is one outlier seed -- the third occurrence of that shape in the later checkpoint after lv-bronte's ckpt925 and lv-hemingway's ckpt1750. Raw artifacts committed at scripts/mccarthy-corpus/gate-results/ so the claims can be re-read without gx10.
16 lines
722 B
Plaintext
16 lines
722 B
Plaintext
measurement surface: raw
|
|
arm n in-band on-beat cover ran-on words
|
|
-------------------------------------------------------------------------
|
|
base 240 0.89 0.71 0.63 0.01 110
|
|
ckpt900 240 0.67 0.42 0.44 0.28 114
|
|
ckpt450 240 0.67 0.45 0.44 0.20 108
|
|
|
|
noise floor (max within-arm spread across seeds): 0.200
|
|
⚠ any between-arm gap at or under that is NOT a finding
|
|
|
|
1. in-band 0.67 vs 0.89 delta -0.22 FAIL (needs > +0.200)
|
|
2. on-beat 0.42 vs 0.71 delta -0.29 FAIL (needs >= -0.200)
|
|
3. ran-on 0.28 vs 0.01 delta +0.27 FAIL (needs <= +0.200)
|
|
|
|
VERDICT: DO-NOT-SCALE
|