720 generations, 3 arms x 60 held-out beats x 4 seeds, against the design frozen in GATE-PREREG.md before any arm existed. AXIS A VOICE -- PASS, both candidates, both reads. Span 0.661 -> 0.370 = 0.291 achievable; ckpt900 closed 59.1% (+0.172, but only 1.2x its floor), ckpt450 52.2% (+0.152 at 2.9x). The normalised secondary read HOLDS at +0.124 / +0.114, so about three quarters of the gain survives stripping every punctuation mark -- the voice is not the cheap win the register made available. AXIS B NOT COPIED -- ckpt450 is the cleanest result in the line. 0.12 hit-rate against the author's own held-out 0.12, and its longest match (11 words) is SHORTER than the author's coincidental longest (12). All 96 matched runs were read: stock grammar in the commonest words, the name-shaped hits are the RENAMED inventions, nothing protectable. The amendment is why this reads as clean -- the defective base control would have shown 0.12 vs 0.00 as a 12x red flag. Separately measured: the "his register makes collisions inevitable" story that was FALSE for Hemingway (0.01) is TRUE for McCarthy (0.12). Neither transfers. AXIS C NO DAMAGE -- FAIL, both, and it survives reading. 20% (ckpt450) / 28% (ckpt900) of generations overshoot the 90-140 band against base's 1%; p90 171/190 words, max 297/279. The worst case is degenerate looping, not a long McCarthy sentence. Base is GOOD on this axis here (0.89 in-band vs Hemingway's 0.05), so the adapter measurably makes instruction-following worse. NOT SHIPPED. Section 7 rule 3 makes axis C disqualifying outright. Recorded honestly: my own prereg's axis C transcribed score_beats.py's v1 criteria, including "in-band up on base", which the operator RETIRED on 2026-09-15 for exactly the reason it fails here -- base maxes it, so it is unsatisfiable on this corpus regardless of adapter quality. Under the operator's v2 (ran-on only) ckpt450 passes by 0.01 against a 0.200 floor. That reading was found AFTER the numbers and was NOT used; lv-bronte's floor defect is the in-line precedent for finding one and declining to exploit it. The prereg gets fixed prospectively for the next author, not re-read for this one. And the finding worth more than the adapter: the two-epoch recipe is now 0 for 3, and this time the loss curve was CONFIDENTLY wrong. On Bronte and Hemingway the epoch-1/epoch-2 checkpoints were tied, so preferring the earlier one cost nothing. Here the curve resolved epoch 2 as better at 4.9x the median neighbour jitter -- and epoch 2 lost every axis that resolves: 4.0x wider seed spread, 1.8x the author's memorisation rate against 1.0x, more ran-on, worse on-beat. Its only win is a 0.019 voice point estimate, inside the floor, and its spread is one outlier seed -- the third occurrence of that shape in the later checkpoint after lv-bronte's ckpt925 and lv-hemingway's ckpt1750. Raw artifacts committed at scripts/mccarthy-corpus/gate-results/ so the claims can be re-read without gx10.
59 lines
3.8 KiB
Plaintext
59 lines
3.8 KiB
Plaintext
reference: held-out McCarthy, 198,486 words, 248 chunks, 400 char-bigram features
|
|
same-author target (held-out McCarthy vs itself): delta_cb = 0.370
|
|
-> the floor of what any arm could reach; lower is more McCarthy-like, this is the best possible
|
|
|
|
arm delta_cb per-seed [words]
|
|
ckpt900 0.490 (0.605 0.457 0.531 0.554) [30935] spread 0.148
|
|
ckpt450 0.509 (0.562 0.530 0.563 0.567) [28494] spread 0.037
|
|
base-unadapted 0.661 (0.711 0.667 0.673 0.659) [26078] spread 0.052
|
|
|
|
all-arms noise floor (largest within-arm seed spread, lv-bronte's rule): 0.148
|
|
PAIRWISE floor is the verdict: max(spread(candidate), spread(base-unadapted) = 0.052)
|
|
|
|
vs base-unadapted control (positive gap = moved toward McCarthy):
|
|
ckpt900 +0.172 (MOVED toward McCarthy (1.2x the pairwise floor 0.148))
|
|
ckpt450 +0.152 (MOVED toward McCarthy (2.9x the pairwise floor 0.052))
|
|
|
|
ordering: ckpt900 < ckpt450 < base-unadapted (lower = more McCarthy-like)
|
|
⚠ RELATIVE reading on one harness: 4 seed group(s) per arm, scored against this corpus's own held-out split. It is not an absolute-band claim and corroborates nothing on its own.
|
|
|
|
PUNCTUATION DENSITY per 10k words -- the confound check, not an axis
|
|
(the eval harness drives EVERY arm with the same register prompt, tics included;
|
|
a compliant base control earns the adapter no delta_cb for them)
|
|
arm quote-marks all-apos contraction-apos dashes
|
|
held-out reference 0.0 121.1 117.6 7.6
|
|
base-unadapted 19.9 204.8 203.6 1.9
|
|
ckpt450 0.0 167.4 166.4 0.0
|
|
ckpt900 0.0 160.7 160.7 0.0
|
|
|
|
[PASS] base control quote density 19.9 <= 100 per 10k: the control complied with the register,
|
|
so the punctuation win is handed to both sides and the primary read stands.
|
|
|
|
==============================================================================
|
|
SECONDARY READ -- PUNCTUATION STRIPPED. Pre-registered, REPORTED, NOT THE VERDICT.
|
|
Every punctuation mark is removed from the reference and from every arm, so a
|
|
gap that survives here is carried by words rather than by marks. It is a LOWER
|
|
BOUND and not a better measurement: stripping terminal punctuation also strips
|
|
sentence-length signal the adapter legitimately learned. Read it as `at least
|
|
this much of the primary gap is not the punctuation trick`.
|
|
==============================================================================
|
|
|
|
reference: held-out McCarthy, 200,682 words, 251 chunks, 400 char-bigram features
|
|
same-author target (held-out McCarthy vs itself): delta_cb = 0.363
|
|
-> the floor of what any arm could reach; lower is more McCarthy-like, this is the best possible
|
|
|
|
arm delta_cb per-seed [words]
|
|
ckpt900 0.464 (0.563 0.460 0.493 0.522) [31440] spread 0.104
|
|
ckpt450 0.474 (0.518 0.506 0.528 0.526) [28975] spread 0.022
|
|
base-unadapted 0.588 (0.634 0.611 0.594 0.588) [26645] spread 0.046
|
|
|
|
all-arms noise floor (largest within-arm seed spread, lv-bronte's rule): 0.104
|
|
PAIRWISE floor is the verdict: max(spread(candidate), spread(base-unadapted) = 0.046)
|
|
|
|
vs base-unadapted control (positive gap = moved toward McCarthy):
|
|
ckpt900 +0.124 (MOVED toward McCarthy (1.2x the pairwise floor 0.104))
|
|
ckpt450 +0.114 (MOVED toward McCarthy (2.5x the pairwise floor 0.046))
|
|
|
|
ordering: ckpt900 < ckpt450 < base-unadapted (lower = more McCarthy-like)
|
|
⚠ RELATIVE reading on one harness: 4 seed group(s) per arm, scored against this corpus's own held-out split. It is not an absolute-band claim and corroborates nothing on its own.
|