1,200 generations across five arms. ckpt300 (epoch 0.652) is the best arm in the
run on every axis that resolves:
VOICE +0.177 at 3.2x its pairwise floor (primary), +0.128 at 2.8x with
every punctuation mark stripped. Best point estimate AND best
margin of any arm, spread 0.055/0.038 with no outlier seed.
MEMORISATION 0.12 against real unseen McCarthy's own 0.12 -- identical -- with
a longest match of 10 words against the author's coincidental 12.
All 31 matches read: stock grammar, names are the renamed
inventions, nothing protectable.
DAMAGE ran-on +0.12. Clears the operator's ratified v2 floor of 0.200 by
40%. FAILS AMENDMENT 3's self-imposed 0.100 bar by 0.02.
NOT SHIPPED, and the reason is the bar rather than the adapter. AMENDMENT 3 fixed
ran-on <= 0.100 before either new arm existed, precisely so a marginal number could
not be talked into a ship, and shipping at 0.12 would make that pre-registration
theatre. But the bar's stated rationale was written against ckpt450's pass by 0.01
-- 5% of the threshold -- and ckpt300 clears by 40%. The number excludes a candidate
the reasoning does not. That is an operator call.
ckpt325/350/375 are on disk and one may sit under 0.100. They were deliberately NOT
gated: searching the checkpoint space until something clears is candidate-shopping,
the same family as threshold-shopping approached from the other side.
THREE CLAIMS FROM EARLIER THIS SESSION ARE REFUTED and are corrected in the record:
1. "The damage is flat across epochs and only rotates direction" -- FALSE. ran-on
is non-monotonic (0.38 -> 0.13 -> 0.20 -> 0.28 across epochs 0.49/0.65/0.98/
1.96) with a real minimum near 0.65, and ckpt225 is 48% out-of-band against
ckpt300's 35%.
2. "ckpt300 runs far too short, ckpt225 will clear ran-on by being short" -- FALSE
on both. ckpt225 runs LONG (median 127, 38% over-band) and is the worst arm in
the run. I generalised from SIX generations of one arm, which is the exact n=1
violation the measurement-discipline rule names, committed in the same breath
as a note about being careful.
3. The original "gate an earlier checkpoint, the overshoot may not have arrived
yet" recommendation was RIGHT. Retracting it an hour later on a three-arm read
was the error, not the recommendation.
What is true and unresolved by any checkpoint choice: 35% of ckpt300's generations
miss the 90-140 band against base's 11%, and in-band is 0.65 against 0.89. An
adapter that buys a voice and costs a third of the length compliance is a trade, not
a defect -- but it is the operator's trade to accept.
Raw artifacts for all five arms at scripts/mccarthy-corpus/gate-results/.
69 lines
4.6 KiB
Plaintext
69 lines
4.6 KiB
Plaintext
reference: held-out McCarthy, 198,486 words, 248 chunks, 400 char-bigram features
|
|
same-author target (held-out McCarthy vs itself): delta_cb = 0.370
|
|
-> the floor of what any arm could reach; lower is more McCarthy-like, this is the best possible
|
|
|
|
arm delta_cb per-seed [words]
|
|
ckpt300 0.484 (0.561 0.509 0.564 0.519) [26444] spread 0.055
|
|
ckpt225 0.485 (0.592 0.493 0.513 0.542) [33153] spread 0.099
|
|
ckpt900 0.490 (0.605 0.457 0.531 0.554) [30935] spread 0.148
|
|
ckpt450 0.509 (0.562 0.530 0.563 0.567) [28494] spread 0.037
|
|
base-unadapted 0.661 (0.711 0.667 0.673 0.659) [26078] spread 0.052
|
|
|
|
all-arms noise floor (largest within-arm seed spread, lv-bronte's rule): 0.148
|
|
PAIRWISE floor is the verdict: max(spread(candidate), spread(base-unadapted) = 0.052)
|
|
|
|
vs base-unadapted control (positive gap = moved toward McCarthy):
|
|
ckpt300 +0.177 (MOVED toward McCarthy (3.2x the pairwise floor 0.055))
|
|
ckpt225 +0.176 (MOVED toward McCarthy (1.8x the pairwise floor 0.099))
|
|
ckpt900 +0.172 (MOVED toward McCarthy (1.2x the pairwise floor 0.148))
|
|
ckpt450 +0.152 (MOVED toward McCarthy (2.9x the pairwise floor 0.052))
|
|
|
|
ordering: ckpt300 < ckpt225 < ckpt900 < ckpt450 < base-unadapted (lower = more McCarthy-like)
|
|
⚠ RELATIVE reading on one harness: 4 seed group(s) per arm, scored against this corpus's own held-out split. It is not an absolute-band claim and corroborates nothing on its own.
|
|
|
|
PUNCTUATION DENSITY per 10k words -- the confound check, not an axis
|
|
(the eval harness drives EVERY arm with the same register prompt, tics included;
|
|
a compliant base control earns the adapter no delta_cb for them)
|
|
arm quote-marks all-apos contraction-apos dashes
|
|
held-out reference 0.0 121.1 117.6 7.6
|
|
base-unadapted 19.9 204.8 203.6 1.9
|
|
ckpt225 0.0 155.3 154.7 0.0
|
|
ckpt300 0.8 188.3 188.3 0.0
|
|
ckpt450 0.0 167.4 166.4 0.0
|
|
ckpt900 0.0 160.7 160.7 0.0
|
|
|
|
[PASS] base control quote density 19.9 <= 100 per 10k: the control complied with the register,
|
|
so the punctuation win is handed to both sides and the primary read stands.
|
|
|
|
==============================================================================
|
|
SECONDARY READ -- PUNCTUATION STRIPPED. Pre-registered, REPORTED, NOT THE VERDICT.
|
|
Every punctuation mark is removed from the reference and from every arm, so a
|
|
gap that survives here is carried by words rather than by marks. It is a LOWER
|
|
BOUND and not a better measurement: stripping terminal punctuation also strips
|
|
sentence-length signal the adapter legitimately learned. Read it as `at least
|
|
this much of the primary gap is not the punctuation trick`.
|
|
==============================================================================
|
|
|
|
reference: held-out McCarthy, 200,682 words, 251 chunks, 400 char-bigram features
|
|
same-author target (held-out McCarthy vs itself): delta_cb = 0.363
|
|
-> the floor of what any arm could reach; lower is more McCarthy-like, this is the best possible
|
|
|
|
arm delta_cb per-seed [words]
|
|
ckpt300 0.460 (0.525 0.510 0.532 0.494) [26956] spread 0.038
|
|
ckpt900 0.464 (0.563 0.460 0.493 0.522) [31440] spread 0.104
|
|
ckpt225 0.466 (0.561 0.481 0.498 0.516) [33678] spread 0.080
|
|
ckpt450 0.474 (0.518 0.506 0.528 0.526) [28975] spread 0.022
|
|
base-unadapted 0.588 (0.634 0.611 0.594 0.588) [26645] spread 0.046
|
|
|
|
all-arms noise floor (largest within-arm seed spread, lv-bronte's rule): 0.104
|
|
PAIRWISE floor is the verdict: max(spread(candidate), spread(base-unadapted) = 0.046)
|
|
|
|
vs base-unadapted control (positive gap = moved toward McCarthy):
|
|
ckpt300 +0.128 (MOVED toward McCarthy (2.8x the pairwise floor 0.046))
|
|
ckpt900 +0.124 (MOVED toward McCarthy (1.2x the pairwise floor 0.104))
|
|
ckpt225 +0.122 (MOVED toward McCarthy (1.5x the pairwise floor 0.080))
|
|
ckpt450 +0.114 (MOVED toward McCarthy (2.5x the pairwise floor 0.046))
|
|
|
|
ordering: ckpt300 < ckpt900 < ckpt225 < ckpt450 < base-unadapted (lower = more McCarthy-like)
|
|
⚠ RELATIVE reading on one harness: 4 seed group(s) per arm, scored against this corpus's own held-out split. It is not an absolute-band claim and corroborates nothing on its own.
|