reference: held-out McCarthy, 198,486 words, 248 chunks, 400 char-bigram features same-author target (held-out McCarthy vs itself): delta_cb = 0.370 -> the floor of what any arm could reach; lower is more McCarthy-like, this is the best possible arm delta_cb per-seed [words] ckpt900 0.490 (0.605 0.457 0.531 0.554) [30935] spread 0.148 ckpt450 0.509 (0.562 0.530 0.563 0.567) [28494] spread 0.037 base-unadapted 0.661 (0.711 0.667 0.673 0.659) [26078] spread 0.052 all-arms noise floor (largest within-arm seed spread, lv-bronte's rule): 0.148 PAIRWISE floor is the verdict: max(spread(candidate), spread(base-unadapted) = 0.052) vs base-unadapted control (positive gap = moved toward McCarthy): ckpt900 +0.172 (MOVED toward McCarthy (1.2x the pairwise floor 0.148)) ckpt450 +0.152 (MOVED toward McCarthy (2.9x the pairwise floor 0.052)) ordering: ckpt900 < ckpt450 < base-unadapted (lower = more McCarthy-like) ⚠ RELATIVE reading on one harness: 4 seed group(s) per arm, scored against this corpus's own held-out split. It is not an absolute-band claim and corroborates nothing on its own. PUNCTUATION DENSITY per 10k words -- the confound check, not an axis (the eval harness drives EVERY arm with the same register prompt, tics included; a compliant base control earns the adapter no delta_cb for them) arm quote-marks all-apos contraction-apos dashes held-out reference 0.0 121.1 117.6 7.6 base-unadapted 19.9 204.8 203.6 1.9 ckpt450 0.0 167.4 166.4 0.0 ckpt900 0.0 160.7 160.7 0.0 [PASS] base control quote density 19.9 <= 100 per 10k: the control complied with the register, so the punctuation win is handed to both sides and the primary read stands. ============================================================================== SECONDARY READ -- PUNCTUATION STRIPPED. Pre-registered, REPORTED, NOT THE VERDICT. Every punctuation mark is removed from the reference and from every arm, so a gap that survives here is carried by words rather than by marks. It is a LOWER BOUND and not a better measurement: stripping terminal punctuation also strips sentence-length signal the adapter legitimately learned. Read it as `at least this much of the primary gap is not the punctuation trick`. ============================================================================== reference: held-out McCarthy, 200,682 words, 251 chunks, 400 char-bigram features same-author target (held-out McCarthy vs itself): delta_cb = 0.363 -> the floor of what any arm could reach; lower is more McCarthy-like, this is the best possible arm delta_cb per-seed [words] ckpt900 0.464 (0.563 0.460 0.493 0.522) [31440] spread 0.104 ckpt450 0.474 (0.518 0.506 0.528 0.526) [28975] spread 0.022 base-unadapted 0.588 (0.634 0.611 0.594 0.588) [26645] spread 0.046 all-arms noise floor (largest within-arm seed spread, lv-bronte's rule): 0.104 PAIRWISE floor is the verdict: max(spread(candidate), spread(base-unadapted) = 0.046) vs base-unadapted control (positive gap = moved toward McCarthy): ckpt900 +0.124 (MOVED toward McCarthy (1.2x the pairwise floor 0.104)) ckpt450 +0.114 (MOVED toward McCarthy (2.5x the pairwise floor 0.046)) ordering: ckpt900 < ckpt450 < base-unadapted (lower = more McCarthy-like) ⚠ RELATIVE reading on one harness: 4 seed group(s) per arm, scored against this corpus's own held-out split. It is not an absolute-band claim and corroborates nothing on its own.