8.9 KiB
[2026-09-17] lv-mccarthy D1→D3 — built, gated, and every stage caught a defect in the stage before it
~/lv-mccarthy/ on pfi-gx10. corpus-clean/ (167 units, 584,716 words),
corpus-renamed/ (6 copies, 1,002 records), scripts/. Commits 705fa3a f3bf3ca
0fa68cb 5aa10bf 5ddb047.
leak gate 0 of 75 renameable and 0 of 37 sub-threshold survive in any copy
positive control 108/108 surfaces found in the unrenamed source
negative control nonce absent from both trees
The shared splitter: choose by SIZE, not by count
scripts/r49-corpus/split_units.py. The inherited rule was "most units above a floor", which
is wrong for any book whose markers are PARTS:
Cities of the Plain 4 roman marks -> 4 units, median 22,312w
The Crossing 4 roman marks -> 4 units, median 37,310w
Four beats one, so it won, and the old guard only fired at exactly one unit. Now: a mode qualifies only if its median unit is inside [600, 12000] AND no unit holds half the work; among qualifying modes priority breaks the tie (contents > chapter-word > roman > bare-numeral > caps-title). Works with no divisions fall back to paragraph-block sections.
⭐⭐ The first version of that rule was WORSE than what it replaced, and a control caught
it. Scoring by "median closest to target" chose caps-title over the real chapters of
Hemingway's True at First Light:
bare-numeral 20 units median 5,337w max 11,155 <- the book's own chapters
caps-title 6 units median 777w max 113,886 <- median looked BETTER
Five stray all-caps lines gave five tiny units beside one holding 97% of the book. A median cannot see that distribution; a max bound can. Controls green both ways afterwards: 8/10 Hemingway works reproduce the shipped mode and count exactly, and 40,000 words with no blank lines returns 1 unit rather than fabricating sections.
⚠ The Hemingway builder is deliberately NOT repointed at this module — its corpus is shipped and its sha is pinned by a live adapter.
D1: the job was protecting a style that reads as damage
quote marks 0.0 per 10k (Hemingway 838)
apostrophes 123 per 10k (Hemingway 241) `dont` `aint` `wont` `didnt`
⚠⚠ repair_typography.py MUST NOT be run on this corpus. It normalises "toward what the
text does" and would put the quotation marks back. The builder runs no normalisation and then
asserts the quote density, so a future well-meaning change fails the build.
⚠ AND IT MAKES THE VOICE GATE EASY TO PASS FOR THE WRONG REASON. voice_distance.py is
Burrows's Delta over CHARACTER BIGRAMS. An adapter that learns only "emit no quotation marks"
moves delta_cb a long way without having learned a sentence. Pre-register a
punctuation-normalised secondary read before gating lv-mccarthy. Tracked in the builder
docstring, commit f3bf3ca.
Also: two truncated catalogue rows dropped for complete mobi siblings; all 15 cross-work containment pairs measured (worst 0.10%); back matter in 4 of 6 works carrying the author's name 26 times → 0; alphabet re-derived at 1,411 non-ASCII letters across 14 Spanish forms.
⚠ The back-matter strip runs BEFORE the split for McCarthy, inverting the Hemingway order: Blood Meridian and The Crossing end with a dumped table of contents made of bare roman numerals on their own lines — the exact shape of a chapter marker.
D2 caught a D1 defect: three small-caps manglings
The entity map returned E, H, T, K as renameable entities with 17–33 capitalised
occurrences each — the G class from Hemingway, where G was about to be renamed to a surname
248 times. Reading them showed the extractor mangled small-caps openings three ways:
1. SPLIT INITIAL `T HE HOUSE was built` -> `The house was built` 32 cases
2. UNMARKED RUN `THEY STOOD in the doorway` -> `They stood in the doorway` 88 cases
3. LOST INITIAL `HE CANDLEFLAME` -> `THE CANDLEFLAME` 1 case
Rule 1 requires a FOLLOWING all-caps word, so A TV was playing and A Mexican was changing
are untouched. Rule 2's [a-z] lookahead is what makes it safe — a genuine shout or sign is
not followed mid-sentence by lowercase. All 23 distinct first words of the 88 were checked.
⚠⚠ A fourth "fix" was nearly shipped that would have CORRUPTED the text. HEY RODE →
THEY RODE looked right from a survey of the BUILT corpus. The raw master has THEY RODE
intact, twice — HEY RODE matched as a SUBSTRING, and the unanchored replace produced
TTHEY RODE, which rule 2 then lowercased to Tthey rode. Caught by the count assertion
(expected 1, replaced 2) and settled by reading the master. ⚠ My first corruption check also
missed it, searching for TTHEY when the pipeline had already lowercased it — check the
shape the pipeline emits, not the shape you imagined.
D2's own gates: and audit_stoplist was scanning its own rationale
⚠⚠ A defect in audit_stoplist.py, latent for every corpus before this one. It built its
surface set from every list value in the stoplist JSON — including _why, which by convention
is a LIST OF PROSE LINES. Its empty separator line matched the honorific pattern 139 times,
printing a flag with no surface name above the one real catch. Now skips _-prefixed keys.
That real catch was a contradiction inside my own file: Franklin sat in the geography list
(the old name for El Paso) while the same file's note recorded "I'm here to see Mr Franklin",
a lawyer in All the Pretty Horses. A second self-inflicted one: a speculative A–Z fragments list
stoplisted I and A, and Sir I dont think I can do that duly tripped the audit. It is now
the four letters actually measured as entities.
Everything ambiguous was read in context: Socorro is the ranch cook, not the New Mexico
town; Niño, Keno and Redbo are HORSES (renameable, the Inglés precedent); Yaqui and Gilenos
are real peoples; Hashknives is a real cattle outfit; Hearst, Trias, Huerta and Madero are real
historical figures on the page under their own names.
Final: 123 map surfaces, 124-surface stoplist, entities.py 27/27 controls, both audits PASS.
The human gender pass is an auditable file
The honorific/window resolver scored 21 correct / 3 held / 1 WRONG against a 26-name
control; the base-rate proximity resolver built for Hemingway scored 18/6/1 and its own guard
correctly REFUSED to write. So the incumbent stands and four entries are fixed by hand in
gender_overrides_mccarthy.json, each carrying its evidence.
⚠ All four are female and all four look male-dominated in raw counts, because this corpus runs 29,144 male pronouns to 5,036 female — a base rate of 85.3% male. Carla Jean Moss at 31m/21f would be 44m/8f at that rate; 21 against an expected 8 is decisive. Same arithmetic that recovered Pilar and Brett on Hemingway. Alfonsa was in the control and is correctly absent from the map at 4 occurrences, below the min-count floor — an error in the control, not the pipeline.
apply_gender_overrides.py refuses twice: a name absent from the map is an error rather than a
silent no-op, and overruling a gender the detector holds needs an explicit "correcting": true
so it cannot look like filling a held entity in a diff.
D3: three calls, and the holdout fix that matters most
--scope corpus, not the per-work default. Nine surfaces appear in more than one work — Parham (The Crossing + Cities of the Plain), Grady and Cole (All the Pretty Horses + Cities of the Plain), Socorro, Héctor. A per-work map gives John Grady a different invented name in each novel, turning one character into two.- A new
mccarthypreset. Hemingway's romance pool carriesit_ITandfr_FRfor his Italian and French casts; McCarthy writes neither language.en_GBgoes for the same reason.en_US+es_MX/es_ESat an even share. --min-cap 5to match the entity map's floor. The first gate run FAILED with 45 survivors:entities.pyadmits cap ≥ 5 whilerename.pyrenamed only cap ≥ 8, so every entity between sat in the map, was never renamed, and counted as a leak. Hemingway never hit it because its map hadsub_threshold_total: 0.
⭐ --holdout-chapter NOW TAKES A LIST. The val split is one chapter index per work, so its
SIZE is set by how many WORKS a corpus has, not how many words:
Hemingway 10 works -> 9 val units -> 36,563 words/copy -> gate DECISIVE
Brontë 4 works -> 4 val units -> 17,043 words/copy -> gate MARGINAL
McCarthy 6 works -> 6 val units -> ~18,000 would have been Brontë's end
Holding out chapters 7 and 17 gives 11 units and 40,653 words per copy — larger than Hemingway's — for 7% of the corpus, on a corpus 40% smaller than his. No amount of corpus size fixes a val split that scales with work count.
Related: 2026-09-17-lv-hemingway-gate, 2026-09-17-mccarthy-krakauer-d1, 2026-09-17-lv-bronte-gate.