fix(lv-mccarthy): the leak gate passed with five protagonist names still in every copy

`leak_gate.py` scans `\b(Surface)\b`. Any character inserted inside a name defeats
that pattern outright, so a mangled occurrence is unrenameable by rename.py AND
unreportable by the gate. lv-mccarthy's 2026-09-17 tree passed at "0 of 75
renameable and 0 of 37 sub-threshold" while carrying 13 occurrences of Bell,
Chigurh, Moss, Toadvine and Glanton in all six copies:

    B ell  C higurh  M oss  T oadvine    a small-caps drop cap kept as its own token
    Toad-vine  Glan-ton                  a print line-break hyphen kept by the extractor

Every visible occurrence HAD been renamed, which is what made the residue invisible
to a spot-read. Fixed at three levels, all three of which must stay:

  build_corpus_mccarthy.py rules 4 and 5 repair the source text — 32 split initials
  with a lowercase remainder, 5 hyphen-split names, each with an expected count so a
  master change fails the build. Rule 4's letter class is consonants only: `I` opens
  1,966 paragraphs, `A` 143 and `Y` 32 (Spanish `y`); folding any would corrupt 2,141
  lines to fix 32.

  leak_gate.py gains a separator-tolerant pass with its own positive and negative
  controls, and it FAILS the gate. Validated against the pre-fix tree: reports all
  five surfaces, exits 1. Its fragment filter is what makes it usable — a naive scan
  returns 18 false positives on Hemingway (`God damn`, `I run`) against 3 real ones;
  requiring one fragment to be a non-word of the corpus cleared all 18 and kept all 3.

  The whole D1→D3 chain is reproduced byte-identically before and after, so the fix
  is the only delta: 6 works, the entity map, the final map and all 36 copy files.

Cross-checked on the shipped corpora: lv-bronte is clean of this class, lv-hemingway
carries 3 (`Primi tivo`, `Pasionar ia`, `Chi cote`) and is live on fv-ml1.

Also in build_sft_pairs.py, both needed before lv-mccarthy's pairs:

  DEFECT 4, hard-wrap reflow. Measured on the SHIPPED lv-bronte adapter, which emits
  mid-sentence line breaks at 12.46 per 1k chars against 0.00 for its own base control
  and 0.00 for every Hemingway arm. McCarthy is the mixed case — The Road is wrapped,
  the other five works are not — so the corpus teaches the break as a coin flip. The
  obvious fix (join every interior newline) corrupts 46 two-speaker exchanges whose
  blank line was lost, and unmarked dialogue is the one thing this adapter exists to
  learn; the rule splits on sentence-final punctuation instead and takes the cheaper
  error. Self-targeting and off by default, so every shipped pair set is unchanged.

  A `mccarthy` register, which names the punctuation deliberately: the eval drives the
  base control arm with this same prompt, so tics left out of it are a surface trick
  only the adapter can perform, and delta_cb is a character-bigram measure.

  drop_leading_heading now also consumes Blood Meridian's dash-separated chapter
  arguments — 131 paragraphs, 0 in every other work of all three corpora.

And a RUNBOOK, because the D1→D3 session recorded nothing and the chain had to be
recovered by rebuilding candidates and matching sha256 against the artifacts on disk.
This commit is contained in:
vh
2026-09-17 11:37:33 -07:00
parent 4dce0d0a43
commit c55966433f
5 changed files with 452 additions and 10 deletions
+108
View File
@@ -0,0 +1,108 @@
# lv-mccarthy — corpus → gate → pairs
Six works, 167 units, 584,684 words. Built on **pfi-gx10** under `~/lv-mccarthy/`
as `infra-ops`, except **D1, which must run on nh3-dev**: the builder reads the
kvasir catalogue at `/home/lkraven/development/kvasir/data/library/catalog.sqlite`
and that path exists only there.
⚠ **This file exists because the 2026-09-17 D1→D3 session recorded nothing.** There
was no runbook, the commands were issued over non-interactive ssh so no shell history
survived, and a later session had to recover the whole chain by rebuilding candidates
and matching sha256 against the artifacts on disk. Every deviation below is now pinned
by a reproducibility control; keep it that way.
## The chain
```bash
R=~/development/eshpfi-management/scripts # nh3-dev for D1, ~/lv-mccarthy/scripts on gx10
# D1 — build. ON nh3-dev (needs the kvasir catalogue), then rsync corpus-clean/ to gx10.
python3 $R/mccarthy-corpus/build_corpus_mccarthy.py --out corpus-clean
# D2 — entity map. --min-mid-ratio is what keeps `Yeah`/`Buenas`/`Shh` out; see below.
python3 $R/r49-corpus/entities.py corpus-clean --out corpus-clean/entities.json \
--stoplist $R/mccarthy-corpus/stoplist_mccarthy.json \
--min-count 5 --fold-clitics --drop-acronyms --min-mid-ratio 0.2 --min-mid 2
# D2c — the four hand-verified genders.
python3 $R/r49-corpus/apply_gender_overrides.py --entities corpus-clean/entities.json \
--overrides $R/mccarthy-corpus/gender_overrides_mccarthy.json \
--out corpus-clean/entities-final.json
# D3 — rename. --scope corpus, --preset mccarthy and --min-cap 5 are all deviations.
python3 $R/r49-corpus/rename.py corpus-clean --entities corpus-clean/entities-final.json \
--dictionary ~/r49-prep/name_dictionary.json --out corpus-renamed \
--preset mccarthy --scope corpus --min-cap 5 --copies 6 --seed 4919 \
--holdout-chapter 7 17
# GATE — must pass before anything is trained. --min-cap 8 here is deliberate; see below.
python3 $R/r49-corpus/leak_gate.py corpus-clean --entities corpus-clean/entities-final.json \
--renamed corpus-renamed --min-cap 8 --report corpus-renamed/leak_gate_report.json
# D4 — instruction pairs. ON gx10; ~24 min against the `gen` seat.
./run-pairs.sh
```
Gate result, 2026-09-17 (after the separator fix): **0 of 75 renameable, 0 of 37
sub-threshold, 0 separator-split surfaces**, all four controls green. Sensitivity
floor: a name under 8 capitals per work is never detected, a phrase under 5
recurrences never audited, and **no phrase map exists for this corpus, so the phrase
audit does not run at all** — the Yarros and Brontë runs both had one.
## The deviations, and what forced each
| deviation | why |
|---|---|
| **D1 runs on nh3-dev** | The builder reads the kvasir catalogue by absolute path. gx10 has no copy. |
| **`--min-count 5`** (was 8) | The entity map's floor has to match rename's, or entities between the two floors sit in the map, are never renamed, and count as leaks. The first D3 gate run failed with 45 survivors for exactly this. |
| **`--min-mid-ratio 0.2 --min-mid 2`** | Without it the map admits `Yeah`, `Buenas`, `Shh`, `Mande`, `Sí`, `Git` — words that are only ever capitalised at a sentence start. ⚠ The map is INSENSITIVE to the exact value: any ratio in **[0.05, 0.3]** with `--min-mid` 1 or 2 reproduces it byte-for-byte. The original run's values are unrecoverable and it does not matter. `--min-mid 3` does NOT reproduce it. |
| **`--fold-clitics --drop-acronyms`** | Both are needed to reproduce the map; `--rescue-honorific` is inert here (0 rescued). |
| **`--scope corpus`** (was `work`) | Nine surfaces appear in more than one work — Parham, Grady, Cole, Socorro, Héctor. A per-work map gives John Grady a different invented name in each Border Trilogy novel. |
| **`--preset mccarthy`** | Hemingway's romance pool carries `it_IT`/`fr_FR` for his Italian and French casts; McCarthy writes neither. `en_US` + `es_MX`/`es_ES` at an even share. |
| **`--holdout-chapter 7 17`** (a LIST) | The val split is one chapter per work, so its size scales with WORK COUNT, not corpus size. Six works would have given a Brontë-class ~18,000-word val reference; two indices give 11 units and 40,653 words per copy, larger than Hemingway's, for 7% of the corpus. |
| **gate at `--min-cap 8`, rename at `--min-cap 5`** | Not a mistake. The gate reports the 37 entities rename never touched *separately*, which is strictly more informative than running both at 5 (where the sub-threshold bucket is empty). |
| **D4 `--reflow-hard-wraps`** | The Road is hard-wrapped at ~76 characters and the other five works are not. See DEFECT 4 in `build_sft_pairs.py`. |
| **D4 `--source-entities`** | Mandatory. The beat generator reads the passage and will supply canonical names from its own memory of the book; the corpus gate never reads the generated beats. |
## The defect the gate could not see, and now can
⚠⚠ **The 2026-09-17 tree passed this gate at "0 of 75 renameable and 0 of 37
sub-threshold" while carrying 13 occurrences of five protagonist names in all six
copies.** `leak_gate.py` scanned `\b(Surface)\b`, and a character inserted inside a
name defeats that pattern outright — so the occurrence was unrenameable by `rename.py`
*and* unreportable by the gate. Two mechanisms, both in the source extraction:
```
B ell C higurh M oss T oadvine a small-caps drop cap survived as its own token
Toad-vine Glan-ton a print line-break hyphen survived the extraction
```
Every *visible* occurrence had been renamed correctly, which is what made the residue
invisible to any spot-read: the names are gone everywhere you look.
Fixed in three places, and all three must stay:
1. **`build_corpus_mccarthy.py` rules 4 and 5** repair the text at its source — 32
split initials with a lowercase remainder, 5 hyphen-split names, both with expected
counts so a master change fails the build. ⚠ Rule 4's letter class is **consonants
only**: `I` opens 1,966 paragraphs, `A` opens 143 and `Y` opens 32 (Spanish *y*), and
folding any of them would corrupt 2,141 lines to fix 32.
2. **`leak_gate.py --split-frag-max`** runs a separator-tolerant scan every time, with
its own positive and negative controls, and **fails the gate**. Verified against the
pre-fix tree: it reports all five surfaces and exits 1.
3. The exact-match passes are unchanged, so the old verdict is reproduced alongside.
Cross-checked on the two shipped corpora: **lv-bronte is clean** of this class;
**lv-hemingway carries 3** (`Primi tivo`, `Pasionar ia`, `Chi cote`) and is live.
## Corpus properties worth knowing before you change anything
- **`repair_typography.py` MUST NOT be run on this corpus.** It normalises "toward what
the text does" and would put quotation marks back into a corpus that measures 0.0 per
10k against Hemingway's 838. The builder runs no normalisation and asserts the density.
- **Back matter is stripped BEFORE the split**, inverting the Hemingway order: Blood
Meridian and The Crossing end with a dumped table of contents made of bare roman
numerals, which is the exact shape of a chapter marker.
- **Blood Meridian sets a dash-separated chapter argument under each roman numeral.**
131 paragraphs; `--drop-leading-heading` now consumes them.
- Two truncated catalogue rows are excluded in favour of complete siblings.
@@ -135,6 +135,60 @@ LOST_INITIALS = {
"all-the-pretty-horses": [(re.compile(r"(?m)^HE CANDLEFLAME"), "THE CANDLEFLAME", 1)],
}
# ⚠⚠⚠ RULES 4 AND 5 ARE A LEAK FIX, NOT A TYPOGRAPHY FIX, AND THE GATE PASSED WITHOUT THEM.
#
# `leak_gate.py` scans `\b(Surface)\b`. ANY character inserted inside a name defeats that
# pattern outright, so a mangled occurrence is not merely unrepaired -- it is UNRENAMEABLE and
# UNREPORTABLE. Measured on the gated 2026-09-17 tree, which had reported "0 of 75 renameable
# and 0 of 37 sub-threshold survive in any copy":
#
# B ell x2 Chigurh cap 119 in the map no-country-for-old-men
# C higurh x3 Bell cap 147 no-country-for-old-men
# M oss x2 Moss cap 126 no-country-for-old-men
# T oadvine x1 Toadvine cap 111 blood-meridian
# Toad-vine x3 Glanton cap 365 blood-meridian
# Glan-ton x2
#
# 13 occurrences of FIVE protagonist names, present in ALL SIX renamed copies, with 0 visible
# survivors -- so the gate's verdict was true for the forms it can see and false overall. Every
# unmangled occurrence WAS renamed (exact-match survivors: 0), which is what makes the residue
# so easy to miss: the names are gone everywhere you look.
#
# Positive control on the probe that found it: the same separator-tolerant scan over the
# UNRENAMED source returns 115-365 hits per name, so it detects these names when present.
# Cross-check on the two SHIPPED corpora: lv-bronte is clean of this class; lv-hemingway carries
# 3 (`Primi tivo`, `Pasionar ia`, `Chi cote`). `leak_gate.py` now runs the scan itself.
#
# 4. SPLIT INITIAL, LOWERCASE REMAINDER `S ee the child` -> `See the child` 32 cases
# The drop cap survived as its own token and the small-caps remainder came through
# LOWERCASE, so rule 1 cannot see it -- rule 1 requires a following ALL-CAPS word, and
# DROPCAP requires two. This is the class both of them leave behind. Confined to
# blood-meridian (16) and no-country-for-old-men (16); the other four works have none.
#
# ⚠ THE VOWELS AND `Y` ARE EXCLUDED BECAUSE THEY ARE REAL WORDS AT A SENTENCE START, and
# this is the `I`/`A` trap that already bit audit_stoplist.py once. Measured in the built
# corpus: `I` opens 1,966 paragraphs (the pronoun), `A` opens 143 (the article, e.g.
# `A blivet is ten pounds of shit in a five pound sack.`), and `Y` opens 32 (Spanish `y`,
# e.g. `Y de los hombres?`). Folding any of those would corrupt 2,141 lines to fix 32.
# The letter class is consonants only, and all 32 survivors were read: the second
# fragments are `ell heir he hen hey higurh ith ive oadvine or oss ow e ee hen`.
#
# ⚠ Runs LAST, after DROPCAP. DROPCAP rewrites `T HE HOUSE was built` to `The house was
# built` -- the space between its groups is consumed, not captured -- so its output is
# never a rule-4 target. Verified: line-anchored and paragraph-anchored counts are both 32,
# i.e. no hit sits mid-paragraph on one of The Road's hard-wrapped lines.
#
# 5. HYPHEN-SPLIT NAME `Toad-vine` -> `Toadvine` 5 cases
# A print edition hyphenated the name across a line break and the extraction kept the
# hyphen while dropping the break. Patched BY NAME with an expected count, never by
# heuristic: McCarthy writes real hyphenated compounds and a rule broad enough to catch
# these would eat them. A master change makes the count wrong and says so.
SPLIT_INITIAL_LOWER = re.compile(r"(?m)^([B-DF-HJ-NP-TV-XZ]) ([a-z][a-z']*)")
HYPHEN_SPLIT_NAMES = {
"blood-meridian": [(re.compile(r"\bToad-vine\b"), "Toadvine", 3),
(re.compile(r"\bGlan-ton\b"), "Glanton", 2)],
}
def restore_smallcaps(line: str) -> str:
if len(SPLIT_INITIAL.findall(line)) < 2:
@@ -143,8 +197,9 @@ def restore_smallcaps(line: str) -> str:
def repair_smallcaps(text: str, slug: str) -> tuple[str, dict]:
"""Run the three repairs in order; the lost initials MUST go first."""
stats = {"lost_initial": 0, "split_initial": 0, "unmarked_run": 0}
"""Run the five repairs in order; the lost initials MUST go first, rule 4 MUST go last."""
stats = {"lost_initial": 0, "split_initial": 0, "unmarked_run": 0,
"split_initial_lower": 0, "hyphen_split_name": 0}
for pat, good, expect in LOST_INITIALS.get(slug, []):
text, n = pat.subn(good, text)
if n != expect:
@@ -157,6 +212,15 @@ def repair_smallcaps(text: str, slug: str) -> tuple[str, dict]:
lambda m: m.group(1) + m.group(2).lower(), text)
text, stats["unmarked_run"] = SMALLCAPS_OPENING.subn(
lambda m: m.group(1)[0] + m.group(1)[1:].lower(), text)
for pat, good, expect in HYPHEN_SPLIT_NAMES.get(slug, []):
text, n = pat.subn(good, text)
if n != expect:
print(f" ⚠ {slug}: expected {expect} occurrence(s) of {pat.pattern!r} and "
f"replaced {n} — the master changed; re-check before trusting this build",
file=sys.stderr)
stats["hyphen_split_name"] += n
text, stats["split_initial_lower"] = SPLIT_INITIAL_LOWER.subn(
lambda m: m.group(1) + m.group(2), text)
return text, stats
+20
View File
@@ -0,0 +1,20 @@
#!/bin/bash
# lv-mccarthy D4 — instruction pairs. Run on pfi-gx10 as infra-ops.
#
# Deviations from the Hemingway D4 recipe, each forced by a measurement:
# --register mccarthy new; it names McCarthy's punctuation deliberately, see REGISTERS
# --source-entities MANDATORY; Hemingway's pairs predate it
# --reflow-hard-wraps DEFECT 4; The Road is hard-wrapped and the other 5 works are not
# --drop-leading-heading now also drops Blood Meridian's dash-separated chapter arguments
# --n 3722 / --n 276 the full passage yield of each split, not a round number
set -e
cd ~/lv-mccarthy
mkdir -p pairs
S=scripts/yarros-corpus/build_sft_pairs.py
COMMON="--corpus corpus-renamed/copies --register mccarthy --model gen
--key-file $HOME/.config/litellm/all-agents-key
--source-entities corpus-clean/entities-final.json
--context-frac 1.0 --drop-leading-heading --reflow-hard-wraps --seed 4919"
python3 $S $COMMON --split train --n 3722 --out pairs/pairs-full.jsonl > pairs/train.log 2>&1
python3 $S $COMMON --split val --n 276 --out pairs/pairs-val.jsonl > pairs/val.log 2>&1
echo DONE > pairs/.complete