From 7964d077deb3c0f3d774ffbb2a14ebe6faa83bb1 Mon Sep 17 00:00:00 2001 From: Vuong Hoang Date: Wed, 16 Sep 2026 21:03:55 -0700 Subject: [PATCH] =?UTF-8?q?bronte-corpus:=20runbook=20=E2=80=94=20the=20fi?= =?UTF-8?q?ve=20deviations=20and=20what=20the=20controls=20caught?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Records the reproducible chain and, more usefully, why it diverges from the Yarros/Hemingway pipeline in five places, each forced by a measurement rather than a preference. Includes the control post-mortem, which is worth keeping because in three of four cases the CONTROL was wrong and the detector was right — the opposite of the reflex. Adele vs Adele-with-a-grave, Hollow at a 0.235 lowercase ratio, and Grace at 0.224 were all correct refusals. Blanche, at 0.0526 against a 0.05 bar, was the one real detector miss. --- scripts/bronte-corpus/RUNBOOK.md | 158 +++++++++++++++++++++++++++++++ 1 file changed, 158 insertions(+) create mode 100644 scripts/bronte-corpus/RUNBOOK.md diff --git a/scripts/bronte-corpus/RUNBOOK.md b/scripts/bronte-corpus/RUNBOOK.md new file mode 100644 index 0000000..d5311ac --- /dev/null +++ b/scripts/bronte-corpus/RUNBOOK.md @@ -0,0 +1,158 @@ +# lv-bronte — corpus pipeline, end to end + +Charlotte Brontë: *Jane Eyre*, *Villette*, *Shirley*, *The Professor*. Four +unrelated novels, 680,291 words as built, **678,367 after repair**. Source is +Project Gutenberg, so unlike the Yarros and Hemingway lines there is no licensed +master and no distribution question. + +`$R` is this repo's `scripts/`. Everything below runs on **pfi-gx10** +(`10.100.50.60`), which also holds the carrier and does the training. + +Every stage refuses to write if its own controls fail, so a silent bad stage is +not a failure mode here. Two stages actually did refuse during the first run, and +both refusals were correct — see *What the controls caught*. + +```bash +# D1b — repair the Gutenberg source. NOT optional; see below. +python3 $R/bronte-corpus/repair_corpus_bronte.py ~/r49-corpus --out corpus-clean + +# D2 — entity map. +python3 $R/r49-corpus/entities.py corpus-clean --out entities.json \ + --stoplist $R/bronte-corpus/stoplist_bronte.json \ + --min-count 3 --min-mid-ratio 0.35 --fold-clitics --drop-acronyms \ + --control "Jane,Rochester,Thornfield,Lowood,Helen,Bessie,Adèle,Diana,Bertha,Lucy,Ginevra,Graham,Shirley,Caroline,Robert,Louis,Hortense,Yorke,Crimsworth,Frances,Hunsden,Pelet,Reed,Brocklehurst,Fairfax,Moore,Keeldar,Helstone,Malone,Sympson,Villette,Paulina,Terrasse,Vashti,Labassecour,Ferndean" \ + --negative-control "Indeed,Perhaps,However,Nevertheless,Certainly,Chapter,Yes,Well,Nothing,Something,Truly,Presently,Meantime,Again,Once,Besides,Moreover,Whether,Mdlle,Rue,Shakespeare,India,Dieu,Jesuit" + +# D2b — gender by proximity, scored against the CORPUS BASE RATE. +python3 $R/hemingway-corpus/gender_by_proximity.py corpus-clean \ + --entities entities.json --out entities-gender.json --control "Jane=f,Rochester=m,…" + +# D2c — pin hand-verified ground truth. Brontë-specific; see below. +python3 $R/bronte-corpus/pin_known_gender.py \ + --entities entities-gender.json --out entities-final.json --pin "Jane=f,Lucy=f,…,Moore=n,Helstone=n" + +# D3 — rename. NOTE --scope corpus and --min-cap 3, both deviations. +python3 $R/r49-corpus/rename.py corpus-clean --entities entities-final.json \ + --dictionary ~/r49-prep/name_dictionary.json --out corpus-renamed \ + --preset bronte --scope corpus --min-cap 3 --copies 6 --seed 4919 \ + --phrase-map $R/bronte-corpus/phrase_map_bronte.json + +# GATE — must pass before anything is trained. +python3 $R/r49-corpus/leak_gate.py corpus-clean --entities entities-final.json \ + --renamed corpus-renamed --min-cap 3 \ + --phrase-map $R/bronte-corpus/phrase_map_bronte.json \ + --report corpus-renamed/leak_gate_report.json + +# D4 — instruction pairs. --source-entities is MANDATORY here; see below. +python3 $R/yarros-corpus/build_sft_pairs.py --corpus corpus-renamed/copies \ + --out pairs/pairs-full.jsonl --register bronte --model gen \ + --source-entities entities-final.json --n 4093 --split train --seed 4919 +# …and the same with --split val --n 200 --out pairs/pairs-val.jsonl +``` + +Gate result, 2026-09-16: **0 of 365 renameable surfaces survive**, positive +control 365/365, negative control clean, phrase audit 0 of 82. Sensitivity floor: +a name under 3 capitals per work is never detected, a phrase under 5 recurrences +never audited. + +## Why this corpus needed its own pipeline + +Five deviations from the Yarros/Hemingway chain, each forced by a measurement. + +| deviation | why | +|---|---| +| **D1b repair exists at all** | 1,922 words of publisher back matter in Shirley's last unit and 1,368 Gutenberg `_italic_` spans. See below. | +| **`--min-count 3` / `--min-cap 3`** (was 8) | At 8, `Bertha`, `Ferndean`, `Rochesters` and `Creemsvort` sat below the floor — detected by nothing, renamed by nothing, reported by nothing. | +| **`--scope corpus`** (was `work`) | `Jane` is above threshold in *Jane Eyre* and *The Professor* but appears 5 times in *Shirley*. Work-scope renames it in two novels and prints it verbatim in the third. The gate is a whole-corpus scan and would have caught it; corpus scope means it never happens. | +| **D2c `pin_known_gender.py`** | The resolver puts `Jane` MALE across 336 occurrences. Structural, not tuning — see below. | +| **`--source-entities` on the pair build** | A leak class the gate cannot see — see below. | + +## The two source defects + +**Publisher back matter, Shirley only, 1,922 words.** A T. Nelson & Sons printer's +block, a `THE NELSON CLASSICS` catalogue advertising Ainsworth, Marryat, Verne, +Kingsley and Dickens, and a Gutenberg transcriber's punctuation list — all inside +the final chapter record. Not Brontë, and the origin of the entity `CHARLES` (12 +occurrences of another author's given name, in the training text). Structurally +identical to the Hemingway back-matter defect and it arises the same way: a +splitter cuts on headings and nothing follows the final one. The other three works +are clean, which the repair script's per-work report shows as an implicit control. + +**Gutenberg `_italic_` markup, 1,368 spans.** Two harms. It teaches the adapter to +emit underscores as prose. And `_` is a **word character**, so the gate's +`\b(Name)\b` scan cannot match inside `_Antigua_` — an entity that appears only in +italics is invisible to the gate, never renamed and never reported, which is the +same failure shape as Yarros's possessive-only `Afendra`. Measured blast radius +here is exactly one entity, but that is a fact about this corpus, not a property of +the defect. ⚠ 24 of the spans wrap across a single newline (`_sotto\nvoce_`, +`_ignis\nfatuus_`); a newline-free pattern silently leaves exactly those behind, +and the script's read-back guard is what caught it. + +## Why gender needs a separate pinning stage + +The inherited honorific/window resolver scores 19 correct / 5 held / **5 wrong**. +Hemingway's base-rate resolver is strictly better on the axis that matters — 17 +correct / 11 held / **1 wrong** — because a wrong gender actively corrupts pronoun +agreement while a held one merely yields a neutral name. Brontë's corpus is nearly +pronoun-balanced (15,588 male / 16,144 female, base_m 0.491), unlike Hemingway's +near-4:1, so the base-rate correction does less work here but still helps. + +The one it still gets wrong is **`Jane`, 341 occurrences, the title character and +narrator**. This is not a tuning miss. Brontë's three narrators are first-person, +so their names appear almost only in *dialogue*, spoken by other characters, +surrounded by those characters' pronouns — proximity inference is structurally +blind to exactly the characters the adapter is being trained on. Yarros solved its +version with the POV chapter header; Brontë's editions have none, so that fix does +not transfer any more than it transferred to Hemingway. + +⚠ **Keep the resolver's evaluation and the shipped map separate.** `pin_known_gender.py` +does not improve the resolver's score and must never be quoted as if it had. It +applies ground truth a human read out of the novels. It also fails loudly on a pin +that matches no entity, which caught a stale `Hall=n` pin on the first run. + +Shared surnames are pinned **neutral** on purpose: `Moore` is Robert, Louis *and* +Hortense; `Helstone` is Matthewson, James and Caroline; `Yorke` is Hiram, Mrs +Yorke, Jessy, Rose and Martin. A gendered rename of a surname shared across sexes +is wrong in one direction no matter which you pick. + +## ⭐ The leak the gate cannot see + +**The corpus gate reads the corpus and the renamed copies. It never reads the +generated beats.** The beat is written by an LLM that just read the passage, and +if it recognises the book it supplies the canonical names from its own training. + +Measured on the first 714 Brontë pairs, before the filter existed: **13 beats +(1.8%) named source characters** — `Rochester` ×6, `Jane` ×3, `Brocklehurst` ×2, +`Beck`, `Fairfax`, `Helen`, `Burns`, `Eyre`, `Reed`, `Rivers` — while **0 of 714 +responses did**. The rename was perfect and the instruction side was not. One beat +read *"Saoirse confirms Rochester's flaws, then agrees in English to marry him"*, +mixing a renamed name and a canonical one in a single sentence. + +The beat is the INSTRUCTION half of the pair, so training on it re-teaches exactly +the inventions the rename exists to remove. + +⚠ **Exposure scales with how well the generator knows the book**, so it is worst +for public-domain classics and mildest for recent work. That is why Yarros and +Hemingway came up clean and Brontë did not — **their clean runs are not evidence +this cannot happen to them.** Pass `--source-entities` on every corpus, and rebuild +those two with it if they are ever regenerated. + +## What the controls caught + +Worth recording, because in three of four cases the control was wrong and the +detector was right — the opposite of the reflex. + +- **`Adele` NOT DETECTED.** The corpus spells it `Adèle` (135×). My control string + was ASCII. Detector right. +- **`Hollow` NOT DETECTED.** 85 capitalised against 20 genuinely lowercase + (`rushing down the hollow`), ratio 0.235. Detector right; `Hollow's Mill` belongs + in the phrase map, which is where it went. +- **`Grace` NOT DETECTED.** 0.224 lowercase ratio — `grace` is a common noun. + Detector right. But sampling all 21 bare capitalised `Grace` found 20 are the + character in direct address and one is theological, so it is token-mapped, with + that single collateral substitution accepted and recorded. +- **`Blanche` NOT DETECTED — and here the detector was wrong.** 19 capitalised + against **one** lowercase gives 0.0526, over the 0.05 bar by a single token. A + named character with 19 mentions, dropped by a hair. Token-mapped. +- **`Antigua` missing from the positive control at gate time.** This is what + exposed the italics defect above.