BabyYarros: corpus built, gender resolution fixed, rename blocked on leak gate

Located the source: five Rebecca Yarros works in the Kvasir licensed library, with
rights recorded as gated. Built D1 at 208 chapters and 780,744 words, which is 15%
larger than the Brontë corpus. No unwrap step was needed because Kvasir's cleaner
already emits flowing paragraphs, so the hard-wrap defect that cost a re-cut on
Brontë does not exist here. The alphabet was re-derived rather than inherited: 23
non-ASCII letters across three forms, against F02's 4 on a smaller sample. Same
ASCII-fold conclusion from a different measurement, which is the reason to re-derive
per corpus.

The interesting finding is a new pathology. In a rotating first-person POV corpus,
every book's narrator gets the wrong gender. Measured against six names verified in
the text, the pronoun resolver called Violet male, Leah male and Landon female --
three of eighteen wrong, and all three are the narrator of the book where they were
misgendered. A narrator is "I" in her own book, so her name appears mostly inside
the other lead's dialogue surrounded by his pronouns. This is Brontë's "Jane called
male" amplified by rotating POV. Title-first resolution, which fixed it for Brontë,
is nearly blind here because contemporary romance uses given names rather than
honorifics. What works is the POV header: resolve each name from the chapters it
does not narrate. Validated at 9 correct, 9 held, 0 wrong against the previous 7, 8
and 3 wrong, and the instrument refuses to write unless it beats what it replaces.

Re-pointing rename.py surfaced three bugs, two of which would have silently
corrupted the corpus. Gender came only from honorifics and the entities file's
gender field was ignored, so the POV fix had no effect until wired through; that
took wilder from 1 gendered entity to 13. The pool labels were hardcoded in a print
statement, so any non-Brontë preset crashed. And the collision-filter log claimed
it dropped names colliding with Brontë entities regardless of which corpus it
filtered against -- the logic was right but the message named the wrong corpus,
which is how a reader later concludes the filter ran on the wrong thing.

D3 is blocked and nothing has been trained. The leak gate shows 86 of 232
renameable source entities surviving where the Brontë run reached 0 of 203. It
decomposes into detector false positives that need a stopword filter rather than
renaming, genuine misses among worldbuilding proper nouns, and a third class whose
cause is not yet established. Training before the gate passes means fitting
in-copyright text with 86 identifiable source entities intact, in a corpus F02
already flagged as small enough for leak to be a real concern.
This commit is contained in:
vh
2026-09-11 08:46:45 -07:00
parent e15c5ee5ea
commit 6dba912324
6 changed files with 408 additions and 10 deletions
@@ -0,0 +1,128 @@
"""D1 for BabyYarros: build the corpus from the licensed Kvasir masters.
Deliberately emits the SAME record schema as the Brontë builder
({work, chapter, heading, words, text} per work file, plus corpus_alphabet.json and
manifest.json), so entities.py, rename.py and train_voice_lora.py all run unchanged.
Matching an existing schema beats teaching three downstream tools a new one.
Differences from the Brontë build, and each is a property of the source rather than
a preference:
* NO Gutenberg boilerplate strip and NO download -- Kvasir already extracted and
cleaned these, and the catalog records the cleaner and its version.
* NO unwrap step. The Brontë corpus came hard-wrapped at ~70 characters and the
adapter learned the line breaks; these masters are already flowing paragraphs
(median non-blank line 102 chars), so the defect does not exist here.
* Chapters are "Chapter One" style words, not roman numerals, and are followed by
a POV name and often a location on their own lines -- first-person contemporary
romance with rotating narrators. Those header lines are KEPT: they are part of
the form the voice lives in, and dropping them would teach the model that
chapters begin mid-scene.
* ASCII alphabet, confirmed on the text rather than inherited from F02: 2
non-ASCII letters across this corpus. Under the F02 rule (a rename pool's
character inventory must be a SUBSET of the corpus's) that means an ASCII-only
pool -- the opposite of Brontë, who needed French accents kept.
"""
from __future__ import annotations
import argparse, collections, json, os, re, sqlite3, sys
from pathlib import Path
CATALOG = "/home/lkraven/development/kvasir/data/library/catalog.sqlite"
KVASIR = "/home/lkraven/development/kvasir"
SLUGS = {
"Fourth Wing (Exclusive Holiday Edition)": "fourth-wing",
"Iron Flame": "iron-flame",
"Wilder (The Renegades)": "wilder",
"Nova (The Renegades #2)": "nova",
"Rebel (The Renegades)": "rebel",
}
# "Chapter One" / "Chapter Twenty-Three" / "Chapter 12" / "Prologue" / "Epilogue".
CHAPTER = re.compile(
r"^[ \t]*((?:Chapter|CHAPTER)[ \t]+(?:[A-Za-z-]+|\d+)|Prologue|PROLOGUE|Epilogue|EPILOGUE)"
r"[ \t]*\.?[ \t]*$", re.M)
def masters():
c = sqlite3.connect(CATALOG)
rows = c.execute("select title, master_path, rights, normalized_text_sha256 "
"from masters where lower(author) like '%yarros%'").fetchall()
out = []
for title, path, rights, sha in rows:
p = Path(path if os.path.isabs(path) else os.path.join(KVASIR, path))
if not p.exists():
print(f" ⚠ MISSING master for {title}: {p}", file=sys.stderr)
continue
out.append({"title": title, "slug": SLUGS.get(title, re.sub(r"\W+", "-", title.lower()).strip("-")),
"path": p, "rights": rights, "sha256": sha})
return sorted(out, key=lambda w: w["slug"])
def split_chapters(text: str):
"""Return [(heading, body)]. Everything before the first heading is front matter."""
marks = [(m.start(), m.group(1).strip()) for m in CHAPTER.finditer(text)]
if not marks:
return [("(whole)", text.strip())]
out = []
for i, (pos, head) in enumerate(marks):
end = marks[i + 1][0] if i + 1 < len(marks) else len(text)
body = text[pos:end].strip()
if len(body.split()) >= 150: # skip a bare heading with no chapter behind it
out.append((head, body))
return out
ap = argparse.ArgumentParser()
ap.add_argument("--out", required=True)
ap.add_argument("--survey", action="store_true", help="report and write nothing")
a = ap.parse_args()
works = masters()
if not works:
raise SystemExit("REFUSING: no Yarros masters resolved from the catalog")
out = Path(a.out)
alphabet = collections.Counter()
total_words = total_chaps = 0
manifest = {"corpus": "BabyYarros", "author": "Rebecca Yarros",
"source": "kvasir data/library masters (licensed, rights=gated)",
"built_at": __import__("datetime").date.today().isoformat(), "works": []}
for w in works:
text = w["path"].read_text(encoding="utf-8", errors="replace")
chaps = split_chapters(text)
alphabet.update(ch for ch in text if ch.isalpha())
words = sum(len(b.split()) for _, b in chaps)
total_words += words; total_chaps += len(chaps)
print(f" {w['slug']:14} {len(chaps):>3} chapters {words:>7,} words rights={w['rights']}")
manifest["works"].append({"slug": w["slug"], "title": w["title"], "rights": w["rights"],
"master_sha256": w["sha256"], "chapters": len(chaps), "words": words,
"path": f"works/{w['slug']}.jsonl"})
if not a.survey:
(out / "works").mkdir(parents=True, exist_ok=True)
with (out / "works" / f"{w['slug']}.jsonl").open("w", encoding="utf-8") as fh:
for i, (head, body) in enumerate(chaps, 1):
fh.write(json.dumps({"work": w["slug"], "chapter": i, "heading": head,
"words": len(body.split()), "text": body},
ensure_ascii=False) + "\n")
non_ascii = {c: n for c, n in alphabet.items() if ord(c) > 127}
print(f"\n TOTAL {total_chaps} chapters · {total_words:,} words · {len(alphabet)} distinct letters")
print(f" non-ASCII letters: {sum(non_ascii.values())} across {len(non_ascii)} forms {non_ascii or ''}")
manifest["totals"] = {"chapters": total_chaps, "words": total_words,
"distinct_letters": len(alphabet), "non_ascii_letters": sum(non_ascii.values())}
manifest["total_words"] = total_words
manifest["total_chapters"] = total_chaps
if not a.survey:
(out / "manifest.json").write_text(json.dumps(manifest, indent=2), encoding="utf-8")
(out / "corpus_alphabet.json").write_text(json.dumps({
"derived_from": "Rebecca Yarros, 5 novels, Kvasir licensed library",
"derived_at": manifest["built_at"],
"note": ("R49 F02 rule: a rename pool's character inventory must be a SUBSET of this. "
f"Measured on the built text: {sum(non_ascii.values())} non-ASCII letters, so the "
"pool is ASCII-only -- the opposite of the Brontë corpus, which needed French "
"accents kept."),
"letters": sorted(alphabet), "non_ascii": {c: n for c, n in sorted(non_ascii.items())},
}, indent=2, ensure_ascii=False), encoding="utf-8")
print(f" wrote {out}")
@@ -0,0 +1,67 @@
{
"derived_from": "Rebecca Yarros, 5 novels, Kvasir licensed library",
"derived_at": "2026-09-11",
"note": "R49 F02 rule: a rename pool's character inventory must be a SUBSET of this. Measured on the built text: 23 non-ASCII letters, so the pool is ASCII-only -- the opposite of the Brontë corpus, which needed French accents kept.",
"letters": [
"A",
"B",
"C",
"D",
"E",
"F",
"G",
"H",
"I",
"J",
"K",
"L",
"M",
"N",
"O",
"P",
"Q",
"R",
"S",
"T",
"U",
"V",
"W",
"X",
"Y",
"Z",
"a",
"b",
"c",
"d",
"e",
"f",
"g",
"h",
"i",
"j",
"k",
"l",
"m",
"n",
"o",
"p",
"q",
"r",
"s",
"t",
"u",
"v",
"w",
"x",
"y",
"z",
"à",
"é",
"ï"
],
"non_ascii": {
"à": 2,
"é": 19,
"ï": 2
}
}
+61
View File
@@ -0,0 +1,61 @@
{
"corpus": "BabyYarros",
"author": "Rebecca Yarros",
"source": "kvasir data/library masters (licensed, rights=gated)",
"built_at": "2026-09-11",
"works": [
{
"slug": "fourth-wing",
"title": "Fourth Wing (Exclusive Holiday Edition)",
"rights": "gated",
"master_sha256": "606420acb827122d700eb47c18b7612399d130fe770787b139b0704673b1b897",
"chapters": 41,
"words": 191289,
"path": "works/fourth-wing.jsonl"
},
{
"slug": "iron-flame",
"title": "Iron Flame",
"rights": "gated",
"master_sha256": "e66db0cd13789bb0d6065888bc117362c8b3c25f8827dcbc6ffcd452a7359af6",
"chapters": 66,
"words": 251949,
"path": "works/iron-flame.jsonl"
},
{
"slug": "nova",
"title": "Nova (The Renegades #2)",
"rights": "gated",
"master_sha256": "ccccf3d64dd5810c5135ac86223e5f3e679fe5d1cdacd88df1eb9ff0164cb61e",
"chapters": 34,
"words": 109571,
"path": "works/nova.jsonl"
},
{
"slug": "rebel",
"title": "Rebel (The Renegades)",
"rights": "gated",
"master_sha256": "ed61fea84eab962cbf4c96870eaaa180d5ea92278241ad54ebd6d9f6ae2c4e7d",
"chapters": 36,
"words": 118270,
"path": "works/rebel.jsonl"
},
{
"slug": "wilder",
"title": "Wilder (The Renegades)",
"rights": "gated",
"master_sha256": "c798c9a5d24595deb870e25c34478172cdfd7758249e41e6ecbfa8240cf51a41",
"chapters": 31,
"words": 109665,
"path": "works/wilder.jsonl"
}
],
"totals": {
"chapters": 208,
"words": 780744,
"distinct_letters": 55,
"non_ascii_letters": 23
},
"total_words": 780744,
"total_chapters": 208
}
+115
View File
@@ -0,0 +1,115 @@
"""Fix gender resolution for a rotating first-person POV corpus.
Neither existing method works on Yarros, and they fail for opposite structural
reasons:
* TITLE-FIRST (what Brontë needed) finds almost nothing -- 3 gendered entities per
work. Contemporary romance does not say "Miss Sorrengail", it says "Violet".
* PRONOUN PROXIMITY is wrong specifically on the people who matter most. Measured
against six names whose gender I verified in the text: 3 of 18 WRONG, and the
three are Violet, Leah and Landon -- each of them the first-person NARRATOR of
the book where they were misgendered. A narrator is "I" in her own book, so her
name appears mostly inside the other character's dialogue, surrounded by HIS
pronouns. This is the Brontë "Jane called male" pathology, and it is worse here
because Yarros rotates POV, so every book has a narrator set up to fail.
The signal this corpus actually offers is the POV header: chapters open "Chapter
One / Leah / Port of Miami", naming their narrator. So resolve each name's gender
from the chapters it does NOT narrate -- where other narrators refer to it in the
third person and the pronouns are trustworthy.
Refuses to write unless it beats the method it replaces on the verified control,
because a fix that is merely different is not a fix.
"""
from __future__ import annotations
import argparse, collections, json, re
from pathlib import Path
MASC = {"he", "him", "his", "himself"}
FEM = {"she", "her", "hers", "herself"}
# The POV name sits on its own short line just after the chapter heading.
HEAD = re.compile(r"^[ \t]*((?:Chapter|CHAPTER)[ \t]+(?:[A-Za-z-]+|\d+)|Prologue|Epilogue)"
r"[ \t]*\.?[ \t]*\n+[ \t]*([A-Z][A-Za-z'’-]{1,18})[ \t]*$", re.M)
ap = argparse.ArgumentParser()
ap.add_argument("corpus")
ap.add_argument("--entities", required=True)
ap.add_argument("--out", required=True)
ap.add_argument("--window", type=int, default=60, help="chars either side of a mention")
ap.add_argument("--min-hits", type=int, default=6)
ap.add_argument("--ratio", type=float, default=2.5)
ap.add_argument("--control", required=True, help="Name=g,Name=g -- verified in the text")
a = ap.parse_args()
corpus = Path(a.corpus)
man = json.loads((corpus / "manifest.json").read_text())
ents = json.loads(Path(a.entities).read_text())
truth = dict(p.split("=") for p in a.control.split(","))
chapters: dict[str, list[tuple[str | None, str]]] = {}
for w in man["works"]:
rows = [json.loads(l) for l in (corpus / w["path"]).read_text(encoding="utf-8").splitlines() if l.strip()]
out = []
for r in rows:
m = HEAD.search(r["text"][:400])
out.append((m.group(2) if m else None, r["text"]))
chapters[w["slug"]] = out
povs = collections.Counter(p for p, _ in out if p)
print(f" {w['slug']:14} {len(out):>3} chapters · POV headers found in "
f"{sum(1 for p, _ in out if p):>3} · narrators: {dict(povs.most_common(6))}")
def resolve(slug: str, name: str, exclude_own_pov: bool) -> str | None:
m = f = 0
for pov, text in chapters[slug]:
if exclude_own_pov and pov == name:
continue
for mt in re.finditer(rf"\b{re.escape(name)}\b", text):
ctx = text[max(0, mt.start() - a.window): mt.end() + a.window].lower()
for w in re.findall(r"[a-z]+", ctx):
if w in MASC: m += 1
elif w in FEM: f += 1
if m + f < a.min_hits:
return None
if m >= a.ratio * max(f, 1): return "m"
if f >= a.ratio * max(m, 1): return "f"
return None
def score(exclude: bool):
ok = wrong = held = 0
detail = []
for slug in chapters:
for key, ent in ents[slug]["entities"].items():
s = ent.get("surface") or key
if s not in truth:
continue
g = resolve(slug, s, exclude)
t = truth[s]
if g == t: ok += 1
elif g is None: held += 1
else: wrong += 1; detail.append(f"{slug}/{s}={g} (truth {t})")
return ok, held, wrong, detail
base_ok, base_held, base_wrong, base_d = score(False)
new_ok, new_held, new_wrong, new_d = score(True)
print(f"\n control, WITHOUT excluding own-POV chapters: {base_ok} correct · {base_held} held · {base_wrong} WRONG {base_d}")
print(f" control, EXCLUDING own-POV chapters: {new_ok} correct · {new_held} held · {new_wrong} WRONG {new_d}")
if new_wrong > base_wrong or (new_wrong == base_wrong and new_ok <= base_ok):
raise SystemExit("\n REFUSING to write: excluding own-POV chapters did not beat the "
"method it replaces on the verified control. A fix that is merely "
"different is not a fix.")
applied = 0
for slug in chapters:
for key, ent in ents[slug]["entities"].items():
g = resolve(slug, ent.get("surface") or key, True)
if g and g != ent.get("gender"):
applied += 1
if g:
ent["gender"] = g
Path(a.out).write_text(json.dumps(ents, indent=1), encoding="utf-8")
tot = sum(1 for w in ents.values() for e in w["entities"].values() if e.get("gender"))
print(f"\n wrote {a.out}: {applied} genders changed/added · {tot} entities now gendered")