~/lv-mccarthy on pfi-gx10: corpus-clean, corpus-renamed (6 copies, 1,002 records), scripts.
leak gate 0 of 75 renameable and 0 of 37 sub-threshold survive in any copy
positive control 108/108 surfaces found in the unrenamed source
negative control nonce absent from both trees
THREE McCARTHY-SPECIFIC DECISIONS, each forced by a measurement.
1. --scope corpus, NOT the default per-work map. The Border Trilogy shares characters
across books -- 9 surfaces appear in more than one work, including Parham (The
Crossing + Cities of the Plain), Grady and Cole (All the Pretty Horses + Cities of
the Plain), Socorro and Héctor. A per-work map would give John Grady a different
invented name in each novel, turning one character into two.
2. A NEW `mccarthy` rename preset rather than reusing `hemingway`. Both are
Spanish-inflected, but Hemingway's romance pool carries it_IT and fr_FR for his
Italian and French casts, and McCarthy writes neither language -- drawing from it
would drop Italian and French surnames into a Texas-Mexico border novel. en_GB goes
for the same reason. en_US + es_MX/es_ES at an even share.
3. --min-cap 5 to MATCH the entity map's floor. The first gate run FAILED with 45
survivors, and the diagnosis is the Brontë lesson exactly: entities.py admits
cap >= 5 while rename.py only renamed cap >= 8, so every entity between 5 and 7 sat
in the map, was never renamed, and was counted as a leak. Hemingway never hit it
because its map had sub_threshold_total 0.
⭐ --holdout-chapter NOW TAKES A LIST, and this is the change with the most downstream
effect. The val split is one chapter index per work, so its SIZE is set by how many
WORKS a corpus has, not how many words:
Hemingway 10 works -> 9 val units -> 36,563 words/copy -> gate DECISIVE
Brontë 4 works -> 4 val units -> 17,043 words/copy -> gate MARGINAL
McCarthy 6 works -> 6 val units -> ~18,000 would have been Brontë's end of that
Holding out chapters 7 AND 17 gives 11 units and 40,653 words per copy -- larger than
Hemingway's, at a cost of 7% of the corpus -- on a corpus 40% smaller than his. No
amount of corpus size fixes a val split that scales with work count.
THE HUMAN GENDER PASS IS NOW AN AUDITABLE FILE, not a hand edit. The honorific/window
resolver scored 21 correct / 3 held / 1 WRONG against a 26-name control; the base-rate
proximity resolver built for Hemingway scored 18/6/1 and its own guard correctly
REFUSED to write. So the incumbent stands and four entries are fixed by hand in
gender_overrides_mccarthy.json, each carrying its evidence.
⚠ All four are female and all four look male-dominated in raw pronoun counts, because
this corpus runs 29,144 male pronouns to 5,036 female -- a base rate of 85.3% male.
Carla Jean Moss at 31m/21f would be 44m/8f at that base rate, so 21 female against an
expected 8 is decisive. Same arithmetic that recovered Pilar and Brett on Hemingway.
Alfonsa was in my control set and is correctly absent from the map at 4 occurrences,
below the min-count floor -- an error in the control, not the pipeline.
apply_gender_overrides.py refuses two ways: a name absent from the map is an error
rather than a silent no-op, and overruling a gender the detector already holds needs
an explicit "correcting": true so it cannot look like filling a held entity in a diff.
87 lines
4.0 KiB
Python
87 lines
4.0 KiB
Python
"""Apply hand-verified genders to an entity map — the human pass the pipeline asks for.
|
|
|
|
`entities.py` says it plainly: "Nothing here guesses. Unresolved entities block corpus
|
|
emission and go to a human pass: held is cheap, wrong is poison -- a silently mis-gendered
|
|
entity scrambles pronoun agreement through every renamed copy and nothing downstream would
|
|
catch it." This is that pass, written down instead of typed into a JSON by hand.
|
|
|
|
⚠ RAW PRONOUN COUNTS ARE THE TRAP, AND THIS IS WHY EVERY OVERRIDE CARRIES ITS EVIDENCE.
|
|
Measured on McCarthy: the corpus runs 29,144 male pronouns to 5,036 female, a base rate of
|
|
**85.3% male**. So a name sitting at 31 male / 21 female nearby is not "male-dominated" — at
|
|
the base rate it would be 44/8, and 21 female against an expected 8 is a strong FEMALE
|
|
signal. Carla Jean Moss was read male by the honorific/window resolver for exactly that
|
|
reason. The same arithmetic recovered Pilar and Brett on Hemingway.
|
|
|
|
TWO REFUSALS, because an override file is a place where a typo is invisible:
|
|
|
|
* a name not present in the entity map is an ERROR, not a no-op. A silent skip means a
|
|
misspelled override looks like it applied and the entity stays mis-gendered.
|
|
* changing a gender the map already holds requires `"correcting": true` on that entry.
|
|
Filling a HELD entity is the ordinary case; overruling the detector is not, and the two
|
|
should not look the same in a diff.
|
|
"""
|
|
from __future__ import annotations
|
|
import argparse, json
|
|
from pathlib import Path
|
|
|
|
|
|
def main() -> int:
|
|
ap = argparse.ArgumentParser()
|
|
ap.add_argument("--entities", required=True)
|
|
ap.add_argument("--overrides", required=True,
|
|
help='JSON: {"<Surface>": {"gender": "f", "why": "...", '
|
|
'"correcting": true}} — `why` is required, `correcting` only when '
|
|
'the map already holds a different gender')
|
|
ap.add_argument("--out", required=True)
|
|
a = ap.parse_args()
|
|
|
|
ents = json.loads(Path(a.entities).read_text())
|
|
blob = json.loads(Path(a.overrides).read_text())
|
|
ov = {k: v for k, v in blob.items() if not k.startswith("_")}
|
|
|
|
# surface -> [(work, key)]
|
|
where: dict[str, list[tuple[str, str]]] = {}
|
|
for w, blk in ents.items():
|
|
for k, e in blk["entities"].items():
|
|
where.setdefault(e.get("surface") or k, []).append((w, k))
|
|
|
|
missing = [n for n in ov if n not in where]
|
|
if missing:
|
|
print(f"== REFUSING: {len(missing)} override(s) name no entity in the map: {missing}")
|
|
print(" A misspelled override that silently does nothing leaves the entity")
|
|
print(" mis-gendered AND looks like it was handled.")
|
|
return 1
|
|
|
|
no_why = [n for n, v in ov.items() if not (v.get("why") or "").strip()]
|
|
if no_why:
|
|
print(f"== REFUSING: no `why` on {no_why}. An override without its evidence is a guess.")
|
|
return 1
|
|
|
|
changed = filled = 0
|
|
for name, spec in sorted(ov.items()):
|
|
g = spec["gender"]
|
|
for w, k in where[name]:
|
|
cur = ents[w]["entities"][k].get("gender")
|
|
if cur == g:
|
|
continue
|
|
if cur and not spec.get("correcting"):
|
|
print(f"== REFUSING: {name} in {w} already reads {cur!r} and the override says "
|
|
f"{g!r} without \"correcting\": true. Overruling the detector is not the "
|
|
f"same act as filling a held entity.")
|
|
return 1
|
|
ents[w]["entities"][k]["gender"] = g
|
|
ents[w]["entities"][k]["gender_source"] = "hand-verified"
|
|
changed += cur is not None
|
|
filled += cur is None
|
|
print(f" {name:<14} {w:<26} {str(cur):>6} -> {g} "
|
|
f"{'CORRECTION' if cur else 'filled held'}")
|
|
|
|
Path(a.out).write_text(json.dumps(ents, ensure_ascii=False, indent=1), encoding="utf-8")
|
|
print(f"\n {filled} held entity/entities filled, {changed} detector reading(s) corrected")
|
|
print(f" wrote {a.out}")
|
|
return 0
|
|
|
|
|
|
if __name__ == "__main__":
|
|
raise SystemExit(main())
|