audit_entity_map: the rename can damage the prose and no gate will ever say so
audit_stoplist.py finds surfaces wrongly held OUT of the entity map -- a stoplisted character is an undetectable leak. This is the mirror: surfaces wrongly held IN it. leak_gate.py only ever asks whether the author's names are GONE, never whether non-names were spared, so renaming `the Chinese` into an invented surname passes it perfectly. Found sideways on Hemingway. The pairs audit reported beats naming African, Chinese, X-ray, Republican and Cezanne as leaks -- correctly, those surfaces really were removed from the corpus. Reading why turned up the larger defect: they should never have been renameable in the first place. Measured on the Hemingway map, both controls green: positive `other` 764/1356 article-preceded = 0.56 negative 100 honorific-confirmed people, highest Inglés at 0.26, bulk 0.00-0.06 FLAGGED 130 of 946 surfaces, 1,616 instances = 0.162% of corpus words The signal is an article in front of the surface: you write `the Frenchman` and `a Martini`, never `the Rinaldi`. It is a heuristic and every hit is reported FOR READING, never auto-removed -- `the Widow` and `the Informer` are genuine Hemingway epithet-names that SHOULD be renamed, and the band's own top entry makes the point, since Inglés at 0.26 is an in-world nickname deliberately kept renameable and sits just under the bar. Initials are excluded from the negative-control band rather than admitted to it. `Mr. P.` is an initial, not a person, so letting it in lets a map defect poison the control that validates the detector -- on Hemingway `P` (0.32, every occurrence `the P. O. U. M.`) was the one surface failing a band whose next highest was 0.26. Initials take no article and are invisible to the scan anyway, so every surface of two characters or fewer is now listed unconditionally. Sixteen of them are in this map, C at 274 occurrences; the same class as the `G` that was caught by hand about to be renamed to a surname 248 times. The unresolved count that drives the exit code is computed over every flagged surface, not the --show slice. Tying a gate's verdict to a display flag is the same defect as a log filter that turns a real event into a clean zero. Also corrects a wrong claim in audit_pairs_sourcenames.py's docstring: the Hemingway rename did not HOLD 591 surfaces. Paris, Madrid and Spain survive because the stoplist keeps them out of the entity map before it is built, so the map is exactly the removed set -- 941 surfaces, 941 removed, 0 kept. Measured per run rather than assumed, because a pipeline that carried kept surfaces into the map would report every `Paris` as a leak.
This commit is contained in:
@@ -0,0 +1,204 @@
|
||||
"""Audit an entity map for entries that are NOT characters — the mirror of audit_stoplist.py.
|
||||
|
||||
⭐ THE TWO INSTRUMENTS COVER OPPOSITE ERRORS AND NEITHER SEES THE OTHER'S.
|
||||
`audit_stoplist.py` finds surfaces wrongly held OUT of the map: a stoplisted character
|
||||
is an undetectable leak, because the gate never scans for what the stoplist removed.
|
||||
This finds surfaces wrongly held IN it: a demonym, a brand, a real place or a medical
|
||||
term that gets renamed into an invented proper noun. `leak_gate.py` is blind to that by
|
||||
construction — it only ever asks whether source names are GONE, never whether non-names
|
||||
were spared, so renaming `the Chinese` to an invented surname passes the gate perfectly.
|
||||
|
||||
Found on Hemingway 2026-09-17, and found sideways: the pairs audit reported beats naming
|
||||
`African`, `Chinese`, `X-ray`, `Republican` and `Cezanne` as leaks, which they are —
|
||||
those surfaces really were removed from the corpus. Reading why revealed the larger
|
||||
defect, that they should never have been renameable. Measured there: 130 of 941 surfaces
|
||||
flagged, 1,616 instances, 0.162% of corpus words.
|
||||
|
||||
THE SIGNAL: a definite or indefinite article in front of it. You write `the Frenchman`,
|
||||
`a Martini`, `the Republican`; you do not write `the Rinaldi`. Measured on Hemingway, a
|
||||
100-name honorific-confirmed band topped out at 0.26 and its bulk sat at 0.00–0.06, while
|
||||
the known non-names ran 0.29–0.96.
|
||||
|
||||
⚠ IT IS A HEURISTIC AND THE HITS MUST BE READ. Hemingway names characters by epithet —
|
||||
`the Widow`, `the Informer` — and those are genuine character designators that SHOULD be
|
||||
renamed. The detector cannot tell an epithet-name from a common noun, and it is blind in
|
||||
the other direction too: `Shakespeare` and `Cezanne` are real people who take no article,
|
||||
so they score 0.00 and this audit will never flag them. It narrows a 941-surface list to
|
||||
something a human can read; it does not decide. The band's own top entry says the same
|
||||
thing: `Inglés` at 0.26 is the gypsies' in-world nickname for Robert Jordan, deliberately
|
||||
kept renameable, and it sits just under the bar rather than safely away from it.
|
||||
|
||||
CONTROLS, both derived from the corpus rather than hand-picked, because a detector
|
||||
validated on names someone chose is validated on the choosing:
|
||||
* POSITIVE -- the corpus's own most frequent article-taking lowercase noun. The ratio
|
||||
function must score it high, or the whole instrument is measuring nothing.
|
||||
* NEGATIVE -- every map surface of 3+ characters that is ever preceded by an honorific
|
||||
is a person; their ratios must stay under the bar. If one flags, the detector is
|
||||
manufacturing signal on this corpus and its output is not usable. See the note at
|
||||
the control itself for why initials are excluded rather than admitted.
|
||||
|
||||
A SECOND CLASS THE SIGNAL CANNOT CARRY: initials. An initial takes no article, so `G` —
|
||||
which on this corpus was caught by hand about to be renamed to a surname 248 times, being
|
||||
only the fragment left by `B.G.` and `G.M.` — scores ~0.00 and is invisible here. Every
|
||||
map surface of two characters or fewer is therefore listed unconditionally for reading.
|
||||
|
||||
Exit 1 when anything is flagged and unread, so it can gate a pipeline.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
import argparse, json, re
|
||||
from collections import Counter
|
||||
from pathlib import Path
|
||||
|
||||
ART = re.compile(r"\b(?:the|a|an|The|A|An)\s+$")
|
||||
HONORIFIC = (r"(?:Mr|Mrs|Miss|Misses|Madame|Mdlle|Mademoiselle|Monsieur|Lord|Lady|Sir|Dr"
|
||||
r"|Doctor|Captain|Colonel|Major|General|Aunt|Uncle|Rev|Reverend|Professor"
|
||||
r"|Master|Saint|St)\.?\s+")
|
||||
|
||||
|
||||
def load_text(corpus: Path) -> str:
|
||||
man = json.loads((corpus / "manifest.json").read_text())
|
||||
parts = []
|
||||
for w in man["works"]:
|
||||
for line in (corpus / w["path"]).read_text(encoding="utf-8").splitlines():
|
||||
if line.strip():
|
||||
parts.append(json.loads(line)["text"])
|
||||
return "\n\n".join(parts)
|
||||
|
||||
|
||||
def article_ratio(text: str, surface: str) -> tuple[int, int, re.Match | None]:
|
||||
"""(occurrences, article-preceded, first article-preceded match)."""
|
||||
pat = re.compile(r"\b" + re.escape(surface) + r"\b")
|
||||
total = preceded = 0
|
||||
first = None
|
||||
for m in pat.finditer(text):
|
||||
total += 1
|
||||
if ART.search(text[max(0, m.start() - 6):m.start()]):
|
||||
preceded += 1
|
||||
if first is None:
|
||||
first = m
|
||||
return total, preceded, first
|
||||
|
||||
|
||||
def main() -> int:
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("corpus", help="source corpus dir (manifest.json + works/)")
|
||||
ap.add_argument("--entities", required=True, help="the entity map rename.py consumes")
|
||||
ap.add_argument("--min-ratio", type=float, default=0.30,
|
||||
help="flag a surface preceded by an article at least this often")
|
||||
ap.add_argument("--min-occ", type=int, default=5,
|
||||
help="below this, the ratio is one or two sentences and means nothing")
|
||||
ap.add_argument("--allow", default="",
|
||||
help="comma-separated surfaces already READ and confirmed to be genuine "
|
||||
"character designators (epithet-names); still shown, do not fail the run")
|
||||
ap.add_argument("--context", type=int, default=46)
|
||||
ap.add_argument("--show", type=int, default=60)
|
||||
a = ap.parse_args()
|
||||
|
||||
text = load_text(Path(a.corpus))
|
||||
ents = json.loads(Path(a.entities).read_text())
|
||||
surfaces = sorted({(e.get("surface") or k)
|
||||
for w in ents.values() for k, e in w["entities"].items()})
|
||||
cleared = {s.strip() for s in a.allow.split(",") if s.strip()}
|
||||
|
||||
# ---- POSITIVE CONTROL: the corpus's own favourite article-taking noun ----
|
||||
after_art = Counter(m.group(1) for m in re.finditer(r"\b(?:the|The)\s+([a-z]{3,})\b", text))
|
||||
if not after_art:
|
||||
print("== no `the <noun>` occurrences at all -- this is not English prose"); return 1
|
||||
probe, _ = after_art.most_common(1)[0]
|
||||
p_tot, p_pre, _ = article_ratio(text, probe)
|
||||
pos_ok = p_tot >= a.min_occ and (p_pre / p_tot) >= a.min_ratio
|
||||
print(f" [{'PASS' if pos_ok else 'FAIL'}] positive control: `{probe}` "
|
||||
f"{p_pre}/{p_tot} article-preceded = {p_pre / p_tot:.2f} "
|
||||
f"(must be >= {a.min_ratio})")
|
||||
|
||||
# ---- NEGATIVE CONTROL: honorific-confirmed people ----
|
||||
# ⚠ INITIALS ARE EXCLUDED FROM THE BAND, and not to make the control pass. An
|
||||
# honorific in front of a one- or two-character surface confirms nothing: `Mr. P.`
|
||||
# is an initial, not a person, so admitting it lets a map DEFECT poison the control
|
||||
# that is supposed to validate the detector. Measured on Hemingway, `P` (0.32, all
|
||||
# of them `the P. O. U. M.`) was the single surface failing a 103-name band whose
|
||||
# next highest was 0.26. Short surfaces are reported below as their own class
|
||||
# instead, because the article heuristic is mostly blind to them.
|
||||
people = []
|
||||
for s in surfaces:
|
||||
if len(s) >= 3 and re.search(HONORIFIC + re.escape(s) + r"\b", text):
|
||||
t, p, _ = article_ratio(text, s)
|
||||
if t >= a.min_occ:
|
||||
people.append((s, t, p / t))
|
||||
people.sort(key=lambda x: -x[2])
|
||||
if people:
|
||||
worst_s, _, worst_r = people[0]
|
||||
neg_ok = worst_r < a.min_ratio
|
||||
print(f" [{'PASS' if neg_ok else 'FAIL'}] negative control: {len(people)} "
|
||||
f"honorific-confirmed people (3+ chars), highest ratio {worst_s} at "
|
||||
f"{worst_r:.2f} (must be < {a.min_ratio})")
|
||||
else:
|
||||
neg_ok = True
|
||||
print(" [ -- ] negative control: no honorific-confirmed people in this map; "
|
||||
"the detector has no character band to be checked against here")
|
||||
if not (pos_ok and neg_ok):
|
||||
print("\n== CONTROLS FAILED -- the flags below are not trustworthy. Do not act on them.")
|
||||
|
||||
# ---- the measurement ----
|
||||
flagged, total_inst, flagged_inst = [], 0, 0
|
||||
for s in surfaces:
|
||||
t, p, first = article_ratio(text, s)
|
||||
total_inst += t
|
||||
if t >= a.min_occ and t and (p / t) >= a.min_ratio:
|
||||
flagged.append((t, p / t, s, first))
|
||||
flagged_inst += t
|
||||
flagged.sort(reverse=True, key=lambda x: x[0])
|
||||
|
||||
words = len(text.split())
|
||||
print(f"\n {len(surfaces)} map surfaces · {total_inst:,} instances · corpus {words:,} words")
|
||||
if not flagged:
|
||||
print(" [PASS] no map surface reads as a common noun")
|
||||
return 0
|
||||
print(f" FLAGGED: {len(flagged)} surfaces ({len(flagged) / len(surfaces):.1%}) · "
|
||||
f"{flagged_inst:,} instances ({flagged_inst / max(1, total_inst):.1%} of renamed text, "
|
||||
f"{flagged_inst / words:.3%} of corpus words)\n")
|
||||
|
||||
# ⚠ unresolved is counted over EVERY flagged surface, not the displayed slice.
|
||||
# Tying the exit code to --show would make a display flag decide whether the gate
|
||||
# passes, which is the same class of defect as a log filter that turns a real event
|
||||
# into a clean zero.
|
||||
unresolved = sum(1 for _, _, s, _ in flagged if s not in cleared)
|
||||
for n, r, s, m in flagged[:a.show]:
|
||||
mark = "cleared" if s in cleared else "⚠ READ THIS"
|
||||
ctx = text[max(0, m.start() - a.context): m.end() + a.context].replace("\n", " ") if m else ""
|
||||
print(f" {mark:<12} {s:<22} {n:>5} occ ratio {r:.2f} …{ctx}…")
|
||||
if len(flagged) > a.show:
|
||||
print(f" ... and {len(flagged) - a.show} more not shown (raise --show); "
|
||||
f"all {len(flagged)} count toward the gate")
|
||||
|
||||
# ---- SHORT SURFACES: a second class the article signal cannot carry ----
|
||||
# `G` was caught by hand on this corpus and would have been renamed to a surname 248
|
||||
# times -- it is the fragment left by `B.G.`, `G.M.`, `G2`. An initial takes no
|
||||
# article, so it scores ~0.00 and the scan above will never raise it. Listed
|
||||
# unconditionally, because the cost of renaming an initial is high and reading
|
||||
# sixteen lines is cheap.
|
||||
short = []
|
||||
for s in surfaces:
|
||||
if len(s) <= 2:
|
||||
t, _, _ = article_ratio(text, s)
|
||||
if t:
|
||||
short.append((t, s))
|
||||
if short:
|
||||
short.sort(reverse=True)
|
||||
shown = ", ".join(f"{s} ({t})" for t, s in short)
|
||||
print(f"\n SHORT surfaces in the map ({len(short)}), occurrences in brackets — an "
|
||||
f"honorific or an\n article cannot tell an initial from a name, so READ these "
|
||||
f"regardless of the scan:\n {shown}")
|
||||
|
||||
if unresolved:
|
||||
print(f"\n== {unresolved} map surface(s) read as common nouns and are not on --allow.")
|
||||
print(" Read each in context. A genuine epithet-name (`the Widow`) belongs in the map")
|
||||
print(" and goes on --allow; a demonym, brand or real place belongs in the STOPLIST,")
|
||||
print(" because renaming it damages the prose and no gate will ever tell you.")
|
||||
return 1
|
||||
print("\n [PASS] every flagged surface is on the read-and-cleared list")
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
Reference in New Issue
Block a user