Files
esh-pfi-infrastructure/scripts/r49-corpus/audit_pairs_sourcenames.py
T
vh 051b99e063 audit_entity_map: the rename can damage the prose and no gate will ever say so
audit_stoplist.py finds surfaces wrongly held OUT of the entity map -- a stoplisted
character is an undetectable leak. This is the mirror: surfaces wrongly held IN it.
leak_gate.py only ever asks whether the author's names are GONE, never whether
non-names were spared, so renaming `the Chinese` into an invented surname passes it
perfectly.

Found sideways on Hemingway. The pairs audit reported beats naming African, Chinese,
X-ray, Republican and Cezanne as leaks -- correctly, those surfaces really were removed
from the corpus. Reading why turned up the larger defect: they should never have been
renameable in the first place.

Measured on the Hemingway map, both controls green:
  positive  `other` 764/1356 article-preceded = 0.56
  negative  100 honorific-confirmed people, highest Inglés at 0.26, bulk 0.00-0.06
  FLAGGED   130 of 946 surfaces, 1,616 instances = 0.162% of corpus words

The signal is an article in front of the surface: you write `the Frenchman` and `a
Martini`, never `the Rinaldi`. It is a heuristic and every hit is reported FOR READING,
never auto-removed -- `the Widow` and `the Informer` are genuine Hemingway epithet-names
that SHOULD be renamed, and the band's own top entry makes the point, since Inglés at
0.26 is an in-world nickname deliberately kept renameable and sits just under the bar.

Initials are excluded from the negative-control band rather than admitted to it. `Mr. P.`
is an initial, not a person, so letting it in lets a map defect poison the control that
validates the detector -- on Hemingway `P` (0.32, every occurrence `the P. O. U. M.`) was
the one surface failing a band whose next highest was 0.26. Initials take no article and
are invisible to the scan anyway, so every surface of two characters or fewer is now
listed unconditionally. Sixteen of them are in this map, C at 274 occurrences; the same
class as the `G` that was caught by hand about to be renamed to a surname 248 times.

The unresolved count that drives the exit code is computed over every flagged surface,
not the --show slice. Tying a gate's verdict to a display flag is the same defect as a
log filter that turns a real event into a clean zero.

Also corrects a wrong claim in audit_pairs_sourcenames.py's docstring: the Hemingway
rename did not HOLD 591 surfaces. Paris, Madrid and Spain survive because the stoplist
keeps them out of the entity map before it is built, so the map is exactly the removed
set -- 941 surfaces, 941 removed, 0 kept. Measured per run rather than assumed, because
a pipeline that carried kept surfaces into the map would report every `Paris` as a leak.
2026-09-17 02:01:01 -07:00

167 lines
8.3 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
"""Do the GENERATED BEATS name characters the rename removed?
`leak_gate.py` reads the corpus and the renamed copies. It never reads the pairs,
so it is structurally blind to the leak found on lv-bronte (2026-09-16): the beat
is written by an LLM that read the passage, and if it recognises the book it
supplies the canonical names from its own memory. The rename can be perfect and
the instruction half of every pair still carry `Rochester`.
`build_sft_pairs.py --source-entities` closes that at BUILD time. This closes it
for pair sets already built — Hemingway's and Yarros's both predate the flag, and
a clean corpus gate is not evidence about them either way.
WHAT COUNTS AS A LEAK, and why the distinction matters. A source surface the beat
names is only a leak if the rename actually took it away — so every matched surface
is classified against the renamed copies first, and only the removed ones count
against the gate. Real-world names the pipeline deliberately keeps (`Paris` 173
occurrences, `Madrid` 108, `Spain` 87, all still present in the renamed copies) are
held back by the stoplist BEFORE the entity map is built, so on Hemingway the map
turns out to be exactly the removed set: 941 surfaces, 941 removed, 0 kept. Do not
assume that holds on another corpus — the classification is measured per run, and a
pipeline that instead carries kept surfaces INTO the map would report every mention
of `Paris` as a leak if this step were skipped.
THE RESPONSE SIDE IS THE DIAGNOSTIC. Beats and responses are scanned separately.
Leaks in the beats with a clean response column is the lv-bronte signature: the
rename worked and the generator undid it on the instruction side. Hits in BOTH
columns mean something upstream is wrong — the pairs were built against an
unrenamed corpus — and that is a different, larger problem.
CONTROLS, every run, because a scanner that only ever sees beats cannot tell
`absent` from `blind`:
* POSITIVE -- the same pattern over the UNRENAMED source works. Every surface
must be found there, or the zeroes downstream are worthless.
* NEGATIVE -- a nonce that appears in no tree. A hit means manufactured signal.
Exit code is the gate: 0 iff the controls pass AND no beat names a removed surface.
"""
from __future__ import annotations
import argparse, json, re, sys
from collections import Counter
from pathlib import Path
sys.path.insert(0, str(Path(__file__).resolve().parent))
from leak_gate import NONCE, load_works, load_copies, scan # noqa: E402
def load_pairs(paths: list[Path]) -> list[dict]:
rows = []
for p in paths:
for line in p.read_text(encoding="utf-8").splitlines():
if line.strip():
r = json.loads(line)
r["_src"] = p.name
rows.append(r)
return rows
def main() -> int:
ap = argparse.ArgumentParser()
ap.add_argument("--pairs", required=True, nargs="+", help="pairs jsonl (train and/or val)")
ap.add_argument("--entities", required=True, help="entities json for the UNRENAMED source")
ap.add_argument("--corpus", required=True, help="source corpus dir (manifest.json + works/)")
ap.add_argument("--renamed", required=True, help="rename.py --out dir, to classify kept vs removed")
ap.add_argument("--min-cap", type=int, default=8,
help="rename.py's renameable threshold; mirrored from leak_gate.py")
ap.add_argument("--report", default=None, help="write the full JSON breakdown here")
ap.add_argument("--show", type=int, default=25, help="example beats to print")
a = ap.parse_args()
ents_all = json.loads(Path(a.entities).read_text())
# Mirror build_sft_pairs.py's --source-entities surface set exactly, so this
# audit answers "would that flag have rejected it", not a near-miss variant.
surfaces = sorted({(e.get("surface") or key)
for w in ents_all.values() for key, e in w["entities"].items()
if "’" not in key and "'" not in key})
if not surfaces:
print("== entity map yields no surfaces"); return 1
source = load_works(Path(a.corpus))
copies = load_copies(Path(a.renamed))
if not copies:
print("== no renamed copies found -- cannot tell a removed name from a kept one"); return 1
# ---- controls --------------------------------------------------------
src_hits = scan(source, surfaces + [NONCE])
missing = [s for s in surfaces if s not in src_hits]
pos_ok = not missing
neg_ok = NONCE not in src_hits
print(f" {len(surfaces)} source surfaces · {len(source)} works · {len(copies)} renamed copies")
print(f" [{'PASS' if pos_ok else 'FAIL'}] positive control: every surface found in the "
f"unrenamed source ({len(surfaces) - len(missing)}/{len(surfaces)})"
+ ("" if pos_ok else f" -- MISSING {missing[:10]}"))
# ---- which surfaces did the rename actually remove? ------------------
copy_hits = scan(copies, surfaces + [NONCE])
neg_ok = neg_ok and NONCE not in copy_hits
print(f" [{'PASS' if neg_ok else 'FAIL'}] negative control: nonce `{NONCE}` absent from both trees")
kept = {s for s in surfaces if s in copy_hits}
removed = [s for s in surfaces if s not in copy_hits]
print(f" rename KEPT {len(kept)} surfaces (real places, allow-listed, sub-threshold) · "
f"REMOVED {len(removed)}")
if not removed:
print("== the rename removed nothing -- this audit has no leak to look for"); return 1
# ---- the measurement -------------------------------------------------
rows = load_pairs([Path(p) for p in a.pairs])
print(f" {len(rows)} pairs from {len({r['_src'] for r in rows})} file(s)")
removed_pat = re.compile(r"\b(" + "|".join(re.escape(s) for s in
sorted(removed, key=len, reverse=True)) + r")\b")
cols = {"beat": Counter(), "response": Counter()}
hit_rows: dict[str, list] = {"beat": [], "response": []}
for r in rows:
for col in cols:
text = r.get(col) or ""
found = sorted(set(removed_pat.findall(text)))
if found:
cols[col].update(found)
hit_rows[col].append({"src": r["_src"], "work": r.get("work"),
"names": found, "text": text})
print()
for col in ("beat", "response"):
n = len(hit_rows[col])
print(f" {col.upper():<9} naming a REMOVED surface: {n} of {len(rows)} "
f"({n / len(rows):.2%}) · {len(cols[col])} distinct names")
for s, c in cols[col].most_common(15):
print(f" {s:<20} x{c}")
# The lv-bronte signature, stated rather than left to be inferred.
nb, nr = len(hit_rows["beat"]), len(hit_rows["response"])
print()
if nb and not nr:
print(" ⭐ BEAT-ONLY leak — the rename held and the beat generator undid it on the "
"instruction side. Regenerate the pairs with --source-entities.")
elif nb and nr:
print(" ⚠⚠ BOTH columns leak — this is NOT the beat-generator class. The pairs were "
"probably built against an unrenamed corpus; check the provenance `corpus` path.")
elif nr:
print(" ⚠⚠ RESPONSE-only leak — the response is copied from the corpus, so a hit here "
"means the renamed copies are not what the pairs were built from.")
for col in ("beat", "response"):
for h in hit_rows[col][:a.show]:
print(f"\n [{col}] {h['src']} · {h['work']} · {h['names']}")
print(f" {h['text'][:300]}")
if a.report:
Path(a.report).write_text(json.dumps({
"pairs": [str(p) for p in a.pairs],
"surfaces_total": len(surfaces), "kept": len(kept), "removed": len(removed),
"controls": {"positive_pass": pos_ok, "negative_pass": neg_ok, "missing": missing[:50]},
"rows": len(rows),
"beat_hits": nb, "beat_names": dict(cols["beat"]),
"response_hits": nr, "response_names": dict(cols["response"]),
"examples": {c: hit_rows[c][:50] for c in hit_rows},
}, ensure_ascii=False, indent=2), encoding="utf-8")
print(f"\n wrote {a.report}")
ok = pos_ok and neg_ok and nb == 0 and nr == 0
print(f"\n GATE: {'PASS' if ok else 'FAIL'}")
return 0 if ok else 2
if __name__ == "__main__":
raise SystemExit(main())