The gate design is written before any generation exists, because lv-bronte's
verdict turned on a choice that was only visible after the numbers printed.
THE FLOOR RULE IS NOW PAIRWISE. lv-bronte computed the noise floor as the largest
within-arm seed spread across ALL arms present. Its ckpt475 shipped at +0.193
against a 0.251 floor set entirely by ckpt925 -- a third arm nobody was shipping,
on one outlier seed. Scored against the arm it was actually compared to, the floor
is 0.092 and the same gap clears at 2.1x. A candidate's verdict must not depend on
which other arms happened to be generated. voice_distance.py now prints both floors
and flags any disagreement, so the lv-bronte record stays comparable.
audit_pairs_sourcenames.py closes the blind spot leak_gate.py has by construction:
it reads the corpus and the renamed copies, never the generated beats, so it cannot
see a beat-writing model restoring the author's real character names. Run over the
Hemingway pairs, which predate build_sft_pairs.py --source-entities:
val 0 of 200 -- the eval fixture is clean, the gate is unconfounded
train 70 of 7,094 (0.96%) -- Santiago x16, Catherine x7, Rinaldi x3, Brett,
Harry, Jake, Pablo, Nick, Maria ...
responses 0 of 7,294 -- the lv-bronte beat-only signature exactly
A matched surface is only counted when the rename actually removed it, verified
against the renamed copies, so a beat naming a held real-world place is not a leak.
Controls run every time: 941/941 surfaces found in the unrenamed source, nonce
absent from both trees, and 6 planted canonical names detected 6/6.
voice_distance.py --author is now REQUIRED. It was hardcoded "Yarros" and printed
"reference: held-out Yarros" over Brontë's numbers into a committed artifact. A
default would have moved the silent-wrong-label failure rather than removed it. The
stale "one seed-pair per arm / corroborates Base < Instruct" footer is replaced with
what the run actually carries.
Gate design: three arms (base-unadapted, ckpt1750, ckpt850), 60 beats, 4 seeds.
ckpt850 is present because the loss curve cannot separate it from ckpt1750 -- +0.0040
against a 0.0044 median neighbour jitter, with three checkpoints inside one jitter of
the minimum. adapter/ is excluded: +0.0762 is 17.4x the jitter and is resolved without
a gate.
165 lines
8.1 KiB
Python
165 lines
8.1 KiB
Python
"""Do the GENERATED BEATS name characters the rename removed?
|
||
|
||
`leak_gate.py` reads the corpus and the renamed copies. It never reads the pairs,
|
||
so it is structurally blind to the leak found on lv-bronte (2026-09-16): the beat
|
||
is written by an LLM that read the passage, and if it recognises the book it
|
||
supplies the canonical names from its own memory. The rename can be perfect and
|
||
the instruction half of every pair still carry `Rochester`.
|
||
|
||
`build_sft_pairs.py --source-entities` closes that at BUILD time. This closes it
|
||
for pair sets already built — Hemingway's and Yarros's both predate the flag, and
|
||
a clean corpus gate is not evidence about them either way.
|
||
|
||
WHAT COUNTS AS A LEAK, and why the distinction matters. A source surface the beat
|
||
names is only a leak if the rename actually took it away. Hemingway's map holds
|
||
941 surfaces and the rename moved 1,097 instances while HOLDING 591 — real places
|
||
(`Paris`, `Madrid`), allow-listed real-world terms, and everything under the
|
||
`--min-cap` threshold. A beat naming `Paris` names something the renamed corpus
|
||
says constantly; a beat naming a removed character restores what the pipeline
|
||
exists to delete. So every matched surface is classified against the renamed
|
||
copies first, and only the removed ones are counted against the gate.
|
||
|
||
THE RESPONSE SIDE IS THE DIAGNOSTIC. Beats and responses are scanned separately.
|
||
Leaks in the beats with a clean response column is the lv-bronte signature: the
|
||
rename worked and the generator undid it on the instruction side. Hits in BOTH
|
||
columns mean something upstream is wrong — the pairs were built against an
|
||
unrenamed corpus — and that is a different, larger problem.
|
||
|
||
CONTROLS, every run, because a scanner that only ever sees beats cannot tell
|
||
`absent` from `blind`:
|
||
* POSITIVE -- the same pattern over the UNRENAMED source works. Every surface
|
||
must be found there, or the zeroes downstream are worthless.
|
||
* NEGATIVE -- a nonce that appears in no tree. A hit means manufactured signal.
|
||
|
||
Exit code is the gate: 0 iff the controls pass AND no beat names a removed surface.
|
||
"""
|
||
from __future__ import annotations
|
||
import argparse, json, re, sys
|
||
from collections import Counter
|
||
from pathlib import Path
|
||
|
||
sys.path.insert(0, str(Path(__file__).resolve().parent))
|
||
from leak_gate import NONCE, load_works, load_copies, scan # noqa: E402
|
||
|
||
|
||
def load_pairs(paths: list[Path]) -> list[dict]:
|
||
rows = []
|
||
for p in paths:
|
||
for line in p.read_text(encoding="utf-8").splitlines():
|
||
if line.strip():
|
||
r = json.loads(line)
|
||
r["_src"] = p.name
|
||
rows.append(r)
|
||
return rows
|
||
|
||
|
||
def main() -> int:
|
||
ap = argparse.ArgumentParser()
|
||
ap.add_argument("--pairs", required=True, nargs="+", help="pairs jsonl (train and/or val)")
|
||
ap.add_argument("--entities", required=True, help="entities json for the UNRENAMED source")
|
||
ap.add_argument("--corpus", required=True, help="source corpus dir (manifest.json + works/)")
|
||
ap.add_argument("--renamed", required=True, help="rename.py --out dir, to classify kept vs removed")
|
||
ap.add_argument("--min-cap", type=int, default=8,
|
||
help="rename.py's renameable threshold; mirrored from leak_gate.py")
|
||
ap.add_argument("--report", default=None, help="write the full JSON breakdown here")
|
||
ap.add_argument("--show", type=int, default=25, help="example beats to print")
|
||
a = ap.parse_args()
|
||
|
||
ents_all = json.loads(Path(a.entities).read_text())
|
||
# Mirror build_sft_pairs.py's --source-entities surface set exactly, so this
|
||
# audit answers "would that flag have rejected it", not a near-miss variant.
|
||
surfaces = sorted({(e.get("surface") or key)
|
||
for w in ents_all.values() for key, e in w["entities"].items()
|
||
if "’" not in key and "'" not in key})
|
||
if not surfaces:
|
||
print("== entity map yields no surfaces"); return 1
|
||
|
||
source = load_works(Path(a.corpus))
|
||
copies = load_copies(Path(a.renamed))
|
||
if not copies:
|
||
print("== no renamed copies found -- cannot tell a removed name from a kept one"); return 1
|
||
|
||
# ---- controls --------------------------------------------------------
|
||
src_hits = scan(source, surfaces + [NONCE])
|
||
missing = [s for s in surfaces if s not in src_hits]
|
||
pos_ok = not missing
|
||
neg_ok = NONCE not in src_hits
|
||
print(f" {len(surfaces)} source surfaces · {len(source)} works · {len(copies)} renamed copies")
|
||
print(f" [{'PASS' if pos_ok else 'FAIL'}] positive control: every surface found in the "
|
||
f"unrenamed source ({len(surfaces) - len(missing)}/{len(surfaces)})"
|
||
+ ("" if pos_ok else f" -- MISSING {missing[:10]}"))
|
||
|
||
# ---- which surfaces did the rename actually remove? ------------------
|
||
copy_hits = scan(copies, surfaces + [NONCE])
|
||
neg_ok = neg_ok and NONCE not in copy_hits
|
||
print(f" [{'PASS' if neg_ok else 'FAIL'}] negative control: nonce `{NONCE}` absent from both trees")
|
||
kept = {s for s in surfaces if s in copy_hits}
|
||
removed = [s for s in surfaces if s not in copy_hits]
|
||
print(f" rename KEPT {len(kept)} surfaces (real places, allow-listed, sub-threshold) · "
|
||
f"REMOVED {len(removed)}")
|
||
if not removed:
|
||
print("== the rename removed nothing -- this audit has no leak to look for"); return 1
|
||
|
||
# ---- the measurement -------------------------------------------------
|
||
rows = load_pairs([Path(p) for p in a.pairs])
|
||
print(f" {len(rows)} pairs from {len({r['_src'] for r in rows})} file(s)")
|
||
removed_pat = re.compile(r"\b(" + "|".join(re.escape(s) for s in
|
||
sorted(removed, key=len, reverse=True)) + r")\b")
|
||
|
||
cols = {"beat": Counter(), "response": Counter()}
|
||
hit_rows: dict[str, list] = {"beat": [], "response": []}
|
||
for r in rows:
|
||
for col in cols:
|
||
text = r.get(col) or ""
|
||
found = sorted(set(removed_pat.findall(text)))
|
||
if found:
|
||
cols[col].update(found)
|
||
hit_rows[col].append({"src": r["_src"], "work": r.get("work"),
|
||
"names": found, "text": text})
|
||
|
||
print()
|
||
for col in ("beat", "response"):
|
||
n = len(hit_rows[col])
|
||
print(f" {col.upper():<9} naming a REMOVED surface: {n} of {len(rows)} "
|
||
f"({n / len(rows):.2%}) · {len(cols[col])} distinct names")
|
||
for s, c in cols[col].most_common(15):
|
||
print(f" {s:<20} x{c}")
|
||
|
||
# The lv-bronte signature, stated rather than left to be inferred.
|
||
nb, nr = len(hit_rows["beat"]), len(hit_rows["response"])
|
||
print()
|
||
if nb and not nr:
|
||
print(" ⭐ BEAT-ONLY leak — the rename held and the beat generator undid it on the "
|
||
"instruction side. Regenerate the pairs with --source-entities.")
|
||
elif nb and nr:
|
||
print(" ⚠⚠ BOTH columns leak — this is NOT the beat-generator class. The pairs were "
|
||
"probably built against an unrenamed corpus; check the provenance `corpus` path.")
|
||
elif nr:
|
||
print(" ⚠⚠ RESPONSE-only leak — the response is copied from the corpus, so a hit here "
|
||
"means the renamed copies are not what the pairs were built from.")
|
||
|
||
for col in ("beat", "response"):
|
||
for h in hit_rows[col][:a.show]:
|
||
print(f"\n [{col}] {h['src']} · {h['work']} · {h['names']}")
|
||
print(f" {h['text'][:300]}")
|
||
|
||
if a.report:
|
||
Path(a.report).write_text(json.dumps({
|
||
"pairs": [str(p) for p in a.pairs],
|
||
"surfaces_total": len(surfaces), "kept": len(kept), "removed": len(removed),
|
||
"controls": {"positive_pass": pos_ok, "negative_pass": neg_ok, "missing": missing[:50]},
|
||
"rows": len(rows),
|
||
"beat_hits": nb, "beat_names": dict(cols["beat"]),
|
||
"response_hits": nr, "response_names": dict(cols["response"]),
|
||
"examples": {c: hit_rows[c][:50] for c in hit_rows},
|
||
}, ensure_ascii=False, indent=2), encoding="utf-8")
|
||
print(f"\n wrote {a.report}")
|
||
|
||
ok = pos_ok and neg_ok and nb == 0 and nr == 0
|
||
print(f"\n GATE: {'PASS' if ok else 'FAIL'}")
|
||
return 0 if ok else 2
|
||
|
||
|
||
if __name__ == "__main__":
|
||
raise SystemExit(main())
|