lv-hemingway: pre-register the v2 gate, and fix the floor rule that decided lv-bronte
The gate design is written before any generation exists, because lv-bronte's
verdict turned on a choice that was only visible after the numbers printed.
THE FLOOR RULE IS NOW PAIRWISE. lv-bronte computed the noise floor as the largest
within-arm seed spread across ALL arms present. Its ckpt475 shipped at +0.193
against a 0.251 floor set entirely by ckpt925 -- a third arm nobody was shipping,
on one outlier seed. Scored against the arm it was actually compared to, the floor
is 0.092 and the same gap clears at 2.1x. A candidate's verdict must not depend on
which other arms happened to be generated. voice_distance.py now prints both floors
and flags any disagreement, so the lv-bronte record stays comparable.
audit_pairs_sourcenames.py closes the blind spot leak_gate.py has by construction:
it reads the corpus and the renamed copies, never the generated beats, so it cannot
see a beat-writing model restoring the author's real character names. Run over the
Hemingway pairs, which predate build_sft_pairs.py --source-entities:
val 0 of 200 -- the eval fixture is clean, the gate is unconfounded
train 70 of 7,094 (0.96%) -- Santiago x16, Catherine x7, Rinaldi x3, Brett,
Harry, Jake, Pablo, Nick, Maria ...
responses 0 of 7,294 -- the lv-bronte beat-only signature exactly
A matched surface is only counted when the rename actually removed it, verified
against the renamed copies, so a beat naming a held real-world place is not a leak.
Controls run every time: 941/941 surfaces found in the unrenamed source, nonce
absent from both trees, and 6 planted canonical names detected 6/6.
voice_distance.py --author is now REQUIRED. It was hardcoded "Yarros" and printed
"reference: held-out Yarros" over Brontë's numbers into a committed artifact. A
default would have moved the silent-wrong-label failure rather than removed it. The
stale "one seed-pair per arm / corroborates Base < Instruct" footer is replaced with
what the run actually carries.
Gate design: three arms (base-unadapted, ckpt1750, ckpt850), 60 beats, 4 seeds.
ckpt850 is present because the loss curve cannot separate it from ckpt1750 -- +0.0040
against a 0.0044 median neighbour jitter, with three checkpoints inside one jitter of
the minimum. adapter/ is excluded: +0.0762 is 17.4x the jitter and is resolved without
a gate.
This commit is contained in:
@@ -0,0 +1,164 @@
|
||||
"""Do the GENERATED BEATS name characters the rename removed?
|
||||
|
||||
`leak_gate.py` reads the corpus and the renamed copies. It never reads the pairs,
|
||||
so it is structurally blind to the leak found on lv-bronte (2026-09-16): the beat
|
||||
is written by an LLM that read the passage, and if it recognises the book it
|
||||
supplies the canonical names from its own memory. The rename can be perfect and
|
||||
the instruction half of every pair still carry `Rochester`.
|
||||
|
||||
`build_sft_pairs.py --source-entities` closes that at BUILD time. This closes it
|
||||
for pair sets already built — Hemingway's and Yarros's both predate the flag, and
|
||||
a clean corpus gate is not evidence about them either way.
|
||||
|
||||
WHAT COUNTS AS A LEAK, and why the distinction matters. A source surface the beat
|
||||
names is only a leak if the rename actually took it away. Hemingway's map holds
|
||||
941 surfaces and the rename moved 1,097 instances while HOLDING 591 — real places
|
||||
(`Paris`, `Madrid`), allow-listed real-world terms, and everything under the
|
||||
`--min-cap` threshold. A beat naming `Paris` names something the renamed corpus
|
||||
says constantly; a beat naming a removed character restores what the pipeline
|
||||
exists to delete. So every matched surface is classified against the renamed
|
||||
copies first, and only the removed ones are counted against the gate.
|
||||
|
||||
THE RESPONSE SIDE IS THE DIAGNOSTIC. Beats and responses are scanned separately.
|
||||
Leaks in the beats with a clean response column is the lv-bronte signature: the
|
||||
rename worked and the generator undid it on the instruction side. Hits in BOTH
|
||||
columns mean something upstream is wrong — the pairs were built against an
|
||||
unrenamed corpus — and that is a different, larger problem.
|
||||
|
||||
CONTROLS, every run, because a scanner that only ever sees beats cannot tell
|
||||
`absent` from `blind`:
|
||||
* POSITIVE -- the same pattern over the UNRENAMED source works. Every surface
|
||||
must be found there, or the zeroes downstream are worthless.
|
||||
* NEGATIVE -- a nonce that appears in no tree. A hit means manufactured signal.
|
||||
|
||||
Exit code is the gate: 0 iff the controls pass AND no beat names a removed surface.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
import argparse, json, re, sys
|
||||
from collections import Counter
|
||||
from pathlib import Path
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent))
|
||||
from leak_gate import NONCE, load_works, load_copies, scan # noqa: E402
|
||||
|
||||
|
||||
def load_pairs(paths: list[Path]) -> list[dict]:
|
||||
rows = []
|
||||
for p in paths:
|
||||
for line in p.read_text(encoding="utf-8").splitlines():
|
||||
if line.strip():
|
||||
r = json.loads(line)
|
||||
r["_src"] = p.name
|
||||
rows.append(r)
|
||||
return rows
|
||||
|
||||
|
||||
def main() -> int:
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("--pairs", required=True, nargs="+", help="pairs jsonl (train and/or val)")
|
||||
ap.add_argument("--entities", required=True, help="entities json for the UNRENAMED source")
|
||||
ap.add_argument("--corpus", required=True, help="source corpus dir (manifest.json + works/)")
|
||||
ap.add_argument("--renamed", required=True, help="rename.py --out dir, to classify kept vs removed")
|
||||
ap.add_argument("--min-cap", type=int, default=8,
|
||||
help="rename.py's renameable threshold; mirrored from leak_gate.py")
|
||||
ap.add_argument("--report", default=None, help="write the full JSON breakdown here")
|
||||
ap.add_argument("--show", type=int, default=25, help="example beats to print")
|
||||
a = ap.parse_args()
|
||||
|
||||
ents_all = json.loads(Path(a.entities).read_text())
|
||||
# Mirror build_sft_pairs.py's --source-entities surface set exactly, so this
|
||||
# audit answers "would that flag have rejected it", not a near-miss variant.
|
||||
surfaces = sorted({(e.get("surface") or key)
|
||||
for w in ents_all.values() for key, e in w["entities"].items()
|
||||
if "’" not in key and "'" not in key})
|
||||
if not surfaces:
|
||||
print("== entity map yields no surfaces"); return 1
|
||||
|
||||
source = load_works(Path(a.corpus))
|
||||
copies = load_copies(Path(a.renamed))
|
||||
if not copies:
|
||||
print("== no renamed copies found -- cannot tell a removed name from a kept one"); return 1
|
||||
|
||||
# ---- controls --------------------------------------------------------
|
||||
src_hits = scan(source, surfaces + [NONCE])
|
||||
missing = [s for s in surfaces if s not in src_hits]
|
||||
pos_ok = not missing
|
||||
neg_ok = NONCE not in src_hits
|
||||
print(f" {len(surfaces)} source surfaces · {len(source)} works · {len(copies)} renamed copies")
|
||||
print(f" [{'PASS' if pos_ok else 'FAIL'}] positive control: every surface found in the "
|
||||
f"unrenamed source ({len(surfaces) - len(missing)}/{len(surfaces)})"
|
||||
+ ("" if pos_ok else f" -- MISSING {missing[:10]}"))
|
||||
|
||||
# ---- which surfaces did the rename actually remove? ------------------
|
||||
copy_hits = scan(copies, surfaces + [NONCE])
|
||||
neg_ok = neg_ok and NONCE not in copy_hits
|
||||
print(f" [{'PASS' if neg_ok else 'FAIL'}] negative control: nonce `{NONCE}` absent from both trees")
|
||||
kept = {s for s in surfaces if s in copy_hits}
|
||||
removed = [s for s in surfaces if s not in copy_hits]
|
||||
print(f" rename KEPT {len(kept)} surfaces (real places, allow-listed, sub-threshold) · "
|
||||
f"REMOVED {len(removed)}")
|
||||
if not removed:
|
||||
print("== the rename removed nothing -- this audit has no leak to look for"); return 1
|
||||
|
||||
# ---- the measurement -------------------------------------------------
|
||||
rows = load_pairs([Path(p) for p in a.pairs])
|
||||
print(f" {len(rows)} pairs from {len({r['_src'] for r in rows})} file(s)")
|
||||
removed_pat = re.compile(r"\b(" + "|".join(re.escape(s) for s in
|
||||
sorted(removed, key=len, reverse=True)) + r")\b")
|
||||
|
||||
cols = {"beat": Counter(), "response": Counter()}
|
||||
hit_rows: dict[str, list] = {"beat": [], "response": []}
|
||||
for r in rows:
|
||||
for col in cols:
|
||||
text = r.get(col) or ""
|
||||
found = sorted(set(removed_pat.findall(text)))
|
||||
if found:
|
||||
cols[col].update(found)
|
||||
hit_rows[col].append({"src": r["_src"], "work": r.get("work"),
|
||||
"names": found, "text": text})
|
||||
|
||||
print()
|
||||
for col in ("beat", "response"):
|
||||
n = len(hit_rows[col])
|
||||
print(f" {col.upper():<9} naming a REMOVED surface: {n} of {len(rows)} "
|
||||
f"({n / len(rows):.2%}) · {len(cols[col])} distinct names")
|
||||
for s, c in cols[col].most_common(15):
|
||||
print(f" {s:<20} x{c}")
|
||||
|
||||
# The lv-bronte signature, stated rather than left to be inferred.
|
||||
nb, nr = len(hit_rows["beat"]), len(hit_rows["response"])
|
||||
print()
|
||||
if nb and not nr:
|
||||
print(" ⭐ BEAT-ONLY leak — the rename held and the beat generator undid it on the "
|
||||
"instruction side. Regenerate the pairs with --source-entities.")
|
||||
elif nb and nr:
|
||||
print(" ⚠⚠ BOTH columns leak — this is NOT the beat-generator class. The pairs were "
|
||||
"probably built against an unrenamed corpus; check the provenance `corpus` path.")
|
||||
elif nr:
|
||||
print(" ⚠⚠ RESPONSE-only leak — the response is copied from the corpus, so a hit here "
|
||||
"means the renamed copies are not what the pairs were built from.")
|
||||
|
||||
for col in ("beat", "response"):
|
||||
for h in hit_rows[col][:a.show]:
|
||||
print(f"\n [{col}] {h['src']} · {h['work']} · {h['names']}")
|
||||
print(f" {h['text'][:300]}")
|
||||
|
||||
if a.report:
|
||||
Path(a.report).write_text(json.dumps({
|
||||
"pairs": [str(p) for p in a.pairs],
|
||||
"surfaces_total": len(surfaces), "kept": len(kept), "removed": len(removed),
|
||||
"controls": {"positive_pass": pos_ok, "negative_pass": neg_ok, "missing": missing[:50]},
|
||||
"rows": len(rows),
|
||||
"beat_hits": nb, "beat_names": dict(cols["beat"]),
|
||||
"response_hits": nr, "response_names": dict(cols["response"]),
|
||||
"examples": {c: hit_rows[c][:50] for c in hit_rows},
|
||||
}, ensure_ascii=False, indent=2), encoding="utf-8")
|
||||
print(f"\n wrote {a.report}")
|
||||
|
||||
ok = pos_ok and neg_ok and nb == 0 and nr == 0
|
||||
print(f"\n GATE: {'PASS' if ok else 'FAIL'}")
|
||||
return 0 if ok else 2
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
Reference in New Issue
Block a user