lv-hemingway: pre-register the v2 gate, and fix the floor rule that decided lv-bronte
The gate design is written before any generation exists, because lv-bronte's
verdict turned on a choice that was only visible after the numbers printed.
THE FLOOR RULE IS NOW PAIRWISE. lv-bronte computed the noise floor as the largest
within-arm seed spread across ALL arms present. Its ckpt475 shipped at +0.193
against a 0.251 floor set entirely by ckpt925 -- a third arm nobody was shipping,
on one outlier seed. Scored against the arm it was actually compared to, the floor
is 0.092 and the same gap clears at 2.1x. A candidate's verdict must not depend on
which other arms happened to be generated. voice_distance.py now prints both floors
and flags any disagreement, so the lv-bronte record stays comparable.
audit_pairs_sourcenames.py closes the blind spot leak_gate.py has by construction:
it reads the corpus and the renamed copies, never the generated beats, so it cannot
see a beat-writing model restoring the author's real character names. Run over the
Hemingway pairs, which predate build_sft_pairs.py --source-entities:
val 0 of 200 -- the eval fixture is clean, the gate is unconfounded
train 70 of 7,094 (0.96%) -- Santiago x16, Catherine x7, Rinaldi x3, Brett,
Harry, Jake, Pablo, Nick, Maria ...
responses 0 of 7,294 -- the lv-bronte beat-only signature exactly
A matched surface is only counted when the rename actually removed it, verified
against the renamed copies, so a beat naming a held real-world place is not a leak.
Controls run every time: 941/941 surfaces found in the unrenamed source, nonce
absent from both trees, and 6 planted canonical names detected 6/6.
voice_distance.py --author is now REQUIRED. It was hardcoded "Yarros" and printed
"reference: held-out Yarros" over Brontë's numbers into a committed artifact. A
default would have moved the silent-wrong-label failure rather than removed it. The
stale "one seed-pair per arm / corroborates Base < Instruct" footer is replaced with
what the run actually carries.
Gate design: three arms (base-unadapted, ckpt1750, ckpt850), 60 beats, 4 seeds.
ckpt850 is present because the loss curve cannot separate it from ckpt1750 -- +0.0040
against a 0.0044 median neighbour jitter, with three checkpoints inside one jitter of
the minimum. adapter/ is excluded: +0.0762 is 17.4x the jitter and is resolved without
a gate.
This commit is contained in:
@@ -0,0 +1,124 @@
|
||||
# lv-hemingway v2 gate — pre-registration
|
||||
|
||||
**Written 2026-09-17, before a single generation existed.** Everything below is
|
||||
fixed. The point of writing it down is that a threshold chosen after seeing the
|
||||
numbers is not a threshold, and this line has already been burned once: lv-bronte's
|
||||
verdict turned on which arms happened to be in the run (see § 4).
|
||||
|
||||
Nothing here may be amended after the first arm is generated. If the design turns
|
||||
out to be wrong, the run is wrong and gets re-run — the rule does not move.
|
||||
|
||||
---
|
||||
|
||||
## 1. What is being gated
|
||||
|
||||
`gx10:~/r49-runs/hemingway-4b-pairs-3ep/` — Qwen3-4B-Instruct + instruction-pair
|
||||
LoRA on the renamed Hemingway corpus (994,760 words, 318 units, 6 renamed copies;
|
||||
leak gate 0 of 941 entities with both controls green). 3-epoch schedule, seed 4919,
|
||||
2,661 steps, finished 2026-09-16 18:27.
|
||||
|
||||
## 2. Arms — three, and why each one is present
|
||||
|
||||
| arm | what it is | role |
|
||||
|---|---|---|
|
||||
| `base-unadapted` | `~/carriers/Qwen3-4B-Instruct`, no adapter | negative control for memorisation (never saw the corpus) **and** the voice baseline |
|
||||
| `ckpt1750` | epoch 1.973, eval_loss **2.2783** — the curve's minimum | ship candidate |
|
||||
| `ckpt850` | epoch 0.959, eval_loss **2.2823** | **tied** candidate, see below |
|
||||
|
||||
⭐ **`ckpt850` is in the run because the loss curve cannot separate it from
|
||||
`ckpt1750`.** Measured on this run's own 54 eval points: median neighbour jitter
|
||||
**0.0044**, and step 850 sits **+0.0040** above the minimum. Three checkpoints are
|
||||
inside one jitter of the best (850, 1700, 1750). "The minimum is at 1750" is true
|
||||
and is *not* a reason to believe 1750 is the better adapter — picking it by staring
|
||||
at the fourth decimal is reading noise, which is exactly what lv-bronte recorded
|
||||
about its own 475-vs-925 pair.
|
||||
|
||||
There is also a prior worth testing rather than assuming: on Brontë the EARLIER,
|
||||
epoch-1 checkpoint won the tiebreak on the axes that resolve (identical-to-control
|
||||
memorisation, 2.7× tighter seed variance). Epoch 2 bought Brontë nothing. If that
|
||||
repeats here, `ckpt850` ships and half an epoch of overfit is avoided.
|
||||
|
||||
**`adapter/` (end of run, epoch 3.0, 2.3546) is NOT gated.** It is +0.0762 worse
|
||||
than the minimum — **17.4× the median jitter** — so unlike the 850/1750 pair this
|
||||
one is resolved by the loss curve alone. The epoch-3 collapse is the robust finding
|
||||
on this schedule; the epoch-2-vs-epoch-1 question is the open one.
|
||||
|
||||
## 3. Fixture and seeds — fixed now
|
||||
|
||||
- **60 beats**, sampled by `build_beat_fixture.py` from `~/hemingway-pairs/pairs-val.jsonl`
|
||||
(held-out `split=val`, 200 pairs, **173 inside the 90–150 word band**), `--seed 4919`,
|
||||
spread across works round-robin.
|
||||
- **4 seeds: 1234 5678 9012 3456.** Same as lv-bronte.
|
||||
- → 240 generations per arm, 720 total.
|
||||
|
||||
**Why 60 and not 30.** The floor is a per-seed range and per-seed variance scales
|
||||
as 1/√N, so doubling lv-bronte's 30-beat fixture buys a ~1.41× tighter floor. 173
|
||||
in-band val pairs make 60 available; Brontë had 44 in-band and *could not* have run
|
||||
this. **More seeds would not help** — the floor is a range statistic and ranges
|
||||
widen with more samples.
|
||||
|
||||
## 4. ⚠ THE FLOOR RULE — changed, and this is the change that matters
|
||||
|
||||
**The noise floor for a candidate is the larger of the within-arm seed spreads of
|
||||
the TWO arms being compared: `max(spread(candidate), spread(base-unadapted))`.**
|
||||
|
||||
This is a deliberate amendment to the rule lv-bronte ran under, made before any
|
||||
Hemingway number exists. Under the old "largest spread across ALL arms" reading:
|
||||
|
||||
```
|
||||
base-unadapted spread 0.062
|
||||
ckpt475 spread 0.092 <- shipped
|
||||
ckpt925 spread 0.251 <- set the floor, on ONE outlier seed
|
||||
```
|
||||
|
||||
`ckpt475`'s **+0.193 was failed by a floor contributed entirely by a third arm
|
||||
nobody was shipping.** Scored pairwise the floor is 0.092 and the same gap clears
|
||||
at 2.1×. A candidate's verdict must not depend on which *other* arms you happened
|
||||
to generate — and with a three-arm design that defect is live again, which is why
|
||||
the rule is settled here rather than after the table prints.
|
||||
|
||||
`voice_distance.py` prints **both** floors and flags any case where the two rules
|
||||
disagree. The pairwise number is the verdict; the all-arms number is kept so the
|
||||
lv-bronte record stays comparable.
|
||||
|
||||
## 5. The three axes, and what passes
|
||||
|
||||
| axis | instrument | passes if |
|
||||
|---|---|---|
|
||||
| **A. VOICE** | `voice_distance.py --author Hemingway` (Burrows's Delta over char bigrams vs the held-out val split) | `delta_cb(base) − delta_cb(candidate)` **exceeds the pairwise floor** |
|
||||
| **B. NOT COPIED** | `memorization_check.py` (longest verbatim n-gram vs the renamed copies, n=8) | hit-rate and max at or near `base-unadapted`, which never saw the corpus. Positive control (corpus vs itself) must saturate, or the zeroes are blind |
|
||||
| **C. NO DAMAGE** | `score_beats.py --metric-source raw` | ran-on / out-of-band rate not worse than base beyond that metric's own floor |
|
||||
|
||||
A **FAIL on A with B and C clean** is the lv-bronte outcome and is NOT
|
||||
automatically a no-ship — it was shipped there as additive, reversible and
|
||||
safety-clean. That call is the operator's, and the difference here is that
|
||||
**Hemingway is not public domain the way Brontë is**, so axis B carries more weight
|
||||
in that conversation, not less.
|
||||
|
||||
## 6. Pre-flight already run — the beat-contamination audit
|
||||
|
||||
`audit_pairs_sourcenames.py` over all 7,294 built pairs, controls green
|
||||
(941/941 surfaces found in the unrenamed source; nonce absent from both trees;
|
||||
6/6 planted names detected in a positive control):
|
||||
|
||||
- **val pairs — 0 of 200.** The fixture is drawn from these, so the eval beats are
|
||||
clean and the gate is unconfounded. This is why the gate runs at all.
|
||||
- **train pairs — 70 of 7,094 (0.96%)** name a character the rename removed
|
||||
(`Santiago` ×16, `Catherine` ×7, `Rinaldi` ×3, `Brett`, `Harry`, `Jake` …),
|
||||
with **0 of 7,294 responses** affected — the lv-bronte beat-only signature
|
||||
exactly. The pairs predate `build_sft_pairs.py --source-entities`.
|
||||
|
||||
**This is a property of the trained adapter and it is on the record before the
|
||||
gate runs, not discovered afterwards.** It does not invalidate any axis: the
|
||||
instruction half is contaminated, the target half is not, and the fixture is clean.
|
||||
Whether to retrain on filtered pairs is a separate decision with its own cost, and
|
||||
it belongs to the operator.
|
||||
|
||||
## 7. Reproduce
|
||||
|
||||
```bash
|
||||
# on pfi-gx10, ~/lv-hemingway
|
||||
./eval-hemingway.sh # three arms, 60 beats, 4 seeds
|
||||
python voice-prep.py # beats5.* -> voice.*, control renamed to *unadapted*
|
||||
scripts/r49-corpus/voice_distance.py corpus-renamed ~/r49-runs/hemingway-eval --author Hemingway
|
||||
```
|
||||
@@ -0,0 +1,164 @@
|
||||
"""Do the GENERATED BEATS name characters the rename removed?
|
||||
|
||||
`leak_gate.py` reads the corpus and the renamed copies. It never reads the pairs,
|
||||
so it is structurally blind to the leak found on lv-bronte (2026-09-16): the beat
|
||||
is written by an LLM that read the passage, and if it recognises the book it
|
||||
supplies the canonical names from its own memory. The rename can be perfect and
|
||||
the instruction half of every pair still carry `Rochester`.
|
||||
|
||||
`build_sft_pairs.py --source-entities` closes that at BUILD time. This closes it
|
||||
for pair sets already built — Hemingway's and Yarros's both predate the flag, and
|
||||
a clean corpus gate is not evidence about them either way.
|
||||
|
||||
WHAT COUNTS AS A LEAK, and why the distinction matters. A source surface the beat
|
||||
names is only a leak if the rename actually took it away. Hemingway's map holds
|
||||
941 surfaces and the rename moved 1,097 instances while HOLDING 591 — real places
|
||||
(`Paris`, `Madrid`), allow-listed real-world terms, and everything under the
|
||||
`--min-cap` threshold. A beat naming `Paris` names something the renamed corpus
|
||||
says constantly; a beat naming a removed character restores what the pipeline
|
||||
exists to delete. So every matched surface is classified against the renamed
|
||||
copies first, and only the removed ones are counted against the gate.
|
||||
|
||||
THE RESPONSE SIDE IS THE DIAGNOSTIC. Beats and responses are scanned separately.
|
||||
Leaks in the beats with a clean response column is the lv-bronte signature: the
|
||||
rename worked and the generator undid it on the instruction side. Hits in BOTH
|
||||
columns mean something upstream is wrong — the pairs were built against an
|
||||
unrenamed corpus — and that is a different, larger problem.
|
||||
|
||||
CONTROLS, every run, because a scanner that only ever sees beats cannot tell
|
||||
`absent` from `blind`:
|
||||
* POSITIVE -- the same pattern over the UNRENAMED source works. Every surface
|
||||
must be found there, or the zeroes downstream are worthless.
|
||||
* NEGATIVE -- a nonce that appears in no tree. A hit means manufactured signal.
|
||||
|
||||
Exit code is the gate: 0 iff the controls pass AND no beat names a removed surface.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
import argparse, json, re, sys
|
||||
from collections import Counter
|
||||
from pathlib import Path
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parent))
|
||||
from leak_gate import NONCE, load_works, load_copies, scan # noqa: E402
|
||||
|
||||
|
||||
def load_pairs(paths: list[Path]) -> list[dict]:
|
||||
rows = []
|
||||
for p in paths:
|
||||
for line in p.read_text(encoding="utf-8").splitlines():
|
||||
if line.strip():
|
||||
r = json.loads(line)
|
||||
r["_src"] = p.name
|
||||
rows.append(r)
|
||||
return rows
|
||||
|
||||
|
||||
def main() -> int:
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("--pairs", required=True, nargs="+", help="pairs jsonl (train and/or val)")
|
||||
ap.add_argument("--entities", required=True, help="entities json for the UNRENAMED source")
|
||||
ap.add_argument("--corpus", required=True, help="source corpus dir (manifest.json + works/)")
|
||||
ap.add_argument("--renamed", required=True, help="rename.py --out dir, to classify kept vs removed")
|
||||
ap.add_argument("--min-cap", type=int, default=8,
|
||||
help="rename.py's renameable threshold; mirrored from leak_gate.py")
|
||||
ap.add_argument("--report", default=None, help="write the full JSON breakdown here")
|
||||
ap.add_argument("--show", type=int, default=25, help="example beats to print")
|
||||
a = ap.parse_args()
|
||||
|
||||
ents_all = json.loads(Path(a.entities).read_text())
|
||||
# Mirror build_sft_pairs.py's --source-entities surface set exactly, so this
|
||||
# audit answers "would that flag have rejected it", not a near-miss variant.
|
||||
surfaces = sorted({(e.get("surface") or key)
|
||||
for w in ents_all.values() for key, e in w["entities"].items()
|
||||
if "’" not in key and "'" not in key})
|
||||
if not surfaces:
|
||||
print("== entity map yields no surfaces"); return 1
|
||||
|
||||
source = load_works(Path(a.corpus))
|
||||
copies = load_copies(Path(a.renamed))
|
||||
if not copies:
|
||||
print("== no renamed copies found -- cannot tell a removed name from a kept one"); return 1
|
||||
|
||||
# ---- controls --------------------------------------------------------
|
||||
src_hits = scan(source, surfaces + [NONCE])
|
||||
missing = [s for s in surfaces if s not in src_hits]
|
||||
pos_ok = not missing
|
||||
neg_ok = NONCE not in src_hits
|
||||
print(f" {len(surfaces)} source surfaces · {len(source)} works · {len(copies)} renamed copies")
|
||||
print(f" [{'PASS' if pos_ok else 'FAIL'}] positive control: every surface found in the "
|
||||
f"unrenamed source ({len(surfaces) - len(missing)}/{len(surfaces)})"
|
||||
+ ("" if pos_ok else f" -- MISSING {missing[:10]}"))
|
||||
|
||||
# ---- which surfaces did the rename actually remove? ------------------
|
||||
copy_hits = scan(copies, surfaces + [NONCE])
|
||||
neg_ok = neg_ok and NONCE not in copy_hits
|
||||
print(f" [{'PASS' if neg_ok else 'FAIL'}] negative control: nonce `{NONCE}` absent from both trees")
|
||||
kept = {s for s in surfaces if s in copy_hits}
|
||||
removed = [s for s in surfaces if s not in copy_hits]
|
||||
print(f" rename KEPT {len(kept)} surfaces (real places, allow-listed, sub-threshold) · "
|
||||
f"REMOVED {len(removed)}")
|
||||
if not removed:
|
||||
print("== the rename removed nothing -- this audit has no leak to look for"); return 1
|
||||
|
||||
# ---- the measurement -------------------------------------------------
|
||||
rows = load_pairs([Path(p) for p in a.pairs])
|
||||
print(f" {len(rows)} pairs from {len({r['_src'] for r in rows})} file(s)")
|
||||
removed_pat = re.compile(r"\b(" + "|".join(re.escape(s) for s in
|
||||
sorted(removed, key=len, reverse=True)) + r")\b")
|
||||
|
||||
cols = {"beat": Counter(), "response": Counter()}
|
||||
hit_rows: dict[str, list] = {"beat": [], "response": []}
|
||||
for r in rows:
|
||||
for col in cols:
|
||||
text = r.get(col) or ""
|
||||
found = sorted(set(removed_pat.findall(text)))
|
||||
if found:
|
||||
cols[col].update(found)
|
||||
hit_rows[col].append({"src": r["_src"], "work": r.get("work"),
|
||||
"names": found, "text": text})
|
||||
|
||||
print()
|
||||
for col in ("beat", "response"):
|
||||
n = len(hit_rows[col])
|
||||
print(f" {col.upper():<9} naming a REMOVED surface: {n} of {len(rows)} "
|
||||
f"({n / len(rows):.2%}) · {len(cols[col])} distinct names")
|
||||
for s, c in cols[col].most_common(15):
|
||||
print(f" {s:<20} x{c}")
|
||||
|
||||
# The lv-bronte signature, stated rather than left to be inferred.
|
||||
nb, nr = len(hit_rows["beat"]), len(hit_rows["response"])
|
||||
print()
|
||||
if nb and not nr:
|
||||
print(" ⭐ BEAT-ONLY leak — the rename held and the beat generator undid it on the "
|
||||
"instruction side. Regenerate the pairs with --source-entities.")
|
||||
elif nb and nr:
|
||||
print(" ⚠⚠ BOTH columns leak — this is NOT the beat-generator class. The pairs were "
|
||||
"probably built against an unrenamed corpus; check the provenance `corpus` path.")
|
||||
elif nr:
|
||||
print(" ⚠⚠ RESPONSE-only leak — the response is copied from the corpus, so a hit here "
|
||||
"means the renamed copies are not what the pairs were built from.")
|
||||
|
||||
for col in ("beat", "response"):
|
||||
for h in hit_rows[col][:a.show]:
|
||||
print(f"\n [{col}] {h['src']} · {h['work']} · {h['names']}")
|
||||
print(f" {h['text'][:300]}")
|
||||
|
||||
if a.report:
|
||||
Path(a.report).write_text(json.dumps({
|
||||
"pairs": [str(p) for p in a.pairs],
|
||||
"surfaces_total": len(surfaces), "kept": len(kept), "removed": len(removed),
|
||||
"controls": {"positive_pass": pos_ok, "negative_pass": neg_ok, "missing": missing[:50]},
|
||||
"rows": len(rows),
|
||||
"beat_hits": nb, "beat_names": dict(cols["beat"]),
|
||||
"response_hits": nr, "response_names": dict(cols["response"]),
|
||||
"examples": {c: hit_rows[c][:50] for c in hit_rows},
|
||||
}, ensure_ascii=False, indent=2), encoding="utf-8")
|
||||
print(f"\n wrote {a.report}")
|
||||
|
||||
ok = pos_ok and neg_ok and nb == 0 and nr == 0
|
||||
print(f"\n GATE: {'PASS' if ok else 'FAIL'}")
|
||||
return 0 if ok else 2
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -1,15 +1,18 @@
|
||||
"""Did the adapter move the voice TOWARD held-out Yarros? A seat-free relative measure.
|
||||
"""Did the adapter move the voice TOWARD the held-out author? A seat-free relative measure.
|
||||
|
||||
NOT the frozen adjudication. That needs a romantasy control-author panel (to place an
|
||||
absolute band and a hard-negative sister), a seed-to-seed spread, and — for BEAT
|
||||
INCUMBENT — the gen seat, none of which are available here. This answers the smaller,
|
||||
honest question the operator can act on: of the three arms generated on ONE harness,
|
||||
which sits closest to real held-out Yarros, and does the adapter beat the base control?
|
||||
Written for Yarros, since used on Brontë and Hemingway. The author is now a REQUIRED
|
||||
argument rather than a hardcoded string — see the note on `--author` in main().
|
||||
|
||||
NOT the frozen adjudication. That needs a control-author panel (to place an absolute
|
||||
band and a hard-negative sister), a seed-to-seed spread, and — for BEAT INCUMBENT — the
|
||||
gen seat, none of which are available here. This answers the smaller, honest question
|
||||
the operator can act on: of the arms generated on ONE harness, which sits closest to
|
||||
the real held-out author, and does the adapter beat the base control?
|
||||
|
||||
Instrument: Burrows's Delta over CHARACTER BIGRAMS (hence delta_cb). Char bigrams are
|
||||
dominated by function-word morphology and rhythm, not proper nouns, so the rename does
|
||||
not move them. Reference profile is the HELD-OUT (val) split — text no arm was trained
|
||||
on. Each arm's pooled generations are scored against it; lower = closer to Yarros.
|
||||
on. Each arm's pooled generations are scored against it; lower = closer to the author.
|
||||
|
||||
Discipline: this is a RELATIVE reading (arms vs each other, same harness), never an
|
||||
absolute-band claim. The A-vs-A floor below is the only thing that makes a between-arm
|
||||
@@ -18,7 +21,7 @@ returns for two samples of the SAME author, so a between-arm gap smaller than th
|
||||
is not a finding.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
import json, re, sys, statistics as st
|
||||
import argparse, json, re, sys, statistics as st
|
||||
from collections import Counter
|
||||
from pathlib import Path
|
||||
|
||||
@@ -40,8 +43,22 @@ def delta(arm_text: str, ref_prof: dict, mu: dict, sd: dict, keys: list[str]) ->
|
||||
|
||||
|
||||
def main() -> int:
|
||||
corpus = Path(sys.argv[1]) # yarros-corpus-renamed (has split=val)
|
||||
evaldir = Path(sys.argv[2]) # dir of voice.*.jsonl
|
||||
# ⚠ --author IS REQUIRED, and that is the fix for a defect this script shipped with.
|
||||
# The reference label was hardcoded "Yarros". Run against Brontë it printed
|
||||
# "reference: held-out Yarros" over Brontë's numbers, and that output is now sitting
|
||||
# in a committed artifact saying the wrong author. A default would have kept the
|
||||
# silent-wrong-label failure and only moved it; naming the author is one word at the
|
||||
# call site and the label can no longer disagree with the data.
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("corpus", help="renamed corpus dir containing copies/ with split=val records")
|
||||
ap.add_argument("evaldir", help="dir of voice.<arm>.jsonl; the control arm's name must "
|
||||
"contain the substring `unadapted`")
|
||||
ap.add_argument("--author", required=True,
|
||||
help="reference author label, e.g. Hemingway. Required: see above.")
|
||||
a = ap.parse_args()
|
||||
corpus = Path(a.corpus)
|
||||
evaldir = Path(a.evaldir)
|
||||
author = a.author
|
||||
# reference = held-out val text
|
||||
val = []
|
||||
for p in sorted((corpus / "copies").glob("*.jsonl")):
|
||||
@@ -62,14 +79,14 @@ def main() -> int:
|
||||
ref_prof = profile(ref_text, keys)
|
||||
|
||||
# SAME-AUTHOR REFERENCE (the target, not a significance threshold): two halves
|
||||
# of held-out Yarros. A perfect mimic scores about this; you cannot get closer
|
||||
# to Yarros than Yarros gets to itself at this sample size.
|
||||
# of the held-out author. A perfect mimic scores about this; you cannot get closer
|
||||
# to the author than the author gets to itself at this sample size.
|
||||
half = len(words) // 2
|
||||
same_author = delta(" ".join(words[:half]), profile(" ".join(words[half:]), keys), mu, sd, keys)
|
||||
|
||||
print(f"reference: held-out Yarros, {len(words):,} words, {len(chunks)} chunks, {len(keys)} char-bigram features")
|
||||
print(f"same-author target (held-out Yarros vs itself): delta_cb = {same_author:.3f}")
|
||||
print(f" -> the floor of what any arm could reach; lower is more Yarros-like, this is the best possible\n")
|
||||
print(f"reference: held-out {author}, {len(words):,} words, {len(chunks)} chunks, {len(keys)} char-bigram features")
|
||||
print(f"same-author target (held-out {author} vs itself): delta_cb = {same_author:.3f}")
|
||||
print(f" -> the floor of what any arm could reach; lower is more {author}-like, this is the best possible\n")
|
||||
|
||||
def arm_texts(f):
|
||||
return [json.loads(l) for l in f.read_text(encoding="utf-8").splitlines()]
|
||||
@@ -93,27 +110,56 @@ def main() -> int:
|
||||
print(" arm delta_cb per-seed [words]")
|
||||
for arm, d, w, sd_, spread in sorted(rows, key=lambda x: x[1]):
|
||||
seeds = " ".join(f"{x:.3f}" for x in sd_)
|
||||
print(f" {arm:20s} {d:.3f} ({seeds}) [{w}]")
|
||||
# the noise floor is the LARGEST within-arm spread across arms
|
||||
floors = [r[4] for r in rows if r[4] == r[4]]
|
||||
noise = max(floors) if floors else float("nan")
|
||||
print(f"\n measured noise floor (largest within-arm seed spread): {noise:.3f}")
|
||||
print(f" -> a between-arm gap must exceed ~{noise:.3f} to be a real difference\n")
|
||||
print(f" {arm:20s} {d:.3f} ({seeds}) [{w}] spread {spread:.3f}")
|
||||
|
||||
base = next((d for a, d, _, _, _ in rows if "unadapted" in a), None)
|
||||
if base is not None:
|
||||
print(" vs base-unadapted control (positive gap = moved toward Yarros):")
|
||||
for arm, d, _, _, _ in sorted(rows, key=lambda x: x[1]):
|
||||
if "unadapted" in arm:
|
||||
continue
|
||||
gap = base - d
|
||||
verdict = ("MOVED toward Yarros (exceeds noise floor)" if gap > noise
|
||||
else "moved toward Yarros, but within the measured noise floor")
|
||||
print(f" {arm:20s} {gap:+.3f} ({verdict})")
|
||||
ordered = [a for a, *_ in sorted(rows, key=lambda x: x[1])]
|
||||
print(f"\n ordering: {' < '.join(ordered)} (lower = more Yarros-like)")
|
||||
print(" ⚠ one seed-pair per arm; this ordering CORROBORATES the independent held-out")
|
||||
print(" loss ordering (Base < Instruct) but is not itself a multi-seed result.")
|
||||
# ⚠⚠ THE FLOOR IS PAIRWISE, and that is a RULE CHANGE made because the all-arms
|
||||
# rule decided lv-bronte. Measured there:
|
||||
# base-unadapted spread 0.062
|
||||
# ckpt475 spread 0.092 <- the candidate that shipped
|
||||
# ckpt925 spread 0.251 <- set the floor, on ONE outlier seed
|
||||
# ckpt475's +0.193 was failed by a floor contributed entirely by a THIRD arm nobody
|
||||
# was shipping. Run as base-vs-ckpt475 the floor is 0.092 and the same gap clears at
|
||||
# 2.1x. A candidate's verdict must not depend on which other arms you happened to
|
||||
# generate, so the comparison's floor is the larger of the TWO arms being compared.
|
||||
# The all-arms number is still printed, because lv-bronte's record used it and a
|
||||
# reader comparing the two runs needs both.
|
||||
floors = [r[4] for r in rows if r[4] == r[4]]
|
||||
noise_all = max(floors) if floors else float("nan")
|
||||
spread_of = {r[0]: r[4] for r in rows}
|
||||
|
||||
base_row = next(((a, d) for a, d, _, _, _ in rows if "unadapted" in a), None)
|
||||
if base_row is None:
|
||||
print(f"\n all-arms noise floor (largest within-arm seed spread): {noise_all:.3f}")
|
||||
print(" ⚠ no arm name contains `unadapted` -- no control identified, no verdict\n")
|
||||
return 0
|
||||
base_arm, base = base_row
|
||||
print(f"\n all-arms noise floor (largest within-arm seed spread, lv-bronte's rule): {noise_all:.3f}")
|
||||
print(f" PAIRWISE floor is the verdict: max(spread(candidate), spread({base_arm}) = "
|
||||
f"{spread_of[base_arm]:.3f})\n")
|
||||
|
||||
print(f" vs {base_arm} control (positive gap = moved toward {author}):")
|
||||
for arm, d, _, _, _ in sorted(rows, key=lambda x: x[1]):
|
||||
if arm == base_arm:
|
||||
continue
|
||||
gap = base - d
|
||||
pair_floor = max(spread_of[arm], spread_of[base_arm])
|
||||
verdict = (f"MOVED toward {author} ({gap / pair_floor:.1f}x the pairwise floor "
|
||||
f"{pair_floor:.3f})" if gap > pair_floor
|
||||
else f"within the pairwise floor {pair_floor:.3f} -- NOT a finding")
|
||||
flag = "" if (gap > noise_all) == (gap > pair_floor) else " <- the two rules DISAGREE"
|
||||
print(f" {arm:20s} {gap:+.3f} ({verdict}){flag}")
|
||||
|
||||
ordered = [a for a, *_ in sorted(rows, key=lambda x: x[1])]
|
||||
print(f"\n ordering: {' < '.join(ordered)} (lower = more {author}-like)")
|
||||
# ⚠ This used to assert "one seed-pair per arm" and claim corroboration from a
|
||||
# "Base < Instruct" held-out loss ordering. Both were Yarros-run facts hardcoded
|
||||
# as if they were properties of the instrument: by lv-bronte every arm carried
|
||||
# four seeds, and no Base-vs-Instruct comparison was in the run at all. Report
|
||||
# what this run actually has instead of a remembered one.
|
||||
nseeds = sorted({len(r[3]) for r in rows})
|
||||
print(f" ⚠ RELATIVE reading on one harness: {nseeds if len(nseeds) > 1 else nseeds[0]} "
|
||||
f"seed group(s) per arm, scored against this corpus's own held-out split. "
|
||||
f"It is not an absolute-band claim and corroborates nothing on its own.")
|
||||
return 0
|
||||
|
||||
|
||||
|
||||
Reference in New Issue
Block a user