eval harness: sample the beat fixture from held-out val, and bind the eval prompt to the trained one
Two harness defects that would each make a voice number uninterpretable. build_beat_fixture.py — the fixture is now SAMPLED from the val split rather than hand-written. The original BabyYarros fixture was five hand-written beats about a stray dog and a kitten: wrong genre, so 'He licked her clean' came back as explicit sex from a romantasy adapter, and n=5 had a noise floor of 0.800 that manufactured a +0.45 result which collapsed to +0.08 at n=120. Sampling from val makes it in-genre and held out by construction, spread across works so a naive head(30) is not one novel. Refuses outright if the pairs carry any split but val, because a fixture drawn from training data makes every downstream number a memorisation measurement wearing a voice label. gen_beats_chat_yarros.py --system-from — the SYS constant in this harness is Yarros's. Driving a Bronte or Hemingway adapter with it measures the arm under a system prompt it was never trained on and confounds the carrier change with a prompt change. Rather than duplicate the register table and rely on whoever runs it to pick the matching one, read the prompt out of the pair build's own provenance, which is the artefact that records what the adapter actually saw.
This commit is contained in:
@@ -44,8 +44,27 @@ ap.add_argument("--top-p", type=float, default=0.95)
|
||||
# arm stays byte-reproducible; the flag exists so the casing can be MEASURED as its own
|
||||
# variable instead of being confounded with the adapter it is being used to judge.
|
||||
ap.add_argument("--user-prefix", default="BEAT: ")
|
||||
# ⭐ TRAIN/EVAL PROMPT PARITY, MADE STRUCTURAL RATHER THAN REMEMBERED.
|
||||
# The hardcoded SYS above is Yarros's. Driving a Brontë or Hemingway adapter with it
|
||||
# would confound the carrier change with a PROMPT change -- the arm would be measured
|
||||
# under a system prompt it was never trained on, and the resulting delta would be
|
||||
# uninterpretable. Rather than duplicate the register table here and rely on whoever
|
||||
# runs this to pick the matching one, read the prompt straight out of the pair build's
|
||||
# own provenance, which is the artefact that records what the adapter actually saw.
|
||||
ap.add_argument("--system-from", default=None,
|
||||
help="pairs provenance json; its `system_prompt` replaces SYS. Use this for "
|
||||
"any corpus but Yarros — it guarantees the eval drives the adapter under "
|
||||
"the prompt it was trained on.")
|
||||
a = ap.parse_args()
|
||||
|
||||
if a.system_from:
|
||||
_prov = json.loads(Path(a.system_from).read_text())
|
||||
_sp = _prov.get("system_prompt")
|
||||
if not _sp:
|
||||
raise SystemExit(f"== {a.system_from} carries no `system_prompt` -- refusing to guess")
|
||||
SYS = _sp
|
||||
print(f" system prompt from {a.system_from}: {SYS[:80]}...")
|
||||
|
||||
beats = json.loads(Path(a.beats).read_text())
|
||||
tok = AutoTokenizer.from_pretrained(a.base)
|
||||
if tok.chat_template is None:
|
||||
|
||||
Reference in New Issue
Block a user