voices-seat: ship lv-hemingway (ckpt850), and replace the memorisation control that passed it

Live on vllm-voices (fv-ml1 GPU0 :8027) beside voices-base, lv-yarros and lv-bronte.
Healthy 190 s after recreate, four models served, GPU0 96,092 -> 96,090 MiB. The adapter
was verified byte-identical to checkpoint-850 by sha256 across both transfer hops, and the
seat was verified by generating, not by reading its config: base emits 170 words of <think>
planning and never writes the passage, lv-hemingway writes the scene.

Gate design was pre-registered before any generation existed (0bb4938). Three arms, 60
held-out beats, 4 seeds, 240 generations per arm.

  A. VOICE   PASS 6.4x   +0.413 delta_cb, pairwise floor 0.064 -- and it clears the OLD
                         all-arms floor (0.113) too, so this verdict does not lean on the
                         rule change. Closes 73.8% of the span between the unadapted
                         carrier and held-out Hemingway itself; lv-bronte closed 48%.
  B. NOT COPIED  see below
  C. NO DAMAGE   PASS    ran-on +0.08, on-beat -0.14, both inside a 0.217 floor

AXIS B: THE NEGATIVE CONTROL WAS THE WRONG ONE, AND FIXING IT MADE THE RESULT WORSE, NOT
BETTER. memorization_check.py uses the base-unadapted arm as its control. Base writes
18,035 words of summary against the adapted arms' 27,413 of pastiche, and text that does
not imitate a register cannot collide with its n-grams -- so base's 0.00 measures "different
register", not "did not memorise". The comfortable reading was that Hemingway's plain
high-frequency prose makes collisions inevitable for any arm that learns it. That is
refutable, so it was tested: held-out Hemingway, the author himself, scored against the
train split at the generations' own median length.

  HELD-OUT HEMINGWAY (never trained)   370 chunks   0.01 hit-rate   mean-longest 0.1   max 10
  base-unadapted                       240 gens     0.00                        0.0        0
  ckpt850 (shipped)                    240 gens     0.07                        0.6        9
  positive control (train vs train)                                             160

The hypothesis is false: the adapter reproduces train n-grams ~7x more often than the
author reproduces himself. That is real and is on the record. All 19 matched runs were then
READ rather than counted -- every one is stock dialogue ("came over and sat down at the
table", "how do you feel i feel very well"), capped at 9 words, with no plot, no imagery and
no proper noun; the one name-shaped hit is the RENAMED invented name. Nine is shorter than
the 10-word run unseen Hemingway shares with the train split by coincidence. Elevated rate,
zero protectable content. Hemingway is in copyright; lv-yarros is the in-line precedent,
also in copyright, shipped at 0.10 against a 0.07 control. Unload is 0.003 s.

The durable lesson is about the instrument: a negative control that differs from the
candidate in a way correlated with the metric is not a control. memorization_selfsim.py and
memorization_dump_matches.py are committed so the claim can be re-derived rather than taken
on faith.

SHIPPED ckpt850, NOT the loss minimum at step 1750. The two are indistinguishable on voice
-- 0.072 apart against a 0.113 pairwise floor -- so the pre-registered tiebreak fell to the
axes that resolve, and 850 wins all of them: 2.3x tighter seed spread (0.050 vs 0.113),
lower memorisation, less ran-on, half an epoch less overfit. ckpt1750's spread is one seed
(0.491, 0.449, 0.468, then 0.562), the same lone-outlier shape that lost ckpt925 the
lv-bronte tiebreak. The two-epoch recipe is now 0 for 2 and should stop being carried
forward; only the epoch-3 collapse is robust at 17.4x jitter.

servers/fv-ml1/ssh-target was a bare IP, so deploy-stack.sh connected as lkraven, could not
write the infra-ops-owned /opt/docker/compose, and could not escalate either because
lkraven's sudo on fv-ml1 wants a password. Now infra-ops@10.251.50.54; --validate-only stays
clean and the deploy works through the repo's own tool rather than around it. Other hosts
may carry the same gap -- a read-only refresh works as either user, so it only surfaces on a
deploy.
This commit is contained in:
vh
2026-09-17 03:36:29 -07:00
parent 5e6611466c
commit 300ecc1276
8 changed files with 462 additions and 29 deletions
@@ -0,0 +1,82 @@
#!/bin/bash
# lv-hemingway v2 gate. Design is FROZEN in scripts/hemingway-corpus/GATE-PREREG.md
# and was written before this script ever ran. Do not edit the arms, the fixture
# size or the seeds to chase a result -- re-run, do not re-tune.
#
# THREE ARMS. ckpt1750 is the eval-loss minimum (2.2783, epoch 1.973). ckpt850
# (2.2823, epoch 0.959) is +0.0040 against a 0.0044 median neighbour jitter, i.e.
# TIED -- the loss curve cannot pick between them, and on Bronte the earlier
# epoch-1 checkpoint won the tiebreak on the axes that do resolve. `base` is the
# negative control for memorisation (it never saw the corpus) and the voice
# baseline. adapter/ (epoch 3.0, +0.0762 = 17.4x jitter) is NOT gated: that one
# the loss curve settles on its own.
set -o pipefail
cd ~/lv-hemingway || exit 1
PY=/home/infra-ops/ml/.venv/bin/python
RUN=~/r49-runs/hemingway-4b-pairs-3ep
OUT=~/r49-runs/hemingway-eval
BEATS=beats-hemingway-60.json
PROV=pairs/pairs-full.jsonl.provenance.json
SEEDS="1234 5678 9012 3456"
mkdir -p "$OUT"
log(){ echo "[eval $(date +%H:%M:%S)] $*"; }
# --system-from is MANDATORY. The harness's built-in SYS is Yarros's; driving a
# Hemingway arm with it would confound the adapter change with a prompt change.
# This binds the eval prompt to the one the run actually trained under (verified:
# run provenance system_prompt == pairs provenance system_prompt).
[ -f "$PROV" ] || { log "MISSING $PROV"; exit 1; }
[ -f "$BEATS" ] || { log "MISSING $BEATS"; exit 1; }
for arm in base:NONE ckpt1750:$RUN/checkpoints/checkpoint-1750 ckpt850:$RUN/checkpoints/checkpoint-850; do
name=${arm%%:*}; path=${arm#*:}
if [ "$name" != "base" ] && [ ! -d "$path" ]; then log "MISSING $path"; exit 1; fi
if [ -s "$OUT/beats5.$name.jsonl" ]; then
log "arm $name already has $(wc -l < "$OUT/beats5.$name.jsonl") generations -- skipping"
continue
fi
log "arm $name"
if [ "$name" = "base" ]; then
"$PY" scripts/r49-corpus/gen_beats_chat_yarros.py \
--base ~/carriers/Qwen3-4B-Instruct --beats "$BEATS" \
--out "$OUT/beats5.$name.jsonl" --arm "$name" --seeds $SEEDS \
--system-from "$PROV" || exit 1
else
"$PY" scripts/r49-corpus/gen_beats_chat_yarros.py \
--base ~/carriers/Qwen3-4B-Instruct --adapter "$path" --beats "$BEATS" \
--out "$OUT/beats5.$name.jsonl" --arm "$name" --seeds $SEEDS \
--system-from "$PROV" || exit 1
fi
log " $(wc -l < "$OUT/beats5.$name.jsonl") generations"
done
log "AXIS B -- MEMORISATION (corpus = the renamed copies the adapter trained on)"
# --corpus and --eval-dir are passed explicitly: the script's Yarros defaults would
# compare a Hemingway arm against the YARROS corpus and report a clean zero that
# means "different book", not "did not memorise".
"$PY" scripts/yarros-corpus/memorization_check.py \
--eval-dir "$OUT" --corpus corpus-renamed/copies --glob 'beats5.*.jsonl' --strip 'beats5.' -n 8 \
2>&1 | tee "$OUT/memorization.txt"
log "AXIS C -- DAMAGE (ran-on / out-of-band), ckpt1750 vs base"
"$PY" scripts/yarros-corpus/score_beats.py \
--arm base="$OUT/beats5.base.jsonl" \
--arm ckpt1750="$OUT/beats5.ckpt1750.jsonl" \
--arm ckpt850="$OUT/beats5.ckpt850.jsonl" \
--baseline base --candidate ckpt1750 --metric-source raw \
--out "$OUT/score.ckpt1750.json" 2>&1 | tee "$OUT/score.ckpt1750.txt"
log "AXIS C -- DAMAGE, ckpt850 vs base"
"$PY" scripts/yarros-corpus/score_beats.py \
--arm base="$OUT/beats5.base.jsonl" \
--arm ckpt1750="$OUT/beats5.ckpt1750.jsonl" \
--arm ckpt850="$OUT/beats5.ckpt850.jsonl" \
--baseline base --candidate ckpt850 --metric-source raw \
--out "$OUT/score.ckpt850.json" 2>&1 | tee "$OUT/score.ckpt850.txt"
log "AXIS A -- VOICE (delta_cb vs held-out Hemingway)"
"$PY" voice-prep.py || exit 1
"$PY" scripts/r49-corpus/voice_distance.py corpus-renamed "$OUT" --author Hemingway \
2>&1 | tee "$OUT/voice_distance.txt"
echo "rc=0" > ~/lv-hemingway/.eval-complete
log "done"
+47
View File
@@ -0,0 +1,47 @@
"""Adapt the lv-hemingway arms for voice_distance.py, and refuse if its reference is empty.
voice_distance.py expects:
- files matching voice.<arm>.jsonl in the eval dir
- a `continuation` field per record (gen_beats_chat writes `raw`)
- a `seed` field (present)
- the control arm's NAME to contain the substring "unadapted" -- it picks the
control by that substring, so `base` alone would leave the control
unidentified and the whole vs-control table would silently be empty.
- a reference built from corpus records whose split == "val"
"""
import json, sys
from collections import Counter
from pathlib import Path
CORP = Path("/home/infra-ops/lv-hemingway/corpus-renamed/copies")
EVAL = Path("/home/infra-ops/r49-runs/hemingway-eval")
c = Counter()
n = 0
for p in sorted(CORP.glob("*.jsonl")):
for line in p.read_text(encoding="utf-8").splitlines():
if not line.strip():
continue
r = json.loads(line)
n += 1
c[r.get("split")] += 1
print(f"corpus records: {n} split values: {dict(c)}")
if c.get("val", 0) == 0:
print("== REFUSING: no split=val records; voice_distance would build an EMPTY reference")
print(" and every delta_cb would be meaningless rather than absent.")
sys.exit(1)
NAMES = {"base": "base-unadapted", "ckpt1750": "ckpt1750", "ckpt850": "ckpt850"}
for src_arm, out_arm in NAMES.items():
src = EVAL / f"beats5.{src_arm}.jsonl"
if not src.exists():
print(f"== missing {src}")
sys.exit(1)
rows = [json.loads(l) for l in src.read_text(encoding="utf-8").splitlines() if l.strip()]
out = EVAL / f"voice.{out_arm}.jsonl"
with out.open("w", encoding="utf-8") as fh:
for r in rows:
fh.write(json.dumps({"id": r["id"], "seed": r["seed"],
"continuation": r["raw"]}, ensure_ascii=False) + "\n")
print(f" {src.name} -> {out.name} ({len(rows)} records)")
print("ready")
@@ -0,0 +1,38 @@
"""What ARE the verbatim 8-gram hits? A rate is not a judgement.
0.09 against a 0.00 control reads alarming; 0.00 against 0.09 could also be an
artefact of the control writing a different register entirely (the base arm wrote
18,035 words of summary against the adapted arms' 27,413 of pastiche, and text that
does not imitate the style trivially fails to match its n-grams). The only way to
tell a memorised passage from a common English run is to read them.
"""
import json, re, sys, pathlib
from collections import Counter
CORP = pathlib.Path(sys.argv[1]); EVAL = pathlib.Path(sys.argv[2]); N = 8
def norm(t): return re.findall(r"[a-z']+", t.lower())
words = []
for f in sorted(CORP.glob("*.copy0.jsonl")):
for l in f.read_text(encoding="utf-8").splitlines():
words.extend(norm(json.loads(l)["text"]))
grams = {" ".join(words[i:i+N]) for i in range(len(words)-N+1)}
print(f"corpus {len(words):,} words, {len(grams):,} distinct {N}-grams\n")
for arm in ("base", "ckpt1750", "ckpt850"):
p = EVAL/f"beats5.{arm}.jsonl"
if not p.exists(): continue
rows = [json.loads(l) for l in p.read_text(encoding="utf-8").splitlines() if l.strip()]
found = Counter()
for r in rows:
w = norm(r["raw"]); i = 0
while i <= len(w)-N:
g = " ".join(w[i:i+N])
if g in grams:
k = N
while i+k < len(w) and " ".join(w[i+k-N+1:i+k+1]) in grams: k += 1
found[" ".join(w[i:i+k])] += 1
i += k
else:
i += 1
print(f"=== {arm}: {len(rows)} gens, {sum(found.values())} matched runs, {len(found)} distinct")
for g, c in found.most_common(40):
print(f" x{c} [{len(g.split())}w] {g}")
print()
@@ -0,0 +1,71 @@
"""THE CONTROL AXIS B WAS MISSING: how much does the author collide with HIMSELF?
memorization_check.py compares each arm against the TRAIN corpus and uses the
base-unadapted arm as the negative control. On Hemingway that control is weak in a way
it was not on Brontë, and the weakness runs one way only -- it makes an innocent arm
look guilty:
* base-unadapted writes 18,035 words of summary prose; the adapted arms write 27,413
of pastiche. Text that does not imitate the register cannot collide with its
n-grams, so 0.00 measures "different register", not "did not memorise".
* Hemingway's register IS short, plain, high-frequency English with heavy unattributed
dialogue. An arm that SUCCEEDS at the voice task must start colliding with 8-grams
built out of the commonest word sequences in the language.
So the honest reference is not the base arm. It is HELD-OUT HEMINGWAY -- val text no arm
trained on, written by the author himself, which by construction did not memorise the
train split. Whatever rate that produces is the floor the metric returns for innocent
text in this register, and only an excess over THAT is evidence of copying.
Chunks are cut to the generations' own length so the comparison is like-for-like: a
longer sample has more chances to collide.
"""
import json, re, sys, pathlib, statistics as st
CORP = pathlib.Path(sys.argv[1]); EVAL = pathlib.Path(sys.argv[2]); N = 8
def norm(t): return re.findall(r"[a-z']+", t.lower())
train_words, val_texts = [], []
for f in sorted(CORP.glob("*.copy0.jsonl")):
for l in f.read_text(encoding="utf-8").splitlines():
r = json.loads(l)
(val_texts.append(r["text"]) if r.get("split") == "val" else train_words.extend(norm(r["text"])))
grams = {" ".join(train_words[i:i+N]) for i in range(len(train_words)-N+1)}
print(f"train (copy0, split=train): {len(train_words):,} words, {len(grams):,} distinct {N}-grams")
def longest(w):
best = 0; i = 0
while i <= len(w)-N:
if " ".join(w[i:i+N]) in grams:
k = N
while i+k < len(w) and " ".join(w[i+k-N+1:i+k+1]) in grams: k += 1
best = max(best, k); i += 1
else: i += 1
return best
arm_lens = []
for arm in ("ckpt1750", "ckpt850", "base"):
p = EVAL/f"beats5.{arm}.jsonl"
if p.exists():
arm_lens += [len(norm(json.loads(l)["raw"]))
for l in p.read_text(encoding="utf-8").splitlines() if l.strip()]
CHUNK = int(st.median(arm_lens))
print(f"median generation length across arms: {CHUNK} words -- val is chunked to match\n")
vw = norm("\n".join(val_texts))
chunks = [vw[i:i+CHUNK] for i in range(0, len(vw)-CHUNK+1, CHUNK)]
longs = [longest(c) for c in chunks]
hits = sum(1 for x in longs if x >= N)
print(f"{'sample':<34} {'n':>5} {'hit-rate':>9} {'mean-longest':>13} {'max':>5}")
print("-"*70)
print(f"{'HELD-OUT HEMINGWAY (never trained)':<34} {len(chunks):>5} {hits/len(chunks):>9.2f} "
f"{sum(longs)/len(longs):>13.1f} {max(longs):>5}")
for arm in ("base", "ckpt1750", "ckpt850"):
p = EVAL/f"beats5.{arm}.jsonl"
if not p.exists(): continue
rows = [json.loads(l) for l in p.read_text(encoding="utf-8").splitlines() if l.strip()]
L = [longest(norm(r["raw"])) for r in rows]
h = sum(1 for x in L if x >= N)
print(f"{arm:<34} {len(rows):>5} {h/len(rows):>9.2f} {sum(L)/len(L):>13.1f} {max(L):>5}")
print(f"\npositive control (a train slice vs train): longest = {longest(train_words[1000:1160])} "
f"(must be large, else blind)")