Live on vllm-voices (fv-ml1 GPU0 :8027) beside voices-base, lv-yarros and lv-bronte.
Healthy 190 s after recreate, four models served, GPU0 96,092 -> 96,090 MiB. The adapter
was verified byte-identical to checkpoint-850 by sha256 across both transfer hops, and the
seat was verified by generating, not by reading its config: base emits 170 words of <think>
planning and never writes the passage, lv-hemingway writes the scene.
Gate design was pre-registered before any generation existed (0bb4938). Three arms, 60
held-out beats, 4 seeds, 240 generations per arm.
A. VOICE PASS 6.4x +0.413 delta_cb, pairwise floor 0.064 -- and it clears the OLD
all-arms floor (0.113) too, so this verdict does not lean on the
rule change. Closes 73.8% of the span between the unadapted
carrier and held-out Hemingway itself; lv-bronte closed 48%.
B. NOT COPIED see below
C. NO DAMAGE PASS ran-on +0.08, on-beat -0.14, both inside a 0.217 floor
AXIS B: THE NEGATIVE CONTROL WAS THE WRONG ONE, AND FIXING IT MADE THE RESULT WORSE, NOT
BETTER. memorization_check.py uses the base-unadapted arm as its control. Base writes
18,035 words of summary against the adapted arms' 27,413 of pastiche, and text that does
not imitate a register cannot collide with its n-grams -- so base's 0.00 measures "different
register", not "did not memorise". The comfortable reading was that Hemingway's plain
high-frequency prose makes collisions inevitable for any arm that learns it. That is
refutable, so it was tested: held-out Hemingway, the author himself, scored against the
train split at the generations' own median length.
HELD-OUT HEMINGWAY (never trained) 370 chunks 0.01 hit-rate mean-longest 0.1 max 10
base-unadapted 240 gens 0.00 0.0 0
ckpt850 (shipped) 240 gens 0.07 0.6 9
positive control (train vs train) 160
The hypothesis is false: the adapter reproduces train n-grams ~7x more often than the
author reproduces himself. That is real and is on the record. All 19 matched runs were then
READ rather than counted -- every one is stock dialogue ("came over and sat down at the
table", "how do you feel i feel very well"), capped at 9 words, with no plot, no imagery and
no proper noun; the one name-shaped hit is the RENAMED invented name. Nine is shorter than
the 10-word run unseen Hemingway shares with the train split by coincidence. Elevated rate,
zero protectable content. Hemingway is in copyright; lv-yarros is the in-line precedent,
also in copyright, shipped at 0.10 against a 0.07 control. Unload is 0.003 s.
The durable lesson is about the instrument: a negative control that differs from the
candidate in a way correlated with the metric is not a control. memorization_selfsim.py and
memorization_dump_matches.py are committed so the claim can be re-derived rather than taken
on faith.
SHIPPED ckpt850, NOT the loss minimum at step 1750. The two are indistinguishable on voice
-- 0.072 apart against a 0.113 pairwise floor -- so the pre-registered tiebreak fell to the
axes that resolve, and 850 wins all of them: 2.3x tighter seed spread (0.050 vs 0.113),
lower memorisation, less ran-on, half an epoch less overfit. ckpt1750's spread is one seed
(0.491, 0.449, 0.468, then 0.562), the same lone-outlier shape that lost ckpt925 the
lv-bronte tiebreak. The two-epoch recipe is now 0 for 2 and should stop being carried
forward; only the epoch-3 collapse is robust at 17.4x jitter.
servers/fv-ml1/ssh-target was a bare IP, so deploy-stack.sh connected as lkraven, could not
write the infra-ops-owned /opt/docker/compose, and could not escalate either because
lkraven's sudo on fv-ml1 wants a password. Now infra-ops@10.251.50.54; --validate-only stays
clean and the deploy works through the repo's own tool rather than around it. Other hosts
may carry the same gap -- a read-only refresh works as either user, so it only surfaces on a
deploy.
39 lines
1.7 KiB
Python
39 lines
1.7 KiB
Python
"""What ARE the verbatim 8-gram hits? A rate is not a judgement.
|
|
|
|
0.09 against a 0.00 control reads alarming; 0.00 against 0.09 could also be an
|
|
artefact of the control writing a different register entirely (the base arm wrote
|
|
18,035 words of summary against the adapted arms' 27,413 of pastiche, and text that
|
|
does not imitate the style trivially fails to match its n-grams). The only way to
|
|
tell a memorised passage from a common English run is to read them.
|
|
"""
|
|
import json, re, sys, pathlib
|
|
from collections import Counter
|
|
CORP = pathlib.Path(sys.argv[1]); EVAL = pathlib.Path(sys.argv[2]); N = 8
|
|
def norm(t): return re.findall(r"[a-z']+", t.lower())
|
|
words = []
|
|
for f in sorted(CORP.glob("*.copy0.jsonl")):
|
|
for l in f.read_text(encoding="utf-8").splitlines():
|
|
words.extend(norm(json.loads(l)["text"]))
|
|
grams = {" ".join(words[i:i+N]) for i in range(len(words)-N+1)}
|
|
print(f"corpus {len(words):,} words, {len(grams):,} distinct {N}-grams\n")
|
|
for arm in ("base", "ckpt1750", "ckpt850"):
|
|
p = EVAL/f"beats5.{arm}.jsonl"
|
|
if not p.exists(): continue
|
|
rows = [json.loads(l) for l in p.read_text(encoding="utf-8").splitlines() if l.strip()]
|
|
found = Counter()
|
|
for r in rows:
|
|
w = norm(r["raw"]); i = 0
|
|
while i <= len(w)-N:
|
|
g = " ".join(w[i:i+N])
|
|
if g in grams:
|
|
k = N
|
|
while i+k < len(w) and " ".join(w[i+k-N+1:i+k+1]) in grams: k += 1
|
|
found[" ".join(w[i:i+k])] += 1
|
|
i += k
|
|
else:
|
|
i += 1
|
|
print(f"=== {arm}: {len(rows)} gens, {sum(found.values())} matched runs, {len(found)} distinct")
|
|
for g, c in found.most_common(40):
|
|
print(f" x{c} [{len(g.split())}w] {g}")
|
|
print()
|