# `[2026-09-17]` lv-hemingway: SHIPPED on ckpt850 — the line's first clean voice pass, and one axis that needs reading **Status: SHIPPED 2026-09-17 03:33 as `lv-hemingway` on `vllm-voices` (fv-ml1 GPU0 :8027), checkpoint-850.** Seat healthy 190 s after recreate, four models served (`voices-base`, `lv-yarros`, `lv-bronte`, `lv-hemingway`), GPU0 96,092 → **96,090 MiB** — a LoRA rides inside the existing seat and costs nothing. Adapter verified byte-identical to the checkpoint by sha256 across two hops. Gate design **pre-registered before any generation existed**: `scripts/hemingway-corpus/GATE-PREREG.md`, commit `0bb4938`. ## The gate result — 3 arms × 60 held-out beats × 4 seeds = 240 generations per arm | axis | result | numbers | |---|---|---| | **A. VOICE** | ✅ **PASS, 6.4×** | +0.413 delta_cb vs base, pairwise floor 0.064. Also clears the OLD all-arms floor (0.113) — **this verdict does not depend on the rule change** | | **B. NOT COPIED** | ⚠ **content clean, rate 7× the author's own** | 0.07 hit-rate, mean-longest 0.6, **max 9 words**. Base 0.00, **held-out Hemingway 0.01** | | **C. NO DAMAGE** | ✅ PASS | ran-on +0.08, on-beat −0.14, both inside a 0.217 floor; in-band 0.79 vs base 0.05 | ``` same-author target (held-out Hemingway vs itself) delta_cb 0.364 <- best achievable ckpt1750 0.439 ckpt850 (SHIPPED) 0.511 base-unadapted 0.924 ``` ⭐ **THE STRONGEST VOICE RESULT IN THE LINE. The span is 0.924 → 0.364 = 0.560 and ckpt850 closed 73.8% of it (ckpt1750 86.6%)**, against lv-bronte's 48%. Power came from the corpus, not from a better method: 173 in-band val pairs allowed a **60-beat** fixture where Brontë had 44 in-band and could only run 30. ## ⚠⚠ AXIS B — THE COMFORTABLE EXPLANATION WAS WRONG, AND THE CONTROL IS THE ARTIFACT `memorization_check.py` uses the **base-unadapted arm** as its negative control, and on this corpus that control is weak in one direction only — **it makes an innocent arm look guilty.** Base writes 18,035 words of *summary*; the adapted arms write 27,413 of *pastiche*. Text that does not imitate the register cannot collide with its n-grams, so base's 0.00 partly measures "different register", not "did not memorise". The obvious hypothesis was that Hemingway's plain, high-frequency register makes 8-gram collisions inevitable for any arm that learns it. **That hypothesis is refutable, was tested, and is FALSE.** New control: **held-out Hemingway — the author himself, val text no arm trained on — scored against the train split**, chunked to the generations' own median length (101 words) so the comparison is like for like. ``` sample n hit-rate mean-longest max HELD-OUT HEMINGWAY (never trained) 370 0.01 0.1 10 base-unadapted 240 0.00 0.0 0 ckpt1750 240 0.08 0.7 9 ckpt850 (SHIPPED) 240 0.07 0.6 9 positive control (train vs train) 160 <- not blind ``` ⭐⭐ **The adapter reproduces train-corpus word sequences ~7× more often than the author reproduces himself.** If the register explained it, real Hemingway would collide at the same rate; it collides at 0.01. ⭐ **And the exposure is still nil, which is a different question from the rate.** All 19 matched runs were READ, not counted. Every one is stock dialogue — `i don t think so the girl said`, `came over and sat down at the table`, `how do you feel i feel very well`. No plot, no imagery, no distinctive phrase, **no proper noun** (the one name-shaped hit, `swift tristan`, is the RENAMED invented name, not Hemingway's). The longest run is **9 words — shorter than the 10-word run genuinely unseen Hemingway shares with the train split by coincidence.** What is being reproduced is the *grammar of his dialogue*, which is the thing the adapter exists to learn, rendered in the commonest words in English. **Elevated rate, zero protectable content.** Hemingway is in copyright; the in-line precedent is lv-yarros, also in copyright, shipped at 0.10 against a 0.07 control. Unload is 0.003 s and one compose line. ⚠ **The durable lesson is about the instrument, not this adapter: a negative control that differs from the candidate in a way CORRELATED with the metric is not a control.** Always ask what the metric returns for a known-innocent sample *in the same register*. ## Why ckpt850 and NOT ckpt1750, the loss minimum ckpt1750 has the better point estimate on voice (0.439 vs 0.511) and **it is not usable**: ``` gap between candidates 0.072 pairwise floor max(0.113, 0.050) 0.113 -> NOT resolvable ``` Indistinguishable, so the pre-registered tiebreak falls to the axes that resolve — and **ckpt850 wins every one**: | | ckpt850 (shipped) | ckpt1750 | |---|---|---| | seed spread | **0.050** | 0.113 — **2.3× wider** | | memorisation hit-rate / mean-longest | **0.07 / 0.6** | 0.08 / 0.7 | | ran-on | **0.08** | 0.12 | | epoch | **0.959** | 1.973 | ckpt1750's spread is one seed: 0.491, 0.449, 0.468, then **0.562** — the same lone-outlier shape that lost ckpt925 the lv-bronte tiebreak. ⭐ **THE TWO-EPOCH RECIPE DID NOT TRANSFER HERE EITHER — it is now 0 for 2.** Hemingway's minimum really is step 1750, but step 850 is **+0.0040 against a 0.0044 median neighbour jitter**, with three checkpoints inside one jitter of the best. Epoch 2 buys nothing that resolves and costs 2.3× the variance. Only the epoch-3 collapse is robust: **+0.0762 = 17.4× jitter**, which is why `adapter/` was never gated. **Stop carrying "two epochs on a three-epoch schedule" forward; read the curve and prefer the earlier tied checkpoint.** ## Pre-flight: the beat leak IS present in Hemingway, and the fixture is clean `audit_pairs_sourcenames.py` (new, commit `0bb4938`) closes the blind spot `leak_gate.py` has by construction. Controls green every run: 941/941 surfaces found in the unrenamed source, nonce absent from both trees, 6/6 planted names detected. ``` train beats 70 of 7,094 (0.96%) Santiago x16, Catherine x7, Rinaldi x3, Brett, Harry, Jake, Pablo, Nick, Maria, Helen ... 36 distinct train responses 0 of 7,294 -- the rename itself held perfectly val beats 0 of 200 -- THE EVAL FIXTURE IS CLEAN; the gate is unconfounded ``` ⭐ The beat-only signature is exactly lv-bronte's. **Yarros's and Hemingway's earlier clean runs were never evidence of immunity** — they predate the detector. **Cross-validated on real data** where the answer was already recorded: the fixed Brontë pairs return **0 of 3,858** (matching "0 leaks across 3,858 pairs"), and `pairs-full.CONTAMINATED.jsonl` returns **15 of 792 = 1.89%** with Rochester ×6, Jane, Brocklehurst ×2, Beck, Fairfax, Burns, Helen, Eyre — against a record of "13 of the first 714 beats (1.8%)" with the same names. An independently written instrument reproducing a documented finding at the right magnitude is what makes its zeroes mean *absent*, not *blind*. `--filter-out` produces a clean **7,024-pair** set in one command (70 dropped, 0.99%), verified by re-audit at 0 of 7,024. **A retrain on it is the operator's call, not done.** ## ⚠ A SECOND corpus defect, measured and NOT acted on `audit_entity_map.py` (new, commit `051b99e`) is the mirror of `audit_stoplist.py`: it finds surfaces wrongly held **IN** the entity map, which `leak_gate.py` cannot see because it only ever asks whether the author's names are GONE, never whether non-names were spared. ``` positive control `other` 764/1356 article-preceded = 0.56 negative control 100 honorific-confirmed people, highest Inglés 0.26, bulk 0.00-0.06 FLAGGED 130 of 946 surfaces · 1,616 instances · 0.162% of corpus words ``` `African`, `Chinese`, `Basques`, `Republican`, `Communist`, `X-ray`, `Coca-Cola`, `Ritz`, `Prado`, `Cezanne` were all renamed into invented proper nouns. **Some flags are correct renames** — `the Widow`, `the Informer` are genuine Hemingway epithet-names — so every hit is reported for reading, never auto-removed. Plus **16 bare initials in the map**, `C` at 274 occurrences: the same class as the `G` caught by hand about to be renamed 248 times. At 0.162% of words this did not block the ship. It is the thing to fix first if a corpus rebuild ever happens. ## Artefacts `gx10:~/lv-hemingway/` (corpus-clean, corpus-renamed, beats-hemingway-60.json + sidecar, eval-hemingway.sh, voice-prep.py, eval.log), `gx10:~/r49-runs/hemingway-4b-pairs-3ep/` (54 checkpoints kept), `gx10:~/r49-runs/hemingway-eval/` (three arms × 240 generations, memorization.txt, voice_distance.txt, score.*.txt). `fv-ml1:/tank/aimodels/voice-adapters/lv-hemingway-4b-v1/` (adapter + a README carrying the axis-B caveat, so it cannot be read as clean by anyone who finds the adapter without this). Commits `0bb4938` `051b99e` `5e66114` `2e9b118`. Related: [[2026-09-17-lv-bronte-gate]], [[2026-09-16-lv-hemingway-corpus]], [[2026-09-16-lv-voices-line]], [[2026-09-16-voices-seat-lora]], [[2026-09-17-beat-contamination-leak]].