diff --git a/scripts/mccarthy-corpus/RUNBOOK.md b/scripts/mccarthy-corpus/RUNBOOK.md index eaba69f..b39f4db 100644 --- a/scripts/mccarthy-corpus/RUNBOOK.md +++ b/scripts/mccarthy-corpus/RUNBOOK.md @@ -1,4 +1,4 @@ -# lv-mccarthy — corpus → gate → pairs +# lv-mccarthy — corpus → leak gate → pairs → train → v2 gate Six works, 167 units, 584,684 words. Built on **pfi-gx10** under `~/lv-mccarthy/` as `infra-ops`, except **D1, which must run on nh3-dev**: the builder reads the @@ -106,3 +106,105 @@ Cross-checked on the two shipped corpora: **lv-bronte is clean** of this class; - **Blood Meridian sets a dash-separated chapter argument under each roman numeral.** 131 paragraphs; `--drop-leading-heading` now consumes them. - Two truncated catalogue rows are excluded in favour of complete siblings. + +## D5 — train the adapter + +```bash +# ON gx10, ~/lv-mccarthy. 1,380 steps, 2h30m47s on the GB10. +./launch-train.sh # scripts/mccarthy-corpus/launch-train.sh +``` + +Qwen3-4B-Instruct, rank 32 / alpha 64, lr 1e-4, seq 1536, batch 1 × accum 8, 3 epochs, +seed 4919 → `~/r49-runs/mccarthy-4b-pairs-3ep/`. The two deliberate deviations from the +script defaults (`--save-total-limit 60`, `--eval-steps/--save-steps 25`) are argued in the +launch script's own header and exist to keep the epoch-1 and epoch-2 checkpoints alive for +the tiebreak; the default limit of 12 would have deleted both. + +**Result, 2026-09-18 00:49 PT:** `train_loss` 2.172, end-of-run `eval_loss` 2.4594, 56 eval +points, **median neighbour jitter 0.00393**. + +``` +ckpt900 2.38706 ep1.958 the minimum 875 +0.4x · 850 +1.1x · 925 +1.5x = tied +ckpt450 2.4063 ep0.980 +4.9x jitter the epoch-1 local minimum, NOT tied +adapter/ 2.4594 ep3.000 +18.4x jitter the epoch-3 collapse, resolved by the curve +``` + +⚠ **The curve does not drift into the epoch-3 collapse, it STEPS**: 2.393 at step 925, +2.457 at 950, then flat for the remaining 430 steps. Same shape as Hemingway's. + +⚠ **`adapter/` is the epoch-3 weights.** Whatever ships is a checkpoint, not `adapter/`. + +### Provenance fields that look wrong and are not + +Checked when the run was first read, three days after it finished. All three reproduce on +the yarros and hemingway runs, so none is specific to this one: + +| field | reading | +|---|---| +| `"run": "r49-babyyarros-pairs-pilot"` | a hardcoded label in `train_pairs_lora.py`, not a stale copy. Cosmetic. | +| `harness_commit: ""` | the trainer never recorded its own commit, on any run. No claim rests on it. | +| `pairs_sha256_16` ≠ `sha256sum` of the pairs file | computed over the loaded records, not the raw bytes. Consistent and per-corpus-unique, so a valid cache key. | +| `pairs` recorded as the relative `pairs/pairs-full.jsonl` | resolves against the launch CWD. **Bind it by COUNT, not path**: `train_pairs` 3673 / `val_pairs_n` 269 match `lv-mccarthy/pairs/` exactly and match no other pair set on the box. | + +## D6 — the gate + +```bash +# ON gx10, ~/lv-mccarthy. 3 arms x 60 beats x 4 seeds = 720 generations, ~2h. +./eval-mccarthy.sh +``` + +**The design is FROZEN in `scripts/mccarthy-corpus/GATE-PREREG.md`, written and committed +before a single generation existed** (`9c8a4e9`, amended `b4ba731`). Do not edit the arms, +the fixture size or the seeds to chase a result — re-run, do not re-tune. + +Three arms: `base` (negative control + voice baseline), `ckpt900` (ship candidate), +`ckpt450` (prior test — it is **not** tied, and the pre-registration says so). Axis A voice +via `voice_distance.py`, axis B memorisation, axis C damage via `score_beats.py`. + +| deviation from the lv-hemingway gate | why | +|---|---| +| **fixture `--max-words 140`** (script default 150) | the mccarthy register asks for 90–140 and `score_beats.py` scores in-band at 90–140. 207 of 269 val pairs are in-band, across all six works. | +| **the run/pairs system-prompt equality is ASSERTED**, not claimed in a comment | Hemingway's script said "verified" in prose, which no reader can check. Here a mismatch exits 1. | +| **`--punct-report --secondary-normalised`** | McCarthy-specific; see below. | +| **axis B uses `--train-only --heldout-reference`** | the base-unadapted control is defective; see below. | + +### ⚠ The voice axis needed settling before it could run + +This corpus measures **0.0 quote marks per 10k** against Hemingway's 838, and `delta_cb` is +Burrows's Delta over CHARACTER BIGRAMS — so "emit no quotation marks" moves the metric a long +way without having learned a sentence. Settled in three parts, all pre-registered: + +1. **PRIMARY** — the `mccarthy` register NAMES the punctuation, and `--system-from` drives + the base control with the same prompt. The cheap win is handed to both sides. +2. **SECONDARY** — `--secondary-normalised` re-runs everything with punctuation stripped. A + conservative LOWER BOUND (it also strips sentence-length signal); reported, never the verdict. +3. **TRIGGER** — `--punct-report` checks whether base actually complied. Base quote density + above **100 per 10k** means it did not, the primary gap is partly the cheap win, and the + normalised read is promoted to load-bearing. The line was fixed before the table printed. + +### ⚠ The memorisation control shipped broken and is now fixed + +`memorization_check.py`'s negative control was `base-unadapted`, and the lv-hemingway gate +established that this is defective: base writes *summary*, the adapted arms write *pastiche*, +and text that does not imitate a register cannot collide with its n-grams. The correct +innocent sample is **the author himself** — held-out text, same register by construction. +That control was computed by hand for Hemingway and never committed. It is now +`--heldout-reference` (requiring `--train-only`, or the val text is scored against a gram set +containing itself, and the script refuses). + +Validated by reproducing Hemingway's hand-computed numbers to the digit — 370 samples, +hit-rate 0.01, mean-longest 0.1, max 10, 101-word median chunks. + +⚠ **An elevated rate is not by itself a no-ship — rate and exposure are different questions.** +Every matched run gets READ. Hemingway shipped at 7× the author's own rate because every match +was stock dialogue, max 9 words, no proper noun. **McCarthy is in copyright with a living +estate**, so a match carrying distinctive imagery or a proper noun is disqualifying in a way a +rate number alone is not. + +### A log tag that reads backwards on this corpus + +`gen_beats_chat_yarros.py` prints `RAN-ON` when its `STOP` regex finds no paragraph break in +the raw output. That heuristic was written for the Yarros register. **McCarthy's register asks +for continuous scene prose, so a single unbroken block is the target, not a defect** — expect +the tag on most generations and do not read it as damage. The axis-C metric is +`score_beats.py --metric-source raw`, where ran-on means `words > 140`, and it is unaffected.