memory: snapshot — lv-mccarthy training launched on gx10, and the next voice seat is measured rather than chosen

In-flight rewritten to the live training run (~150/1380, ETA ~00:45 PT) with the
--save-total-limit finding that would otherwise have deleted the epoch-1/epoch-2
checkpoints both prior gates were decided on.

Two decisions added: the next-seat ranking (Faulkner, Morrison, Chandler -- and the
finding that the corpus size ranking inverts the voice ranking, with King and Christie
as the two biggest non-candidates), and the romantasy register measured on the gate's
own char-bigram instrument (Yarros is the cluster outlier we already shipped; Maas is
the centroid and so the worst pick; Kenyon at 27 val units if the lane gets a seat).

Auto-archival: 4 entries moved to archival-memory.md; 4 held back by the open-deferred
guard.
This commit is contained in:
vh
2026-09-17 22:38:13 -07:00
parent 36f2e4dbbc
commit d94b5a1934
9 changed files with 380 additions and 174 deletions
+37
View File
@@ -0,0 +1,37 @@
#!/bin/bash
# lv-mccarthy D5 — pairs LoRA. Run on pfi-gx10 as infra-ops.
#
# Mirrors the lv-hemingway recipe deliberately: same carrier, seq-len 1536, rank 32 (alpha 64),
# lr 1e-4, 3 epochs, batch 1 x accum 8, seed 4919. Those are the script defaults; they are
# written out explicitly anyway so this file IS the record of what ran.
#
# ⚠ TWO DELIBERATE DEVIATIONS FROM THE DEFAULTS, both about the checkpoint tiebreak:
#
# --save-total-limit 60 The default is 12. At a 25-step cadence over ~1,377 planned steps
# that is 55 checkpoints, so a limit of 12 keeps only epochs ~2.35-3.0
# and DELETES the epoch-1 and epoch-2 checkpoints. Both lv-bronte and
# lv-hemingway were decided on exactly those: Hemingway's ckpt850
# (epoch 0.96) tied its eval-loss minimum ckpt1750 (epoch 1.97) to
# within the neighbour jitter, and on Brontë the earlier epoch-1
# checkpoint won the tiebreak on the axes that resolve. Losing them
# would remove the choice before it could be made. 55 x 757 MB is
# ~42 GB against 442 GB free.
#
# --eval-steps 25 Hemingway ran 50 on 887 steps/epoch = 0.056 epoch granularity.
# --save-steps 25 McCarthy has 459 steps/epoch, so 25 gives 0.054 -- matching the
# EPOCH resolution rather than the step number, which is what the
# tiebreak is actually expressed in.
set -e
cd ~/lv-mccarthy
PY=/home/infra-ops/ml/.venv/bin/python
OUT=~/r49-runs/mccarthy-4b-pairs-3ep
mkdir -p "$OUT"
"$PY" scripts/yarros-corpus/train_pairs_lora.py \
--pairs pairs/pairs-full.jsonl \
--val-pairs pairs/pairs-val.jsonl \
--base /home/infra-ops/carriers/Qwen3-4B-Instruct \
--out "$OUT" \
--seq-len 1536 --rank 32 --lr 1e-4 --epochs 3 --batch 1 --accum 8 --seed 4919 \
--eval-steps 25 --save-steps 25 --save-total-limit 60 \
> "$OUT/train.log" 2>&1
echo DONE > "$OUT/.complete"