diff --git a/archival-memory.md b/archival-memory.md index e6113c3..a0afb95 100644 --- a/archival-memory.md +++ b/archival-memory.md @@ -12092,3 +12092,1154 @@ Related: [[2026-09-15-fv-cross-site-snat]] - `[2026-09-15]` **Remote-site MASQUERADE rules on nh3-scale** for the asymmetric-return theory — they fired (counters incremented) but were not the fix. Reverted rather than left to accumulate as NAT achieving nothing. _Archived 2026-09-30._ + +## Recent decisions (archived) + +# `[2026-09-17]` Which voices earn a training seat next — measured against the catalogue, not chosen by taste + +Method: rank every author in the kvasir catalogue by **usable extracted** works, then apply the +selection criterion the lv-krakauer parking established — *does the author have a voice*, asked +before any corpus work, and specifically **does that voice live where the instrument looks**. +`voice_distance.py` is Burrows's Delta over CHARACTER BIGRAMS, so it sees function-word morphology, +punctuation and sentence rhythm. A writer whose distinction is plot, research or subject matter is +invisible to it — an adapter cannot carry that, and the gate cannot measure it. + +⚠ `triage.length` is in **CHARACTERS**, ~5.2 chars/word calibrated against builds we did ourselves +(The Crossing mobi 777,420 chars = our measured 149,985 words). Dedup by title taking the max across +formats, and floor at 100,000 chars — that is what excludes the `accepted`-but-truncated rows +(Blood Meridian epub at 6,031 chars beside the mobi's 623,849). + +## ⭐ The size ranking INVERTS the voice ranking at the top + +``` +Stephen King 76 works 12,133,529 w <- biggest, and NOT a candidate +Agatha Christie 72 5,451,377 <- second biggest, the Krakauer case exactly +Terry Pratchett 52 4,821,474 +Georgette Heyer 30 3,416,867 +Graham Greene 45 3,037,425 +William Faulkner 25 2,981,183 <- the pick +``` + +Christie is the whole lesson in one row: a superb writer whose genius is plot architecture, in prose +deliberately kept transparent. Nothing for a char-bigram Delta to grip. King is the softer version — +distinctive in pacing and brand-name texture, not in syntax. + +## The three that clear both bars + +**1. William Faulkner — 25 catalogue rows, ~15 pure novels, ~1.6M words.** +*The voice in one sentence:* sentences that defer their main clause through stacked subordination +and coined compounds until the reader is held inside a single unbroken perception. +About as char-bigram-legible as English gets — the voice IS the clause-joining morphology and the +`and`/`which`/`that` density. ⭐ **And he is McCarthy's stylistic ancestor, which is the real +argument:** the Brontë gate record states the frozen adjudication needs "a control-author panel (to +place an absolute band and a hard-negative sister)" and notes we have none. Faulkner beside McCarthy +makes each the other's hard negative — a METHOD upgrade, not just another roster entry. +⚠ Messiest corpus of the three: a 446k-word `Snopes: The Hamlet, The Town, The Mansion` omnibus +duplicates novels also present individually, and `Three Famous Short Novels` overlaps it again. That +is the Hemingway trap (169,759 words of measured 90-96% collection duplication) — containment pass +before anything else. + +**2. Toni Morrison — 13 rows, 11 novels after pruning, ~818k words.** +*The voice:* free-indirect discourse sliding between narrator and character mid-sentence, carried on +incantatory repetition and deliberate fragments. +Cleanest corpus shape on the list: **11 novels → 11 val units, beating Hemingway's 10.** Val units +scale with WORK COUNT, which is the structural reason Brontë's voice axis came back underpowered at +4 with no cheap fix. ⚠ Drop `Burn This Book` (anthology she edited) and `Playing in the Dark` +(criticism) — same reason Krakauer's reporting does not transfer. + +**3. Raymond Chandler — 9 rows, 7 novels + a 409k short-story omnibus, ~970k words.** +*The voice:* clipped first-person declaratives that periodically detonate into one baroque simile, +with dialogue carrying most of the scene. +Fills the register gap nobody else fills — **first-person hardboiled**; the line has no first-person +male narrator at all. Corpus is almost exactly Hemingway-sized (997k vs 958k), which was the +decisive gate. ⚠ Drop `Essays and Reviews` — non-fiction. + +## Held, and why + +**Conrad** (31 works, 2.4M) is a genuine tier-1.5 if a fourth is wanted. **Melville** (10, 1.9M) has +a superb voice but a mixed-register corpus — the cetology chapters are a different book from the +narrative. **Austen** (12, 1.18M) is worth noting because Burrows's Delta was developed on her, so +the instrument is known to resolve her. The romantasy cluster is a separate question entirely — +see [[2026-09-17-romantasy-register-measured]]. + +Related: [[2026-09-17-mccarthy-split-name-leak]], [[2026-09-17-lv-bronte-gate]], +[[2026-09-17-lv-hemingway-gate]]. + _Archived 2026-10-01._ + +# `[2026-09-17]` Romantasy measured as a register — it is real, we already took its best voice, and the obvious next pick is its worst + +Prompted by the operator pushing back on a one-clause dismissal of the lane as "depth behind +Yarros". The dismissal was taste; this is a measurement, on the gate's own instrument. + +**Method.** Char-bigram Burrows's Delta, the same measure `voice_distance.py` gates on. ~120k words +per author, sampled from the MIDDLE quartile of each author's largest works (front and back matter +are not the voice), equalised so a bigger sample is not a different measurement. 400 most-frequent +bigrams as the feature set, z-scored over 4,000-word chunks pooled across all authors. + +**Controls first, because a between-author number without a within-author floor is unfalsifiable.** + +``` +A-vs-A floor (two halves of the SAME author) + Yarros 0.285 Maas 0.298 Armentrout 0.314 St. Clair 0.322 Cole 0.338 + Kenyon 0.375 Reyne 0.391 + McCarthy 0.303 Morrison 0.327 Brontë 0.209 Hemingway 0.454 <- worst, used as the bar + +positive controls (known-distinct pairs — the instrument must separate these) + Yarros vs McCarthy 0.862 1.9x + Hemingway vs Brontë 0.773 1.7x + McCarthy vs Morrison 0.675 1.5x + Hemingway vs McCarthy 0.655 1.4x + +romantasy, all 21 pairs median 0.537 1.2x floor (range 0.465 - 0.674) +``` + +**The register is real but tight.** 1.2x floor against controls at 1.4-1.9x. Only one pair falls to +1.0x, so it is not seven names for one voice. + +⚠ **Sensitivity floor, stated because a result without one is unfalsifiable.** The 0.454 bar is +Hemingway's, inflated by his own heterogeneous corpus (1920s-1960s, novels + stories + posthumous). +Against the romantasy authors' OWN floors (~0.34) the same pairs read ~1.6x — control-grade. The +truth sits between those readings and **this method cannot split it finer**. One sample per pair, no +repeat draws: read the rank ordering as indicative, do not read small gaps at all. + +## Two findings that survive either floor reading + +⭐ **Yarros is the cluster OUTLIER, not a typical member.** Four of the five largest distances in the +matrix involve her — Yarros-Kenyon 0.674, Yarros-St. Clair 0.644, Yarros-Reyne 0.637, Yarros-Maas +0.567. **We already trained the most distinctive romantasy voice we hold**, so a second seat in the +lane buys measurably less than the first did. That is the actual answer to "what about romantasy". + +⭐ **Maas is the centroid, so the obvious commercial pick is the least distinctive.** Maas-Reyne +0.465 and Maas-Cole 0.470 are the two SMALLEST distances in the whole matrix. She is the biggest +name available (922k words) and measurably the most generic of the seven in char-bigram terms. +Picking by sales rank picks the worst adapter. + +## If the lane gets a second seat it is Kenyon + +Furthest from the shipped Yarros (0.674), so it adds the most new signal — **and 27 works means 27 +val units, the best-powered gate the line could build** (Hemingway 10, McCarthy 6, Brontë 4, where 4 +is the documented structural cause of an underpowered voice axis with no cheap fix). + +⚠ Two costs: the 27 are one series (Dark-Hunter), so the shared proper-noun space makes +`--scope corpus` mandatory rather than optional; and a "Dark Hunter - The Dark Hunter Complete" +omnibus sits in the catalogue rows, so the containment pass runs first. + +**Corpus shapes for the lane** (works ≥100k chars, deduped by title): +``` +Sherrilyn Kenyon 27 2,368,396 w Scarlett St. Clair 11 1,133,067 +Opal Reyne 14 2,466,307 Kresley Cole 10 1,037,914 +Jennifer Armentrout 6 1,179,953 Sarah J. Maas 5 922,711 +Rebecca Yarros 5 820,425 <- SHIPPED on this +``` +⭐ Worth noting for any future bar-setting: **Yarros shipped on 5 works / 820k words.** The corpus +bar is lower than it looks. + +Instrument: `scratchpad/regdist.py` (screening tool, not the gate). + +Related: [[2026-09-17-next-voice-seats]], [[2026-09-16-lv-voices-line]], [[2026-09-17-lv-bronte-gate]]. + _Archived 2026-10-01._ + +# `[2026-09-17]` headscale now split-DNSes `nh3.phasefinal.com` to the three AdGuards, so mesh clients can resolve the internal-only wildcard + +**headscale now split-DNSes `nh3.phasefinal.com` to the three AdGuards, so mesh clients can resolve the internal-only wildcard** (`talk`, `booth` — public DNS has no record for them; the fleet AdGuard answers 10.100.10.50). Operator-approved, scoped to nh3 rather than all of `phasefinal.com`. Config `/etc/headscale/config.yaml` in CT 106 on nh3-pve, backup `config.yaml.bak-2026-09-17-splitdns`, restarted, and the new route **read back from a node's netmap** rather than assumed. ⚠ Two things worth knowing: split DNS works fine here with `global: []` — headscale issue #1161's "split ignored without global" does NOT apply to v0.29.3, verified on the live mesh — and `override_local_dns: true` would REQUIRE global, which is the config that makes a roaming laptop lose ALL DNS when the mesh is down. That is why split, not global. Routing was never the problem: nh3-scale already serves 10.100.0.0/16. + _Archived 2026-10-01._ + +# `[2026-09-17]` ESH: Cityside Fiber failed twice in six hours; site ran on Verizon failover + +**Timeline (PDT).** + +``` +19:09:07 Cityside dies. UDM fails over to Verizon 5G (WAN2). Site stays up at ~33 ms. +19:51 Verified healthy on failover. +20:01:37 esh-scale drops off the headscale mesh; 10.0.0.0/16 withdrawn; whole site dark + from the colo. Beszel fires on all five ESH hosts. +20:11-15 Service restored. Operator had switched WAN1 to DHCP to get Cityside working at + all; it came back on CGNAT, not the static. Latency back to 9 ms. +01:06:23 Cityside fails AGAIN. Failover to Verizon. Site up, ~37 ms. +``` + +⭐ **The 20:01 blackout was most likely the operator's own WAN reconfiguration**, not ISP +instability — switching the WAN type bounces the interface, esh-scale loses its path, +headscale withdraws the route, and the site vanishes from the colo's view until it settles. +An earlier session theory ("Cityside came back half-provisioned and the UDM failed back into +an unstable session") is retired. + +⚠ **The diagnostic that mattered: physical link stayed UP at 2.5 GE with zero errors +throughout, while the ISP's next-hop `128.177.138.181` was unresponsive.** So "the ONT is +fine, it is upstream of the ONT" — the line to give Cityside. Traceroute from NH3 reached +`209.249.146.170` (one hop short) before dying, so the prefix was still routed. + +⚠ **CROWDSEC ROTATION FRAGILITY IS LIVE.** The `esh` allowlist on ana-docker carries the now- +dark static `128.177.138.182` (never-expiry), plus `97.190.18.88` (Verizon failover) and +`23.164.40.174` (Cityside CGNAT), both **7-day expiry**. ESH is on a rotating carrier-NAT +egress until the static is restored — the exact regime the 09-08 static purchase was meant to +end, and the class that once blackholed the whole site via a false ban. **If ESH loses colo +access, check `curl -s4 ifconfig.me` from esh-docker-vm FIRST** and allowlist the new address. + +**Still pinned to the dark static and broken until it returns:** FortiGate `infra-ops` +trusthost3 = `128.177.138.182`, so logins to ana-gw from ESH are refused. The dormant +`esh-ana` IPsec is bound to wan1/static (disabled, so no impact). + +⭐ **The mesh was NOT degraded on CGNAT** — tailscale hole-punched straight through +(`direct 23.164.40.174:41641`), which is why latency read 9 ms rather than a DERP figure. An +expectation of relay-on-CGNAT was wrong. + +## RESOLVED 2026-09-17 ~12:30 PT — the static is back, confirmed on four axes + +Not one check, because egress alone cannot tell a static WAN from a carrier NAT that happens +to answer (see auto-memory `feedback_egress_ip_cannot_detect_cgnat`): + +``` +config UDM WAN1 `wan_type = static`, ip 128.177.138.182, mask /30, gw 128.177.138.181 + — switched BACK from the DHCP the operator set at 20:11 during the outage +active stat/health: isp_name "Cityside Fiber", ASN 18731, num_disconnected 0. + WAN2 Verizon-5G is failover-only at priority 2 and idle. +egress esh-docker-vm sees 128.177.138.182 — EQUAL to the WAN ip, so not behind CGNAT +perf 2005/2142 Mbps symmetric; colo -> ESH 5.0 ms, 0% loss over 4 hosts-worth of pings + (Cityside CGNAT read 9 ms, Verizon failover 33-37 ms) +``` + +⭐ **The FortiGate pin un-broke itself and that was verified, not inferred.** `infra-ops` +trusthost3 is `128.177.138.182`; from esh-docker-vm, ana-gw `tcp/22` is OPEN and the +FortiGate offers a password prompt rather than dropping the connection — a trusthost +mismatch refuses outright, so reaching auth *is* the trusthost passing. The dormant +`esh-ana` IPsec bind to wan1/static is correct again (still disabled, still no impact). + +⚠ **The crowdsec temporary allowlist entries are being LEFT to expire on their own** +(2026-09-23): `97.190.18.88` Verizon and `23.164.40.174` Cityside CGNAT. Cityside failed +twice in six hours on 09-17, so until the line has earned some confidence those two are +cheap insurance against the exact false-ban blackout this rotation-fragility caused before. +`128.177.138.182` stays never-expiry. + +Related: [[2026-09-06-headscale-cutover]], [[2026-09-08-esh-static-wan-followups-and-ytvc]]. + _Archived 2026-10-01._ + +- `[2026-09-17]` **Operator ruled "leave it" on lv-hemingway's 3 separator-hidden names.** So `leak_gate.py` exits 1 on a SHIPPED tree by design; a future session seeing that red result should read this line, not start fixing. lv-bronte re-ran clean. + _Archived 2026-10-01._ + +# `[2026-09-17]` The leak gate passed with five protagonist names still in every copy + +Found during D4 pre-flight, three stages downstream of where it happened. Commit `c559664`. + +``` +leak gate, 2026-09-17 morning 0 of 75 renameable, 0 of 37 sub-threshold, both controls green +actually present, all 6 copies Bell x2 Chigurh x3 Moss x2 Toadvine x4 Glanton x2 +``` + +## The mechanism + +`leak_gate.py` scans `\b(Surface)\b`. **A character inserted inside a name defeats that +pattern outright**, so a mangled occurrence is not merely unrepaired — it is *unrenameable* +by `rename.py` and *unreportable* by the gate, and the gate prints a clean zero over it. +Two extraction artifacts produce exactly that: + +``` +B ell C higurh M oss T oadvine a small-caps drop cap kept as its own token +Toad-vine Glan-ton a print line-break hyphen kept by the extractor +``` + +⭐ **Every VISIBLE occurrence had been renamed correctly** — exact-match survivors were 0, +as the gate said. That is what makes this residue invisible to a spot-read: the names are +gone everywhere you look. `Bell` sits in the entity map at 147 capitals, `Glanton` at 365. + +This is the third member of a family. lv-bronte's was `_Antigua_` (`_` is a word character, +so `\bAntigua\b` cannot match inside it), found by hand in 2026-09-16 and never generalised. +**The generalisation is the point: any separator inside a name blinds a word-boundary scan.** + +## Fixed at three levels, and all three must stay + +1. **`build_corpus_mccarthy.py` rules 4 and 5** repair the source text — 32 split initials + with a *lowercase* remainder (rule 1 requires a following ALL-CAPS word and DROPCAP + requires two, so this is the class both leave behind), 5 hyphen-split names by name. + Both carry expected counts so a master change fails the build. + ⚠ **Rule 4's letter class is consonants only.** `I` opens **1,966** paragraphs (the + pronoun), `A` opens 143 (the article), `Y` opens 32 (Spanish *y*). Folding any of them + would corrupt 2,141 lines to fix 32 — the same `I`/`A` trap that bit `audit_stoplist.py`. +2. **`leak_gate.py` runs a separator-tolerant pass every time**, with its own positive and + negative controls, and **it fails the gate**. Validated against the pre-fix tree: reports + all five surfaces, exits 1. +3. The exact-match passes are untouched, so the old verdict is reproduced alongside the new. + +⚠ **The fragment filter is what makes the new pass usable.** A naive separator-tolerant +scan is dominated by false positives — on Hemingway it returns 21 hits of which **18 are +ordinary text** (`God damn` for the surface `Goddamn` ×14, plus `I run`, `On an`, `Do me`, +`Si le`). The discriminator, with no dictionary: in a genuine split at least one FRAGMENT +is not a word of this corpus. `God` and `damn` occur constantly; `Primi`, `tivo`, `ell`, +`higurh`, `Toad` do not. That one test cleared all 18 and kept all 3 real ones. + +⚠ **My first negative control could not pass.** It planted the split nonce in its own probe +text and then asserted the nonce was absent — an alarm wired to itself, failing on every +run. It now hunts the split nonce in the *real* copies. A control that cannot pass is not a +control. + +## The shipped corpora, checked with the committed instrument + +Re-derived with the COMMITTED gate, not a scratch probe: + +``` +lv-bronte GATE PASSED 0 separator-split survivors (9.7 s) +lv-hemingway GATE FAILED Pasionaria, Primitivo, Chicote (34.6 s) + 1 occurrence each per copy, in all 6 copies — SHIPPED and LIVE +``` + +⚠ The first version of this scan was **too slow to run** on Hemingway — per-surface scanning +is O(surfaces x copies x corpus) and 881 surfaces x 10 copies was still going at 5 minutes +when it was killed. Rebuilt as one alternation pass, same trick `scan()` already used: 35 s, +identical verdict and identical hit counts on both McCarthy trees. **A gate too slow to run +is not a gate.** + +**Operator call outstanding** on whether 3 names in a 958k-word corpus warrant re-gating and +retraining a live adapter. Not acted on. + +## The chain was recovered, not remembered — and is now written down + +There was **no McCarthy runbook**, and the D1→D3 session issued its commands over +non-interactive ssh so no shell history survived. The chain was recovered by rebuilding +candidates and matching sha256 against the artifacts on disk, then pinned: + +``` +D1 build_corpus_mccarthy.py 6 works byte-identical +D2 entities.py --min-count 5 --fold-clitics --drop-acronyms --min-mid-ratio 0.2 --min-mid 2 +D2c apply_gender_overrides.py entities-final.json byte-identical +D3 rename.py --preset mccarthy --scope corpus --min-cap 5 --copies 6 --seed 4919 + --holdout-chapter 7 17 all 36 copy files byte-identical +``` + +⚠ `--min-mid-ratio` is what keeps `Yeah`/`Buenas`/`Shh`/`Sí` out of the map. The map is +**insensitive** to it: any value in [0.05, 0.3] with `--min-mid` 1 or 2 reproduces byte-for- +byte; `--min-mid 3` does not. The original values are unrecoverable and it does not matter — +which is worth saying, because an exact-looking recipe that was never pinned invites a +false claim of reproduction. Full recipe and every deviation: `scripts/mccarthy-corpus/RUNBOOK.md`. + +⚠ **D1 must run on nh3-dev** — the builder reads the kvasir catalogue by absolute path and +gx10 has no copy. The previous session's "on gx10" note is true of D2 onward only. + +⚠ **No phrase map exists for this corpus**, so the gate's phrase audit does not run at all. +Yarros and Brontë both had one. Not closed. + +Rollback: `~/lv-mccarthy/corpus-{clean,renamed}.pre-splitfix` on gx10. + +Related: [[2026-09-17-mccarthy-d1-d3]], [[2026-09-17-beat-contamination-leak]], +[[2026-09-17-lv-hemingway-gate]], [[2026-09-17-lv-bronte-gate]]. + _Archived 2026-10-01._ + +# `[2026-09-17]` The SHIPPED lv-bronte adapter emits mid-sentence line breaks at 12.46 per 1k chars, and nothing downstream looks for it. + +⭐⭐ **The SHIPPED lv-bronte adapter emits mid-sentence line breaks at 12.46 per 1k chars, and nothing downstream looks for it.** Its corpus is 100% hard-wrapped at ~68 chars (Gutenberg plain text) and the wrap transfers: base control 0.00, ckpt475 (shipped) 12.46, ckpt925 11.79, every Hemingway arm 0.00 on a 0%-wrapped corpus. Both controls fire. `score_beats.py` passed Brontë's damage axis anyway. McCarthy is the MIXED case — The Road wrapped, the other five works not — which is worse to learn than either pure one, so `build_sft_pairs.py --reflow-hard-wraps` (DEFECT 4) fixes it at pair time, off by default. ⚠ The obvious fix, joining every interior newline, CORRUPTS 46 two-speaker exchanges whose blank line was lost — and unmarked dialogue is the one thing this adapter exists to learn. The rule splits on sentence-final punctuation and takes the cheaper error deliberately. + _Archived 2026-10-01._ + +# `[2026-09-17]` The `mccarthy` register names the punctuation ON PURPOSE, and that is a gate-design call made before any McCarthy number existed. + +⭐ **The `mccarthy` register names the punctuation ON PURPOSE, and that is a gate-design call made before any McCarthy number existed.** `eval-*.sh` drives the base control arm with the SAME system prompt via `--system-from`, and `voice_distance.py` is Burrows's Delta over character bigrams — so a tic left OUT of the register is a cheap win only the adapter can take, on a corpus measuring 0.0 quote marks per 10k against Hemingway's 838. Stating them hands them to the control too. Cost stated up front: the voice axis gets harder, and McCarthy's 276-passage val split (against Brontë's 44) is why that trade is affordable here and was not there. + _Archived 2026-10-01._ + +# `[2026-09-17]` lv-mccarthy's D1→D3 chain was RECOVERED, not remembered — there was no runbook and the commands went over non-interactive ssh, so no history survived. + +**lv-mccarthy's D1→D3 chain was RECOVERED, not remembered — there was no runbook and the commands went over non-interactive ssh, so no history survived.** Rebuilt candidates and matched sha256 against the artifacts on disk: 6 works, the entity map, the final map and all 36 copy files byte-identical. Now pinned in `scripts/mccarthy-corpus/RUNBOOK.md` with every deviation. ⚠ **D1 must run on nh3-dev** (the builder reads the kvasir catalogue by absolute path); the prior "on gx10" note is true of D2 onward only. ⚠ No phrase map exists for this corpus, so the gate's phrase audit never ran — Yarros and Brontë both had one. + _Archived 2026-10-01._ + +# `[2026-09-17]` Measured and DELIBERATELY not changed, three of them. + +**Measured and DELIBERATELY not changed, three of them.** The oversize-passage drop is 13.9% of McCarthy's train words, between Hemingway's 10.0% and the shipped Brontë's 18.3% — in range, no change. `BEAT_PROMPT` asserts the passage is first-person and McCarthy is third; measured inert (**0** narrator-retries against Hemingway's 615 of 7,094), so the prompt was left alone. Blood Meridian's 131 dash-separated chapter-argument paragraphs DID warrant a change and `--drop-leading-heading` now eats them (0 in every other work of all three corpora). + _Archived 2026-10-01._ + +# `[2026-09-17]` lv-mccarthy D1→D3 — built, gated, and every stage caught a defect in the stage before it + +**`~/lv-mccarthy/` on pfi-gx10.** `corpus-clean/` (167 units, 584,716 words), +`corpus-renamed/` (6 copies, 1,002 records), `scripts/`. Commits `705fa3a` `f3bf3ca` +`0fa68cb` `5aa10bf` `5ddb047`. + +``` +leak gate 0 of 75 renameable and 0 of 37 sub-threshold survive in any copy + positive control 108/108 surfaces found in the unrenamed source + negative control nonce absent from both trees +``` + +## The shared splitter: choose by SIZE, not by count + +`scripts/r49-corpus/split_units.py`. The inherited rule was "most units above a floor", which +is wrong for any book whose markers are PARTS: + +``` +Cities of the Plain 4 roman marks -> 4 units, median 22,312w +The Crossing 4 roman marks -> 4 units, median 37,310w +``` + +Four beats one, so it won, and the old guard only fired at exactly one unit. Now: a mode +qualifies only if its median unit is inside **[600, 12000]** AND no unit holds half the work; +among qualifying modes **priority** breaks the tie (contents > chapter-word > roman > +bare-numeral > caps-title). Works with no divisions fall back to **paragraph-block sections**. + +⭐⭐ **The first version of that rule was WORSE than what it replaced, and a control caught +it.** Scoring by "median closest to target" chose `caps-title` over the real chapters of +Hemingway's *True at First Light*: + +``` +bare-numeral 20 units median 5,337w max 11,155 <- the book's own chapters +caps-title 6 units median 777w max 113,886 <- median looked BETTER +``` + +Five stray all-caps lines gave five tiny units beside **one holding 97% of the book**. A median +cannot see that distribution; a max bound can. Controls green both ways afterwards: 8/10 +Hemingway works reproduce the shipped mode and count exactly, and 40,000 words with no blank +lines returns **1 unit** rather than fabricating sections. + +⚠ The Hemingway builder is deliberately NOT repointed at this module — its corpus is shipped +and its sha is pinned by a live adapter. + +## D1: the job was protecting a style that reads as damage + +``` +quote marks 0.0 per 10k (Hemingway 838) +apostrophes 123 per 10k (Hemingway 241) `dont` `aint` `wont` `didnt` +``` + +⚠⚠ **`repair_typography.py` MUST NOT be run on this corpus.** It normalises "toward what the +text does" and would put the quotation marks back. The builder runs no normalisation and then +**asserts** the quote density, so a future well-meaning change fails the build. + +⚠ **AND IT MAKES THE VOICE GATE EASY TO PASS FOR THE WRONG REASON.** `voice_distance.py` is +Burrows's Delta over CHARACTER BIGRAMS. An adapter that learns only "emit no quotation marks" +moves delta_cb a long way without having learned a sentence. **Pre-register a +punctuation-normalised secondary read before gating lv-mccarthy.** Tracked in the builder +docstring, commit `f3bf3ca`. + +Also: two truncated catalogue rows dropped for complete mobi siblings; all 15 cross-work +containment pairs measured (worst **0.10%**); back matter in 4 of 6 works carrying the author's +name 26 times → **0**; alphabet re-derived at 1,411 non-ASCII letters across 14 Spanish forms. + +⚠ The back-matter strip runs **BEFORE** the split for McCarthy, inverting the Hemingway order: +Blood Meridian and The Crossing end with a dumped table of contents made of bare roman numerals +on their own lines — the exact shape of a chapter marker. + +## D2 caught a D1 defect: three small-caps manglings + +The entity map returned `E`, `H`, `T`, `K` as renameable entities with 17–33 capitalised +occurrences each — the `G` class from Hemingway, where `G` was about to be renamed to a surname +248 times. Reading them showed the extractor mangled small-caps openings three ways: + +``` +1. SPLIT INITIAL `T HE HOUSE was built` -> `The house was built` 32 cases +2. UNMARKED RUN `THEY STOOD in the doorway` -> `They stood in the doorway` 88 cases +3. LOST INITIAL `HE CANDLEFLAME` -> `THE CANDLEFLAME` 1 case +``` + +Rule 1 requires a FOLLOWING all-caps word, so `A TV was playing` and `A Mexican was changing` +are untouched. Rule 2's `[a-z]` lookahead is what makes it safe — a genuine shout or sign is +not followed mid-sentence by lowercase. All 23 distinct first words of the 88 were checked. + +⚠⚠ **A fourth "fix" was nearly shipped that would have CORRUPTED the text.** `HEY RODE` → +`THEY RODE` looked right from a survey of the BUILT corpus. The raw master has `THEY RODE` +intact, twice — `HEY RODE` matched as a SUBSTRING, and the unanchored replace produced +`TTHEY RODE`, which rule 2 then lowercased to `Tthey rode`. Caught by the count assertion +(expected 1, replaced 2) and settled by reading the master. ⚠ My first corruption check also +missed it, searching for `TTHEY` when the pipeline had already lowercased it — **check the +shape the pipeline emits, not the shape you imagined.** + +## D2's own gates: and `audit_stoplist` was scanning its own rationale + +⚠⚠ **A defect in `audit_stoplist.py`, latent for every corpus before this one.** It built its +surface set from every list value in the stoplist JSON — including `_why`, which by convention +is a LIST OF PROSE LINES. Its empty separator line matched the honorific pattern **139 times**, +printing a flag with no surface name above the one real catch. Now skips `_`-prefixed keys. + +That real catch was a contradiction **inside my own file**: `Franklin` sat in the geography list +(the old name for El Paso) while the same file's note recorded *"I'm here to see Mr Franklin"*, +a lawyer in All the Pretty Horses. A second self-inflicted one: a speculative A–Z fragments list +stoplisted `I` and `A`, and `Sir I dont think I can do that` duly tripped the audit. It is now +the four letters actually measured as entities. + +Everything ambiguous was read in context: **Socorro is the ranch cook, not the New Mexico +town**; Niño, Keno and Redbo are HORSES (renameable, the `Inglés` precedent); Yaqui and Gilenos +are real peoples; Hashknives is a real cattle outfit; Hearst, Trias, Huerta and Madero are real +historical figures on the page under their own names. + +Final: 123 map surfaces, 124-surface stoplist, `entities.py` 27/27 controls, both audits PASS. + +## The human gender pass is an auditable file + +The honorific/window resolver scored **21 correct / 3 held / 1 WRONG** against a 26-name +control; the base-rate proximity resolver built for Hemingway scored 18/6/1 and **its own guard +correctly REFUSED to write**. So the incumbent stands and four entries are fixed by hand in +`gender_overrides_mccarthy.json`, each carrying its evidence. + +⚠ All four are female and all four look male-dominated in raw counts, because this corpus runs +**29,144 male pronouns to 5,036 female — a base rate of 85.3% male**. Carla Jean Moss at +31m/21f would be 44m/8f at that rate; 21 against an expected 8 is decisive. Same arithmetic that +recovered Pilar and Brett on Hemingway. Alfonsa was in the control and is correctly absent from +the map at 4 occurrences, below the min-count floor — an error in the control, not the pipeline. + +`apply_gender_overrides.py` refuses twice: a name absent from the map is an error rather than a +silent no-op, and overruling a gender the detector holds needs an explicit `"correcting": true` +so it cannot look like filling a held entity in a diff. + +## D3: three calls, and the holdout fix that matters most + +1. **`--scope corpus`**, not the per-work default. Nine surfaces appear in more than one work — + Parham (The Crossing + Cities of the Plain), Grady and Cole (All the Pretty Horses + Cities + of the Plain), Socorro, Héctor. A per-work map gives John Grady a different invented name in + each novel, turning one character into two. +2. **A new `mccarthy` preset.** Hemingway's romance pool carries `it_IT` and `fr_FR` for his + Italian and French casts; McCarthy writes neither language. `en_GB` goes for the same reason. + `en_US` + `es_MX`/`es_ES` at an even share. +3. **`--min-cap 5` to match the entity map's floor.** The first gate run FAILED with 45 + survivors: `entities.py` admits cap ≥ 5 while `rename.py` renamed only cap ≥ 8, so every + entity between sat in the map, was never renamed, and counted as a leak. Hemingway never hit + it because its map had `sub_threshold_total: 0`. + +⭐ **`--holdout-chapter` NOW TAKES A LIST.** The val split is one chapter index per work, so its +SIZE is set by how many WORKS a corpus has, not how many words: + +``` +Hemingway 10 works -> 9 val units -> 36,563 words/copy -> gate DECISIVE +Brontë 4 works -> 4 val units -> 17,043 words/copy -> gate MARGINAL +McCarthy 6 works -> 6 val units -> ~18,000 would have been Brontë's end +``` + +Holding out chapters **7 and 17** gives **11 units and 40,653 words per copy — larger than +Hemingway's** — for 7% of the corpus, on a corpus 40% smaller than his. No amount of corpus size +fixes a val split that scales with work count. + +Related: [[2026-09-17-lv-hemingway-gate]], [[2026-09-17-mccarthy-krakauer-d1]], +[[2026-09-17-lv-bronte-gate]]. + _Archived 2026-10-01._ + +- `[2026-09-17]` ⭐ **lv-mccarthy D1–D3 complete on gx10, leak gate PASSED (0 of 75 renameable, 0 of 37 sub-threshold, both controls green).** Three McCarthy-specific calls, each forced by a measurement: corpus-scoped rename (the Border Trilogy shares 9 surfaces across books), a new `mccarthy` name preset (Hemingway's carries it_IT/fr_FR and McCarthy writes neither), and `--min-cap 5` to match the entity map's floor — the first gate run failed with 45 survivors purely because rename's floor was 8 and the map's was 5. → `persistent-memory.d/2026-09-17-mccarthy-d1-d3.md` + _Archived 2026-10-01._ + +# `[2026-09-17]` Hemingway ships as-is: operator ruled "ship stands" on both measured corpus defects + +**Hemingway ships as-is: operator ruled "ship stands" on both measured corpus defects** — the 0.96% beat contamination and the 130 non-name entity-map surfaces. `audit_pairs_sourcenames.py --filter-out` and `audit_entity_map.py` exist and are the instruments if that is ever revisited; neither was run against the shipped adapter. + _Archived 2026-10-01._ + +# `[2026-09-17]` A unit splitter must choose by SIZE, not by count — the inherited rule silently produced 22,000-word "chapters". + +⭐⭐ **A unit splitter must choose by SIZE, not by count — the inherited rule silently produced 22,000-word "chapters".** `scripts/r49-corpus/split_units.py`: a marker mode qualifies only if its median unit is inside [600, 12000] AND no unit holds half the work; among qualifying modes PRIORITY breaks the tie (contents > chapter-word > roman > bare-numeral > caps-title), and paragraph-block sections are the fallback for works with no divisions. ⭐ **Both rules exist because a control caught them**: scoring by "median closest to target" chose `caps-title` (6 units, one holding **97%** of the book) over True at First Light's real 20 chapters, because a median cannot see that distribution and a max bound can. Positive control: 8/10 Hemingway works reproduce the shipped mode and count exactly. Negative control: 40,000 words with no blank lines → 1 unit, refuses to fabricate divisions. Commit `705fa3a`. + _Archived 2026-10-01._ + +# `[2026-09-17]` lv-mccarthy D1 built — 167 units, 584,756 words — and the whole job was protecting a style that reads as damage. + +⭐ **lv-mccarthy D1 built — 167 units, 584,756 words — and the whole job was protecting a style that reads as damage.** 0.0 quote marks per 10k (Hemingway 838), `dont`/`aint`/`wont`. The builder runs NO typography normalisation and asserts the quote density afterwards. Two truncated catalogue rows dropped for complete mobi siblings; all 15 containment pairs measured (worst 0.10%); back matter in 4 of 6 works carried the author's name 26 times → 0. ⚠ The back-matter strip runs BEFORE the split here — Blood Meridian and The Crossing end with a dumped TOC of bare roman numerals, the exact shape of a chapter marker. Commit `f3bf3ca`. + _Archived 2026-10-01._ + +# `[2026-09-17]` lv-krakauer D1 built — 126 units, 422,880 words — and its name guard caught three defects nothing else would have reported. + +⭐ **lv-krakauer D1 built — 126 units, 422,880 words — and its name guard caught three defects nothing else would have reported.** Back matter searched only the LAST unit while the apparatus sat in unit 37 of 41; relying on the splitter to drop front matter failed because the ebook TOC sits above the author's note and gave it a `Chapter Thirty-Two` to start on; and **zero was the wrong bar** — 2 survivors are Krakauer writing about his own father in Into the Wild's autobiographical chapters, so the allowance is pinned at 2 with every survivor printed. ⚠ Both strips are windowed in the OPPOSITE direction from McCarthy's, because Krakauer's `ALSO BY`/`Copyright`/`About the Author` sit at 0.0–0.6% of the file. Commit `4be0630`. + _Archived 2026-10-01._ + +# `[2026-09-17]` `triage_disposition = 'accepted'` in the Kvasir catalogue does NOT mean the extraction succeeded. + +⚠ **`triage_disposition = 'accepted'` in the Kvasir catalogue does NOT mean the extraction succeeded.** Blood Meridian's epub row holds 1,167 words of a 117,000-word book, The Crossing's 222 of 150,000 — real prose, real titles, accepted. Faulkner's *The Mansion* is 39 words. `near_dup_pairs` holds ONE row in the entire 1,284-work library and is blind to a fragment beside its full sibling. **Word-count every master before trusting a row**, and note that word count alone cannot tell a truncated novel from a legitimately short work. + _Archived 2026-10-01._ + +# `[2026-09-17]` lv-hemingway: SHIPPED on ckpt850 — the line's first clean voice pass, and one axis that needs reading + +**Status: SHIPPED 2026-09-17 03:33 as `lv-hemingway` on `vllm-voices` (fv-ml1 GPU0 :8027), +checkpoint-850.** Seat healthy 190 s after recreate, four models served +(`voices-base`, `lv-yarros`, `lv-bronte`, `lv-hemingway`), GPU0 96,092 → **96,090 MiB** — a +LoRA rides inside the existing seat and costs nothing. Adapter verified byte-identical to +the checkpoint by sha256 across two hops. + +Gate design **pre-registered before any generation existed**: +`scripts/hemingway-corpus/GATE-PREREG.md`, commit `0bb4938`. + +## The gate result — 3 arms × 60 held-out beats × 4 seeds = 240 generations per arm + +| axis | result | numbers | +|---|---|---| +| **A. VOICE** | ✅ **PASS, 6.4×** | +0.413 delta_cb vs base, pairwise floor 0.064. Also clears the OLD all-arms floor (0.113) — **this verdict does not depend on the rule change** | +| **B. NOT COPIED** | ⚠ **content clean, rate 7× the author's own** | 0.07 hit-rate, mean-longest 0.6, **max 9 words**. Base 0.00, **held-out Hemingway 0.01** | +| **C. NO DAMAGE** | ✅ PASS | ran-on +0.08, on-beat −0.14, both inside a 0.217 floor; in-band 0.79 vs base 0.05 | + +``` +same-author target (held-out Hemingway vs itself) delta_cb 0.364 <- best achievable +ckpt1750 0.439 +ckpt850 (SHIPPED) 0.511 +base-unadapted 0.924 +``` + +⭐ **THE STRONGEST VOICE RESULT IN THE LINE. The span is 0.924 → 0.364 = 0.560 and ckpt850 +closed 73.8% of it (ckpt1750 86.6%)**, against lv-bronte's 48%. Power came from the corpus, +not from a better method: 173 in-band val pairs allowed a **60-beat** fixture where Brontë +had 44 in-band and could only run 30. + +## ⚠⚠ AXIS B — THE COMFORTABLE EXPLANATION WAS WRONG, AND THE CONTROL IS THE ARTIFACT + +`memorization_check.py` uses the **base-unadapted arm** as its negative control, and on this +corpus that control is weak in one direction only — **it makes an innocent arm look guilty.** +Base writes 18,035 words of *summary*; the adapted arms write 27,413 of *pastiche*. Text that +does not imitate the register cannot collide with its n-grams, so base's 0.00 partly measures +"different register", not "did not memorise". + +The obvious hypothesis was that Hemingway's plain, high-frequency register makes 8-gram +collisions inevitable for any arm that learns it. **That hypothesis is refutable, was tested, +and is FALSE.** New control: **held-out Hemingway — the author himself, val text no arm +trained on — scored against the train split**, chunked to the generations' own median length +(101 words) so the comparison is like for like. + +``` +sample n hit-rate mean-longest max +HELD-OUT HEMINGWAY (never trained) 370 0.01 0.1 10 +base-unadapted 240 0.00 0.0 0 +ckpt1750 240 0.08 0.7 9 +ckpt850 (SHIPPED) 240 0.07 0.6 9 +positive control (train vs train) 160 <- not blind +``` + +⭐⭐ **The adapter reproduces train-corpus word sequences ~7× more often than the author +reproduces himself.** If the register explained it, real Hemingway would collide at the same +rate; it collides at 0.01. + +⭐ **And the exposure is still nil, which is a different question from the rate.** All 19 +matched runs were READ, not counted. Every one is stock dialogue — `i don t think so the girl +said`, `came over and sat down at the table`, `how do you feel i feel very well`. No plot, no +imagery, no distinctive phrase, **no proper noun** (the one name-shaped hit, `swift tristan`, +is the RENAMED invented name, not Hemingway's). The longest run is **9 words — shorter than +the 10-word run genuinely unseen Hemingway shares with the train split by coincidence.** + +What is being reproduced is the *grammar of his dialogue*, which is the thing the adapter +exists to learn, rendered in the commonest words in English. **Elevated rate, zero +protectable content.** Hemingway is in copyright; the in-line precedent is lv-yarros, also in +copyright, shipped at 0.10 against a 0.07 control. Unload is 0.003 s and one compose line. + +⚠ **The durable lesson is about the instrument, not this adapter: a negative control that +differs from the candidate in a way CORRELATED with the metric is not a control.** Always ask +what the metric returns for a known-innocent sample *in the same register*. + +## Why ckpt850 and NOT ckpt1750, the loss minimum + +ckpt1750 has the better point estimate on voice (0.439 vs 0.511) and **it is not usable**: + +``` +gap between candidates 0.072 +pairwise floor max(0.113, 0.050) 0.113 -> NOT resolvable +``` + +Indistinguishable, so the pre-registered tiebreak falls to the axes that resolve — and +**ckpt850 wins every one**: + +| | ckpt850 (shipped) | ckpt1750 | +|---|---|---| +| seed spread | **0.050** | 0.113 — **2.3× wider** | +| memorisation hit-rate / mean-longest | **0.07 / 0.6** | 0.08 / 0.7 | +| ran-on | **0.08** | 0.12 | +| epoch | **0.959** | 1.973 | + +ckpt1750's spread is one seed: 0.491, 0.449, 0.468, then **0.562** — the same lone-outlier +shape that lost ckpt925 the lv-bronte tiebreak. + +⭐ **THE TWO-EPOCH RECIPE DID NOT TRANSFER HERE EITHER — it is now 0 for 2.** Hemingway's +minimum really is step 1750, but step 850 is **+0.0040 against a 0.0044 median neighbour +jitter**, with three checkpoints inside one jitter of the best. Epoch 2 buys nothing that +resolves and costs 2.3× the variance. Only the epoch-3 collapse is robust: **+0.0762 = 17.4× +jitter**, which is why `adapter/` was never gated. **Stop carrying "two epochs on a +three-epoch schedule" forward; read the curve and prefer the earlier tied checkpoint.** + +## Pre-flight: the beat leak IS present in Hemingway, and the fixture is clean + +`audit_pairs_sourcenames.py` (new, commit `0bb4938`) closes the blind spot `leak_gate.py` has +by construction. Controls green every run: 941/941 surfaces found in the unrenamed source, +nonce absent from both trees, 6/6 planted names detected. + +``` +train beats 70 of 7,094 (0.96%) Santiago x16, Catherine x7, Rinaldi x3, Brett, Harry, + Jake, Pablo, Nick, Maria, Helen ... 36 distinct +train responses 0 of 7,294 -- the rename itself held perfectly +val beats 0 of 200 -- THE EVAL FIXTURE IS CLEAN; the gate is unconfounded +``` + +⭐ The beat-only signature is exactly lv-bronte's. **Yarros's and Hemingway's earlier clean +runs were never evidence of immunity** — they predate the detector. + +**Cross-validated on real data** where the answer was already recorded: the fixed Brontë +pairs return **0 of 3,858** (matching "0 leaks across 3,858 pairs"), and +`pairs-full.CONTAMINATED.jsonl` returns **15 of 792 = 1.89%** with Rochester ×6, Jane, +Brocklehurst ×2, Beck, Fairfax, Burns, Helen, Eyre — against a record of "13 of the first 714 +beats (1.8%)" with the same names. An independently written instrument reproducing a +documented finding at the right magnitude is what makes its zeroes mean *absent*, not *blind*. + +`--filter-out` produces a clean **7,024-pair** set in one command (70 dropped, 0.99%), +verified by re-audit at 0 of 7,024. **A retrain on it is the operator's call, not done.** + +## ⚠ A SECOND corpus defect, measured and NOT acted on + +`audit_entity_map.py` (new, commit `051b99e`) is the mirror of `audit_stoplist.py`: it finds +surfaces wrongly held **IN** the entity map, which `leak_gate.py` cannot see because it only +ever asks whether the author's names are GONE, never whether non-names were spared. + +``` + positive control `other` 764/1356 article-preceded = 0.56 + negative control 100 honorific-confirmed people, highest Inglés 0.26, bulk 0.00-0.06 + FLAGGED 130 of 946 surfaces · 1,616 instances · 0.162% of corpus words +``` + +`African`, `Chinese`, `Basques`, `Republican`, `Communist`, `X-ray`, `Coca-Cola`, `Ritz`, +`Prado`, `Cezanne` were all renamed into invented proper nouns. **Some flags are correct +renames** — `the Widow`, `the Informer` are genuine Hemingway epithet-names — so every hit is +reported for reading, never auto-removed. Plus **16 bare initials in the map**, `C` at 274 +occurrences: the same class as the `G` caught by hand about to be renamed 248 times. + +At 0.162% of words this did not block the ship. It is the thing to fix first if a corpus +rebuild ever happens. + +## Artefacts + +`gx10:~/lv-hemingway/` (corpus-clean, corpus-renamed, beats-hemingway-60.json + sidecar, +eval-hemingway.sh, voice-prep.py, eval.log), `gx10:~/r49-runs/hemingway-4b-pairs-3ep/` +(54 checkpoints kept), `gx10:~/r49-runs/hemingway-eval/` (three arms × 240 generations, +memorization.txt, voice_distance.txt, score.*.txt). +`fv-ml1:/tank/aimodels/voice-adapters/lv-hemingway-4b-v1/` (adapter + a README carrying the +axis-B caveat, so it cannot be read as clean by anyone who finds the adapter without this). +Commits `0bb4938` `051b99e` `5e66114` `2e9b118`. + +Related: [[2026-09-17-lv-bronte-gate]], [[2026-09-16-lv-hemingway-corpus]], +[[2026-09-16-lv-voices-line]], [[2026-09-16-voices-seat-lora]], +[[2026-09-17-beat-contamination-leak]]. + _Archived 2026-10-01._ + +# `[2026-09-17]` The v2 voice floor is now PAIRWISE, and it retroactively passes lv-bronte. + +⭐⭐ **The v2 voice floor is now PAIRWISE, and it retroactively passes lv-bronte.** lv-bronte's ckpt475 shipped as a voice-axis FAILURE at +0.193 against a 0.251 floor contributed entirely by ckpt925 — a third arm nobody was shipping, on one outlier seed. Scored against the arm it was actually compared to the floor is 0.091 and it clears at **2.1x**. The rule was changed **prospectively**, pre-registered for lv-hemingway before any Hemingway number existed, on an argument independent of the answer: the sampling variability of a difference A−B depends on A and B, never on a third arm C. The previous session found the defect and deliberately declined to exploit it; this follows from fixing it. lv-hemingway passes under **both** rules, so its verdict does not lean on the change. Caveats amended append-only in the compose, the NFS README and the gate record. Commits `0bb4938` `2e9b118`. + _Archived 2026-10-01._ + +# `[2026-09-17]` The beat-contamination leak IS present in Hemingway — 70 of 7,094 train beats (0.96%), 0 of 200 val. + +⭐ **The beat-contamination leak IS present in Hemingway — 70 of 7,094 train beats (0.96%), 0 of 200 val.** `scripts/r49-corpus/audit_pairs_sourcenames.py` closes the blind spot `leak_gate.py` has by construction (it reads the corpus and the renamed copies, never the generated beats). Cross-validated on real data: the fixed Brontë pairs return 0 of 3,858 and `pairs-full.CONTAMINATED` returns 15 of 792 = 1.89% with the recorded names. `--filter-out` yields a verified-clean 7,024-pair set in one command; the retrain is the operator's call. **The val split being clean is why the gate could run at all.** + _Archived 2026-10-01._ + +# `[2026-09-17]` `audit_entity_map.py` — the rename can DAMAGE the prose and no gate will ever say so. + +⭐ **`audit_entity_map.py` — the rename can DAMAGE the prose and no gate will ever say so.** Mirror of `audit_stoplist.py`: surfaces wrongly held IN the map rather than out of it. 130 of 946 Hemingway surfaces flagged (1,616 instances, 0.162% of words) — `African`, `Chinese`, `X-ray`, `Coca-Cola`, `Ritz`, `Prado` renamed into invented names — plus 16 bare initials incl. `C` at 274 occurrences. Signal is a preceding article; controls derived from the corpus, not hand-picked. Every hit reported for READING: `the Widow` and `the Informer` are genuine epithet-names that should be renamed. Commit `051b99e`. + _Archived 2026-10-01._ + +# `[2026-09-17]` The two-epoch recipe is now 0 for 2 and should stop being carried forward. + +**The two-epoch recipe is now 0 for 2 and should stop being carried forward.** Hemingway's eval minimum is step 1750, but step 850 is +0.0040 against a 0.0044 median neighbour jitter — three checkpoints inside one jitter — and 850 won every resolving axis (2.3x tighter seed spread, lower memorisation, less ran-on). Same outcome as Brontë. What IS robust on this schedule is the epoch-3 collapse: +0.0762 = **17.4x jitter**. + _Archived 2026-10-01._ + +# `[2026-09-17]` gitea was reaching the PUBLIC route from every repo on nh3-dev + +**gitea was reaching the PUBLIC route from every repo on nh3-dev** — brokkr-smithy, sleipnir, Galdrabok, kvasir — and brokkr-smithy is pushed several times a week, so the fail2ban trigger was live, not dormant. Measured before acting (no split-horizon rewrite, no ssh alias, `ssh -G` confirmed port 22 to 38.120.12.44). Fixed by overriding the NAME once in `~/.ssh/config` rather than rewriting N remotes, so fresh clones and unaudited repos are covered too. Verified with a real `git ls-remote`, not by inspection. Commit `dcc1abc`. Flagged by brokkr-smithy-dev; `vh/imogen` created for them the same session. + _Archived 2026-10-01._ + +# `[2026-09-17]` `servers/fv-ml1/ssh-target` was bare `10.251.50.54`, so `deploy-stack.sh` connected as `lkraven` and could not write the infra-ops-owned `/opt/docker/compose/` + +**`servers/fv-ml1/ssh-target` was bare `10.251.50.54`, so `deploy-stack.sh` connected as `lkraven` and could not write the infra-ops-owned `/opt/docker/compose/`** — and lkraven's sudo on fv-ml1 needs a password, so `DEPLOY_SUDO=1` failed too. Now `infra-ops@10.251.50.54`; `--validate-only` still clean, deploy works. ⚠ Other hosts' `ssh-target` files may carry the same gap — a read-only refresh works as either user, so the fault only surfaces on a deploy. + _Archived 2026-10-01._ + +# `[2026-09-17]` The beat-writing model restores the author's real names — a leak the corpus gate structurally cannot see + +⭐⭐ **THE RENAME AND ITS GATE ARE BOTH WORKING, AND THE TRAINING DATA IS STILL CONTAMINATED.** +The corpus gate reads the corpus and the renamed copies. **It never reads the generated +instruction beats.** Those are written by an LLM that just read the passage — and if it +recognises the book, it supplies the canonical names out of its own training. + +**Measured on the first 714 lv-bronte pairs, before the filter existed:** + +- **13 beats (1.8%)** named source characters — `Rochester` ×6, `Jane` ×3, `Brocklehurst` ×2, + `Beck`, `Fairfax`, `Helen`, `Burns`, `Eyre`, `Reed`, `Rivers`. +- **0 of 714 RESPONSES did.** The rename was perfect; the instruction side was not. +- One beat read *"Saoirse confirms Rochester's flaws, then agrees in English to marry him"* — + a renamed name and a canonical one in the same sentence, which is the mechanism in miniature. + +**Why it matters more than 1.8% sounds:** the beat is the INSTRUCTION half of the pair, so +training on it re-teaches exactly the inventions the rename pipeline exists to remove. + +⚠⚠ **EXPOSURE SCALES WITH HOW WELL THE GENERATOR KNOWS THE BOOK.** It is worst for +public-domain classics and mildest for recent work. That is precisely why the Yarros and +Hemingway runs came up clean and Brontë did not — **their clean runs are NOT evidence they are +immune.** Both should be re-verified, and regenerated with `--source-entities`, before their +pairs are trusted again. + +**The fix.** `vet()` in `scripts/yarros-corpus/build_sft_pairs.py` gained a `sourcename` +reject plus `--source-entities `, taking the UNRENAMED entity map. Fired at +~3% of attempts on the Brontë rebuild. Commit `533cc0c`. + +**The end-to-end guard that proves it.** The chain now verifies every built pair — beat, +response and context — against every source surface before spending GPU hours: +`[verify] 3858 pairs vs 368 source surfaces -> 0 leaks`. + +⚠ A guard stricter than the gate cries wolf. The first verify pass excluded nothing and +flagged `Monsieur’` ×14 and `Qu’est-ce` ×4 — French grammar, not leaks — because +`--fold-clitics` leaves apostrophe keys the gate deliberately skips. Mirror `leak_gate.py`'s +own predicate; a guard that fails on false positives gets disabled, which is worse than the +leak it guarded. + +Related: [[2026-09-17-lv-bronte-gate]], [[2026-09-16-lv-hemingway-corpus]], [[2026-09-16-lv-voices-line]]. + _Archived 2026-10-01._ + +# `[2026-09-17]` A stoplist entry is an assertion the leak gate can no longer check + +⭐ **A stoplist entry is an assertion the leak gate can no longer check** — stoplisting removes a surface from the entity map, so a wrongly stoplisted CHARACTER is an undetectable leak. Three were wrong on Brontë (Leaven, Pierrot, Samuel); `scripts/r49-corpus/audit_stoplist.py` finds them by honorific and now gates the pipeline. Commit `8bb7686`. + _Archived 2026-10-01._ + +- `[2026-09-17]` **ESH: Cityside Fiber failed TWICE (19:09 and ~01:06); operator switched WAN1 to DHCP to restore service and has a ticket for the static.** crowdsec `esh` allowlist carries both failover egresses with 7-day expiries — the rotation-fragility is live. → `persistent-memory.d/2026-09-17-esh-fiber-outages.md` + _Archived 2026-10-01._ + +# `[2026-09-17]` lv-bronte: corpus gated for real, adapter trained, SHIPPED with a FAILED voice axis on the record + +**Status: SHIPPED 2026-09-17 01:24 as `lv-bronte` on `vllm-voices` (fv-ml1 GPU0 :8027), ckpt475 — +and it did NOT pass its voice gate.** Shipped because it is additive (one more named LoRA beside +`voices-base` and `lv-yarros`, reached only by requesting it), reversible (one compose line; hot-unload +measures 0.003 s), and clean on the SAFETY axis — 8-gram overlap identical to the never-saw-it control, +on a public-domain corpus. VRAM cost was nil: GPU0 96092 -> 96090 MiB. The caveat is written into the +compose file and into `/tank/aimodels/voice-adapters/lv-bronte-4b-v1/README.md` so it cannot be read +as a clean pass by anyone who finds the adapter without finding this note. + +⚠ **Do NOT cite lv-bronte as evidence pair-SFT works for this author.** The voice axis is unresolved, +not passed. + +## The gate result, in full + +| axis | result | numbers | +|---|---|---| +| **A. VOICE** | ❌ **FAIL** (both candidates) | ckpt925 +0.210, ckpt475 +0.193 vs base — both **under** the 0.251 measured noise floor | +| **B. NOT COPIED** | ✅ PASS | ckpt475 **0.00 hit-rate, max 0 — identical to the never-saw-it control**; ckpt925 0.01, max 8. Positive control saturates at 160, so the detector is not blind | +| **C. NO DAMAGE** | ✅ PASS | ran-on +0.15 against a 0.400 floor | + +``` +same-author target (held-out Brontë vs itself) delta_cb 0.338 <- best achievable +ckpt925 0.531 +ckpt475 0.548 +base-unadapted 0.741 +``` + +⭐ **THE EFFECT LOOKS REAL AND SUBSTANTIAL; THE INSTRUMENT CANNOT CERTIFY IT.** The reachable +span is 0.741 → 0.338 = 0.403, and the adapters closed **48–52% of everything achievable**. +Both beat base on *every individual seed*. This is an UNDERPOWERED result, not a null one — +and a "no effect" without its floor is unfalsifiable, so: **this method cannot resolve a voice +improvement smaller than ~0.251 delta_cb at 30 beats × 4 seeds on this corpus.** + +⭐⭐ **THE CAUSE IS STRUCTURAL: Brontë's val split yields 81 pairs (44 in-band) against +Hemingway's 200**, because the corpus is 678k words against 994k. Maxing the fixture 30 → 44 +beats would shrink the floor by only ~√1.47 ≈ 1.2× (to ~0.21, against a 0.21 gap) — still +marginal. **More SEEDS would not help either: the floor is a RANGE statistic, and ranges widen +with more samples.** There is no cheap fix. + +## ⚠ A DEFECT IN THE v2 RULE ITSELF, worth fixing before the next author + +The floor is defined as the **largest within-arm seed spread across ALL arms**. Measured here: + +``` +base-unadapted 0.772 0.813 0.751 0.772 spread 0.062 +ckpt475 0.670 0.631 0.604 0.578 spread 0.092 +ckpt925 0.776 0.584 0.525 0.620 spread 0.251 <- sets the floor, on ONE seed +``` + +So **adding a third, noisier arm raised the bar that failed the clean one.** Run as the +two-arm gate (base + ckpt475) the floor would have been 0.092 and +0.193 would have cleared +at 2.1×. This was NOT exploited — picking the floor that passes your preferred answer is the +threshold-chosen-after-seeing-the-numbers failure the pre-registration exists to prevent — but +the rule should say whether the floor is computed over the compared pair or over every arm +present. As written, a candidate's verdict depends on which *other* arms you happened to run. + +**The outlier was diagnosed, not waved away.** Degeneracy probe (fraction of a generation made +of its most repeated 5-gram) is uniform across every seed and both arms, 0.0078–0.0102. Seed +1234 is not a collapsed generation; delta_cb genuinely has that variance for that arm. + +## Which checkpoint, if it ships: **ckpt475** + +The two are 0.017 apart on voice — far inside any floor, i.e. indistinguishable. On the axes +that DO resolve, ckpt475 wins both: memorisation identical to the control (ckpt925 has a +verbatim 8-gram hit), and **2.7× tighter seed-to-seed variance** (0.092 vs 0.251) with no +degeneracy to explain the difference — consistent with ckpt925 sitting nearer the overfit +boundary. Given a coin-flip on voice, take the one that provably did not memorise. + +⭐ **THE RECIPE DID NOT TRANSFER.** Yarros and Hemingway both found their minimum inside +epoch two. Brontë's minima are step 475 (ep 1.00, 2.6107) and step 925 (ep 1.96, 2.6129) — +**0.0022 apart against a 0.0046 median neighbour jitter, i.e. indistinguishable**. Epoch 2 +buys Brontë NOTHING over epoch 1. What IS robust is the epoch-3 collapse: +0.075, ~16× jitter. +Do not carry "two epochs on a three-epoch schedule" to a new author as settled. + +## Artefacts + +`gx10:~/lv-bronte/` (corpus-clean, corpus-renamed, entities-final.json, pairs/, beats-bronte-30.json), +`gx10:~/r49-runs/bronte-4b-pairs-3ep/` (57 checkpoints kept), `gx10:~/r49-runs/bronte-eval/` +(three arms × 120 generations, memorization.txt, voice_distance.txt, score.*.txt). +Commits `fc834a8` `533cc0c` `7964d07` `e9e8c40` `8bb7686`. + +⚠ Two output labels in `voice_distance.py` are hardcoded Yarros strings — it prints +"reference: held-out Yarros" and a boilerplate "Base < Instruct" corroboration line. The +NUMBERS are Brontë's; those two labels are not. Not yet fixed. + +Related: [[2026-09-16-lv-voices-line]], [[2026-09-16-lv-hemingway-corpus]], [[2026-09-16-voices-seat-lora]]. + +--- + +## ⚠⚠ AMENDED 2026-09-17 — THE VOICE AXIS PASSES UNDER THE CORRECTED FLOOR RULE + +Everything above is left verbatim; it is what was believed at ship time. This section is +the correction, not a rewrite. + +**The defect this file itself named was fixed, and fixing it flips ckpt475's verdict.** +The section "⚠ A DEFECT IN THE v2 RULE ITSELF" above says the rule should state whether +the floor is computed over the compared pair or over every arm present. It is now +**pairwise**, pre-registered in `scripts/hemingway-corpus/GATE-PREREG.md` before a single +lv-hemingway number existed. Re-scoring the SAME 360 generations — no re-run, no changed +delta_cb: + +``` + arm delta_cb per-seed spread + ckpt925 0.531 (0.776 0.584 0.525 0.620) 0.251 + ckpt475 0.548 (0.670 0.631 0.604 0.578) 0.091 + base-unadapted 0.741 (0.772 0.813 0.751 0.772) 0.062 + + all-arms floor (as run) 0.251 + ckpt475 +0.193 vs pairwise floor 0.091 -> MOVED toward Brontë, 2.1x <- the two rules DISAGREE + ckpt925 +0.210 vs pairwise floor 0.251 -> within the floor, NOT a finding +``` + +⭐ **The sequence matters and is the reason this is not threshold-shopping.** The previous +session found the defect, recorded it, and explicitly declined to exploit it. The rule was +then changed prospectively on a structural argument independent of the answer it produces — +the sampling variability of a difference A−B depends on A and B, never on a third arm C, so +a candidate's verdict must not depend on which other arms were generated. `voice_distance.py` +prints both floors and flags disagreement, so neither number can be quoted alone. + +**Consequences:** +- lv-bronte's voice axis is a **PASS at 2.1x**, not a fail. The caveat is amended in place + (append-only) in `stacks/voices-seat/compose.yaml` and + `/tank/aimodels/voice-adapters/lv-bronte-4b-v1/README.md`. +- The sensitivity floor for that measurement is **0.091**, not 0.251. +- "Do not cite lv-bronte as evidence pair-SFT works for this author" is **WITHDRAWN**. +- ckpt475 over ckpt925 is unchanged and for unchanged reasons (no verbatim 8-gram hit, + 2.7x tighter seed variance). +- The "no cheap fix for the underpowered result" analysis above is superseded for Brontë: + it was underpowered against an inflated floor, not against its own. + +**Also amended:** the two hardcoded Yarros labels flagged at the end of this file are fixed. +`voice_distance.py --author` is now REQUIRED — the committed Brontë output literally reads +"reference: held-out Yarros" over Brontë's numbers — and the stale "one seed-pair per arm / +corroborates Base < Instruct" footer now reports what the run actually carries. + _Archived 2026-10-01._ + +# `[2026-09-16]` The lv-* voice line: Option C proved, lv-yarros shipped, lv-hemingway training + +⭐⭐ **INSTRUCTION-PAIR SFT BEATS RAW-TEXT TRAINING FOR AUTHOR VOICE, AND THE INCUMBENT NEVER +CLEARED ITS OWN CONTROL.** Measured n=120 per arm, 30 in-genre beats from HELD-OUT val passages +× 4 seeds, all arms re-measured in one session on one box: + +| arm | delta_cb (lower = more Yarros) | vs base control | 8-gram overlap | +|---|---|---|---| +| pairs 2ep ckpt-1650 | **0.410** | +0.289 ✅ | 0.12 | +| pairs 3ep ckpt-1650 (**shipped**) | 0.438 | +0.262 ✅ | **0.09** | +| raw-text instruct (incumbent) | 0.558 | +0.141 ❌ **inside the 0.153 floor** | 0.14 | +| base-unadapted (control) | 0.700 | — | 0.07 | + +Same-author target 0.463 (held-out Yarros vs itself). ⚠ **The two pair arms are NOT +distinguishable on voice** — 0.028 against a 0.153 floor. The 3ep checkpoint was chosen on the +axes that ARE resolvable: better held-out fit (2.3126 vs 2.3264), less overshoot (0.06 vs 0.10), +and verbatim overlap nearest the never-saw-it control. + +⭐ **THE RECIPE IS TWO EPOCHS ON A THREE-EPOCH SCHEDULE, not three epochs.** Launch `--epochs 3`; +the minimum lands at step 1650 **inside epoch two** and epoch three overfits (2.3126 → 2.3882, +flat). The entire gain over a 2-epoch run came from the stretched cosine keeping the LR alive — +at step 1600 the 3ep run was at 2.9e-05 where the 2ep run had annealed to 2e-07. ⚠⚠ **A +resume-and-append-one-epoch is a NO-OP for exactly that reason** (lr 2.3e-09 at step 1670): it +must be a fresh run with the longer schedule. + +⭐ **THE SAFETY PROPERTY: the model writes the INSTRUCTION, never the RESPONSE.** Every response +is real renamed prose; only the beat is machine-written, so voice is inherited rather than +synthesised. Memorisation checked with both controls (positive control saturates at 160): the +shipped arm sits at 0.09 against a 0.07 never-saw-it baseline and BELOW the raw-text arm's 0.14. + +⚠ **THE v1 DECISION RULE WAS WELL-FORMED AND MEASURED THE WRONG THING**, and the amendment is +recorded in `scripts/yarros-corpus/score_beats.py` with v1 retained verbatim. It gated on +in-band / on-beat / ran-on — and **base-unadapted scores in-band 0.96**. Instruction-following is +something Qwen3-4B-Instruct ships with, so those axes detect only DAMAGE, never the benefit an +adapter exists to buy. v2 gates on voice (delta_cb vs control beyond the floor) + not-copied +(8-gram overlap near control) + no-damage (overshoot). ⚠ on-beat's −0.27 was outside the floor +and is dropped from the gate, **not explained away** — the keyword proxy punishes prose that +DRAMATISES "she mocks him" rather than echoing the word, but three read samples is an anecdote. + +⭐ **A 5-BEAT FIXTURE HAD A NOISE FLOOR OF 0.800 AND MANUFACTURED A +0.45 RESULT.** At n=20 the +pilot looked like a clear in-band win; at n=120 the same gap was +0.08, inside a 0.233 floor. +One sample moves a rate by 0.2 when there are five. The 30-beat in-genre fixture (built from +held-out val pairs, `~/beats-yarros-30.json`) is the instrument; the Brontë stray-dog/kitten +fixture was also the wrong GENRE — "He licked her clean" came back as explicit sex. + +⚠ **THE HARNESS TRUNCATES AT THE FIRST BLANK LINE and that surface reported the pair arm as +"19 words, off-beat 0.10"** when the untruncated output was 90–132 words with the beat rendered +in a later block. `score_beats.py --metric-source raw|paragraph` keeps both views and the verdict +names which it used. Same family as `feedback_filters_that_silently_narrow_the_window`. + +**Artefacts.** `scripts/yarros-corpus/{build_sft_pairs,train_pairs_lora,score_beats, +memorization_check}.py`; commits `9b3d3c8` `90ed506` `713e83d` `efb7345` `7505124`. Booth +(24h TTL) was `http://10.100.10.50:8090/b/babyyarros-beats/` — six beats × four arms, blind-labelled. + _Archived 2026-10-01._ + +# `[2026-09-16]` voices-seat: LoRA over merge, measured — and GPU 0 is now full + +**`vllm-voices` live on fv-ml1 GPU 0 :8027**, one Qwen3-4B-Instruct carrier serving +`voices-base` plus `lv-` LoRA adapters. `stacks/voices-seat/`, commit `d17bd3d`. + +⭐ **LORA COSTS 24.3% OF DECODE THROUGHPUT AND IT IS WORTH PAYING.** n=30 per arm, interleaved, +A-vs-A noise floor **0.1%**: base **143.0 tok/s** median vs adapter **108.2**. The measurement is +unusually clean because `--enable-lora` serves BOTH the base name and the adapter name from ONE +process — the arm is a per-request field, so no restart, no second seat, no cold-vs-warm confound. +Arms were **interleaved rather than blocked** because the card's co-tenants take traffic this seat +does not control, and a block design would alias their load onto one arm. + +**Why pay it:** 3 authors cost 8.4 GB as adapters against ~23 GB merged; 6 cost 9.2 vs ~46. On a +card with 1.8 GB free afterwards that is the whole argument. If a voice ever lands on a latency +path, merge THAT one and serve it separately. + +⭐ **ADAPTER HOT-SWAP IS REAL AND FAST — MEASURED, not read from docs.** +`POST /v1/load_lora_adapter` **200 in 0.24 s**, `POST /v1/unload_lora_adapter` **200 in 0.003 s**, +VRAM unchanged, container stayed healthy. Proven by performing the `babyyarros`→`lv-yarros` +rename through it with no restart. ⚠ **A runtime-loaded adapter is GONE on the next +`compose up -d`** unless it is also in `--lora-modules` (which costs a recreate + ~3 min reload). +Runtime load is for TRYING a voice; the compose list is what persists. Switching between loaded +voices is just the `model` field — **not** a LiteLLM alias; LiteLLM is a thinner layer on top, +one alias entry per voice, no new deployment. + +⚠⚠ **`--gpu-memory-utilization` IS A REQUEST AGAINST *TOTAL* VRAM THAT THE CARD MUST ALREADY BE +ABLE TO HONOUR — not a share of what is free.** First bring-up REFUSED: *"Free memory on device +cuda:0 (11.16/94.97 GiB) is less than desired GPU memory utilization (0.12, 11.4 GiB)"*. Refusing +was the right outcome — it protected `cyberprev`, `gen-small` and the Parakeet STT seat rather +than squeezing them. + +⭐ **PINNING `--kv-cache-memory` IN BYTES MAKES THE FRACTION PREDICTIVE.** Requested 0.11 +(10,700 MiB), got **10,740 MiB** resident — a 40 MiB miss on a box where the fraction has been +wrong by **8–10 GB in BOTH directions** (cyberprev 0.40→47.1 GB, gen-small 0.48→36.9 GB). Second +seat to prove it after `gen-small`. Do not remove the pin. + +⚠ **fv-ml1 GPU 0 is now 96.0 of 97.9 GB.** GPU 1 ~5.7 free, GPU 2 ~2.4, GPU 3 is a HELD RESERVE +for a future full-card seat (`flash-next` alone needs 93 of 96 GiB). **There is no room for +another seat on fv-ml1 without a placement decision.** + +⚠ **SUPPORT WAS CHECKED, NOT ASSUMED**, per the training playbook's own lesson that LoRA support +is per-ARCHITECTURE not per-family: `vllm/model_executor/models/qwen3.py:271` declares +`Qwen3ForCausalLM` with `SupportsLoRA` plus `packed_modules_mapping` and `embedding_modules`. +**Do not transplant this compose onto an MoE carrier without re-running that grep** — the +playbook records a LoRA refusal on a Qwen3 MoE arch. + +**Naming (operator, 2026-09-16):** `lv-` — lv for **lang-voice**, retiring `baby*`, which +read fine for one experiment and invites confusion across a family. The adapter NAME is the +request's `model` field, so it is the public API of a voice. Historical persistent-memory entries +still say BabyYarros/BabyHemingway and were deliberately left as dated records. + _Archived 2026-10-01._ + +# `[2026-09-16]` lv-hemingway corpus: half the work was EXCLUSION, and the gate found what a hand count would not + +**994,760 words · 318 units · 6 renamed copies · leak gate PASSED 0 of 941 entities and 0 of 117 +audited phrases, both controls green.** `~/hemingway-corpus{,-renamed}`, builder +`scripts/hemingway-corpus/build_corpus_hemingway.py`, commits `9598d0b` `03b4a3f`. + +⭐ **THE CATALOGUE HOLDS 2,105,679 WORDS AND ROUGHLY HALF MUST NOT BE TRAINED ON.** Operator +scoped it to fiction only. Three exclusion passes, each measured or voice-specific: + +1. **Non-fiction, 8 works ~911k words** — By-Line, Dateline: Toronto, Death in the Afternoon, + Green Hills of Africa, The Dangerous Summer, the three posthumous "Hemingway on X" anthologies. +2. ⭐ **Four story collections, 169,759 words — MEASURED, not assumed.** `Short Stories` is the + First Forty-Nine and CONTAINS the others. 8-gram containment of the smaller work: Winner Take + Nothing **96.0%**, Snows of Kilimanjaro **95.2%**, Men Without Women **92.9%**, In Our Time + **90.6%**. ⚠⚠ **The catalogue's own `near_dup_pairs` table is BLIND to this** — it holds + whole-document simhashes (ONE row in the entire 1,284-work library) and this is PARTIAL + containment. Whole-document dedup cannot see a collection inside a larger collection. +3. ⭐ **`The Torrents of Spring` — excluded for a reason no word count could justify.** It is a + deliberate PARODY of Sherwood Anderson: the target author's name on a different author's + style, i.e. mislabelled data for a voice adapter. + +⚠ **THE AUTHOR'S OWN NAME WAS IN THE TRAINING TEXT 95 TIMES ACROSS 7 WORKS** — publisher back +matter ("Ernest Hemingway was one of America's foremost journalists… died in 1961") riding inside +the last unit, because a splitter cuts on headings and nothing follows the final one. **Identical +to the Yarros defect; nothing about the source changed to cause it.** Stripping the publisher +block left 18, all in `true-at-first-light`, inside a **CAST OF CHARACTERS and SWAHILI GLOSSARY +written by Patrick Hemingway** — an editor describing the author's real household. Markers are +matched in file order, earliest wins. Now 0. + +⚠ **`G` WAS ABOUT TO BE RENAMED TO A SURNAME, 248 TIMES.** Not a name: the fragment left by +`B.G.`, `G.M.`, `G2`, `G3`. Caught by reading surfaces IN CONTEXT, which is the Yarros lesson +repeating. Also read in context: `Gran` (fragment of `Gran Sasso`/`Gran Italia`/`Gran Hotel`), +`Shamba` (Swahili common noun), and `Inglés` — **kept renameable deliberately**, the gypsies' +in-world nickname for Robert Jordan, exactly parallel to Yarros's `Violence`. + +⭐⭐ **THE GENDER RESOLVER HAD TO BE REBUILT AND ITS OWN GATE CAUGHT THE FIRST ATTEMPT.** The +inherited one returned **397 male / 20 female** across 1,102 records with Catherine Barkley, +Brett Ashley, Pilar, Maria and Mary all held neutral. A plain majority vote over nearby pronouns +scored 18 correct but **5 WRONG** against the incumbent's 1 — and **every error was +female-read-as-male** (Pilar m=426 f=243, Brett m=249 f=137). The refuse-unless-better guard +rejected it, correctly. ⭐ **Cause, measured: the corpus base rate is 34,315 male pronouns to +8,699 female, nearly 4:1.** Pilar's "male-dominated" 426:243 is strongly FEMALE against that +background. Scoring each name's local mix against the corpus base rate instead of 50:50 gives +**18 correct / 11 held / 0 WRONG**, distribution 275m / 120f. +`scripts/hemingway-corpus/gender_by_proximity.py`. ⚠ Yarros solved its version with the POV +chapter header; Hemingway's editions have none, so that fix does NOT transfer. + +⚠ **THE ALPHABET DOES NOT TRANSFER EITHER: 1,496 non-ASCII letters across 23 forms** against +Yarros's 2. Hemingway writes Spanish, French and Italian constantly, so the rename pool needs +accents (new `hemingway` preset in `rename.py`). The Yarros ASCII-only conclusion would have +stranded every Spanish and Italian name in the cast — which is why F02 says re-derive per corpus. + +**Three source defects the splitter surfaced.** `Islands in the Stream` came out as ONE +143k-word record (roman numerals, unhandled). `Short Stories` came out as 5 units then 27, +because the edition carries a **SECOND contents listing** and first-occurrence matching resolved +31 of 58 titles to an index entry — keeping every occurrence and letting the word floor decide is +self-correcting; now 57. ⚠⚠ **And the drop-cap defect is in the HEADINGS here** (`T HE O LD M AN +AND THE S EA`), which **INVERTS the Yarros pipeline order: repair must run BEFORE the split**, or +the splitter cannot see the headings it needs. + +⭐ **Hemingway needed THREE mapped phrases where Yarros needed 48** (`Gran Maestro`, +`Unknown Tongue`, `Sin House`) and a 45-entry allow list — the whole difference being that Yarros +invented a world and Hemingway named the real one. Several allow entries were non-obvious and +required reading: `Royal Game` is a real colonial-Kenyan legal category, `White Heather` a Scotch +brand, `Bwana Game` a job title, `Roman Soldier`/`Wine Seller` stage-direction labels from the +one-act play `Today is Friday`. + _Archived 2026-10-01._ + +# `[2026-09-16]` Grok token broker — built, then shelved by the transport ruling. Do NOT arm the probe. + +**`services/grok-token-broker/` — seeded, committed, DISARMED, no consumer.** Commits `ebc4dac` +`b907a0e` `cf9d167`. ⛔ **Do not arm `probe-rotation`.** This is a finished resting place, not a +half-built tool: the gate works and the thing it gated for went away. + +**Operator ruling, relayed by heid:** *"keep the jail stop the a/b"* +(`heid dispatch-log/2026-09.jsonl#groa-transport-20260916-operator-keeps-the-jail`, alongside +`#groa-transport-ab-20260916-operator-stop`). Gróa dispatches through the read jail; +`groa_http_dispatch.py` is a documented fallback with no scheduled use. **Nothing in the fleet +wants a renewable xAI session.** + +⭐ **THE CODE-PLAN ENDPOINT EXISTS and I was one message away from telling the operator it did +not.** `https://cli-chat-proxy.grok.com/v1` serving **grok-4.6** (500,000 context) and grok-4.5, +`agent_type: grok-build-plan`, `auth_method: session`, `api_key`/`env_key`/`api_base_url` all +null. ⚠⚠ **It is in `~/.grok/models_cache.json` — the Grok CLI's own config, on nh3-dev.** I had +swept heid's repo, the gateway `.env`, the LiteLLM config and Vaultwarden, all correctly, and +concluded "does not exist". ⭐ **heid's line, taken: absence from the places you searched is not +absence.** The check that separates the two states is a LIVE REQUEST, not a grep. + +⚠⚠ **NOT WRITING `~/.grok/auth.json` IS NECESSARY AND NOT SUFFICIENT.** The refresh grant at +`https://auth.x.ai/oauth2/token` may ROTATE the refresh token, and many OIDC providers invalidate +the old one SERVER-SIDE. A broker refreshing the same credential kills the CLI login even though +it never touches the file. heid's module header reasoned about the WRITE; they amended it to name +invalidation, credited. This correction went infra-ops→heid the same day heid's went the other +way — **neither of us reaches the right answer alone.** + +⚠⚠ **THE PROBE'S BLAST RADIUS IS BOTH GRÓA TRANSPORTS, which is not visible from the infra side.** +`heid/scripts/groa_dispatch.py` builds `argv = ["grok", "-p", prompt, "--cwd", jail, ...]` and +shells the CLI, which authenticates from the same `~/.grok/auth.json`. The bwrap in the process +table is grok's own Landlock sandbox, not something Heid wraps. **One session, two ways of +reaching it** — an invalidating probe takes Gróa down on EVERY path until an interactive re-login. +🔴 **I had recommended "run the probe now while the CLI is idle" and withdrew it in writing**; +"idle" was a convenient assumption I never checked, on a day that had already taken eight panels. + +**Why the jail won, and it was not performance.** HTTP is faster (~523 s median vs ~890 s), +simpler, and arguably SAFER on confinement (no tools, so the 2026-06-10 escape class is +structurally impossible). It lost on FAILURE MODE: HTTP fails by returning a fast, confident, +well-formatted review that found nothing — indistinguishable from a clean bill. The jail fails by +timing out, which you can see. ⚠ **Do NOT quote a per-transport finding rate from this**: heid +states the 0/0/0-vs-5/7/3 numbers are confounded with bundle size (the zeros were all huge inline +bundles; the one HTTP round at jail-comparable size produced Gróa's leading solo), n=3–4 per cell, +no noise floor. Asymmetric-risk argument, **not** a resolved measurement. + +⚠ **Still unmeasured, and it is a billing question:** the jail reaches the coding plan already +paid for; the HTTP path reaches the METERED API and its responses carry `cost_in_usd_ticks`. +Whether that bills on top of the plan was never part of the ruling. One look at the xAI billing +console — **this fleet holds no xAI credential**, so it needs the operator's account access. + +⚠ The coding plan speaks the **Responses API** (`api_backend: "responses"`), not +`/chat/completions` — a second, independent obstacle to any LiteLLM alias. Moot while the jail is +ruled. heid also found and killed two live instructions in their own persistent-memory telling a +fresh session to dispatch `--groa-transport http`; either would have resumed a stopped experiment +after a context reset. + _Archived 2026-10-01._ diff --git a/persistent-memory.d/2026-09-16-grok-broker-shelved.md b/persistent-memory.d/2026-09-16-grok-broker-shelved.md deleted file mode 100644 index 7768c10..0000000 --- a/persistent-memory.d/2026-09-16-grok-broker-shelved.md +++ /dev/null @@ -1,54 +0,0 @@ -# `[2026-09-16]` Grok token broker — built, then shelved by the transport ruling. Do NOT arm the probe. - -**`services/grok-token-broker/` — seeded, committed, DISARMED, no consumer.** Commits `ebc4dac` -`b907a0e` `cf9d167`. ⛔ **Do not arm `probe-rotation`.** This is a finished resting place, not a -half-built tool: the gate works and the thing it gated for went away. - -**Operator ruling, relayed by heid:** *"keep the jail stop the a/b"* -(`heid dispatch-log/2026-09.jsonl#groa-transport-20260916-operator-keeps-the-jail`, alongside -`#groa-transport-ab-20260916-operator-stop`). Gróa dispatches through the read jail; -`groa_http_dispatch.py` is a documented fallback with no scheduled use. **Nothing in the fleet -wants a renewable xAI session.** - -⭐ **THE CODE-PLAN ENDPOINT EXISTS and I was one message away from telling the operator it did -not.** `https://cli-chat-proxy.grok.com/v1` serving **grok-4.6** (500,000 context) and grok-4.5, -`agent_type: grok-build-plan`, `auth_method: session`, `api_key`/`env_key`/`api_base_url` all -null. ⚠⚠ **It is in `~/.grok/models_cache.json` — the Grok CLI's own config, on nh3-dev.** I had -swept heid's repo, the gateway `.env`, the LiteLLM config and Vaultwarden, all correctly, and -concluded "does not exist". ⭐ **heid's line, taken: absence from the places you searched is not -absence.** The check that separates the two states is a LIVE REQUEST, not a grep. - -⚠⚠ **NOT WRITING `~/.grok/auth.json` IS NECESSARY AND NOT SUFFICIENT.** The refresh grant at -`https://auth.x.ai/oauth2/token` may ROTATE the refresh token, and many OIDC providers invalidate -the old one SERVER-SIDE. A broker refreshing the same credential kills the CLI login even though -it never touches the file. heid's module header reasoned about the WRITE; they amended it to name -invalidation, credited. This correction went infra-ops→heid the same day heid's went the other -way — **neither of us reaches the right answer alone.** - -⚠⚠ **THE PROBE'S BLAST RADIUS IS BOTH GRÓA TRANSPORTS, which is not visible from the infra side.** -`heid/scripts/groa_dispatch.py` builds `argv = ["grok", "-p", prompt, "--cwd", jail, ...]` and -shells the CLI, which authenticates from the same `~/.grok/auth.json`. The bwrap in the process -table is grok's own Landlock sandbox, not something Heid wraps. **One session, two ways of -reaching it** — an invalidating probe takes Gróa down on EVERY path until an interactive re-login. -🔴 **I had recommended "run the probe now while the CLI is idle" and withdrew it in writing**; -"idle" was a convenient assumption I never checked, on a day that had already taken eight panels. - -**Why the jail won, and it was not performance.** HTTP is faster (~523 s median vs ~890 s), -simpler, and arguably SAFER on confinement (no tools, so the 2026-06-10 escape class is -structurally impossible). It lost on FAILURE MODE: HTTP fails by returning a fast, confident, -well-formatted review that found nothing — indistinguishable from a clean bill. The jail fails by -timing out, which you can see. ⚠ **Do NOT quote a per-transport finding rate from this**: heid -states the 0/0/0-vs-5/7/3 numbers are confounded with bundle size (the zeros were all huge inline -bundles; the one HTTP round at jail-comparable size produced Gróa's leading solo), n=3–4 per cell, -no noise floor. Asymmetric-risk argument, **not** a resolved measurement. - -⚠ **Still unmeasured, and it is a billing question:** the jail reaches the coding plan already -paid for; the HTTP path reaches the METERED API and its responses carry `cost_in_usd_ticks`. -Whether that bills on top of the plan was never part of the ruling. One look at the xAI billing -console — **this fleet holds no xAI credential**, so it needs the operator's account access. - -⚠ The coding plan speaks the **Responses API** (`api_backend: "responses"`), not -`/chat/completions` — a second, independent obstacle to any LiteLLM alias. Moot while the jail is -ruled. heid also found and killed two live instructions in their own persistent-memory telling a -fresh session to dispatch `--groa-transport http`; either would have resumed a stopped experiment -after a context reset. diff --git a/persistent-memory.d/2026-09-16-lv-hemingway-corpus.md b/persistent-memory.d/2026-09-16-lv-hemingway-corpus.md deleted file mode 100644 index 232a7f5..0000000 --- a/persistent-memory.d/2026-09-16-lv-hemingway-corpus.md +++ /dev/null @@ -1,66 +0,0 @@ -# `[2026-09-16]` lv-hemingway corpus: half the work was EXCLUSION, and the gate found what a hand count would not - -**994,760 words · 318 units · 6 renamed copies · leak gate PASSED 0 of 941 entities and 0 of 117 -audited phrases, both controls green.** `~/hemingway-corpus{,-renamed}`, builder -`scripts/hemingway-corpus/build_corpus_hemingway.py`, commits `9598d0b` `03b4a3f`. - -⭐ **THE CATALOGUE HOLDS 2,105,679 WORDS AND ROUGHLY HALF MUST NOT BE TRAINED ON.** Operator -scoped it to fiction only. Three exclusion passes, each measured or voice-specific: - -1. **Non-fiction, 8 works ~911k words** — By-Line, Dateline: Toronto, Death in the Afternoon, - Green Hills of Africa, The Dangerous Summer, the three posthumous "Hemingway on X" anthologies. -2. ⭐ **Four story collections, 169,759 words — MEASURED, not assumed.** `Short Stories` is the - First Forty-Nine and CONTAINS the others. 8-gram containment of the smaller work: Winner Take - Nothing **96.0%**, Snows of Kilimanjaro **95.2%**, Men Without Women **92.9%**, In Our Time - **90.6%**. ⚠⚠ **The catalogue's own `near_dup_pairs` table is BLIND to this** — it holds - whole-document simhashes (ONE row in the entire 1,284-work library) and this is PARTIAL - containment. Whole-document dedup cannot see a collection inside a larger collection. -3. ⭐ **`The Torrents of Spring` — excluded for a reason no word count could justify.** It is a - deliberate PARODY of Sherwood Anderson: the target author's name on a different author's - style, i.e. mislabelled data for a voice adapter. - -⚠ **THE AUTHOR'S OWN NAME WAS IN THE TRAINING TEXT 95 TIMES ACROSS 7 WORKS** — publisher back -matter ("Ernest Hemingway was one of America's foremost journalists… died in 1961") riding inside -the last unit, because a splitter cuts on headings and nothing follows the final one. **Identical -to the Yarros defect; nothing about the source changed to cause it.** Stripping the publisher -block left 18, all in `true-at-first-light`, inside a **CAST OF CHARACTERS and SWAHILI GLOSSARY -written by Patrick Hemingway** — an editor describing the author's real household. Markers are -matched in file order, earliest wins. Now 0. - -⚠ **`G` WAS ABOUT TO BE RENAMED TO A SURNAME, 248 TIMES.** Not a name: the fragment left by -`B.G.`, `G.M.`, `G2`, `G3`. Caught by reading surfaces IN CONTEXT, which is the Yarros lesson -repeating. Also read in context: `Gran` (fragment of `Gran Sasso`/`Gran Italia`/`Gran Hotel`), -`Shamba` (Swahili common noun), and `Inglés` — **kept renameable deliberately**, the gypsies' -in-world nickname for Robert Jordan, exactly parallel to Yarros's `Violence`. - -⭐⭐ **THE GENDER RESOLVER HAD TO BE REBUILT AND ITS OWN GATE CAUGHT THE FIRST ATTEMPT.** The -inherited one returned **397 male / 20 female** across 1,102 records with Catherine Barkley, -Brett Ashley, Pilar, Maria and Mary all held neutral. A plain majority vote over nearby pronouns -scored 18 correct but **5 WRONG** against the incumbent's 1 — and **every error was -female-read-as-male** (Pilar m=426 f=243, Brett m=249 f=137). The refuse-unless-better guard -rejected it, correctly. ⭐ **Cause, measured: the corpus base rate is 34,315 male pronouns to -8,699 female, nearly 4:1.** Pilar's "male-dominated" 426:243 is strongly FEMALE against that -background. Scoring each name's local mix against the corpus base rate instead of 50:50 gives -**18 correct / 11 held / 0 WRONG**, distribution 275m / 120f. -`scripts/hemingway-corpus/gender_by_proximity.py`. ⚠ Yarros solved its version with the POV -chapter header; Hemingway's editions have none, so that fix does NOT transfer. - -⚠ **THE ALPHABET DOES NOT TRANSFER EITHER: 1,496 non-ASCII letters across 23 forms** against -Yarros's 2. Hemingway writes Spanish, French and Italian constantly, so the rename pool needs -accents (new `hemingway` preset in `rename.py`). The Yarros ASCII-only conclusion would have -stranded every Spanish and Italian name in the cast — which is why F02 says re-derive per corpus. - -**Three source defects the splitter surfaced.** `Islands in the Stream` came out as ONE -143k-word record (roman numerals, unhandled). `Short Stories` came out as 5 units then 27, -because the edition carries a **SECOND contents listing** and first-occurrence matching resolved -31 of 58 titles to an index entry — keeping every occurrence and letting the word floor decide is -self-correcting; now 57. ⚠⚠ **And the drop-cap defect is in the HEADINGS here** (`T HE O LD M AN -AND THE S EA`), which **INVERTS the Yarros pipeline order: repair must run BEFORE the split**, or -the splitter cannot see the headings it needs. - -⭐ **Hemingway needed THREE mapped phrases where Yarros needed 48** (`Gran Maestro`, -`Unknown Tongue`, `Sin House`) and a 45-entry allow list — the whole difference being that Yarros -invented a world and Hemingway named the real one. Several allow entries were non-obvious and -required reading: `Royal Game` is a real colonial-Kenyan legal category, `White Heather` a Scotch -brand, `Bwana Game` a job title, `Roman Soldier`/`Wine Seller` stage-direction labels from the -one-act play `Today is Friday`. diff --git a/persistent-memory.d/2026-09-16-lv-voices-line.md b/persistent-memory.d/2026-09-16-lv-voices-line.md deleted file mode 100644 index 1662d04..0000000 --- a/persistent-memory.d/2026-09-16-lv-voices-line.md +++ /dev/null @@ -1,53 +0,0 @@ -# `[2026-09-16]` The lv-* voice line: Option C proved, lv-yarros shipped, lv-hemingway training - -⭐⭐ **INSTRUCTION-PAIR SFT BEATS RAW-TEXT TRAINING FOR AUTHOR VOICE, AND THE INCUMBENT NEVER -CLEARED ITS OWN CONTROL.** Measured n=120 per arm, 30 in-genre beats from HELD-OUT val passages -× 4 seeds, all arms re-measured in one session on one box: - -| arm | delta_cb (lower = more Yarros) | vs base control | 8-gram overlap | -|---|---|---|---| -| pairs 2ep ckpt-1650 | **0.410** | +0.289 ✅ | 0.12 | -| pairs 3ep ckpt-1650 (**shipped**) | 0.438 | +0.262 ✅ | **0.09** | -| raw-text instruct (incumbent) | 0.558 | +0.141 ❌ **inside the 0.153 floor** | 0.14 | -| base-unadapted (control) | 0.700 | — | 0.07 | - -Same-author target 0.463 (held-out Yarros vs itself). ⚠ **The two pair arms are NOT -distinguishable on voice** — 0.028 against a 0.153 floor. The 3ep checkpoint was chosen on the -axes that ARE resolvable: better held-out fit (2.3126 vs 2.3264), less overshoot (0.06 vs 0.10), -and verbatim overlap nearest the never-saw-it control. - -⭐ **THE RECIPE IS TWO EPOCHS ON A THREE-EPOCH SCHEDULE, not three epochs.** Launch `--epochs 3`; -the minimum lands at step 1650 **inside epoch two** and epoch three overfits (2.3126 → 2.3882, -flat). The entire gain over a 2-epoch run came from the stretched cosine keeping the LR alive — -at step 1600 the 3ep run was at 2.9e-05 where the 2ep run had annealed to 2e-07. ⚠⚠ **A -resume-and-append-one-epoch is a NO-OP for exactly that reason** (lr 2.3e-09 at step 1670): it -must be a fresh run with the longer schedule. - -⭐ **THE SAFETY PROPERTY: the model writes the INSTRUCTION, never the RESPONSE.** Every response -is real renamed prose; only the beat is machine-written, so voice is inherited rather than -synthesised. Memorisation checked with both controls (positive control saturates at 160): the -shipped arm sits at 0.09 against a 0.07 never-saw-it baseline and BELOW the raw-text arm's 0.14. - -⚠ **THE v1 DECISION RULE WAS WELL-FORMED AND MEASURED THE WRONG THING**, and the amendment is -recorded in `scripts/yarros-corpus/score_beats.py` with v1 retained verbatim. It gated on -in-band / on-beat / ran-on — and **base-unadapted scores in-band 0.96**. Instruction-following is -something Qwen3-4B-Instruct ships with, so those axes detect only DAMAGE, never the benefit an -adapter exists to buy. v2 gates on voice (delta_cb vs control beyond the floor) + not-copied -(8-gram overlap near control) + no-damage (overshoot). ⚠ on-beat's −0.27 was outside the floor -and is dropped from the gate, **not explained away** — the keyword proxy punishes prose that -DRAMATISES "she mocks him" rather than echoing the word, but three read samples is an anecdote. - -⭐ **A 5-BEAT FIXTURE HAD A NOISE FLOOR OF 0.800 AND MANUFACTURED A +0.45 RESULT.** At n=20 the -pilot looked like a clear in-band win; at n=120 the same gap was +0.08, inside a 0.233 floor. -One sample moves a rate by 0.2 when there are five. The 30-beat in-genre fixture (built from -held-out val pairs, `~/beats-yarros-30.json`) is the instrument; the Brontë stray-dog/kitten -fixture was also the wrong GENRE — "He licked her clean" came back as explicit sex. - -⚠ **THE HARNESS TRUNCATES AT THE FIRST BLANK LINE and that surface reported the pair arm as -"19 words, off-beat 0.10"** when the untruncated output was 90–132 words with the beat rendered -in a later block. `score_beats.py --metric-source raw|paragraph` keeps both views and the verdict -names which it used. Same family as `feedback_filters_that_silently_narrow_the_window`. - -**Artefacts.** `scripts/yarros-corpus/{build_sft_pairs,train_pairs_lora,score_beats, -memorization_check}.py`; commits `9b3d3c8` `90ed506` `713e83d` `efb7345` `7505124`. Booth -(24h TTL) was `http://10.100.10.50:8090/b/babyyarros-beats/` — six beats × four arms, blind-labelled. diff --git a/persistent-memory.d/2026-09-16-voices-seat-lora.md b/persistent-memory.d/2026-09-16-voices-seat-lora.md deleted file mode 100644 index f6a9d7a..0000000 --- a/persistent-memory.d/2026-09-16-voices-seat-lora.md +++ /dev/null @@ -1,50 +0,0 @@ -# `[2026-09-16]` voices-seat: LoRA over merge, measured — and GPU 0 is now full - -**`vllm-voices` live on fv-ml1 GPU 0 :8027**, one Qwen3-4B-Instruct carrier serving -`voices-base` plus `lv-` LoRA adapters. `stacks/voices-seat/`, commit `d17bd3d`. - -⭐ **LORA COSTS 24.3% OF DECODE THROUGHPUT AND IT IS WORTH PAYING.** n=30 per arm, interleaved, -A-vs-A noise floor **0.1%**: base **143.0 tok/s** median vs adapter **108.2**. The measurement is -unusually clean because `--enable-lora` serves BOTH the base name and the adapter name from ONE -process — the arm is a per-request field, so no restart, no second seat, no cold-vs-warm confound. -Arms were **interleaved rather than blocked** because the card's co-tenants take traffic this seat -does not control, and a block design would alias their load onto one arm. - -**Why pay it:** 3 authors cost 8.4 GB as adapters against ~23 GB merged; 6 cost 9.2 vs ~46. On a -card with 1.8 GB free afterwards that is the whole argument. If a voice ever lands on a latency -path, merge THAT one and serve it separately. - -⭐ **ADAPTER HOT-SWAP IS REAL AND FAST — MEASURED, not read from docs.** -`POST /v1/load_lora_adapter` **200 in 0.24 s**, `POST /v1/unload_lora_adapter` **200 in 0.003 s**, -VRAM unchanged, container stayed healthy. Proven by performing the `babyyarros`→`lv-yarros` -rename through it with no restart. ⚠ **A runtime-loaded adapter is GONE on the next -`compose up -d`** unless it is also in `--lora-modules` (which costs a recreate + ~3 min reload). -Runtime load is for TRYING a voice; the compose list is what persists. Switching between loaded -voices is just the `model` field — **not** a LiteLLM alias; LiteLLM is a thinner layer on top, -one alias entry per voice, no new deployment. - -⚠⚠ **`--gpu-memory-utilization` IS A REQUEST AGAINST *TOTAL* VRAM THAT THE CARD MUST ALREADY BE -ABLE TO HONOUR — not a share of what is free.** First bring-up REFUSED: *"Free memory on device -cuda:0 (11.16/94.97 GiB) is less than desired GPU memory utilization (0.12, 11.4 GiB)"*. Refusing -was the right outcome — it protected `cyberprev`, `gen-small` and the Parakeet STT seat rather -than squeezing them. - -⭐ **PINNING `--kv-cache-memory` IN BYTES MAKES THE FRACTION PREDICTIVE.** Requested 0.11 -(10,700 MiB), got **10,740 MiB** resident — a 40 MiB miss on a box where the fraction has been -wrong by **8–10 GB in BOTH directions** (cyberprev 0.40→47.1 GB, gen-small 0.48→36.9 GB). Second -seat to prove it after `gen-small`. Do not remove the pin. - -⚠ **fv-ml1 GPU 0 is now 96.0 of 97.9 GB.** GPU 1 ~5.7 free, GPU 2 ~2.4, GPU 3 is a HELD RESERVE -for a future full-card seat (`flash-next` alone needs 93 of 96 GiB). **There is no room for -another seat on fv-ml1 without a placement decision.** - -⚠ **SUPPORT WAS CHECKED, NOT ASSUMED**, per the training playbook's own lesson that LoRA support -is per-ARCHITECTURE not per-family: `vllm/model_executor/models/qwen3.py:271` declares -`Qwen3ForCausalLM` with `SupportsLoRA` plus `packed_modules_mapping` and `embedding_modules`. -**Do not transplant this compose onto an MoE carrier without re-running that grep** — the -playbook records a LoRA refusal on a Qwen3 MoE arch. - -**Naming (operator, 2026-09-16):** `lv-` — lv for **lang-voice**, retiring `baby*`, which -read fine for one experiment and invites confusion across a family. The adapter NAME is the -request's `model` field, so it is the public API of a voice. Historical persistent-memory entries -still say BabyYarros/BabyHemingway and were deliberately left as dated records. diff --git a/persistent-memory.d/2026-09-17-a-stoplist-entry-is-an-assertion-the-leak-gate-can-no.md b/persistent-memory.d/2026-09-17-a-stoplist-entry-is-an-assertion-the-leak-gate-can-no.md deleted file mode 100644 index 118b4f2..0000000 --- a/persistent-memory.d/2026-09-17-a-stoplist-entry-is-an-assertion-the-leak-gate-can-no.md +++ /dev/null @@ -1,3 +0,0 @@ -# `[2026-09-17]` A stoplist entry is an assertion the leak gate can no longer check - -⭐ **A stoplist entry is an assertion the leak gate can no longer check** — stoplisting removes a surface from the entity map, so a wrongly stoplisted CHARACTER is an undetectable leak. Three were wrong on Brontë (Leaven, Pierrot, Samuel); `scripts/r49-corpus/audit_stoplist.py` finds them by honorific and now gates the pipeline. Commit `8bb7686`. diff --git a/persistent-memory.d/2026-09-17-a-unit-splitter-must-choose-by-size-not-by-count-the.md b/persistent-memory.d/2026-09-17-a-unit-splitter-must-choose-by-size-not-by-count-the.md deleted file mode 100644 index d6726db..0000000 --- a/persistent-memory.d/2026-09-17-a-unit-splitter-must-choose-by-size-not-by-count-the.md +++ /dev/null @@ -1,3 +0,0 @@ -# `[2026-09-17]` A unit splitter must choose by SIZE, not by count — the inherited rule silently produced 22,000-word "chapters". - -⭐⭐ **A unit splitter must choose by SIZE, not by count — the inherited rule silently produced 22,000-word "chapters".** `scripts/r49-corpus/split_units.py`: a marker mode qualifies only if its median unit is inside [600, 12000] AND no unit holds half the work; among qualifying modes PRIORITY breaks the tie (contents > chapter-word > roman > bare-numeral > caps-title), and paragraph-block sections are the fallback for works with no divisions. ⭐ **Both rules exist because a control caught them**: scoring by "median closest to target" chose `caps-title` (6 units, one holding **97%** of the book) over True at First Light's real 20 chapters, because a median cannot see that distribution and a max bound can. Positive control: 8/10 Hemingway works reproduce the shipped mode and count exactly. Negative control: 40,000 words with no blank lines → 1 unit, refuses to fabricate divisions. Commit `705fa3a`. diff --git a/persistent-memory.d/2026-09-17-auditentitymap-py-the-rename-can-damage-the-prose-and-no.md b/persistent-memory.d/2026-09-17-auditentitymap-py-the-rename-can-damage-the-prose-and-no.md deleted file mode 100644 index 7442b9f..0000000 --- a/persistent-memory.d/2026-09-17-auditentitymap-py-the-rename-can-damage-the-prose-and-no.md +++ /dev/null @@ -1,3 +0,0 @@ -# `[2026-09-17]` `audit_entity_map.py` — the rename can DAMAGE the prose and no gate will ever say so. - -⭐ **`audit_entity_map.py` — the rename can DAMAGE the prose and no gate will ever say so.** Mirror of `audit_stoplist.py`: surfaces wrongly held IN the map rather than out of it. 130 of 946 Hemingway surfaces flagged (1,616 instances, 0.162% of words) — `African`, `Chinese`, `X-ray`, `Coca-Cola`, `Ritz`, `Prado` renamed into invented names — plus 16 bare initials incl. `C` at 274 occurrences. Signal is a preceding article; controls derived from the corpus, not hand-picked. Every hit reported for READING: `the Widow` and `the Informer` are genuine epithet-names that should be renamed. Commit `051b99e`. diff --git a/persistent-memory.d/2026-09-17-beat-contamination-leak.md b/persistent-memory.d/2026-09-17-beat-contamination-leak.md deleted file mode 100644 index 1e95342..0000000 --- a/persistent-memory.d/2026-09-17-beat-contamination-leak.md +++ /dev/null @@ -1,39 +0,0 @@ -# `[2026-09-17]` The beat-writing model restores the author's real names — a leak the corpus gate structurally cannot see - -⭐⭐ **THE RENAME AND ITS GATE ARE BOTH WORKING, AND THE TRAINING DATA IS STILL CONTAMINATED.** -The corpus gate reads the corpus and the renamed copies. **It never reads the generated -instruction beats.** Those are written by an LLM that just read the passage — and if it -recognises the book, it supplies the canonical names out of its own training. - -**Measured on the first 714 lv-bronte pairs, before the filter existed:** - -- **13 beats (1.8%)** named source characters — `Rochester` ×6, `Jane` ×3, `Brocklehurst` ×2, - `Beck`, `Fairfax`, `Helen`, `Burns`, `Eyre`, `Reed`, `Rivers`. -- **0 of 714 RESPONSES did.** The rename was perfect; the instruction side was not. -- One beat read *"Saoirse confirms Rochester's flaws, then agrees in English to marry him"* — - a renamed name and a canonical one in the same sentence, which is the mechanism in miniature. - -**Why it matters more than 1.8% sounds:** the beat is the INSTRUCTION half of the pair, so -training on it re-teaches exactly the inventions the rename pipeline exists to remove. - -⚠⚠ **EXPOSURE SCALES WITH HOW WELL THE GENERATOR KNOWS THE BOOK.** It is worst for -public-domain classics and mildest for recent work. That is precisely why the Yarros and -Hemingway runs came up clean and Brontë did not — **their clean runs are NOT evidence they are -immune.** Both should be re-verified, and regenerated with `--source-entities`, before their -pairs are trusted again. - -**The fix.** `vet()` in `scripts/yarros-corpus/build_sft_pairs.py` gained a `sourcename` -reject plus `--source-entities `, taking the UNRENAMED entity map. Fired at -~3% of attempts on the Brontë rebuild. Commit `533cc0c`. - -**The end-to-end guard that proves it.** The chain now verifies every built pair — beat, -response and context — against every source surface before spending GPU hours: -`[verify] 3858 pairs vs 368 source surfaces -> 0 leaks`. - -⚠ A guard stricter than the gate cries wolf. The first verify pass excluded nothing and -flagged `Monsieur’` ×14 and `Qu’est-ce` ×4 — French grammar, not leaks — because -`--fold-clitics` leaves apostrophe keys the gate deliberately skips. Mirror `leak_gate.py`'s -own predicate; a guard that fails on false positives gets disabled, which is worse than the -leak it guarded. - -Related: [[2026-09-17-lv-bronte-gate]], [[2026-09-16-lv-hemingway-corpus]], [[2026-09-16-lv-voices-line]]. diff --git a/persistent-memory.d/2026-09-17-esh-fiber-outages.md b/persistent-memory.d/2026-09-17-esh-fiber-outages.md deleted file mode 100644 index 4e15fcf..0000000 --- a/persistent-memory.d/2026-09-17-esh-fiber-outages.md +++ /dev/null @@ -1,68 +0,0 @@ -# `[2026-09-17]` ESH: Cityside Fiber failed twice in six hours; site ran on Verizon failover - -**Timeline (PDT).** - -``` -19:09:07 Cityside dies. UDM fails over to Verizon 5G (WAN2). Site stays up at ~33 ms. -19:51 Verified healthy on failover. -20:01:37 esh-scale drops off the headscale mesh; 10.0.0.0/16 withdrawn; whole site dark - from the colo. Beszel fires on all five ESH hosts. -20:11-15 Service restored. Operator had switched WAN1 to DHCP to get Cityside working at - all; it came back on CGNAT, not the static. Latency back to 9 ms. -01:06:23 Cityside fails AGAIN. Failover to Verizon. Site up, ~37 ms. -``` - -⭐ **The 20:01 blackout was most likely the operator's own WAN reconfiguration**, not ISP -instability — switching the WAN type bounces the interface, esh-scale loses its path, -headscale withdraws the route, and the site vanishes from the colo's view until it settles. -An earlier session theory ("Cityside came back half-provisioned and the UDM failed back into -an unstable session") is retired. - -⚠ **The diagnostic that mattered: physical link stayed UP at 2.5 GE with zero errors -throughout, while the ISP's next-hop `128.177.138.181` was unresponsive.** So "the ONT is -fine, it is upstream of the ONT" — the line to give Cityside. Traceroute from NH3 reached -`209.249.146.170` (one hop short) before dying, so the prefix was still routed. - -⚠ **CROWDSEC ROTATION FRAGILITY IS LIVE.** The `esh` allowlist on ana-docker carries the now- -dark static `128.177.138.182` (never-expiry), plus `97.190.18.88` (Verizon failover) and -`23.164.40.174` (Cityside CGNAT), both **7-day expiry**. ESH is on a rotating carrier-NAT -egress until the static is restored — the exact regime the 09-08 static purchase was meant to -end, and the class that once blackholed the whole site via a false ban. **If ESH loses colo -access, check `curl -s4 ifconfig.me` from esh-docker-vm FIRST** and allowlist the new address. - -**Still pinned to the dark static and broken until it returns:** FortiGate `infra-ops` -trusthost3 = `128.177.138.182`, so logins to ana-gw from ESH are refused. The dormant -`esh-ana` IPsec is bound to wan1/static (disabled, so no impact). - -⭐ **The mesh was NOT degraded on CGNAT** — tailscale hole-punched straight through -(`direct 23.164.40.174:41641`), which is why latency read 9 ms rather than a DERP figure. An -expectation of relay-on-CGNAT was wrong. - -## RESOLVED 2026-09-17 ~12:30 PT — the static is back, confirmed on four axes - -Not one check, because egress alone cannot tell a static WAN from a carrier NAT that happens -to answer (see auto-memory `feedback_egress_ip_cannot_detect_cgnat`): - -``` -config UDM WAN1 `wan_type = static`, ip 128.177.138.182, mask /30, gw 128.177.138.181 - — switched BACK from the DHCP the operator set at 20:11 during the outage -active stat/health: isp_name "Cityside Fiber", ASN 18731, num_disconnected 0. - WAN2 Verizon-5G is failover-only at priority 2 and idle. -egress esh-docker-vm sees 128.177.138.182 — EQUAL to the WAN ip, so not behind CGNAT -perf 2005/2142 Mbps symmetric; colo -> ESH 5.0 ms, 0% loss over 4 hosts-worth of pings - (Cityside CGNAT read 9 ms, Verizon failover 33-37 ms) -``` - -⭐ **The FortiGate pin un-broke itself and that was verified, not inferred.** `infra-ops` -trusthost3 is `128.177.138.182`; from esh-docker-vm, ana-gw `tcp/22` is OPEN and the -FortiGate offers a password prompt rather than dropping the connection — a trusthost -mismatch refuses outright, so reaching auth *is* the trusthost passing. The dormant -`esh-ana` IPsec bind to wan1/static is correct again (still disabled, still no impact). - -⚠ **The crowdsec temporary allowlist entries are being LEFT to expire on their own** -(2026-09-23): `97.190.18.88` Verizon and `23.164.40.174` Cityside CGNAT. Cityside failed -twice in six hours on 09-17, so until the line has earned some confidence those two are -cheap insurance against the exact false-ban blackout this rotation-fragility caused before. -`128.177.138.182` stays never-expiry. - -Related: [[2026-09-06-headscale-cutover]], [[2026-09-08-esh-static-wan-followups-and-ytvc]]. diff --git a/persistent-memory.d/2026-09-17-gitea-was-reaching-the-public-route-from-every-repo-on-nh3.md b/persistent-memory.d/2026-09-17-gitea-was-reaching-the-public-route-from-every-repo-on-nh3.md deleted file mode 100644 index 52aa3f6..0000000 --- a/persistent-memory.d/2026-09-17-gitea-was-reaching-the-public-route-from-every-repo-on-nh3.md +++ /dev/null @@ -1,3 +0,0 @@ -# `[2026-09-17]` gitea was reaching the PUBLIC route from every repo on nh3-dev - -**gitea was reaching the PUBLIC route from every repo on nh3-dev** — brokkr-smithy, sleipnir, Galdrabok, kvasir — and brokkr-smithy is pushed several times a week, so the fail2ban trigger was live, not dormant. Measured before acting (no split-horizon rewrite, no ssh alias, `ssh -G` confirmed port 22 to 38.120.12.44). Fixed by overriding the NAME once in `~/.ssh/config` rather than rewriting N remotes, so fresh clones and unaudited repos are covered too. Verified with a real `git ls-remote`, not by inspection. Commit `dcc1abc`. Flagged by brokkr-smithy-dev; `vh/imogen` created for them the same session. diff --git a/persistent-memory.d/2026-09-17-headscale-now-split-dnses-nh3-phasefinal-com-to-the-three.md b/persistent-memory.d/2026-09-17-headscale-now-split-dnses-nh3-phasefinal-com-to-the-three.md deleted file mode 100644 index 75a3ae9..0000000 --- a/persistent-memory.d/2026-09-17-headscale-now-split-dnses-nh3-phasefinal-com-to-the-three.md +++ /dev/null @@ -1,3 +0,0 @@ -# `[2026-09-17]` headscale now split-DNSes `nh3.phasefinal.com` to the three AdGuards, so mesh clients can resolve the internal-only wildcard - -**headscale now split-DNSes `nh3.phasefinal.com` to the three AdGuards, so mesh clients can resolve the internal-only wildcard** (`talk`, `booth` — public DNS has no record for them; the fleet AdGuard answers 10.100.10.50). Operator-approved, scoped to nh3 rather than all of `phasefinal.com`. Config `/etc/headscale/config.yaml` in CT 106 on nh3-pve, backup `config.yaml.bak-2026-09-17-splitdns`, restarted, and the new route **read back from a node's netmap** rather than assumed. ⚠ Two things worth knowing: split DNS works fine here with `global: []` — headscale issue #1161's "split ignored without global" does NOT apply to v0.29.3, verified on the live mesh — and `override_local_dns: true` would REQUIRE global, which is the config that makes a roaming laptop lose ALL DNS when the mesh is down. That is why split, not global. Routing was never the problem: nh3-scale already serves 10.100.0.0/16. diff --git a/persistent-memory.d/2026-09-17-hemingway-ships-as-is-operator-ruled-ship-stands-on-both.md b/persistent-memory.d/2026-09-17-hemingway-ships-as-is-operator-ruled-ship-stands-on-both.md deleted file mode 100644 index 03f844b..0000000 --- a/persistent-memory.d/2026-09-17-hemingway-ships-as-is-operator-ruled-ship-stands-on-both.md +++ /dev/null @@ -1,3 +0,0 @@ -# `[2026-09-17]` Hemingway ships as-is: operator ruled "ship stands" on both measured corpus defects - -**Hemingway ships as-is: operator ruled "ship stands" on both measured corpus defects** — the 0.96% beat contamination and the 130 non-name entity-map surfaces. `audit_pairs_sourcenames.py --filter-out` and `audit_entity_map.py` exist and are the instruments if that is ever revisited; neither was run against the shipped adapter. diff --git a/persistent-memory.d/2026-09-17-lv-bronte-gate.md b/persistent-memory.d/2026-09-17-lv-bronte-gate.md deleted file mode 100644 index 491a316..0000000 --- a/persistent-memory.d/2026-09-17-lv-bronte-gate.md +++ /dev/null @@ -1,135 +0,0 @@ -# `[2026-09-17]` lv-bronte: corpus gated for real, adapter trained, SHIPPED with a FAILED voice axis on the record - -**Status: SHIPPED 2026-09-17 01:24 as `lv-bronte` on `vllm-voices` (fv-ml1 GPU0 :8027), ckpt475 — -and it did NOT pass its voice gate.** Shipped because it is additive (one more named LoRA beside -`voices-base` and `lv-yarros`, reached only by requesting it), reversible (one compose line; hot-unload -measures 0.003 s), and clean on the SAFETY axis — 8-gram overlap identical to the never-saw-it control, -on a public-domain corpus. VRAM cost was nil: GPU0 96092 -> 96090 MiB. The caveat is written into the -compose file and into `/tank/aimodels/voice-adapters/lv-bronte-4b-v1/README.md` so it cannot be read -as a clean pass by anyone who finds the adapter without finding this note. - -⚠ **Do NOT cite lv-bronte as evidence pair-SFT works for this author.** The voice axis is unresolved, -not passed. - -## The gate result, in full - -| axis | result | numbers | -|---|---|---| -| **A. VOICE** | ❌ **FAIL** (both candidates) | ckpt925 +0.210, ckpt475 +0.193 vs base — both **under** the 0.251 measured noise floor | -| **B. NOT COPIED** | ✅ PASS | ckpt475 **0.00 hit-rate, max 0 — identical to the never-saw-it control**; ckpt925 0.01, max 8. Positive control saturates at 160, so the detector is not blind | -| **C. NO DAMAGE** | ✅ PASS | ran-on +0.15 against a 0.400 floor | - -``` -same-author target (held-out Brontë vs itself) delta_cb 0.338 <- best achievable -ckpt925 0.531 -ckpt475 0.548 -base-unadapted 0.741 -``` - -⭐ **THE EFFECT LOOKS REAL AND SUBSTANTIAL; THE INSTRUMENT CANNOT CERTIFY IT.** The reachable -span is 0.741 → 0.338 = 0.403, and the adapters closed **48–52% of everything achievable**. -Both beat base on *every individual seed*. This is an UNDERPOWERED result, not a null one — -and a "no effect" without its floor is unfalsifiable, so: **this method cannot resolve a voice -improvement smaller than ~0.251 delta_cb at 30 beats × 4 seeds on this corpus.** - -⭐⭐ **THE CAUSE IS STRUCTURAL: Brontë's val split yields 81 pairs (44 in-band) against -Hemingway's 200**, because the corpus is 678k words against 994k. Maxing the fixture 30 → 44 -beats would shrink the floor by only ~√1.47 ≈ 1.2× (to ~0.21, against a 0.21 gap) — still -marginal. **More SEEDS would not help either: the floor is a RANGE statistic, and ranges widen -with more samples.** There is no cheap fix. - -## ⚠ A DEFECT IN THE v2 RULE ITSELF, worth fixing before the next author - -The floor is defined as the **largest within-arm seed spread across ALL arms**. Measured here: - -``` -base-unadapted 0.772 0.813 0.751 0.772 spread 0.062 -ckpt475 0.670 0.631 0.604 0.578 spread 0.092 -ckpt925 0.776 0.584 0.525 0.620 spread 0.251 <- sets the floor, on ONE seed -``` - -So **adding a third, noisier arm raised the bar that failed the clean one.** Run as the -two-arm gate (base + ckpt475) the floor would have been 0.092 and +0.193 would have cleared -at 2.1×. This was NOT exploited — picking the floor that passes your preferred answer is the -threshold-chosen-after-seeing-the-numbers failure the pre-registration exists to prevent — but -the rule should say whether the floor is computed over the compared pair or over every arm -present. As written, a candidate's verdict depends on which *other* arms you happened to run. - -**The outlier was diagnosed, not waved away.** Degeneracy probe (fraction of a generation made -of its most repeated 5-gram) is uniform across every seed and both arms, 0.0078–0.0102. Seed -1234 is not a collapsed generation; delta_cb genuinely has that variance for that arm. - -## Which checkpoint, if it ships: **ckpt475** - -The two are 0.017 apart on voice — far inside any floor, i.e. indistinguishable. On the axes -that DO resolve, ckpt475 wins both: memorisation identical to the control (ckpt925 has a -verbatim 8-gram hit), and **2.7× tighter seed-to-seed variance** (0.092 vs 0.251) with no -degeneracy to explain the difference — consistent with ckpt925 sitting nearer the overfit -boundary. Given a coin-flip on voice, take the one that provably did not memorise. - -⭐ **THE RECIPE DID NOT TRANSFER.** Yarros and Hemingway both found their minimum inside -epoch two. Brontë's minima are step 475 (ep 1.00, 2.6107) and step 925 (ep 1.96, 2.6129) — -**0.0022 apart against a 0.0046 median neighbour jitter, i.e. indistinguishable**. Epoch 2 -buys Brontë NOTHING over epoch 1. What IS robust is the epoch-3 collapse: +0.075, ~16× jitter. -Do not carry "two epochs on a three-epoch schedule" to a new author as settled. - -## Artefacts - -`gx10:~/lv-bronte/` (corpus-clean, corpus-renamed, entities-final.json, pairs/, beats-bronte-30.json), -`gx10:~/r49-runs/bronte-4b-pairs-3ep/` (57 checkpoints kept), `gx10:~/r49-runs/bronte-eval/` -(three arms × 120 generations, memorization.txt, voice_distance.txt, score.*.txt). -Commits `fc834a8` `533cc0c` `7964d07` `e9e8c40` `8bb7686`. - -⚠ Two output labels in `voice_distance.py` are hardcoded Yarros strings — it prints -"reference: held-out Yarros" and a boilerplate "Base < Instruct" corroboration line. The -NUMBERS are Brontë's; those two labels are not. Not yet fixed. - -Related: [[2026-09-16-lv-voices-line]], [[2026-09-16-lv-hemingway-corpus]], [[2026-09-16-voices-seat-lora]]. - ---- - -## ⚠⚠ AMENDED 2026-09-17 — THE VOICE AXIS PASSES UNDER THE CORRECTED FLOOR RULE - -Everything above is left verbatim; it is what was believed at ship time. This section is -the correction, not a rewrite. - -**The defect this file itself named was fixed, and fixing it flips ckpt475's verdict.** -The section "⚠ A DEFECT IN THE v2 RULE ITSELF" above says the rule should state whether -the floor is computed over the compared pair or over every arm present. It is now -**pairwise**, pre-registered in `scripts/hemingway-corpus/GATE-PREREG.md` before a single -lv-hemingway number existed. Re-scoring the SAME 360 generations — no re-run, no changed -delta_cb: - -``` - arm delta_cb per-seed spread - ckpt925 0.531 (0.776 0.584 0.525 0.620) 0.251 - ckpt475 0.548 (0.670 0.631 0.604 0.578) 0.091 - base-unadapted 0.741 (0.772 0.813 0.751 0.772) 0.062 - - all-arms floor (as run) 0.251 - ckpt475 +0.193 vs pairwise floor 0.091 -> MOVED toward Brontë, 2.1x <- the two rules DISAGREE - ckpt925 +0.210 vs pairwise floor 0.251 -> within the floor, NOT a finding -``` - -⭐ **The sequence matters and is the reason this is not threshold-shopping.** The previous -session found the defect, recorded it, and explicitly declined to exploit it. The rule was -then changed prospectively on a structural argument independent of the answer it produces — -the sampling variability of a difference A−B depends on A and B, never on a third arm C, so -a candidate's verdict must not depend on which other arms were generated. `voice_distance.py` -prints both floors and flags disagreement, so neither number can be quoted alone. - -**Consequences:** -- lv-bronte's voice axis is a **PASS at 2.1x**, not a fail. The caveat is amended in place - (append-only) in `stacks/voices-seat/compose.yaml` and - `/tank/aimodels/voice-adapters/lv-bronte-4b-v1/README.md`. -- The sensitivity floor for that measurement is **0.091**, not 0.251. -- "Do not cite lv-bronte as evidence pair-SFT works for this author" is **WITHDRAWN**. -- ckpt475 over ckpt925 is unchanged and for unchanged reasons (no verbatim 8-gram hit, - 2.7x tighter seed variance). -- The "no cheap fix for the underpowered result" analysis above is superseded for Brontë: - it was underpowered against an inflated floor, not against its own. - -**Also amended:** the two hardcoded Yarros labels flagged at the end of this file are fixed. -`voice_distance.py --author` is now REQUIRED — the committed Brontë output literally reads -"reference: held-out Yarros" over Brontë's numbers — and the stale "one seed-pair per arm / -corroborates Base < Instruct" footer now reports what the run actually carries. diff --git a/persistent-memory.d/2026-09-17-lv-hemingway-gate.md b/persistent-memory.d/2026-09-17-lv-hemingway-gate.md deleted file mode 100644 index df7be7d..0000000 --- a/persistent-memory.d/2026-09-17-lv-hemingway-gate.md +++ /dev/null @@ -1,163 +0,0 @@ -# `[2026-09-17]` lv-hemingway: SHIPPED on ckpt850 — the line's first clean voice pass, and one axis that needs reading - -**Status: SHIPPED 2026-09-17 03:33 as `lv-hemingway` on `vllm-voices` (fv-ml1 GPU0 :8027), -checkpoint-850.** Seat healthy 190 s after recreate, four models served -(`voices-base`, `lv-yarros`, `lv-bronte`, `lv-hemingway`), GPU0 96,092 → **96,090 MiB** — a -LoRA rides inside the existing seat and costs nothing. Adapter verified byte-identical to -the checkpoint by sha256 across two hops. - -Gate design **pre-registered before any generation existed**: -`scripts/hemingway-corpus/GATE-PREREG.md`, commit `0bb4938`. - -## The gate result — 3 arms × 60 held-out beats × 4 seeds = 240 generations per arm - -| axis | result | numbers | -|---|---|---| -| **A. VOICE** | ✅ **PASS, 6.4×** | +0.413 delta_cb vs base, pairwise floor 0.064. Also clears the OLD all-arms floor (0.113) — **this verdict does not depend on the rule change** | -| **B. NOT COPIED** | ⚠ **content clean, rate 7× the author's own** | 0.07 hit-rate, mean-longest 0.6, **max 9 words**. Base 0.00, **held-out Hemingway 0.01** | -| **C. NO DAMAGE** | ✅ PASS | ran-on +0.08, on-beat −0.14, both inside a 0.217 floor; in-band 0.79 vs base 0.05 | - -``` -same-author target (held-out Hemingway vs itself) delta_cb 0.364 <- best achievable -ckpt1750 0.439 -ckpt850 (SHIPPED) 0.511 -base-unadapted 0.924 -``` - -⭐ **THE STRONGEST VOICE RESULT IN THE LINE. The span is 0.924 → 0.364 = 0.560 and ckpt850 -closed 73.8% of it (ckpt1750 86.6%)**, against lv-bronte's 48%. Power came from the corpus, -not from a better method: 173 in-band val pairs allowed a **60-beat** fixture where Brontë -had 44 in-band and could only run 30. - -## ⚠⚠ AXIS B — THE COMFORTABLE EXPLANATION WAS WRONG, AND THE CONTROL IS THE ARTIFACT - -`memorization_check.py` uses the **base-unadapted arm** as its negative control, and on this -corpus that control is weak in one direction only — **it makes an innocent arm look guilty.** -Base writes 18,035 words of *summary*; the adapted arms write 27,413 of *pastiche*. Text that -does not imitate the register cannot collide with its n-grams, so base's 0.00 partly measures -"different register", not "did not memorise". - -The obvious hypothesis was that Hemingway's plain, high-frequency register makes 8-gram -collisions inevitable for any arm that learns it. **That hypothesis is refutable, was tested, -and is FALSE.** New control: **held-out Hemingway — the author himself, val text no arm -trained on — scored against the train split**, chunked to the generations' own median length -(101 words) so the comparison is like for like. - -``` -sample n hit-rate mean-longest max -HELD-OUT HEMINGWAY (never trained) 370 0.01 0.1 10 -base-unadapted 240 0.00 0.0 0 -ckpt1750 240 0.08 0.7 9 -ckpt850 (SHIPPED) 240 0.07 0.6 9 -positive control (train vs train) 160 <- not blind -``` - -⭐⭐ **The adapter reproduces train-corpus word sequences ~7× more often than the author -reproduces himself.** If the register explained it, real Hemingway would collide at the same -rate; it collides at 0.01. - -⭐ **And the exposure is still nil, which is a different question from the rate.** All 19 -matched runs were READ, not counted. Every one is stock dialogue — `i don t think so the girl -said`, `came over and sat down at the table`, `how do you feel i feel very well`. No plot, no -imagery, no distinctive phrase, **no proper noun** (the one name-shaped hit, `swift tristan`, -is the RENAMED invented name, not Hemingway's). The longest run is **9 words — shorter than -the 10-word run genuinely unseen Hemingway shares with the train split by coincidence.** - -What is being reproduced is the *grammar of his dialogue*, which is the thing the adapter -exists to learn, rendered in the commonest words in English. **Elevated rate, zero -protectable content.** Hemingway is in copyright; the in-line precedent is lv-yarros, also in -copyright, shipped at 0.10 against a 0.07 control. Unload is 0.003 s and one compose line. - -⚠ **The durable lesson is about the instrument, not this adapter: a negative control that -differs from the candidate in a way CORRELATED with the metric is not a control.** Always ask -what the metric returns for a known-innocent sample *in the same register*. - -## Why ckpt850 and NOT ckpt1750, the loss minimum - -ckpt1750 has the better point estimate on voice (0.439 vs 0.511) and **it is not usable**: - -``` -gap between candidates 0.072 -pairwise floor max(0.113, 0.050) 0.113 -> NOT resolvable -``` - -Indistinguishable, so the pre-registered tiebreak falls to the axes that resolve — and -**ckpt850 wins every one**: - -| | ckpt850 (shipped) | ckpt1750 | -|---|---|---| -| seed spread | **0.050** | 0.113 — **2.3× wider** | -| memorisation hit-rate / mean-longest | **0.07 / 0.6** | 0.08 / 0.7 | -| ran-on | **0.08** | 0.12 | -| epoch | **0.959** | 1.973 | - -ckpt1750's spread is one seed: 0.491, 0.449, 0.468, then **0.562** — the same lone-outlier -shape that lost ckpt925 the lv-bronte tiebreak. - -⭐ **THE TWO-EPOCH RECIPE DID NOT TRANSFER HERE EITHER — it is now 0 for 2.** Hemingway's -minimum really is step 1750, but step 850 is **+0.0040 against a 0.0044 median neighbour -jitter**, with three checkpoints inside one jitter of the best. Epoch 2 buys nothing that -resolves and costs 2.3× the variance. Only the epoch-3 collapse is robust: **+0.0762 = 17.4× -jitter**, which is why `adapter/` was never gated. **Stop carrying "two epochs on a -three-epoch schedule" forward; read the curve and prefer the earlier tied checkpoint.** - -## Pre-flight: the beat leak IS present in Hemingway, and the fixture is clean - -`audit_pairs_sourcenames.py` (new, commit `0bb4938`) closes the blind spot `leak_gate.py` has -by construction. Controls green every run: 941/941 surfaces found in the unrenamed source, -nonce absent from both trees, 6/6 planted names detected. - -``` -train beats 70 of 7,094 (0.96%) Santiago x16, Catherine x7, Rinaldi x3, Brett, Harry, - Jake, Pablo, Nick, Maria, Helen ... 36 distinct -train responses 0 of 7,294 -- the rename itself held perfectly -val beats 0 of 200 -- THE EVAL FIXTURE IS CLEAN; the gate is unconfounded -``` - -⭐ The beat-only signature is exactly lv-bronte's. **Yarros's and Hemingway's earlier clean -runs were never evidence of immunity** — they predate the detector. - -**Cross-validated on real data** where the answer was already recorded: the fixed Brontë -pairs return **0 of 3,858** (matching "0 leaks across 3,858 pairs"), and -`pairs-full.CONTAMINATED.jsonl` returns **15 of 792 = 1.89%** with Rochester ×6, Jane, -Brocklehurst ×2, Beck, Fairfax, Burns, Helen, Eyre — against a record of "13 of the first 714 -beats (1.8%)" with the same names. An independently written instrument reproducing a -documented finding at the right magnitude is what makes its zeroes mean *absent*, not *blind*. - -`--filter-out` produces a clean **7,024-pair** set in one command (70 dropped, 0.99%), -verified by re-audit at 0 of 7,024. **A retrain on it is the operator's call, not done.** - -## ⚠ A SECOND corpus defect, measured and NOT acted on - -`audit_entity_map.py` (new, commit `051b99e`) is the mirror of `audit_stoplist.py`: it finds -surfaces wrongly held **IN** the entity map, which `leak_gate.py` cannot see because it only -ever asks whether the author's names are GONE, never whether non-names were spared. - -``` - positive control `other` 764/1356 article-preceded = 0.56 - negative control 100 honorific-confirmed people, highest Inglés 0.26, bulk 0.00-0.06 - FLAGGED 130 of 946 surfaces · 1,616 instances · 0.162% of corpus words -``` - -`African`, `Chinese`, `Basques`, `Republican`, `Communist`, `X-ray`, `Coca-Cola`, `Ritz`, -`Prado`, `Cezanne` were all renamed into invented proper nouns. **Some flags are correct -renames** — `the Widow`, `the Informer` are genuine Hemingway epithet-names — so every hit is -reported for reading, never auto-removed. Plus **16 bare initials in the map**, `C` at 274 -occurrences: the same class as the `G` caught by hand about to be renamed 248 times. - -At 0.162% of words this did not block the ship. It is the thing to fix first if a corpus -rebuild ever happens. - -## Artefacts - -`gx10:~/lv-hemingway/` (corpus-clean, corpus-renamed, beats-hemingway-60.json + sidecar, -eval-hemingway.sh, voice-prep.py, eval.log), `gx10:~/r49-runs/hemingway-4b-pairs-3ep/` -(54 checkpoints kept), `gx10:~/r49-runs/hemingway-eval/` (three arms × 240 generations, -memorization.txt, voice_distance.txt, score.*.txt). -`fv-ml1:/tank/aimodels/voice-adapters/lv-hemingway-4b-v1/` (adapter + a README carrying the -axis-B caveat, so it cannot be read as clean by anyone who finds the adapter without this). -Commits `0bb4938` `051b99e` `5e66114` `2e9b118`. - -Related: [[2026-09-17-lv-bronte-gate]], [[2026-09-16-lv-hemingway-corpus]], -[[2026-09-16-lv-voices-line]], [[2026-09-16-voices-seat-lora]], -[[2026-09-17-beat-contamination-leak]]. diff --git a/persistent-memory.d/2026-09-17-lv-krakauer-d1-built-126-units-422-880-words-and-its-name.md b/persistent-memory.d/2026-09-17-lv-krakauer-d1-built-126-units-422-880-words-and-its-name.md deleted file mode 100644 index 8511c00..0000000 --- a/persistent-memory.d/2026-09-17-lv-krakauer-d1-built-126-units-422-880-words-and-its-name.md +++ /dev/null @@ -1,3 +0,0 @@ -# `[2026-09-17]` lv-krakauer D1 built — 126 units, 422,880 words — and its name guard caught three defects nothing else would have reported. - -⭐ **lv-krakauer D1 built — 126 units, 422,880 words — and its name guard caught three defects nothing else would have reported.** Back matter searched only the LAST unit while the apparatus sat in unit 37 of 41; relying on the splitter to drop front matter failed because the ebook TOC sits above the author's note and gave it a `Chapter Thirty-Two` to start on; and **zero was the wrong bar** — 2 survivors are Krakauer writing about his own father in Into the Wild's autobiographical chapters, so the allowance is pinned at 2 with every survivor printed. ⚠ Both strips are windowed in the OPPOSITE direction from McCarthy's, because Krakauer's `ALSO BY`/`Copyright`/`About the Author` sit at 0.0–0.6% of the file. Commit `4be0630`. diff --git a/persistent-memory.d/2026-09-17-lv-mccarthy-d1-built-167-units-584-756-words-and-the-whole.md b/persistent-memory.d/2026-09-17-lv-mccarthy-d1-built-167-units-584-756-words-and-the-whole.md deleted file mode 100644 index 4329a4b..0000000 --- a/persistent-memory.d/2026-09-17-lv-mccarthy-d1-built-167-units-584-756-words-and-the-whole.md +++ /dev/null @@ -1,3 +0,0 @@ -# `[2026-09-17]` lv-mccarthy D1 built — 167 units, 584,756 words — and the whole job was protecting a style that reads as damage. - -⭐ **lv-mccarthy D1 built — 167 units, 584,756 words — and the whole job was protecting a style that reads as damage.** 0.0 quote marks per 10k (Hemingway 838), `dont`/`aint`/`wont`. The builder runs NO typography normalisation and asserts the quote density afterwards. Two truncated catalogue rows dropped for complete mobi siblings; all 15 containment pairs measured (worst 0.10%); back matter in 4 of 6 works carried the author's name 26 times → 0. ⚠ The back-matter strip runs BEFORE the split here — Blood Meridian and The Crossing end with a dumped TOC of bare roman numerals, the exact shape of a chapter marker. Commit `f3bf3ca`. diff --git a/persistent-memory.d/2026-09-17-lv-mccarthy-s-d1d3-chain-was-recovered-not-remembered-there.md b/persistent-memory.d/2026-09-17-lv-mccarthy-s-d1d3-chain-was-recovered-not-remembered-there.md deleted file mode 100644 index 24f2ece..0000000 --- a/persistent-memory.d/2026-09-17-lv-mccarthy-s-d1d3-chain-was-recovered-not-remembered-there.md +++ /dev/null @@ -1,3 +0,0 @@ -# `[2026-09-17]` lv-mccarthy's D1→D3 chain was RECOVERED, not remembered — there was no runbook and the commands went over non-interactive ssh, so no history survived. - -**lv-mccarthy's D1→D3 chain was RECOVERED, not remembered — there was no runbook and the commands went over non-interactive ssh, so no history survived.** Rebuilt candidates and matched sha256 against the artifacts on disk: 6 works, the entity map, the final map and all 36 copy files byte-identical. Now pinned in `scripts/mccarthy-corpus/RUNBOOK.md` with every deviation. ⚠ **D1 must run on nh3-dev** (the builder reads the kvasir catalogue by absolute path); the prior "on gx10" note is true of D2 onward only. ⚠ No phrase map exists for this corpus, so the gate's phrase audit never ran — Yarros and Brontë both had one. diff --git a/persistent-memory.d/2026-09-17-mccarthy-d1-d3.md b/persistent-memory.d/2026-09-17-mccarthy-d1-d3.md deleted file mode 100644 index d260984..0000000 --- a/persistent-memory.d/2026-09-17-mccarthy-d1-d3.md +++ /dev/null @@ -1,159 +0,0 @@ -# `[2026-09-17]` lv-mccarthy D1→D3 — built, gated, and every stage caught a defect in the stage before it - -**`~/lv-mccarthy/` on pfi-gx10.** `corpus-clean/` (167 units, 584,716 words), -`corpus-renamed/` (6 copies, 1,002 records), `scripts/`. Commits `705fa3a` `f3bf3ca` -`0fa68cb` `5aa10bf` `5ddb047`. - -``` -leak gate 0 of 75 renameable and 0 of 37 sub-threshold survive in any copy - positive control 108/108 surfaces found in the unrenamed source - negative control nonce absent from both trees -``` - -## The shared splitter: choose by SIZE, not by count - -`scripts/r49-corpus/split_units.py`. The inherited rule was "most units above a floor", which -is wrong for any book whose markers are PARTS: - -``` -Cities of the Plain 4 roman marks -> 4 units, median 22,312w -The Crossing 4 roman marks -> 4 units, median 37,310w -``` - -Four beats one, so it won, and the old guard only fired at exactly one unit. Now: a mode -qualifies only if its median unit is inside **[600, 12000]** AND no unit holds half the work; -among qualifying modes **priority** breaks the tie (contents > chapter-word > roman > -bare-numeral > caps-title). Works with no divisions fall back to **paragraph-block sections**. - -⭐⭐ **The first version of that rule was WORSE than what it replaced, and a control caught -it.** Scoring by "median closest to target" chose `caps-title` over the real chapters of -Hemingway's *True at First Light*: - -``` -bare-numeral 20 units median 5,337w max 11,155 <- the book's own chapters -caps-title 6 units median 777w max 113,886 <- median looked BETTER -``` - -Five stray all-caps lines gave five tiny units beside **one holding 97% of the book**. A median -cannot see that distribution; a max bound can. Controls green both ways afterwards: 8/10 -Hemingway works reproduce the shipped mode and count exactly, and 40,000 words with no blank -lines returns **1 unit** rather than fabricating sections. - -⚠ The Hemingway builder is deliberately NOT repointed at this module — its corpus is shipped -and its sha is pinned by a live adapter. - -## D1: the job was protecting a style that reads as damage - -``` -quote marks 0.0 per 10k (Hemingway 838) -apostrophes 123 per 10k (Hemingway 241) `dont` `aint` `wont` `didnt` -``` - -⚠⚠ **`repair_typography.py` MUST NOT be run on this corpus.** It normalises "toward what the -text does" and would put the quotation marks back. The builder runs no normalisation and then -**asserts** the quote density, so a future well-meaning change fails the build. - -⚠ **AND IT MAKES THE VOICE GATE EASY TO PASS FOR THE WRONG REASON.** `voice_distance.py` is -Burrows's Delta over CHARACTER BIGRAMS. An adapter that learns only "emit no quotation marks" -moves delta_cb a long way without having learned a sentence. **Pre-register a -punctuation-normalised secondary read before gating lv-mccarthy.** Tracked in the builder -docstring, commit `f3bf3ca`. - -Also: two truncated catalogue rows dropped for complete mobi siblings; all 15 cross-work -containment pairs measured (worst **0.10%**); back matter in 4 of 6 works carrying the author's -name 26 times → **0**; alphabet re-derived at 1,411 non-ASCII letters across 14 Spanish forms. - -⚠ The back-matter strip runs **BEFORE** the split for McCarthy, inverting the Hemingway order: -Blood Meridian and The Crossing end with a dumped table of contents made of bare roman numerals -on their own lines — the exact shape of a chapter marker. - -## D2 caught a D1 defect: three small-caps manglings - -The entity map returned `E`, `H`, `T`, `K` as renameable entities with 17–33 capitalised -occurrences each — the `G` class from Hemingway, where `G` was about to be renamed to a surname -248 times. Reading them showed the extractor mangled small-caps openings three ways: - -``` -1. SPLIT INITIAL `T HE HOUSE was built` -> `The house was built` 32 cases -2. UNMARKED RUN `THEY STOOD in the doorway` -> `They stood in the doorway` 88 cases -3. LOST INITIAL `HE CANDLEFLAME` -> `THE CANDLEFLAME` 1 case -``` - -Rule 1 requires a FOLLOWING all-caps word, so `A TV was playing` and `A Mexican was changing` -are untouched. Rule 2's `[a-z]` lookahead is what makes it safe — a genuine shout or sign is -not followed mid-sentence by lowercase. All 23 distinct first words of the 88 were checked. - -⚠⚠ **A fourth "fix" was nearly shipped that would have CORRUPTED the text.** `HEY RODE` → -`THEY RODE` looked right from a survey of the BUILT corpus. The raw master has `THEY RODE` -intact, twice — `HEY RODE` matched as a SUBSTRING, and the unanchored replace produced -`TTHEY RODE`, which rule 2 then lowercased to `Tthey rode`. Caught by the count assertion -(expected 1, replaced 2) and settled by reading the master. ⚠ My first corruption check also -missed it, searching for `TTHEY` when the pipeline had already lowercased it — **check the -shape the pipeline emits, not the shape you imagined.** - -## D2's own gates: and `audit_stoplist` was scanning its own rationale - -⚠⚠ **A defect in `audit_stoplist.py`, latent for every corpus before this one.** It built its -surface set from every list value in the stoplist JSON — including `_why`, which by convention -is a LIST OF PROSE LINES. Its empty separator line matched the honorific pattern **139 times**, -printing a flag with no surface name above the one real catch. Now skips `_`-prefixed keys. - -That real catch was a contradiction **inside my own file**: `Franklin` sat in the geography list -(the old name for El Paso) while the same file's note recorded *"I'm here to see Mr Franklin"*, -a lawyer in All the Pretty Horses. A second self-inflicted one: a speculative A–Z fragments list -stoplisted `I` and `A`, and `Sir I dont think I can do that` duly tripped the audit. It is now -the four letters actually measured as entities. - -Everything ambiguous was read in context: **Socorro is the ranch cook, not the New Mexico -town**; Niño, Keno and Redbo are HORSES (renameable, the `Inglés` precedent); Yaqui and Gilenos -are real peoples; Hashknives is a real cattle outfit; Hearst, Trias, Huerta and Madero are real -historical figures on the page under their own names. - -Final: 123 map surfaces, 124-surface stoplist, `entities.py` 27/27 controls, both audits PASS. - -## The human gender pass is an auditable file - -The honorific/window resolver scored **21 correct / 3 held / 1 WRONG** against a 26-name -control; the base-rate proximity resolver built for Hemingway scored 18/6/1 and **its own guard -correctly REFUSED to write**. So the incumbent stands and four entries are fixed by hand in -`gender_overrides_mccarthy.json`, each carrying its evidence. - -⚠ All four are female and all four look male-dominated in raw counts, because this corpus runs -**29,144 male pronouns to 5,036 female — a base rate of 85.3% male**. Carla Jean Moss at -31m/21f would be 44m/8f at that rate; 21 against an expected 8 is decisive. Same arithmetic that -recovered Pilar and Brett on Hemingway. Alfonsa was in the control and is correctly absent from -the map at 4 occurrences, below the min-count floor — an error in the control, not the pipeline. - -`apply_gender_overrides.py` refuses twice: a name absent from the map is an error rather than a -silent no-op, and overruling a gender the detector holds needs an explicit `"correcting": true` -so it cannot look like filling a held entity in a diff. - -## D3: three calls, and the holdout fix that matters most - -1. **`--scope corpus`**, not the per-work default. Nine surfaces appear in more than one work — - Parham (The Crossing + Cities of the Plain), Grady and Cole (All the Pretty Horses + Cities - of the Plain), Socorro, Héctor. A per-work map gives John Grady a different invented name in - each novel, turning one character into two. -2. **A new `mccarthy` preset.** Hemingway's romance pool carries `it_IT` and `fr_FR` for his - Italian and French casts; McCarthy writes neither language. `en_GB` goes for the same reason. - `en_US` + `es_MX`/`es_ES` at an even share. -3. **`--min-cap 5` to match the entity map's floor.** The first gate run FAILED with 45 - survivors: `entities.py` admits cap ≥ 5 while `rename.py` renamed only cap ≥ 8, so every - entity between sat in the map, was never renamed, and counted as a leak. Hemingway never hit - it because its map had `sub_threshold_total: 0`. - -⭐ **`--holdout-chapter` NOW TAKES A LIST.** The val split is one chapter index per work, so its -SIZE is set by how many WORKS a corpus has, not how many words: - -``` -Hemingway 10 works -> 9 val units -> 36,563 words/copy -> gate DECISIVE -Brontë 4 works -> 4 val units -> 17,043 words/copy -> gate MARGINAL -McCarthy 6 works -> 6 val units -> ~18,000 would have been Brontë's end -``` - -Holding out chapters **7 and 17** gives **11 units and 40,653 words per copy — larger than -Hemingway's** — for 7% of the corpus, on a corpus 40% smaller than his. No amount of corpus size -fixes a val split that scales with work count. - -Related: [[2026-09-17-lv-hemingway-gate]], [[2026-09-17-mccarthy-krakauer-d1]], -[[2026-09-17-lv-bronte-gate]]. diff --git a/persistent-memory.d/2026-09-17-mccarthy-split-name-leak.md b/persistent-memory.d/2026-09-17-mccarthy-split-name-leak.md deleted file mode 100644 index 12bbe0f..0000000 --- a/persistent-memory.d/2026-09-17-mccarthy-split-name-leak.md +++ /dev/null @@ -1,104 +0,0 @@ -# `[2026-09-17]` The leak gate passed with five protagonist names still in every copy - -Found during D4 pre-flight, three stages downstream of where it happened. Commit `c559664`. - -``` -leak gate, 2026-09-17 morning 0 of 75 renameable, 0 of 37 sub-threshold, both controls green -actually present, all 6 copies Bell x2 Chigurh x3 Moss x2 Toadvine x4 Glanton x2 -``` - -## The mechanism - -`leak_gate.py` scans `\b(Surface)\b`. **A character inserted inside a name defeats that -pattern outright**, so a mangled occurrence is not merely unrepaired — it is *unrenameable* -by `rename.py` and *unreportable* by the gate, and the gate prints a clean zero over it. -Two extraction artifacts produce exactly that: - -``` -B ell C higurh M oss T oadvine a small-caps drop cap kept as its own token -Toad-vine Glan-ton a print line-break hyphen kept by the extractor -``` - -⭐ **Every VISIBLE occurrence had been renamed correctly** — exact-match survivors were 0, -as the gate said. That is what makes this residue invisible to a spot-read: the names are -gone everywhere you look. `Bell` sits in the entity map at 147 capitals, `Glanton` at 365. - -This is the third member of a family. lv-bronte's was `_Antigua_` (`_` is a word character, -so `\bAntigua\b` cannot match inside it), found by hand in 2026-09-16 and never generalised. -**The generalisation is the point: any separator inside a name blinds a word-boundary scan.** - -## Fixed at three levels, and all three must stay - -1. **`build_corpus_mccarthy.py` rules 4 and 5** repair the source text — 32 split initials - with a *lowercase* remainder (rule 1 requires a following ALL-CAPS word and DROPCAP - requires two, so this is the class both leave behind), 5 hyphen-split names by name. - Both carry expected counts so a master change fails the build. - ⚠ **Rule 4's letter class is consonants only.** `I` opens **1,966** paragraphs (the - pronoun), `A` opens 143 (the article), `Y` opens 32 (Spanish *y*). Folding any of them - would corrupt 2,141 lines to fix 32 — the same `I`/`A` trap that bit `audit_stoplist.py`. -2. **`leak_gate.py` runs a separator-tolerant pass every time**, with its own positive and - negative controls, and **it fails the gate**. Validated against the pre-fix tree: reports - all five surfaces, exits 1. -3. The exact-match passes are untouched, so the old verdict is reproduced alongside the new. - -⚠ **The fragment filter is what makes the new pass usable.** A naive separator-tolerant -scan is dominated by false positives — on Hemingway it returns 21 hits of which **18 are -ordinary text** (`God damn` for the surface `Goddamn` ×14, plus `I run`, `On an`, `Do me`, -`Si le`). The discriminator, with no dictionary: in a genuine split at least one FRAGMENT -is not a word of this corpus. `God` and `damn` occur constantly; `Primi`, `tivo`, `ell`, -`higurh`, `Toad` do not. That one test cleared all 18 and kept all 3 real ones. - -⚠ **My first negative control could not pass.** It planted the split nonce in its own probe -text and then asserted the nonce was absent — an alarm wired to itself, failing on every -run. It now hunts the split nonce in the *real* copies. A control that cannot pass is not a -control. - -## The shipped corpora, checked with the committed instrument - -Re-derived with the COMMITTED gate, not a scratch probe: - -``` -lv-bronte GATE PASSED 0 separator-split survivors (9.7 s) -lv-hemingway GATE FAILED Pasionaria, Primitivo, Chicote (34.6 s) - 1 occurrence each per copy, in all 6 copies — SHIPPED and LIVE -``` - -⚠ The first version of this scan was **too slow to run** on Hemingway — per-surface scanning -is O(surfaces x copies x corpus) and 881 surfaces x 10 copies was still going at 5 minutes -when it was killed. Rebuilt as one alternation pass, same trick `scan()` already used: 35 s, -identical verdict and identical hit counts on both McCarthy trees. **A gate too slow to run -is not a gate.** - -**Operator call outstanding** on whether 3 names in a 958k-word corpus warrant re-gating and -retraining a live adapter. Not acted on. - -## The chain was recovered, not remembered — and is now written down - -There was **no McCarthy runbook**, and the D1→D3 session issued its commands over -non-interactive ssh so no shell history survived. The chain was recovered by rebuilding -candidates and matching sha256 against the artifacts on disk, then pinned: - -``` -D1 build_corpus_mccarthy.py 6 works byte-identical -D2 entities.py --min-count 5 --fold-clitics --drop-acronyms --min-mid-ratio 0.2 --min-mid 2 -D2c apply_gender_overrides.py entities-final.json byte-identical -D3 rename.py --preset mccarthy --scope corpus --min-cap 5 --copies 6 --seed 4919 - --holdout-chapter 7 17 all 36 copy files byte-identical -``` - -⚠ `--min-mid-ratio` is what keeps `Yeah`/`Buenas`/`Shh`/`Sí` out of the map. The map is -**insensitive** to it: any value in [0.05, 0.3] with `--min-mid` 1 or 2 reproduces byte-for- -byte; `--min-mid 3` does not. The original values are unrecoverable and it does not matter — -which is worth saying, because an exact-looking recipe that was never pinned invites a -false claim of reproduction. Full recipe and every deviation: `scripts/mccarthy-corpus/RUNBOOK.md`. - -⚠ **D1 must run on nh3-dev** — the builder reads the kvasir catalogue by absolute path and -gx10 has no copy. The previous session's "on gx10" note is true of D2 onward only. - -⚠ **No phrase map exists for this corpus**, so the gate's phrase audit does not run at all. -Yarros and Brontë both had one. Not closed. - -Rollback: `~/lv-mccarthy/corpus-{clean,renamed}.pre-splitfix` on gx10. - -Related: [[2026-09-17-mccarthy-d1-d3]], [[2026-09-17-beat-contamination-leak]], -[[2026-09-17-lv-hemingway-gate]], [[2026-09-17-lv-bronte-gate]]. diff --git a/persistent-memory.d/2026-09-17-measured-and-deliberately-not-changed-three-of-them.md b/persistent-memory.d/2026-09-17-measured-and-deliberately-not-changed-three-of-them.md deleted file mode 100644 index be336d7..0000000 --- a/persistent-memory.d/2026-09-17-measured-and-deliberately-not-changed-three-of-them.md +++ /dev/null @@ -1,3 +0,0 @@ -# `[2026-09-17]` Measured and DELIBERATELY not changed, three of them. - -**Measured and DELIBERATELY not changed, three of them.** The oversize-passage drop is 13.9% of McCarthy's train words, between Hemingway's 10.0% and the shipped Brontë's 18.3% — in range, no change. `BEAT_PROMPT` asserts the passage is first-person and McCarthy is third; measured inert (**0** narrator-retries against Hemingway's 615 of 7,094), so the prompt was left alone. Blood Meridian's 131 dash-separated chapter-argument paragraphs DID warrant a change and `--drop-leading-heading` now eats them (0 in every other work of all three corpora). diff --git a/persistent-memory.d/2026-09-17-next-voice-seats.md b/persistent-memory.d/2026-09-17-next-voice-seats.md deleted file mode 100644 index 003da11..0000000 --- a/persistent-memory.d/2026-09-17-next-voice-seats.md +++ /dev/null @@ -1,69 +0,0 @@ -# `[2026-09-17]` Which voices earn a training seat next — measured against the catalogue, not chosen by taste - -Method: rank every author in the kvasir catalogue by **usable extracted** works, then apply the -selection criterion the lv-krakauer parking established — *does the author have a voice*, asked -before any corpus work, and specifically **does that voice live where the instrument looks**. -`voice_distance.py` is Burrows's Delta over CHARACTER BIGRAMS, so it sees function-word morphology, -punctuation and sentence rhythm. A writer whose distinction is plot, research or subject matter is -invisible to it — an adapter cannot carry that, and the gate cannot measure it. - -⚠ `triage.length` is in **CHARACTERS**, ~5.2 chars/word calibrated against builds we did ourselves -(The Crossing mobi 777,420 chars = our measured 149,985 words). Dedup by title taking the max across -formats, and floor at 100,000 chars — that is what excludes the `accepted`-but-truncated rows -(Blood Meridian epub at 6,031 chars beside the mobi's 623,849). - -## ⭐ The size ranking INVERTS the voice ranking at the top - -``` -Stephen King 76 works 12,133,529 w <- biggest, and NOT a candidate -Agatha Christie 72 5,451,377 <- second biggest, the Krakauer case exactly -Terry Pratchett 52 4,821,474 -Georgette Heyer 30 3,416,867 -Graham Greene 45 3,037,425 -William Faulkner 25 2,981,183 <- the pick -``` - -Christie is the whole lesson in one row: a superb writer whose genius is plot architecture, in prose -deliberately kept transparent. Nothing for a char-bigram Delta to grip. King is the softer version — -distinctive in pacing and brand-name texture, not in syntax. - -## The three that clear both bars - -**1. William Faulkner — 25 catalogue rows, ~15 pure novels, ~1.6M words.** -*The voice in one sentence:* sentences that defer their main clause through stacked subordination -and coined compounds until the reader is held inside a single unbroken perception. -About as char-bigram-legible as English gets — the voice IS the clause-joining morphology and the -`and`/`which`/`that` density. ⭐ **And he is McCarthy's stylistic ancestor, which is the real -argument:** the Brontë gate record states the frozen adjudication needs "a control-author panel (to -place an absolute band and a hard-negative sister)" and notes we have none. Faulkner beside McCarthy -makes each the other's hard negative — a METHOD upgrade, not just another roster entry. -⚠ Messiest corpus of the three: a 446k-word `Snopes: The Hamlet, The Town, The Mansion` omnibus -duplicates novels also present individually, and `Three Famous Short Novels` overlaps it again. That -is the Hemingway trap (169,759 words of measured 90-96% collection duplication) — containment pass -before anything else. - -**2. Toni Morrison — 13 rows, 11 novels after pruning, ~818k words.** -*The voice:* free-indirect discourse sliding between narrator and character mid-sentence, carried on -incantatory repetition and deliberate fragments. -Cleanest corpus shape on the list: **11 novels → 11 val units, beating Hemingway's 10.** Val units -scale with WORK COUNT, which is the structural reason Brontë's voice axis came back underpowered at -4 with no cheap fix. ⚠ Drop `Burn This Book` (anthology she edited) and `Playing in the Dark` -(criticism) — same reason Krakauer's reporting does not transfer. - -**3. Raymond Chandler — 9 rows, 7 novels + a 409k short-story omnibus, ~970k words.** -*The voice:* clipped first-person declaratives that periodically detonate into one baroque simile, -with dialogue carrying most of the scene. -Fills the register gap nobody else fills — **first-person hardboiled**; the line has no first-person -male narrator at all. Corpus is almost exactly Hemingway-sized (997k vs 958k), which was the -decisive gate. ⚠ Drop `Essays and Reviews` — non-fiction. - -## Held, and why - -**Conrad** (31 works, 2.4M) is a genuine tier-1.5 if a fourth is wanted. **Melville** (10, 1.9M) has -a superb voice but a mixed-register corpus — the cetology chapters are a different book from the -narrative. **Austen** (12, 1.18M) is worth noting because Burrows's Delta was developed on her, so -the instrument is known to resolve her. The romantasy cluster is a separate question entirely — -see [[2026-09-17-romantasy-register-measured]]. - -Related: [[2026-09-17-mccarthy-split-name-leak]], [[2026-09-17-lv-bronte-gate]], -[[2026-09-17-lv-hemingway-gate]]. diff --git a/persistent-memory.d/2026-09-17-romantasy-register-measured.md b/persistent-memory.d/2026-09-17-romantasy-register-measured.md deleted file mode 100644 index 519a91a..0000000 --- a/persistent-memory.d/2026-09-17-romantasy-register-measured.md +++ /dev/null @@ -1,71 +0,0 @@ -# `[2026-09-17]` Romantasy measured as a register — it is real, we already took its best voice, and the obvious next pick is its worst - -Prompted by the operator pushing back on a one-clause dismissal of the lane as "depth behind -Yarros". The dismissal was taste; this is a measurement, on the gate's own instrument. - -**Method.** Char-bigram Burrows's Delta, the same measure `voice_distance.py` gates on. ~120k words -per author, sampled from the MIDDLE quartile of each author's largest works (front and back matter -are not the voice), equalised so a bigger sample is not a different measurement. 400 most-frequent -bigrams as the feature set, z-scored over 4,000-word chunks pooled across all authors. - -**Controls first, because a between-author number without a within-author floor is unfalsifiable.** - -``` -A-vs-A floor (two halves of the SAME author) - Yarros 0.285 Maas 0.298 Armentrout 0.314 St. Clair 0.322 Cole 0.338 - Kenyon 0.375 Reyne 0.391 - McCarthy 0.303 Morrison 0.327 Brontë 0.209 Hemingway 0.454 <- worst, used as the bar - -positive controls (known-distinct pairs — the instrument must separate these) - Yarros vs McCarthy 0.862 1.9x - Hemingway vs Brontë 0.773 1.7x - McCarthy vs Morrison 0.675 1.5x - Hemingway vs McCarthy 0.655 1.4x - -romantasy, all 21 pairs median 0.537 1.2x floor (range 0.465 - 0.674) -``` - -**The register is real but tight.** 1.2x floor against controls at 1.4-1.9x. Only one pair falls to -1.0x, so it is not seven names for one voice. - -⚠ **Sensitivity floor, stated because a result without one is unfalsifiable.** The 0.454 bar is -Hemingway's, inflated by his own heterogeneous corpus (1920s-1960s, novels + stories + posthumous). -Against the romantasy authors' OWN floors (~0.34) the same pairs read ~1.6x — control-grade. The -truth sits between those readings and **this method cannot split it finer**. One sample per pair, no -repeat draws: read the rank ordering as indicative, do not read small gaps at all. - -## Two findings that survive either floor reading - -⭐ **Yarros is the cluster OUTLIER, not a typical member.** Four of the five largest distances in the -matrix involve her — Yarros-Kenyon 0.674, Yarros-St. Clair 0.644, Yarros-Reyne 0.637, Yarros-Maas -0.567. **We already trained the most distinctive romantasy voice we hold**, so a second seat in the -lane buys measurably less than the first did. That is the actual answer to "what about romantasy". - -⭐ **Maas is the centroid, so the obvious commercial pick is the least distinctive.** Maas-Reyne -0.465 and Maas-Cole 0.470 are the two SMALLEST distances in the whole matrix. She is the biggest -name available (922k words) and measurably the most generic of the seven in char-bigram terms. -Picking by sales rank picks the worst adapter. - -## If the lane gets a second seat it is Kenyon - -Furthest from the shipped Yarros (0.674), so it adds the most new signal — **and 27 works means 27 -val units, the best-powered gate the line could build** (Hemingway 10, McCarthy 6, Brontë 4, where 4 -is the documented structural cause of an underpowered voice axis with no cheap fix). - -⚠ Two costs: the 27 are one series (Dark-Hunter), so the shared proper-noun space makes -`--scope corpus` mandatory rather than optional; and a "Dark Hunter - The Dark Hunter Complete" -omnibus sits in the catalogue rows, so the containment pass runs first. - -**Corpus shapes for the lane** (works ≥100k chars, deduped by title): -``` -Sherrilyn Kenyon 27 2,368,396 w Scarlett St. Clair 11 1,133,067 -Opal Reyne 14 2,466,307 Kresley Cole 10 1,037,914 -Jennifer Armentrout 6 1,179,953 Sarah J. Maas 5 922,711 -Rebecca Yarros 5 820,425 <- SHIPPED on this -``` -⭐ Worth noting for any future bar-setting: **Yarros shipped on 5 works / 820k words.** The corpus -bar is lower than it looks. - -Instrument: `scratchpad/regdist.py` (screening tool, not the gate). - -Related: [[2026-09-17-next-voice-seats]], [[2026-09-16-lv-voices-line]], [[2026-09-17-lv-bronte-gate]]. diff --git a/persistent-memory.d/2026-09-17-servers-fv-ml1-ssh-target-was-bare-10-251-50-54-so-deploy.md b/persistent-memory.d/2026-09-17-servers-fv-ml1-ssh-target-was-bare-10-251-50-54-so-deploy.md deleted file mode 100644 index aea1dfe..0000000 --- a/persistent-memory.d/2026-09-17-servers-fv-ml1-ssh-target-was-bare-10-251-50-54-so-deploy.md +++ /dev/null @@ -1,3 +0,0 @@ -# `[2026-09-17]` `servers/fv-ml1/ssh-target` was bare `10.251.50.54`, so `deploy-stack.sh` connected as `lkraven` and could not write the infra-ops-owned `/opt/docker/compose/` - -**`servers/fv-ml1/ssh-target` was bare `10.251.50.54`, so `deploy-stack.sh` connected as `lkraven` and could not write the infra-ops-owned `/opt/docker/compose/`** — and lkraven's sudo on fv-ml1 needs a password, so `DEPLOY_SUDO=1` failed too. Now `infra-ops@10.251.50.54`; `--validate-only` still clean, deploy works. ⚠ Other hosts' `ssh-target` files may carry the same gap — a read-only refresh works as either user, so the fault only surfaces on a deploy. diff --git a/persistent-memory.d/2026-09-17-the-beat-contamination-leak-is-present-in-hemingway-70-of-7.md b/persistent-memory.d/2026-09-17-the-beat-contamination-leak-is-present-in-hemingway-70-of-7.md deleted file mode 100644 index b111d8f..0000000 --- a/persistent-memory.d/2026-09-17-the-beat-contamination-leak-is-present-in-hemingway-70-of-7.md +++ /dev/null @@ -1,3 +0,0 @@ -# `[2026-09-17]` The beat-contamination leak IS present in Hemingway — 70 of 7,094 train beats (0.96%), 0 of 200 val. - -⭐ **The beat-contamination leak IS present in Hemingway — 70 of 7,094 train beats (0.96%), 0 of 200 val.** `scripts/r49-corpus/audit_pairs_sourcenames.py` closes the blind spot `leak_gate.py` has by construction (it reads the corpus and the renamed copies, never the generated beats). Cross-validated on real data: the fixed Brontë pairs return 0 of 3,858 and `pairs-full.CONTAMINATED` returns 15 of 792 = 1.89% with the recorded names. `--filter-out` yields a verified-clean 7,024-pair set in one command; the retrain is the operator's call. **The val split being clean is why the gate could run at all.** diff --git a/persistent-memory.d/2026-09-17-the-mccarthy-register-names-the-punctuation-on-purpose-and.md b/persistent-memory.d/2026-09-17-the-mccarthy-register-names-the-punctuation-on-purpose-and.md deleted file mode 100644 index 9378643..0000000 --- a/persistent-memory.d/2026-09-17-the-mccarthy-register-names-the-punctuation-on-purpose-and.md +++ /dev/null @@ -1,3 +0,0 @@ -# `[2026-09-17]` The `mccarthy` register names the punctuation ON PURPOSE, and that is a gate-design call made before any McCarthy number existed. - -⭐ **The `mccarthy` register names the punctuation ON PURPOSE, and that is a gate-design call made before any McCarthy number existed.** `eval-*.sh` drives the base control arm with the SAME system prompt via `--system-from`, and `voice_distance.py` is Burrows's Delta over character bigrams — so a tic left OUT of the register is a cheap win only the adapter can take, on a corpus measuring 0.0 quote marks per 10k against Hemingway's 838. Stating them hands them to the control too. Cost stated up front: the voice axis gets harder, and McCarthy's 276-passage val split (against Brontë's 44) is why that trade is affordable here and was not there. diff --git a/persistent-memory.d/2026-09-17-the-shipped-lv-bronte-adapter-emits-mid-sentence-line.md b/persistent-memory.d/2026-09-17-the-shipped-lv-bronte-adapter-emits-mid-sentence-line.md deleted file mode 100644 index b9d84fd..0000000 --- a/persistent-memory.d/2026-09-17-the-shipped-lv-bronte-adapter-emits-mid-sentence-line.md +++ /dev/null @@ -1,3 +0,0 @@ -# `[2026-09-17]` The SHIPPED lv-bronte adapter emits mid-sentence line breaks at 12.46 per 1k chars, and nothing downstream looks for it. - -⭐⭐ **The SHIPPED lv-bronte adapter emits mid-sentence line breaks at 12.46 per 1k chars, and nothing downstream looks for it.** Its corpus is 100% hard-wrapped at ~68 chars (Gutenberg plain text) and the wrap transfers: base control 0.00, ckpt475 (shipped) 12.46, ckpt925 11.79, every Hemingway arm 0.00 on a 0%-wrapped corpus. Both controls fire. `score_beats.py` passed Brontë's damage axis anyway. McCarthy is the MIXED case — The Road wrapped, the other five works not — which is worse to learn than either pure one, so `build_sft_pairs.py --reflow-hard-wraps` (DEFECT 4) fixes it at pair time, off by default. ⚠ The obvious fix, joining every interior newline, CORRUPTS 46 two-speaker exchanges whose blank line was lost — and unmarked dialogue is the one thing this adapter exists to learn. The rule splits on sentence-final punctuation and takes the cheaper error deliberately. diff --git a/persistent-memory.d/2026-09-17-the-two-epoch-recipe-is-now-0-for-2-and-should-stop-being.md b/persistent-memory.d/2026-09-17-the-two-epoch-recipe-is-now-0-for-2-and-should-stop-being.md deleted file mode 100644 index 47bda6e..0000000 --- a/persistent-memory.d/2026-09-17-the-two-epoch-recipe-is-now-0-for-2-and-should-stop-being.md +++ /dev/null @@ -1,3 +0,0 @@ -# `[2026-09-17]` The two-epoch recipe is now 0 for 2 and should stop being carried forward. - -**The two-epoch recipe is now 0 for 2 and should stop being carried forward.** Hemingway's eval minimum is step 1750, but step 850 is +0.0040 against a 0.0044 median neighbour jitter — three checkpoints inside one jitter — and 850 won every resolving axis (2.3x tighter seed spread, lower memorisation, less ran-on). Same outcome as Brontë. What IS robust on this schedule is the epoch-3 collapse: +0.0762 = **17.4x jitter**. diff --git a/persistent-memory.d/2026-09-17-the-v2-voice-floor-is-now-pairwise-and-it-retroactively.md b/persistent-memory.d/2026-09-17-the-v2-voice-floor-is-now-pairwise-and-it-retroactively.md deleted file mode 100644 index b456dd1..0000000 --- a/persistent-memory.d/2026-09-17-the-v2-voice-floor-is-now-pairwise-and-it-retroactively.md +++ /dev/null @@ -1,3 +0,0 @@ -# `[2026-09-17]` The v2 voice floor is now PAIRWISE, and it retroactively passes lv-bronte. - -⭐⭐ **The v2 voice floor is now PAIRWISE, and it retroactively passes lv-bronte.** lv-bronte's ckpt475 shipped as a voice-axis FAILURE at +0.193 against a 0.251 floor contributed entirely by ckpt925 — a third arm nobody was shipping, on one outlier seed. Scored against the arm it was actually compared to the floor is 0.091 and it clears at **2.1x**. The rule was changed **prospectively**, pre-registered for lv-hemingway before any Hemingway number existed, on an argument independent of the answer: the sampling variability of a difference A−B depends on A and B, never on a third arm C. The previous session found the defect and deliberately declined to exploit it; this follows from fixing it. lv-hemingway passes under **both** rules, so its verdict does not lean on the change. Caveats amended append-only in the compose, the NFS README and the gate record. Commits `0bb4938` `2e9b118`. diff --git a/persistent-memory.d/2026-09-17-triagedisposition-accepted-in-the-kvasir-catalogue-does-not.md b/persistent-memory.d/2026-09-17-triagedisposition-accepted-in-the-kvasir-catalogue-does-not.md deleted file mode 100644 index 257bb59..0000000 --- a/persistent-memory.d/2026-09-17-triagedisposition-accepted-in-the-kvasir-catalogue-does-not.md +++ /dev/null @@ -1,3 +0,0 @@ -# `[2026-09-17]` `triage_disposition = 'accepted'` in the Kvasir catalogue does NOT mean the extraction succeeded. - -⚠ **`triage_disposition = 'accepted'` in the Kvasir catalogue does NOT mean the extraction succeeded.** Blood Meridian's epub row holds 1,167 words of a 117,000-word book, The Crossing's 222 of 150,000 — real prose, real titles, accepted. Faulkner's *The Mansion* is 39 words. `near_dup_pairs` holds ONE row in the entire 1,284-work library and is blind to a fragment beside its full sibling. **Word-count every master before trusting a row**, and note that word count alone cannot tell a truncated novel from a legitimately short work. diff --git a/persistent-memory.d/2026-09-30-parakeet-seat-switch-approved.md b/persistent-memory.d/2026-09-30-parakeet-seat-switch-approved.md index 26de617..ea90883 100644 --- a/persistent-memory.d/2026-09-30-parakeet-seat-switch-approved.md +++ b/persistent-memory.d/2026-09-30-parakeet-seat-switch-approved.md @@ -18,3 +18,9 @@ 2. Free ~+1.1–1.5 GB on GPU 0 by trimming the `vllm-gen-small` util. ⚠ MEASURE the resulting free memory; util does not predict resident VRAM. 3. Cut over with the old seat kept as the rollback. 4. Re-measure live on GPU 0. + +**DONE 2026-10-01 ~0126 PT (infra-hermes; `stacks/parakeet-nemo`, de6ea32), infra-ops audit PASSED 0137.** +- p50 on GPU 0: 33 / 36 / 42 / 71 ms, against 187 / 308 / 626. WER: clean 1.965, other 3.026. +- A 714 s file returns 200 in 0.85 s (360 s windows). The steady state is 3,582 MiB, and GPU 0 Free is 385. +- gen-small util 0.48 → 0.36: it took three boots and ~34 min of downtime at midnight, with zero LiteLLM errors. Its `.env` holds 0.33 for the next restart. The KV is byte-pinned and unchanged. +- infra-hermes caught that httptools 0.8.0 emits `HTTP/1.1 200\x00OK`, which LiteLLM rejects; it is pinned to `--http h11`. diff --git a/persistent-memory.md b/persistent-memory.md index 9153e70..d269bac 100644 --- a/persistent-memory.md +++ b/persistent-memory.md @@ -1,6 +1,6 @@ # Persistent memory — eshpfi-management -_Last updated: 2026-09-30 ~1800 PT (U11a legacy off on both Worldtree instances, U11b gate armed; SemIf → intern-decision with Jev /v1/systemone at 32k; Scriberr → GPU 3 with overlap slicer + gap retry; Parakeet seat switch to unified-en APPROVED, next session; 26 old entries archived.)_ +_Last updated: 2026-10-01 ~0420 PT (Parakeet seat → unified-en under NeMo LIVE + audited; gen-small util 0.36 → .env 0.33; leftover bench weights + spike dirs deleted; Scriberr no-upstream; eshpfi + worldtree-instance-configs pushed. Prior: U11a off, SemIf → intern-decision, Scriberr GPU 3 + patches.)_ > **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its > `Written:` stamp is under **8 hours** old, read it (it carries the in-flight @@ -115,53 +115,44 @@ no longer deployed sidecars here. See Recent decisions.) ## Current state / in-flight -_As of 2026-09-30 ~1800 PT._ +_As of 2026-10-01 ~0420 PT._ -### Parakeet speech seat switched to unified-en under NeMo (Prime 2026-09-30 ~1755: "reasonable terms, ship the switch") +### Parakeet speech seat: unified-en under NeMo, LIVE (2026-10-01) -- **DONE: LIVE since ~0126 PT 2026-10-01** as `parakeet-nemo` (`stacks/parakeet-nemo`, image `local/parakeet-nemo:nemo-0.1.0`, infra-hermes de6ea32 + 41d2014) on :8300; LiteLLM untouched. **infra-ops AUDIT PASSED 0137.** - - Latency on GPU 0, p50 for 1–3 / 3–8 / 8–20 / 20–60 s: 33 / 36 / 42 / 71 ms, against 187 / 308 / 626 ms before. Through LiteLLM a 2.5 s clip takes 80–98 ms. +- **⚠ INCIDENT 04:21 PT 2026-10-01, MITIGATED, root fix in flight:** + - `vllm-gen-small`'s EngineCore CUDA-OOM'd when it needed a 394 MiB runtime workspace and GPU 0 had 388 MiB Free. The parakeet seat was parked at its 3,582 MiB window cache. + - vLLM grows ~0.8 GB at runtime beyond its preallocation; my audit checked gen-small's boot margin, not its runtime growth. gen-small auto-restarted, healthy at 04:23. + - I restarted parakeet-nemo to drop its cache (rest 2,084 MiB). gen-small then answered 3/3 via LiteLLM, and GPU 0 Free is ~1,075 MiB. + - **infra-hermes is tasked with nemo-0.1.1:** `empty_cache` after windowed requests, plus a hard memory ceiling so the seat 503s instead of starving gen-small. It also measures gen-small's runtime growth. + - **Awaiting Prime:** trim gen-small's KV pin (8 → 7 GiB frees ~1 GiB; 670k → ~586k tokens), or move the seat to GPU 3. + - Until fixed, a long transcription can re-grow the seat's cache and starve gen-small. +- **LIVE since ~0126 PT 2026-10-01** as `parakeet-nemo` (`stacks/parakeet-nemo`, `local/parakeet-nemo:nemo-0.1.0`, built by infra-hermes) on fv-ml1 GPU 0 :8300, with LiteLLM `ext-stt`/`whisper-1` unchanged. **infra-ops audit PASSED 0137.** + - p50 on GPU 0 for 1–3 / 3–8 / 8–20 / 20–60 s: 33 / 36 / 42 / 71 ms, against 187 / 308 / 626 ms for the old seat. - WER: LibriSpeech clean 1.965, other 3.026. - - Fixed: a 714–726 s file returns 200 (long files go in 360 s windows, because NeMo builds the full T×T attention mask even under local attention); no drop after a pause. - - Rollback: `docker stop parakeet-nemo && docker start parakeet` (the old container is stopped, not removed). - - **gen-small: util 0.48 → 0.36** (0.46 and 0.40 failed its boot check). Its KV is byte-pinned (`--kv-cache-memory 8 GiB`), still 670,142 tokens / 2.56×, so there was no KV cost. It was down ~34 min (0012–0046 PDT) while the util was iterated; zero LiteLLM errors. - - ⚠ **Steady state is 3,582 MiB for the seat (its cached window peak); GPU 0 Free is 385 MiB.** That leaves gen-small's restart boot-check margin at only ~0.45 GiB. infra-hermes was asked to set gen-small util to 0.33 in the .env WITHOUT restarting, so it applies at the next restart. **Nothing else fits on GPU 0.** - - Seat invariants are in its README: cast to bf16 AFTER change_attention_model; uvicorn pinned with `--http h11`, because httptools 0.8.0 writes `HTTP/1.1 200\x00OK` and LiteLLM/httpx rejects it; the 360 s window. -- The original brief, for reference: - -- **The task:** replace the live speech seat with `nvidia/parakeet-unified-en-0.6b` under NeMo 3.0.0 with bf16 weights. The **NVIDIA Open Model License is ACCEPTED** for internal use. - - The live seat today: container `parakeet` on fv-ml1 GPU 0, port :8300, sherpa-onnx int8 `parakeet-tdt-0.6b-v3`, reached through LiteLLM as `ext-stt` and `whisper-1`; its caller is `talk`. - - Evidence, `docs/pfi/parakeet-seat-ab-2026-09-30.md` (a6c1d3c): end-to-end p50 for 1–3 / 3–8 / 8–20 s clips goes 144 / 260 / 565 → 23 / 27 / 33 ms, and WER is lower on every set. -- **Kit:** the wrapper `services/parakeet-ab-2026-09-30/code/serve_nemo.py` keeps the seat's endpoints and text, and matched NeMo's own transcribe on 400/400. The weights are pinned on fv-ml1 in `/tank/aimodels/huggingface` (rev `fe53cd88`). A working NeMo 3.0.0 env for reference is under `/tank/spikes/parakeet-ab`. **No image is built yet.** -- **What the image needs:** - - a warm-up at the longest served length; - - a bf16 cast BEFORE `.to(cuda)`, which avoids a +1.5 GB load spike; - - local attention for long files (a 30-min file took 2.6 s in one request). -- **Room:** it needs about +1.1 GB while serving (+1.5 GB at load) over the seat's 1,690 MiB, and GPU 0 has ~100 MiB free. - - The plan is to trim `vllm-gen-small` `--gpu-memory-utilization` from 0.48 to about 0.46 at a quiet moment (a 2–3 min restart). - - ⚠ The util value does NOT predict resident VRAM: on 09-15 gen-small at 0.48 held 36,942 MiB and cyberprev at 0.40 held 47,124. MEASURE nvidia-smi Free after the change; do not compute it. My 09-30 "0.01 ≈ 0.95 GB" estimate is unverified. - - GPU 1's ~6.6 GB free is intern-decision's 32k headroom, so it is not available. -- **Cut-over:** keep the old seat as the rollback, and leave LiteLLM alone unless the port changes. Re-measure live on GPU 0: latency per length bin against the old seat, a WER spot-check, and memory. -- **Live-seat defects until then:** - - HTTP 500 above ~400 s of audio; - - long-form dropouts; - - after a 1.5 s digital-silence pause it can drop the rest of the utterance (6 of 40). + - Files longer than 6 min run in 360 s windows. That avoids NeMo's T×T attention mask; a seam can lose a space or a word. +- **Rollback:** `docker stop parakeet-nemo && docker start parakeet`. The old container and image are kept. +- **GPU 0 is FULL:** + - The seat's steady state is **3,582 MiB** (its cached window peak); Free is **385 MiB**. + - `vllm-gen-small` runs at util 0.36, and its `.env` holds **0.33** for the next restart (~3 GiB of boot-check margin). Its KV is byte-pinned: 670,142 tokens / 2.56×. + - Before restarting any vLLM seat on this card, check that util × 95.6 GiB ≤ measured Free + the seat's own resident memory. Do not trial-boot. A trial-boot sequence took gen-small down for 34 min on 2026-10-01. +- Seat invariants (in its README): cast to bf16 AFTER change_attention_model; uvicorn pinned with `--http h11` (httptools 0.8.0 emits `HTTP/1.1 200\x00OK`, which LiteLLM/httpx rejects). +- NVIDIA Open Model License accepted for internal use. → `persistent-memory.d/2026-09-30-parakeet-seat-switch-approved.md` ### Worldtree U11 memory cutover (demo + personal) -- **Legacy plane OFF since 0115/0120 PT 2026-09-30** (config repo 63cf268; personal /metrics in b6fdd81). → `persistent-memory.d/2026-09-30-worldtree-u11a-off-u11b-gate.md` -- **Daily gate batches:** infra-hermes runs them from 2026-10-01 with `scripts/wt-memory-gate-batch` and copies me on every verdict. The count is **1 of 3** consecutive PASS at off (20260930T090608Z); a FAIL restarts it. +- **Legacy plane OFF since 0115/0120 PT 2026-09-30** (config repo 63cf268; personal /metrics b6fdd81; the repo is pushed). → `persistent-memory.d/2026-09-30-worldtree-u11a-off-u11b-gate.md` +- **Daily gate batches:** infra-hermes runs `scripts/wt-memory-gate-batch` from 2026-10-01 and copies me on every verdict. The count is **1 of 3** consecutive PASS at off (20260930T090608Z); a FAIL restarts it. - **⚠ U11b STEP 5 IS MINE, triggered by the 3rd consecutive PASS:** 1. Run `docker exec -i python - < scripts/wt-h2-count.py` VERBATIM, right before each instance's deletion. Exit 2 means STOP and send worldtree-dev the output. 2. Delete LIVE, with the api running, using literal paths only: `agents/{forseti,lofn,mimir}/memory/.chroma` and `memory/context_promotion`, on BOTH instances. 3. Send worldtree-dev the stamp; b193 ships after it. - - A b192 restart re-creates an empty schema-only `ledger.db`. That is residue, not memory data: say so in the stamp, and remove it after b193. -- **Legacy archive:** DESTROY it whole by 2026-10-30, or at retirement-done, or on any subject-erasure request, whichever comes first. The runbook is in the detail file. -- **TODO:** re-sweep both api logs after real traffic, grepping `legacy executor|Traceback|ERROR|will not be remembered`. After the b193 push, remove the retired config keys at my pace. + - A b192 restart re-creates an empty schema-only `ledger.db`. That is residue: say so in the stamp and remove it after b193. +- **Legacy archive:** DESTROY it whole by **2026-10-30**, or at retirement-done, or on a subject-erasure request, whichever comes first. The runbook is in the detail file. +- **TODO:** re-sweep both api logs after real traffic. After the b193 push, remove the retired config keys. -### fv-ml1 GPU layout (as of 2026-09-30) +### fv-ml1 GPU layout (as of 2026-10-01) -- **GPU 0:** cyberprev (47.1 GB), gen-small (37.5 GB), voices (10.8 GB), the parakeet seat (1.7 GB); ~100 MiB free. +- **GPU 0:** cyberprev (47.1 GB), gen-small (35.3 GB), voices (10.8 GB), parakeet-nemo (3.6 GB steady). Free 385 MiB, FULL. - **GPU 1:** vllm-coder, erp-seat, meromero-rp, plus intern-decision (cap 14.4 GiB, 32k tokens, peak 15,220 of a 15,437 MiB budget). FULL. - **GPU 3:** the full-size-seat reserve (Flash-Next is parked). On-demand tenants: Blender, and Scriberr (0 idle, ~5.5 GB per job). When a full-size seat claims GPU 3, Scriberr steps aside to **irv-ml1's A6000**, not back to GPU 1. @@ -172,17 +163,7 @@ _As of 2026-09-30 ~1800 PT._ ### Scriberr (fv-ml1 GPU 3) -- **LIVE `scriberr:local-blackwell-a353078-dropout2`:** upstream a353078 plus patch 0001 (overlap slicer) and patch 0002 (gap retry, `PARAKEET_MODEL_PATH`). v3 stays (Prime: no NeMo 3.0.0 surgery). → `persistent-memory.d/2026-09-30-scriberr-slicer-gap-retry-gpu3.md` -- **Awaiting Prime:** - - Open the slicer upstream PR, and choose which GitHub account (`stacks/scriberr/patches/upstream-pr/`). - - Delete the leftovers: - - the candidate weights in `/tank/aimodels/huggingface`, EXCEPT unified-en, which the seat switch needs; - - `/tank/spikes/scriberr-slicer`, including `private/`, which holds Prime's recordings (mode 700); - - `/tank/spikes/parakeet-ab` (~25 GB), but only after the switch. - -### irv-ml1 /storetank: CLOSED (2026-10-01) - -- Prime ruled, via comfy-dev: "delete unused weights + old staged files". comfy-dev executed it himself, 84 files, logged at irv-ml1 `~/3d-dl/deleted-2026-10-01.log`. Free went from 260 to 299 GB (86% → 84%). Tiers B/C/D got no ruling. infra-hermes closed it (thread `01M3TCSYRSFNAPA9BPQTMFQ6KJ`). +- **LIVE `scriberr:local-blackwell-a353078-dropout2`:** upstream a353078 plus patch 0001 (overlap slicer) and patch 0002 (gap retry, `PARAKEET_MODEL_PATH`), carried LOCALLY ONLY (Prime 2026-10-01: no upstream). v3 stays. `scripts/scriberr-rebuild` re-applies both. → `persistent-memory.d/2026-09-30-scriberr-slicer-gap-retry-gpu3.md` ### nh3-pve + nh3-ml1: post-visit, all live (2026-09-25/26) @@ -305,7 +286,7 @@ _As of 2026-09-30 ~1800 PT._ ### Live threads -- git: origin/main is at `128d1d8`, pushed 2026-09-30 1047 by someone other than infra-ops (presumably Prime). Local is ahead with unpushed commits, this snapshot included. `worldtree-instance-configs` has 3 unpushed commits (0a1387e, 63cf268, b6fdd81). Pushing is Prime's call. ⚠ The working tree AND index are shared with infra-hermes and subagents: commit with `git commit -- ` (auto-memory `feedback_shared_git_index_commit_pathspecs`). `graphify-out/GRAPH_REPORT.md` stays modified and uncommitted on purpose: it is auto-regenerated. +- git: **eshpfi-management pushed to `6b66207` and worldtree-instance-configs to `b6fdd81` (2026-10-01 ~0418, Prime's go).** Anything after that is unpushed. ⚠ The working tree AND index are shared with infra-hermes and subagents: commit with `git commit -- ` (auto-memory `feedback_shared_git_index_commit_pathspecs`). `graphify-out/GRAPH_REPORT.md` stays modified and uncommitted on purpose: it is auto-regenerated. - nh3-dev root disk was cleaned 2026-09-30 1704 (uv prune, dangling images, old build cache): 86% → 82%. The Beszel 85% alert flaps near the line. - Booth submit-all fix (Prime's report) is LIVE since 2026-09-27 1705, via booth-dev (booth `50bfc7b`). Pushing it is booth's call, per Prime; it is not ours. @@ -313,6 +294,8 @@ _As of 2026-09-30 ~1800 PT._ ## Recent decisions +- `[2026-10-01]` **Prime: delete the bench leftovers, no upstream for Scriberr, push.** DONE: 7 HF revisions deleted through huggingface_hub's cache API (25.1 GB; the parakeet 1.1B/ctc/v2 models and whisper-large-v3; **unified-en KEPT, the live seat mounts it**), plus `/tank/spikes/scriberr-slicer` (including the private copies of Prime's recordings) and `/tank/spikes/parakeet-ab`. The Scriberr upstream PR text was dropped (6b66207). Both repos pushed. +- `[2026-10-01]` **irv-ml1 /storetank reclaim done:** Prime ruled through comfy-dev, which deleted 84 files of its own (260 → 299 GB free). Tiers B/C/D got no ruling (infra-hermes thread `01M3TCSYRSFNAPA9BPQTMFQ6KJ`). - `[2026-09-30]` **Parakeet speech seat → `parakeet-unified-en-0.6b` under NeMo (bf16) APPROVED by Prime, NVIDIA Open Model License accepted. DONE 2026-10-01 0126 by infra-hermes; infra-ops audit passed 0137.** → `persistent-memory.d/2026-09-30-parakeet-seat-switch-approved.md` - `[2026-09-30]` **Worldtree U11a: legacy memory plane OFF on demo and personal. The U11b data deletion is gated on 3 consecutive PASS and step 5 is mine; the legacy archive must be destroyed by 2026-10-30.** → `persistent-memory.d/2026-09-30-worldtree-u11a-off-u11b-gate.md` - `[2026-09-30]` **SemIf replaced by intern-decision (Intern-Decision-4B, the Jev bench pick): semif-compatible plus Jev `/v1/systemone` at 32k tokens on GPU 1, with a Triton warm-up cache volume.** → `persistent-memory.d/2026-09-30-semif-replaced-by-intern-decision.md` @@ -432,49 +415,15 @@ _As of 2026-09-30 ~1800 PT._ - `[2026-09-18]` **Miranda's Hermes plugin install is now a SYMLINK to the svos repo, not a copy** — (operator-approved). → `persistent-memory.d/2026-09-18-miranda-s-hermes-plugin-install-is-now-a-symlink-to-the.md` - `[2026-09-18]` **Worldtree's `env.sh` secrets are vaulted** — 10 entries under `worldtree/` (gitea, matrix as/hs, openai, uv-index, vastblueai, wt-admin demo+personal… → `persistent-memory.d/2026-09-18-worldtree-s-env-sh-secrets-are-vaulted.md` -- `[2026-09-17]` ⭐⭐ **The next voice seat was MEASURED, not chosen by taste — and the corpus size ranking INVERTS the voice ranking at the top.** Our two largest authors are Stephen King (76 works, 12.1M words) and Agatha Christie (72, 5.5M); neither should get a seat, Christie being the Krakauer failure mode exactly (genius in plot architecture, prose deliberately transparent, invisible to a char-bigram Delta). Picks, in order: **Faulkner** (~15 pure novels, ~1.6M words — highest voice signal in the catalogue, AND he is McCarthy's stylistic ancestor, so training him next supplies the **hard-negative sister the gate has lacked since the Brontë record named it missing**); **Morrison** (11 novels after pruning criticism/anthology, ~818k — 11 val units, beating Hemingway's 10); **Chandler** (7 novels + a 409k short-story omnibus — fills the first-person hardboiled gap, Hemingway-class corpus size). ⚠ Faulkner's catalogue rows carry a 446k-word Snopes omnibus that duplicates novels also present individually — the Hemingway 90-96% collection-duplication trap, needs the containment pass first. → `persistent-memory.d/2026-09-17-next-voice-seats.md` -- `[2026-09-17]` ⭐⭐ **Romantasy IS a real register, we already trained its most distinctive member, and the obvious next pick is its worst.** Measured on the gate's own instrument (char-bigram Burrows's Delta, ~120k words/author from mid-work), with within-author floors and cross-genre positive controls. Cluster median pair **0.537 = 1.2x the worst floor** against controls at 1.4-1.9x — tighter than cross-genre but NOT collapsed. Two findings survive either floor reading: **Yarros is the cluster OUTLIER** (4 of the 5 largest pair distances involve her), so a second romantasy seat buys measurably less than the first did; and **Maas is the centroid** (the two smallest distances in the matrix are hers), so the obvious commercial pick is the least distinctive. If the lane gets a seat it is **Kenyon** — furthest from Yarros at 0.674 and **27 works = 27 val units, the best-powered gate the line could build** (Hemingway 10, McCarthy 6, Brontë 4). ⚠ Sensitivity floor stated: one sample per pair, no repeat draws; the rank ordering is indicative, fine gaps are not resolvable. → `persistent-memory.d/2026-09-17-romantasy-register-measured.md` - `[2026-09-17]` **`dragonfireacoustics.com` expires 2026-10-30 — six weeks — at eNom with NO transfer lock, and its sibling domain was already lost exactly this way.** → `persistent-memory.d/2026-09-17-dragonfireacoustics-com-expires-2026-10-30-six-weeks-at.md` ⚠ **If it is ever transferred, DNS does NOT come with the registration** — the nameservers are eNom's `name-services.com` and the zone must be recreated first or mail dies. The whole zone is two facts plus a landmine: `*` (WILDCARD) → 199.250.192.76 which is **dead** (no HTTP at all, and it is what the apex/mail/webmail/admin/ftp all answer with), `www` → 38.120.12.45 (us), and **7 Google Workspace MX records that must not be lost**. No DNSSEC (`delegationSigned: false`), so no transfer complication. ⚠ Also found: **no SPF and no DMARC** at the apex on a Google Workspace domain — a live deliverability problem independent of everything else. Transfer gate is the **TAC/EPP code from the eNom account**, not the lock; the missing lock is not authorization. 60-day rule is satisfied (last changed 2025-10-24). - `[2026-09-17]` **`dragonfireacoustics.com` IS configured on `pfi-ana-webhost`, and the whole thing is dead — a forgotten public-facing VM.** → `persistent-memory.d/2026-09-17-dragonfireacoustics-com-is-configured-on-pfi-ana-webhost.md` -- `[2026-09-17]` **headscale now split-DNSes `nh3.phasefinal.com` to the three AdGuards, so mesh clients can resolve the internal-only wildcard** → `persistent-memory.d/2026-09-17-headscale-now-split-dnses-nh3-phasefinal-com-to-the-three.md` -- `[2026-09-17]` **ESH is back on the Cityside static `128.177.138.182/30` and the site is healthy — confirmed on four axes, not one.** UDM WAN1 `wan_type` is `static` again (switched back from the DHCP set during the 09-17 outage), `stat/health` names Cityside Fiber with 0 disconnected and Verizon-5G idle at failover priority 2, esh-docker-vm's egress EQUALS the WAN ip so nothing is behind CGNAT, and colo→ESH reads **5.0 ms / 0% loss** at 2005/2142 Mbps (Cityside CGNAT was 9 ms, Verizon failover 33–37 ms). ⭐ The FortiGate `infra-ops` trusthost3 pin un-broke itself and that was VERIFIED: from ESH, ana-gw tcp/22 is open and offers a password prompt, which a trusthost mismatch would never do. ⚠ The two 7-day crowdsec entries are being left to expire 2026-09-23 on purpose — Cityside failed twice in six hours, so they are cheap insurance. → `persistent-memory.d/2026-09-17-esh-fiber-outages.md` -- `[2026-09-17]` **Operator ruled "leave it" on lv-hemingway's 3 separator-hidden names.** So `leak_gate.py` exits 1 on a SHIPPED tree by design; a future session seeing that red result should read this line, not start fixing. lv-bronte re-ran clean. -- `[2026-09-17]` ⭐⭐⭐ **The leak gate PASSED lv-mccarthy while five protagonist names sat in all six copies, and the blind spot generalises to every corpus in the line.** `\b(Surface)\b` cannot match a name with a character inserted in it, so a mangled occurrence is unrenameable AND unreportable: `B ell`, `C higurh`, `M oss`, `T oadvine` (a small-caps drop cap kept as its own token) and `Toad-vine`, `Glan-ton` (a print line-break hyphen). Every VISIBLE occurrence had been renamed, which is what made it invisible. Same family as lv-bronte's `_Antigua_`, now generalised: **any separator inside a name blinds a word-boundary scan.** Fixed in the corpus builder (rules 4+5, counted), and `leak_gate.py` now runs a separator-tolerant pass with its own controls that FAILS the gate — validated against the pre-fix tree. ⚠ Its fragment filter is load-bearing: a naive scan returns 18 false positives on Hemingway (`God damn`, `I run`) against 3 real. Whole D1→D3 chain reproduced byte-identically before and after. Commit `c559664`. → `persistent-memory.d/2026-09-17-mccarthy-split-name-leak.md` -- `[2026-09-17]` **The SHIPPED lv-bronte adapter emits mid-sentence line breaks at 12.46 per 1k chars, and nothing downstream looks for it.** → `persistent-memory.d/2026-09-17-the-shipped-lv-bronte-adapter-emits-mid-sentence-line.md` -- `[2026-09-17]` **The `mccarthy` register names the punctuation ON PURPOSE, and that is a gate-design call made before any McCarthy number existed.** → `persistent-memory.d/2026-09-17-the-mccarthy-register-names-the-punctuation-on-purpose-and.md` -- `[2026-09-17]` **lv-mccarthy's D1→D3 chain was RECOVERED, not remembered — there was no runbook and the commands went over non-interactive ssh, so no history survived.** → `persistent-memory.d/2026-09-17-lv-mccarthy-s-d1d3-chain-was-recovered-not-remembered-there.md` -- `[2026-09-17]` **Measured and DELIBERATELY not changed, three of them.** — The oversize-passage drop is 13.9% of McCarthy's train words, between Hemingway's 10.0% and the shipped… → `persistent-memory.d/2026-09-17-measured-and-deliberately-not-changed-three-of-them.md` -- `[2026-09-17]` ⭐⭐ **A unit splitter must choose by SIZE, not count — and the val split scales with WORK COUNT, not corpus size.** `scripts/r49-corpus/split_units.py` + a multi-index `--holdout-chapter`. The inherited most-units rule gave Cities of the Plain 4 units of 22,312w (the book's PARTS); the single-index holdout would have given McCarthy a Brontë-class 18k-word val reference on a 588k corpus. Both fixed, both caught by controls. → `persistent-memory.d/2026-09-17-mccarthy-d1-d3.md` -- `[2026-09-17]` ⭐ **lv-mccarthy D1–D3 complete on gx10, leak gate PASSED (0 of 75 renameable, 0 of 37 sub-threshold, both controls green).** Three McCarthy-specific calls, each forced by a measurement: corpus-scoped rename (the Border Trilogy shares 9 surfaces across books), a new `mccarthy` name preset (Hemingway's carries it_IT/fr_FR and McCarthy writes neither), and `--min-cap 5` to match the entity map's floor — the first gate run failed with 45 survivors purely because rename's floor was 8 and the map's was 5. → `persistent-memory.d/2026-09-17-mccarthy-d1-d3.md` - `[2026-09-17]` **PARKED lv-krakauer, and the reason is a selection criterion the line was missing: ask whether the author HAS a voice before investigating whether the…** → `persistent-memory.d/2026-09-17-parked-lv-krakauer-and-the-reason-is-a-selection-criterion.md` - `[2026-09-17]` **The althing route-declaring SessionStart hook is documented but NOT installed on nh3-dev** — `dev_launch.py` has zero occurrences of "route", no hook declares one, and every live route was hand-declared… → `persistent-memory.d/2026-09-17-the-althing-route-declaring-sessionstart-hook-is-documented.md` -- `[2026-09-17]` **Hemingway ships as-is: operator ruled "ship stands" on both measured corpus defects** — the 0.96% beat contamination and the 130 non-name entity-map surfaces. → `persistent-memory.d/2026-09-17-hemingway-ships-as-is-operator-ruled-ship-stands-on-both.md` - `[2026-09-17]` **PARKED lv-krakauer, and the reason is a selection criterion the line was missing: ask whether the author HAS a voice before investigating whether the…** → `persistent-memory.d/2026-09-17-parked-lv-krakauer-and-the-reason-is-a-selection-criterion-2.md` -- `[2026-09-17]` **A unit splitter must choose by SIZE, not by count — the inherited rule silently produced 22,000-word "chapters".** → `persistent-memory.d/2026-09-17-a-unit-splitter-must-choose-by-size-not-by-count-the.md` -- `[2026-09-17]` **lv-mccarthy D1 built — 167 units, 584,756 words — and the whole job was protecting a style that reads as damage.** → `persistent-memory.d/2026-09-17-lv-mccarthy-d1-built-167-units-584-756-words-and-the-whole.md` -- `[2026-09-17]` **lv-krakauer D1 built — 126 units, 422,880 words — and its name guard caught three defects nothing else would have reported.** → `persistent-memory.d/2026-09-17-lv-krakauer-d1-built-126-units-422-880-words-and-its-name.md` -- `[2026-09-17]` **`triage_disposition = 'accepted'` in the Kvasir catalogue does NOT mean the extraction succeeded.** — Blood Meridian's epub row holds 1,167 words of a 117,000-word book, The Crossing's 222 of 150,000 — real… → `persistent-memory.d/2026-09-17-triagedisposition-accepted-in-the-kvasir-catalogue-does-not.md` - -- `[2026-09-17]` ⭐⭐⭐ **lv-hemingway SHIPPED (ckpt850) with the line's strongest voice result — and the memorisation control it passed turned out to be the WRONG control.** Voice +0.413 delta_cb at 6.4x the floor, closing 73.8% of the achievable span (lv-bronte closed 48%). ⚠ `memorization_check.py` uses the base-unadapted arm as its negative control, but base writes 18,035 words of summary against the adapted arms' 27,413 of pastiche — **text that does not imitate the register cannot collide with its n-grams**, so a 0.00 there means "different register", not "did not memorise". The right reference is the author himself: **held-out Hemingway against the train split collides at 0.01 while the adapter does at 0.07**, so the comfortable "his plain register makes collisions inevitable" story is FALSE and was refuted rather than assumed. All 19 matched runs were READ: stock dialogue, max **9 words**, no proper noun — shorter than the 10-word run unseen Hemingway shares with the train split by coincidence. ⭐ **A negative control that differs from the candidate in a way correlated with the metric is not a control.** → `persistent-memory.d/2026-09-17-lv-hemingway-gate.md` -- `[2026-09-17]` **The v2 voice floor is now PAIRWISE, and it retroactively passes lv-bronte.** — lv-bronte's ckpt475 shipped as a voice-axis FAILURE at +0.193 against a 0.251 floor contributed entirely by… → `persistent-memory.d/2026-09-17-the-v2-voice-floor-is-now-pairwise-and-it-retroactively.md` -- `[2026-09-17]` **The beat-contamination leak IS present in Hemingway — 70 of 7,094 train beats (0.96%), 0 of 200 val.** → `persistent-memory.d/2026-09-17-the-beat-contamination-leak-is-present-in-hemingway-70-of-7.md` -- `[2026-09-17]` **`audit_entity_map.py` — the rename can DAMAGE the prose and no gate will ever say so.** — Mirror of `audit_stoplist.py`: surfaces wrongly held IN the map rather than out of it. → `persistent-memory.d/2026-09-17-auditentitymap-py-the-rename-can-damage-the-prose-and-no.md` -- `[2026-09-17]` **The two-epoch recipe is now 0 for 2 and should stop being carried forward.** — Hemingway's eval minimum is step 1750, but step 850 is +0.0040 against a 0.0044 median neighbour jitter … → `persistent-memory.d/2026-09-17-the-two-epoch-recipe-is-now-0-for-2-and-should-stop-being.md` -- `[2026-09-17]` **gitea was reaching the PUBLIC route from every repo on nh3-dev** — brokkr-smithy, sleipnir, Galdrabok, kvasir — and brokkr-smithy is pushed several times a week, so the… → `persistent-memory.d/2026-09-17-gitea-was-reaching-the-public-route-from-every-repo-on-nh3.md` -- `[2026-09-17]` **`servers/fv-ml1/ssh-target` was bare `10.251.50.54`, so `deploy-stack.sh` connected as `lkraven` and could not write the infra-ops-owned…** → `persistent-memory.d/2026-09-17-servers-fv-ml1-ssh-target-was-bare-10-251-50-54-so-deploy.md` - -- `[2026-09-17]` ⭐⭐ **A leak class the corpus gate structurally CANNOT see: the beat-writing model recognises the book and restores the author's real character names.** 1.8% of Brontë beats named Rochester/Jane/Brocklehurst while 0 responses did. Worst for public-domain classics; Yarros and Hemingway's clean runs are NOT evidence they are immune. Patched as a `sourcename` reject + `--source-entities`. → `persistent-memory.d/2026-09-17-beat-contamination-leak.md` -- `[2026-09-17]` **A stoplist entry is an assertion the leak gate can no longer check** — stoplisting removes a surface from the entity map, so a wrongly stoplisted CHARACTER is an undetectable leak. → `persistent-memory.d/2026-09-17-a-stoplist-entry-is-an-assertion-the-leak-gate-can-no.md` -- `[2026-09-17]` **ESH: Cityside Fiber failed TWICE (19:09 and ~01:06); operator switched WAN1 to DHCP to restore service and has a ticket for the static.** crowdsec `esh` allowlist carries both failover egresses with 7-day expiries — the rotation-fragility is live. → `persistent-memory.d/2026-09-17-esh-fiber-outages.md` -- `[2026-09-17]` ⭐⭐ **lv-bronte SHIPPED on voices-seat (ckpt475) DESPITE failing the v2 VOICE axis — additive, reversible, safety-axis clean.** Both candidates closed 48–52% of the achievable distance to Brontë but +0.193/+0.210 sit under a 0.251 noise floor set by ONE outlier seed in the arm not being shipped; cause is structural (81 val pairs vs Hemingway's 200) and not cheaply fixable. ckpt475 is the pick if it ships. The two-epoch recipe did NOT transfer. → `persistent-memory.d/2026-09-17-lv-bronte-gate.md` -- `[2026-09-16]` ⭐⭐ **Instruction-pair SFT BEATS raw-text for author voice, and the raw-text incumbent never cleared its own control (+0.141 against a 0.153 floor).** `lv-yarros` shipped; the v1 decision rule was amended by the operator after measurement showed it gated on axes the unadapted carrier already maxes. → `persistent-memory.d/2026-09-16-lv-voices-line.md` -- `[2026-09-16]` ⭐ **voices-seat live: one carrier, N `lv-` LoRA adapters, hot-swap measured at 0.24 s.** LoRA costs 24.3% of decode against a 0.1% A-vs-A floor and is worth paying; `--gpu-memory-utilization` is a request against TOTAL VRAM and only a pinned KV makes it predictive. → `persistent-memory.d/2026-09-16-voices-seat-lora.md` -- `[2026-09-16]` ⭐ **lv-hemingway corpus gated at 994,760 words — and half the catalogue had to be EXCLUDED.** 169,759 words of measured 90–96% collection duplication, a Sherwood Anderson parody, and the author's own name 95 times in publisher back matter; the gender resolver needed a corpus base-rate correction to stop reading women as men. → `persistent-memory.d/2026-09-16-lv-hemingway-corpus.md` -- `[2026-09-16]` **Grok token broker built then SHELVED — operator ruled "keep the jail stop the a/b", so the renewal feature has no consumer.** ⛔ Do NOT arm `probe-rotation`: the risk did not shrink (it reaches BOTH Gróa transports through one shared session) and the payoff went to zero. → `persistent-memory.d/2026-09-16-grok-broker-shelved.md` - `[2026-09-15]` ⚠⚠ **`--gpu-memory-utilization` DOES NOT PREDICT RESIDENT VRAM — measure it, never compute it.** Wrong in **both** directions on fv-ml1: `vllm-cyberprev` util 0.40 (expect ~39,155 MiB) holds **47,124** (+8 GB over); `vllm-gen-small` util 0.48 (expect ~46,986) holds **36,942** (−10 GB under). Planning a placement off the fractions would have been 8 GB wrong. Read `nvidia-smi --query-compute-apps`. Full per-seat residency table + the breeze shuffle arithmetic → `persistent-memory.d/2026-09-15-breeze-placement-sizing.md` - `[2026-09-15]` **breeze-tts stays on irv-ml1; the TTS-stack move to fv-ml1 is PARKED (park id 75, `move-the-tts-stack-breeze-tts-bragi-tts-gateway`), triggered on evacuating embed/rerank/reward.** ⚠ Trigger as stated says "gpu0" but those three are on **GPU 1** (~0.16 util, ~15.7 GB; GPU 1 is the tight card at 0.975 / 4,336 MiB free) — confirm which he meant before executing. All three services move together because only `breeze-tts` is GPU-resident (~10.3 GiB, **growing**) while `bragi` and `tts-gateway` are CPU proxies, and co-location is what avoids a cross-site hop per TTS call. **breeze-tts sizing — original recommendation NOT to move it.** ~**10.3 GiB** measured under load at 53 min uptime, **up from 9.2 GiB** shortly after warm-up (it grows; n=2, plateau unmeasured) — so GPU 0's 11,982 MiB free is a **1.7 GB margin and shrinking**, on the live chat serving path. ⚠ Two measurement traps: it reports **nothing at idle on the wrong card** (`BREEZE_GPU_DEVICES=0` = the **3090**, not the A6000), and an early reading understates it. ⭐ The real objection is **topology**: `tts-gateway` is on irv-ml1 and reaches it same-box, so moving breeze alone adds a cross-site hop to every TTS call against a 478 ms first-sample budget. GPU 3 would fit it but spends the reserve. → `persistent-memory.d/2026-09-15-breeze-placement-sizing.md` @@ -487,7 +436,7 @@ _As of 2026-09-30 ~1800 PT._ - `[2026-08-24]` **`nconnect=8` on `/mnt/smithy` — approved but DEFERRED at operator instruction.** brokkr-smithy-dev pre-approved it for "once the FortiGate work settles" and does not need re-asking; the operator declined it in this session's scope. Tracked at althing thread `01M0R46SFYF83099N16WD67KGD`. -_135 older entries archived to archival-memory.md._ +_167 older entries archived to archival-memory.md._ ## Tried and abandoned