memory: snapshot — speech seat live + gen-small OOM incident (mitigated, fix tasked); leftovers deleted; no Scriberr upstream; repos pushed; 32 entries archived

This commit is contained in:
vh
2026-10-01 04:24:44 -07:00
parent 6b66207b0c
commit f83e35b94b
32 changed files with 1189 additions and 1165 deletions
@@ -1,54 +0,0 @@
# `[2026-09-16]` Grok token broker — built, then shelved by the transport ruling. Do NOT arm the probe.
**`services/grok-token-broker/` — seeded, committed, DISARMED, no consumer.** Commits `ebc4dac`
`b907a0e` `cf9d167`. ⛔ **Do not arm `probe-rotation`.** This is a finished resting place, not a
half-built tool: the gate works and the thing it gated for went away.
**Operator ruling, relayed by heid:** *"keep the jail stop the a/b"*
(`heid dispatch-log/2026-09.jsonl#groa-transport-20260916-operator-keeps-the-jail`, alongside
`#groa-transport-ab-20260916-operator-stop`). Gróa dispatches through the read jail;
`groa_http_dispatch.py` is a documented fallback with no scheduled use. **Nothing in the fleet
wants a renewable xAI session.**
⭐ **THE CODE-PLAN ENDPOINT EXISTS and I was one message away from telling the operator it did
not.** `https://cli-chat-proxy.grok.com/v1` serving **grok-4.6** (500,000 context) and grok-4.5,
`agent_type: grok-build-plan`, `auth_method: session`, `api_key`/`env_key`/`api_base_url` all
null. ⚠⚠ **It is in `~/.grok/models_cache.json` — the Grok CLI's own config, on nh3-dev.** I had
swept heid's repo, the gateway `.env`, the LiteLLM config and Vaultwarden, all correctly, and
concluded "does not exist". ⭐ **heid's line, taken: absence from the places you searched is not
absence.** The check that separates the two states is a LIVE REQUEST, not a grep.
⚠⚠ **NOT WRITING `~/.grok/auth.json` IS NECESSARY AND NOT SUFFICIENT.** The refresh grant at
`https://auth.x.ai/oauth2/token` may ROTATE the refresh token, and many OIDC providers invalidate
the old one SERVER-SIDE. A broker refreshing the same credential kills the CLI login even though
it never touches the file. heid's module header reasoned about the WRITE; they amended it to name
invalidation, credited. This correction went infra-ops→heid the same day heid's went the other
way — **neither of us reaches the right answer alone.**
⚠⚠ **THE PROBE'S BLAST RADIUS IS BOTH GRÓA TRANSPORTS, which is not visible from the infra side.**
`heid/scripts/groa_dispatch.py` builds `argv = ["grok", "-p", prompt, "--cwd", jail, ...]` and
shells the CLI, which authenticates from the same `~/.grok/auth.json`. The bwrap in the process
table is grok's own Landlock sandbox, not something Heid wraps. **One session, two ways of
reaching it** — an invalidating probe takes Gróa down on EVERY path until an interactive re-login.
🔴 **I had recommended "run the probe now while the CLI is idle" and withdrew it in writing**;
"idle" was a convenient assumption I never checked, on a day that had already taken eight panels.
**Why the jail won, and it was not performance.** HTTP is faster (~523 s median vs ~890 s),
simpler, and arguably SAFER on confinement (no tools, so the 2026-06-10 escape class is
structurally impossible). It lost on FAILURE MODE: HTTP fails by returning a fast, confident,
well-formatted review that found nothing — indistinguishable from a clean bill. The jail fails by
timing out, which you can see. ⚠ **Do NOT quote a per-transport finding rate from this**: heid
states the 0/0/0-vs-5/7/3 numbers are confounded with bundle size (the zeros were all huge inline
bundles; the one HTTP round at jail-comparable size produced Gróa's leading solo), n=3–4 per cell,
no noise floor. Asymmetric-risk argument, **not** a resolved measurement.
⚠ **Still unmeasured, and it is a billing question:** the jail reaches the coding plan already
paid for; the HTTP path reaches the METERED API and its responses carry `cost_in_usd_ticks`.
Whether that bills on top of the plan was never part of the ruling. One look at the xAI billing
console — **this fleet holds no xAI credential**, so it needs the operator's account access.
⚠ The coding plan speaks the **Responses API** (`api_backend: "responses"`), not
`/chat/completions` — a second, independent obstacle to any LiteLLM alias. Moot while the jail is
ruled. heid also found and killed two live instructions in their own persistent-memory telling a
fresh session to dispatch `--groa-transport http`; either would have resumed a stopped experiment
after a context reset.
@@ -1,66 +0,0 @@
# `[2026-09-16]` lv-hemingway corpus: half the work was EXCLUSION, and the gate found what a hand count would not
**994,760 words · 318 units · 6 renamed copies · leak gate PASSED 0 of 941 entities and 0 of 117
audited phrases, both controls green.** `~/hemingway-corpus{,-renamed}`, builder
`scripts/hemingway-corpus/build_corpus_hemingway.py`, commits `9598d0b` `03b4a3f`.
⭐ **THE CATALOGUE HOLDS 2,105,679 WORDS AND ROUGHLY HALF MUST NOT BE TRAINED ON.** Operator
scoped it to fiction only. Three exclusion passes, each measured or voice-specific:
1. **Non-fiction, 8 works ~911k words** — By-Line, Dateline: Toronto, Death in the Afternoon,
Green Hills of Africa, The Dangerous Summer, the three posthumous "Hemingway on X" anthologies.
2. ⭐ **Four story collections, 169,759 words — MEASURED, not assumed.** `Short Stories` is the
First Forty-Nine and CONTAINS the others. 8-gram containment of the smaller work: Winner Take
Nothing **96.0%**, Snows of Kilimanjaro **95.2%**, Men Without Women **92.9%**, In Our Time
**90.6%**. ⚠⚠ **The catalogue's own `near_dup_pairs` table is BLIND to this** — it holds
whole-document simhashes (ONE row in the entire 1,284-work library) and this is PARTIAL
containment. Whole-document dedup cannot see a collection inside a larger collection.
3. ⭐ **`The Torrents of Spring` — excluded for a reason no word count could justify.** It is a
deliberate PARODY of Sherwood Anderson: the target author's name on a different author's
style, i.e. mislabelled data for a voice adapter.
⚠ **THE AUTHOR'S OWN NAME WAS IN THE TRAINING TEXT 95 TIMES ACROSS 7 WORKS** — publisher back
matter ("Ernest Hemingway was one of America's foremost journalists… died in 1961") riding inside
the last unit, because a splitter cuts on headings and nothing follows the final one. **Identical
to the Yarros defect; nothing about the source changed to cause it.** Stripping the publisher
block left 18, all in `true-at-first-light`, inside a **CAST OF CHARACTERS and SWAHILI GLOSSARY
written by Patrick Hemingway** — an editor describing the author's real household. Markers are
matched in file order, earliest wins. Now 0.
⚠ **`G` WAS ABOUT TO BE RENAMED TO A SURNAME, 248 TIMES.** Not a name: the fragment left by
`B.G.`, `G.M.`, `G2`, `G3`. Caught by reading surfaces IN CONTEXT, which is the Yarros lesson
repeating. Also read in context: `Gran` (fragment of `Gran Sasso`/`Gran Italia`/`Gran Hotel`),
`Shamba` (Swahili common noun), and `Inglés` — **kept renameable deliberately**, the gypsies'
in-world nickname for Robert Jordan, exactly parallel to Yarros's `Violence`.
⭐⭐ **THE GENDER RESOLVER HAD TO BE REBUILT AND ITS OWN GATE CAUGHT THE FIRST ATTEMPT.** The
inherited one returned **397 male / 20 female** across 1,102 records with Catherine Barkley,
Brett Ashley, Pilar, Maria and Mary all held neutral. A plain majority vote over nearby pronouns
scored 18 correct but **5 WRONG** against the incumbent's 1 — and **every error was
female-read-as-male** (Pilar m=426 f=243, Brett m=249 f=137). The refuse-unless-better guard
rejected it, correctly. ⭐ **Cause, measured: the corpus base rate is 34,315 male pronouns to
8,699 female, nearly 4:1.** Pilar's "male-dominated" 426:243 is strongly FEMALE against that
background. Scoring each name's local mix against the corpus base rate instead of 50:50 gives
**18 correct / 11 held / 0 WRONG**, distribution 275m / 120f.
`scripts/hemingway-corpus/gender_by_proximity.py`. ⚠ Yarros solved its version with the POV
chapter header; Hemingway's editions have none, so that fix does NOT transfer.
⚠ **THE ALPHABET DOES NOT TRANSFER EITHER: 1,496 non-ASCII letters across 23 forms** against
Yarros's 2. Hemingway writes Spanish, French and Italian constantly, so the rename pool needs
accents (new `hemingway` preset in `rename.py`). The Yarros ASCII-only conclusion would have
stranded every Spanish and Italian name in the cast — which is why F02 says re-derive per corpus.
**Three source defects the splitter surfaced.** `Islands in the Stream` came out as ONE
143k-word record (roman numerals, unhandled). `Short Stories` came out as 5 units then 27,
because the edition carries a **SECOND contents listing** and first-occurrence matching resolved
31 of 58 titles to an index entry — keeping every occurrence and letting the word floor decide is
self-correcting; now 57. ⚠⚠ **And the drop-cap defect is in the HEADINGS here** (`T HE O LD M AN
AND THE S EA`), which **INVERTS the Yarros pipeline order: repair must run BEFORE the split**, or
the splitter cannot see the headings it needs.
⭐ **Hemingway needed THREE mapped phrases where Yarros needed 48** (`Gran Maestro`,
`Unknown Tongue`, `Sin House`) and a 45-entry allow list — the whole difference being that Yarros
invented a world and Hemingway named the real one. Several allow entries were non-obvious and
required reading: `Royal Game` is a real colonial-Kenyan legal category, `White Heather` a Scotch
brand, `Bwana Game` a job title, `Roman Soldier`/`Wine Seller` stage-direction labels from the
one-act play `Today is Friday`.
@@ -1,53 +0,0 @@
# `[2026-09-16]` The lv-* voice line: Option C proved, lv-yarros shipped, lv-hemingway training
⭐⭐ **INSTRUCTION-PAIR SFT BEATS RAW-TEXT TRAINING FOR AUTHOR VOICE, AND THE INCUMBENT NEVER
CLEARED ITS OWN CONTROL.** Measured n=120 per arm, 30 in-genre beats from HELD-OUT val passages
× 4 seeds, all arms re-measured in one session on one box:
| arm | delta_cb (lower = more Yarros) | vs base control | 8-gram overlap |
|---|---|---|---|
| pairs 2ep ckpt-1650 | **0.410** | +0.289 ✅ | 0.12 |
| pairs 3ep ckpt-1650 (**shipped**) | 0.438 | +0.262 ✅ | **0.09** |
| raw-text instruct (incumbent) | 0.558 | +0.141 ❌ **inside the 0.153 floor** | 0.14 |
| base-unadapted (control) | 0.700 | — | 0.07 |
Same-author target 0.463 (held-out Yarros vs itself). ⚠ **The two pair arms are NOT
distinguishable on voice** — 0.028 against a 0.153 floor. The 3ep checkpoint was chosen on the
axes that ARE resolvable: better held-out fit (2.3126 vs 2.3264), less overshoot (0.06 vs 0.10),
and verbatim overlap nearest the never-saw-it control.
⭐ **THE RECIPE IS TWO EPOCHS ON A THREE-EPOCH SCHEDULE, not three epochs.** Launch `--epochs 3`;
the minimum lands at step 1650 **inside epoch two** and epoch three overfits (2.3126 → 2.3882,
flat). The entire gain over a 2-epoch run came from the stretched cosine keeping the LR alive —
at step 1600 the 3ep run was at 2.9e-05 where the 2ep run had annealed to 2e-07. ⚠⚠ **A
resume-and-append-one-epoch is a NO-OP for exactly that reason** (lr 2.3e-09 at step 1670): it
must be a fresh run with the longer schedule.
⭐ **THE SAFETY PROPERTY: the model writes the INSTRUCTION, never the RESPONSE.** Every response
is real renamed prose; only the beat is machine-written, so voice is inherited rather than
synthesised. Memorisation checked with both controls (positive control saturates at 160): the
shipped arm sits at 0.09 against a 0.07 never-saw-it baseline and BELOW the raw-text arm's 0.14.
⚠ **THE v1 DECISION RULE WAS WELL-FORMED AND MEASURED THE WRONG THING**, and the amendment is
recorded in `scripts/yarros-corpus/score_beats.py` with v1 retained verbatim. It gated on
in-band / on-beat / ran-on — and **base-unadapted scores in-band 0.96**. Instruction-following is
something Qwen3-4B-Instruct ships with, so those axes detect only DAMAGE, never the benefit an
adapter exists to buy. v2 gates on voice (delta_cb vs control beyond the floor) + not-copied
(8-gram overlap near control) + no-damage (overshoot). ⚠ on-beat's −0.27 was outside the floor
and is dropped from the gate, **not explained away** — the keyword proxy punishes prose that
DRAMATISES "she mocks him" rather than echoing the word, but three read samples is an anecdote.
⭐ **A 5-BEAT FIXTURE HAD A NOISE FLOOR OF 0.800 AND MANUFACTURED A +0.45 RESULT.** At n=20 the
pilot looked like a clear in-band win; at n=120 the same gap was +0.08, inside a 0.233 floor.
One sample moves a rate by 0.2 when there are five. The 30-beat in-genre fixture (built from
held-out val pairs, `~/beats-yarros-30.json`) is the instrument; the Brontë stray-dog/kitten
fixture was also the wrong GENRE — "He licked her clean" came back as explicit sex.
⚠ **THE HARNESS TRUNCATES AT THE FIRST BLANK LINE and that surface reported the pair arm as
"19 words, off-beat 0.10"** when the untruncated output was 90–132 words with the beat rendered
in a later block. `score_beats.py --metric-source raw|paragraph` keeps both views and the verdict
names which it used. Same family as `feedback_filters_that_silently_narrow_the_window`.
**Artefacts.** `scripts/yarros-corpus/{build_sft_pairs,train_pairs_lora,score_beats,
memorization_check}.py`; commits `9b3d3c8` `90ed506` `713e83d` `efb7345` `7505124`. Booth
(24h TTL) was `http://10.100.10.50:8090/b/babyyarros-beats/` — six beats × four arms, blind-labelled.
@@ -1,50 +0,0 @@
# `[2026-09-16]` voices-seat: LoRA over merge, measured — and GPU 0 is now full
**`vllm-voices` live on fv-ml1 GPU 0 :8027**, one Qwen3-4B-Instruct carrier serving
`voices-base` plus `lv-<author>` LoRA adapters. `stacks/voices-seat/`, commit `d17bd3d`.
⭐ **LORA COSTS 24.3% OF DECODE THROUGHPUT AND IT IS WORTH PAYING.** n=30 per arm, interleaved,
A-vs-A noise floor **0.1%**: base **143.0 tok/s** median vs adapter **108.2**. The measurement is
unusually clean because `--enable-lora` serves BOTH the base name and the adapter name from ONE
process — the arm is a per-request field, so no restart, no second seat, no cold-vs-warm confound.
Arms were **interleaved rather than blocked** because the card's co-tenants take traffic this seat
does not control, and a block design would alias their load onto one arm.
**Why pay it:** 3 authors cost 8.4 GB as adapters against ~23 GB merged; 6 cost 9.2 vs ~46. On a
card with 1.8 GB free afterwards that is the whole argument. If a voice ever lands on a latency
path, merge THAT one and serve it separately.
⭐ **ADAPTER HOT-SWAP IS REAL AND FAST — MEASURED, not read from docs.**
`POST /v1/load_lora_adapter` **200 in 0.24 s**, `POST /v1/unload_lora_adapter` **200 in 0.003 s**,
VRAM unchanged, container stayed healthy. Proven by performing the `babyyarros`→`lv-yarros`
rename through it with no restart. ⚠ **A runtime-loaded adapter is GONE on the next
`compose up -d`** unless it is also in `--lora-modules` (which costs a recreate + ~3 min reload).
Runtime load is for TRYING a voice; the compose list is what persists. Switching between loaded
voices is just the `model` field — **not** a LiteLLM alias; LiteLLM is a thinner layer on top,
one alias entry per voice, no new deployment.
⚠⚠ **`--gpu-memory-utilization` IS A REQUEST AGAINST *TOTAL* VRAM THAT THE CARD MUST ALREADY BE
ABLE TO HONOUR — not a share of what is free.** First bring-up REFUSED: *"Free memory on device
cuda:0 (11.16/94.97 GiB) is less than desired GPU memory utilization (0.12, 11.4 GiB)"*. Refusing
was the right outcome — it protected `cyberprev`, `gen-small` and the Parakeet STT seat rather
than squeezing them.
⭐ **PINNING `--kv-cache-memory` IN BYTES MAKES THE FRACTION PREDICTIVE.** Requested 0.11
(10,700 MiB), got **10,740 MiB** resident — a 40 MiB miss on a box where the fraction has been
wrong by **8–10 GB in BOTH directions** (cyberprev 0.40→47.1 GB, gen-small 0.48→36.9 GB). Second
seat to prove it after `gen-small`. Do not remove the pin.
⚠ **fv-ml1 GPU 0 is now 96.0 of 97.9 GB.** GPU 1 ~5.7 free, GPU 2 ~2.4, GPU 3 is a HELD RESERVE
for a future full-card seat (`flash-next` alone needs 93 of 96 GiB). **There is no room for
another seat on fv-ml1 without a placement decision.**
⚠ **SUPPORT WAS CHECKED, NOT ASSUMED**, per the training playbook's own lesson that LoRA support
is per-ARCHITECTURE not per-family: `vllm/model_executor/models/qwen3.py:271` declares
`Qwen3ForCausalLM` with `SupportsLoRA` plus `packed_modules_mapping` and `embedding_modules`.
**Do not transplant this compose onto an MoE carrier without re-running that grep** — the
playbook records a LoRA refusal on a Qwen3 MoE arch.
**Naming (operator, 2026-09-16):** `lv-<author>` — lv for **lang-voice**, retiring `baby*`, which
read fine for one experiment and invites confusion across a family. The adapter NAME is the
request's `model` field, so it is the public API of a voice. Historical persistent-memory entries
still say BabyYarros/BabyHemingway and were deliberately left as dated records.
@@ -1,3 +0,0 @@
# `[2026-09-17]` A stoplist entry is an assertion the leak gate can no longer check
⭐ **A stoplist entry is an assertion the leak gate can no longer check** — stoplisting removes a surface from the entity map, so a wrongly stoplisted CHARACTER is an undetectable leak. Three were wrong on Brontë (Leaven, Pierrot, Samuel); `scripts/r49-corpus/audit_stoplist.py` finds them by honorific and now gates the pipeline. Commit `8bb7686`.
@@ -1,3 +0,0 @@
# `[2026-09-17]` A unit splitter must choose by SIZE, not by count — the inherited rule silently produced 22,000-word "chapters".
⭐⭐ **A unit splitter must choose by SIZE, not by count — the inherited rule silently produced 22,000-word "chapters".** `scripts/r49-corpus/split_units.py`: a marker mode qualifies only if its median unit is inside [600, 12000] AND no unit holds half the work; among qualifying modes PRIORITY breaks the tie (contents > chapter-word > roman > bare-numeral > caps-title), and paragraph-block sections are the fallback for works with no divisions. ⭐ **Both rules exist because a control caught them**: scoring by "median closest to target" chose `caps-title` (6 units, one holding **97%** of the book) over True at First Light's real 20 chapters, because a median cannot see that distribution and a max bound can. Positive control: 8/10 Hemingway works reproduce the shipped mode and count exactly. Negative control: 40,000 words with no blank lines → 1 unit, refuses to fabricate divisions. Commit `705fa3a`.
@@ -1,3 +0,0 @@
# `[2026-09-17]` `audit_entity_map.py` — the rename can DAMAGE the prose and no gate will ever say so.
⭐ **`audit_entity_map.py` — the rename can DAMAGE the prose and no gate will ever say so.** Mirror of `audit_stoplist.py`: surfaces wrongly held IN the map rather than out of it. 130 of 946 Hemingway surfaces flagged (1,616 instances, 0.162% of words) — `African`, `Chinese`, `X-ray`, `Coca-Cola`, `Ritz`, `Prado` renamed into invented names — plus 16 bare initials incl. `C` at 274 occurrences. Signal is a preceding article; controls derived from the corpus, not hand-picked. Every hit reported for READING: `the Widow` and `the Informer` are genuine epithet-names that should be renamed. Commit `051b99e`.
@@ -1,39 +0,0 @@
# `[2026-09-17]` The beat-writing model restores the author's real names — a leak the corpus gate structurally cannot see
⭐⭐ **THE RENAME AND ITS GATE ARE BOTH WORKING, AND THE TRAINING DATA IS STILL CONTAMINATED.**
The corpus gate reads the corpus and the renamed copies. **It never reads the generated
instruction beats.** Those are written by an LLM that just read the passage — and if it
recognises the book, it supplies the canonical names out of its own training.
**Measured on the first 714 lv-bronte pairs, before the filter existed:**
- **13 beats (1.8%)** named source characters — `Rochester` ×6, `Jane` ×3, `Brocklehurst` ×2,
`Beck`, `Fairfax`, `Helen`, `Burns`, `Eyre`, `Reed`, `Rivers`.
- **0 of 714 RESPONSES did.** The rename was perfect; the instruction side was not.
- One beat read *"Saoirse confirms Rochester's flaws, then agrees in English to marry him"* —
a renamed name and a canonical one in the same sentence, which is the mechanism in miniature.
**Why it matters more than 1.8% sounds:** the beat is the INSTRUCTION half of the pair, so
training on it re-teaches exactly the inventions the rename pipeline exists to remove.
⚠⚠ **EXPOSURE SCALES WITH HOW WELL THE GENERATOR KNOWS THE BOOK.** It is worst for
public-domain classics and mildest for recent work. That is precisely why the Yarros and
Hemingway runs came up clean and Brontë did not — **their clean runs are NOT evidence they are
immune.** Both should be re-verified, and regenerated with `--source-entities`, before their
pairs are trusted again.
**The fix.** `vet()` in `scripts/yarros-corpus/build_sft_pairs.py` gained a `sourcename`
reject plus `--source-entities <entities.json>`, taking the UNRENAMED entity map. Fired at
~3% of attempts on the Brontë rebuild. Commit `533cc0c`.
**The end-to-end guard that proves it.** The chain now verifies every built pair — beat,
response and context — against every source surface before spending GPU hours:
`[verify] 3858 pairs vs 368 source surfaces -> 0 leaks`.
⚠ A guard stricter than the gate cries wolf. The first verify pass excluded nothing and
flagged `Monsieur’` ×14 and `Qu’est-ce` ×4 — French grammar, not leaks — because
`--fold-clitics` leaves apostrophe keys the gate deliberately skips. Mirror `leak_gate.py`'s
own predicate; a guard that fails on false positives gets disabled, which is worse than the
leak it guarded.
Related: [[2026-09-17-lv-bronte-gate]], [[2026-09-16-lv-hemingway-corpus]], [[2026-09-16-lv-voices-line]].
@@ -1,68 +0,0 @@
# `[2026-09-17]` ESH: Cityside Fiber failed twice in six hours; site ran on Verizon failover
**Timeline (PDT).**
```
19:09:07 Cityside dies. UDM fails over to Verizon 5G (WAN2). Site stays up at ~33 ms.
19:51 Verified healthy on failover.
20:01:37 esh-scale drops off the headscale mesh; 10.0.0.0/16 withdrawn; whole site dark
from the colo. Beszel fires on all five ESH hosts.
20:11-15 Service restored. Operator had switched WAN1 to DHCP to get Cityside working at
all; it came back on CGNAT, not the static. Latency back to 9 ms.
01:06:23 Cityside fails AGAIN. Failover to Verizon. Site up, ~37 ms.
```
⭐ **The 20:01 blackout was most likely the operator's own WAN reconfiguration**, not ISP
instability — switching the WAN type bounces the interface, esh-scale loses its path,
headscale withdraws the route, and the site vanishes from the colo's view until it settles.
An earlier session theory ("Cityside came back half-provisioned and the UDM failed back into
an unstable session") is retired.
⚠ **The diagnostic that mattered: physical link stayed UP at 2.5 GE with zero errors
throughout, while the ISP's next-hop `128.177.138.181` was unresponsive.** So "the ONT is
fine, it is upstream of the ONT" — the line to give Cityside. Traceroute from NH3 reached
`209.249.146.170` (one hop short) before dying, so the prefix was still routed.
⚠ **CROWDSEC ROTATION FRAGILITY IS LIVE.** The `esh` allowlist on ana-docker carries the now-
dark static `128.177.138.182` (never-expiry), plus `97.190.18.88` (Verizon failover) and
`23.164.40.174` (Cityside CGNAT), both **7-day expiry**. ESH is on a rotating carrier-NAT
egress until the static is restored — the exact regime the 09-08 static purchase was meant to
end, and the class that once blackholed the whole site via a false ban. **If ESH loses colo
access, check `curl -s4 ifconfig.me` from esh-docker-vm FIRST** and allowlist the new address.
**Still pinned to the dark static and broken until it returns:** FortiGate `infra-ops`
trusthost3 = `128.177.138.182`, so logins to ana-gw from ESH are refused. The dormant
`esh-ana` IPsec is bound to wan1/static (disabled, so no impact).
⭐ **The mesh was NOT degraded on CGNAT** — tailscale hole-punched straight through
(`direct 23.164.40.174:41641`), which is why latency read 9 ms rather than a DERP figure. An
expectation of relay-on-CGNAT was wrong.
## RESOLVED 2026-09-17 ~12:30 PT — the static is back, confirmed on four axes
Not one check, because egress alone cannot tell a static WAN from a carrier NAT that happens
to answer (see auto-memory `feedback_egress_ip_cannot_detect_cgnat`):
```
config UDM WAN1 `wan_type = static`, ip 128.177.138.182, mask /30, gw 128.177.138.181
— switched BACK from the DHCP the operator set at 20:11 during the outage
active stat/health: isp_name "Cityside Fiber", ASN 18731, num_disconnected 0.
WAN2 Verizon-5G is failover-only at priority 2 and idle.
egress esh-docker-vm sees 128.177.138.182 — EQUAL to the WAN ip, so not behind CGNAT
perf 2005/2142 Mbps symmetric; colo -> ESH 5.0 ms, 0% loss over 4 hosts-worth of pings
(Cityside CGNAT read 9 ms, Verizon failover 33-37 ms)
```
⭐ **The FortiGate pin un-broke itself and that was verified, not inferred.** `infra-ops`
trusthost3 is `128.177.138.182`; from esh-docker-vm, ana-gw `tcp/22` is OPEN and the
FortiGate offers a password prompt rather than dropping the connection — a trusthost
mismatch refuses outright, so reaching auth *is* the trusthost passing. The dormant
`esh-ana` IPsec bind to wan1/static is correct again (still disabled, still no impact).
⚠ **The crowdsec temporary allowlist entries are being LEFT to expire on their own**
(2026-09-23): `97.190.18.88` Verizon and `23.164.40.174` Cityside CGNAT. Cityside failed
twice in six hours on 09-17, so until the line has earned some confidence those two are
cheap insurance against the exact false-ban blackout this rotation-fragility caused before.
`128.177.138.182` stays never-expiry.
Related: [[2026-09-06-headscale-cutover]], [[2026-09-08-esh-static-wan-followups-and-ytvc]].
@@ -1,3 +0,0 @@
# `[2026-09-17]` gitea was reaching the PUBLIC route from every repo on nh3-dev
**gitea was reaching the PUBLIC route from every repo on nh3-dev** — brokkr-smithy, sleipnir, Galdrabok, kvasir — and brokkr-smithy is pushed several times a week, so the fail2ban trigger was live, not dormant. Measured before acting (no split-horizon rewrite, no ssh alias, `ssh -G` confirmed port 22 to 38.120.12.44). Fixed by overriding the NAME once in `~/.ssh/config` rather than rewriting N remotes, so fresh clones and unaudited repos are covered too. Verified with a real `git ls-remote`, not by inspection. Commit `dcc1abc`. Flagged by brokkr-smithy-dev; `vh/imogen` created for them the same session.
@@ -1,3 +0,0 @@
# `[2026-09-17]` headscale now split-DNSes `nh3.phasefinal.com` to the three AdGuards, so mesh clients can resolve the internal-only wildcard
**headscale now split-DNSes `nh3.phasefinal.com` to the three AdGuards, so mesh clients can resolve the internal-only wildcard** (`talk`, `booth` — public DNS has no record for them; the fleet AdGuard answers 10.100.10.50). Operator-approved, scoped to nh3 rather than all of `phasefinal.com`. Config `/etc/headscale/config.yaml` in CT 106 on nh3-pve, backup `config.yaml.bak-2026-09-17-splitdns`, restarted, and the new route **read back from a node's netmap** rather than assumed. ⚠ Two things worth knowing: split DNS works fine here with `global: []` — headscale issue #1161's "split ignored without global" does NOT apply to v0.29.3, verified on the live mesh — and `override_local_dns: true` would REQUIRE global, which is the config that makes a roaming laptop lose ALL DNS when the mesh is down. That is why split, not global. Routing was never the problem: nh3-scale already serves 10.100.0.0/16.
@@ -1,3 +0,0 @@
# `[2026-09-17]` Hemingway ships as-is: operator ruled "ship stands" on both measured corpus defects
**Hemingway ships as-is: operator ruled "ship stands" on both measured corpus defects** — the 0.96% beat contamination and the 130 non-name entity-map surfaces. `audit_pairs_sourcenames.py --filter-out` and `audit_entity_map.py` exist and are the instruments if that is ever revisited; neither was run against the shipped adapter.
@@ -1,135 +0,0 @@
# `[2026-09-17]` lv-bronte: corpus gated for real, adapter trained, SHIPPED with a FAILED voice axis on the record
**Status: SHIPPED 2026-09-17 01:24 as `lv-bronte` on `vllm-voices` (fv-ml1 GPU0 :8027), ckpt475 —
and it did NOT pass its voice gate.** Shipped because it is additive (one more named LoRA beside
`voices-base` and `lv-yarros`, reached only by requesting it), reversible (one compose line; hot-unload
measures 0.003 s), and clean on the SAFETY axis — 8-gram overlap identical to the never-saw-it control,
on a public-domain corpus. VRAM cost was nil: GPU0 96092 -> 96090 MiB. The caveat is written into the
compose file and into `/tank/aimodels/voice-adapters/lv-bronte-4b-v1/README.md` so it cannot be read
as a clean pass by anyone who finds the adapter without finding this note.
⚠ **Do NOT cite lv-bronte as evidence pair-SFT works for this author.** The voice axis is unresolved,
not passed.
## The gate result, in full
| axis | result | numbers |
|---|---|---|
| **A. VOICE** | ❌ **FAIL** (both candidates) | ckpt925 +0.210, ckpt475 +0.193 vs base — both **under** the 0.251 measured noise floor |
| **B. NOT COPIED** | ✅ PASS | ckpt475 **0.00 hit-rate, max 0 — identical to the never-saw-it control**; ckpt925 0.01, max 8. Positive control saturates at 160, so the detector is not blind |
| **C. NO DAMAGE** | ✅ PASS | ran-on +0.15 against a 0.400 floor |
```
same-author target (held-out Brontë vs itself) delta_cb 0.338 <- best achievable
ckpt925 0.531
ckpt475 0.548
base-unadapted 0.741
```
⭐ **THE EFFECT LOOKS REAL AND SUBSTANTIAL; THE INSTRUMENT CANNOT CERTIFY IT.** The reachable
span is 0.741 → 0.338 = 0.403, and the adapters closed **48–52% of everything achievable**.
Both beat base on *every individual seed*. This is an UNDERPOWERED result, not a null one —
and a "no effect" without its floor is unfalsifiable, so: **this method cannot resolve a voice
improvement smaller than ~0.251 delta_cb at 30 beats × 4 seeds on this corpus.**
⭐⭐ **THE CAUSE IS STRUCTURAL: Brontë's val split yields 81 pairs (44 in-band) against
Hemingway's 200**, because the corpus is 678k words against 994k. Maxing the fixture 30 → 44
beats would shrink the floor by only ~√1.47 ≈ 1.2× (to ~0.21, against a 0.21 gap) — still
marginal. **More SEEDS would not help either: the floor is a RANGE statistic, and ranges widen
with more samples.** There is no cheap fix.
## ⚠ A DEFECT IN THE v2 RULE ITSELF, worth fixing before the next author
The floor is defined as the **largest within-arm seed spread across ALL arms**. Measured here:
```
base-unadapted 0.772 0.813 0.751 0.772 spread 0.062
ckpt475 0.670 0.631 0.604 0.578 spread 0.092
ckpt925 0.776 0.584 0.525 0.620 spread 0.251 <- sets the floor, on ONE seed
```
So **adding a third, noisier arm raised the bar that failed the clean one.** Run as the
two-arm gate (base + ckpt475) the floor would have been 0.092 and +0.193 would have cleared
at 2.1×. This was NOT exploited — picking the floor that passes your preferred answer is the
threshold-chosen-after-seeing-the-numbers failure the pre-registration exists to prevent — but
the rule should say whether the floor is computed over the compared pair or over every arm
present. As written, a candidate's verdict depends on which *other* arms you happened to run.
**The outlier was diagnosed, not waved away.** Degeneracy probe (fraction of a generation made
of its most repeated 5-gram) is uniform across every seed and both arms, 0.0078–0.0102. Seed
1234 is not a collapsed generation; delta_cb genuinely has that variance for that arm.
## Which checkpoint, if it ships: **ckpt475**
The two are 0.017 apart on voice — far inside any floor, i.e. indistinguishable. On the axes
that DO resolve, ckpt475 wins both: memorisation identical to the control (ckpt925 has a
verbatim 8-gram hit), and **2.7× tighter seed-to-seed variance** (0.092 vs 0.251) with no
degeneracy to explain the difference — consistent with ckpt925 sitting nearer the overfit
boundary. Given a coin-flip on voice, take the one that provably did not memorise.
⭐ **THE RECIPE DID NOT TRANSFER.** Yarros and Hemingway both found their minimum inside
epoch two. Brontë's minima are step 475 (ep 1.00, 2.6107) and step 925 (ep 1.96, 2.6129) —
**0.0022 apart against a 0.0046 median neighbour jitter, i.e. indistinguishable**. Epoch 2
buys Brontë NOTHING over epoch 1. What IS robust is the epoch-3 collapse: +0.075, ~16× jitter.
Do not carry "two epochs on a three-epoch schedule" to a new author as settled.
## Artefacts
`gx10:~/lv-bronte/` (corpus-clean, corpus-renamed, entities-final.json, pairs/, beats-bronte-30.json),
`gx10:~/r49-runs/bronte-4b-pairs-3ep/` (57 checkpoints kept), `gx10:~/r49-runs/bronte-eval/`
(three arms × 120 generations, memorization.txt, voice_distance.txt, score.*.txt).
Commits `fc834a8` `533cc0c` `7964d07` `e9e8c40` `8bb7686`.
⚠ Two output labels in `voice_distance.py` are hardcoded Yarros strings — it prints
"reference: held-out Yarros" and a boilerplate "Base < Instruct" corroboration line. The
NUMBERS are Brontë's; those two labels are not. Not yet fixed.
Related: [[2026-09-16-lv-voices-line]], [[2026-09-16-lv-hemingway-corpus]], [[2026-09-16-voices-seat-lora]].
---
## ⚠⚠ AMENDED 2026-09-17 — THE VOICE AXIS PASSES UNDER THE CORRECTED FLOOR RULE
Everything above is left verbatim; it is what was believed at ship time. This section is
the correction, not a rewrite.
**The defect this file itself named was fixed, and fixing it flips ckpt475's verdict.**
The section "⚠ A DEFECT IN THE v2 RULE ITSELF" above says the rule should state whether
the floor is computed over the compared pair or over every arm present. It is now
**pairwise**, pre-registered in `scripts/hemingway-corpus/GATE-PREREG.md` before a single
lv-hemingway number existed. Re-scoring the SAME 360 generations — no re-run, no changed
delta_cb:
```
arm delta_cb per-seed spread
ckpt925 0.531 (0.776 0.584 0.525 0.620) 0.251
ckpt475 0.548 (0.670 0.631 0.604 0.578) 0.091
base-unadapted 0.741 (0.772 0.813 0.751 0.772) 0.062
all-arms floor (as run) 0.251
ckpt475 +0.193 vs pairwise floor 0.091 -> MOVED toward Brontë, 2.1x <- the two rules DISAGREE
ckpt925 +0.210 vs pairwise floor 0.251 -> within the floor, NOT a finding
```
⭐ **The sequence matters and is the reason this is not threshold-shopping.** The previous
session found the defect, recorded it, and explicitly declined to exploit it. The rule was
then changed prospectively on a structural argument independent of the answer it produces —
the sampling variability of a difference A−B depends on A and B, never on a third arm C, so
a candidate's verdict must not depend on which other arms were generated. `voice_distance.py`
prints both floors and flags disagreement, so neither number can be quoted alone.
**Consequences:**
- lv-bronte's voice axis is a **PASS at 2.1x**, not a fail. The caveat is amended in place
(append-only) in `stacks/voices-seat/compose.yaml` and
`/tank/aimodels/voice-adapters/lv-bronte-4b-v1/README.md`.
- The sensitivity floor for that measurement is **0.091**, not 0.251.
- "Do not cite lv-bronte as evidence pair-SFT works for this author" is **WITHDRAWN**.
- ckpt475 over ckpt925 is unchanged and for unchanged reasons (no verbatim 8-gram hit,
2.7x tighter seed variance).
- The "no cheap fix for the underpowered result" analysis above is superseded for Brontë:
it was underpowered against an inflated floor, not against its own.
**Also amended:** the two hardcoded Yarros labels flagged at the end of this file are fixed.
`voice_distance.py --author` is now REQUIRED — the committed Brontë output literally reads
"reference: held-out Yarros" over Brontë's numbers — and the stale "one seed-pair per arm /
corroborates Base < Instruct" footer now reports what the run actually carries.
@@ -1,163 +0,0 @@
# `[2026-09-17]` lv-hemingway: SHIPPED on ckpt850 — the line's first clean voice pass, and one axis that needs reading
**Status: SHIPPED 2026-09-17 03:33 as `lv-hemingway` on `vllm-voices` (fv-ml1 GPU0 :8027),
checkpoint-850.** Seat healthy 190 s after recreate, four models served
(`voices-base`, `lv-yarros`, `lv-bronte`, `lv-hemingway`), GPU0 96,092 → **96,090 MiB** — a
LoRA rides inside the existing seat and costs nothing. Adapter verified byte-identical to
the checkpoint by sha256 across two hops.
Gate design **pre-registered before any generation existed**:
`scripts/hemingway-corpus/GATE-PREREG.md`, commit `0bb4938`.
## The gate result — 3 arms × 60 held-out beats × 4 seeds = 240 generations per arm
| axis | result | numbers |
|---|---|---|
| **A. VOICE** | ✅ **PASS, 6.4×** | +0.413 delta_cb vs base, pairwise floor 0.064. Also clears the OLD all-arms floor (0.113) — **this verdict does not depend on the rule change** |
| **B. NOT COPIED** | ⚠ **content clean, rate 7× the author's own** | 0.07 hit-rate, mean-longest 0.6, **max 9 words**. Base 0.00, **held-out Hemingway 0.01** |
| **C. NO DAMAGE** | ✅ PASS | ran-on +0.08, on-beat −0.14, both inside a 0.217 floor; in-band 0.79 vs base 0.05 |
```
same-author target (held-out Hemingway vs itself) delta_cb 0.364 <- best achievable
ckpt1750 0.439
ckpt850 (SHIPPED) 0.511
base-unadapted 0.924
```
⭐ **THE STRONGEST VOICE RESULT IN THE LINE. The span is 0.924 → 0.364 = 0.560 and ckpt850
closed 73.8% of it (ckpt1750 86.6%)**, against lv-bronte's 48%. Power came from the corpus,
not from a better method: 173 in-band val pairs allowed a **60-beat** fixture where Brontë
had 44 in-band and could only run 30.
## ⚠⚠ AXIS B — THE COMFORTABLE EXPLANATION WAS WRONG, AND THE CONTROL IS THE ARTIFACT
`memorization_check.py` uses the **base-unadapted arm** as its negative control, and on this
corpus that control is weak in one direction only — **it makes an innocent arm look guilty.**
Base writes 18,035 words of *summary*; the adapted arms write 27,413 of *pastiche*. Text that
does not imitate the register cannot collide with its n-grams, so base's 0.00 partly measures
"different register", not "did not memorise".
The obvious hypothesis was that Hemingway's plain, high-frequency register makes 8-gram
collisions inevitable for any arm that learns it. **That hypothesis is refutable, was tested,
and is FALSE.** New control: **held-out Hemingway — the author himself, val text no arm
trained on — scored against the train split**, chunked to the generations' own median length
(101 words) so the comparison is like for like.
```
sample n hit-rate mean-longest max
HELD-OUT HEMINGWAY (never trained) 370 0.01 0.1 10
base-unadapted 240 0.00 0.0 0
ckpt1750 240 0.08 0.7 9
ckpt850 (SHIPPED) 240 0.07 0.6 9
positive control (train vs train) 160 <- not blind
```
⭐⭐ **The adapter reproduces train-corpus word sequences ~7× more often than the author
reproduces himself.** If the register explained it, real Hemingway would collide at the same
rate; it collides at 0.01.
⭐ **And the exposure is still nil, which is a different question from the rate.** All 19
matched runs were READ, not counted. Every one is stock dialogue — `i don t think so the girl
said`, `came over and sat down at the table`, `how do you feel i feel very well`. No plot, no
imagery, no distinctive phrase, **no proper noun** (the one name-shaped hit, `swift tristan`,
is the RENAMED invented name, not Hemingway's). The longest run is **9 words — shorter than
the 10-word run genuinely unseen Hemingway shares with the train split by coincidence.**
What is being reproduced is the *grammar of his dialogue*, which is the thing the adapter
exists to learn, rendered in the commonest words in English. **Elevated rate, zero
protectable content.** Hemingway is in copyright; the in-line precedent is lv-yarros, also in
copyright, shipped at 0.10 against a 0.07 control. Unload is 0.003 s and one compose line.
⚠ **The durable lesson is about the instrument, not this adapter: a negative control that
differs from the candidate in a way CORRELATED with the metric is not a control.** Always ask
what the metric returns for a known-innocent sample *in the same register*.
## Why ckpt850 and NOT ckpt1750, the loss minimum
ckpt1750 has the better point estimate on voice (0.439 vs 0.511) and **it is not usable**:
```
gap between candidates 0.072
pairwise floor max(0.113, 0.050) 0.113 -> NOT resolvable
```
Indistinguishable, so the pre-registered tiebreak falls to the axes that resolve — and
**ckpt850 wins every one**:
| | ckpt850 (shipped) | ckpt1750 |
|---|---|---|
| seed spread | **0.050** | 0.113 — **2.3× wider** |
| memorisation hit-rate / mean-longest | **0.07 / 0.6** | 0.08 / 0.7 |
| ran-on | **0.08** | 0.12 |
| epoch | **0.959** | 1.973 |
ckpt1750's spread is one seed: 0.491, 0.449, 0.468, then **0.562** — the same lone-outlier
shape that lost ckpt925 the lv-bronte tiebreak.
⭐ **THE TWO-EPOCH RECIPE DID NOT TRANSFER HERE EITHER — it is now 0 for 2.** Hemingway's
minimum really is step 1750, but step 850 is **+0.0040 against a 0.0044 median neighbour
jitter**, with three checkpoints inside one jitter of the best. Epoch 2 buys nothing that
resolves and costs 2.3× the variance. Only the epoch-3 collapse is robust: **+0.0762 = 17.4×
jitter**, which is why `adapter/` was never gated. **Stop carrying "two epochs on a
three-epoch schedule" forward; read the curve and prefer the earlier tied checkpoint.**
## Pre-flight: the beat leak IS present in Hemingway, and the fixture is clean
`audit_pairs_sourcenames.py` (new, commit `0bb4938`) closes the blind spot `leak_gate.py` has
by construction. Controls green every run: 941/941 surfaces found in the unrenamed source,
nonce absent from both trees, 6/6 planted names detected.
```
train beats 70 of 7,094 (0.96%) Santiago x16, Catherine x7, Rinaldi x3, Brett, Harry,
Jake, Pablo, Nick, Maria, Helen ... 36 distinct
train responses 0 of 7,294 -- the rename itself held perfectly
val beats 0 of 200 -- THE EVAL FIXTURE IS CLEAN; the gate is unconfounded
```
⭐ The beat-only signature is exactly lv-bronte's. **Yarros's and Hemingway's earlier clean
runs were never evidence of immunity** — they predate the detector.
**Cross-validated on real data** where the answer was already recorded: the fixed Brontë
pairs return **0 of 3,858** (matching "0 leaks across 3,858 pairs"), and
`pairs-full.CONTAMINATED.jsonl` returns **15 of 792 = 1.89%** with Rochester ×6, Jane,
Brocklehurst ×2, Beck, Fairfax, Burns, Helen, Eyre — against a record of "13 of the first 714
beats (1.8%)" with the same names. An independently written instrument reproducing a
documented finding at the right magnitude is what makes its zeroes mean *absent*, not *blind*.
`--filter-out` produces a clean **7,024-pair** set in one command (70 dropped, 0.99%),
verified by re-audit at 0 of 7,024. **A retrain on it is the operator's call, not done.**
## ⚠ A SECOND corpus defect, measured and NOT acted on
`audit_entity_map.py` (new, commit `051b99e`) is the mirror of `audit_stoplist.py`: it finds
surfaces wrongly held **IN** the entity map, which `leak_gate.py` cannot see because it only
ever asks whether the author's names are GONE, never whether non-names were spared.
```
positive control `other` 764/1356 article-preceded = 0.56
negative control 100 honorific-confirmed people, highest Inglés 0.26, bulk 0.00-0.06
FLAGGED 130 of 946 surfaces · 1,616 instances · 0.162% of corpus words
```
`African`, `Chinese`, `Basques`, `Republican`, `Communist`, `X-ray`, `Coca-Cola`, `Ritz`,
`Prado`, `Cezanne` were all renamed into invented proper nouns. **Some flags are correct
renames** — `the Widow`, `the Informer` are genuine Hemingway epithet-names — so every hit is
reported for reading, never auto-removed. Plus **16 bare initials in the map**, `C` at 274
occurrences: the same class as the `G` caught by hand about to be renamed 248 times.
At 0.162% of words this did not block the ship. It is the thing to fix first if a corpus
rebuild ever happens.
## Artefacts
`gx10:~/lv-hemingway/` (corpus-clean, corpus-renamed, beats-hemingway-60.json + sidecar,
eval-hemingway.sh, voice-prep.py, eval.log), `gx10:~/r49-runs/hemingway-4b-pairs-3ep/`
(54 checkpoints kept), `gx10:~/r49-runs/hemingway-eval/` (three arms × 240 generations,
memorization.txt, voice_distance.txt, score.*.txt).
`fv-ml1:/tank/aimodels/voice-adapters/lv-hemingway-4b-v1/` (adapter + a README carrying the
axis-B caveat, so it cannot be read as clean by anyone who finds the adapter without this).
Commits `0bb4938` `051b99e` `5e66114` `2e9b118`.
Related: [[2026-09-17-lv-bronte-gate]], [[2026-09-16-lv-hemingway-corpus]],
[[2026-09-16-lv-voices-line]], [[2026-09-16-voices-seat-lora]],
[[2026-09-17-beat-contamination-leak]].
@@ -1,3 +0,0 @@
# `[2026-09-17]` lv-krakauer D1 built — 126 units, 422,880 words — and its name guard caught three defects nothing else would have reported.
⭐ **lv-krakauer D1 built — 126 units, 422,880 words — and its name guard caught three defects nothing else would have reported.** Back matter searched only the LAST unit while the apparatus sat in unit 37 of 41; relying on the splitter to drop front matter failed because the ebook TOC sits above the author's note and gave it a `Chapter Thirty-Two` to start on; and **zero was the wrong bar** — 2 survivors are Krakauer writing about his own father in Into the Wild's autobiographical chapters, so the allowance is pinned at 2 with every survivor printed. ⚠ Both strips are windowed in the OPPOSITE direction from McCarthy's, because Krakauer's `ALSO BY`/`Copyright`/`About the Author` sit at 0.0–0.6% of the file. Commit `4be0630`.
@@ -1,3 +0,0 @@
# `[2026-09-17]` lv-mccarthy D1 built — 167 units, 584,756 words — and the whole job was protecting a style that reads as damage.
⭐ **lv-mccarthy D1 built — 167 units, 584,756 words — and the whole job was protecting a style that reads as damage.** 0.0 quote marks per 10k (Hemingway 838), `dont`/`aint`/`wont`. The builder runs NO typography normalisation and asserts the quote density afterwards. Two truncated catalogue rows dropped for complete mobi siblings; all 15 containment pairs measured (worst 0.10%); back matter in 4 of 6 works carried the author's name 26 times → 0. ⚠ The back-matter strip runs BEFORE the split here — Blood Meridian and The Crossing end with a dumped TOC of bare roman numerals, the exact shape of a chapter marker. Commit `f3bf3ca`.
@@ -1,3 +0,0 @@
# `[2026-09-17]` lv-mccarthy's D1→D3 chain was RECOVERED, not remembered — there was no runbook and the commands went over non-interactive ssh, so no history survived.
**lv-mccarthy's D1→D3 chain was RECOVERED, not remembered — there was no runbook and the commands went over non-interactive ssh, so no history survived.** Rebuilt candidates and matched sha256 against the artifacts on disk: 6 works, the entity map, the final map and all 36 copy files byte-identical. Now pinned in `scripts/mccarthy-corpus/RUNBOOK.md` with every deviation. ⚠ **D1 must run on nh3-dev** (the builder reads the kvasir catalogue by absolute path); the prior "on gx10" note is true of D2 onward only. ⚠ No phrase map exists for this corpus, so the gate's phrase audit never ran — Yarros and Brontë both had one.
@@ -1,159 +0,0 @@
# `[2026-09-17]` lv-mccarthy D1→D3 — built, gated, and every stage caught a defect in the stage before it
**`~/lv-mccarthy/` on pfi-gx10.** `corpus-clean/` (167 units, 584,716 words),
`corpus-renamed/` (6 copies, 1,002 records), `scripts/`. Commits `705fa3a` `f3bf3ca`
`0fa68cb` `5aa10bf` `5ddb047`.
```
leak gate 0 of 75 renameable and 0 of 37 sub-threshold survive in any copy
positive control 108/108 surfaces found in the unrenamed source
negative control nonce absent from both trees
```
## The shared splitter: choose by SIZE, not by count
`scripts/r49-corpus/split_units.py`. The inherited rule was "most units above a floor", which
is wrong for any book whose markers are PARTS:
```
Cities of the Plain 4 roman marks -> 4 units, median 22,312w
The Crossing 4 roman marks -> 4 units, median 37,310w
```
Four beats one, so it won, and the old guard only fired at exactly one unit. Now: a mode
qualifies only if its median unit is inside **[600, 12000]** AND no unit holds half the work;
among qualifying modes **priority** breaks the tie (contents > chapter-word > roman >
bare-numeral > caps-title). Works with no divisions fall back to **paragraph-block sections**.
⭐⭐ **The first version of that rule was WORSE than what it replaced, and a control caught
it.** Scoring by "median closest to target" chose `caps-title` over the real chapters of
Hemingway's *True at First Light*:
```
bare-numeral 20 units median 5,337w max 11,155 <- the book's own chapters
caps-title 6 units median 777w max 113,886 <- median looked BETTER
```
Five stray all-caps lines gave five tiny units beside **one holding 97% of the book**. A median
cannot see that distribution; a max bound can. Controls green both ways afterwards: 8/10
Hemingway works reproduce the shipped mode and count exactly, and 40,000 words with no blank
lines returns **1 unit** rather than fabricating sections.
⚠ The Hemingway builder is deliberately NOT repointed at this module — its corpus is shipped
and its sha is pinned by a live adapter.
## D1: the job was protecting a style that reads as damage
```
quote marks 0.0 per 10k (Hemingway 838)
apostrophes 123 per 10k (Hemingway 241) `dont` `aint` `wont` `didnt`
```
⚠⚠ **`repair_typography.py` MUST NOT be run on this corpus.** It normalises "toward what the
text does" and would put the quotation marks back. The builder runs no normalisation and then
**asserts** the quote density, so a future well-meaning change fails the build.
⚠ **AND IT MAKES THE VOICE GATE EASY TO PASS FOR THE WRONG REASON.** `voice_distance.py` is
Burrows's Delta over CHARACTER BIGRAMS. An adapter that learns only "emit no quotation marks"
moves delta_cb a long way without having learned a sentence. **Pre-register a
punctuation-normalised secondary read before gating lv-mccarthy.** Tracked in the builder
docstring, commit `f3bf3ca`.
Also: two truncated catalogue rows dropped for complete mobi siblings; all 15 cross-work
containment pairs measured (worst **0.10%**); back matter in 4 of 6 works carrying the author's
name 26 times → **0**; alphabet re-derived at 1,411 non-ASCII letters across 14 Spanish forms.
⚠ The back-matter strip runs **BEFORE** the split for McCarthy, inverting the Hemingway order:
Blood Meridian and The Crossing end with a dumped table of contents made of bare roman numerals
on their own lines — the exact shape of a chapter marker.
## D2 caught a D1 defect: three small-caps manglings
The entity map returned `E`, `H`, `T`, `K` as renameable entities with 17–33 capitalised
occurrences each — the `G` class from Hemingway, where `G` was about to be renamed to a surname
248 times. Reading them showed the extractor mangled small-caps openings three ways:
```
1. SPLIT INITIAL `T HE HOUSE was built` -> `The house was built` 32 cases
2. UNMARKED RUN `THEY STOOD in the doorway` -> `They stood in the doorway` 88 cases
3. LOST INITIAL `HE CANDLEFLAME` -> `THE CANDLEFLAME` 1 case
```
Rule 1 requires a FOLLOWING all-caps word, so `A TV was playing` and `A Mexican was changing`
are untouched. Rule 2's `[a-z]` lookahead is what makes it safe — a genuine shout or sign is
not followed mid-sentence by lowercase. All 23 distinct first words of the 88 were checked.
⚠⚠ **A fourth "fix" was nearly shipped that would have CORRUPTED the text.** `HEY RODE` →
`THEY RODE` looked right from a survey of the BUILT corpus. The raw master has `THEY RODE`
intact, twice — `HEY RODE` matched as a SUBSTRING, and the unanchored replace produced
`TTHEY RODE`, which rule 2 then lowercased to `Tthey rode`. Caught by the count assertion
(expected 1, replaced 2) and settled by reading the master. ⚠ My first corruption check also
missed it, searching for `TTHEY` when the pipeline had already lowercased it — **check the
shape the pipeline emits, not the shape you imagined.**
## D2's own gates: and `audit_stoplist` was scanning its own rationale
⚠⚠ **A defect in `audit_stoplist.py`, latent for every corpus before this one.** It built its
surface set from every list value in the stoplist JSON — including `_why`, which by convention
is a LIST OF PROSE LINES. Its empty separator line matched the honorific pattern **139 times**,
printing a flag with no surface name above the one real catch. Now skips `_`-prefixed keys.
That real catch was a contradiction **inside my own file**: `Franklin` sat in the geography list
(the old name for El Paso) while the same file's note recorded *"I'm here to see Mr Franklin"*,
a lawyer in All the Pretty Horses. A second self-inflicted one: a speculative A–Z fragments list
stoplisted `I` and `A`, and `Sir I dont think I can do that` duly tripped the audit. It is now
the four letters actually measured as entities.
Everything ambiguous was read in context: **Socorro is the ranch cook, not the New Mexico
town**; Niño, Keno and Redbo are HORSES (renameable, the `Inglés` precedent); Yaqui and Gilenos
are real peoples; Hashknives is a real cattle outfit; Hearst, Trias, Huerta and Madero are real
historical figures on the page under their own names.
Final: 123 map surfaces, 124-surface stoplist, `entities.py` 27/27 controls, both audits PASS.
## The human gender pass is an auditable file
The honorific/window resolver scored **21 correct / 3 held / 1 WRONG** against a 26-name
control; the base-rate proximity resolver built for Hemingway scored 18/6/1 and **its own guard
correctly REFUSED to write**. So the incumbent stands and four entries are fixed by hand in
`gender_overrides_mccarthy.json`, each carrying its evidence.
⚠ All four are female and all four look male-dominated in raw counts, because this corpus runs
**29,144 male pronouns to 5,036 female — a base rate of 85.3% male**. Carla Jean Moss at
31m/21f would be 44m/8f at that rate; 21 against an expected 8 is decisive. Same arithmetic that
recovered Pilar and Brett on Hemingway. Alfonsa was in the control and is correctly absent from
the map at 4 occurrences, below the min-count floor — an error in the control, not the pipeline.
`apply_gender_overrides.py` refuses twice: a name absent from the map is an error rather than a
silent no-op, and overruling a gender the detector holds needs an explicit `"correcting": true`
so it cannot look like filling a held entity in a diff.
## D3: three calls, and the holdout fix that matters most
1. **`--scope corpus`**, not the per-work default. Nine surfaces appear in more than one work —
Parham (The Crossing + Cities of the Plain), Grady and Cole (All the Pretty Horses + Cities
of the Plain), Socorro, Héctor. A per-work map gives John Grady a different invented name in
each novel, turning one character into two.
2. **A new `mccarthy` preset.** Hemingway's romance pool carries `it_IT` and `fr_FR` for his
Italian and French casts; McCarthy writes neither language. `en_GB` goes for the same reason.
`en_US` + `es_MX`/`es_ES` at an even share.
3. **`--min-cap 5` to match the entity map's floor.** The first gate run FAILED with 45
survivors: `entities.py` admits cap ≥ 5 while `rename.py` renamed only cap ≥ 8, so every
entity between sat in the map, was never renamed, and counted as a leak. Hemingway never hit
it because its map had `sub_threshold_total: 0`.
⭐ **`--holdout-chapter` NOW TAKES A LIST.** The val split is one chapter index per work, so its
SIZE is set by how many WORKS a corpus has, not how many words:
```
Hemingway 10 works -> 9 val units -> 36,563 words/copy -> gate DECISIVE
Brontë 4 works -> 4 val units -> 17,043 words/copy -> gate MARGINAL
McCarthy 6 works -> 6 val units -> ~18,000 would have been Brontë's end
```
Holding out chapters **7 and 17** gives **11 units and 40,653 words per copy — larger than
Hemingway's** — for 7% of the corpus, on a corpus 40% smaller than his. No amount of corpus size
fixes a val split that scales with work count.
Related: [[2026-09-17-lv-hemingway-gate]], [[2026-09-17-mccarthy-krakauer-d1]],
[[2026-09-17-lv-bronte-gate]].
@@ -1,104 +0,0 @@
# `[2026-09-17]` The leak gate passed with five protagonist names still in every copy
Found during D4 pre-flight, three stages downstream of where it happened. Commit `c559664`.
```
leak gate, 2026-09-17 morning 0 of 75 renameable, 0 of 37 sub-threshold, both controls green
actually present, all 6 copies Bell x2 Chigurh x3 Moss x2 Toadvine x4 Glanton x2
```
## The mechanism
`leak_gate.py` scans `\b(Surface)\b`. **A character inserted inside a name defeats that
pattern outright**, so a mangled occurrence is not merely unrepaired — it is *unrenameable*
by `rename.py` and *unreportable* by the gate, and the gate prints a clean zero over it.
Two extraction artifacts produce exactly that:
```
B ell C higurh M oss T oadvine a small-caps drop cap kept as its own token
Toad-vine Glan-ton a print line-break hyphen kept by the extractor
```
⭐ **Every VISIBLE occurrence had been renamed correctly** — exact-match survivors were 0,
as the gate said. That is what makes this residue invisible to a spot-read: the names are
gone everywhere you look. `Bell` sits in the entity map at 147 capitals, `Glanton` at 365.
This is the third member of a family. lv-bronte's was `_Antigua_` (`_` is a word character,
so `\bAntigua\b` cannot match inside it), found by hand in 2026-09-16 and never generalised.
**The generalisation is the point: any separator inside a name blinds a word-boundary scan.**
## Fixed at three levels, and all three must stay
1. **`build_corpus_mccarthy.py` rules 4 and 5** repair the source text — 32 split initials
with a *lowercase* remainder (rule 1 requires a following ALL-CAPS word and DROPCAP
requires two, so this is the class both leave behind), 5 hyphen-split names by name.
Both carry expected counts so a master change fails the build.
⚠ **Rule 4's letter class is consonants only.** `I` opens **1,966** paragraphs (the
pronoun), `A` opens 143 (the article), `Y` opens 32 (Spanish *y*). Folding any of them
would corrupt 2,141 lines to fix 32 — the same `I`/`A` trap that bit `audit_stoplist.py`.
2. **`leak_gate.py` runs a separator-tolerant pass every time**, with its own positive and
negative controls, and **it fails the gate**. Validated against the pre-fix tree: reports
all five surfaces, exits 1.
3. The exact-match passes are untouched, so the old verdict is reproduced alongside the new.
⚠ **The fragment filter is what makes the new pass usable.** A naive separator-tolerant
scan is dominated by false positives — on Hemingway it returns 21 hits of which **18 are
ordinary text** (`God damn` for the surface `Goddamn` ×14, plus `I run`, `On an`, `Do me`,
`Si le`). The discriminator, with no dictionary: in a genuine split at least one FRAGMENT
is not a word of this corpus. `God` and `damn` occur constantly; `Primi`, `tivo`, `ell`,
`higurh`, `Toad` do not. That one test cleared all 18 and kept all 3 real ones.
⚠ **My first negative control could not pass.** It planted the split nonce in its own probe
text and then asserted the nonce was absent — an alarm wired to itself, failing on every
run. It now hunts the split nonce in the *real* copies. A control that cannot pass is not a
control.
## The shipped corpora, checked with the committed instrument
Re-derived with the COMMITTED gate, not a scratch probe:
```
lv-bronte GATE PASSED 0 separator-split survivors (9.7 s)
lv-hemingway GATE FAILED Pasionaria, Primitivo, Chicote (34.6 s)
1 occurrence each per copy, in all 6 copies — SHIPPED and LIVE
```
⚠ The first version of this scan was **too slow to run** on Hemingway — per-surface scanning
is O(surfaces x copies x corpus) and 881 surfaces x 10 copies was still going at 5 minutes
when it was killed. Rebuilt as one alternation pass, same trick `scan()` already used: 35 s,
identical verdict and identical hit counts on both McCarthy trees. **A gate too slow to run
is not a gate.**
**Operator call outstanding** on whether 3 names in a 958k-word corpus warrant re-gating and
retraining a live adapter. Not acted on.
## The chain was recovered, not remembered — and is now written down
There was **no McCarthy runbook**, and the D1→D3 session issued its commands over
non-interactive ssh so no shell history survived. The chain was recovered by rebuilding
candidates and matching sha256 against the artifacts on disk, then pinned:
```
D1 build_corpus_mccarthy.py 6 works byte-identical
D2 entities.py --min-count 5 --fold-clitics --drop-acronyms --min-mid-ratio 0.2 --min-mid 2
D2c apply_gender_overrides.py entities-final.json byte-identical
D3 rename.py --preset mccarthy --scope corpus --min-cap 5 --copies 6 --seed 4919
--holdout-chapter 7 17 all 36 copy files byte-identical
```
⚠ `--min-mid-ratio` is what keeps `Yeah`/`Buenas`/`Shh`/`Sí` out of the map. The map is
**insensitive** to it: any value in [0.05, 0.3] with `--min-mid` 1 or 2 reproduces byte-for-
byte; `--min-mid 3` does not. The original values are unrecoverable and it does not matter —
which is worth saying, because an exact-looking recipe that was never pinned invites a
false claim of reproduction. Full recipe and every deviation: `scripts/mccarthy-corpus/RUNBOOK.md`.
⚠ **D1 must run on nh3-dev** — the builder reads the kvasir catalogue by absolute path and
gx10 has no copy. The previous session's "on gx10" note is true of D2 onward only.
⚠ **No phrase map exists for this corpus**, so the gate's phrase audit does not run at all.
Yarros and Brontë both had one. Not closed.
Rollback: `~/lv-mccarthy/corpus-{clean,renamed}.pre-splitfix` on gx10.
Related: [[2026-09-17-mccarthy-d1-d3]], [[2026-09-17-beat-contamination-leak]],
[[2026-09-17-lv-hemingway-gate]], [[2026-09-17-lv-bronte-gate]].
@@ -1,3 +0,0 @@
# `[2026-09-17]` Measured and DELIBERATELY not changed, three of them.
**Measured and DELIBERATELY not changed, three of them.** The oversize-passage drop is 13.9% of McCarthy's train words, between Hemingway's 10.0% and the shipped Brontë's 18.3% — in range, no change. `BEAT_PROMPT` asserts the passage is first-person and McCarthy is third; measured inert (**0** narrator-retries against Hemingway's 615 of 7,094), so the prompt was left alone. Blood Meridian's 131 dash-separated chapter-argument paragraphs DID warrant a change and `--drop-leading-heading` now eats them (0 in every other work of all three corpora).
@@ -1,69 +0,0 @@
# `[2026-09-17]` Which voices earn a training seat next — measured against the catalogue, not chosen by taste
Method: rank every author in the kvasir catalogue by **usable extracted** works, then apply the
selection criterion the lv-krakauer parking established — *does the author have a voice*, asked
before any corpus work, and specifically **does that voice live where the instrument looks**.
`voice_distance.py` is Burrows's Delta over CHARACTER BIGRAMS, so it sees function-word morphology,
punctuation and sentence rhythm. A writer whose distinction is plot, research or subject matter is
invisible to it — an adapter cannot carry that, and the gate cannot measure it.
⚠ `triage.length` is in **CHARACTERS**, ~5.2 chars/word calibrated against builds we did ourselves
(The Crossing mobi 777,420 chars = our measured 149,985 words). Dedup by title taking the max across
formats, and floor at 100,000 chars — that is what excludes the `accepted`-but-truncated rows
(Blood Meridian epub at 6,031 chars beside the mobi's 623,849).
## ⭐ The size ranking INVERTS the voice ranking at the top
```
Stephen King 76 works 12,133,529 w <- biggest, and NOT a candidate
Agatha Christie 72 5,451,377 <- second biggest, the Krakauer case exactly
Terry Pratchett 52 4,821,474
Georgette Heyer 30 3,416,867
Graham Greene 45 3,037,425
William Faulkner 25 2,981,183 <- the pick
```
Christie is the whole lesson in one row: a superb writer whose genius is plot architecture, in prose
deliberately kept transparent. Nothing for a char-bigram Delta to grip. King is the softer version —
distinctive in pacing and brand-name texture, not in syntax.
## The three that clear both bars
**1. William Faulkner — 25 catalogue rows, ~15 pure novels, ~1.6M words.**
*The voice in one sentence:* sentences that defer their main clause through stacked subordination
and coined compounds until the reader is held inside a single unbroken perception.
About as char-bigram-legible as English gets — the voice IS the clause-joining morphology and the
`and`/`which`/`that` density. ⭐ **And he is McCarthy's stylistic ancestor, which is the real
argument:** the Brontë gate record states the frozen adjudication needs "a control-author panel (to
place an absolute band and a hard-negative sister)" and notes we have none. Faulkner beside McCarthy
makes each the other's hard negative — a METHOD upgrade, not just another roster entry.
⚠ Messiest corpus of the three: a 446k-word `Snopes: The Hamlet, The Town, The Mansion` omnibus
duplicates novels also present individually, and `Three Famous Short Novels` overlaps it again. That
is the Hemingway trap (169,759 words of measured 90-96% collection duplication) — containment pass
before anything else.
**2. Toni Morrison — 13 rows, 11 novels after pruning, ~818k words.**
*The voice:* free-indirect discourse sliding between narrator and character mid-sentence, carried on
incantatory repetition and deliberate fragments.
Cleanest corpus shape on the list: **11 novels → 11 val units, beating Hemingway's 10.** Val units
scale with WORK COUNT, which is the structural reason Brontë's voice axis came back underpowered at
4 with no cheap fix. ⚠ Drop `Burn This Book` (anthology she edited) and `Playing in the Dark`
(criticism) — same reason Krakauer's reporting does not transfer.
**3. Raymond Chandler — 9 rows, 7 novels + a 409k short-story omnibus, ~970k words.**
*The voice:* clipped first-person declaratives that periodically detonate into one baroque simile,
with dialogue carrying most of the scene.
Fills the register gap nobody else fills — **first-person hardboiled**; the line has no first-person
male narrator at all. Corpus is almost exactly Hemingway-sized (997k vs 958k), which was the
decisive gate. ⚠ Drop `Essays and Reviews` — non-fiction.
## Held, and why
**Conrad** (31 works, 2.4M) is a genuine tier-1.5 if a fourth is wanted. **Melville** (10, 1.9M) has
a superb voice but a mixed-register corpus — the cetology chapters are a different book from the
narrative. **Austen** (12, 1.18M) is worth noting because Burrows's Delta was developed on her, so
the instrument is known to resolve her. The romantasy cluster is a separate question entirely —
see [[2026-09-17-romantasy-register-measured]].
Related: [[2026-09-17-mccarthy-split-name-leak]], [[2026-09-17-lv-bronte-gate]],
[[2026-09-17-lv-hemingway-gate]].
@@ -1,71 +0,0 @@
# `[2026-09-17]` Romantasy measured as a register — it is real, we already took its best voice, and the obvious next pick is its worst
Prompted by the operator pushing back on a one-clause dismissal of the lane as "depth behind
Yarros". The dismissal was taste; this is a measurement, on the gate's own instrument.
**Method.** Char-bigram Burrows's Delta, the same measure `voice_distance.py` gates on. ~120k words
per author, sampled from the MIDDLE quartile of each author's largest works (front and back matter
are not the voice), equalised so a bigger sample is not a different measurement. 400 most-frequent
bigrams as the feature set, z-scored over 4,000-word chunks pooled across all authors.
**Controls first, because a between-author number without a within-author floor is unfalsifiable.**
```
A-vs-A floor (two halves of the SAME author)
Yarros 0.285 Maas 0.298 Armentrout 0.314 St. Clair 0.322 Cole 0.338
Kenyon 0.375 Reyne 0.391
McCarthy 0.303 Morrison 0.327 Brontë 0.209 Hemingway 0.454 <- worst, used as the bar
positive controls (known-distinct pairs — the instrument must separate these)
Yarros vs McCarthy 0.862 1.9x
Hemingway vs Brontë 0.773 1.7x
McCarthy vs Morrison 0.675 1.5x
Hemingway vs McCarthy 0.655 1.4x
romantasy, all 21 pairs median 0.537 1.2x floor (range 0.465 - 0.674)
```
**The register is real but tight.** 1.2x floor against controls at 1.4-1.9x. Only one pair falls to
1.0x, so it is not seven names for one voice.
⚠ **Sensitivity floor, stated because a result without one is unfalsifiable.** The 0.454 bar is
Hemingway's, inflated by his own heterogeneous corpus (1920s-1960s, novels + stories + posthumous).
Against the romantasy authors' OWN floors (~0.34) the same pairs read ~1.6x — control-grade. The
truth sits between those readings and **this method cannot split it finer**. One sample per pair, no
repeat draws: read the rank ordering as indicative, do not read small gaps at all.
## Two findings that survive either floor reading
⭐ **Yarros is the cluster OUTLIER, not a typical member.** Four of the five largest distances in the
matrix involve her — Yarros-Kenyon 0.674, Yarros-St. Clair 0.644, Yarros-Reyne 0.637, Yarros-Maas
0.567. **We already trained the most distinctive romantasy voice we hold**, so a second seat in the
lane buys measurably less than the first did. That is the actual answer to "what about romantasy".
⭐ **Maas is the centroid, so the obvious commercial pick is the least distinctive.** Maas-Reyne
0.465 and Maas-Cole 0.470 are the two SMALLEST distances in the whole matrix. She is the biggest
name available (922k words) and measurably the most generic of the seven in char-bigram terms.
Picking by sales rank picks the worst adapter.
## If the lane gets a second seat it is Kenyon
Furthest from the shipped Yarros (0.674), so it adds the most new signal — **and 27 works means 27
val units, the best-powered gate the line could build** (Hemingway 10, McCarthy 6, Brontë 4, where 4
is the documented structural cause of an underpowered voice axis with no cheap fix).
⚠ Two costs: the 27 are one series (Dark-Hunter), so the shared proper-noun space makes
`--scope corpus` mandatory rather than optional; and a "Dark Hunter - The Dark Hunter Complete"
omnibus sits in the catalogue rows, so the containment pass runs first.
**Corpus shapes for the lane** (works ≥100k chars, deduped by title):
```
Sherrilyn Kenyon 27 2,368,396 w Scarlett St. Clair 11 1,133,067
Opal Reyne 14 2,466,307 Kresley Cole 10 1,037,914
Jennifer Armentrout 6 1,179,953 Sarah J. Maas 5 922,711
Rebecca Yarros 5 820,425 <- SHIPPED on this
```
⭐ Worth noting for any future bar-setting: **Yarros shipped on 5 works / 820k words.** The corpus
bar is lower than it looks.
Instrument: `scratchpad/regdist.py` (screening tool, not the gate).
Related: [[2026-09-17-next-voice-seats]], [[2026-09-16-lv-voices-line]], [[2026-09-17-lv-bronte-gate]].
@@ -1,3 +0,0 @@
# `[2026-09-17]` `servers/fv-ml1/ssh-target` was bare `10.251.50.54`, so `deploy-stack.sh` connected as `lkraven` and could not write the infra-ops-owned `/opt/docker/compose/`
**`servers/fv-ml1/ssh-target` was bare `10.251.50.54`, so `deploy-stack.sh` connected as `lkraven` and could not write the infra-ops-owned `/opt/docker/compose/`** — and lkraven's sudo on fv-ml1 needs a password, so `DEPLOY_SUDO=1` failed too. Now `infra-ops@10.251.50.54`; `--validate-only` still clean, deploy works. ⚠ Other hosts' `ssh-target` files may carry the same gap — a read-only refresh works as either user, so the fault only surfaces on a deploy.
@@ -1,3 +0,0 @@
# `[2026-09-17]` The beat-contamination leak IS present in Hemingway — 70 of 7,094 train beats (0.96%), 0 of 200 val.
⭐ **The beat-contamination leak IS present in Hemingway — 70 of 7,094 train beats (0.96%), 0 of 200 val.** `scripts/r49-corpus/audit_pairs_sourcenames.py` closes the blind spot `leak_gate.py` has by construction (it reads the corpus and the renamed copies, never the generated beats). Cross-validated on real data: the fixed Brontë pairs return 0 of 3,858 and `pairs-full.CONTAMINATED` returns 15 of 792 = 1.89% with the recorded names. `--filter-out` yields a verified-clean 7,024-pair set in one command; the retrain is the operator's call. **The val split being clean is why the gate could run at all.**
@@ -1,3 +0,0 @@
# `[2026-09-17]` The `mccarthy` register names the punctuation ON PURPOSE, and that is a gate-design call made before any McCarthy number existed.
⭐ **The `mccarthy` register names the punctuation ON PURPOSE, and that is a gate-design call made before any McCarthy number existed.** `eval-*.sh` drives the base control arm with the SAME system prompt via `--system-from`, and `voice_distance.py` is Burrows's Delta over character bigrams — so a tic left OUT of the register is a cheap win only the adapter can take, on a corpus measuring 0.0 quote marks per 10k against Hemingway's 838. Stating them hands them to the control too. Cost stated up front: the voice axis gets harder, and McCarthy's 276-passage val split (against Brontë's 44) is why that trade is affordable here and was not there.
@@ -1,3 +0,0 @@
# `[2026-09-17]` The SHIPPED lv-bronte adapter emits mid-sentence line breaks at 12.46 per 1k chars, and nothing downstream looks for it.
⭐⭐ **The SHIPPED lv-bronte adapter emits mid-sentence line breaks at 12.46 per 1k chars, and nothing downstream looks for it.** Its corpus is 100% hard-wrapped at ~68 chars (Gutenberg plain text) and the wrap transfers: base control 0.00, ckpt475 (shipped) 12.46, ckpt925 11.79, every Hemingway arm 0.00 on a 0%-wrapped corpus. Both controls fire. `score_beats.py` passed Brontë's damage axis anyway. McCarthy is the MIXED case — The Road wrapped, the other five works not — which is worse to learn than either pure one, so `build_sft_pairs.py --reflow-hard-wraps` (DEFECT 4) fixes it at pair time, off by default. ⚠ The obvious fix, joining every interior newline, CORRUPTS 46 two-speaker exchanges whose blank line was lost — and unmarked dialogue is the one thing this adapter exists to learn. The rule splits on sentence-final punctuation and takes the cheaper error deliberately.
@@ -1,3 +0,0 @@
# `[2026-09-17]` The two-epoch recipe is now 0 for 2 and should stop being carried forward.
**The two-epoch recipe is now 0 for 2 and should stop being carried forward.** Hemingway's eval minimum is step 1750, but step 850 is +0.0040 against a 0.0044 median neighbour jitter — three checkpoints inside one jitter — and 850 won every resolving axis (2.3x tighter seed spread, lower memorisation, less ran-on). Same outcome as Brontë. What IS robust on this schedule is the epoch-3 collapse: +0.0762 = **17.4x jitter**.
@@ -1,3 +0,0 @@
# `[2026-09-17]` The v2 voice floor is now PAIRWISE, and it retroactively passes lv-bronte.
⭐⭐ **The v2 voice floor is now PAIRWISE, and it retroactively passes lv-bronte.** lv-bronte's ckpt475 shipped as a voice-axis FAILURE at +0.193 against a 0.251 floor contributed entirely by ckpt925 — a third arm nobody was shipping, on one outlier seed. Scored against the arm it was actually compared to the floor is 0.091 and it clears at **2.1x**. The rule was changed **prospectively**, pre-registered for lv-hemingway before any Hemingway number existed, on an argument independent of the answer: the sampling variability of a difference A−B depends on A and B, never on a third arm C. The previous session found the defect and deliberately declined to exploit it; this follows from fixing it. lv-hemingway passes under **both** rules, so its verdict does not lean on the change. Caveats amended append-only in the compose, the NFS README and the gate record. Commits `0bb4938` `2e9b118`.
@@ -1,3 +0,0 @@
# `[2026-09-17]` `triage_disposition = 'accepted'` in the Kvasir catalogue does NOT mean the extraction succeeded.
⚠ **`triage_disposition = 'accepted'` in the Kvasir catalogue does NOT mean the extraction succeeded.** Blood Meridian's epub row holds 1,167 words of a 117,000-word book, The Crossing's 222 of 150,000 — real prose, real titles, accepted. Faulkner's *The Mansion* is 39 words. `near_dup_pairs` holds ONE row in the entire 1,284-work library and is blind to a fragment beside its full sibling. **Word-count every master before trusting a row**, and note that word count alone cannot tell a truncated novel from a legitimately short work.
@@ -18,3 +18,9 @@
2. Free ~+1.1–1.5 GB on GPU 0 by trimming the `vllm-gen-small` util. ⚠ MEASURE the resulting free memory; util does not predict resident VRAM.
3. Cut over with the old seat kept as the rollback.
4. Re-measure live on GPU 0.
**DONE 2026-10-01 ~0126 PT (infra-hermes; `stacks/parakeet-nemo`, de6ea32), infra-ops audit PASSED 0137.**
- p50 on GPU 0: 33 / 36 / 42 / 71 ms, against 187 / 308 / 626. WER: clean 1.965, other 3.026.
- A 714 s file returns 200 in 0.85 s (360 s windows). The steady state is 3,582 MiB, and GPU 0 Free is 385.
- gen-small util 0.48 → 0.36: it took three boots and ~34 min of downtime at midnight, with zero LiteLLM errors. Its `.env` holds 0.33 for the next restart. The KV is byte-pinned and unchanged.
- infra-hermes caught that httptools 0.8.0 emits `HTTP/1.1 200\x00OK`, which LiteLLM rejects; it is pinned to `--http h11`.