memory: snapshot — speech seat live + gen-small OOM incident (mitigated, fix tasked); leftovers deleted; no Scriberr upstream; repos pushed; 32 entries archived
This commit is contained in:
@@ -1,54 +0,0 @@
|
||||
# `[2026-09-16]` Grok token broker — built, then shelved by the transport ruling. Do NOT arm the probe.
|
||||
|
||||
**`services/grok-token-broker/` — seeded, committed, DISARMED, no consumer.** Commits `ebc4dac`
|
||||
`b907a0e` `cf9d167`. ⛔ **Do not arm `probe-rotation`.** This is a finished resting place, not a
|
||||
half-built tool: the gate works and the thing it gated for went away.
|
||||
|
||||
**Operator ruling, relayed by heid:** *"keep the jail stop the a/b"*
|
||||
(`heid dispatch-log/2026-09.jsonl#groa-transport-20260916-operator-keeps-the-jail`, alongside
|
||||
`#groa-transport-ab-20260916-operator-stop`). Gróa dispatches through the read jail;
|
||||
`groa_http_dispatch.py` is a documented fallback with no scheduled use. **Nothing in the fleet
|
||||
wants a renewable xAI session.**
|
||||
|
||||
⭐ **THE CODE-PLAN ENDPOINT EXISTS and I was one message away from telling the operator it did
|
||||
not.** `https://cli-chat-proxy.grok.com/v1` serving **grok-4.6** (500,000 context) and grok-4.5,
|
||||
`agent_type: grok-build-plan`, `auth_method: session`, `api_key`/`env_key`/`api_base_url` all
|
||||
null. ⚠⚠ **It is in `~/.grok/models_cache.json` — the Grok CLI's own config, on nh3-dev.** I had
|
||||
swept heid's repo, the gateway `.env`, the LiteLLM config and Vaultwarden, all correctly, and
|
||||
concluded "does not exist". ⭐ **heid's line, taken: absence from the places you searched is not
|
||||
absence.** The check that separates the two states is a LIVE REQUEST, not a grep.
|
||||
|
||||
⚠⚠ **NOT WRITING `~/.grok/auth.json` IS NECESSARY AND NOT SUFFICIENT.** The refresh grant at
|
||||
`https://auth.x.ai/oauth2/token` may ROTATE the refresh token, and many OIDC providers invalidate
|
||||
the old one SERVER-SIDE. A broker refreshing the same credential kills the CLI login even though
|
||||
it never touches the file. heid's module header reasoned about the WRITE; they amended it to name
|
||||
invalidation, credited. This correction went infra-ops→heid the same day heid's went the other
|
||||
way — **neither of us reaches the right answer alone.**
|
||||
|
||||
⚠⚠ **THE PROBE'S BLAST RADIUS IS BOTH GRÓA TRANSPORTS, which is not visible from the infra side.**
|
||||
`heid/scripts/groa_dispatch.py` builds `argv = ["grok", "-p", prompt, "--cwd", jail, ...]` and
|
||||
shells the CLI, which authenticates from the same `~/.grok/auth.json`. The bwrap in the process
|
||||
table is grok's own Landlock sandbox, not something Heid wraps. **One session, two ways of
|
||||
reaching it** — an invalidating probe takes Gróa down on EVERY path until an interactive re-login.
|
||||
🔴 **I had recommended "run the probe now while the CLI is idle" and withdrew it in writing**;
|
||||
"idle" was a convenient assumption I never checked, on a day that had already taken eight panels.
|
||||
|
||||
**Why the jail won, and it was not performance.** HTTP is faster (~523 s median vs ~890 s),
|
||||
simpler, and arguably SAFER on confinement (no tools, so the 2026-06-10 escape class is
|
||||
structurally impossible). It lost on FAILURE MODE: HTTP fails by returning a fast, confident,
|
||||
well-formatted review that found nothing — indistinguishable from a clean bill. The jail fails by
|
||||
timing out, which you can see. ⚠ **Do NOT quote a per-transport finding rate from this**: heid
|
||||
states the 0/0/0-vs-5/7/3 numbers are confounded with bundle size (the zeros were all huge inline
|
||||
bundles; the one HTTP round at jail-comparable size produced Gróa's leading solo), n=3–4 per cell,
|
||||
no noise floor. Asymmetric-risk argument, **not** a resolved measurement.
|
||||
|
||||
⚠ **Still unmeasured, and it is a billing question:** the jail reaches the coding plan already
|
||||
paid for; the HTTP path reaches the METERED API and its responses carry `cost_in_usd_ticks`.
|
||||
Whether that bills on top of the plan was never part of the ruling. One look at the xAI billing
|
||||
console — **this fleet holds no xAI credential**, so it needs the operator's account access.
|
||||
|
||||
⚠ The coding plan speaks the **Responses API** (`api_backend: "responses"`), not
|
||||
`/chat/completions` — a second, independent obstacle to any LiteLLM alias. Moot while the jail is
|
||||
ruled. heid also found and killed two live instructions in their own persistent-memory telling a
|
||||
fresh session to dispatch `--groa-transport http`; either would have resumed a stopped experiment
|
||||
after a context reset.
|
||||
@@ -1,66 +0,0 @@
|
||||
# `[2026-09-16]` lv-hemingway corpus: half the work was EXCLUSION, and the gate found what a hand count would not
|
||||
|
||||
**994,760 words · 318 units · 6 renamed copies · leak gate PASSED 0 of 941 entities and 0 of 117
|
||||
audited phrases, both controls green.** `~/hemingway-corpus{,-renamed}`, builder
|
||||
`scripts/hemingway-corpus/build_corpus_hemingway.py`, commits `9598d0b` `03b4a3f`.
|
||||
|
||||
⭐ **THE CATALOGUE HOLDS 2,105,679 WORDS AND ROUGHLY HALF MUST NOT BE TRAINED ON.** Operator
|
||||
scoped it to fiction only. Three exclusion passes, each measured or voice-specific:
|
||||
|
||||
1. **Non-fiction, 8 works ~911k words** — By-Line, Dateline: Toronto, Death in the Afternoon,
|
||||
Green Hills of Africa, The Dangerous Summer, the three posthumous "Hemingway on X" anthologies.
|
||||
2. ⭐ **Four story collections, 169,759 words — MEASURED, not assumed.** `Short Stories` is the
|
||||
First Forty-Nine and CONTAINS the others. 8-gram containment of the smaller work: Winner Take
|
||||
Nothing **96.0%**, Snows of Kilimanjaro **95.2%**, Men Without Women **92.9%**, In Our Time
|
||||
**90.6%**. ⚠⚠ **The catalogue's own `near_dup_pairs` table is BLIND to this** — it holds
|
||||
whole-document simhashes (ONE row in the entire 1,284-work library) and this is PARTIAL
|
||||
containment. Whole-document dedup cannot see a collection inside a larger collection.
|
||||
3. ⭐ **`The Torrents of Spring` — excluded for a reason no word count could justify.** It is a
|
||||
deliberate PARODY of Sherwood Anderson: the target author's name on a different author's
|
||||
style, i.e. mislabelled data for a voice adapter.
|
||||
|
||||
⚠ **THE AUTHOR'S OWN NAME WAS IN THE TRAINING TEXT 95 TIMES ACROSS 7 WORKS** — publisher back
|
||||
matter ("Ernest Hemingway was one of America's foremost journalists… died in 1961") riding inside
|
||||
the last unit, because a splitter cuts on headings and nothing follows the final one. **Identical
|
||||
to the Yarros defect; nothing about the source changed to cause it.** Stripping the publisher
|
||||
block left 18, all in `true-at-first-light`, inside a **CAST OF CHARACTERS and SWAHILI GLOSSARY
|
||||
written by Patrick Hemingway** — an editor describing the author's real household. Markers are
|
||||
matched in file order, earliest wins. Now 0.
|
||||
|
||||
⚠ **`G` WAS ABOUT TO BE RENAMED TO A SURNAME, 248 TIMES.** Not a name: the fragment left by
|
||||
`B.G.`, `G.M.`, `G2`, `G3`. Caught by reading surfaces IN CONTEXT, which is the Yarros lesson
|
||||
repeating. Also read in context: `Gran` (fragment of `Gran Sasso`/`Gran Italia`/`Gran Hotel`),
|
||||
`Shamba` (Swahili common noun), and `Inglés` — **kept renameable deliberately**, the gypsies'
|
||||
in-world nickname for Robert Jordan, exactly parallel to Yarros's `Violence`.
|
||||
|
||||
⭐⭐ **THE GENDER RESOLVER HAD TO BE REBUILT AND ITS OWN GATE CAUGHT THE FIRST ATTEMPT.** The
|
||||
inherited one returned **397 male / 20 female** across 1,102 records with Catherine Barkley,
|
||||
Brett Ashley, Pilar, Maria and Mary all held neutral. A plain majority vote over nearby pronouns
|
||||
scored 18 correct but **5 WRONG** against the incumbent's 1 — and **every error was
|
||||
female-read-as-male** (Pilar m=426 f=243, Brett m=249 f=137). The refuse-unless-better guard
|
||||
rejected it, correctly. ⭐ **Cause, measured: the corpus base rate is 34,315 male pronouns to
|
||||
8,699 female, nearly 4:1.** Pilar's "male-dominated" 426:243 is strongly FEMALE against that
|
||||
background. Scoring each name's local mix against the corpus base rate instead of 50:50 gives
|
||||
**18 correct / 11 held / 0 WRONG**, distribution 275m / 120f.
|
||||
`scripts/hemingway-corpus/gender_by_proximity.py`. ⚠ Yarros solved its version with the POV
|
||||
chapter header; Hemingway's editions have none, so that fix does NOT transfer.
|
||||
|
||||
⚠ **THE ALPHABET DOES NOT TRANSFER EITHER: 1,496 non-ASCII letters across 23 forms** against
|
||||
Yarros's 2. Hemingway writes Spanish, French and Italian constantly, so the rename pool needs
|
||||
accents (new `hemingway` preset in `rename.py`). The Yarros ASCII-only conclusion would have
|
||||
stranded every Spanish and Italian name in the cast — which is why F02 says re-derive per corpus.
|
||||
|
||||
**Three source defects the splitter surfaced.** `Islands in the Stream` came out as ONE
|
||||
143k-word record (roman numerals, unhandled). `Short Stories` came out as 5 units then 27,
|
||||
because the edition carries a **SECOND contents listing** and first-occurrence matching resolved
|
||||
31 of 58 titles to an index entry — keeping every occurrence and letting the word floor decide is
|
||||
self-correcting; now 57. ⚠⚠ **And the drop-cap defect is in the HEADINGS here** (`T HE O LD M AN
|
||||
AND THE S EA`), which **INVERTS the Yarros pipeline order: repair must run BEFORE the split**, or
|
||||
the splitter cannot see the headings it needs.
|
||||
|
||||
⭐ **Hemingway needed THREE mapped phrases where Yarros needed 48** (`Gran Maestro`,
|
||||
`Unknown Tongue`, `Sin House`) and a 45-entry allow list — the whole difference being that Yarros
|
||||
invented a world and Hemingway named the real one. Several allow entries were non-obvious and
|
||||
required reading: `Royal Game` is a real colonial-Kenyan legal category, `White Heather` a Scotch
|
||||
brand, `Bwana Game` a job title, `Roman Soldier`/`Wine Seller` stage-direction labels from the
|
||||
one-act play `Today is Friday`.
|
||||
@@ -1,53 +0,0 @@
|
||||
# `[2026-09-16]` The lv-* voice line: Option C proved, lv-yarros shipped, lv-hemingway training
|
||||
|
||||
⭐⭐ **INSTRUCTION-PAIR SFT BEATS RAW-TEXT TRAINING FOR AUTHOR VOICE, AND THE INCUMBENT NEVER
|
||||
CLEARED ITS OWN CONTROL.** Measured n=120 per arm, 30 in-genre beats from HELD-OUT val passages
|
||||
× 4 seeds, all arms re-measured in one session on one box:
|
||||
|
||||
| arm | delta_cb (lower = more Yarros) | vs base control | 8-gram overlap |
|
||||
|---|---|---|---|
|
||||
| pairs 2ep ckpt-1650 | **0.410** | +0.289 ✅ | 0.12 |
|
||||
| pairs 3ep ckpt-1650 (**shipped**) | 0.438 | +0.262 ✅ | **0.09** |
|
||||
| raw-text instruct (incumbent) | 0.558 | +0.141 ❌ **inside the 0.153 floor** | 0.14 |
|
||||
| base-unadapted (control) | 0.700 | — | 0.07 |
|
||||
|
||||
Same-author target 0.463 (held-out Yarros vs itself). ⚠ **The two pair arms are NOT
|
||||
distinguishable on voice** — 0.028 against a 0.153 floor. The 3ep checkpoint was chosen on the
|
||||
axes that ARE resolvable: better held-out fit (2.3126 vs 2.3264), less overshoot (0.06 vs 0.10),
|
||||
and verbatim overlap nearest the never-saw-it control.
|
||||
|
||||
⭐ **THE RECIPE IS TWO EPOCHS ON A THREE-EPOCH SCHEDULE, not three epochs.** Launch `--epochs 3`;
|
||||
the minimum lands at step 1650 **inside epoch two** and epoch three overfits (2.3126 → 2.3882,
|
||||
flat). The entire gain over a 2-epoch run came from the stretched cosine keeping the LR alive —
|
||||
at step 1600 the 3ep run was at 2.9e-05 where the 2ep run had annealed to 2e-07. ⚠⚠ **A
|
||||
resume-and-append-one-epoch is a NO-OP for exactly that reason** (lr 2.3e-09 at step 1670): it
|
||||
must be a fresh run with the longer schedule.
|
||||
|
||||
⭐ **THE SAFETY PROPERTY: the model writes the INSTRUCTION, never the RESPONSE.** Every response
|
||||
is real renamed prose; only the beat is machine-written, so voice is inherited rather than
|
||||
synthesised. Memorisation checked with both controls (positive control saturates at 160): the
|
||||
shipped arm sits at 0.09 against a 0.07 never-saw-it baseline and BELOW the raw-text arm's 0.14.
|
||||
|
||||
⚠ **THE v1 DECISION RULE WAS WELL-FORMED AND MEASURED THE WRONG THING**, and the amendment is
|
||||
recorded in `scripts/yarros-corpus/score_beats.py` with v1 retained verbatim. It gated on
|
||||
in-band / on-beat / ran-on — and **base-unadapted scores in-band 0.96**. Instruction-following is
|
||||
something Qwen3-4B-Instruct ships with, so those axes detect only DAMAGE, never the benefit an
|
||||
adapter exists to buy. v2 gates on voice (delta_cb vs control beyond the floor) + not-copied
|
||||
(8-gram overlap near control) + no-damage (overshoot). ⚠ on-beat's −0.27 was outside the floor
|
||||
and is dropped from the gate, **not explained away** — the keyword proxy punishes prose that
|
||||
DRAMATISES "she mocks him" rather than echoing the word, but three read samples is an anecdote.
|
||||
|
||||
⭐ **A 5-BEAT FIXTURE HAD A NOISE FLOOR OF 0.800 AND MANUFACTURED A +0.45 RESULT.** At n=20 the
|
||||
pilot looked like a clear in-band win; at n=120 the same gap was +0.08, inside a 0.233 floor.
|
||||
One sample moves a rate by 0.2 when there are five. The 30-beat in-genre fixture (built from
|
||||
held-out val pairs, `~/beats-yarros-30.json`) is the instrument; the Brontë stray-dog/kitten
|
||||
fixture was also the wrong GENRE — "He licked her clean" came back as explicit sex.
|
||||
|
||||
⚠ **THE HARNESS TRUNCATES AT THE FIRST BLANK LINE and that surface reported the pair arm as
|
||||
"19 words, off-beat 0.10"** when the untruncated output was 90–132 words with the beat rendered
|
||||
in a later block. `score_beats.py --metric-source raw|paragraph` keeps both views and the verdict
|
||||
names which it used. Same family as `feedback_filters_that_silently_narrow_the_window`.
|
||||
|
||||
**Artefacts.** `scripts/yarros-corpus/{build_sft_pairs,train_pairs_lora,score_beats,
|
||||
memorization_check}.py`; commits `9b3d3c8` `90ed506` `713e83d` `efb7345` `7505124`. Booth
|
||||
(24h TTL) was `http://10.100.10.50:8090/b/babyyarros-beats/` — six beats × four arms, blind-labelled.
|
||||
@@ -1,50 +0,0 @@
|
||||
# `[2026-09-16]` voices-seat: LoRA over merge, measured — and GPU 0 is now full
|
||||
|
||||
**`vllm-voices` live on fv-ml1 GPU 0 :8027**, one Qwen3-4B-Instruct carrier serving
|
||||
`voices-base` plus `lv-<author>` LoRA adapters. `stacks/voices-seat/`, commit `d17bd3d`.
|
||||
|
||||
⭐ **LORA COSTS 24.3% OF DECODE THROUGHPUT AND IT IS WORTH PAYING.** n=30 per arm, interleaved,
|
||||
A-vs-A noise floor **0.1%**: base **143.0 tok/s** median vs adapter **108.2**. The measurement is
|
||||
unusually clean because `--enable-lora` serves BOTH the base name and the adapter name from ONE
|
||||
process — the arm is a per-request field, so no restart, no second seat, no cold-vs-warm confound.
|
||||
Arms were **interleaved rather than blocked** because the card's co-tenants take traffic this seat
|
||||
does not control, and a block design would alias their load onto one arm.
|
||||
|
||||
**Why pay it:** 3 authors cost 8.4 GB as adapters against ~23 GB merged; 6 cost 9.2 vs ~46. On a
|
||||
card with 1.8 GB free afterwards that is the whole argument. If a voice ever lands on a latency
|
||||
path, merge THAT one and serve it separately.
|
||||
|
||||
⭐ **ADAPTER HOT-SWAP IS REAL AND FAST — MEASURED, not read from docs.**
|
||||
`POST /v1/load_lora_adapter` **200 in 0.24 s**, `POST /v1/unload_lora_adapter` **200 in 0.003 s**,
|
||||
VRAM unchanged, container stayed healthy. Proven by performing the `babyyarros`→`lv-yarros`
|
||||
rename through it with no restart. ⚠ **A runtime-loaded adapter is GONE on the next
|
||||
`compose up -d`** unless it is also in `--lora-modules` (which costs a recreate + ~3 min reload).
|
||||
Runtime load is for TRYING a voice; the compose list is what persists. Switching between loaded
|
||||
voices is just the `model` field — **not** a LiteLLM alias; LiteLLM is a thinner layer on top,
|
||||
one alias entry per voice, no new deployment.
|
||||
|
||||
⚠⚠ **`--gpu-memory-utilization` IS A REQUEST AGAINST *TOTAL* VRAM THAT THE CARD MUST ALREADY BE
|
||||
ABLE TO HONOUR — not a share of what is free.** First bring-up REFUSED: *"Free memory on device
|
||||
cuda:0 (11.16/94.97 GiB) is less than desired GPU memory utilization (0.12, 11.4 GiB)"*. Refusing
|
||||
was the right outcome — it protected `cyberprev`, `gen-small` and the Parakeet STT seat rather
|
||||
than squeezing them.
|
||||
|
||||
⭐ **PINNING `--kv-cache-memory` IN BYTES MAKES THE FRACTION PREDICTIVE.** Requested 0.11
|
||||
(10,700 MiB), got **10,740 MiB** resident — a 40 MiB miss on a box where the fraction has been
|
||||
wrong by **8–10 GB in BOTH directions** (cyberprev 0.40→47.1 GB, gen-small 0.48→36.9 GB). Second
|
||||
seat to prove it after `gen-small`. Do not remove the pin.
|
||||
|
||||
⚠ **fv-ml1 GPU 0 is now 96.0 of 97.9 GB.** GPU 1 ~5.7 free, GPU 2 ~2.4, GPU 3 is a HELD RESERVE
|
||||
for a future full-card seat (`flash-next` alone needs 93 of 96 GiB). **There is no room for
|
||||
another seat on fv-ml1 without a placement decision.**
|
||||
|
||||
⚠ **SUPPORT WAS CHECKED, NOT ASSUMED**, per the training playbook's own lesson that LoRA support
|
||||
is per-ARCHITECTURE not per-family: `vllm/model_executor/models/qwen3.py:271` declares
|
||||
`Qwen3ForCausalLM` with `SupportsLoRA` plus `packed_modules_mapping` and `embedding_modules`.
|
||||
**Do not transplant this compose onto an MoE carrier without re-running that grep** — the
|
||||
playbook records a LoRA refusal on a Qwen3 MoE arch.
|
||||
|
||||
**Naming (operator, 2026-09-16):** `lv-<author>` — lv for **lang-voice**, retiring `baby*`, which
|
||||
read fine for one experiment and invites confusion across a family. The adapter NAME is the
|
||||
request's `model` field, so it is the public API of a voice. Historical persistent-memory entries
|
||||
still say BabyYarros/BabyHemingway and were deliberately left as dated records.
|
||||
-3
@@ -1,3 +0,0 @@
|
||||
# `[2026-09-17]` A stoplist entry is an assertion the leak gate can no longer check
|
||||
|
||||
⭐ **A stoplist entry is an assertion the leak gate can no longer check** — stoplisting removes a surface from the entity map, so a wrongly stoplisted CHARACTER is an undetectable leak. Three were wrong on Brontë (Leaven, Pierrot, Samuel); `scripts/r49-corpus/audit_stoplist.py` finds them by honorific and now gates the pipeline. Commit `8bb7686`.
|
||||
@@ -1,3 +0,0 @@
|
||||
# `[2026-09-17]` A unit splitter must choose by SIZE, not by count — the inherited rule silently produced 22,000-word "chapters".
|
||||
|
||||
⭐⭐ **A unit splitter must choose by SIZE, not by count — the inherited rule silently produced 22,000-word "chapters".** `scripts/r49-corpus/split_units.py`: a marker mode qualifies only if its median unit is inside [600, 12000] AND no unit holds half the work; among qualifying modes PRIORITY breaks the tie (contents > chapter-word > roman > bare-numeral > caps-title), and paragraph-block sections are the fallback for works with no divisions. ⭐ **Both rules exist because a control caught them**: scoring by "median closest to target" chose `caps-title` (6 units, one holding **97%** of the book) over True at First Light's real 20 chapters, because a median cannot see that distribution and a max bound can. Positive control: 8/10 Hemingway works reproduce the shipped mode and count exactly. Negative control: 40,000 words with no blank lines → 1 unit, refuses to fabricate divisions. Commit `705fa3a`.
|
||||
-3
@@ -1,3 +0,0 @@
|
||||
# `[2026-09-17]` `audit_entity_map.py` — the rename can DAMAGE the prose and no gate will ever say so.
|
||||
|
||||
⭐ **`audit_entity_map.py` — the rename can DAMAGE the prose and no gate will ever say so.** Mirror of `audit_stoplist.py`: surfaces wrongly held IN the map rather than out of it. 130 of 946 Hemingway surfaces flagged (1,616 instances, 0.162% of words) — `African`, `Chinese`, `X-ray`, `Coca-Cola`, `Ritz`, `Prado` renamed into invented names — plus 16 bare initials incl. `C` at 274 occurrences. Signal is a preceding article; controls derived from the corpus, not hand-picked. Every hit reported for READING: `the Widow` and `the Informer` are genuine epithet-names that should be renamed. Commit `051b99e`.
|
||||
@@ -1,39 +0,0 @@
|
||||
# `[2026-09-17]` The beat-writing model restores the author's real names — a leak the corpus gate structurally cannot see
|
||||
|
||||
⭐⭐ **THE RENAME AND ITS GATE ARE BOTH WORKING, AND THE TRAINING DATA IS STILL CONTAMINATED.**
|
||||
The corpus gate reads the corpus and the renamed copies. **It never reads the generated
|
||||
instruction beats.** Those are written by an LLM that just read the passage — and if it
|
||||
recognises the book, it supplies the canonical names out of its own training.
|
||||
|
||||
**Measured on the first 714 lv-bronte pairs, before the filter existed:**
|
||||
|
||||
- **13 beats (1.8%)** named source characters — `Rochester` ×6, `Jane` ×3, `Brocklehurst` ×2,
|
||||
`Beck`, `Fairfax`, `Helen`, `Burns`, `Eyre`, `Reed`, `Rivers`.
|
||||
- **0 of 714 RESPONSES did.** The rename was perfect; the instruction side was not.
|
||||
- One beat read *"Saoirse confirms Rochester's flaws, then agrees in English to marry him"* —
|
||||
a renamed name and a canonical one in the same sentence, which is the mechanism in miniature.
|
||||
|
||||
**Why it matters more than 1.8% sounds:** the beat is the INSTRUCTION half of the pair, so
|
||||
training on it re-teaches exactly the inventions the rename pipeline exists to remove.
|
||||
|
||||
⚠⚠ **EXPOSURE SCALES WITH HOW WELL THE GENERATOR KNOWS THE BOOK.** It is worst for
|
||||
public-domain classics and mildest for recent work. That is precisely why the Yarros and
|
||||
Hemingway runs came up clean and Brontë did not — **their clean runs are NOT evidence they are
|
||||
immune.** Both should be re-verified, and regenerated with `--source-entities`, before their
|
||||
pairs are trusted again.
|
||||
|
||||
**The fix.** `vet()` in `scripts/yarros-corpus/build_sft_pairs.py` gained a `sourcename`
|
||||
reject plus `--source-entities <entities.json>`, taking the UNRENAMED entity map. Fired at
|
||||
~3% of attempts on the Brontë rebuild. Commit `533cc0c`.
|
||||
|
||||
**The end-to-end guard that proves it.** The chain now verifies every built pair — beat,
|
||||
response and context — against every source surface before spending GPU hours:
|
||||
`[verify] 3858 pairs vs 368 source surfaces -> 0 leaks`.
|
||||
|
||||
⚠ A guard stricter than the gate cries wolf. The first verify pass excluded nothing and
|
||||
flagged `Monsieur’` ×14 and `Qu’est-ce` ×4 — French grammar, not leaks — because
|
||||
`--fold-clitics` leaves apostrophe keys the gate deliberately skips. Mirror `leak_gate.py`'s
|
||||
own predicate; a guard that fails on false positives gets disabled, which is worse than the
|
||||
leak it guarded.
|
||||
|
||||
Related: [[2026-09-17-lv-bronte-gate]], [[2026-09-16-lv-hemingway-corpus]], [[2026-09-16-lv-voices-line]].
|
||||
@@ -1,68 +0,0 @@
|
||||
# `[2026-09-17]` ESH: Cityside Fiber failed twice in six hours; site ran on Verizon failover
|
||||
|
||||
**Timeline (PDT).**
|
||||
|
||||
```
|
||||
19:09:07 Cityside dies. UDM fails over to Verizon 5G (WAN2). Site stays up at ~33 ms.
|
||||
19:51 Verified healthy on failover.
|
||||
20:01:37 esh-scale drops off the headscale mesh; 10.0.0.0/16 withdrawn; whole site dark
|
||||
from the colo. Beszel fires on all five ESH hosts.
|
||||
20:11-15 Service restored. Operator had switched WAN1 to DHCP to get Cityside working at
|
||||
all; it came back on CGNAT, not the static. Latency back to 9 ms.
|
||||
01:06:23 Cityside fails AGAIN. Failover to Verizon. Site up, ~37 ms.
|
||||
```
|
||||
|
||||
⭐ **The 20:01 blackout was most likely the operator's own WAN reconfiguration**, not ISP
|
||||
instability — switching the WAN type bounces the interface, esh-scale loses its path,
|
||||
headscale withdraws the route, and the site vanishes from the colo's view until it settles.
|
||||
An earlier session theory ("Cityside came back half-provisioned and the UDM failed back into
|
||||
an unstable session") is retired.
|
||||
|
||||
⚠ **The diagnostic that mattered: physical link stayed UP at 2.5 GE with zero errors
|
||||
throughout, while the ISP's next-hop `128.177.138.181` was unresponsive.** So "the ONT is
|
||||
fine, it is upstream of the ONT" — the line to give Cityside. Traceroute from NH3 reached
|
||||
`209.249.146.170` (one hop short) before dying, so the prefix was still routed.
|
||||
|
||||
⚠ **CROWDSEC ROTATION FRAGILITY IS LIVE.** The `esh` allowlist on ana-docker carries the now-
|
||||
dark static `128.177.138.182` (never-expiry), plus `97.190.18.88` (Verizon failover) and
|
||||
`23.164.40.174` (Cityside CGNAT), both **7-day expiry**. ESH is on a rotating carrier-NAT
|
||||
egress until the static is restored — the exact regime the 09-08 static purchase was meant to
|
||||
end, and the class that once blackholed the whole site via a false ban. **If ESH loses colo
|
||||
access, check `curl -s4 ifconfig.me` from esh-docker-vm FIRST** and allowlist the new address.
|
||||
|
||||
**Still pinned to the dark static and broken until it returns:** FortiGate `infra-ops`
|
||||
trusthost3 = `128.177.138.182`, so logins to ana-gw from ESH are refused. The dormant
|
||||
`esh-ana` IPsec is bound to wan1/static (disabled, so no impact).
|
||||
|
||||
⭐ **The mesh was NOT degraded on CGNAT** — tailscale hole-punched straight through
|
||||
(`direct 23.164.40.174:41641`), which is why latency read 9 ms rather than a DERP figure. An
|
||||
expectation of relay-on-CGNAT was wrong.
|
||||
|
||||
## RESOLVED 2026-09-17 ~12:30 PT — the static is back, confirmed on four axes
|
||||
|
||||
Not one check, because egress alone cannot tell a static WAN from a carrier NAT that happens
|
||||
to answer (see auto-memory `feedback_egress_ip_cannot_detect_cgnat`):
|
||||
|
||||
```
|
||||
config UDM WAN1 `wan_type = static`, ip 128.177.138.182, mask /30, gw 128.177.138.181
|
||||
— switched BACK from the DHCP the operator set at 20:11 during the outage
|
||||
active stat/health: isp_name "Cityside Fiber", ASN 18731, num_disconnected 0.
|
||||
WAN2 Verizon-5G is failover-only at priority 2 and idle.
|
||||
egress esh-docker-vm sees 128.177.138.182 — EQUAL to the WAN ip, so not behind CGNAT
|
||||
perf 2005/2142 Mbps symmetric; colo -> ESH 5.0 ms, 0% loss over 4 hosts-worth of pings
|
||||
(Cityside CGNAT read 9 ms, Verizon failover 33-37 ms)
|
||||
```
|
||||
|
||||
⭐ **The FortiGate pin un-broke itself and that was verified, not inferred.** `infra-ops`
|
||||
trusthost3 is `128.177.138.182`; from esh-docker-vm, ana-gw `tcp/22` is OPEN and the
|
||||
FortiGate offers a password prompt rather than dropping the connection — a trusthost
|
||||
mismatch refuses outright, so reaching auth *is* the trusthost passing. The dormant
|
||||
`esh-ana` IPsec bind to wan1/static is correct again (still disabled, still no impact).
|
||||
|
||||
⚠ **The crowdsec temporary allowlist entries are being LEFT to expire on their own**
|
||||
(2026-09-23): `97.190.18.88` Verizon and `23.164.40.174` Cityside CGNAT. Cityside failed
|
||||
twice in six hours on 09-17, so until the line has earned some confidence those two are
|
||||
cheap insurance against the exact false-ban blackout this rotation-fragility caused before.
|
||||
`128.177.138.182` stays never-expiry.
|
||||
|
||||
Related: [[2026-09-06-headscale-cutover]], [[2026-09-08-esh-static-wan-followups-and-ytvc]].
|
||||
-3
@@ -1,3 +0,0 @@
|
||||
# `[2026-09-17]` gitea was reaching the PUBLIC route from every repo on nh3-dev
|
||||
|
||||
**gitea was reaching the PUBLIC route from every repo on nh3-dev** — brokkr-smithy, sleipnir, Galdrabok, kvasir — and brokkr-smithy is pushed several times a week, so the fail2ban trigger was live, not dormant. Measured before acting (no split-horizon rewrite, no ssh alias, `ssh -G` confirmed port 22 to 38.120.12.44). Fixed by overriding the NAME once in `~/.ssh/config` rather than rewriting N remotes, so fresh clones and unaudited repos are covered too. Verified with a real `git ls-remote`, not by inspection. Commit `dcc1abc`. Flagged by brokkr-smithy-dev; `vh/imogen` created for them the same session.
|
||||
-3
@@ -1,3 +0,0 @@
|
||||
# `[2026-09-17]` headscale now split-DNSes `nh3.phasefinal.com` to the three AdGuards, so mesh clients can resolve the internal-only wildcard
|
||||
|
||||
**headscale now split-DNSes `nh3.phasefinal.com` to the three AdGuards, so mesh clients can resolve the internal-only wildcard** (`talk`, `booth` — public DNS has no record for them; the fleet AdGuard answers 10.100.10.50). Operator-approved, scoped to nh3 rather than all of `phasefinal.com`. Config `/etc/headscale/config.yaml` in CT 106 on nh3-pve, backup `config.yaml.bak-2026-09-17-splitdns`, restarted, and the new route **read back from a node's netmap** rather than assumed. ⚠ Two things worth knowing: split DNS works fine here with `global: []` — headscale issue #1161's "split ignored without global" does NOT apply to v0.29.3, verified on the live mesh — and `override_local_dns: true` would REQUIRE global, which is the config that makes a roaming laptop lose ALL DNS when the mesh is down. That is why split, not global. Routing was never the problem: nh3-scale already serves 10.100.0.0/16.
|
||||
-3
@@ -1,3 +0,0 @@
|
||||
# `[2026-09-17]` Hemingway ships as-is: operator ruled "ship stands" on both measured corpus defects
|
||||
|
||||
**Hemingway ships as-is: operator ruled "ship stands" on both measured corpus defects** — the 0.96% beat contamination and the 130 non-name entity-map surfaces. `audit_pairs_sourcenames.py --filter-out` and `audit_entity_map.py` exist and are the instruments if that is ever revisited; neither was run against the shipped adapter.
|
||||
@@ -1,135 +0,0 @@
|
||||
# `[2026-09-17]` lv-bronte: corpus gated for real, adapter trained, SHIPPED with a FAILED voice axis on the record
|
||||
|
||||
**Status: SHIPPED 2026-09-17 01:24 as `lv-bronte` on `vllm-voices` (fv-ml1 GPU0 :8027), ckpt475 —
|
||||
and it did NOT pass its voice gate.** Shipped because it is additive (one more named LoRA beside
|
||||
`voices-base` and `lv-yarros`, reached only by requesting it), reversible (one compose line; hot-unload
|
||||
measures 0.003 s), and clean on the SAFETY axis — 8-gram overlap identical to the never-saw-it control,
|
||||
on a public-domain corpus. VRAM cost was nil: GPU0 96092 -> 96090 MiB. The caveat is written into the
|
||||
compose file and into `/tank/aimodels/voice-adapters/lv-bronte-4b-v1/README.md` so it cannot be read
|
||||
as a clean pass by anyone who finds the adapter without finding this note.
|
||||
|
||||
⚠ **Do NOT cite lv-bronte as evidence pair-SFT works for this author.** The voice axis is unresolved,
|
||||
not passed.
|
||||
|
||||
## The gate result, in full
|
||||
|
||||
| axis | result | numbers |
|
||||
|---|---|---|
|
||||
| **A. VOICE** | ❌ **FAIL** (both candidates) | ckpt925 +0.210, ckpt475 +0.193 vs base — both **under** the 0.251 measured noise floor |
|
||||
| **B. NOT COPIED** | ✅ PASS | ckpt475 **0.00 hit-rate, max 0 — identical to the never-saw-it control**; ckpt925 0.01, max 8. Positive control saturates at 160, so the detector is not blind |
|
||||
| **C. NO DAMAGE** | ✅ PASS | ran-on +0.15 against a 0.400 floor |
|
||||
|
||||
```
|
||||
same-author target (held-out Brontë vs itself) delta_cb 0.338 <- best achievable
|
||||
ckpt925 0.531
|
||||
ckpt475 0.548
|
||||
base-unadapted 0.741
|
||||
```
|
||||
|
||||
⭐ **THE EFFECT LOOKS REAL AND SUBSTANTIAL; THE INSTRUMENT CANNOT CERTIFY IT.** The reachable
|
||||
span is 0.741 → 0.338 = 0.403, and the adapters closed **48–52% of everything achievable**.
|
||||
Both beat base on *every individual seed*. This is an UNDERPOWERED result, not a null one —
|
||||
and a "no effect" without its floor is unfalsifiable, so: **this method cannot resolve a voice
|
||||
improvement smaller than ~0.251 delta_cb at 30 beats × 4 seeds on this corpus.**
|
||||
|
||||
⭐⭐ **THE CAUSE IS STRUCTURAL: Brontë's val split yields 81 pairs (44 in-band) against
|
||||
Hemingway's 200**, because the corpus is 678k words against 994k. Maxing the fixture 30 → 44
|
||||
beats would shrink the floor by only ~√1.47 ≈ 1.2× (to ~0.21, against a 0.21 gap) — still
|
||||
marginal. **More SEEDS would not help either: the floor is a RANGE statistic, and ranges widen
|
||||
with more samples.** There is no cheap fix.
|
||||
|
||||
## ⚠ A DEFECT IN THE v2 RULE ITSELF, worth fixing before the next author
|
||||
|
||||
The floor is defined as the **largest within-arm seed spread across ALL arms**. Measured here:
|
||||
|
||||
```
|
||||
base-unadapted 0.772 0.813 0.751 0.772 spread 0.062
|
||||
ckpt475 0.670 0.631 0.604 0.578 spread 0.092
|
||||
ckpt925 0.776 0.584 0.525 0.620 spread 0.251 <- sets the floor, on ONE seed
|
||||
```
|
||||
|
||||
So **adding a third, noisier arm raised the bar that failed the clean one.** Run as the
|
||||
two-arm gate (base + ckpt475) the floor would have been 0.092 and +0.193 would have cleared
|
||||
at 2.1×. This was NOT exploited — picking the floor that passes your preferred answer is the
|
||||
threshold-chosen-after-seeing-the-numbers failure the pre-registration exists to prevent — but
|
||||
the rule should say whether the floor is computed over the compared pair or over every arm
|
||||
present. As written, a candidate's verdict depends on which *other* arms you happened to run.
|
||||
|
||||
**The outlier was diagnosed, not waved away.** Degeneracy probe (fraction of a generation made
|
||||
of its most repeated 5-gram) is uniform across every seed and both arms, 0.0078–0.0102. Seed
|
||||
1234 is not a collapsed generation; delta_cb genuinely has that variance for that arm.
|
||||
|
||||
## Which checkpoint, if it ships: **ckpt475**
|
||||
|
||||
The two are 0.017 apart on voice — far inside any floor, i.e. indistinguishable. On the axes
|
||||
that DO resolve, ckpt475 wins both: memorisation identical to the control (ckpt925 has a
|
||||
verbatim 8-gram hit), and **2.7× tighter seed-to-seed variance** (0.092 vs 0.251) with no
|
||||
degeneracy to explain the difference — consistent with ckpt925 sitting nearer the overfit
|
||||
boundary. Given a coin-flip on voice, take the one that provably did not memorise.
|
||||
|
||||
⭐ **THE RECIPE DID NOT TRANSFER.** Yarros and Hemingway both found their minimum inside
|
||||
epoch two. Brontë's minima are step 475 (ep 1.00, 2.6107) and step 925 (ep 1.96, 2.6129) —
|
||||
**0.0022 apart against a 0.0046 median neighbour jitter, i.e. indistinguishable**. Epoch 2
|
||||
buys Brontë NOTHING over epoch 1. What IS robust is the epoch-3 collapse: +0.075, ~16× jitter.
|
||||
Do not carry "two epochs on a three-epoch schedule" to a new author as settled.
|
||||
|
||||
## Artefacts
|
||||
|
||||
`gx10:~/lv-bronte/` (corpus-clean, corpus-renamed, entities-final.json, pairs/, beats-bronte-30.json),
|
||||
`gx10:~/r49-runs/bronte-4b-pairs-3ep/` (57 checkpoints kept), `gx10:~/r49-runs/bronte-eval/`
|
||||
(three arms × 120 generations, memorization.txt, voice_distance.txt, score.*.txt).
|
||||
Commits `fc834a8` `533cc0c` `7964d07` `e9e8c40` `8bb7686`.
|
||||
|
||||
⚠ Two output labels in `voice_distance.py` are hardcoded Yarros strings — it prints
|
||||
"reference: held-out Yarros" and a boilerplate "Base < Instruct" corroboration line. The
|
||||
NUMBERS are Brontë's; those two labels are not. Not yet fixed.
|
||||
|
||||
Related: [[2026-09-16-lv-voices-line]], [[2026-09-16-lv-hemingway-corpus]], [[2026-09-16-voices-seat-lora]].
|
||||
|
||||
---
|
||||
|
||||
## ⚠⚠ AMENDED 2026-09-17 — THE VOICE AXIS PASSES UNDER THE CORRECTED FLOOR RULE
|
||||
|
||||
Everything above is left verbatim; it is what was believed at ship time. This section is
|
||||
the correction, not a rewrite.
|
||||
|
||||
**The defect this file itself named was fixed, and fixing it flips ckpt475's verdict.**
|
||||
The section "⚠ A DEFECT IN THE v2 RULE ITSELF" above says the rule should state whether
|
||||
the floor is computed over the compared pair or over every arm present. It is now
|
||||
**pairwise**, pre-registered in `scripts/hemingway-corpus/GATE-PREREG.md` before a single
|
||||
lv-hemingway number existed. Re-scoring the SAME 360 generations — no re-run, no changed
|
||||
delta_cb:
|
||||
|
||||
```
|
||||
arm delta_cb per-seed spread
|
||||
ckpt925 0.531 (0.776 0.584 0.525 0.620) 0.251
|
||||
ckpt475 0.548 (0.670 0.631 0.604 0.578) 0.091
|
||||
base-unadapted 0.741 (0.772 0.813 0.751 0.772) 0.062
|
||||
|
||||
all-arms floor (as run) 0.251
|
||||
ckpt475 +0.193 vs pairwise floor 0.091 -> MOVED toward Brontë, 2.1x <- the two rules DISAGREE
|
||||
ckpt925 +0.210 vs pairwise floor 0.251 -> within the floor, NOT a finding
|
||||
```
|
||||
|
||||
⭐ **The sequence matters and is the reason this is not threshold-shopping.** The previous
|
||||
session found the defect, recorded it, and explicitly declined to exploit it. The rule was
|
||||
then changed prospectively on a structural argument independent of the answer it produces —
|
||||
the sampling variability of a difference A−B depends on A and B, never on a third arm C, so
|
||||
a candidate's verdict must not depend on which other arms were generated. `voice_distance.py`
|
||||
prints both floors and flags disagreement, so neither number can be quoted alone.
|
||||
|
||||
**Consequences:**
|
||||
- lv-bronte's voice axis is a **PASS at 2.1x**, not a fail. The caveat is amended in place
|
||||
(append-only) in `stacks/voices-seat/compose.yaml` and
|
||||
`/tank/aimodels/voice-adapters/lv-bronte-4b-v1/README.md`.
|
||||
- The sensitivity floor for that measurement is **0.091**, not 0.251.
|
||||
- "Do not cite lv-bronte as evidence pair-SFT works for this author" is **WITHDRAWN**.
|
||||
- ckpt475 over ckpt925 is unchanged and for unchanged reasons (no verbatim 8-gram hit,
|
||||
2.7x tighter seed variance).
|
||||
- The "no cheap fix for the underpowered result" analysis above is superseded for Brontë:
|
||||
it was underpowered against an inflated floor, not against its own.
|
||||
|
||||
**Also amended:** the two hardcoded Yarros labels flagged at the end of this file are fixed.
|
||||
`voice_distance.py --author` is now REQUIRED — the committed Brontë output literally reads
|
||||
"reference: held-out Yarros" over Brontë's numbers — and the stale "one seed-pair per arm /
|
||||
corroborates Base < Instruct" footer now reports what the run actually carries.
|
||||
@@ -1,163 +0,0 @@
|
||||
# `[2026-09-17]` lv-hemingway: SHIPPED on ckpt850 — the line's first clean voice pass, and one axis that needs reading
|
||||
|
||||
**Status: SHIPPED 2026-09-17 03:33 as `lv-hemingway` on `vllm-voices` (fv-ml1 GPU0 :8027),
|
||||
checkpoint-850.** Seat healthy 190 s after recreate, four models served
|
||||
(`voices-base`, `lv-yarros`, `lv-bronte`, `lv-hemingway`), GPU0 96,092 → **96,090 MiB** — a
|
||||
LoRA rides inside the existing seat and costs nothing. Adapter verified byte-identical to
|
||||
the checkpoint by sha256 across two hops.
|
||||
|
||||
Gate design **pre-registered before any generation existed**:
|
||||
`scripts/hemingway-corpus/GATE-PREREG.md`, commit `0bb4938`.
|
||||
|
||||
## The gate result — 3 arms × 60 held-out beats × 4 seeds = 240 generations per arm
|
||||
|
||||
| axis | result | numbers |
|
||||
|---|---|---|
|
||||
| **A. VOICE** | ✅ **PASS, 6.4×** | +0.413 delta_cb vs base, pairwise floor 0.064. Also clears the OLD all-arms floor (0.113) — **this verdict does not depend on the rule change** |
|
||||
| **B. NOT COPIED** | ⚠ **content clean, rate 7× the author's own** | 0.07 hit-rate, mean-longest 0.6, **max 9 words**. Base 0.00, **held-out Hemingway 0.01** |
|
||||
| **C. NO DAMAGE** | ✅ PASS | ran-on +0.08, on-beat −0.14, both inside a 0.217 floor; in-band 0.79 vs base 0.05 |
|
||||
|
||||
```
|
||||
same-author target (held-out Hemingway vs itself) delta_cb 0.364 <- best achievable
|
||||
ckpt1750 0.439
|
||||
ckpt850 (SHIPPED) 0.511
|
||||
base-unadapted 0.924
|
||||
```
|
||||
|
||||
⭐ **THE STRONGEST VOICE RESULT IN THE LINE. The span is 0.924 → 0.364 = 0.560 and ckpt850
|
||||
closed 73.8% of it (ckpt1750 86.6%)**, against lv-bronte's 48%. Power came from the corpus,
|
||||
not from a better method: 173 in-band val pairs allowed a **60-beat** fixture where Brontë
|
||||
had 44 in-band and could only run 30.
|
||||
|
||||
## ⚠⚠ AXIS B — THE COMFORTABLE EXPLANATION WAS WRONG, AND THE CONTROL IS THE ARTIFACT
|
||||
|
||||
`memorization_check.py` uses the **base-unadapted arm** as its negative control, and on this
|
||||
corpus that control is weak in one direction only — **it makes an innocent arm look guilty.**
|
||||
Base writes 18,035 words of *summary*; the adapted arms write 27,413 of *pastiche*. Text that
|
||||
does not imitate the register cannot collide with its n-grams, so base's 0.00 partly measures
|
||||
"different register", not "did not memorise".
|
||||
|
||||
The obvious hypothesis was that Hemingway's plain, high-frequency register makes 8-gram
|
||||
collisions inevitable for any arm that learns it. **That hypothesis is refutable, was tested,
|
||||
and is FALSE.** New control: **held-out Hemingway — the author himself, val text no arm
|
||||
trained on — scored against the train split**, chunked to the generations' own median length
|
||||
(101 words) so the comparison is like for like.
|
||||
|
||||
```
|
||||
sample n hit-rate mean-longest max
|
||||
HELD-OUT HEMINGWAY (never trained) 370 0.01 0.1 10
|
||||
base-unadapted 240 0.00 0.0 0
|
||||
ckpt1750 240 0.08 0.7 9
|
||||
ckpt850 (SHIPPED) 240 0.07 0.6 9
|
||||
positive control (train vs train) 160 <- not blind
|
||||
```
|
||||
|
||||
⭐⭐ **The adapter reproduces train-corpus word sequences ~7× more often than the author
|
||||
reproduces himself.** If the register explained it, real Hemingway would collide at the same
|
||||
rate; it collides at 0.01.
|
||||
|
||||
⭐ **And the exposure is still nil, which is a different question from the rate.** All 19
|
||||
matched runs were READ, not counted. Every one is stock dialogue — `i don t think so the girl
|
||||
said`, `came over and sat down at the table`, `how do you feel i feel very well`. No plot, no
|
||||
imagery, no distinctive phrase, **no proper noun** (the one name-shaped hit, `swift tristan`,
|
||||
is the RENAMED invented name, not Hemingway's). The longest run is **9 words — shorter than
|
||||
the 10-word run genuinely unseen Hemingway shares with the train split by coincidence.**
|
||||
|
||||
What is being reproduced is the *grammar of his dialogue*, which is the thing the adapter
|
||||
exists to learn, rendered in the commonest words in English. **Elevated rate, zero
|
||||
protectable content.** Hemingway is in copyright; the in-line precedent is lv-yarros, also in
|
||||
copyright, shipped at 0.10 against a 0.07 control. Unload is 0.003 s and one compose line.
|
||||
|
||||
⚠ **The durable lesson is about the instrument, not this adapter: a negative control that
|
||||
differs from the candidate in a way CORRELATED with the metric is not a control.** Always ask
|
||||
what the metric returns for a known-innocent sample *in the same register*.
|
||||
|
||||
## Why ckpt850 and NOT ckpt1750, the loss minimum
|
||||
|
||||
ckpt1750 has the better point estimate on voice (0.439 vs 0.511) and **it is not usable**:
|
||||
|
||||
```
|
||||
gap between candidates 0.072
|
||||
pairwise floor max(0.113, 0.050) 0.113 -> NOT resolvable
|
||||
```
|
||||
|
||||
Indistinguishable, so the pre-registered tiebreak falls to the axes that resolve — and
|
||||
**ckpt850 wins every one**:
|
||||
|
||||
| | ckpt850 (shipped) | ckpt1750 |
|
||||
|---|---|---|
|
||||
| seed spread | **0.050** | 0.113 — **2.3× wider** |
|
||||
| memorisation hit-rate / mean-longest | **0.07 / 0.6** | 0.08 / 0.7 |
|
||||
| ran-on | **0.08** | 0.12 |
|
||||
| epoch | **0.959** | 1.973 |
|
||||
|
||||
ckpt1750's spread is one seed: 0.491, 0.449, 0.468, then **0.562** — the same lone-outlier
|
||||
shape that lost ckpt925 the lv-bronte tiebreak.
|
||||
|
||||
⭐ **THE TWO-EPOCH RECIPE DID NOT TRANSFER HERE EITHER — it is now 0 for 2.** Hemingway's
|
||||
minimum really is step 1750, but step 850 is **+0.0040 against a 0.0044 median neighbour
|
||||
jitter**, with three checkpoints inside one jitter of the best. Epoch 2 buys nothing that
|
||||
resolves and costs 2.3× the variance. Only the epoch-3 collapse is robust: **+0.0762 = 17.4×
|
||||
jitter**, which is why `adapter/` was never gated. **Stop carrying "two epochs on a
|
||||
three-epoch schedule" forward; read the curve and prefer the earlier tied checkpoint.**
|
||||
|
||||
## Pre-flight: the beat leak IS present in Hemingway, and the fixture is clean
|
||||
|
||||
`audit_pairs_sourcenames.py` (new, commit `0bb4938`) closes the blind spot `leak_gate.py` has
|
||||
by construction. Controls green every run: 941/941 surfaces found in the unrenamed source,
|
||||
nonce absent from both trees, 6/6 planted names detected.
|
||||
|
||||
```
|
||||
train beats 70 of 7,094 (0.96%) Santiago x16, Catherine x7, Rinaldi x3, Brett, Harry,
|
||||
Jake, Pablo, Nick, Maria, Helen ... 36 distinct
|
||||
train responses 0 of 7,294 -- the rename itself held perfectly
|
||||
val beats 0 of 200 -- THE EVAL FIXTURE IS CLEAN; the gate is unconfounded
|
||||
```
|
||||
|
||||
⭐ The beat-only signature is exactly lv-bronte's. **Yarros's and Hemingway's earlier clean
|
||||
runs were never evidence of immunity** — they predate the detector.
|
||||
|
||||
**Cross-validated on real data** where the answer was already recorded: the fixed Brontë
|
||||
pairs return **0 of 3,858** (matching "0 leaks across 3,858 pairs"), and
|
||||
`pairs-full.CONTAMINATED.jsonl` returns **15 of 792 = 1.89%** with Rochester ×6, Jane,
|
||||
Brocklehurst ×2, Beck, Fairfax, Burns, Helen, Eyre — against a record of "13 of the first 714
|
||||
beats (1.8%)" with the same names. An independently written instrument reproducing a
|
||||
documented finding at the right magnitude is what makes its zeroes mean *absent*, not *blind*.
|
||||
|
||||
`--filter-out` produces a clean **7,024-pair** set in one command (70 dropped, 0.99%),
|
||||
verified by re-audit at 0 of 7,024. **A retrain on it is the operator's call, not done.**
|
||||
|
||||
## ⚠ A SECOND corpus defect, measured and NOT acted on
|
||||
|
||||
`audit_entity_map.py` (new, commit `051b99e`) is the mirror of `audit_stoplist.py`: it finds
|
||||
surfaces wrongly held **IN** the entity map, which `leak_gate.py` cannot see because it only
|
||||
ever asks whether the author's names are GONE, never whether non-names were spared.
|
||||
|
||||
```
|
||||
positive control `other` 764/1356 article-preceded = 0.56
|
||||
negative control 100 honorific-confirmed people, highest Inglés 0.26, bulk 0.00-0.06
|
||||
FLAGGED 130 of 946 surfaces · 1,616 instances · 0.162% of corpus words
|
||||
```
|
||||
|
||||
`African`, `Chinese`, `Basques`, `Republican`, `Communist`, `X-ray`, `Coca-Cola`, `Ritz`,
|
||||
`Prado`, `Cezanne` were all renamed into invented proper nouns. **Some flags are correct
|
||||
renames** — `the Widow`, `the Informer` are genuine Hemingway epithet-names — so every hit is
|
||||
reported for reading, never auto-removed. Plus **16 bare initials in the map**, `C` at 274
|
||||
occurrences: the same class as the `G` caught by hand about to be renamed 248 times.
|
||||
|
||||
At 0.162% of words this did not block the ship. It is the thing to fix first if a corpus
|
||||
rebuild ever happens.
|
||||
|
||||
## Artefacts
|
||||
|
||||
`gx10:~/lv-hemingway/` (corpus-clean, corpus-renamed, beats-hemingway-60.json + sidecar,
|
||||
eval-hemingway.sh, voice-prep.py, eval.log), `gx10:~/r49-runs/hemingway-4b-pairs-3ep/`
|
||||
(54 checkpoints kept), `gx10:~/r49-runs/hemingway-eval/` (three arms × 240 generations,
|
||||
memorization.txt, voice_distance.txt, score.*.txt).
|
||||
`fv-ml1:/tank/aimodels/voice-adapters/lv-hemingway-4b-v1/` (adapter + a README carrying the
|
||||
axis-B caveat, so it cannot be read as clean by anyone who finds the adapter without this).
|
||||
Commits `0bb4938` `051b99e` `5e66114` `2e9b118`.
|
||||
|
||||
Related: [[2026-09-17-lv-bronte-gate]], [[2026-09-16-lv-hemingway-corpus]],
|
||||
[[2026-09-16-lv-voices-line]], [[2026-09-16-voices-seat-lora]],
|
||||
[[2026-09-17-beat-contamination-leak]].
|
||||
-3
@@ -1,3 +0,0 @@
|
||||
# `[2026-09-17]` lv-krakauer D1 built — 126 units, 422,880 words — and its name guard caught three defects nothing else would have reported.
|
||||
|
||||
⭐ **lv-krakauer D1 built — 126 units, 422,880 words — and its name guard caught three defects nothing else would have reported.** Back matter searched only the LAST unit while the apparatus sat in unit 37 of 41; relying on the splitter to drop front matter failed because the ebook TOC sits above the author's note and gave it a `Chapter Thirty-Two` to start on; and **zero was the wrong bar** — 2 survivors are Krakauer writing about his own father in Into the Wild's autobiographical chapters, so the allowance is pinned at 2 with every survivor printed. ⚠ Both strips are windowed in the OPPOSITE direction from McCarthy's, because Krakauer's `ALSO BY`/`Copyright`/`About the Author` sit at 0.0–0.6% of the file. Commit `4be0630`.
|
||||
-3
@@ -1,3 +0,0 @@
|
||||
# `[2026-09-17]` lv-mccarthy D1 built — 167 units, 584,756 words — and the whole job was protecting a style that reads as damage.
|
||||
|
||||
⭐ **lv-mccarthy D1 built — 167 units, 584,756 words — and the whole job was protecting a style that reads as damage.** 0.0 quote marks per 10k (Hemingway 838), `dont`/`aint`/`wont`. The builder runs NO typography normalisation and asserts the quote density afterwards. Two truncated catalogue rows dropped for complete mobi siblings; all 15 containment pairs measured (worst 0.10%); back matter in 4 of 6 works carried the author's name 26 times → 0. ⚠ The back-matter strip runs BEFORE the split here — Blood Meridian and The Crossing end with a dumped TOC of bare roman numerals, the exact shape of a chapter marker. Commit `f3bf3ca`.
|
||||
-3
@@ -1,3 +0,0 @@
|
||||
# `[2026-09-17]` lv-mccarthy's D1→D3 chain was RECOVERED, not remembered — there was no runbook and the commands went over non-interactive ssh, so no history survived.
|
||||
|
||||
**lv-mccarthy's D1→D3 chain was RECOVERED, not remembered — there was no runbook and the commands went over non-interactive ssh, so no history survived.** Rebuilt candidates and matched sha256 against the artifacts on disk: 6 works, the entity map, the final map and all 36 copy files byte-identical. Now pinned in `scripts/mccarthy-corpus/RUNBOOK.md` with every deviation. ⚠ **D1 must run on nh3-dev** (the builder reads the kvasir catalogue by absolute path); the prior "on gx10" note is true of D2 onward only. ⚠ No phrase map exists for this corpus, so the gate's phrase audit never ran — Yarros and Brontë both had one.
|
||||
@@ -1,159 +0,0 @@
|
||||
# `[2026-09-17]` lv-mccarthy D1→D3 — built, gated, and every stage caught a defect in the stage before it
|
||||
|
||||
**`~/lv-mccarthy/` on pfi-gx10.** `corpus-clean/` (167 units, 584,716 words),
|
||||
`corpus-renamed/` (6 copies, 1,002 records), `scripts/`. Commits `705fa3a` `f3bf3ca`
|
||||
`0fa68cb` `5aa10bf` `5ddb047`.
|
||||
|
||||
```
|
||||
leak gate 0 of 75 renameable and 0 of 37 sub-threshold survive in any copy
|
||||
positive control 108/108 surfaces found in the unrenamed source
|
||||
negative control nonce absent from both trees
|
||||
```
|
||||
|
||||
## The shared splitter: choose by SIZE, not by count
|
||||
|
||||
`scripts/r49-corpus/split_units.py`. The inherited rule was "most units above a floor", which
|
||||
is wrong for any book whose markers are PARTS:
|
||||
|
||||
```
|
||||
Cities of the Plain 4 roman marks -> 4 units, median 22,312w
|
||||
The Crossing 4 roman marks -> 4 units, median 37,310w
|
||||
```
|
||||
|
||||
Four beats one, so it won, and the old guard only fired at exactly one unit. Now: a mode
|
||||
qualifies only if its median unit is inside **[600, 12000]** AND no unit holds half the work;
|
||||
among qualifying modes **priority** breaks the tie (contents > chapter-word > roman >
|
||||
bare-numeral > caps-title). Works with no divisions fall back to **paragraph-block sections**.
|
||||
|
||||
⭐⭐ **The first version of that rule was WORSE than what it replaced, and a control caught
|
||||
it.** Scoring by "median closest to target" chose `caps-title` over the real chapters of
|
||||
Hemingway's *True at First Light*:
|
||||
|
||||
```
|
||||
bare-numeral 20 units median 5,337w max 11,155 <- the book's own chapters
|
||||
caps-title 6 units median 777w max 113,886 <- median looked BETTER
|
||||
```
|
||||
|
||||
Five stray all-caps lines gave five tiny units beside **one holding 97% of the book**. A median
|
||||
cannot see that distribution; a max bound can. Controls green both ways afterwards: 8/10
|
||||
Hemingway works reproduce the shipped mode and count exactly, and 40,000 words with no blank
|
||||
lines returns **1 unit** rather than fabricating sections.
|
||||
|
||||
⚠ The Hemingway builder is deliberately NOT repointed at this module — its corpus is shipped
|
||||
and its sha is pinned by a live adapter.
|
||||
|
||||
## D1: the job was protecting a style that reads as damage
|
||||
|
||||
```
|
||||
quote marks 0.0 per 10k (Hemingway 838)
|
||||
apostrophes 123 per 10k (Hemingway 241) `dont` `aint` `wont` `didnt`
|
||||
```
|
||||
|
||||
⚠⚠ **`repair_typography.py` MUST NOT be run on this corpus.** It normalises "toward what the
|
||||
text does" and would put the quotation marks back. The builder runs no normalisation and then
|
||||
**asserts** the quote density, so a future well-meaning change fails the build.
|
||||
|
||||
⚠ **AND IT MAKES THE VOICE GATE EASY TO PASS FOR THE WRONG REASON.** `voice_distance.py` is
|
||||
Burrows's Delta over CHARACTER BIGRAMS. An adapter that learns only "emit no quotation marks"
|
||||
moves delta_cb a long way without having learned a sentence. **Pre-register a
|
||||
punctuation-normalised secondary read before gating lv-mccarthy.** Tracked in the builder
|
||||
docstring, commit `f3bf3ca`.
|
||||
|
||||
Also: two truncated catalogue rows dropped for complete mobi siblings; all 15 cross-work
|
||||
containment pairs measured (worst **0.10%**); back matter in 4 of 6 works carrying the author's
|
||||
name 26 times → **0**; alphabet re-derived at 1,411 non-ASCII letters across 14 Spanish forms.
|
||||
|
||||
⚠ The back-matter strip runs **BEFORE** the split for McCarthy, inverting the Hemingway order:
|
||||
Blood Meridian and The Crossing end with a dumped table of contents made of bare roman numerals
|
||||
on their own lines — the exact shape of a chapter marker.
|
||||
|
||||
## D2 caught a D1 defect: three small-caps manglings
|
||||
|
||||
The entity map returned `E`, `H`, `T`, `K` as renameable entities with 17–33 capitalised
|
||||
occurrences each — the `G` class from Hemingway, where `G` was about to be renamed to a surname
|
||||
248 times. Reading them showed the extractor mangled small-caps openings three ways:
|
||||
|
||||
```
|
||||
1. SPLIT INITIAL `T HE HOUSE was built` -> `The house was built` 32 cases
|
||||
2. UNMARKED RUN `THEY STOOD in the doorway` -> `They stood in the doorway` 88 cases
|
||||
3. LOST INITIAL `HE CANDLEFLAME` -> `THE CANDLEFLAME` 1 case
|
||||
```
|
||||
|
||||
Rule 1 requires a FOLLOWING all-caps word, so `A TV was playing` and `A Mexican was changing`
|
||||
are untouched. Rule 2's `[a-z]` lookahead is what makes it safe — a genuine shout or sign is
|
||||
not followed mid-sentence by lowercase. All 23 distinct first words of the 88 were checked.
|
||||
|
||||
⚠⚠ **A fourth "fix" was nearly shipped that would have CORRUPTED the text.** `HEY RODE` →
|
||||
`THEY RODE` looked right from a survey of the BUILT corpus. The raw master has `THEY RODE`
|
||||
intact, twice — `HEY RODE` matched as a SUBSTRING, and the unanchored replace produced
|
||||
`TTHEY RODE`, which rule 2 then lowercased to `Tthey rode`. Caught by the count assertion
|
||||
(expected 1, replaced 2) and settled by reading the master. ⚠ My first corruption check also
|
||||
missed it, searching for `TTHEY` when the pipeline had already lowercased it — **check the
|
||||
shape the pipeline emits, not the shape you imagined.**
|
||||
|
||||
## D2's own gates: and `audit_stoplist` was scanning its own rationale
|
||||
|
||||
⚠⚠ **A defect in `audit_stoplist.py`, latent for every corpus before this one.** It built its
|
||||
surface set from every list value in the stoplist JSON — including `_why`, which by convention
|
||||
is a LIST OF PROSE LINES. Its empty separator line matched the honorific pattern **139 times**,
|
||||
printing a flag with no surface name above the one real catch. Now skips `_`-prefixed keys.
|
||||
|
||||
That real catch was a contradiction **inside my own file**: `Franklin` sat in the geography list
|
||||
(the old name for El Paso) while the same file's note recorded *"I'm here to see Mr Franklin"*,
|
||||
a lawyer in All the Pretty Horses. A second self-inflicted one: a speculative A–Z fragments list
|
||||
stoplisted `I` and `A`, and `Sir I dont think I can do that` duly tripped the audit. It is now
|
||||
the four letters actually measured as entities.
|
||||
|
||||
Everything ambiguous was read in context: **Socorro is the ranch cook, not the New Mexico
|
||||
town**; Niño, Keno and Redbo are HORSES (renameable, the `Inglés` precedent); Yaqui and Gilenos
|
||||
are real peoples; Hashknives is a real cattle outfit; Hearst, Trias, Huerta and Madero are real
|
||||
historical figures on the page under their own names.
|
||||
|
||||
Final: 123 map surfaces, 124-surface stoplist, `entities.py` 27/27 controls, both audits PASS.
|
||||
|
||||
## The human gender pass is an auditable file
|
||||
|
||||
The honorific/window resolver scored **21 correct / 3 held / 1 WRONG** against a 26-name
|
||||
control; the base-rate proximity resolver built for Hemingway scored 18/6/1 and **its own guard
|
||||
correctly REFUSED to write**. So the incumbent stands and four entries are fixed by hand in
|
||||
`gender_overrides_mccarthy.json`, each carrying its evidence.
|
||||
|
||||
⚠ All four are female and all four look male-dominated in raw counts, because this corpus runs
|
||||
**29,144 male pronouns to 5,036 female — a base rate of 85.3% male**. Carla Jean Moss at
|
||||
31m/21f would be 44m/8f at that rate; 21 against an expected 8 is decisive. Same arithmetic that
|
||||
recovered Pilar and Brett on Hemingway. Alfonsa was in the control and is correctly absent from
|
||||
the map at 4 occurrences, below the min-count floor — an error in the control, not the pipeline.
|
||||
|
||||
`apply_gender_overrides.py` refuses twice: a name absent from the map is an error rather than a
|
||||
silent no-op, and overruling a gender the detector holds needs an explicit `"correcting": true`
|
||||
so it cannot look like filling a held entity in a diff.
|
||||
|
||||
## D3: three calls, and the holdout fix that matters most
|
||||
|
||||
1. **`--scope corpus`**, not the per-work default. Nine surfaces appear in more than one work —
|
||||
Parham (The Crossing + Cities of the Plain), Grady and Cole (All the Pretty Horses + Cities
|
||||
of the Plain), Socorro, Héctor. A per-work map gives John Grady a different invented name in
|
||||
each novel, turning one character into two.
|
||||
2. **A new `mccarthy` preset.** Hemingway's romance pool carries `it_IT` and `fr_FR` for his
|
||||
Italian and French casts; McCarthy writes neither language. `en_GB` goes for the same reason.
|
||||
`en_US` + `es_MX`/`es_ES` at an even share.
|
||||
3. **`--min-cap 5` to match the entity map's floor.** The first gate run FAILED with 45
|
||||
survivors: `entities.py` admits cap ≥ 5 while `rename.py` renamed only cap ≥ 8, so every
|
||||
entity between sat in the map, was never renamed, and counted as a leak. Hemingway never hit
|
||||
it because its map had `sub_threshold_total: 0`.
|
||||
|
||||
⭐ **`--holdout-chapter` NOW TAKES A LIST.** The val split is one chapter index per work, so its
|
||||
SIZE is set by how many WORKS a corpus has, not how many words:
|
||||
|
||||
```
|
||||
Hemingway 10 works -> 9 val units -> 36,563 words/copy -> gate DECISIVE
|
||||
Brontë 4 works -> 4 val units -> 17,043 words/copy -> gate MARGINAL
|
||||
McCarthy 6 works -> 6 val units -> ~18,000 would have been Brontë's end
|
||||
```
|
||||
|
||||
Holding out chapters **7 and 17** gives **11 units and 40,653 words per copy — larger than
|
||||
Hemingway's** — for 7% of the corpus, on a corpus 40% smaller than his. No amount of corpus size
|
||||
fixes a val split that scales with work count.
|
||||
|
||||
Related: [[2026-09-17-lv-hemingway-gate]], [[2026-09-17-mccarthy-krakauer-d1]],
|
||||
[[2026-09-17-lv-bronte-gate]].
|
||||
@@ -1,104 +0,0 @@
|
||||
# `[2026-09-17]` The leak gate passed with five protagonist names still in every copy
|
||||
|
||||
Found during D4 pre-flight, three stages downstream of where it happened. Commit `c559664`.
|
||||
|
||||
```
|
||||
leak gate, 2026-09-17 morning 0 of 75 renameable, 0 of 37 sub-threshold, both controls green
|
||||
actually present, all 6 copies Bell x2 Chigurh x3 Moss x2 Toadvine x4 Glanton x2
|
||||
```
|
||||
|
||||
## The mechanism
|
||||
|
||||
`leak_gate.py` scans `\b(Surface)\b`. **A character inserted inside a name defeats that
|
||||
pattern outright**, so a mangled occurrence is not merely unrepaired — it is *unrenameable*
|
||||
by `rename.py` and *unreportable* by the gate, and the gate prints a clean zero over it.
|
||||
Two extraction artifacts produce exactly that:
|
||||
|
||||
```
|
||||
B ell C higurh M oss T oadvine a small-caps drop cap kept as its own token
|
||||
Toad-vine Glan-ton a print line-break hyphen kept by the extractor
|
||||
```
|
||||
|
||||
⭐ **Every VISIBLE occurrence had been renamed correctly** — exact-match survivors were 0,
|
||||
as the gate said. That is what makes this residue invisible to a spot-read: the names are
|
||||
gone everywhere you look. `Bell` sits in the entity map at 147 capitals, `Glanton` at 365.
|
||||
|
||||
This is the third member of a family. lv-bronte's was `_Antigua_` (`_` is a word character,
|
||||
so `\bAntigua\b` cannot match inside it), found by hand in 2026-09-16 and never generalised.
|
||||
**The generalisation is the point: any separator inside a name blinds a word-boundary scan.**
|
||||
|
||||
## Fixed at three levels, and all three must stay
|
||||
|
||||
1. **`build_corpus_mccarthy.py` rules 4 and 5** repair the source text — 32 split initials
|
||||
with a *lowercase* remainder (rule 1 requires a following ALL-CAPS word and DROPCAP
|
||||
requires two, so this is the class both leave behind), 5 hyphen-split names by name.
|
||||
Both carry expected counts so a master change fails the build.
|
||||
⚠ **Rule 4's letter class is consonants only.** `I` opens **1,966** paragraphs (the
|
||||
pronoun), `A` opens 143 (the article), `Y` opens 32 (Spanish *y*). Folding any of them
|
||||
would corrupt 2,141 lines to fix 32 — the same `I`/`A` trap that bit `audit_stoplist.py`.
|
||||
2. **`leak_gate.py` runs a separator-tolerant pass every time**, with its own positive and
|
||||
negative controls, and **it fails the gate**. Validated against the pre-fix tree: reports
|
||||
all five surfaces, exits 1.
|
||||
3. The exact-match passes are untouched, so the old verdict is reproduced alongside the new.
|
||||
|
||||
⚠ **The fragment filter is what makes the new pass usable.** A naive separator-tolerant
|
||||
scan is dominated by false positives — on Hemingway it returns 21 hits of which **18 are
|
||||
ordinary text** (`God damn` for the surface `Goddamn` ×14, plus `I run`, `On an`, `Do me`,
|
||||
`Si le`). The discriminator, with no dictionary: in a genuine split at least one FRAGMENT
|
||||
is not a word of this corpus. `God` and `damn` occur constantly; `Primi`, `tivo`, `ell`,
|
||||
`higurh`, `Toad` do not. That one test cleared all 18 and kept all 3 real ones.
|
||||
|
||||
⚠ **My first negative control could not pass.** It planted the split nonce in its own probe
|
||||
text and then asserted the nonce was absent — an alarm wired to itself, failing on every
|
||||
run. It now hunts the split nonce in the *real* copies. A control that cannot pass is not a
|
||||
control.
|
||||
|
||||
## The shipped corpora, checked with the committed instrument
|
||||
|
||||
Re-derived with the COMMITTED gate, not a scratch probe:
|
||||
|
||||
```
|
||||
lv-bronte GATE PASSED 0 separator-split survivors (9.7 s)
|
||||
lv-hemingway GATE FAILED Pasionaria, Primitivo, Chicote (34.6 s)
|
||||
1 occurrence each per copy, in all 6 copies — SHIPPED and LIVE
|
||||
```
|
||||
|
||||
⚠ The first version of this scan was **too slow to run** on Hemingway — per-surface scanning
|
||||
is O(surfaces x copies x corpus) and 881 surfaces x 10 copies was still going at 5 minutes
|
||||
when it was killed. Rebuilt as one alternation pass, same trick `scan()` already used: 35 s,
|
||||
identical verdict and identical hit counts on both McCarthy trees. **A gate too slow to run
|
||||
is not a gate.**
|
||||
|
||||
**Operator call outstanding** on whether 3 names in a 958k-word corpus warrant re-gating and
|
||||
retraining a live adapter. Not acted on.
|
||||
|
||||
## The chain was recovered, not remembered — and is now written down
|
||||
|
||||
There was **no McCarthy runbook**, and the D1→D3 session issued its commands over
|
||||
non-interactive ssh so no shell history survived. The chain was recovered by rebuilding
|
||||
candidates and matching sha256 against the artifacts on disk, then pinned:
|
||||
|
||||
```
|
||||
D1 build_corpus_mccarthy.py 6 works byte-identical
|
||||
D2 entities.py --min-count 5 --fold-clitics --drop-acronyms --min-mid-ratio 0.2 --min-mid 2
|
||||
D2c apply_gender_overrides.py entities-final.json byte-identical
|
||||
D3 rename.py --preset mccarthy --scope corpus --min-cap 5 --copies 6 --seed 4919
|
||||
--holdout-chapter 7 17 all 36 copy files byte-identical
|
||||
```
|
||||
|
||||
⚠ `--min-mid-ratio` is what keeps `Yeah`/`Buenas`/`Shh`/`Sí` out of the map. The map is
|
||||
**insensitive** to it: any value in [0.05, 0.3] with `--min-mid` 1 or 2 reproduces byte-for-
|
||||
byte; `--min-mid 3` does not. The original values are unrecoverable and it does not matter —
|
||||
which is worth saying, because an exact-looking recipe that was never pinned invites a
|
||||
false claim of reproduction. Full recipe and every deviation: `scripts/mccarthy-corpus/RUNBOOK.md`.
|
||||
|
||||
⚠ **D1 must run on nh3-dev** — the builder reads the kvasir catalogue by absolute path and
|
||||
gx10 has no copy. The previous session's "on gx10" note is true of D2 onward only.
|
||||
|
||||
⚠ **No phrase map exists for this corpus**, so the gate's phrase audit does not run at all.
|
||||
Yarros and Brontë both had one. Not closed.
|
||||
|
||||
Rollback: `~/lv-mccarthy/corpus-{clean,renamed}.pre-splitfix` on gx10.
|
||||
|
||||
Related: [[2026-09-17-mccarthy-d1-d3]], [[2026-09-17-beat-contamination-leak]],
|
||||
[[2026-09-17-lv-hemingway-gate]], [[2026-09-17-lv-bronte-gate]].
|
||||
@@ -1,3 +0,0 @@
|
||||
# `[2026-09-17]` Measured and DELIBERATELY not changed, three of them.
|
||||
|
||||
**Measured and DELIBERATELY not changed, three of them.** The oversize-passage drop is 13.9% of McCarthy's train words, between Hemingway's 10.0% and the shipped Brontë's 18.3% — in range, no change. `BEAT_PROMPT` asserts the passage is first-person and McCarthy is third; measured inert (**0** narrator-retries against Hemingway's 615 of 7,094), so the prompt was left alone. Blood Meridian's 131 dash-separated chapter-argument paragraphs DID warrant a change and `--drop-leading-heading` now eats them (0 in every other work of all three corpora).
|
||||
@@ -1,69 +0,0 @@
|
||||
# `[2026-09-17]` Which voices earn a training seat next — measured against the catalogue, not chosen by taste
|
||||
|
||||
Method: rank every author in the kvasir catalogue by **usable extracted** works, then apply the
|
||||
selection criterion the lv-krakauer parking established — *does the author have a voice*, asked
|
||||
before any corpus work, and specifically **does that voice live where the instrument looks**.
|
||||
`voice_distance.py` is Burrows's Delta over CHARACTER BIGRAMS, so it sees function-word morphology,
|
||||
punctuation and sentence rhythm. A writer whose distinction is plot, research or subject matter is
|
||||
invisible to it — an adapter cannot carry that, and the gate cannot measure it.
|
||||
|
||||
⚠ `triage.length` is in **CHARACTERS**, ~5.2 chars/word calibrated against builds we did ourselves
|
||||
(The Crossing mobi 777,420 chars = our measured 149,985 words). Dedup by title taking the max across
|
||||
formats, and floor at 100,000 chars — that is what excludes the `accepted`-but-truncated rows
|
||||
(Blood Meridian epub at 6,031 chars beside the mobi's 623,849).
|
||||
|
||||
## ⭐ The size ranking INVERTS the voice ranking at the top
|
||||
|
||||
```
|
||||
Stephen King 76 works 12,133,529 w <- biggest, and NOT a candidate
|
||||
Agatha Christie 72 5,451,377 <- second biggest, the Krakauer case exactly
|
||||
Terry Pratchett 52 4,821,474
|
||||
Georgette Heyer 30 3,416,867
|
||||
Graham Greene 45 3,037,425
|
||||
William Faulkner 25 2,981,183 <- the pick
|
||||
```
|
||||
|
||||
Christie is the whole lesson in one row: a superb writer whose genius is plot architecture, in prose
|
||||
deliberately kept transparent. Nothing for a char-bigram Delta to grip. King is the softer version —
|
||||
distinctive in pacing and brand-name texture, not in syntax.
|
||||
|
||||
## The three that clear both bars
|
||||
|
||||
**1. William Faulkner — 25 catalogue rows, ~15 pure novels, ~1.6M words.**
|
||||
*The voice in one sentence:* sentences that defer their main clause through stacked subordination
|
||||
and coined compounds until the reader is held inside a single unbroken perception.
|
||||
About as char-bigram-legible as English gets — the voice IS the clause-joining morphology and the
|
||||
`and`/`which`/`that` density. ⭐ **And he is McCarthy's stylistic ancestor, which is the real
|
||||
argument:** the Brontë gate record states the frozen adjudication needs "a control-author panel (to
|
||||
place an absolute band and a hard-negative sister)" and notes we have none. Faulkner beside McCarthy
|
||||
makes each the other's hard negative — a METHOD upgrade, not just another roster entry.
|
||||
⚠ Messiest corpus of the three: a 446k-word `Snopes: The Hamlet, The Town, The Mansion` omnibus
|
||||
duplicates novels also present individually, and `Three Famous Short Novels` overlaps it again. That
|
||||
is the Hemingway trap (169,759 words of measured 90-96% collection duplication) — containment pass
|
||||
before anything else.
|
||||
|
||||
**2. Toni Morrison — 13 rows, 11 novels after pruning, ~818k words.**
|
||||
*The voice:* free-indirect discourse sliding between narrator and character mid-sentence, carried on
|
||||
incantatory repetition and deliberate fragments.
|
||||
Cleanest corpus shape on the list: **11 novels → 11 val units, beating Hemingway's 10.** Val units
|
||||
scale with WORK COUNT, which is the structural reason Brontë's voice axis came back underpowered at
|
||||
4 with no cheap fix. ⚠ Drop `Burn This Book` (anthology she edited) and `Playing in the Dark`
|
||||
(criticism) — same reason Krakauer's reporting does not transfer.
|
||||
|
||||
**3. Raymond Chandler — 9 rows, 7 novels + a 409k short-story omnibus, ~970k words.**
|
||||
*The voice:* clipped first-person declaratives that periodically detonate into one baroque simile,
|
||||
with dialogue carrying most of the scene.
|
||||
Fills the register gap nobody else fills — **first-person hardboiled**; the line has no first-person
|
||||
male narrator at all. Corpus is almost exactly Hemingway-sized (997k vs 958k), which was the
|
||||
decisive gate. ⚠ Drop `Essays and Reviews` — non-fiction.
|
||||
|
||||
## Held, and why
|
||||
|
||||
**Conrad** (31 works, 2.4M) is a genuine tier-1.5 if a fourth is wanted. **Melville** (10, 1.9M) has
|
||||
a superb voice but a mixed-register corpus — the cetology chapters are a different book from the
|
||||
narrative. **Austen** (12, 1.18M) is worth noting because Burrows's Delta was developed on her, so
|
||||
the instrument is known to resolve her. The romantasy cluster is a separate question entirely —
|
||||
see [[2026-09-17-romantasy-register-measured]].
|
||||
|
||||
Related: [[2026-09-17-mccarthy-split-name-leak]], [[2026-09-17-lv-bronte-gate]],
|
||||
[[2026-09-17-lv-hemingway-gate]].
|
||||
@@ -1,71 +0,0 @@
|
||||
# `[2026-09-17]` Romantasy measured as a register — it is real, we already took its best voice, and the obvious next pick is its worst
|
||||
|
||||
Prompted by the operator pushing back on a one-clause dismissal of the lane as "depth behind
|
||||
Yarros". The dismissal was taste; this is a measurement, on the gate's own instrument.
|
||||
|
||||
**Method.** Char-bigram Burrows's Delta, the same measure `voice_distance.py` gates on. ~120k words
|
||||
per author, sampled from the MIDDLE quartile of each author's largest works (front and back matter
|
||||
are not the voice), equalised so a bigger sample is not a different measurement. 400 most-frequent
|
||||
bigrams as the feature set, z-scored over 4,000-word chunks pooled across all authors.
|
||||
|
||||
**Controls first, because a between-author number without a within-author floor is unfalsifiable.**
|
||||
|
||||
```
|
||||
A-vs-A floor (two halves of the SAME author)
|
||||
Yarros 0.285 Maas 0.298 Armentrout 0.314 St. Clair 0.322 Cole 0.338
|
||||
Kenyon 0.375 Reyne 0.391
|
||||
McCarthy 0.303 Morrison 0.327 Brontë 0.209 Hemingway 0.454 <- worst, used as the bar
|
||||
|
||||
positive controls (known-distinct pairs — the instrument must separate these)
|
||||
Yarros vs McCarthy 0.862 1.9x
|
||||
Hemingway vs Brontë 0.773 1.7x
|
||||
McCarthy vs Morrison 0.675 1.5x
|
||||
Hemingway vs McCarthy 0.655 1.4x
|
||||
|
||||
romantasy, all 21 pairs median 0.537 1.2x floor (range 0.465 - 0.674)
|
||||
```
|
||||
|
||||
**The register is real but tight.** 1.2x floor against controls at 1.4-1.9x. Only one pair falls to
|
||||
1.0x, so it is not seven names for one voice.
|
||||
|
||||
⚠ **Sensitivity floor, stated because a result without one is unfalsifiable.** The 0.454 bar is
|
||||
Hemingway's, inflated by his own heterogeneous corpus (1920s-1960s, novels + stories + posthumous).
|
||||
Against the romantasy authors' OWN floors (~0.34) the same pairs read ~1.6x — control-grade. The
|
||||
truth sits between those readings and **this method cannot split it finer**. One sample per pair, no
|
||||
repeat draws: read the rank ordering as indicative, do not read small gaps at all.
|
||||
|
||||
## Two findings that survive either floor reading
|
||||
|
||||
⭐ **Yarros is the cluster OUTLIER, not a typical member.** Four of the five largest distances in the
|
||||
matrix involve her — Yarros-Kenyon 0.674, Yarros-St. Clair 0.644, Yarros-Reyne 0.637, Yarros-Maas
|
||||
0.567. **We already trained the most distinctive romantasy voice we hold**, so a second seat in the
|
||||
lane buys measurably less than the first did. That is the actual answer to "what about romantasy".
|
||||
|
||||
⭐ **Maas is the centroid, so the obvious commercial pick is the least distinctive.** Maas-Reyne
|
||||
0.465 and Maas-Cole 0.470 are the two SMALLEST distances in the whole matrix. She is the biggest
|
||||
name available (922k words) and measurably the most generic of the seven in char-bigram terms.
|
||||
Picking by sales rank picks the worst adapter.
|
||||
|
||||
## If the lane gets a second seat it is Kenyon
|
||||
|
||||
Furthest from the shipped Yarros (0.674), so it adds the most new signal — **and 27 works means 27
|
||||
val units, the best-powered gate the line could build** (Hemingway 10, McCarthy 6, Brontë 4, where 4
|
||||
is the documented structural cause of an underpowered voice axis with no cheap fix).
|
||||
|
||||
⚠ Two costs: the 27 are one series (Dark-Hunter), so the shared proper-noun space makes
|
||||
`--scope corpus` mandatory rather than optional; and a "Dark Hunter - The Dark Hunter Complete"
|
||||
omnibus sits in the catalogue rows, so the containment pass runs first.
|
||||
|
||||
**Corpus shapes for the lane** (works ≥100k chars, deduped by title):
|
||||
```
|
||||
Sherrilyn Kenyon 27 2,368,396 w Scarlett St. Clair 11 1,133,067
|
||||
Opal Reyne 14 2,466,307 Kresley Cole 10 1,037,914
|
||||
Jennifer Armentrout 6 1,179,953 Sarah J. Maas 5 922,711
|
||||
Rebecca Yarros 5 820,425 <- SHIPPED on this
|
||||
```
|
||||
⭐ Worth noting for any future bar-setting: **Yarros shipped on 5 works / 820k words.** The corpus
|
||||
bar is lower than it looks.
|
||||
|
||||
Instrument: `scratchpad/regdist.py` (screening tool, not the gate).
|
||||
|
||||
Related: [[2026-09-17-next-voice-seats]], [[2026-09-16-lv-voices-line]], [[2026-09-17-lv-bronte-gate]].
|
||||
-3
@@ -1,3 +0,0 @@
|
||||
# `[2026-09-17]` `servers/fv-ml1/ssh-target` was bare `10.251.50.54`, so `deploy-stack.sh` connected as `lkraven` and could not write the infra-ops-owned `/opt/docker/compose/`
|
||||
|
||||
**`servers/fv-ml1/ssh-target` was bare `10.251.50.54`, so `deploy-stack.sh` connected as `lkraven` and could not write the infra-ops-owned `/opt/docker/compose/`** — and lkraven's sudo on fv-ml1 needs a password, so `DEPLOY_SUDO=1` failed too. Now `infra-ops@10.251.50.54`; `--validate-only` still clean, deploy works. ⚠ Other hosts' `ssh-target` files may carry the same gap — a read-only refresh works as either user, so the fault only surfaces on a deploy.
|
||||
-3
@@ -1,3 +0,0 @@
|
||||
# `[2026-09-17]` The beat-contamination leak IS present in Hemingway — 70 of 7,094 train beats (0.96%), 0 of 200 val.
|
||||
|
||||
⭐ **The beat-contamination leak IS present in Hemingway — 70 of 7,094 train beats (0.96%), 0 of 200 val.** `scripts/r49-corpus/audit_pairs_sourcenames.py` closes the blind spot `leak_gate.py` has by construction (it reads the corpus and the renamed copies, never the generated beats). Cross-validated on real data: the fixed Brontë pairs return 0 of 3,858 and `pairs-full.CONTAMINATED` returns 15 of 792 = 1.89% with the recorded names. `--filter-out` yields a verified-clean 7,024-pair set in one command; the retrain is the operator's call. **The val split being clean is why the gate could run at all.**
|
||||
-3
@@ -1,3 +0,0 @@
|
||||
# `[2026-09-17]` The `mccarthy` register names the punctuation ON PURPOSE, and that is a gate-design call made before any McCarthy number existed.
|
||||
|
||||
⭐ **The `mccarthy` register names the punctuation ON PURPOSE, and that is a gate-design call made before any McCarthy number existed.** `eval-*.sh` drives the base control arm with the SAME system prompt via `--system-from`, and `voice_distance.py` is Burrows's Delta over character bigrams — so a tic left OUT of the register is a cheap win only the adapter can take, on a corpus measuring 0.0 quote marks per 10k against Hemingway's 838. Stating them hands them to the control too. Cost stated up front: the voice axis gets harder, and McCarthy's 276-passage val split (against Brontë's 44) is why that trade is affordable here and was not there.
|
||||
-3
@@ -1,3 +0,0 @@
|
||||
# `[2026-09-17]` The SHIPPED lv-bronte adapter emits mid-sentence line breaks at 12.46 per 1k chars, and nothing downstream looks for it.
|
||||
|
||||
⭐⭐ **The SHIPPED lv-bronte adapter emits mid-sentence line breaks at 12.46 per 1k chars, and nothing downstream looks for it.** Its corpus is 100% hard-wrapped at ~68 chars (Gutenberg plain text) and the wrap transfers: base control 0.00, ckpt475 (shipped) 12.46, ckpt925 11.79, every Hemingway arm 0.00 on a 0%-wrapped corpus. Both controls fire. `score_beats.py` passed Brontë's damage axis anyway. McCarthy is the MIXED case — The Road wrapped, the other five works not — which is worse to learn than either pure one, so `build_sft_pairs.py --reflow-hard-wraps` (DEFECT 4) fixes it at pair time, off by default. ⚠ The obvious fix, joining every interior newline, CORRUPTS 46 two-speaker exchanges whose blank line was lost — and unmarked dialogue is the one thing this adapter exists to learn. The rule splits on sentence-final punctuation and takes the cheaper error deliberately.
|
||||
-3
@@ -1,3 +0,0 @@
|
||||
# `[2026-09-17]` The two-epoch recipe is now 0 for 2 and should stop being carried forward.
|
||||
|
||||
**The two-epoch recipe is now 0 for 2 and should stop being carried forward.** Hemingway's eval minimum is step 1750, but step 850 is +0.0040 against a 0.0044 median neighbour jitter — three checkpoints inside one jitter — and 850 won every resolving axis (2.3x tighter seed spread, lower memorisation, less ran-on). Same outcome as Brontë. What IS robust on this schedule is the epoch-3 collapse: +0.0762 = **17.4x jitter**.
|
||||
-3
@@ -1,3 +0,0 @@
|
||||
# `[2026-09-17]` The v2 voice floor is now PAIRWISE, and it retroactively passes lv-bronte.
|
||||
|
||||
⭐⭐ **The v2 voice floor is now PAIRWISE, and it retroactively passes lv-bronte.** lv-bronte's ckpt475 shipped as a voice-axis FAILURE at +0.193 against a 0.251 floor contributed entirely by ckpt925 — a third arm nobody was shipping, on one outlier seed. Scored against the arm it was actually compared to the floor is 0.091 and it clears at **2.1x**. The rule was changed **prospectively**, pre-registered for lv-hemingway before any Hemingway number existed, on an argument independent of the answer: the sampling variability of a difference A−B depends on A and B, never on a third arm C. The previous session found the defect and deliberately declined to exploit it; this follows from fixing it. lv-hemingway passes under **both** rules, so its verdict does not lean on the change. Caveats amended append-only in the compose, the NFS README and the gate record. Commits `0bb4938` `2e9b118`.
|
||||
-3
@@ -1,3 +0,0 @@
|
||||
# `[2026-09-17]` `triage_disposition = 'accepted'` in the Kvasir catalogue does NOT mean the extraction succeeded.
|
||||
|
||||
⚠ **`triage_disposition = 'accepted'` in the Kvasir catalogue does NOT mean the extraction succeeded.** Blood Meridian's epub row holds 1,167 words of a 117,000-word book, The Crossing's 222 of 150,000 — real prose, real titles, accepted. Faulkner's *The Mansion* is 39 words. `near_dup_pairs` holds ONE row in the entire 1,284-work library and is blind to a fragment beside its full sibling. **Word-count every master before trusting a row**, and note that word count alone cannot tell a truncated novel from a legitimately short work.
|
||||
@@ -18,3 +18,9 @@
|
||||
2. Free ~+1.1–1.5 GB on GPU 0 by trimming the `vllm-gen-small` util. ⚠ MEASURE the resulting free memory; util does not predict resident VRAM.
|
||||
3. Cut over with the old seat kept as the rollback.
|
||||
4. Re-measure live on GPU 0.
|
||||
|
||||
**DONE 2026-10-01 ~0126 PT (infra-hermes; `stacks/parakeet-nemo`, de6ea32), infra-ops audit PASSED 0137.**
|
||||
- p50 on GPU 0: 33 / 36 / 42 / 71 ms, against 187 / 308 / 626. WER: clean 1.965, other 3.026.
|
||||
- A 714 s file returns 200 in 0.85 s (360 s windows). The steady state is 3,582 MiB, and GPU 0 Free is 385.
|
||||
- gen-small util 0.48 → 0.36: it took three boots and ~34 min of downtime at midnight, with zero LiteLLM errors. Its `.env` holds 0.33 for the next restart. The KV is byte-pinned and unchanged.
|
||||
- infra-hermes caught that httptools 0.8.0 emits `HTTP/1.1 200\x00OK`, which LiteLLM rejects; it is pinned to `--http h11`.
|
||||
|
||||
Reference in New Issue
Block a user