`leak_gate.py` scans `\b(Surface)\b`. Any character inserted inside a name defeats
that pattern outright, so a mangled occurrence is unrenameable by rename.py AND
unreportable by the gate. lv-mccarthy's 2026-09-17 tree passed at "0 of 75
renameable and 0 of 37 sub-threshold" while carrying 13 occurrences of Bell,
Chigurh, Moss, Toadvine and Glanton in all six copies:
B ell C higurh M oss T oadvine a small-caps drop cap kept as its own token
Toad-vine Glan-ton a print line-break hyphen kept by the extractor
Every visible occurrence HAD been renamed, which is what made the residue invisible
to a spot-read. Fixed at three levels, all three of which must stay:
build_corpus_mccarthy.py rules 4 and 5 repair the source text — 32 split initials
with a lowercase remainder, 5 hyphen-split names, each with an expected count so a
master change fails the build. Rule 4's letter class is consonants only: `I` opens
1,966 paragraphs, `A` 143 and `Y` 32 (Spanish `y`); folding any would corrupt 2,141
lines to fix 32.
leak_gate.py gains a separator-tolerant pass with its own positive and negative
controls, and it FAILS the gate. Validated against the pre-fix tree: reports all
five surfaces, exits 1. Its fragment filter is what makes it usable — a naive scan
returns 18 false positives on Hemingway (`God damn`, `I run`) against 3 real ones;
requiring one fragment to be a non-word of the corpus cleared all 18 and kept all 3.
The whole D1→D3 chain is reproduced byte-identically before and after, so the fix
is the only delta: 6 works, the entity map, the final map and all 36 copy files.
Cross-checked on the shipped corpora: lv-bronte is clean of this class, lv-hemingway
carries 3 (`Primi tivo`, `Pasionar ia`, `Chi cote`) and is live on fv-ml1.
Also in build_sft_pairs.py, both needed before lv-mccarthy's pairs:
DEFECT 4, hard-wrap reflow. Measured on the SHIPPED lv-bronte adapter, which emits
mid-sentence line breaks at 12.46 per 1k chars against 0.00 for its own base control
and 0.00 for every Hemingway arm. McCarthy is the mixed case — The Road is wrapped,
the other five works are not — so the corpus teaches the break as a coin flip. The
obvious fix (join every interior newline) corrupts 46 two-speaker exchanges whose
blank line was lost, and unmarked dialogue is the one thing this adapter exists to
learn; the rule splits on sentence-final punctuation instead and takes the cheaper
error. Self-targeting and off by default, so every shipped pair set is unchanged.
A `mccarthy` register, which names the punctuation deliberately: the eval drives the
base control arm with this same prompt, so tics left out of it are a surface trick
only the adapter can perform, and delta_cb is a character-bigram measure.
drop_leading_heading now also consumes Blood Meridian's dash-separated chapter
arguments — 131 paragraphs, 0 in every other work of all three corpora.
And a RUNBOOK, because the D1→D3 session recorded nothing and the chain had to be
recovered by rebuilding candidates and matching sha256 against the artifacts on disk.
637 lines
36 KiB
Python
637 lines
36 KiB
Python
"""BabyYarros Option C — build instruction→response SFT pairs from the gated corpus.
|
||
|
||
The architecture question was settled on 2026-09-11: the adapted completion carrier has
|
||
the voice and cannot take direction; the instruct model takes direction and has no voice.
|
||
Skaldsong needs both, which means the corpus has to be rebuilt as instruction→response
|
||
pairs and trained through the chat template rather than as raw continuation text.
|
||
|
||
⭐ THE PROPERTY THAT MAKES THIS SAFE: the model writes the INSTRUCTION, never the
|
||
RESPONSE. Every response is real (renamed) Yarros prose, so the voice is inherited from
|
||
the corpus and not synthesised. Only the beat is machine-made, and a beat is a summary —
|
||
the easiest job in the building. A pipeline that generated the prose would be training
|
||
the carrier on a 27B's pastiche of Yarros, which is the opposite of the point.
|
||
|
||
⚠ PASSAGES, NOT PARAGRAPHS. Measured on copy0: the median Yarros paragraph is 39 words
|
||
(p90 67) against the product's 90-140 band, so one paragraph is nowhere near one product
|
||
output. Accumulating consecutive paragraphs to the band gives 6,992 passages at median
|
||
104 words / 4 paragraphs — in-band by construction, and multi-paragraph, which is also
|
||
the continuity requirement the stitching failure raised (independently-generated
|
||
paragraphs drift in POV; the pair corpus has to contain multi-paragraph examples).
|
||
|
||
⚠ ONE COPY ONLY. The rename produced 6 copies under different name maps. Six copies of
|
||
the same passage would be six near-duplicate responses differing only in proper nouns,
|
||
which is corpus inflation, not augmentation, at pair-training scale.
|
||
|
||
⚠ TRAIN SPLIT ONLY. The val split stays untouched so held-out loss remains comparable to
|
||
the raw-text arms.
|
||
|
||
Generator is the local `gen` seat: free, and abliterated, which matters here for a
|
||
non-obvious reason — the corpus contains sex and violence, and a refusing generator would
|
||
silently drop exactly the passages where the voice is most distinctive, biasing the
|
||
dataset toward its tamest regions. A refusal is a sampling bias, not just a gap.
|
||
"""
|
||
from __future__ import annotations
|
||
import argparse, json, random, re, sys, time, urllib.error, urllib.request
|
||
from pathlib import Path
|
||
|
||
GATEWAY = "http://10.250.50.70:4000"
|
||
|
||
# Matches gen_beats_chat_yarros.py verbatim so the pairs are trained in the SAME shape the
|
||
# evaluation harness measures. A pair corpus trained under a different system prompt than
|
||
# the eval drives would confound the carrier change with a prompt change.
|
||
# ⚠ DEFECT 1 of the BabyYarros pair build: the system prompt said "ONE paragraph" while every
|
||
# response was a MEDIAN OF FOUR. The carrier believed the data, correctly, and then looked
|
||
# broken against a harness that truncated at the first blank line. The prompt now describes
|
||
# what the response actually is. Both are kept: `SYS` is frozen so the shipped BabyYarros
|
||
# adapter can still be reproduced and re-evaluated under the prompt it was trained on.
|
||
SYS = ("You expand a single story beat into ONE paragraph of prose in the manner of Rebecca "
|
||
"Yarros — contemporary first-person PRESENT-tense narration, emotionally charged, sensory "
|
||
"and physical, the voice of new-adult romantasy. Render the beat itself; do not move past "
|
||
"it, do not add a new scene, do not comment. Output the paragraph only, 90–140 words.")
|
||
|
||
SYS_PASSAGE = ("You expand a single story beat into a SHORT PASSAGE of prose in the manner of "
|
||
"{author} — {register}. The passage may run to several paragraphs and should read "
|
||
"as continuous scene, not a summary. Render the beat itself; do not move past it, "
|
||
"do not begin a new scene, do not comment, do not write a chapter heading. Output "
|
||
"the prose only, 90–140 words.")
|
||
|
||
REGISTERS = {
|
||
"yarros": ("Rebecca Yarros", "contemporary first-person PRESENT-tense narration, emotionally "
|
||
"charged, sensory and physical, the voice of new-adult romantasy"),
|
||
"hemingway": ("Ernest Hemingway", "spare declarative sentences, concrete physical detail, "
|
||
"heavy unattributed dialogue, feeling shown through action and omission rather "
|
||
"than stated"),
|
||
# Brontë is the far end of the same axis from Hemingway, and the register has to say
|
||
# so or the beat-writer produces modern summary prose that the passages never match:
|
||
# long periodic sentences, an explicitly retrospective first person, and moral
|
||
# weather carried in the landscape rather than in the dialogue.
|
||
"bronte": ("Charlotte Brontë", "mid-nineteenth-century first-person retrospective narration, "
|
||
"long periodic sentences built on semicolons and dashes, heightened interior "
|
||
"analysis and moral reflection, occasional direct address to the reader, and "
|
||
"physical setting — Yorkshire weather, schoolrooms, Belgian pensionnats — "
|
||
"rendered with the feeling it carries"),
|
||
# ⭐ McCarthy is the first register in this map that names PUNCTUATION, and that is a
|
||
# GATE-DESIGN choice made before any McCarthy number existed, not a description choice.
|
||
# The eval harness drives the base (unadapted) control arm with this same prompt via
|
||
# `gen_beats_chat_yarros.py --system-from <pairs provenance>`, and `voice_distance.py` is
|
||
# Burrows's Delta over CHARACTER BIGRAMS. An adapter that learns only "emit no quotation
|
||
# marks" moves delta_cb a long way without having learned a sentence -- and on a corpus
|
||
# measuring 0.0 quote marks per 10k words against Hemingway's 838 that is the single
|
||
# cheapest available trick. Stating the tics here hands them to the control arm too, so
|
||
# the adapter earns no delta for them and the remaining gap is attributable to sentence
|
||
# structure, which is what the axis claims to measure. lv-mccarthy D1 pre-registered a
|
||
# punctuation-normalised SECONDARY read for exactly this risk; this closes it in the
|
||
# PRIMARY read as well. Cost, stated up front: the voice axis gets harder, and on an
|
||
# underpowered fixture that risks a Brontë-style marginal result. McCarthy's val split
|
||
# yields 269 in-band passages against Brontë's 44, which is the reason that trade is
|
||
# affordable here and was not there.
|
||
# ⚠ Deliberately NOT mentioned: the untranslated Spanish dialogue of the Border Trilogy.
|
||
# It is a property of 3 of the 6 works, not of the voice, and inviting an LLM to produce
|
||
# Spanish on a beat that has none is damage rather than register.
|
||
"mccarthy": ("Cormac McCarthy", "third-person past-tense narration held at the surface of "
|
||
"things, with no access to what anyone thinks; long sentences strung together "
|
||
"on `and`, set against clipped fragment paragraphs; dialogue carried WITHOUT "
|
||
"quotation marks, each speech its own paragraph, attributed plainly or not at "
|
||
"all; contractions written with no apostrophe — `dont`, `aint`, `wont`, "
|
||
"`didnt`; terrain, animals, weather and tools named concretely and "
|
||
"technically; violence and landscape rendered flatly and without comment"),
|
||
}
|
||
|
||
CONTEXT_BLOCK = """The passage is preceded by this, for reference only. Do NOT write a beat for it — it is
|
||
here so you resolve names and pronouns correctly.
|
||
|
||
PRECEDING:
|
||
{prev}
|
||
|
||
"""
|
||
|
||
BEAT_PROMPT = """{context}Below is a passage from a novel. Write the single story BEAT that a writer would have been given to produce it.
|
||
|
||
Rules:
|
||
- ONE sentence, 8 to 30 words, present tense, third person.
|
||
- Name the characters who appear, using the names exactly as the passage spells them.
|
||
- The passage is first-person and its narrator is usually UNNAMED in it. Call that person
|
||
"she" or "he" as the passage implies. Never write "the narrator", "the speaker" or "I".
|
||
- State WHAT HAPPENS — the action, the turn, the decision, the reveal. Not how it is written.
|
||
- Do NOT describe the passage ("this excerpt shows..."), do NOT mention prose, style, tone or the author.
|
||
- Do NOT reuse distinctive phrases from the passage. Say the event in your own plain words.
|
||
- Output the sentence alone, with no label, no quotes and no preamble.
|
||
|
||
PASSAGE:
|
||
{passage}"""
|
||
|
||
ABBREV = re.compile(r"\b(Mr|Mrs|Ms|Dr|St|Sr|Jr|Lt|Col|Gen|Capt|Sgt|Prof|vs|etc|No)\.$")
|
||
|
||
META = re.compile(r"\b(passage|excerpt|paragraph|prose|narrat(?:or|ion)|the (?:author|text|scene) (?:is|describes)|this (?:scene|chapter))\b", re.I)
|
||
|
||
# ⚠ MEASURED, and it was a composition bias rather than a nuisance: of 23 `meta` rejects in
|
||
# a 150-passage diagnostic, **22 were the single word "narrator"** and 1 was the story-word
|
||
# "passage". Those beats were otherwise clean ("Miguel shouts a warning as the venin
|
||
# advances, while the dark wielder taunts the narrator..."). The model reaches for "the
|
||
# narrator" precisely when the POV character is an ACTOR, so rejecting on it silently drops
|
||
# action scenes and keeps the ones where she only observes -- a systematic skew in what the
|
||
# carrier would learn to render, invisible in any spot-read of the kept pairs. The fix is a
|
||
# single bounded RETRY with the correction restated, not a substitution: the generator knows
|
||
# the character's gender from the prose and this session does not, and guessing it wrong
|
||
# across 780k words of first-person narration is the worse failure.
|
||
NARRATOR_ONLY = re.compile(r"\bnarrat(?:or|ion)\b", re.I)
|
||
RETRY_NOTE = ("\n\nYour previous answer used the word \"narrator\". Rewrite it referring to that "
|
||
"person as \"she\" or \"he\" — whichever the passage implies — and change nothing else.")
|
||
|
||
|
||
def ctx_block(prev: str | None) -> str:
|
||
"""The beat generator sees the PRECEDING passage even when the pair will not carry it.
|
||
|
||
Found in the positive control: passage [4] of nova ch16 is genuinely ambiguous about
|
||
who is buried and who is digging, and the hand beat and the model beat disagreed for
|
||
exactly that reason. A mid-scene passage can underdetermine its own referents; showing
|
||
the generator the previous passage fixes the pronouns without putting anything in the
|
||
instruction that the response does not support.
|
||
"""
|
||
return CONTEXT_BLOCK.format(prev=prev) if prev else ""
|
||
|
||
|
||
def post(path: str, payload: dict, key: str, timeout: int = 120) -> dict:
|
||
req = urllib.request.Request(
|
||
GATEWAY + path, data=json.dumps(payload).encode(),
|
||
headers={"Content-Type": "application/json", "Authorization": f"Bearer {key}"})
|
||
with urllib.request.urlopen(req, timeout=timeout) as r:
|
||
return json.loads(r.read())
|
||
|
||
|
||
HEADING_MAX_WORDS = 6
|
||
# A dash-separated fragment list: the shape of a chapter argument, not of prose.
|
||
ARGUMENT_LIST = re.compile(r"\s[-\u2013\u2014]\s")
|
||
|
||
|
||
def drop_leading_heading(paras: list[str], enabled: bool) -> list[str]:
|
||
"""Drop a unit's opening block when it is a bare chapter heading.
|
||
|
||
⚠ DEFECT 3 of the BabyYarros pair build, and it is far worse on Hemingway: 20 of 600
|
||
Yarros responses carried a chapter heading (3.3%), but **315 of 318 Hemingway units open
|
||
with one** because these editions set `I`, `27` or a story title on its own line. The
|
||
corpus builder KEEPS those headings deliberately -- they are part of the form the voice
|
||
lives in and dropping them would teach the model that chapters begin mid-scene -- so the
|
||
removal belongs here, at pair construction, not upstream in the corpus.
|
||
|
||
A response beginning `CHAPTER SIXTY-SIX` teaches the carrier to emit chapter headings when
|
||
asked for prose, which is exactly what the pilot adapter did.
|
||
"""
|
||
if not enabled or not paras:
|
||
return paras
|
||
if len(paras[0].split()) <= HEADING_MAX_WORDS:
|
||
paras = paras[1:]
|
||
# ⭐ AND THEN THE CHAPTER ARGUMENT, which is a heading that is 40 words long.
|
||
# Blood Meridian sets each chapter's argument as a dash-separated list of title-case
|
||
# fragments under the roman numeral -- `Desert castaways - The backtrack - A hideout -
|
||
# The wind takes a side - The judge returns` -- hard-wrapped across several short
|
||
# paragraphs. The <=6-word rule drops the `XXI` above it and leaves the argument, so a
|
||
# chapter-opening passage trains the carrier to emit a dash-separated summary before the
|
||
# prose. Same harm as the bare heading, a different shape.
|
||
# Measured: 131 paragraphs, ALL in blood-meridian; 0 in the other five McCarthy works,
|
||
# 0 in all ten Hemingway works and 0 in all four Brontë works, so this extension is a
|
||
# byte no-op on every pair set already shipped under this flag.
|
||
while paras and ARGUMENT_LIST.search(paras[0]) and len(paras[0].split()) <= 20:
|
||
paras = paras[1:]
|
||
return paras
|
||
|
||
|
||
# A line that ends a sentence, after detached terminal punctuation is closed up: this corpus
|
||
# carries 151 occurrences of `boxcutter .` in The Road, and a naive test reads those as
|
||
# mid-sentence and joins across a real paragraph break.
|
||
SENT_FINAL = re.compile(r'[.!?"\u201d]$')
|
||
DETACHED_PUNCT = re.compile(r"\s+([.,;:!?])")
|
||
|
||
|
||
def reflow_hard_wraps(paras: list[str], enabled: bool) -> tuple[list[str], int, int]:
|
||
"""Undo an extraction's HARD LINE WRAPPING inside a paragraph. Returns (paras, joined, split).
|
||
|
||
⚠ DEFECT 4 of this pair build, and unlike the first three it was MEASURED ON A SHIPPED
|
||
ADAPTER before it was fixed here. `lv-bronte` trained on a corpus whose four works are
|
||
100% hard-wrapped at ~68 characters (Gutenberg plain text), and the wrap transfers
|
||
straight through to the product:
|
||
|
||
arm intra-para breaks MID-SENTENCE per 1k chars
|
||
bronte base (control) 0 0 0.00
|
||
bronte ckpt475 (SHIPPED) 1,251 1,219 12.46
|
||
bronte ckpt925 1,134 1,087 11.79
|
||
hemingway base/ckpt*/all 58/0/0 0 0.00 <- 0% wrapped corpus
|
||
|
||
Both controls fire: the base arms emit none, so the carrier is not the source, and the
|
||
Hemingway arms emit none, so the instrument is not manufacturing signal. `score_beats.py`
|
||
does not look for this and passed Brontë's damage axis anyway (ran-on +0.15 / 0.400 floor),
|
||
so nothing downstream would have reported it.
|
||
|
||
McCarthy is the MIXED case, which is worse to learn than either pure one: The Road is
|
||
hard-wrapped (3,587 intra-paragraph newlines, 60.9 per 1k words) and the other five works
|
||
have exactly ZERO, so the corpus teaches the break as a coin flip. The transform is
|
||
therefore SELF-TARGETING -- a paragraph with no interior newline is returned untouched --
|
||
and needs no per-work special-casing.
|
||
|
||
⚠⚠ THE OBVIOUS FIX -- join every interior newline with a space -- CORRUPTS THE THING THIS
|
||
ADAPTER EXISTS TO LEARN. 46 paragraphs in The Road are two-speaker exchanges whose blank
|
||
line was lost, and McCarthy's dialogue is unmarked, so the speaker boundary IS the
|
||
paragraph boundary:
|
||
|
||
"And we're carrying the fire." || 'Yes.'
|
||
"I'm going to blow out the lamp." || 'Is that okay?'
|
||
'Take me with you, the boy said.' || 'He looked as if he was going to cry.'
|
||
|
||
Joining those puts two speakers in one paragraph and unmarked dialogue stops parsing.
|
||
|
||
⚠ A WRAP-WIDTH test cannot separate them either, and this was measured rather than
|
||
assumed: the wrap is POSITION-DEPENDENT -- first lines of a paragraph break at ~30
|
||
characters and later lines at ~78 -- so `We're not the first ones here.` (29) is
|
||
geometrically indistinguishable from a full line, and a max-line-length rule splits 347
|
||
paragraphs of which the majority are genuine wraps (`If you died I would want to die` ||
|
||
`too.`).
|
||
|
||
So the discriminator is PUNCTUATION, and the error budget is deliberately asymmetric:
|
||
|
||
- Li does NOT end sentence-final -> a wrap. JOIN with a space. 3,315 of 3,587 (92.4%),
|
||
and this is the ENTIRE defect class: a mid-sentence newline is the thing the Brontë
|
||
adapter learned to emit.
|
||
- Li DOES end sentence-final -> promote the newline to a paragraph BREAK. 272 cases.
|
||
Correct for all 46 dialogue exchanges; for some fragmentary narration it inserts a
|
||
paragraph break the book does not have.
|
||
|
||
That residual is the cheap error on purpose. A spurious break inside McCarthy narration is
|
||
invisible -- the page is already full of one-line fragment paragraphs -- while a merged
|
||
pair of speakers is a form error in the corpus's most distinctive feature. Splitting
|
||
where the book does not costs paragraphing; joining where the book does not costs the
|
||
voice.
|
||
"""
|
||
if not enabled:
|
||
return paras, 0, 0
|
||
out: list[str] = []
|
||
joined = promoted = 0
|
||
for para in paras:
|
||
lines = [x.strip() for x in para.split("\n") if x.strip()]
|
||
if len(lines) < 2:
|
||
out.append(para)
|
||
continue
|
||
buf = [lines[0]]
|
||
for nxt in lines[1:]:
|
||
if SENT_FINAL.search(DETACHED_PUNCT.sub(r"\1", buf[-1])):
|
||
out.append(" ".join(buf))
|
||
buf = [nxt]
|
||
promoted += 1
|
||
else:
|
||
buf.append(nxt)
|
||
joined += 1
|
||
out.append(" ".join(buf))
|
||
return out, joined, promoted
|
||
|
||
|
||
def chunk(corpus: Path, lo: int, hi: int, split: str, drop_heading: bool = False,
|
||
reflow: bool = False) -> list[dict]:
|
||
"""Accumulate consecutive paragraphs into product-band passages, per chapter.
|
||
|
||
Chapter-bounded so a passage never straddles a chapter break. The trailing buffer of
|
||
each chapter is DROPPED rather than emitted short -- an out-of-band response would
|
||
teach the length the product is trying to hold.
|
||
"""
|
||
out = []
|
||
reflow_joined = reflow_split = 0
|
||
for f in sorted(corpus.glob("*.copy0.jsonl")):
|
||
for line in f.read_text(encoding="utf-8").splitlines():
|
||
d = json.loads(line)
|
||
if d.get("split") != split:
|
||
continue
|
||
paras = [p.strip() for p in re.split(r"\n\s*\n", d["text"]) if p.strip()]
|
||
paras, _j, _s = reflow_hard_wraps(paras, reflow)
|
||
reflow_joined += _j
|
||
reflow_split += _s
|
||
paras = drop_leading_heading(paras, drop_heading)
|
||
buf, n, prev = [], 0, None
|
||
for p in paras:
|
||
buf.append(p); n += len(p.split())
|
||
if n >= lo:
|
||
if n <= int(hi * 1.6):
|
||
out.append({"work": d["work"], "chapter": d["chapter"], "split": d["split"],
|
||
"response": "\n\n".join(buf), "words": n, "prev": prev})
|
||
prev = "\n\n".join(buf)
|
||
else:
|
||
prev = None # oversized run dropped; context would be a lie
|
||
buf, n = [], 0
|
||
if reflow:
|
||
print(f"[reflow] {reflow_joined} wrapped lines rejoined, "
|
||
f"{reflow_split} interior newlines promoted to paragraph breaks", flush=True)
|
||
return out
|
||
|
||
|
||
def ngrams(text: str, n: int) -> set[str]:
|
||
w = re.findall(r"[a-z']+", text.lower())
|
||
return {" ".join(w[i:i + n]) for i in range(max(0, len(w) - n + 1))}
|
||
|
||
|
||
def vet(beat: str, passage: str, overlap_n: int,
|
||
source_pat: "re.Pattern[str] | None" = None) -> tuple[str | None, str]:
|
||
"""Return (clean_beat, reason). reason is '' on accept.
|
||
|
||
Every rejection is a dataset defect that would otherwise train silently:
|
||
echo -- a beat quoting the passage teaches the model to copy its instruction
|
||
back, not to expand it. This is the one that would look fine in a
|
||
spot-read and poison the whole run.
|
||
meta -- 'this passage shows...' is a description of text, not a story beat
|
||
length -- a 60-word beat is a summary; a 4-word beat is a title
|
||
sourcename -- ⭐ the beat names a character the RENAME removed. See below.
|
||
|
||
⚠⭐ `sourcename` IS A LEAK THE CORPUS GATE STRUCTURALLY CANNOT SEE, and it was found
|
||
on lv-bronte (2026-09-16). The rename strips the author's names from the prose and
|
||
leak_gate.py proves they are gone — but the beat is written by an LLM that READ THE
|
||
PASSAGE, and if it recognises the book it supplies the canonical names from its own
|
||
training. Measured on Brontë: 13 of the first 714 beats (1.8%) named Rochester (x6),
|
||
Jane (x3), Brocklehurst (x2), Beck, Fairfax, Helen, Burns, Eyre, Reed and Rivers,
|
||
while 0 of 714 RESPONSES did — the rename was perfect and the instruction side was
|
||
not. One beat read `Saoirse confirms Rochester's flaws`, mixing a renamed name and a
|
||
canonical one in a single sentence, which is the mechanism in miniature.
|
||
|
||
The beat is the INSTRUCTION half of the pair, so training on it re-teaches exactly the
|
||
inventions the rename pipeline exists to remove, and the corpus gate never looks at it:
|
||
the gate reads the corpus and the renamed copies, never the generated beats.
|
||
|
||
⚠ Exposure scales with how well the generator knows the book, so it is WORST for
|
||
public-domain classics and mildest for recent work — which is precisely why Yarros and
|
||
Hemingway did not surface it and Brontë did. Do not read their clean runs as evidence
|
||
this cannot happen; pass --source-entities on every corpus.
|
||
"""
|
||
beat = " ".join(beat.strip().split())
|
||
beat = re.sub(r'^(?:beat|answer)\s*[:\-]\s*', '', beat, flags=re.I).strip().strip('"“”')
|
||
if not beat:
|
||
return None, "empty"
|
||
# First sentence only -- the model sometimes adds a second.
|
||
# ⚠ An abbreviation is not a sentence end. The naive `.+?[.!?]\s` truncated
|
||
# "Jeremias arrives with Mr. Derrick and he sizes him up" to "...with Mr." -- caught in the
|
||
# Hemingway positive control, and it would have been near-invisible in the built pairs
|
||
# because the result is still a short grammatical-looking fragment. Hemingway is full of
|
||
# `Mr. Singh`, `Mr. Bobby`, `Mr. Derrick`, so this fires constantly on this corpus and
|
||
# essentially never on Yarros -- a defect one corpus exposes and another hides.
|
||
m = re.match(r"^(.+?[.!?])(?:\s|$)", beat)
|
||
while m and ABBREV.search(m.group(1)):
|
||
nxt = re.match(r"^(.+?[.!?])(?:\s|$)", beat[m.end():])
|
||
if not nxt:
|
||
m = None
|
||
break
|
||
m = re.match(r"^(.{%d,}?[.!?])(?:\s|$)" % (m.end() + 1), beat)
|
||
if not m:
|
||
break
|
||
if m:
|
||
beat = m.group(1).strip()
|
||
nw = len(beat.split())
|
||
if nw < 6 or nw > 34:
|
||
return None, f"length({nw})"
|
||
if META.search(beat):
|
||
return None, "meta"
|
||
if ngrams(beat, overlap_n) & ngrams(passage, overlap_n):
|
||
return None, f"echo({overlap_n}gram)"
|
||
if source_pat is not None:
|
||
m = source_pat.search(beat)
|
||
if m:
|
||
return None, f"sourcename({m.group(1)})"
|
||
return beat, ""
|
||
|
||
|
||
NAME = re.compile(r"(?<![.!?“\"]\s)(?<!^)\b([A-Z][a-z]{2,})\b")
|
||
|
||
|
||
def names_not_in_response(beat: str, response: str) -> list[str]:
|
||
"""Names the BEAT introduces that the RESPONSE never spells out.
|
||
|
||
Reported, deliberately NOT rejected. Giving the generator the preceding passage fixed
|
||
the referent ambiguity the positive control exposed -- it corrected a hand-written beat
|
||
that had the buried and the digging characters backwards -- but it also lets a beat name
|
||
someone who appears only in that preceding text. Rejecting on this would drop exactly the
|
||
passages whose POV character is unnamed, which is a sampling bias dressed as a guard
|
||
(same argument as the refusing-generator note above). So it is counted and surfaced in
|
||
provenance, and the rate decides whether it needs a rule.
|
||
"""
|
||
return sorted({n for n in NAME.findall(beat) if n not in response})
|
||
|
||
|
||
def main() -> int:
|
||
ap = argparse.ArgumentParser()
|
||
ap.add_argument("--corpus", required=True, help="dir holding *.copy0.jsonl")
|
||
ap.add_argument("--out", required=True)
|
||
ap.add_argument("--key-file", default=None, help="file holding the gateway key")
|
||
ap.add_argument("--key", default=None)
|
||
ap.add_argument("--model", default="gen")
|
||
ap.add_argument("--n", type=int, default=600)
|
||
ap.add_argument("--lo", type=int, default=90)
|
||
ap.add_argument("--hi", type=int, default=150)
|
||
ap.add_argument("--split", default="train")
|
||
ap.add_argument("--seed", type=int, default=4919)
|
||
# ⚠ DEFECT 2 of the BabyYarros pair build, and the one that shaped its output most.
|
||
# The chunker starts each passage where the last one ended, so a passage opens MID-SCENE
|
||
# with lead-in the beat does not describe -- the beat summarises the whole span. At 0.4 the
|
||
# pair usually withheld the preceding text, so the carrier learned "open somewhere
|
||
# unrelated, then drift toward the beat", which is exactly what the pilot generations did:
|
||
# the beat material landed in block 3 or 4.
|
||
# Setting this to 1.0 motivates the opening instead of hiding it, and it matches how
|
||
# Skaldsong actually calls the model -- a stitcher always has the previous passage. A
|
||
# passage at the START of a unit has no predecessor and stays a genuine cold open, which is
|
||
# correct: a chapter's first passage IS one, and its beat legitimately describes its start.
|
||
ap.add_argument("--context-frac", type=float, default=0.4,
|
||
help="fraction of pairs carrying the preceding passage as context; "
|
||
"use 1.0 for the corrected recipe (see DEFECT 2)")
|
||
ap.add_argument("--overlap-n", type=int, default=6)
|
||
ap.add_argument("--source-entities", default=None,
|
||
help="entities json for the UNRENAMED source. Any beat naming a surface "
|
||
"from it is rejected as `sourcename`. Pass this on every corpus: the "
|
||
"generator reads the passage and will supply canonical names from its "
|
||
"own memory of the book if it recognises it, which the corpus leak "
|
||
"gate cannot see because it never reads the generated beats.")
|
||
ap.add_argument("--register", choices=sorted(REGISTERS), default=None,
|
||
help="DEFECT 1 fix: emit the PASSAGE system prompt for this author instead "
|
||
"of the frozen one-paragraph Yarros prompt. Omit to keep the original.")
|
||
ap.add_argument("--temperature", type=float, default=0.3)
|
||
ap.add_argument("--control-out", default=None,
|
||
help="write the first --control-n passages out for hand-written beats")
|
||
ap.add_argument("--control-n", type=int, default=10)
|
||
ap.add_argument("--dump-passages", action="store_true", help="chunk and report, generate nothing")
|
||
ap.add_argument("--drop-leading-heading", action="store_true",
|
||
help="DEFECT 3 fix: drop a unit's opening block when it is a bare chapter "
|
||
"heading. OFF by default so the BabyYarros pair set stays byte-reproducible.")
|
||
ap.add_argument("--reflow-hard-wraps", action="store_true",
|
||
help="DEFECT 4 fix: rejoin lines an extraction hard-wrapped mid-sentence "
|
||
"inside a paragraph. MEASURED on the shipped lv-bronte adapter, which "
|
||
"emits mid-sentence line breaks at 12.5 per 1k chars against 0.00 for "
|
||
"its own base control. Self-targeting -- a no-op on any work whose "
|
||
"paragraphs hold no interior newline, which is 5 of 6 McCarthy works "
|
||
"and all 10 Hemingway works. OFF by default so the Yarros, Hemingway "
|
||
"and Brontë pair sets stay byte-reproducible.")
|
||
ap.add_argument("--control-in", default=None,
|
||
help="JSON of hand-written beats [{idx,beat}]; generate beats for the SAME "
|
||
"passages and print side by side. The positive control -- a beat "
|
||
"generator nobody checked against a human can produce a clean-looking "
|
||
"dataset that teaches the wrong mapping, and nothing downstream would show it.")
|
||
a = ap.parse_args()
|
||
|
||
sys_prompt = SYS
|
||
if a.register:
|
||
author, register = REGISTERS[a.register]
|
||
sys_prompt = SYS_PASSAGE.format(author=author, register=register)
|
||
key = a.key or (Path(a.key_file).read_text().strip() if a.key_file else None)
|
||
passages = chunk(Path(a.corpus), a.lo, a.hi, a.split, a.drop_leading_heading,
|
||
a.reflow_hard_wraps)
|
||
print(f"[chunk] {len(passages)} passages in split={a.split}, band {a.lo}-{a.hi}", flush=True)
|
||
if not passages:
|
||
print("REFUSING: no passages -- wrong corpus dir or split", file=sys.stderr)
|
||
return 2
|
||
|
||
random.seed(a.seed)
|
||
random.shuffle(passages)
|
||
|
||
if a.control_out:
|
||
ctl = [{"idx": i, "work": p["work"], "chapter": p["chapter"], "response": p["response"],
|
||
"beat_handwritten": ""} for i, p in enumerate(passages[:a.control_n])]
|
||
Path(a.control_out).write_text(json.dumps(ctl, indent=2, ensure_ascii=False), encoding="utf-8")
|
||
print(f"[control] wrote {len(ctl)} passages to {a.control_out} for hand-written beats")
|
||
|
||
if a.dump_passages:
|
||
return 0
|
||
if not key:
|
||
print("REFUSING: no gateway key (--key or --key-file)", file=sys.stderr)
|
||
return 2
|
||
|
||
# The alias is not the model. `gen` has pointed at different concrete backends over
|
||
# time; counting by an alias once inflated an exposure figure 4.7x on this fleet. The
|
||
# build fingerprint is resolved at START and again at END and both go in provenance --
|
||
# a seat repointed mid-run would otherwise be invisible in the artefact.
|
||
def fingerprint() -> str:
|
||
try:
|
||
r = post("/v1/chat/completions", {"model": a.model, "max_tokens": 1,
|
||
"messages": [{"role": "user", "content": "ok"}]}, key, 60)
|
||
return r.get("system_fingerprint") or "unknown"
|
||
except Exception as e:
|
||
return f"unresolved:{e}"
|
||
|
||
fp_start = fingerprint()
|
||
print(f"[seat] {a.model} fingerprint at start: {fp_start}", flush=True)
|
||
|
||
if a.control_in:
|
||
hand = {h["idx"]: h["beat"] for h in json.loads(Path(a.control_in).read_text(encoding="utf-8"))}
|
||
agree = 0
|
||
for i in sorted(hand):
|
||
psg = passages[i]
|
||
r = post("/v1/chat/completions", {
|
||
"model": a.model, "temperature": a.temperature, "max_tokens": 80,
|
||
"messages": [{"role": "user", "content": BEAT_PROMPT.format(passage=psg["response"], context=ctx_block(psg["prev"]))}],
|
||
}, key)
|
||
beat, why = vet(r["choices"][0]["message"]["content"] or "", psg["response"], a.overlap_n)
|
||
print("=" * 72)
|
||
print(f"[{i}] {psg['work']} ch{psg['chapter']}")
|
||
print(f" HAND : {hand[i]}")
|
||
print(f" MODEL: {beat if beat else '<<REJECTED: ' + why + '>>'}")
|
||
if beat:
|
||
agree += 1
|
||
print("=" * 72)
|
||
print(f"[control] {agree}/{len(hand)} model beats passed the guards; "
|
||
f"judge the EVENT match by reading, not by this count")
|
||
return 0
|
||
|
||
out_path = Path(a.out)
|
||
source_pat = None
|
||
if a.source_entities:
|
||
_e = json.loads(Path(a.source_entities).read_text())
|
||
_surf = sorted({(v.get("surface") or k)
|
||
for w in _e.values() for k, v in w["entities"].items()
|
||
if "\u2019" not in k and "'" not in k},
|
||
key=len, reverse=True)
|
||
if _surf:
|
||
source_pat = re.compile(r"\b(" + "|".join(re.escape(x) for x in _surf) + r")\b")
|
||
print(f" source-entity filter: {len(_surf)} surfaces the beat may not name")
|
||
|
||
rejects: dict[str, int] = {}
|
||
kept = 0
|
||
name_drift = 0
|
||
retried = 0
|
||
retry_saved = 0
|
||
drift_examples: list[dict] = []
|
||
t0 = time.time()
|
||
with out_path.open("w", encoding="utf-8") as fh:
|
||
for p in passages:
|
||
if kept >= a.n:
|
||
break
|
||
try:
|
||
r = post("/v1/chat/completions", {
|
||
"model": a.model, "temperature": a.temperature, "max_tokens": 80,
|
||
"messages": [{"role": "user", "content": BEAT_PROMPT.format(passage=p["response"], context=ctx_block(p["prev"]))}],
|
||
}, key)
|
||
raw = r["choices"][0]["message"]["content"] or ""
|
||
except (urllib.error.URLError, urllib.error.HTTPError, KeyError, TimeoutError) as e:
|
||
rejects["transport"] = rejects.get("transport", 0) + 1
|
||
print(f"[warn] {type(e).__name__}: {e}", flush=True)
|
||
continue
|
||
beat, why = vet(raw, p["response"], a.overlap_n, source_pat)
|
||
if beat is None and why == "meta" and NARRATOR_ONLY.search(raw):
|
||
retried += 1
|
||
try:
|
||
r2 = post("/v1/chat/completions", {
|
||
"model": a.model, "temperature": a.temperature, "max_tokens": 80,
|
||
"messages": [
|
||
{"role": "user", "content": BEAT_PROMPT.format(
|
||
passage=p["response"], context=ctx_block(p["prev"]))},
|
||
{"role": "assistant", "content": raw},
|
||
{"role": "user", "content": RETRY_NOTE.strip()},
|
||
]}, key)
|
||
beat, why = vet(r2["choices"][0]["message"]["content"] or "",
|
||
p["response"], a.overlap_n, source_pat)
|
||
if beat is not None:
|
||
retry_saved += 1
|
||
except (urllib.error.URLError, urllib.error.HTTPError, KeyError, TimeoutError):
|
||
pass
|
||
if beat is None:
|
||
rejects[why.split("(")[0]] = rejects.get(why.split("(")[0], 0) + 1
|
||
continue
|
||
drift = names_not_in_response(beat, p["response"])
|
||
if drift:
|
||
name_drift += 1
|
||
drift_examples.append({"beat": beat, "names": drift})
|
||
use_ctx = p["prev"] is not None and random.random() < a.context_frac
|
||
fh.write(json.dumps({
|
||
"work": p["work"], "chapter": p["chapter"], "split": p["split"],
|
||
"beat": beat, "context": p["prev"] if use_ctx else None,
|
||
"response": p["response"], "words": p["words"],
|
||
}, ensure_ascii=False) + "\n")
|
||
kept += 1
|
||
if kept % 50 == 0:
|
||
print(f"[gen] {kept}/{a.n} {time.time()-t0:.0f}s rejects={rejects}", flush=True)
|
||
|
||
fp_end = fingerprint()
|
||
prov = {
|
||
"generated_at": time.strftime("%Y-%m-%dT%H:%M:%S%z"),
|
||
"generator_model_alias": a.model,
|
||
"generator_fingerprint_start": fp_start,
|
||
"generator_fingerprint_end": fp_end,
|
||
"generator_repointed_midrun": fp_start != fp_end,
|
||
"gateway": GATEWAY, "temperature": a.temperature,
|
||
"corpus": str(a.corpus), "split": a.split, "copy": "copy0",
|
||
"band_words": [a.lo, a.hi], "seed": a.seed, "context_frac": a.context_frac,
|
||
"overlap_guard_n": a.overlap_n,
|
||
"passages_available": len(passages), "pairs_kept": kept, "rejects": rejects,
|
||
"narrator_retries": retried, "narrator_retries_recovered": retry_saved,
|
||
"beats_naming_absent_entity": name_drift,
|
||
"beats_naming_absent_entity_rate": round(name_drift / max(1, kept), 4),
|
||
"beats_naming_absent_entity_examples": drift_examples[:15],
|
||
"reject_rate": round(sum(rejects.values()) / max(1, kept + sum(rejects.values())), 4),
|
||
"elapsed_s": round(time.time() - t0, 1),
|
||
"system_prompt": sys_prompt,
|
||
"register": a.register or "yarros-frozen-one-paragraph",
|
||
"drop_leading_heading": a.drop_leading_heading,
|
||
"reflow_hard_wraps": a.reflow_hard_wraps,
|
||
}
|
||
Path(str(out_path) + ".provenance.json").write_text(json.dumps(prov, indent=2), encoding="utf-8")
|
||
print(f"[done] {kept} pairs -> {out_path}")
|
||
print(f"[done] rejects {rejects} (rate {prov['reject_rate']})")
|
||
print(f"[done] narrator-retries {retried}, recovered {retry_saved}")
|
||
print(f"[done] beats naming an entity absent from their response: {name_drift}/{kept} "
|
||
f"({prov['beats_naming_absent_entity_rate']}) -- reported, not rejected")
|
||
if prov["generator_repointed_midrun"]:
|
||
print("⚠ SEAT REPOINTED MID-RUN -- the pairs are not from one generator", file=sys.stderr)
|
||
return 0
|
||
|
||
|
||
if __name__ == "__main__":
|
||
raise SystemExit(main())
|