split_units: choose a unit mode by SIZE, not by count, and fall back to paragraph blocks

McCarthy and Krakauer both need this before a corpus can be built, so it is a shared module
rather than a third copy of the Hemingway splitter.

THE INHERITED RULE IS "MOST UNITS ABOVE A FLOOR" AND IT BREAKS ON PART MARKERS. Measured:

  Cities of the Plain   4 roman marks -> 4 units, median 22,312w   <- the book's PARTS
  The Crossing          4 roman marks -> 4 units, median 37,310w   <- same

"Most units" scores 4 over the 1 that finding-nothing gives, so it wins, and the existing
guard only fires at exactly one unit. A 37,000-word "chapter" sails through and every
downstream tool accepts it. Size is now the eligibility test: a mode qualifies only if its
median unit is inside [600, 12000] AND no single unit holds half the work.

TWO THINGS A CONTROL RUN CAUGHT, BOTH NOW FIXED IN THE RULE. The first version scored
eligible modes by "median closest to target". Run over Hemingway, whose markers are known
good, it chose caps-title over the book's own chapters on True at First Light:

  bare-numeral  20 units  median 5,337w  max  11,155   <- the real chapters
  caps-title     6 units  median   777w  max 113,886   <- median looked BETTER

caps-title matched five stray all-caps lines, so five tiny units sat beside one holding 97%
of the book. A median cannot see that distribution; a max bound can. And caps-title is the
weakest of the four signals, which is why the tiebreak among eligible modes is now PRIORITY
(contents > chapter-word > roman > bare-numeral > caps-title), not size.

CONTROLS, both green after the fix:
  positive  Hemingway's ten works, markers known good -> 8/10 reproduce the shipped mode and
            unit count exactly. The two differences are explained, neither is a mode error:
            short-stories used `contents`, which the harness does not supply, and The Old Man
            and the Sea was deliberately kept whole as CONTINUOUS.
  negative  40,000 words with no blank lines -> 1 unit. It refuses to fabricate divisions
            out of unstructured text rather than returning a plausible section count.

Result on the two new authors: McCarthy 167 units / 587,233 words, Krakauer 135 / 431,938,
both median ~3,200-3,500w against Hemingway's 3,128.

⚠ CORRECTION TO AN EARLIER SURVEY. I reported that all four Krakauer works carry zero
chapter markers. That was wrong and it was my regex, not the books: the survey pattern
required "Chapter" followed by a numeral, and Krakauer writes "CHAPTER ONE". Missoula and
Where Men Win Glory split on chapter-word (33 and 41 units); Into the Wild and Under the
Banner of Heaven on caps-title (28 and 33). Only McCarthy's All the Pretty Horses, Cities of
the Plain, The Crossing and The Road actually need the fallback.

The Hemingway builder is deliberately NOT repointed at this module. Its corpus is shipped and
its provenance sha is pinned by a live adapter; the one behavioural difference (The Old Man
and the Sea would section into 9 rather than stay whole) is an improvement nobody asked for
on a corpus nobody should churn.
This commit is contained in:
2026-09-17 08:12:09 -07:00
parent 9f35c8d659
commit 705fa3a65b
+165
View File
@@ -0,0 +1,165 @@
"""Cut a work into units, and REFUSE to do it badly quietly.
A "unit" is a chapter-shaped record. Three things downstream depend on it and none of
them will complain if it is wrong: passages are chunked inside a unit, `chapter` is the
field `rename.py` holds a train/val split out on, and the manifest's unit count is the
only number anyone looks at to decide the split worked.
⚠⚠ THE INHERITED SELECTION RULE IS "MOST UNITS ABOVE A FLOOR", AND IT IS WRONG FOR ANY
BOOK WHOSE MARKERS ARE PARTS RATHER THAN CHAPTERS. Measured on McCarthy before this module
existed:
Cities of the Plain 4 roman-numeral marks -> 4 units of ~22,500 words each
The Crossing 8 roman-numeral marks -> 8 units of ~18,750 words each
Those marks are the book's four PARTS, not its chapters. "Most units" scores 4 over the
1 that finding-nothing gives, so it wins, and the existing guard only fires at exactly one
unit — so a 22,500-word "chapter" sails through. The Hemingway builder's own warning says a
splitter that silently finds one chapter in a novel is the failure to fear; this is the same
failure wearing a plausible number.
⭐ SO SIZE IS THE ELIGIBILITY TEST AND PRIORITY IS THE TIEBREAK. A mode qualifies only if
its median unit lands inside a sane band AND no single unit swallows half the work; among
qualifying modes the most trustworthy marker form wins, in the inherited order
contents > chapter-word > roman > bare-numeral > caps-title. Every candidate is reported
with its stats whether it won or lost, because the interesting question when a corpus looks
wrong later is which modes were considered and why they lost.
⚠ THE MAX BOUND AND THE PRIORITY ORDER ARE BOTH THERE BECAUSE A CONTROL RUN CAUGHT THEM.
Scoring eligible modes by "median closest to target" was the first rule written here, and on
Hemingway's `True at First Light` it chose `caps-title` (6 units) over the book's own
`bare-numeral` chapters (20 units), changing a shipped corpus for the worse:
bare-numeral 20 units median 5,337w max 11,155
caps-title 6 units median 777w max 113,886 <- median looked BETTER
caps-title matched five stray all-caps lines, so five tiny units sat beside one unit holding
**97% of the book**, and the median was computed over a distribution that pathological. A
median alone cannot see that; a max bound can. `caps-title` is also the weakest signal of the
four — it matches any all-caps line — which is why priority, not size, breaks the tie.
⭐ AND THERE IS A FALLBACK, because some books genuinely have no divisions. Krakauer's four
works carry ZERO chapter markers of any kind between them, and so do McCarthy's The Road and
All the Pretty Horses. Without a fallback each of those becomes one 59k129k-word record and
the train/val holdout has nothing to hold out. `paragraph-blocks` accumulates blank-line
separated blocks up to `target` and cuts on a block boundary — synthetic sections, labelled
as such so nobody mistakes them for the author's own divisions.
⚠ THE FALLBACK IS LAST, NOT BEST. It can always hit `target` exactly, so scoring it against
the marker modes on size would let it win everywhere. Real divisions are semantically
coherent and synthetic ones are not, so a sane marker mode beats the fallback by rule rather
than by score.
"""
from __future__ import annotations
import re, statistics as st
# The heading forms, in the order the Hemingway builder tried them. Kept byte-identical so
# a work that split one way there splits the same way here.
HEAD_PATTERNS = [
("chapter-word", re.compile(
r"^[ \t]*((?:Chapter|CHAPTER)[ \t]+(?:[A-Za-z-]+|\d+)|Prologue|PROLOGUE|Epilogue|EPILOGUE)"
r"[ \t]*\.?[ \t]*$", re.M)),
("roman-numeral", re.compile(r"^[ \t]*((?=[IVXL])[IVXL]{1,7})[ \t]*\.?[ \t]*$", re.M)),
("bare-numeral", re.compile(r"^[ \t]*(\d{1,3})[ \t]*\.?[ \t]*$", re.M)),
("caps-title", re.compile(r"^[ \t]*([A-Z][A-Z '\-,!\.]{4,60})[ \t]*$", re.M)),
]
PRIORITY = ["contents", "chapter-word", "roman-numeral", "bare-numeral", "caps-title"]
MIN_UNIT_WORDS = 150 # below this a "unit" is a fragment, not a chapter
TARGET_WORDS = 3000 # Hemingway's built corpus averages 3,128 words per unit
SANE_LO = 600 # a 400-word "chapter" means the marker is a scene break
SANE_HI = 12000 # a 22,500-word "chapter" means the marker is a PART
def _units_from_marks(text: str, pat: "re.Pattern[str]") -> list[tuple[str, str]]:
marks = [(m.start(), m.group(1).strip()) for m in pat.finditer(text)]
units = []
for i, (pos, head) in enumerate(marks):
end = marks[i + 1][0] if i + 1 < len(marks) else len(text)
body = text[pos:end].strip()
if len(body.split()) >= MIN_UNIT_WORDS:
units.append((head, body))
return units
def paragraph_blocks(text: str, target: int = TARGET_WORDS) -> list[tuple[str, str]]:
"""Synthetic sections on blank-line boundaries, ~`target` words each.
Cuts only between blocks, never inside one, so no sentence is ever split. A trailing
remainder shorter than half the target is merged back into the previous section rather
than emitted as a runt — a 300-word final "chapter" would be held out as a val unit at
the same weight as a 3,000-word one.
"""
blocks = [b.strip() for b in re.split(r"\n\s*\n", text) if b.strip()]
units, cur, n = [], [], 0
for b in blocks:
cur.append(b)
n += len(b.split())
if n >= target:
units.append(cur)
cur, n = [], 0
if cur:
if units and n < target // 2:
units[-1].extend(cur)
else:
units.append(cur)
return [(f"§{i}", "\n\n".join(u)) for i, u in enumerate(units, 1)]
def _stats(units: list[tuple[str, str]]) -> dict:
w = [len(b.split()) for _, b in units]
if not w:
return {"n": 0, "median": 0, "min": 0, "max": 0, "words": 0}
return {"n": len(w), "median": int(st.median(w)), "min": min(w), "max": max(w),
"words": sum(w)}
def choose_units(text: str, *, contents_units: list[tuple[str, str]] | None = None,
target: int = TARGET_WORDS, sane_lo: int = SANE_LO,
sane_hi: int = SANE_HI) -> tuple[str, list[tuple[str, str]], list[dict]]:
"""Return (mode, units, report). `report` holds every candidate, won or lost."""
candidates: list[tuple[str, list[tuple[str, str]]]] = []
if contents_units:
candidates.append(("contents", contents_units))
for name, pat in HEAD_PATTERNS:
u = _units_from_marks(text, pat)
if u:
candidates.append((name, u))
total = len(text.split()) or 1
report = []
eligible = []
for name, u in candidates:
st_ = _stats(u)
hogs = st_["max"] > 0.5 * total
ok = (len(u) >= 2 and sane_lo <= st_["median"] <= sane_hi and not hogs)
why = ("eligible" if ok else
"only one unit" if len(u) < 2 else
f"one unit holds {st_['max']/total:.0%} of the work — the marks do not partition it"
if hogs else
f"median {st_['median']:,}w below {sane_lo:,} — marks are scene breaks"
if st_["median"] < sane_lo else
f"median {st_['median']:,}w above {sane_hi:,} — marks are PARTS, not chapters")
report.append({"mode": name, **st_, "eligible": ok, "verdict": why})
if ok:
eligible.append((PRIORITY.index(name), name, u))
fb = paragraph_blocks(text, target)
report.append({"mode": "paragraph-blocks", **_stats(fb), "eligible": True,
"verdict": "fallback — only used when no marker mode is eligible"})
if eligible:
eligible.sort(key=lambda x: x[0]) # by PRIORITY, not by size -- see the docstring
_, name, u = eligible[0]
return name, u, report
return "paragraph-blocks", fb, report
def format_report(report: list[dict], chosen: str) -> str:
lines = []
for r in report:
mark = " <- CHOSEN" if r["mode"] == chosen else ""
lines.append(f" {r['mode']:<18}{r['n']:>5} units median {r['median']:>7,}w "
f"min {r['min']:>6,} max {r['max']:>7,} {r['verdict']}{mark}")
return "\n".join(lines)