Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-17-romantasy-register-measured.md
T
vh d94b5a1934 memory: snapshot — lv-mccarthy training launched on gx10, and the next voice seat is measured rather than chosen
In-flight rewritten to the live training run (~150/1380, ETA ~00:45 PT) with the
--save-total-limit finding that would otherwise have deleted the epoch-1/epoch-2
checkpoints both prior gates were decided on.

Two decisions added: the next-seat ranking (Faulkner, Morrison, Chandler -- and the
finding that the corpus size ranking inverts the voice ranking, with King and Christie
as the two biggest non-candidates), and the romantasy register measured on the gate's
own char-bigram instrument (Yarros is the cluster outlier we already shipped; Maas is
the centroid and so the worst pick; Kenyon at 27 val units if the lane gets a seat).

Auto-archival: 4 entries moved to archival-memory.md; 4 held back by the open-deferred
guard.
2026-09-17 22:38:13 -07:00

4.0 KiB

[2026-09-17] Romantasy measured as a register — it is real, we already took its best voice, and the obvious next pick is its worst

Prompted by the operator pushing back on a one-clause dismissal of the lane as "depth behind Yarros". The dismissal was taste; this is a measurement, on the gate's own instrument.

Method. Char-bigram Burrows's Delta, the same measure voice_distance.py gates on. ~120k words per author, sampled from the MIDDLE quartile of each author's largest works (front and back matter are not the voice), equalised so a bigger sample is not a different measurement. 400 most-frequent bigrams as the feature set, z-scored over 4,000-word chunks pooled across all authors.

Controls first, because a between-author number without a within-author floor is unfalsifiable.

A-vs-A floor (two halves of the SAME author)
  Yarros 0.285  Maas 0.298  Armentrout 0.314  St. Clair 0.322  Cole 0.338
  Kenyon 0.375  Reyne 0.391
  McCarthy 0.303  Morrison 0.327  Brontë 0.209   Hemingway 0.454  <- worst, used as the bar

positive controls (known-distinct pairs — the instrument must separate these)
  Yarros    vs McCarthy   0.862   1.9x
  Hemingway vs Brontë     0.773   1.7x
  McCarthy  vs Morrison   0.675   1.5x
  Hemingway vs McCarthy   0.655   1.4x

romantasy, all 21 pairs   median 0.537   1.2x floor   (range 0.465 - 0.674)

The register is real but tight. 1.2x floor against controls at 1.4-1.9x. Only one pair falls to 1.0x, so it is not seven names for one voice.

⚠ Sensitivity floor, stated because a result without one is unfalsifiable. The 0.454 bar is Hemingway's, inflated by his own heterogeneous corpus (1920s-1960s, novels + stories + posthumous). Against the romantasy authors' OWN floors (~0.34) the same pairs read ~1.6x — control-grade. The truth sits between those readings and this method cannot split it finer. One sample per pair, no repeat draws: read the rank ordering as indicative, do not read small gaps at all.

Two findings that survive either floor reading

⭐ Yarros is the cluster OUTLIER, not a typical member. Four of the five largest distances in the matrix involve her — Yarros-Kenyon 0.674, Yarros-St. Clair 0.644, Yarros-Reyne 0.637, Yarros-Maas 0.567. We already trained the most distinctive romantasy voice we hold, so a second seat in the lane buys measurably less than the first did. That is the actual answer to "what about romantasy".

⭐ Maas is the centroid, so the obvious commercial pick is the least distinctive. Maas-Reyne 0.465 and Maas-Cole 0.470 are the two SMALLEST distances in the whole matrix. She is the biggest name available (922k words) and measurably the most generic of the seven in char-bigram terms. Picking by sales rank picks the worst adapter.

If the lane gets a second seat it is Kenyon

Furthest from the shipped Yarros (0.674), so it adds the most new signal — and 27 works means 27 val units, the best-powered gate the line could build (Hemingway 10, McCarthy 6, Brontë 4, where 4 is the documented structural cause of an underpowered voice axis with no cheap fix).

⚠ Two costs: the 27 are one series (Dark-Hunter), so the shared proper-noun space makes --scope corpus mandatory rather than optional; and a "Dark Hunter - The Dark Hunter Complete" omnibus sits in the catalogue rows, so the containment pass runs first.

Corpus shapes for the lane (works ≥100k chars, deduped by title):

Sherrilyn Kenyon   27   2,368,396 w      Scarlett St. Clair  11   1,133,067
Opal Reyne         14   2,466,307        Kresley Cole        10   1,037,914
Jennifer Armentrout 6   1,179,953        Sarah J. Maas        5     922,711
Rebecca Yarros      5     820,425   <- SHIPPED on this

⭐ Worth noting for any future bar-setting: Yarros shipped on 5 works / 820k words. The corpus bar is lower than it looks.

Instrument: scratchpad/regdist.py (screening tool, not the gate).

Related: 2026-09-17-next-voice-seats, 2026-09-16-lv-voices-line, 2026-09-17-lv-bronte-gate.