audit_entity_map: the rename can damage the prose and no gate will ever say so

audit_stoplist.py finds surfaces wrongly held OUT of the entity map -- a stoplisted
character is an undetectable leak. This is the mirror: surfaces wrongly held IN it.
leak_gate.py only ever asks whether the author's names are GONE, never whether
non-names were spared, so renaming `the Chinese` into an invented surname passes it
perfectly.

Found sideways on Hemingway. The pairs audit reported beats naming African, Chinese,
X-ray, Republican and Cezanne as leaks -- correctly, those surfaces really were removed
from the corpus. Reading why turned up the larger defect: they should never have been
renameable in the first place.

Measured on the Hemingway map, both controls green:
  positive  `other` 764/1356 article-preceded = 0.56
  negative  100 honorific-confirmed people, highest Inglés at 0.26, bulk 0.00-0.06
  FLAGGED   130 of 946 surfaces, 1,616 instances = 0.162% of corpus words

The signal is an article in front of the surface: you write `the Frenchman` and `a
Martini`, never `the Rinaldi`. It is a heuristic and every hit is reported FOR READING,
never auto-removed -- `the Widow` and `the Informer` are genuine Hemingway epithet-names
that SHOULD be renamed, and the band's own top entry makes the point, since Inglés at
0.26 is an in-world nickname deliberately kept renameable and sits just under the bar.

Initials are excluded from the negative-control band rather than admitted to it. `Mr. P.`
is an initial, not a person, so letting it in lets a map defect poison the control that
validates the detector -- on Hemingway `P` (0.32, every occurrence `the P. O. U. M.`) was
the one surface failing a band whose next highest was 0.26. Initials take no article and
are invisible to the scan anyway, so every surface of two characters or fewer is now
listed unconditionally. Sixteen of them are in this map, C at 274 occurrences; the same
class as the `G` that was caught by hand about to be renamed to a surname 248 times.

The unresolved count that drives the exit code is computed over every flagged surface,
not the --show slice. Tying a gate's verdict to a display flag is the same defect as a
log filter that turns a real event into a clean zero.

Also corrects a wrong claim in audit_pairs_sourcenames.py's docstring: the Hemingway
rename did not HOLD 591 surfaces. Paris, Madrid and Spain survive because the stoplist
keeps them out of the entity map before it is built, so the map is exactly the removed
set -- 941 surfaces, 941 removed, 0 kept. Measured per run rather than assumed, because
a pipeline that carried kept surfaces into the map would report every `Paris` as a leak.
This commit is contained in:
vh
2026-09-17 02:01:01 -07:00
parent 0bb4938518
commit 051b99e063
2 changed files with 213 additions and 7 deletions
@@ -11,13 +11,15 @@ for pair sets already built — Hemingway's and Yarros's both predate the flag,
a clean corpus gate is not evidence about them either way.
WHAT COUNTS AS A LEAK, and why the distinction matters. A source surface the beat
names is only a leak if the rename actually took it away. Hemingway's map holds
941 surfaces and the rename moved 1,097 instances while HOLDING 591 — real places
(`Paris`, `Madrid`), allow-listed real-world terms, and everything under the
`--min-cap` threshold. A beat naming `Paris` names something the renamed corpus
says constantly; a beat naming a removed character restores what the pipeline
exists to delete. So every matched surface is classified against the renamed
copies first, and only the removed ones are counted against the gate.
names is only a leak if the rename actually took it away — so every matched surface
is classified against the renamed copies first, and only the removed ones count
against the gate. Real-world names the pipeline deliberately keeps (`Paris` 173
occurrences, `Madrid` 108, `Spain` 87, all still present in the renamed copies) are
held back by the stoplist BEFORE the entity map is built, so on Hemingway the map
turns out to be exactly the removed set: 941 surfaces, 941 removed, 0 kept. Do not
assume that holds on another corpus — the classification is measured per run, and a
pipeline that instead carries kept surfaces INTO the map would report every mention
of `Paris` as a leak if this step were skipped.
THE RESPONSE SIDE IS THE DIAGNOSTIC. Beats and responses are scanned separately.
Leaks in the beats with a clean response column is the lv-bronte signature: the