Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-21-my-confound-detector-was-counting-apostrophes-as-quote.md
T

4 lines
1.0 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# `[2026-09-21]` My confound detector was counting apostrophes as quote marks, and I saw it FIRE before I saw the bug.
⚠⚠ **My confound detector was counting apostrophes as quote marks, and I saw it FIRE before I saw the bug.** `voice_distance.py`'s quote class shipped with `'` and `’` in it — on the one corpus whose signature is `dont`/`aint`/`wont`. It reported the held-out reference at 121.1 "quote marks" per 10k for a corpus whose builder **ASSERTS 0.0**, and fired the pre-registered trigger on a base arm whose true density is 19.9. Fixing a detector to measure the quantity the frozen rule names is not moving the rule, but the fix un-fires the trigger, which is indistinguishable from shopping — **so the trigger was made MOOT instead of adjudicated**: the normalised read is load-bearing unconditionally, both columns reported, zero verdict effect. ⚠ The lesson: I controlled `strip_punct` (2500→0) and the byte-identity of the default path, and never asked the quote counter for a value whose answer I already knew. Commit `0d80e49`.