Two bugs R42 (brokkr-smithy-dev) surfaced on first live-index contact:
1. Session-reuse degradation. run_yardstick/run_term reused one mimir
session across terms; mimir returns EMPTY search_library results after
a session's first query (Worldtree #391), silently scoring every later
term a false-MISS. Fixed by making search_library and reference_knowledge
self-session (fresh session per call) so no caller can re-hoist it. Live
yardstick now reproduces all four anchors HIT top-10. Fresh-session-per-
query is the pinned arm-2 protocol; folded into the conventions docstring.
2. Curly-vs-ASCII apostrophe. _on_target substring-matched raw ASCII while
the b170 extraction stores U+2019, so possessive-named subjects
false-MISSed. _on_target now NFKC-normalizes + quote-folds both sides
(NFKC alone does not fold U+2019, so the explicit fold is load-bearing).
Adds tests/test_fiction_wing_probe.py covering the apostrophe fold both
directions with a negative control.
Self-contained, re-runnable probe requested by brokkr-smithy-dev for R42 (fiction-wing
retrieval characterization) and the standing #389 ranking acceptance gate. Two paths kept
separate by noise property: search_library (mimir, fixed-string, deterministic — ranking
arm) and reference_knowledge (donut, captures her reformulated tool_query — the query-
formulation/arm-4 surface). Scoring conventions baked in: high/medium/low RRF buckets
(0.030/0.016), on-target = a row whose excerpt names the subject, bucket-distribution over
N runs. Carries the frozen artifact yardstick (4 source-verified items + epithet-dropped
variants). Config from env (no secrets). Smoke-verified live: reproduces the Crown-HIT /
other-three-MISS baseline and the near-floor bucket flips.