Companion to lexical_recall_gate.py for the #397 order_by="chapter" flag (deployed
personal b184). Drives narrative/temporal queries and measures three axes end-to-end:
- ADOPTION: does the agent invoke order_by="chapter" for a temporal query? (schema
teaches it; usage varies — the #397 analog of query-formulation variance)
- MECHANISM (flag applied): are served hits' provenance.chapter monotonically
non-decreasing (earliest first)?
- VALUE (flag not applied): the relevance baseline is NOT chapter-sorted — the
applied-vs-not monotonicity gap is the flag's payoff.
Built against the real live shapes (order_by enum ["chapter"], result carries
ordered_by, provenance.chapter), not guessed. Baseline @ b184 (--runs=3, 15 trials):
adoption 47%, flag-applied->monotone 100%, not-applied->monotone 0%. So the mechanism
is a clean discriminator; the residual is adoption (same class as the crown's
query-formulation variance — the irreducible prompt-side gap).
Diagnostics fixture, no production runtime — no version bump. persistent-memory
snapshot alongside (commit-along).