From 17ae1558f9d984f04fff8c50244aad3aa03d1a83 Mon Sep 17 00:00:00 2001 From: Vuong Hoang Date: Fri, 7 Aug 2026 19:07:05 -0700 Subject: [PATCH] =?UTF-8?q?docs(diagnostics):=20add=20lexical-recall=20gat?= =?UTF-8?q?e=20=E2=80=94=20class=20acceptance=20instrument=20for=20exact-t?= =?UTF-8?q?erm=20recall?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Generalizes the crown repro (Worldtree #400 / thread 01KZETD98T) beyond its anchor into a before/after regression instrument for the class property: when the corpus holds a chunk whose text literally carries a queried surface form, a natural query should serve >=1 such chunk at a usable rank. - Anchors tagged stress (common word + competing dense cluster, e.g. crown) vs control (distinctive name — should sit ~0% miss). - Binary per trial: does a natural query serve >=1 term-containing chunk within top-10 (USABLE_K)? Ranks >=8 flagged KNIFE-EDGE (the RRF fused-rank 9-11 window residual worldtree-dev's decomposition measured). - Real-world end-to-end: drives the agent (it composes its own reference_knowledge query, as in production); --runs samples query-formulation variance to estimate a true miss-rate. - Extensible anchor list; --anchor filters. This is the deciding instrument for the rerank_hybrid_floor lever: its stress-class miss-rate (alongside brokkr's fleet demotion rate) rules the floor in or out after the BM25 stemming fold deploys. Pre-fold baseline captured today (the "before"): control 0% miss / stress[crown] 100% miss / 0% knife-edge, 11 trials. Diagnostics fixture, no production runtime — no version bump. persistent-memory snapshot committed alongside (commit-along). --- docs/diagnostics/lexical_recall_gate.py | 145 ++++++++++++++++++++++++ persistent-memory.md | 3 + 2 files changed, 148 insertions(+) create mode 100644 docs/diagnostics/lexical_recall_gate.py diff --git a/docs/diagnostics/lexical_recall_gate.py b/docs/diagnostics/lexical_recall_gate.py new file mode 100644 index 0000000..9d5e39f --- /dev/null +++ b/docs/diagnostics/lexical_recall_gate.py @@ -0,0 +1,145 @@ +"""Lexical-recall gate — the class acceptance instrument for exact-term recall survival. + +Generalizes the crown repro (Worldtree #400 / thread 01KZETD98T) beyond its anchor. The +class property under test: when the corpus contains a chunk whose text literally carries a +queried surface form, a natural query for that entity should serve >= 1 such chunk at a +USABLE rank (inside the top-K window). The crown ("Crown of the Sepsis Whore") is the +motivating STRESS case — a common word with a dense-similar vanity cluster that buries the +exact-lexical match; distinctive names (Krakaren, Vine Creeper) are CONTROLs that should +always pass. The gap lives on the stress class, not the controls. + +This is a before/after regression instrument, NOT a fix: run it pre-deploy and post-deploy +(the BM25 stemming fold, then any rerank_hybrid_floor lever) to measure whether the served +miss-rate on the stress class actually moves. Real-world by design — it drives the agent +end-to-end (the agent composes its own reference_knowledge query, as in production), and +--runs samples that query-formulation variance to estimate a true miss-rate. + +Self-contained (httpx only). Config from env (source env.sh first): + WORLDTREE_API_URL (default personal :8081), WORLDTREE_API_KEY (required), + RATATOSKR_END_USER_ID (default ratatoskr-tui), RATATOSKR_TTS_AGENT unused here. + + uv run --with httpx python docs/diagnostics/lexical_recall_gate.py + uv run --with httpx python docs/diagnostics/lexical_recall_gate.py --runs 5 + uv run --with httpx python docs/diagnostics/lexical_recall_gate.py --anchor crown +""" +from __future__ import annotations + +import argparse +import json +import os +import re + +import httpx + +# Served window: a hit past this rank is not "usable" (Worldtree serves ~top-10; a row at +# rank 9-11 is the RRF knife-edge worldtree-dev identified — treated as a KNIFE-EDGE pass). +USABLE_K = 10 +KNIFE_EDGE_FROM = 8 # ranks >= this inside the window are fragile (one-rank-edge residual class) + +AGENT = "ratatoskr:donut" + +# (label, kind, term-regex the served chunk's excerpt must contain, [natural user messages]). +# kind: "stress" = common word + competing dense cluster; "control" = distinctive name. +# Controls should pass every run; the class limitation shows as stress-class misses / knife-edges. +ANCHORS = [ + ("crown", "stress", r"\bcrown", + ["What crown do you own?", "Do you have a crown?", "Tell me about your crown."]), + ("vine-creeper", "control", r"vine creeper", + ["Tell me about the Vine Creeper.", "What is the Vine Creeper?"]), + ("danger-dingo", "control", r"danger dingo|\bdingo", + ["What is the Danger Dingo?", "Describe the Danger Dingo."]), + ("pedicure-kit", "control", r"pedicure", + ["What does the Pedicure Kit do?", "Tell me about the Pedicure Kit."]), + ("neighborhood-map", "control", r"neighborhood map", + ["What is the Neighborhood Map?", "Describe the Neighborhood Map."]), +] + + +def _cfg() -> tuple[str, dict, str]: + base = os.environ.get("WORLDTREE_API_URL", "http://10.250.50.152:8081") + key = os.environ.get("WORLDTREE_API_KEY") + if not key: + raise SystemExit("WORLDTREE_API_KEY unset — source env.sh first.") + return base, {"Authorization": f"Bearer {key}"}, os.environ.get("RATATOSKR_END_USER_ID", "ratatoskr-tui") + + +def _session(base: str, headers: dict, end_user: str) -> str: + r = httpx.post(f"{base}/sessions", json={"agent_id": AGENT, "end_user_id": end_user}, + headers=headers, timeout=30) + r.raise_for_status() + return r.json()["session_id"] + + +def _drive(base: str, headers: dict, sid: str, content: str) -> tuple[str | None, list]: + """POST a turn; return (actual reference_knowledge query, served hits list).""" + q, hits = None, [] + with httpx.stream("POST", f"{base}/sessions/{sid}/messages", json={"content": content}, + headers=headers, timeout=180) as r: + for line in r.iter_lines(): + if not line.startswith("data: "): + continue + ev = json.loads(line[6:]) + t = ev.get("type") + if t == "tool_start" and q is None: + q = (ev.get("arguments") or {}).get("query") + elif t == "tool_result" and not hits: + res = ev.get("result") + if isinstance(res, dict): + hits = res.get("hits", res.get("results", [])) or [] + elif t == "done": + break + return q, hits + + +def _served_rank(hits: list, term_re: str) -> int | None: + """Rank of the first served hit whose excerpt literally contains the term (None = miss).""" + for i, h in enumerate(hits[:USABLE_K]): + if isinstance(h, dict) and re.search(term_re, h.get("excerpt", ""), re.I): + return i + return None + + +def main() -> None: + ap = argparse.ArgumentParser() + ap.add_argument("--runs", type=int, default=1, help="repeats per message (samples query variance)") + ap.add_argument("--anchor", default=None, help="run only this anchor label") + args = ap.parse_args() + base, headers, end_user = _cfg() + + anchors = [a for a in ANCHORS if args.anchor is None or a[0] == args.anchor] + totals = {"trials": 0, "miss": 0, "knife": 0} + by_kind: dict[str, dict] = {} + + for label, kind, term_re, messages in anchors: + print(f"\n[{label}] ({kind}) term=/{term_re}/") + agg = by_kind.setdefault(kind, {"trials": 0, "miss": 0, "knife": 0}) + for msg in messages: + for _ in range(args.runs): + sid = _session(base, headers, end_user) # fresh session per trial + q, hits = _drive(base, headers, sid, msg) + rank = _served_rank(hits, term_re) + miss = rank is None + knife = rank is not None and rank >= KNIFE_EDGE_FROM + for d in (totals, agg): + d["trials"] += 1 + d["miss"] += int(miss) + d["knife"] += int(knife) + tag = "MISS" if miss else (f"knife@{rank}" if knife else f"ok@{rank}") + print(f" {tag:9} msg={msg!r:42} q={q!r}") + + def pct(n: int, d: int) -> str: + return f"{(100*n/d):.0f}%" if d else "n/a" + + print("\n=== SUMMARY ===") + for kind, d in sorted(by_kind.items()): + print(f" {kind:8} trials={d['trials']:3} miss={pct(d['miss'], d['trials'])} " + f"knife-edge={pct(d['knife'], d['trials'])}") + t = totals + print(f" {'ALL':8} trials={t['trials']:3} miss={pct(t['miss'], t['trials'])} " + f"knife-edge={pct(t['knife'], t['trials'])}") + print("\nGate: stress-class miss-rate is the deciding signal for the rerank_hybrid_floor lever.") + print("Controls should sit at ~0% miss; a stress miss/knife-edge is the residual class to weigh.") + + +if __name__ == "__main__": + main() diff --git a/persistent-memory.md b/persistent-memory.md index 73e82da..101c75f 100644 --- a/persistent-memory.md +++ b/persistent-memory.md @@ -46,6 +46,8 @@ upstream API key stays server-side (INV-003). _As of 2026-08-07:_ +**✅ CROWN-RECALL investigation → a RETRIEVAL-CLASS limitation; gate shipped; awaiting deploy pre/post (this session).** Operator: "shouldn't Donut know about her crown?" Traced (with worldtree-dev, thread `01KZETD98T` / WT **#400**) to a GENERAL retrieval property — NOT persona, extraction, or synonymy: an exact-lexical match on an INDEXED surface form ("Crown of the Sepsis Whore", chunks `ch11_s022_c04`/`ch11_s023_c01` — literally contain "crown") loses to dense-semantic neighbors (the Donut-vanity cluster) on ownership-phrased queries. Three mechanisms (worldtree-dev decomposition vs a pulled index): (1) BM25 tokenizer has **no stemming** → "crowns"≠"crown" (plural miss); (2) RRF arithmetic parks single-arm-strong rows on the top-10 **knife edge** (fused rank 9-11); (3) the **reranker** prefers status chunks for ownership phrasings. **Operator-greenlit evidence-gated sequence:** ship the stemming fold (fixes the plural CLASS — crowns/swords/potions, runtime, re-runs R42 floor bracketed) + measure the residual + **HOLD the `rerank_hybrid_floor` lever** until the gate reports a real-world miss-rate. **Consumer deliverable: `docs/diagnostics/lexical_recall_gate.py`** — class acceptance instrument (stress vs control anchors; binary = a natural query serves ≥1 term-containing chunk in top-10; knife-edge≥8; real-world end-to-end + `--runs` samples query-formulation variance). **Pre-fold baseline (the "before"): control 0% miss / stress[crown] 100% miss / 0% knife-edge, 11 trials.** PENDING: worldtree-dev pings on the fold+#397 deploy → run pre/post at higher `--runs`, drop the stress-class miss-rate delta on #400 = the number that rules the floor lever in/out (decided alongside brokkr's fleet demotion rate — two independent instruments). Sibling items from this thread: **WT #399** (expose rerank-relevance in served results — FILED), **WT #397** (`order_by=chapter`, committed `c79ceaa` rev 1.3, NOT deployed — point the temporal fixture at it on land). Explicitly NOT a persona fix (operator rejected fact-stashing) and NOT a synonym band-aid (operator rejected overfitting the symptom). + **✅ TTS MIGRATED off Zonos → chatterbox-fast (this session; COMMITTED, not pushed).** `tts.py` repointed from the Zonos gateway (:8890 `/v1/audio/speech`) to **chatterbox-fast** (`http://10.100.79.3:8197/tts` — bespoke non-OpenAI `{text,voice,format,stream}` schema, no auth, 24kHz, infra-ops-verified against image `local/chatterbox-fast:v1`). Three subsystems DELETED: (1) **affect** — Turbo has no emotion knob, so `PadState`/`EmotionDials`/`pad_to_dials` + `/api/tts` `p`/`a` fields + browser `pad` arg are gone (DEC-7 retired; operator-directed "drop it for chatterbox"); (2) **client-side chunking** — no per-synth cap (gateway chunks internally), so `chunk_text`/`tts_stream_long`/`_pcm_after_header` gone, one `tts_stream` call voices a whole turn (DEC-10 retired, the mid-stream `yielded_any` degrade folded into `tts_stream`); (3) **language pin** — English-only, no `language` field (DEC-9 re-purposed, below). **Browser SR 44100→24000** (load-bearing correctness fix). Default voice `Cora`→`glados_25s`; `donut` registered lowercase at `/refs/donut.wav`. 518 suite green, live-smoked (real 24kHz synth + endpoint proxy + bounced `ratatoskr-web`). Contract `donut_voiced_interview.contract.md` amended (migration banner; DEC-1/3/8 amended; DEC-7/9/10 retired w/ historical notes). `tts.py` is the single swap seam; `RATATOSKR_TTS_URL` overrides (no env pin, uses the code default). **✅ ENGLISH-DRIFT — RESOLVED SERVER-SIDE (this session; fix deployed + operator ear-confirmed).** Operator: Donut "swaps to German halfway through." TWO wrong hypotheses before the real cause (infra-ops, thread `01KZEDMJ…`): (W1) "English-only, nothing to drift" — falsified; (W2) "multilingual leak → tighten sampling" — I shipped `top_k 1000→80`/`top_p 0.95→0.85`/`temp 0.8→0.5` and it made it WORSE (tight sampling pulls the garble onset to ~200 chars vs ~300 at default). **REAL CAUSE (signal-measured): the Turbo model OVER-RUNS its generation TAIL** — garble/dead-air in the final ~2-3s of a long single generation (voiced-tail ZCR 1.58x the middle). The scheduler's unbounded ratchet built 300-600 char mega-chunks landing in that zone; the "German" was tail garble mis-heard (+ shared-3090 OOM garbage, now gone — Zonos moved off the 3090). **FIX = infra-ops server-side `max_chunk_chars=250` (image :v2), operator ear-confirmed clean audio + clean joins (ZCR 1.58x→0.64x).** CONSUMER SIDE (shipped): **REVERTED the W2 sampling knobs** — `gateway_body` back to `{text,voice,format,stream}`, send FULL text + gateway DEFAULT sampling (the curbs fought the :v2 cap); **KEPT** the `/api/tts` empty-200→503 guard as hygiene (DEC-9a; OOM itself resolved). Per-request `max_chunk_chars` override available if per-call tuning ever wanted. **Loop CLOSED — nothing pending.** @@ -197,6 +199,7 @@ relational-dynamics verify (bind `--bifrost-url :8392`); WT #356 resume-durabili Chronological log of decisions with `[YYYY-MM-DD]` prefix. One line per decision. Captures rationale that won't be obvious from code alone. +- `[2026-08-07]` **Crown-recall traced to a general RETRIEVAL-CLASS limitation (exact-lexical match on an indexed surface form buried by dense-semantic neighbors), NOT persona/extraction/synonymy.** worldtree-dev decomposition (WT #400): no-stemming tokenizer + RRF top-10 knife-edge + reranker ownership-demotion. Operator-greenlit evidence-gated sequence (ship stemming fold → gate measures residual → hold `rerank_hybrid_floor` lever). Shipped `docs/diagnostics/lexical_recall_gate.py` (class acceptance instrument); pre-fold baseline control 0% / stress[crown] 100% miss. Awaiting worldtree-dev fold+#397 deploy ping for the pre/post delta. Thread `01KZETD98T`; siblings WT #399 (filed), #397 (committed, not deployed). - `[2026-08-07]` **TTS migrated Zonos→chatterbox-fast (`:8197` bespoke schema); affect DROPPED (Turbo has no emotion knob, operator "drop it for chatterbox"), client-chunking DROPPED (no per-synth cap), language pin DROPPED, browser SR 44100→24000.** Pushed `19b499a`. English "German drift" real cause (after 2 wrong hypotheses) = Turbo model OVER-RUNS its generation TAIL on long single generations (garble in final ~2-3s, ZCR 1.58x); **fixed SERVER-SIDE by infra-ops (`max_chunk_chars=250`, image :v2, operator ear-confirmed).** Consumer: **REVERTED my interim `top_k/top_p/temp` curbs** (they made it WORSE — pulled garble onset earlier), send full text + default sampling; KEPT the `/api/tts` empty-200→503 guard as hygiene. Contract `donut_voiced_interview.contract.md` amended. Loop closed. - `[2026-08-07]` **order_by=chapter tool flag → FILED as Worldtree #397 (DEFERRED to next session's contract pass).** Narrative/temporal-query gap ("first encounter in the dungeon"): `reference_knowledge` sorts by relevance not chronology; `provenance.chapter` is on every chunk but the consumer can't reorder native results (kb_bridge retired). Operator ruled the upstream sort flag the clean fix; worldtree-dev accepted, our fixture is the measurement instrument. Tracked at **Worldtree #397** (+ althing thread `01KZED2T3XHJ2WMS5NCYK42W6R`). - `[2026-08-07]` **#393 (descriptive-query subject binding) CLOSED — persona-expand lever the win (4/10→9/10), tool directive the fleet floor.** `docs/diagnostics/descriptive_query_binding.py` is the canonical #393 fixture; two-mechanism split (cross-wing dilution vs fiction-scope selection). Commits `6c83a3b`/`4f4b5ad`/`2cc670e`.