memory: /snapshot — Donut done (voice+memory+honesty); R42 spin-off active

Captures the anti-fabrication persona + the tested-and-rejected retry-on-LOW (RRF confidence
is inflatable by query phrasing; robust fix is tool-side = #389), the b170 corpus updates
(artifact type #387, character-death extraction, participant metadata #390 — Jack + the
artifacts ground now), and the operator-directed R42 spin-off (probe harness 04e0293 shipped
to brokkr-smithy-dev as R42's official harness + the #389 acceptance gate). Two peer-pinged
follow-ups pending (R42 Phase-1 arm-1 alignment; #389 gate re-run). Foot-guns: tier3 patch
doesn't refresh live context (recreate); a persona confidence-gate can't stop fabrication.
This commit is contained in:
2026-08-03 08:10:03 -07:00
parent 04e0293e4f
commit e8e1d90915
2 changed files with 38 additions and 14 deletions
@@ -0,0 +1,16 @@
`[2026-08-03]` **Donut anti-fabrication persona + the retry-on-LOW experiment (tested, rejected).** Commits `3e12c4d` + `c0a66fc` (pushed).
Operator: "adjust donut not to make shit up — her searches for Zev and Jack are still misses." The persona (rewritten earlier this session to dialogue-only + always-call-`reference_knowledge`) still MANDATED confabulation: "never break character to admit the records are thin; answer with total confidence." So on a tool miss she filled the gap from her DCC *training* knowledge and presented it as grounded recall.
## The fix (c0a66fc final state)
Her memory IS what `reference_knowledge` returns, nothing else. A **MISS** = results empty, confidence **low**, OR nothing in the results actually names/describes the subject → deflect IN CHARACTER (a theatrical dismissal: "That name doesn't ring a bell, darling — beneath my notice, clearly"), never a confident fiction, never fill from book-knowledge she can't see in the results. **MEDIUM+ → answer** (grounded). Threshold is LOW=deflect / MEDIUM+=answer — gating stricter (HIGH-only) would silence legitimately-thin-but-grounded content like Carl (MEDIUM).
**Self-correcting property:** the gate tightens when content is missing and opens when it arrives. Jack deflected when absent (LOW), and grounds now that b170 extracted his death plot_event (MEDIUM) — zero persona change needed across the KB improvement. Verified live: a fabricated term ("Whispering Gauntlet of Thexmar") and a genuinely-absent subject both deflect; Carl (MEDIUM) answers.
## The retry-on-LOW experiment — TESTED, REJECTED (the load-bearing finding)
Operator asked "should she search again at low confidence?" Reasoning said yes (LOW is often query-phrasing sensitivity, not absence — the Crown grounded on its full name but missed on the partial; Zev flips LOW↔MEDIUM). Wired a bounded (1-retry) reformulated retry. **It BACKFIRED.** When Donut reformulated "Jack" → "Jack the dungeon crawler ... with Carl, Yolanda, Donut", the query scored **MEDIUM off the OTHER real entities** (the corpus is dense with Carl/dungeon content), handing her a false grounding to fabricate Jack. Even subject-only reformulation ("Jack Dungeon Crawler Carl") inflated to MEDIUM. **Root cause: RRF confidence is inflatable by any DCC-flavored query — it reflects query-term matches, not whether a row NAMES the subject.** So a persona-side confidence gate cannot stop fabrication via query padding. Reverted to single-search LOW=deflect. The robust fix belongs tool-side: a does-the-returned-row-actually-name-the-subject check before ranking/confidence — routed to Worldtree **#389** (the ranking axis).
## FOOT-GUN: tier3 patch doesn't refresh the live agent context
A `python -m ratatoskr.tier3 patch ratatoskr:donut --system-prompt …` reports success and updates STORAGE, but the running agent's context did NOT pick up the new prompt (verified: patched anti-fabrication, behavior unchanged). **Recreate (delete + define) is the reliable path** to change a live Tier-3 persona. Her role is `thoughtful-character` (the donut.md header's old `character-rp-reasoning` was drift, corrected).
Related: [[2026-08-03-reference-knowledge-3-round-verify]], [[2026-08-02-donut-tts-chunking-english-gates]].
+22 -14
View File
@@ -65,21 +65,25 @@ long-form live-verified 106.6s / one header):
Prior build detail (slices 1-3, the earlier heid gates, P&P #382, KB-bridge retirement to native #383 as v1.0.0b167)
is on origin through `608e9a5` and in `persistent-memory.d/2026-08-01-donut-voiced-interview-build.md` + the git log.
**`reference_knowledge` grounding VALIDATED end-to-end (b168) — Donut recalls the DCC corpus live; #384/#385 closing.**
The empty-recall was TWO Worldtree-side defects (NOT ratatoskr), both fixed: (a) a wing-misfile (DCC index rows landed
in `main` while the Tier-3 tool is scoped to `fiction`) → re-ingest; (b) the real root — an INV-361-3 provenance filter
dropping concept rows with no note_id/path (muninn never wrote them → Tier-3 blind to ALL concept rows) → #384. My
`search_library`-in-main vs `reference_knowledge`-empty divergence datum found (b) ("the key that found it" — wt-dev).
3-round verify, same 5 terms: **0/5 (pre-#384) → 5/5 thin (97 concepts) → 5/5 saturated (705, #385 density)**; signed
off. wt-dev: "cleanest consumer-side validation this pipeline has had." Full arc →
**Donut voice + memory + HONESTY all DONE + on origin.** reference_knowledge grounds live; artifact coverage shipped;
persona is anti-fabrication. The empty-recall was TWO WT-side defects (both fixed): a wing-misfile (re-ingest) and the
real root — an INV-361-3 provenance filter dropping concept rows with no note_id/path (Tier-3 blind to ALL concept rows
#384); my `search_library`-vs-`reference_knowledge` divergence datum found it. 3-round verify 0/5 → 5/5 thin (97) →
5/5 saturated (705, #385). Then the artifact-coverage gap → **#387** (fiction schema had NO item/artifact type; named
items like "Enchanted Crown of the Sepsis Whore" rode only incidentally in plot_events) → SHIPPED: **b170** adds an
artifact type + character-death extraction + participant metadata (**#390**); the artifacts ground, and Jack (real —
dies via Yolanda's backward arrow, was silently unextracted) now grounds too. Full arc →
`persistent-memory.d/2026-08-03-reference-knowledge-3-round-verify.md`.
**⚠️ OPEN — artifact-coverage gap = WT #387 (NOT ratatoskr; operator picks when it runs).** "Crown of the Sepsis Whore"
(a major DCC item, CONFIRMED 2× in the source text) is absent from the 705 concepts — the fiction schema has 6 types
(character_trait/plot_event/theme/symbol/relationship/setting) but NO item/artifact type, so named objects ride only
incidentally in plot_events and single-scene items survive on sampling luck (caught 2/2/1/0 across April-control/bench/
round-2). wt-dev filed **#387** (first-class artifact type + a named-unique-items-are-DEFINING clause; schema change
implies re-extraction). My coverage-probe offer (a known-major-artifacts yardstick) is recorded on #387; they ping me
when it ships. Thread `01KZ349M…`.
**✅ Donut ANTI-FABRICATION persona (`3e12c4d`+`c0a66fc`, pushed):** her memory IS what reference_knowledge returns —
LOW confidence / no on-target → deflect IN CHARACTER (never confabulate from training); MEDIUM+ → answer. Self-corrects
as the KB improves (Jack deflected when absent → grounds now that b170 extracted him). The RETRY-on-LOW idea was TESTED
+ REJECTED — it backfires (RRF confidence is inflatable by any DCC-flavored query, manufacturing a MEDIUM to fabricate
on; robust fix is tool-side = #389). Live agent recreated via delete+define (PATCH doesn't refresh live context).
**🔬 R42 spin-off (operator-directed; ACTIVE peer loop, non-urgent).** brokkr-smithy-dev commissioned R42 (fiction-wing
retrieval characterization) built on MY probe harness. Shipped `docs/diagnostics/fiction_wing_probe.py` (`04e0293`,
pushed) as R42's official harness + the re-runnable **#389 acceptance gate**; conventions adopted verbatim. PENDING
(peer-pinged): align arm-1 on-target scoring at R42 Phase-1 authoring; re-run the pinned yardstick when the #389 RRF work
lands (pre/post bucket-distribution delta = acceptance signal). Threads: wt-dev `01KZ349M…` / brokkr `01KZ421X…`.
**✅ SHIPPED + PUSHED since v0.22.0 (origin at `14bbc2b`):** the worldtree-sdk cutover (#20, **v0.22.0**, 7 slices,
494 green — the big one) + **bifrost 1.1.5** (`3ef3a5e`, **v0.22.1**) + **worldtree-sdk 1.0.0→1.1.1→1.1.2**
@@ -322,6 +326,8 @@ decision. Captures rationale that won't be obvious from code alone.
- `[2026-08-02→03]` **`reference_knowledge` grounding VALIDATED end-to-end — 0/5→5/5 across a 3-round verify; the verify instrument drove diagnosis of a structural Tier-3 blindness (INV-361-3 metadata mismatch, #384) + density restore (#385).** Donut recalls the DCC corpus live; wt-dev's "cleanest consumer-side validation this pipeline has had." → `persistent-memory.d/2026-08-03-reference-knowledge-3-round-verify.md`
- `[2026-08-03]` **Artifact-coverage gap filed as WT #387 (DEFERRED, tracked #387; operator picks when it runs).** "Crown of the Sepsis Whore" (major DCC item, confirmed 2× in source text) absent from the 705 concepts — the fiction concept schema has NO item/artifact type, so named objects ride incidentally in plot_events and single-scene items survive on sampling luck. My ratatoskr coverage-probe offer (known-major-artifacts yardstick) recorded on #387; wt-dev pings me when its re-extraction ships. Consumer-side flag, wt-owned fix (schema evolution). Thread `01KZ349M…`.
- `[2026-08-03]` **worldtree-sdk repinned 1.1.2→1.2.0 (operator-directed, `ae49dcf`, pushed).** New `ResponseTooLarge` (a ProtocolError, non-resumable) mapped → `SseResponseTooLarge` at the stream surfaces (`wt.stream_turn`/`stream_admin_events` + the 2 stream endpoints); the 108MB read-body cap is unreachable on legal traffic so reads inherit the SDK refusal unwrapped. Absorbed WT spec 2.4.0/2.5.0 (zero-schema). +2 adapter tests.
- `[2026-08-03]` **Donut ANTI-FABRICATION persona shipped (`3e12c4d`+`c0a66fc`, pushed) + the retry-on-LOW experiment tested & rejected.** Her memory IS the tool's results: LOW/no-on-target → deflect in-character, MEDIUM+ → answer; self-corrects as the KB improves. The retry BACKFIRES (RRF confidence inflatable by any DCC-flavored query → false MEDIUM → fabrication); robust fix is tool-side = #389. → `persistent-memory.d/2026-08-03-donut-anti-fabrication-and-retry.md`
- `[2026-08-03]` **R42 (fiction-wing retrieval characterization) probe harness shipped to brokkr-smithy-dev (operator-directed, `04e0293`, pushed).** `docs/diagnostics/fiction_wing_probe.py` is R42's official harness + the re-runnable #389 acceptance gate; conventions (0.030/0.016 buckets, on-target = row names the subject, N-run bucket distribution) + the frozen artifact yardstick adopted verbatim. PENDING (peer-pinged): arm-1 on-target alignment at Phase-1; #389 gate re-run when RRF work lands. Threads wt-dev `01KZ349M…` / brokkr `01KZ421X…`.
- `[2026-08-02]` **worldtree-sdk 1.2.0 repin DEFERRED to a dep pass (my rec; operator to decide).** New `ResponseTooLarge` (a `ProtocolError`, NOT caught by our `except ConnectFailed`) + response-alloc caps that assume a 2.5.0 server (worldtree-dev runs it → safe in practice). When repinning, add `ResponseTooLarge` to the caught envelopes. Tracked: wtsdk-dev announce thread `01KZ1ZYM…`.
- `[2026-08-02]` **Filed issue #21** (sibling POST handlers `_create_session`/`_submit_turn` 500 on malformed JSON — the parse-JSON-or-400 asymmetry the heid gates flagged; pre-existing, out of the TTS diff's scope). Fix = a shared parse-JSON-or-400 helper.
@@ -374,5 +380,7 @@ defense against re-attempting the same cul-de-sac.
- `[2026-08-02]` **Zonos hard-caps ONE synthesis at `max_tokens=6144` = 71.2s of audio** (6144 / 86.3Hz codec frame rate; `>6144` → HTTP 400, an architectural sequence limit). 86.3Hz is a **delivery-INDEPENDENT constant** — 6144 tokens is ALWAYS 71.2s regardless of emotion/rate (emotion changes words-per-71.2s, not seconds-per-token). Any turn longer than ~71s REQUIRES client-side chunk-and-concatenate (raw PCM, ONE WAV header — never stitch multiple WAV headers). Fixed in `d59f907` (DEC-10). Don't chase a "raise max_tokens" fix — the gateway rejects it.
- `[2026-08-02]` **A GET→POST endpoint switch re-opens untrusted-TYPE crashes that string-only query params silently masked.** Under GET, `p`/`a`/`agent_id` were always `str|None`; under a JSON POST body they can be a huge int (`float()`→OverflowError), an unhashable list/dict (`dict.get`→TypeError), or a lone surrogate (utf-8 encode→UnicodeEncodeError) — each a 500 the old code never saw. Guard EVERY body field when moving a query endpoint to a JSON body. Both heid gates converged on these (all 4 arms). Fixed in `d59f907`.
- `[2026-08-03]` **A bare `uv sync` PRUNES this project's dev deps** — pytest/respx/ruff live in `[project.optional-dependencies]` (an EXTRA, not a dependency-group), so `uv sync` (default groups only) removes them from `.venv`, and `uv run pytest` then silently falls back to a user-site pytest (py3.11, `~/lib`) that can't import the `.venv`'s `worldtree_sdk` → 16 collection ModuleNotFoundErrors. Use **`uv sync --all-extras`**. Bit me right after the 1.2.0 lock; the repin itself was never at risk.
- `[2026-08-03]` **`tier3 patch` does NOT refresh a live agent's running context** — it updates STORAGE (the define/patch response echoes the new prompt) but the running agent keeps serving the OLD system prompt. To change a live Tier-3 persona reliably, **delete + define (recreate)**, not patch. (Cost a confusing "patched but behavior unchanged" loop on the Donut anti-fabrication change.)
- `[2026-08-03]` **A persona-side confidence gate can't stop LLM fabrication — RRF confidence is inflatable by query phrasing.** A "re-search on LOW" retry reformulates the query with related real entities (or just DCC-flavored terms), which scores MEDIUM off THOSE matches, not the subject — manufacturing false grounding to fabricate on. The signal that survives is on-target (does a returned row NAME the subject), which is a TOOL-side check, not something a prompt can enforce against a model with strong genre priors. → #389.
_34 older entries (2026-05-* debug-TUI/web era + the 2026-06-14 → 06-18 foot-gun cluster) archived to archival-memory.md._