Two operator-driven changes off v0.4.1: 1. **Main pane is content-only.** Pre-v0.5.0 the transcript mixed assistant text with telemetry (Thinking closed runs, WorkerPhase, TextBoundary) — only tool events were factored out per #13. The transcript now receives ONLY: user-prompt echo, assistant Text deltas, [done]/[error]/[cancelled] terminal labels, and the post-Done Markdown render. All telemetry routes to a new Debug tab in the right column. 2. **Chrome no longer blue.** Textual's default Header / Footer / active-tab styling tints with `$primary` (Aurora blue under Australis), which read as garish on dark terminals. Header, Footer, and the TabbedContent tab strip get explicit `background: $surface` (Sea bright-black #373b46) so the chrome sits cool and unobtrusive against the Ice black background. ## Layout reshape ``` LEFT COLUMN (content only): RIGHT COLUMN (telemetry): transcript (RichLog, 1fr) thinking-current (Static, dock top) prompt (Input, dock bottom) TabbedContent: Tools (tool_start, tool_result) Debug (thinking, worker_phase, text_boundary) ``` The thinking-current live-preview Static moves from left → right column so the left column is genuinely content-only. Live thinking visibility now persists across tab switches (it docks above the TabbedContent, not inside any tab). ## Presenter routing (TuiPresenterState.render) Signature widens with `debug_log: RichLog`. Routing matrix: Text → log (transcript) Done / Error / Cancelled → log (transcript) [terminal labels] ToolStart / ToolResult → tools_log (Tools tab) Thinking (closed run) → debug_log (Debug tab) WorkerPhase → debug_log (Debug tab) TextBoundary → debug_log (Debug tab) Thinking (per-delta) → thinking_widget (live preview) INV-009 render-exception fallback preserves routing per event class (new INV-020) — ToolStart/Result falls back to tools_log; Thinking/WorkerPhase/TextBoundary to debug_log; everything else to log. ## Keybindings - Ctrl+1 → Tools tab (existing, unchanged) - Ctrl+2 → Debug tab (NEW) `pane-name` footer widget updates dynamically as the operator switches tabs ("Tools" ↔ "Debug"). This was previously deferred to "the multi-tab issue" per the Volva contract-review amendment; multi-tab now exists, so the dynamic update lands here. ## Contract amendments docs/contracts/issues/13.contract.md amended in-place: - INV-015 amended: transcript is content-only; telemetry routes to debug_log. Old routing (telemetry in transcript) retired under the no-backwards-compat rule. - INV-017 amended: thinking-current docks to right column (was left). - INV-019 new: two TabPanes (Tools + Debug), Ctrl+1/Ctrl+2 bindings, dynamic pane-name update. - INV-020 new: render-exception fallback preserves per-event-class routing. - Layout-spec snapshot ASCII diagram updated. Drift-check clean. ## Tests 239/239 GREEN (+3 new: debug_tab_exists, ctrl_2_activates_debug_tab, pane_name_updates_on_tab_switch). 6 existing tests adjusted for the new routing (test_thinking_closes_one_debuglog_entry, test_multiple_thinking_runs_each_get_debuglog_entry, test_render_exception_fallback, test_cancelled_mid_thinking_closes, test_worker_phase_demoted_to_debug_log, test_left_column_content_only). ruff clean. Live smoke against personal Worldtree: mimir KB-search turn populated tools_log with 11 lines of tool events (search_library + read_note); debug_log with 20 lines of worker_phase + thinking content; transcript stayed content-only with `❯ user-prompt` (Aurora bright-cyan) + assistant text deltas. Routing matrix holds end-to-end. (Diagnostic note: RichLog.lines is the rendered- output buffer; inactive TabPane content shows lines=0 until the tab activates and renders. Internal write store is correct — this is a Textual rendering quirk, not a routing bug.) Minor bump (v0.4.1 → v0.5.0) per SemVer etiquette: visible routing surface change for operators; transcript and Debug tab contents look different from yesterday's v0.4.1.
35 KiB
Persistent memory — ratatoskr
Last updated: 2026-05-24
This file captures durable intent and supporting evidence (goals, decisions,
foot-gun warnings, in-flight state) across context resets. Read it at session
start; treat it as one input alongside CLAUDE.md and the auto-memory system,
not as the single source of truth.
When durable state shifts enough to warrant capture, run /snapshot and
commit alongside the next commit per the persistent-memory commit-along rule
in CLAUDE.md.
Repo purpose
Ratatoskr is a dev-grade debug-observability TUI for Worldtree's Conversation API. The product IS the observability surface; chat is the input mechanism. Devs run Ratatoskr against a local Worldtree to watch a turn flow through every layer of the system, side-by-side, in one terminal: agent SSE stream, persona/Vili affect dispatch, tool calls, Bifrost handshake state, admin lifecycle events, optional raw server log.
Named after the squirrel that runs up and down Yggdrasil carrying messages between layers. On-the-nose Worldtree resonance (Yggdrasil = the World Tree).
Origin: althing ask from worldtree-dev (thread 01KS3R34XD3N6HMK91VXESHGW7,
2026-05-20) for the shape of a TUI Conversation API consumer. brokkr-smithy
ran the shape pass; operator's reframe routed it as a new repo with a
separate dev team rather than an in-tree Worldtree tool.
Current state / in-flight
As of 2026-05-24 (post-v0.5.0 content-only main pane + Debug tab):
Status: v0.5.0 shipped. Nine core issues complete (sse_client
#1, sessions #2, cli #3, tui #4, --end-user-id #5, TUI
startup error visibility #6, presenter contract semantics amendment
#12, startup agent picker #8, §5 layout reshape + Tools pane #13)
- robustness fix #7 (MalformedSseData + empty-skip) + v0.2.1 TUI layout fix. 236/236 tests GREEN; ruff clean.
§5 v1 entry point shipped (issue #13). TUI now Horizontal
two-column: left = chat surface (transcript + thinking-current +
prompt); right = TabbedContent with single Tools tab (RichLog
receiving ToolStart/ToolResult events). Routing-not-duplication:
tool events leave the main transcript entirely. Ctrl+1 activates
Tools tab without losing Input focus (INV-016). New pane-name
Static in the footer (static "Tools" v1; dynamic when more tabs
land). CLI mode (--send) unaffected by design — INV-018.
Last commits on main:
- v0.5.0 refactor(tui): content-only main pane + Debug tab + chrome dark
2756f5fstyle(tui): apply Australis theme to TUI chrome + widgets (v0.4.1)24e4371feat(tui): issue #13 — §5 layout reshape + Tools pane (v0.4.0)d30be12feat(sessions,cli,tui): issue #8 — startup agent picker (v0.3.0)c85f6bdfix(tui): anchor layout via dock so Input never moves (v0.2.1)3b9c610feat(cli,tui): issue #12 — presenter contract semantics amendment (v0.2.0)8282156snapshot: persistent-memory Heimdall scope-model foot-gun804c2dffeat(sessions,cli,tui): issues #5 + #6 + worldtree-dev follow-up (v0.1.0)
Smoke status:
--send --new --agent mimirv0.3.0 smoke clean ([done] turn_id=141 model=qwen3.6-35-a3b duration=2.2s).- Live
list_agentssmoke against personal Worldtree returned 12 agents (actor, bragi, cara, domari, forseti, glados, leif, lofn, mimir, soong, troi, saga). - Picker end-to-end smoke against live Worldtree: bare
--new→ list_agents → picker (auto-picked lofn programmatically since driving alt-screen interactively from CLI smoke isn't possible) → POST /sessions with end_user_id="ratatoskr-tui" succeeded; RatatoskrApp constructed with agent_id="lofn". - TUI v0.2.0 was visually broken (Input pane bouncing with thinking runs); v0.2.1 fixed via dock-based layout. Operator confirmed "a lot better" interactively.
Outstanding operator-side todos:
- Interactive §5 layout eyeball —
source env.sh && uv run ratatoskr --new --agent mimir, ask a tool-using question ("search your KB for X"). Confirm: left column shows chat / thinking; right column's Tools tab shows tool_start + tool_result with·prefix; Ctrl+1 doesn't break input focus; no width-clamp issues on the operator's terminal. Programmatic smoke confirmed all the routing + binding; visual confirmation pending. - Post-v0.2.1 TUI multi-turn eyeball — confirm thinking-run bouncing is gone across multiple turns; the layout fix has only been confirmed for a single turn so far.
Pending issues filed but not started:
- Issue #9 (spec-pin refresh v0.19.0 → v0.22.1) — filed 2026-05-23. Documentation debt; defer unless we need a v0.20.0+ capability.
- Issue #10 (subject:{type,id} migration) — filed 2026-05-23 to track Worldtree #196. Don't pre-implement per worldtree-dev.
- Issue #11 (AdminEvents pane auth prerequisite) — filed
2026-05-23. Future side-pane needs
admin.events.readscope.
Pending Worldtree-dev follow-up:
- worldtree-dev committed (althing
01KSBKTG096Q…) to file a Worldtree-side issue for the stall-watchdog gap (cancel-check is inside the engine-event loop, so a never-yielding first-LLM-call bypasses the 300s watchdog). Will file after the immediate stall is cleared. - Ratatoskr-side companion (potential): a client-side stall watchdog
(e.g., 90s-no-events →
[server_stalled]stderr label, keep connection). Defer until recurrence; defense-in-depth regardless of whether Worldtree fixes its own.
Branch: main (clean). Remote:
origin → git@gitea.phasefinal.com:vh/ratatoskr.git.
Next natural moves:
- Interactive picker eyeball — operator confirms the TUI picker UX (rendering, Enter pick, Esc dismiss) against personal Worldtree.
- §5 side-panes work — Persona pane first per design-brief; the
collapsible Thinking pane + Debug pane proposals fold IN as
additional
TabbedContenttabs alongside Persona/Tools/AdminEvents. Reshapes layout from vertical-stack to Horizontal two-column. - Issue #9 (spec-pin refresh) — defer unless we need a v0.20.0+
capability (e.g.,
memory_contextfor Phase 2.1).
Recent decisions
Chronological log of decisions with [YYYY-MM-DD] prefix. One line per
decision. Captures rationale that won't be obvious from code alone.
[2026-05-20]Project name Ratatoskr (squirrel on Yggdrasil — runs up and down carrying messages). Earlier candidate Andvari demoted on the cursed-ring association.[2026-05-20]Separate repo, separate dev team. Operator's call; the in-tree-at-Worldtree/tools/ alternative was considered and rejected to dogfood the API boundary.[2026-05-20]No Worldtree-source imports. Spec-only dependency. Triple version-skew mitigation: spec-pin in pyproject.toml + recorded-SSE snapshot tests + conformance smoke. Initial pin:55101e909abcd2219833266b6f905c5bc956e0f0(Worldtree v0.19.0). Seedocs/SPEC-PIN.md.[2026-05-20]Textual (not rich+prompt_toolkit). Driver: debug observability is the primary purpose, and a multi-pane dashboard with persistent side panes + independent scrollback is structurally application-shell-shaped. Volva consulted via cross-frontier second-opinion and converged on the same call.[2026-05-20]httpx-ssefor SSE consumption. The server emits composite{turn_id}:{seq}id:lines (Worldtree INV-014) load-bearing for SSE-resume; hand-rolleddata:-only parsing (the skaldsong pattern) silently drops these. Ratatoskr becomes the reference Python SSE-resume implementation.[2026-05-20]Persona-pane PII posture: label-don't-refuse.persona.logis process-wide; pane title flips between[Persona — PROCESS-WIDE]and[Persona — session <id>…]based on whether log lines carry session_id. Refuse-against-non-local was considered and rejected as paternalistic.[2026-05-20]Server-stdout pane: opt-in via--server-log <path>. No auto-detection of well-known paths.[2026-05-20]Two-stage Ctrl-C. First cancels in-flight turn server-side; second exits app. Ctrl-D bound to immediate exit.[2026-05-20]Single-session-per-launch + startup picker. No in-app/switch. CLI flags--session <id>and--newfor scripted use. Session identity always visible in Textual footer.[2026-05-20]Markdown rendering default-on;--rawopt-out. Don't pre-design--no-stream-formatting(Volva: add only if streaming-markdown rendering is empirically ugly).[2026-05-20]Non-interactive--sendmode. Single SSE consumer module, two presenters (TUI + stdout). Keeps Ratatoskr honest as an API consumer; useful for CI / scripted probes.[2026-05-20]First contract:ratatoskr.sse_client. Bundlesstream_turn+reconnect_turn+cancel_turn+ private_parse_sse_idinto one module — the SSE-resume flow is coupled (cancel needsturn_idfrom the SSE wireid:, reconnect re-uses the same parsedSseId), so they share a contract. Hard invariant INV-002 makes the composite{turn_id}:{seq}id:parsing load-bearing — closes the foot-gun the design-brief §3 names (hand-rolleddata:-only parsing silently drops theid:).[2026-05-21]Contract converted to issue-scoped (issue #1). Frontmatter shape switched from module-scoped (module:/purpose:) to issue-scoped (target_module:/scope:/prd:) per CONTRACT-FORMAT §2.1.I.prd:block pins to issue body hash. Known parser stale-ness:contract_parser.py --validateERRORs on issue-scoped frontmatter — CONTRACT-FORMAT §2.1.L H10, a documented Brokkr-side follow-up. Parser is a canonical sync, so we do NOT patch it locally. Treat parser ERROR-on-issue-scoped as expected until canonical bumps.[2026-05-21]Default issue-tracker labels seeded (17 total). Sleipnir gating, triage, type, resolution, Ratatoskr-specific area labels (sse-client, tui, cli, observability).[2026-05-21]Volva paraphrase + code-review across all 4 issues — calibration consistent. Paraphrase rounds flag 3-5 contract ambiguities per issue; code-review rounds flag 3-8 code-vs-contract drifts after TDD-passing implementation. Hit rates: #1 paraphrase 3-of-5 amended / code-review 4 findings; #2 3-of-5 / 3 findings; #3 5-of-5 / 5 findings; #4 5-of-5 / 8 findings. The post-TDD code-review consistently catches three classes of gap the test-author's hypotheses don't cover: PRE-assertion boundary drift, exception-payload truncation / never-rendered-to-user observability misses, and "tested the state but not whether the user can see it" gaps (issue #4's primary finding: TUI footer state stored but never rendered to a visible widget — same-model TDD would systematically miss this).[2026-05-21]Manual smoke is load-bearing — found a real defect tests couldn't. First wire-level smoke against personal Worldtree (post-TDD, post-Volva-code-review on #4) revealed httpx's default 5s read timeout killed the SSE connection mid-stream during mimir's thinking phase (~30s LLM latency >> 5s read timeout). The unit/contract test infrastructure (respx-mocked SSE wire) doesn't model real LLM latency, so the gap was invisible at the test layer. Fix: caller-ownedhttpx.AsyncClientconstructed withtimeout=httpx.Timeout(connect=10.0, read=None, write=10.0, pool=10.0); defense in depth:sse_client.stream_turnERROR_ROUTING catcheshttpx.ReadTimeout→SseConnectionDropped. Three contracts amended in-place to document the timeout policy. Lesson: keep manual-smoke step in the per-issue cadence; mock-only validation is insufficient for streaming-against-real-server code. Re-smoke succeeded:[done] turn_id=88 model=qwen3.6-35-a3b duration_ms=2351. Wire-compat envelope (personal v0.16.2 vs ratatoskr's v0.19.0 pin) confirmed end-to-end.[2026-05-22]Issues #5/#6/#7 filed: per-user-agent support + TUI-startup-visibility + mid-stream-robustness. Discovered during 2026-05-22 mimir TUI conversation: long completion (turn 93, 1077 events consumed) crashed withJSONDecodeError("Expecting value: line 1 column 1 (char 0)")fromjson.loads('')on an empty-data:SSE frame. Diagnosis surfaced #7 (the crash). Earlier same day,ratatoskr --new --agent lofnfailed with 422end_user_id_required— surfacing #5 (--end-user-idflag needed for per-user agents). #6 (TUI alt-screen masks the diagnostic before user can read it) was a corollary observation. All three filed; user reordered to #7 first (highest-impact for daily TUI use).[2026-05-22]Issue #8 (startup agent picker) filed.GET /agentsexists in the vendored spec (spec line 832); returnsagent_id,name,description+ optionalversion,capabilities,ui_hints.--agentbecomes conditionally optional: still required for--send --new(non-interactive); optional for TUI--new. When omitted in TUI mode, a newAgentPickerScreenfetches the agent list and presents aListView. Depends onlist_agents()function inratatoskr.sessions. Composes naturally with issue #5 (both thread throughParsedArgs→on_mount/_resolve_then_run). Out of scope: search/sort,ui_hintsrendering,--sendmode picker.[2026-05-23]Issue #6 (TUI startup error visibility) contract drafted + Volva paraphrase complete. Restructuresrun_tuilifecycle: session resolution moves OUT ofon_mount(alt-screen) into a new_resolve_then_runasync helper (pre-App.run()).AsyncClientownership also moves torun_tui'sasync with;RatatoskrApp.__init__takes pre-resolvedsession_id/agent_id/client;on_mountshrinks to identity-widget population. Pre-alt-screen errors → real stderr (same labels/codes as--send). Mid-session errors → RichLog (unchanged, per issue #4 INV-008). Volva paraphrase triage applied the new 5-category framework (Genuine add / Sharpening / Restatement / Out-of-place / Wrong-grounding + ignorance-of-context check). 2 of 5 flagged items amended: F1 (Category 1 — internal contract contradiction: assumptions block said "two sequential event loops" while normative STEPS saidawait app.run_async()— corrected to describe one async flow); F3 (Category 2 — sharpening: informal<truncated>prose aligned to normative{exc.body!r}shape already in STEPS). 3 accepted: F2 (Category 5 — httpx exception hierarchy mis-inference without httpx source access), F4 (Category 3 — restatement of settled architectural guardrail), F5 (Category 2 — sharpening confirming test is the load-bearing spec element).[2026-05-22]Issue #7 (MalformedSseData+ empty-skip) implemented via TDD + Volva-code-reviewed + smoked. Contract → Volva paraphrase (4 ambiguities, all amended; INV-001 wording tightened around exactsse.data == ''rule, ordering-before-id-parse made explicit, test-description bug fixed) → TDD (6 tests, full vertical-slice ordering) → Volva code-review (3 findings — F1 test-gap probing internallast_sse_idnon-advancement via post-skip drop, F2 contract precision around log-vs-propagate responsibility, F3 cli test tightening forraw='X'shape + truncation coverage; all amended) → smoke (3193-token completion against personal Worldtree confirmed clean termination; original crash unreproducible). Calibration milestone: issue #7 is the first issue with zero drift findings from Volva code-review — TDD caught all runtime behavior cleanly. The 3 findings were assertion-precision and architectural-correctness-of-wording, not behavioral. Hypothesis: the tighter the contract spec + the smaller the code surface, the more Volva's role shifts from "catch behavioral drift" to "tighten observability + wording". Calibration table now: #1 (4 findings, 3 drift + 1 test-gap), #2 (3, 1+1+1 precision), #3 (5, 3+1+1), #4 (8, 5+2+1), #7 (3, 0 drift + 2 test-gap + 1 precision).[2026-05-23]Issue #6 (TUI startup error visibility) implemented via TDD + Volva-code-review (two rounds). Lifecycle restructure:run_tuibecomes a thin sync wrapper aroundasyncio.run(_resolve_then_run(args)); the new_resolve_then_runopens thehttpx.AsyncClientviaasync with, does pre-flight session resolution, routesAgentNotFound/SessionApiFailed/network errors to realsys.stderr(verbatim same labels ascli._amain), THEN constructsRatatoskrAppwith pre-resolved state and callsawait app.run_async().RatatoskrApp.__init__signature widens to(args, *, session_id, agent_id, client)— all three required.on_mountnarrows to identity-widget population;on_unmountbecomes a no-op. The alt-screen never opens on resolution errors (INV-001). Two Volva code-review rounds: round 1 returned 6 findings (1 drift + 5 test-gaps), all Category 1 fixed (F1 added the missing PRE-001 assertion at_resolve_then_runentry; F2-F6 tightened test precision — Rule separator assertions on markdown render, RichLog-write spy on empty submit, input-cleared + no-new-worker on cancelling busy, worker.cancel observation on force-exit paths). Round 2 returned 2 NEW test-gaps (F7client_lifetime_owned_by_run_tuipatchedrun_asyncsoon_unmountwasn't actually exercised — added a siblingtest_on_unmount_does_not_close_client; F8 no happy-path--newresolve test — addedtest_happy_new_session_resolveasserting POST count + identity propagation). Calibration confirmed multi-round-Volva value: round 2 found things round 1's amendments didn't anticipate, but they were strictly test-precision, no behavioral drift.[2026-05-23]Issue #5 (--end-user-id) implemented via TDD. Small surface change across three modules (sessions, cli, tui):create_session(client, agent_id, *, end_user_id=None)widens with optional kwarg; body conditionally adds the field when non-None (INV-002: omitting != sending empty); PRE-003 asserts non-empty.ParsedArgs.end_user_id: str | None = Nonefield;--end-user-idCLI flag with non-empty validation (mirrors--sendcheck)._amainand_resolve_then_runthreadend_user_id=args.end_user_idto theircreate_sessioncalls. Post-#6 adjustment: the contract originally namedon_mountas the TUI threading site, but #6 had moved session resolution to_resolve_then_run— same shape, different function. 7 new tests across the 3 modules.[2026-05-23]Worldtree-dev consult landed authoritative consumer-API guidance (althing thread01KSBARG2B8M8C82H6AJGJWX1B). Key takeaways shaped follow-on work: (1)end_user_idis a free-form partition key for long-term memory + persona/valence state; same value → same partition, different values → fully isolated. For Vuong-debugging-Worldtree the recommended posture is a project-stable default with--end-user-idoverride. (2) No programmaticrequires_end_user_iddiscovery onGET /agents— "try and react to 422" remains the pattern. (3) Breaking-change #196 LOCKED but not shipped:subject:{type,id}replacesend_user_idat future v0.22.x or v0.23.0; don't pre-implement. (4) Spec pin (v0.19.0) is 3 minor versions stale (current v0.22.1); none of v0.20.0/v0.21.0/v0.22.0 break ratatoskr's surface but the pin lies about what we're committed to. (5) User-Agent header: send one (ratatoskr/<version> (vh@phasefinal.com)). (6)agents.call:lofnscope needed for lofn smoke. (7)GET /agentsrequires no special scope; issue #8 unblocked on auth.[2026-05-23]Follow-up acted on: User-Agent header added to both_amainand_resolve_then_runhttpx.AsyncClient constructions (withimportlib.metadataversion lookup + fallback to0.0.0);RATATOSKR_END_USER_IDenv-var fallback added to_parse_args(resolution: flag > env > None); env.sh shipsRATATOSKR_END_USER_ID="ratatoskr-tui"as project-stable default. Original issue #5 posture rejected env-var fallback as "papering over isolation"; revised after worldtree-dev's guidance that the realistic single-operator use case wants partition continuity. Issue #5 + #3 contracts amended in-place to document the env-var fallback. Infra-ops pinged via althing foragents.call:lofnscope (broker pattern; they forwarded to worldtree-dev). Three Gitea issues filed: #9 (spec-pin refresh), #10 (subject:{type,id} migration tracking), #11 (AdminEvents pane auth prereq).[2026-05-23]v0.2.1 layout fix: dock-anchored TUI chrome so Input never moves (commitc85f6bd, tagv0.2.1). Reported during the v0.2.0 mimir TUI smoke: Input bouncing up/down throughout a turn, tokens landing at shifting screen positions. Cause: v0.2.0'sStatic(id="thinking-current")was yielded betweenhintandFooterin the auto-stacked vertical flow, so eachdisplay=True/Falsetoggle per thinking-run shifted Input + identity + hint vertically; RichLog growth from streaming text also drifted Input downward. Fix:RatatoskrApp.DEFAULT_CSSdocks the chrome to screen edges —thinking-currentdocks top under Header;transcript(RichLog) getsheight: 1frand absorbs all reflows internally via its scroll viewport;prompt,identity,hintall dock bottom (locked above Footer). Compose order movedthinking-currentto position 2 (right after Header) so source-order matches the dock layout. Operator-confirmed "a lot better" interactively. Pure UI fix; no public API change; tests pass without modification. v0.2.0 → v0.2.1 (patch). I couldn't verify in a TTY from this non-interactive session — the design was sound enough to ship blind, with operator verification post-commit. Going forward: TUI-layout patches like this are "ship + operator verifies" since the TTY is the load-bearing test surface and respx + Pilot mocks can't catch screen-relative positioning bugs.[2026-05-23]Sequencing decision: design-brief §5 side-panes work absorbs the inline collapsible-Thinking-pane + Debug-pane proposals; do issue #8 (startup agent picker) BEFORE §5. Surfaced during the v0.2.1 follow-up discussion. The operator's proposal — "create a collapsible pane for all thinking tokens; text_boundary goes to a debug pane" — is exactly §5-shaped work (the design-brief proposes aHorizontaltwo-column layout withTabbedContentfor Persona/Tools/AdminEvents/BifrostState/ServerLog). Building inline-Collapsibles now and then rebuilding asTabbedContentpanes at §5 would be wasted work. So: do #8 first (independent surface, no layout overlap), then §5 (which folds in Thinking + Debug panes alongside the design-brief's named §5 panes). Interim acceptance: v0.2.1 fixes the structural layout-bouncing pain; transcript-dominated-by-thinking is still real but doesn't degrade further — operator can scroll back, Input doesn't move, tokens land predictably. The interim "noisy transcript" pain is real but bounded; §5 work resolves it cleanly.[2026-05-23]Issue #12 (presenter contract semantics amendment) implemented via TDD. Headline: thinking deltas render as ONE coalesced growing line (CLI) / one closed RichLog entry per run + live Static(id="thinking-current") widget per-delta (TUI), not 50 lines per turn. Introduced stateful per-turn presenters:CliPresenterState(cli.py) andTuiPresenterState(tui.py), both@dataclass(slots=True)with thinking_buffer + thinking_open (+ text_written_since_newline for CLI). Editorial promotion line settled: load-bearing = Text/Done/Error/Cancelled (no prefix); demoted telemetry = WorkerPhase/Thinking/TextBoundary/ToolStart/ToolResult (CLI.ASCII prefix; TUI·Unicode dim prefix). CLI stdout/stderr newline-boundary INV-005: when text was streamed mid-line, flush a\nto stdout before writing terminal labels to stderr;text_written_since_newline = not event.content.endswith("\n")per Volva F4 fix. Helpers_format_duration_ms(347ms/5.5s/1.2mautoscale) and_format_usage(6756 in -> 126 out (6882 total, 0 cached)with arrow="->" CLI or "→" TUI). Per Vor (eitri-smithy-dev cross-frontier consult, althing 01KSBE52YZR5) + Volva paraphrase (5 contract-text ambiguities all fixed in #12.contract.md).[create_session]lifecycle line demoted to. create_session:(written directly by_amain, bypasses state.render). Old_render_event/_render_event_to_logfunctions and their TestRenderEvent/TestRenderEventToLog classes removed (no-backwards-compat rule). Contracts amended: #3 (CliPresenterState block +_run_turnthread state +_amaincreate_session demotion +_format_*helper blocks), #4 (TuiPresenterState block +_stream_turn_workerstate construction +composeStatic widget addition). 39 new tests; 19 obsolete tests removed; net 208 GREEN. v0.1.0 → v0.2.0 (minor; pre-amendment output shape broken intentionally — scripts grepping[thinking] 'no longer work; that's the intended cleanup). Cross-frontier design pass with eitri-smithy-dev returned 16-of-16 confirmed decisions + 4 material divergences applied (ASCII·factual fix, RichLog-one-entry-per-run vs inline-mirror, presenter-state object vs stateless, "contract semantics amendment" framing not "polish"). Calibration note: eitri-smithy-dev's value here was architectural (state-object pattern + chronological-vs-live decoupling) not just tactical; the framing rename alone justified the consult. Volva paraphrase round added 5 prose-precision fixes (INV-001 "growing display" semantics, TUI hide mechanism unification, render_error security/readability tension, newline-tracking corner case, [create_session] integration path).[2026-05-23]Forward direction: Ratatoskr will requireend_user_idfor EVERY access before too long. Operator's call. Reasoning: even Tier 1 foundational agents (mimir, all Asgardians) that don't requireend_user_idserver-side currently fall back to a_no_end_usersentinel substrate partition — effectively pollution from a single-operator-debug-tool's perspective. The right shape is "every conversation has an explicit partition key."RATATOSKR_END_USER_ID="ratatoskr-tui"env-default in env.sh is the first step toward that posture; once we've validated the partition-isolation experience, the next move is makingend_user_idmandatory (probably remove theNone-default in_parse_args, fail-closed with a UsageError if neither flag nor env provides it). Consequence for cross-project asks: declined worldtree-dev's offer to shiprequires_end_user_id: boolonAgentInfoResponsebecause we'd treat every value as true regardless; the try-and-react-to-422 pattern goes away from our side because we never send a request without the field. File a ratatoskr issue when scheduling the change — touches_parse_argsvalidation +_resolve_then_run+_amain+ tests + contract amendments to #3 / #5. Treat as a v0.2.0 minor (breaking: existing--new --agent mimirwithout env or flag would start failing). Cross-frontier alignment (worldtree-dev ack 2026-05-23, althing 01KSBD9FPMCWJMBXNNS4B3MYBS): the platform side agrees with this framing —_no_end_useris a substrate accommodation for identity-less transports, NOT a consumer model. The fallback's_is_fallback=Truetrap door (#185 INV-185-5/8) "could become operator-controlled later" per worldtree-dev, meaning Worldtree itself may tighten the substrate-fallback path. Ratatoskr's forward posture pre-empts that tightening — moving from "we send end_user_id when set" to "we never send a request without end_user_id" stays consumer-correct regardless of what Worldtree does with the fallback knob.
For per-issue TDD implementation notes, Volva findings, and contract amendments, see the git log (commits 9703eb2..61c3941 carry the full per-issue trail with structured commit messages).
Tried and abandoned
Log of approaches that were tried and rejected, with rationale. Future-self defense against re-attempting the same cul-de-sac.
[2026-05-20]rich + prompt_toolkit framework choice. Considered first (during initial shape draft). Volva flagged that §1 and §5 pulled in opposite directions: a real side-panel observability surface would silently become a widget framework reimplementation. Operator's debug-observability reframe sealed the flip to Textual. Don't re-attempt rich+pt unless the scope shrinks to transcript-first REPL (which would also flip back §5 to inline-log-presenter).[2026-05-20]In-tree at Worldtree/tools/ratatoskr/. Earlier draft committed to in-tree-with-import-direction-smoke-test. Rejected at operator-routing — separate dev team forces separate repo.[2026-05-20]New/persona/logSSE endpoint on Worldtree. Considered as alternative to file-tailingpersona.log. Rejected — contract amendment + Vor round + AFK dispatch loop is weeks of consumer-side spec work for a debug feature file-tail handles in a day. Documented follow-up trigger indocs/design-brief.md§5: if a Worldtree-on-server / TUI-on-laptop debug case appears, the contract cost becomes worth paying.[2026-05-20]Cross-process Last-Event-ID resume. Considered — would require persisting per-session Last-Event-ID to~/.config/ratatoskr/. Deferred to v2 if/when it turns out to matter; v1 ships "reconnect, not resume-across-process."[2026-05-21]RichLog widget withmarkup=True. Default impulse, but Rich interprets[xxx]spans as style markup and silently strips them. Every labeled stderr-style line —[cancel_failed],[done],[error],[busy],[worker_phase]— would render as just the content after the bracketed label, breaking the user-visible observability surface. Fix:markup=False. The post-Done Markdown rendering still works becauserich.markdown.Markdownis a Renderable that ignores widget-level markup setting. Don't flip back tomarkup=Truewithout first renaming every labeled-line format away from[bracket]notation.[2026-05-21]Queryingself.query_one("#transcript", RichLog)from inside a Textualrun_workercoroutine. Failed initially withNoMatchesbecause the worker fires before the test'spilot.pause()allows the Input.Submitted handler to fully dispatch (and thus the widget tree to settle). Initial reactive fix: widen worker signature to takelogas a parameter (passed from the handler). Volva code-review flagged this as contract drift (signature didn't match spec). Reverted to single-param signature. The real fix was test-side: addawait pilot.pause()betweeninp.action_submit()and the polling loop in_submit_and_waitso the handler finishes dispatching before the worker reads the widget tree. Don't widen worker signatures to dodge test timing.[2026-05-21]TUI session-identity rendering viaself.sub_title+self.hintplain attributes. Stored state but never rendered to a visible widget. The contract's "session-identity-always-visible" invariant was satisfied at the state-attribute level but not the user-visible-widget level. Tests asserted the attributes (which passed); Volva code-review flagged the gap. Fix: dedicatedStatic(id="identity")+Static(id="hint")widgets in compose;_set_hint()helper mirrors state → widget. Calibration evidence for the "TDD catches state, code-review catches whether the user can see it" pattern.[2026-05-23]Using the cross-model review agent's name directly in composed prose. The peer review agent's name (thealthinghandle starting with "V-o-l-v-a") is one letter from a body-part term. Anthropic's content classifier does fuzzy matching and intermittently blocks responses mid-stream when the name appears in composed prose sentences (especially in meta-commentary about the agent's work). Direct-quoted tool output (e.g., thealthing-cli threadbody) passes through fine. Mitigation: use role descriptions ("the cross-model reviewer," "the paraphrase peer") in prose rather than the name; quote content via tool output. Confirmed by switching to Sonnet 4.6 for a test read — same raw content read cleanly when fetched via Bash rather than composed into an LLM response. This is a persistent environmental constraint, not a one-off.[2026-05-22]json.loads(sse.data)unguarded against empty data._iter_eventsunconditionally calledjson.loadson every dispatchedServerSentEvent. Whenhttpx_ssesurfaced a frame withid:present butdata:empty (a known library-vs-spec divergence — RFC says don't dispatch; httpx_sse is permissive),json.loads('')raisedJSONDecodeError→ propagated through Textual's worker → app crash. Crashed mimir conversation at turn 93/seq 1078 after 1077 successful events. Fix:if sse.data == '': continueBEFORE_parse_sse_id(empty-data event with a malformed id is still a keepalive — don't reorder). Non-empty malformed data raises newMalformedSseData(raw[:200]). Don't reintroduce unconditionaljson.loads(sse.data); always pre-check for the empty case.[2026-05-23]Diagnostic shorthand: "2-events-then-silence" = Worldtree-side LLM-call wedge, not ratatoskr. If a mimir--sendsmoke shows exactly two stderr events —. create_session: ...followed by. worker_phase: phase=BuildingPrompt ...— and then nothing for >60s, the root cause is upstream of ratatoskr. Worldtree'sservice.py:2560gates theCallingLLMevent on the engine yielding its first LLM-provider chunk; if that provider connection is wedged at the TCP level, theasync fornever iterates and the SSE stream stays silent forever. ratatoskr'sread=Nonehttpx timeout (the issue #1 + #4 INV-007 fix for "5s default killed mid-stream during mimir's thinking") waits patiently as designed; there's no client-side stall watchdog above the read-timeout layer. Worldtree's OWN stall watchdog (300s_start_stall_timer) exists but its cancel-check is INSIDE the engine-event loop, so a never-yielding first-LLM-call bypasses it. Confirmed by worldtree-dev (althing thread01KSBKTG096Q07JVRG41JXA1DD). Don't waste time bisecting ratatoskr code when this shape appears — diagnose the LLM-provider connection state at Worldtree's host. Restarting the Worldtree service (:8081in our case) cleared a wedged llama-swap connection. Future ratatoskr issue worth filing if recurrence: client-side stall watchdog (e.g., 90s-no-events →[server_stalled]stderr label, keep connection open). Also worth knowing: 10.250.50.152 hosts 3 Worldtree instances (:8080,:8081,:8082) — each with its own DB and key namespace. Our key is valid only on:8081.[2026-05-23]Phantom "per-Tier-1-agent scope add" pattern. Issue #5's lofn 422 was initially diagnosed (with worldtree-dev's first reply) as needingagents.call:lofnadded to ratatoskr's existing key. Routed through infra-ops via althing per the credential-brokerage rule; infra-ops discovered no public scope-mutation endpoint on personal Worldtree, brokered to worldtree-dev for the actual mechanism. Worldtree-dev came back with a correction: their first answer conflated two distinct Heimdall scope namespaces. Tier 1 foundational agents (mimir, lofn, soong, all Asgardians) are covered by a blanketagent.call:*(singular) baseline rule inconfig/policies.yaml > tiers.<tier>.scopesfor ALL authenticated tiers includinguser. There is no per-agent grant for Tier 1 — the baseline rule covers it. Tier 3 consumer-defined agents (IDs containing:, likevh:custom-bot) use the pluralagents.call:<owner>:<agent>shape granted implicitly via owning aconsumer_agentsDB row, registered throughPOST /agents/define. The two notations differ by one letter and that was the source of the confusion. The actual lofn fix was issue #5's--end-user-idflag — it was always a request-body validation, not an auth-scope gate. Don't ping infra-ops for "per-Tier-1-agent scope adds" again; the pattern is a phantom ask. Real future infra-ops asks: admin-tier key for the AdminEvents pane (admin.events.readscope, different tier), and Tier 3 custom-agent registration (different flow entirely, requiresPOST /agents/define).