vh c17af18351 fix(sse_client): address Volva code-vs-contract drift (issue #1)
Volva's code-spec review (thread 01KS4CP6ZZ1F) surfaced four code-vs-
contract drift findings on the TDD-passing implementation. All four
addressed here; no contract amendments required.

1. _iter_events fell off the end of aiter_sse() normally on clean EOF
   before any Done/Error/Cancelled. Per INV-001 the iterator MUST NOT
   raise StopAsyncIteration before a terminal event unless the HTTP
   connection drops, in which case it raises SseConnectionDropped.
   Clean EOF before terminal is the same semantic — the stream ended
   without delivering its contracted invariant. Fix: track terminal_seen
   inside _iter_events; after the async-for completes, if not seen,
   raise SseConnectionDropped(last_seen_sse_id=...). Two new tests:
   test_clean_eof_before_terminal (one text then EOF) and
   test_zero_event_eof (empty stream — last_seen_sse_id is None).

2. SseConnectFailed and CancelFailed both store .body without
   truncation; ERROR_ROUTING specifies resp.read()[:1024]. Fix
   truncates in each exception's __init__ before storing. New test
   test_connect_failed_body_truncated (503 + 5000-byte body → 1024)
   and test_cancel_failed_truncates_body (same shape on cancel).

3. _parse_sse_id PRE-001 specifies `assert isinstance(raw, str)`.
   Previous code called raw.split(":") directly, which raises an
   incidental AttributeError on non-str inputs — not the contracted
   precondition path. Fix adds the assert. New test
   test_non_string_input covers int and None.

4. Cancel ERROR_ROUTING said httpx.HTTPStatusError other status →
   CancelFailed, but no test exercised the branch. test_cancel_failed_
   truncates_body covers this (above) — single test double-covers
   findings 2 and 4.

43 tests GREEN (42 sse_client + boundary smoke); ruff clean.

Meta-note from Volva: TDD caught the main happy/adversarial SSE shape,
resume header/body, turn-id flip, and cancel races. The remaining
misses were "negative space" cases (clean premature EOF, exception
payload truncation, untested generic cancel branch). Calibration
evidence that cross-model review pulls weight on what same-model
TDD's hypothesis-space doesn't probe.
2026-05-20 21:33:32 -07:00

Ratatoskr

A Worldtree Conversation API debug TUI. Runs up and down Worldtree's API surface — sessions, turns, persona, tools, admin events, Bifrost state — carrying messages between layers. Like the squirrel.

The product is the observability surface; chat is the input mechanism. Devs run Ratatoskr against a local Worldtree to watch a turn flow through every layer of the system, side-by-side, in one terminal.

Status

v0 scaffold. Design locked; implementation hasn't started. The dev team owns the implementation pass.

Read in this order

  1. docs/design-brief.md — the locked design. Read this first. Every architectural decision is recorded with its rationale, the alternatives considered, and (where relevant) the operator's lock-in moment.
  2. docs/SPEC-PIN.md — what Worldtree spec version Ratatoskr is built against, where the vendored snapshot lives, and how to bump the pin.
  3. docs/conversation-api-spec.md — the vendored Worldtree spec snapshot. Read this to understand the API surface Ratatoskr consumes. Do not import anything from a Worldtree checkout — the boundary is the spec, not the code. See docs/design-brief.md §2.
  4. CLAUDE.md — Claude Code conventions for this repo (mostly inherited from the Corviduo template).
  5. persistent-memory.md — durable intent across context resets. Update as decisions and state evolve.

Quickstart

# 1. Environment
uv venv && source .venv/bin/activate
uv pip install -e ".[dev]"

# 2. Verify the spec pin
cat docs/SPEC-PIN.md   # documented Worldtree SHA + bump procedure

# 3. Tests (none yet; scaffold only)
uv run pytest

# 4. Run against a local Worldtree (once implementation lands)
# Worldtree must be running:  python -m core.conversation_api
ratatoskr --agent mimir

What this repo is NOT

  • NOT a polished consumer for Worldtree's end-users — that's the web app.
  • NOT a Worldtree admin tool — admin CLI is separate (sessions_cli.py lives in Worldtree).
  • NOT a featuretracking shadow of the web app — when the Conversation API surface grows, Ratatoskr does not necessarily grow with it.
  • NOT a remote-Worldtree debug client — file-tail surfaces (persona log, server log) assume local-dev posture.

The full negative-clause list lives in docs/design-brief.md §6.

Boundary rule

Ratatoskr depends on three things only:

  • httpx + httpx-sse (network layer)
  • textual (TUI framework)
  • Worldtree's published Conversation API spec at the pinned SHA

Hard rule: no imports from a Worldtree checkout. No core.* imports, no from worldtree.*, no shared models, no submodule of Worldtree. The spec is the entire surface. Boundary smoke test at tests/test_no_worldtree_imports.py enforces this in CI.

Version-skew strategy

The dev team is decoupled from Worldtree's dev team. To detect drift when Worldtree changes the API:

  1. Spec-version pin in pyproject.toml (worldtree-spec-rev). Bump explicitly; bumps are a tracked action.
  2. Recorded-SSE snapshot tests at tests/snapshots/. Captured against a live Worldtree; replayed in CI. Re-record after every pin bump.
  3. Conformance smoke test — boots Worldtree via Docker compose in CI, runs a one-turn happy path. Catches integration-level drift.

See docs/SPEC-PIN.md for the bump procedure.

Consumer-side discoveries

If you find a spec gap, ambiguity, or missing-but-needed endpoint while working on Ratatoskr: route the discovery back to worldtree-dev via althing rather than via PR on Worldtree directly. The separate-team boundary is intentional and helps catch spec gaps that an in-tree consumer would paper over.

  • Worldtree — the API Ratatoskr consumes
  • brokkr-smithy — authored this design brief
  • Skaldsong — another Worldtree API consumer (Python; render-and-aggregate pattern)
  • mead-hall — another Worldtree API consumer (TypeScript; server-side broker)
S
Description
No description provided
Readme 20 MiB
Languages
Python 82%
HTML 17.6%
Shell 0.4%