Compare commits

..

4 Commits

Author SHA1 Message Date
vh 139771c8d8 feat(tui): live Markdown rendering during text streaming (v0.9.0)
Replaces v0.8.2's drop-Markdown patch with proper in-place Markdown
rendering. The transcript becomes a VerticalScroll container; each
turn's response body lives as a single Static widget whose content
is updated as Text deltas arrive — Markdown is re-rendered in place
rather than re-printed on Done. Eliminates the v0.8.x double-print
without sacrificing rich formatting.

- transcript: RichLog → VerticalScroll (#transcript-scroll)
- Text deltas: mount Static(Markdown(buffer)) on first delta;
  Static.update(Markdown(buffer)) on subsequent deltas
- --raw mode: bypass Markdown, mount Static(plain_str) for the same
  in-place update semantics
- Terminal events (Done/Error/Cancelled) mount styled label Statics
- _cancel_via_sse: write → mount Static on the new container
- _write_turn_headers: transcript gets a styled RichText Static
  ("── turn N ──"); other panes still receive Rule renderables
- Test suite reshape: bulk rename `log` → `transcript` for the
  presenter contract, `_mounted_renderables` helper extracts
  Static.content for assertion, `_spy_writes` captures both
  RichLog.write and VerticalScroll.mount
2026-05-24 22:18:45 -07:00
vh 489cfee1f0 fix(tui): drop post-Done Markdown body re-render (v0.8.2)
Operator: "first turn double prints agent's turn."

Root cause: v0.8.1 wrote both the streamed Text lines AND the post-
Done `Markdown(event.response)` body into the transcript. Same
content rendered twice — once as plain streaming, once as a full
markdown re-render. The v0.8.1 commit message documented this as
"some duplication is acceptable" but the live UX read as a bug.

## Fix

Drop the post-Done `Rule + Markdown(response)` writes in non-raw
mode. The streamed text IS the response; whatever the model emitted
flows into the transcript line-by-line via coalesce-on-newline.
Markdown formatting (bold, lists, code blocks) renders as plain
text — a known regression from v0.8.1's polished output but the
right tradeoff vs the duplication bug.

## What this loses temporarily

Pre-v0.8.2 (after Done):
  [done] turn_id=... ───
  ─── (Rule separator) ───
  **Bold text** rendered bold, `code` highlighted, lists as bullets, etc.

v0.8.2 (after Done):
  [done] turn_id=... ───
  **Bold text** as plain asterisks, `code` as backticks, lists as plain dashes

## v0.9.0 plan

Restore markdown rendering via LIVE rendering during the stream
(not post-Done re-render). Replace `RichLog#transcript` with a
`VerticalScroll` container that mounts a fresh `Markdown` widget
per turn; Text deltas update the widget; markdown renders as
content arrives. No duplication, no snap, full formatting.
Operator-confirmed direction (2026-05-25 AskUserQuestion).

## Tests

287/287 GREEN; ruff clean. Two tests updated for the new shape:
- test_done_renders_markdown_after_label → renamed
  test_done_flushes_tail_and_writes_label; asserts NO Markdown, NO
  Rule (post-Done) in the writes.
- test_happy_text_done_renders_markdown → renamed
  test_happy_text_done_no_double_print; asserts NO Markdown in the
  spy.

Patch bump (v0.8.1 → v0.8.2): bug fix; no public API change.
2026-05-24 21:53:20 -07:00
vh 11ef6830ab fix(tui,sse): inline Text streaming + empty-id keepalive skip (v0.8.1)
Two related fixes for the same user-reported bug pattern from a
running session against ratatoskr:sindra (qwen3.6-35-a3b-heretic):

## 1. Streaming text overlapping the transcript

Operator: "new text comes at the bottom and overwrites the existing
pane information instead of pushing it up naturally."

Root cause: the v0.6.0 `#current-text` Static was `dock: bottom`
with `height: auto`, sitting between the transcript RichLog (1fr)
and the prompt Input (dock: bottom). As text streamed, the Static
grew UPWARD but Textual didn't dynamically resize the 1fr transcript
to accommodate — the growing Static visually OVERLAPPED the
transcript's bottom rows. On Done, `current_text.update("")` snapped
it to height 0 and the transcript re-laid-out — "boom, everything
updates."

Fix: remove `#current-text` Static entirely. Apply the same
coalesce-on-newline pattern v0.7.1 used for thinking — Text deltas
accumulate in `TuiPresenterState.text_chunk_buffer`, flushing whole
lines (each `\n` boundary) directly to `log` (transcript). On Done:
flush remaining tail, then [done] label + Rule + Markdown body.

Trade-off accepted: streamed lines + post-Done Markdown body are
both in the transcript (some content duplication). The Markdown
body re-renders the same content with proper formatting (lists,
bold, code blocks). Acceptable — operator gets both the live-progress
streaming AND the canonical rendered version.

## 2. MalformedSseId raw='' crashing every turn

Operator: "current session is erroring on every turn with
[malformed_sse_id] raw=''"

Worldtree's qwen3.6-35-a3b-heretic provider emits some events
without `id:` lines (observed 2026-05-25 mid-stream). When the FIRST
such event arrives before any prior id has been seen, httpx_sse's
`ServerSentEvent.id` is `""`. `_parse_sse_id('')` raised ValueError
→ MalformedSseId → turn worker bailed → operator saw the label
every turn.

Per SSE RFC, events without `id:` are legitimate (they just don't
update Last-Event-ID). Issue #7 already covered the empty-DATA
keepalive case with skip-silently semantics. Empty-id is the same
shape of wire weirdness; same fix shape:

  if sse.id == "":
      continue  # treat as keepalive

Ordered AFTER the empty-data branch so an empty-data + empty-id
event still gets skipped on the data check.

## Tests + smoke

287/287 GREEN (was 286, +1 for empty-id skip; +1 net Text-flow test
adjustments). Ruff clean.

Verified Worldtree alive when the user hit the empty-id bug
(/healthz returned ok in 18ms) — not a server-down issue, just
wire-format mid-stream.

## Caveats

The fix doesn't recover content from the dropped empty-id event.
If the event happened to carry meaningful data (not a true
keepalive), we silently lose it. Acceptable trade-off: pre-v0.8.1
EVERY turn died on the offending agent; post-v0.8.1 the turn
continues and any single dropped frame is recoverable from logs if
debugging. Worldtree-side fix (always emit ids) is the right
upstream answer; ratatoskr just stops panicking on wire weirdness.

Patch bump (v0.8.0 → v0.8.1) — both fixes are bug fixes; no public
API change. The `TuiPresenterState.render` signature loses the
`current_text` parameter (was added v0.6.0), but presenter is an
internal contract; no external callers.
2026-05-24 21:39:02 -07:00
vh 9fade55901 feat(local_agents): tier-3 index + picker merge (v0.8.0)
Worldtree's GET /agents doesn't return consumer-defined (tier-3)
agents — the public list excludes them by design. Confirmed live in
v0.7.0's smoke. Without server-side knowledge, ratatoskr's picker
couldn't show tier-3 agents the operator had defined; the workflow
was "remember the agent_id, pass --agent ratatoskr:<name>
explicitly." Friction grows with every tier-3 agent.

## Fix: client-side index, merged at picker time

New module `ratatoskr.local_agents` maintains a JSON-backed index at
$XDG_CONFIG_HOME/ratatoskr/local_agents.json (override via
$RATATOSKR_LOCAL_AGENTS). `tier3` CLI define / patch / delete update
the index as side-effects. `tui._resolve_then_run` loads the index
after `list_agents(client)` and appends entries not already in the
remote list (dedup by agent_id; remote wins on conflict).

Library-level `tier3.define_agent` / `patch_agent` / `delete_agent`
stay pure — local persistence lives in the CLI layer (`_run_define`
etc.), not in the library functions. Tests of the library don't
touch the filesystem.

## Public surface

  ratatoskr.local_agents:
    LocalAgentEntry (frozen dataclass)
    load_local_agents() -> list[LocalAgentEntry]
    add_local_agent(entry)
    update_local_agent(entry)  # same semantics as add (agent_id key)
    remove_local_agent(agent_id)
    make_description(system_prompt) -> str  # synthetic picker label

Failure modes are lenient: missing file → empty index; corrupt JSON
or schema mismatch → empty index (no crash). The picker continues
to show foundational agents either way; tier-3 surface degrades to
the pre-v0.8.0 workflow.

## Picker integration

Local entries convert to ratatoskr.sessions.AgentInfo with synthetic
fields:
  name        = agent_name (from LocalAgentEntry)
  description = "(tier 3) <first non-empty line of system prompt>"
  version, capabilities, supported_models, persona_traits, ui_hints
    = None / [] / [] / {} / {}

If Worldtree later starts returning tier-3 in GET /agents, this
module's role narrows to redundant local cache; can be removed
cleanly since the dedup-by-agent-id keeps remote-wins behavior.

## Tests

286/286 GREEN (was 265, +21: 20 local_agents + 1 picker-merge
integration). Ruff clean. Tests isolate the index via
$RATATOSKR_LOCAL_AGENTS pointed at pytest's tmp_path — no pollution
of operator's real ~/.config/ratatoskr/.

## Manual smoke

Sindra-like define against personal Worldtree:
  python -m ratatoskr.tier3 define --name foo --system-prompt "..." --model X
  cat ~/.config/ratatoskr/local_agents.json
  # ratatoskr --new picker now shows ratatoskr:foo alongside mimir et al.

Cross-machine: the file is per-host. Operator can sync via dotfiles
if needed; out of scope for this commit.

Minor bump (v0.7.1 → v0.8.0) — new public module + new picker
behavior (more agents shown). No caller-side breaking changes.
2026-05-24 21:13:30 -07:00
11 changed files with 1031 additions and 299 deletions
+6 -3
View File
@@ -32,9 +32,9 @@ separate dev team rather than an in-tree Worldtree tool.
## Current state / in-flight
_As of 2026-05-25 (post-v0.7.1 thinking coalesce-by-newline):_
_As of 2026-05-25 (post-v0.8.2 drop double-print; v0.9.0 live-md next):_
**Status: v0.7.1 shipped.** Ten core features complete (`sse_client`
**Status: v0.8.2 shipped.** Eleven core features complete (`sse_client`
#1, `sessions` #2, `cli` #3, `tui` #4, `--end-user-id` #5, TUI
startup error visibility #6, presenter contract semantics amendment
#12, startup agent picker #8, §5 layout reshape + Tools pane #13)
@@ -51,7 +51,10 @@ Static in the footer (static "Tools" v1; dynamic when more tabs
land). CLI mode (--send) unaffected by design — INV-018.
Last commits on `main`:
- v0.7.1 fix(tui): coalesce thinking deltas on `\n` — no more per-token newlines
- v0.8.2 fix(tui): drop post-Done Markdown body re-render (no double-print)
- `11ef683` fix(tui,sse): inline Text streaming + empty-id keepalive skip (v0.8.1)
- `9fade55` feat(local_agents): JSON-backed local tier-3 index + picker merge (v0.8.0)
- `9918c10` fix(tui): coalesce thinking deltas on `\n` (v0.7.1)
- `c086ae2` feat(tier3): ratatoskr.tier3 module + CLI (v0.7.0)
- `d356990` refactor(tui): thinking streams into thinking-log (v0.6.5)
- `82437bd` style(tui): picker highlighted item → Aurora blue (v0.6.4)
+1 -1
View File
@@ -4,7 +4,7 @@ build-backend = "hatchling.build"
[project]
name = "ratatoskr"
version = "0.7.1"
version = "0.9.0"
description = "Worldtree Conversation API debug TUI — multi-pane observability dashboard"
readme = "README.md"
requires-python = ">=3.12"
+134
View File
@@ -0,0 +1,134 @@
"""Local index of tier-3 agents defined via `python -m ratatoskr.tier3`.
Workaround for Worldtree's ``GET /agents`` not returning consumer-defined
agents (the public list excludes tier-3 per-spec; see issue #15 smoke
findings). Local file maintains a list of agent_ids + display metadata so
the picker can show them alongside foundational agents.
Storage shape: JSON at ``$XDG_CONFIG_HOME/ratatoskr/local_agents.json``
(default ``~/.config/ratatoskr/local_agents.json``). Override via
``$RATATOSKR_LOCAL_AGENTS`` env var for tests / per-machine isolation.
If Worldtree later starts returning tier-3 agents in ``GET /agents``, this
module's role narrows to redundant local cache; can be removed cleanly
since the picker's dedup-by-agent-id keeps remote-wins behavior.
Failure modes are lenient: missing file → empty index; corrupt JSON or
schema mismatch → empty index (no crash). The picker continues to show
foundational agents either way; the local-tier-3 surface degrades to
"operator passes --agent ratatoskr:<name> explicitly" — the
pre-v0.8.0 workflow.
"""
from __future__ import annotations
import json
import os
from dataclasses import asdict, dataclass
from pathlib import Path
_SCHEMA_VERSION = 1
@dataclass(frozen=True)
class LocalAgentEntry:
"""One row in the local tier-3 agent index.
Schema:
- ``agent_id``: full "user_id:agent_name" string (Worldtree-owned).
- ``agent_name``: slug from define (display name).
- ``model``: provider model ID at last define/patch.
- ``description``: synthetic display string (typically derived from
the system_prompt's first line + a "(tier 3)" prefix; the picker
uses this in its ``{id} · {name}{description}`` rendering).
- ``defined_at``: ISO-8601 timestamp from the Tier3AgentInfo response.
"""
agent_id: str
agent_name: str
model: str
description: str
defined_at: str
def _local_agents_path() -> Path:
"""Resolve the local index file path with XDG + env-var override."""
override = os.environ.get("RATATOSKR_LOCAL_AGENTS")
if override:
return Path(override)
xdg = os.environ.get("XDG_CONFIG_HOME")
base = Path(xdg) if xdg else (Path.home() / ".config")
return base / "ratatoskr" / "local_agents.json"
def load_local_agents() -> list[LocalAgentEntry]:
"""Read the local index. Returns ``[]`` on missing file, corrupt JSON,
schema mismatch, or any read error — never raises.
"""
path = _local_agents_path()
if not path.exists():
return []
try:
raw = json.loads(path.read_text())
except (json.JSONDecodeError, OSError):
return []
if not isinstance(raw, dict) or raw.get("version") != _SCHEMA_VERSION:
return []
agents = raw.get("agents", [])
if not isinstance(agents, list):
return []
out: list[LocalAgentEntry] = []
for item in agents:
if not isinstance(item, dict):
continue
try:
out.append(LocalAgentEntry(**item))
except TypeError:
# Malformed row (missing/extra fields) — skip silently.
continue
return out
def _save_local_agents(agents: list[LocalAgentEntry]) -> None:
"""Persist the index. Creates parent dir as needed."""
path = _local_agents_path()
path.parent.mkdir(parents=True, exist_ok=True)
payload = {"version": _SCHEMA_VERSION, "agents": [asdict(a) for a in agents]}
path.write_text(json.dumps(payload, indent=2))
def add_local_agent(entry: LocalAgentEntry) -> None:
"""Add (or replace) an agent in the local index. agent_id is the key."""
agents = [a for a in load_local_agents() if a.agent_id != entry.agent_id]
agents.append(entry)
_save_local_agents(agents)
def update_local_agent(entry: LocalAgentEntry) -> None:
"""Update an existing entry. Identical semantics to ``add_local_agent``
(agent_id is the dedup key), exposed separately so callers can
self-document intent.
"""
add_local_agent(entry)
def remove_local_agent(agent_id: str) -> None:
"""Remove an entry by agent_id. No-op if absent (idempotent)."""
agents = [a for a in load_local_agents() if a.agent_id != agent_id]
_save_local_agents(agents)
def make_description(system_prompt: str) -> str:
"""Synthesize a one-line description for the picker from a system prompt.
Strategy: first non-empty line, stripped of leading markdown heading
markers and whitespace, prefixed with "(tier 3) ", truncated to 80
chars. Falls back to "(tier 3) custom system prompt" if the prompt is
empty (defensive — define rejects empty prompts at PRE-002).
"""
for line in system_prompt.splitlines():
stripped = line.lstrip("# ").strip()
if stripped:
label = f"(tier 3) {stripped}"
return label[:80] + ("" if len(label) > 80 else "")
return "(tier 3) custom system prompt"
+8
View File
@@ -309,6 +309,14 @@ async def _iter_events(
# with a bad id is still a keepalive). Don't reorder.
if sse.data == "":
continue
# v0.8.1: empty-id frames are also treated as keepalives. Worldtree
# SOMETIMES emits events without an `id:` line (observed mid-stream
# on the qwen3.6-35-a3b-heretic provider, 2026-05-25). Per the SSE
# RFC, events without ids are legitimate (they just don't update
# Last-Event-ID); the previous strict behavior crashed every turn
# on the offending agent. Treat same as empty-data: skip silently.
if sse.id == "":
continue
try:
sse_id = _parse_sse_id(sse.id)
except ValueError as exc:
+33
View File
@@ -309,6 +309,11 @@ def _resolve_auth(ns: argparse.Namespace) -> tuple[str, str]:
async def _run_define(ns: argparse.Namespace) -> int:
api_key, server_url = _resolve_auth(ns)
from ratatoskr.cli import USER_AGENT
from ratatoskr.local_agents import (
LocalAgentEntry,
add_local_agent,
make_description,
)
async with httpx.AsyncClient(
base_url=server_url,
@@ -324,6 +329,16 @@ async def _run_define(ns: argparse.Namespace) -> int:
system_prompt=ns.system_prompt,
model=ns.model,
)
# v0.8.0: persist to local index so the picker can show it.
add_local_agent(
LocalAgentEntry(
agent_id=info.agent_id,
agent_name=info.agent_name,
model=info.model,
description=make_description(info.system_prompt),
defined_at=info.created_at,
)
)
print(f"defined {info.agent_id} ({info.model})")
return 0
@@ -331,6 +346,11 @@ async def _run_define(ns: argparse.Namespace) -> int:
async def _run_patch(ns: argparse.Namespace) -> int:
api_key, server_url = _resolve_auth(ns)
from ratatoskr.cli import USER_AGENT
from ratatoskr.local_agents import (
LocalAgentEntry,
make_description,
update_local_agent,
)
if ns.system_prompt is None and ns.model is None:
raise _Tier3UsageError(
@@ -350,6 +370,16 @@ async def _run_patch(ns: argparse.Namespace) -> int:
system_prompt=ns.system_prompt,
model=ns.model,
)
# v0.8.0: refresh local index with the post-patch state.
update_local_agent(
LocalAgentEntry(
agent_id=info.agent_id,
agent_name=info.agent_name,
model=info.model,
description=make_description(info.system_prompt),
defined_at=info.updated_at,
)
)
print(f"patched {info.agent_id}")
return 0
@@ -357,6 +387,7 @@ async def _run_patch(ns: argparse.Namespace) -> int:
async def _run_delete(ns: argparse.Namespace) -> int:
api_key, server_url = _resolve_auth(ns)
from ratatoskr.cli import USER_AGENT
from ratatoskr.local_agents import remove_local_agent
async with httpx.AsyncClient(
base_url=server_url,
@@ -367,6 +398,8 @@ async def _run_delete(ns: argparse.Namespace) -> int:
timeout=httpx.Timeout(connect=10.0, read=30.0, write=10.0, pool=10.0),
) as client:
await delete_agent(client, ns.agent_id)
# v0.8.0: drop from local index so the picker stops listing it.
remove_local_agent(ns.agent_id)
print(f"deleted {ns.agent_id}")
return 0
+230 -105
View File
@@ -10,13 +10,13 @@ from __future__ import annotations
import asyncio
import sys
from dataclasses import dataclass, field
from dataclasses import dataclass
from typing import ClassVar, Literal
import httpx
from textual.app import App, ComposeResult
from textual.binding import Binding
from textual.containers import Horizontal, Vertical
from textual.containers import Horizontal, Vertical, VerticalScroll
from textual.theme import Theme
from textual.widgets import (
Footer,
@@ -186,9 +186,6 @@ class TuiPresenterState:
"""
thinking_open: bool = False
# v0.6.0: per-turn streaming text buffer. Text deltas accumulate here
# and update `current_text` Static in place — no per-token RichLog spam.
text_buffer: list[str] = field(default_factory=list)
# Thinking-run counter for turn-scoped start/end markers.
thinking_run_index: int = 0
# v0.7.1: thinking-content accumulator. Worldtree emits Thinking deltas
@@ -197,13 +194,20 @@ class TuiPresenterState:
# only on `\n` boundaries (one written line per natural paragraph) or
# when the run closes (any leftover tail).
thinking_chunk_buffer: str = ""
# v0.9.0: Text accumulator for live Markdown rendering. Worldtree emits
# Text deltas at token granularity; each delta appends to this buffer
# and the current_response_widget re-renders Markdown(text_chunk_buffer)
# in place. On terminal event the widget is finalized + reference clears.
text_chunk_buffer: str = ""
# v0.9.0: reference to the Static widget holding the current turn's
# response Markdown Renderable. None between turns.
current_response_widget: object = None
def render(
self,
event: Event,
*,
log: RichLog,
current_text: Static,
transcript: "VerticalScroll",
tools_log: RichLog,
debug_log: RichLog,
thinking_log: RichLog,
@@ -211,18 +215,14 @@ class TuiPresenterState:
) -> None:
"""Render one Worldtree SSE event with the TUI hierarchy + coalescing.
v0.6.5 routing:
- `log` (transcript) = content only: user-prompt echo (written
outside the presenter), terminal labels, post-Done Markdown body.
- `current_text` (Static below transcript) = live-streaming Text
deltas accumulated into one growing line; cleared on terminal.
- `tools_log` = ToolStart + ToolResult.
- `debug_log` = WorkerPhase + TextBoundary.
- `thinking_log` = streaming Thinking deltas inline (each chunk =
one line in the scrollable log). Rule(start)/Rule(end) markers
wrap each run. The whole pane scrolls naturally — no separate
tail-scrolling Static at the bottom (v0.6.5 removed
`thinking-current`).
v0.9.0 routing:
- `transcript` (VerticalScroll) = chat content: each turn mounts
child widgets (turn-header / prompt-echo / response Markdown /
done-label). Live Markdown rendering during Text streaming.
- `tools_log` (RichLog) = ToolStart + ToolResult.
- `debug_log` (RichLog) = WorkerPhase + TextBoundary.
- `thinking_log` (RichLog) = streaming Thinking deltas inline
(coalesced on `\n`); Rule(start)/Rule(end) wrap each run.
Exceptions caught at the presenter boundary (INV-009 fallback).
"""
@@ -284,51 +284,77 @@ class TuiPresenterState:
self.thinking_open = False
# Now render the non-thinking event itself.
if isinstance(event, Text):
# v0.6.0: streaming text accumulates into current_text Static
# — one growing live line, NOT per-delta RichLog entries.
self.text_buffer.append(event.content)
current_text.update("".join(self.text_buffer))
# v0.8.1: stream Text deltas into transcript directly,
# coalesced on `\n`. Same pattern as Thinking (v0.7.1).
# The pre-v0.8.1 #current-text Static is gone — its dock-
# bottom growth was overlapping the transcript visually.
#
# v0.9.0: Text deltas accumulate in text_chunk_buffer and
# the current_response_widget renders Markdown(buffer) in
# place. First Text delta of the turn mounts a fresh Static
# holding the Markdown Renderable; subsequent deltas update
# the same widget. Live markdown rendering — no post-Done
# re-render needed.
from rich.markdown import Markdown
self.text_chunk_buffer += event.content
# --raw bypasses Markdown rendering — useful for debugging
# the raw text stream surface, and matches the pre-v0.9.0
# --raw semantics (which dropped the post-Done Markdown re-
# render). In raw mode the response widget holds plain str.
rendered = (
self.text_chunk_buffer if raw else Markdown(self.text_chunk_buffer)
)
if self.current_response_widget is None:
self.current_response_widget = Static(
rendered, classes="response-md"
)
transcript.mount(self.current_response_widget)
else:
self.current_response_widget.update(rendered)
transcript.scroll_end(animate=False)
return
if isinstance(event, (Done, Error, Cancelled)):
# Terminal event: clear the streaming Static first so the
# live-preview band collapses. Then write the colored label
# + (non-raw) Markdown body / (raw) accumulated plain text
# to the transcript.
accumulated = "".join(self.text_buffer)
self.text_buffer.clear()
current_text.update("")
# Terminal labels tinted per outcome (Aurora green / Dawn red
# / Dawn yellow) for at-a-glance scanning.
# Terminal event: finalize the response widget (clear ref so
# the next turn mounts a fresh one). The accumulated text is
# already rendered as Markdown in the widget — no post-Done
# re-render, no double-print.
self.text_chunk_buffer = ""
self.current_response_widget = None
# Terminal labels mount as styled Statics. Tinted per outcome
# (Aurora green / Dawn red / Dawn yellow) for at-a-glance
# scanning.
if isinstance(event, Done):
log.write(RichText(
f"[done] turn_id={event.sse_id.turn_id} model={event.model} "
f"duration={_format_duration_ms(event.duration_ms)} "
f"usage {_format_usage(event.usage, arrow='')}",
style=_AU_SUCCESS,
transcript.mount(Static(
RichText(
f"[done] turn_id={event.sse_id.turn_id} "
f"model={event.model} "
f"duration={_format_duration_ms(event.duration_ms)} "
f"usage {_format_usage(event.usage, arrow='')}",
style=_AU_SUCCESS,
),
classes="done-label",
))
if raw:
# Raw mode: emit the accumulated streamed text verbatim
# so the operator has a record after the Static clears.
if accumulated:
log.write(accumulated)
else:
from rich.markdown import Markdown
from rich.rule import Rule
log.write(Rule(style=_AU_DEMOTED))
log.write(Markdown(event.response))
elif isinstance(event, Error):
log.write(RichText(
f"[error] turn_id={event.sse_id.turn_id} code={event.error_code} "
f"message={event.message!r}",
style=_AU_ERROR,
transcript.mount(Static(
RichText(
f"[error] turn_id={event.sse_id.turn_id} "
f"code={event.error_code} message={event.message!r}",
style=_AU_ERROR,
),
classes="error-label",
))
else: # Cancelled
log.write(RichText(
f"[cancelled] turn_id={event.turn_id} reason={event.reason!r} "
f"partial_message_id={event.partial_message_id}",
style=_AU_WARNING,
transcript.mount(Static(
RichText(
f"[cancelled] turn_id={event.turn_id} "
f"reason={event.reason!r} "
f"partial_message_id={event.partial_message_id}",
style=_AU_WARNING,
),
classes="cancelled-label",
))
transcript.scroll_end(animate=False)
return
if isinstance(event, WorkerPhase):
# v0.5.0: telemetry → Debug pane, not transcript.
@@ -360,22 +386,24 @@ class TuiPresenterState:
# the original event AND a render_error line with the class name only
# (NO exception message — security clause). Volva F1 fix.
#
# v0.6.0 routing-under-failure preservation — fallback writes go
# to the same destination the successful render would have used:
# - ToolStart/ToolResult → tools_log
# - Thinking → thinking_log
# - WorkerPhase/TextBoundary → debug_log
# - everything else → log
# v0.9.0 routing-under-failure: panes (RichLog) still write Strip
# lines; transcript (VerticalScroll) mounts a Static instead.
if isinstance(event, (ToolStart, ToolResult)):
target = tools_log
tools_log.write(_plain_label(event))
tools_log.write(f"[render_error] {type(exc).__name__}")
elif isinstance(event, Thinking):
target = thinking_log
thinking_log.write(_plain_label(event))
thinking_log.write(f"[render_error] {type(exc).__name__}")
elif isinstance(event, (WorkerPhase, TextBoundary)):
target = debug_log
debug_log.write(_plain_label(event))
debug_log.write(f"[render_error] {type(exc).__name__}")
else:
target = log
target.write(_plain_label(event))
target.write(f"[render_error] {type(exc).__name__}")
# Transcript-bound event (Text / Done / Error / Cancelled).
transcript.mount(Static(_plain_label(event), classes="error-label"))
transcript.mount(
Static(f"[render_error] {type(exc).__name__}", classes="error-label")
)
transcript.scroll_end(animate=False)
class AgentPickerApp(App[str | None]):
@@ -564,21 +592,43 @@ class RatatoskrApp(App[int]):
}
/* v0.6.5: thinking-current Static removed; thinking now streams
directly into thinking-log so the whole pane scrolls naturally. */
#transcript {
/* v0.9.0: transcript is a VerticalScroll container holding dynamically
mounted Statics + Markdown widgets per turn. Live Markdown rendering
replaces the v0.8.x RichLog approach which couldn't render Markdown
in-flight (only on Done as a re-render → double-print bug). */
#transcript-scroll {
height: 1fr;
background: $background;
padding: 0 1;
}
/* v0.6.0: streaming-text Static carries in-flight assistant tokens.
Replaces per-token RichLog spam — one growing line that updates in
place. Cleared on terminal event; final Markdown body lands in the
transcript. */
#current-text {
dock: bottom;
/* Per-turn mounted widgets carry id-prefix conventions:
- .turn-header "── turn N ──" (dim)
- .prompt-echo " user input" (aurora bright cyan)
- .response-md Markdown(accumulated_text) — updated live
- .done-label "[done] turn_id=…" (aurora green)
- .error-label "[error] …" (dawn red)
- .cancelled-label "[cancelled] …" (dawn yellow)
*/
.turn-header {
height: auto;
padding: 0 1;
color: $au-dark-60;
}
.prompt-echo {
height: auto;
background: $background;
padding: 0 1;
}
.response-md {
height: auto;
padding: 0 1;
}
.done-label, .error-label, .cancelled-label {
height: auto;
padding: 0 1;
}
/* v0.8.1: #current-text Static removed. Streaming text now coalesces
on `\n` and writes directly to #transcript (same pattern as v0.7.1
thinking fix). Eliminates the dock-bottom-growth-overlap bug. */
#tools-log, #debug-log, #thinking-log {
background: $background;
padding: 0 1;
@@ -677,8 +727,12 @@ class RatatoskrApp(App[int]):
# work without widget-level markup=True.
with Horizontal(id="main-row"):
with Vertical(id="left-column"):
yield RichLog(id="transcript", wrap=True, markup=False, highlight=False)
yield Static("", id="current-text")
# v0.9.0: transcript is a VerticalScroll holding per-turn
# mounted widgets (turn header, prompt echo, response Markdown,
# done label). Live Markdown rendering happens via Static
# widgets holding `Markdown` Renderables, updated as Text
# deltas arrive.
yield VerticalScroll(id="transcript-scroll")
yield Input(id="prompt", placeholder="Type a message and press Enter")
with Vertical(id="right-column"):
with TabbedContent(id="side-panes"):
@@ -741,20 +795,30 @@ class RatatoskrApp(App[int]):
self._set_hint(self.HINT_IDLE)
def _write_turn_headers(self, turn_id: int) -> None:
"""v0.6.0: Write `── turn N ──` Rule headers across every pane so
operators can visually correlate sections during cross-pane
debugging. Called from `_stream_turn_worker` on first event of
each new turn (idempotent per turn via active_turn_id guard).
"""v0.6.0: turn-ID headers across every pane for cross-pane
correlation. v0.9.0: transcript is a VerticalScroll; mounts a
Static with rule-style text instead of writing a Rule Renderable
to RichLog. Other panes still use RichLog.write(Rule).
"""
from rich.rule import Rule
from rich.text import Text as RichText
title = f"turn {turn_id}"
rule = Rule(title=title, style=_AU_DEMOTED)
try:
self.query_one("#transcript", RichLog).write(rule)
# Transcript (VerticalScroll): mount a styled Static.
transcript = self.query_one("#transcript-scroll", VerticalScroll)
transcript.mount(
Static(
RichText(f"── turn {turn_id} ──", style=_AU_DEMOTED),
classes="turn-header",
)
)
# Other panes (RichLog): write the Rule Renderable.
self.query_one("#tools-log", RichLog).write(rule)
self.query_one("#debug-log", RichLog).write(rule)
self.query_one("#thinking-log", RichLog).write(rule)
transcript.scroll_end(animate=False)
except Exception:
# Defensive: widget tree may be tearing down — never let a
# turn-header write block the SSE consumer.
@@ -770,12 +834,19 @@ class RatatoskrApp(App[int]):
pass
async def on_input_submitted(self, event: Input.Submitted) -> None:
"""Echo user prompt, spawn stream worker; busy notice if not idle."""
"""Echo user prompt, spawn stream worker; busy notice if not idle.
v0.9.0: prompt echo mounts as a Static in the transcript VerticalScroll
(was log.write to RichLog).
"""
if event.input.id != "prompt":
return
log = self.query_one("#transcript", RichLog)
transcript = self.query_one("#transcript-scroll", VerticalScroll)
if self.state != "idle":
log.write("[busy] turn in flight; input ignored")
transcript.mount(
Static("[busy] turn in flight; input ignored", classes="error-label")
)
transcript.scroll_end(animate=False)
event.input.value = ""
return
content = event.input.value.strip()
@@ -784,7 +855,13 @@ class RatatoskrApp(App[int]):
# v0.4.1 retheme: operator's voice gets Australis bright cyan so it
# stands out against the default-foreground assistant text below it.
from rich.text import Text as RichText
log.write(RichText(f" {content}", style=_AU_USER_ECHO)) # noqa: RUF001
transcript.mount(
Static(
RichText(f" {content}", style=_AU_USER_ECHO), # noqa: RUF001
classes="prompt-echo",
)
)
transcript.scroll_end(animate=False)
event.input.value = ""
self.state = "streaming"
self._set_hint(self.HINT_STREAMING)
@@ -793,28 +870,37 @@ class RatatoskrApp(App[int]):
)
async def _stream_turn_worker(self, content: str) -> None:
"""Drive stream_turn, render events via TuiPresenterState (issue #12)."""
"""Drive stream_turn, render events via TuiPresenterState.
v0.9.0: transcript is a VerticalScroll; the presenter's `transcript`
argument is the container, and the presenter mounts Static / Markdown-
backed widgets directly. Wire-error labels mount as `error-label`
Statics into the transcript-scroll.
"""
assert self.state == "streaming"
assert self.client is not None
assert content
log = self.query_one("#transcript", RichLog)
current_text = self.query_one("#current-text", Static)
transcript = self.query_one("#transcript-scroll", VerticalScroll)
tools_log = self.query_one("#tools-log", RichLog)
debug_log = self.query_one("#debug-log", RichLog)
thinking_log = self.query_one("#thinking-log", RichLog)
presenter = TuiPresenterState()
def _mount_wire_error(label: str) -> None:
try:
transcript.mount(Static(label, classes="error-label"))
transcript.scroll_end(animate=False)
except Exception:
pass
try:
async for event in stream_turn(self.client, self.session_id, content):
if self.active_turn_id is None:
self.active_turn_id = event.sse_id.turn_id
# v0.6.0: turn-ID headers across all panes so the
# operator can visually correlate sections during
# cross-pane debugging.
self._write_turn_headers(self.active_turn_id)
presenter.render(
event,
log=log,
current_text=current_text,
transcript=transcript,
tools_log=tools_log,
debug_log=debug_log,
thinking_log=thinking_log,
@@ -823,15 +909,15 @@ class RatatoskrApp(App[int]):
if isinstance(event, (Done, Error, Cancelled)):
break
except SseConnectFailed as exc:
log.write(f"[sse_connect_failed] status={exc.status} body={exc.body!r}")
_mount_wire_error(f"[sse_connect_failed] status={exc.status} body={exc.body!r}")
except SseConnectionDropped as exc:
log.write(f"[connection_dropped] last_seen={exc.last_seen_sse_id}")
_mount_wire_error(f"[connection_dropped] last_seen={exc.last_seen_sse_id}")
except MalformedSseId as exc:
log.write(f"[malformed_sse_id] raw={exc.raw!r}")
_mount_wire_error(f"[malformed_sse_id] raw={exc.raw!r}")
except MalformedSseData as exc:
log.write(f"[malformed_sse_data] raw={exc.raw!r}")
_mount_wire_error(f"[malformed_sse_data] raw={exc.raw!r}")
except TurnIdFlip as exc:
log.write(f"[turn_id_flip] expected={exc.established} got={exc.got}")
_mount_wire_error(f"[turn_id_flip] expected={exc.established} got={exc.got}")
finally:
self.state = "idle"
self.active_turn_id = None
@@ -854,9 +940,12 @@ class RatatoskrApp(App[int]):
return
self.state = "cancelling"
self._set_hint(self.HINT_CANCELLING)
log = self.query_one("#transcript", RichLog)
transcript = self.query_one("#transcript-scroll", VerticalScroll)
self.run_worker(
_cancel_via_sse(self.client, self.session_id, self.active_turn_id, log=log)
_cancel_via_sse(
self.client, self.session_id, self.active_turn_id,
transcript=transcript,
)
)
elif self.state == "cancelling":
if self.stream_worker is not None:
@@ -928,6 +1017,14 @@ async def _resolve_then_run(args: ParsedArgs) -> int:
# Issue #8: startup agent picker — fetch GET /agents and prompt when
# --new is passed without --agent. list_agents errors land on real
# stderr before any alt-screen opens (preserves issue #6 INV-001).
#
# v0.8.0: merge in local tier-3 agent index. Worldtree's GET /agents
# doesn't return consumer-defined agents (issue #15 smoke finding);
# ratatoskr keeps its own JSON-backed index of agents the operator
# defined via `python -m ratatoskr.tier3 define`. Merged here so the
# picker shows foundational + local-tier-3 in one list. Dedup by
# agent_id (remote wins on conflict, since a server-listed agent
# is the authoritative source).
chosen_agent_id: str | None = args.agent_id
if args.new and args.agent_id is None:
try:
@@ -940,6 +1037,23 @@ async def _resolve_then_run(args: ParsedArgs) -> int:
except (httpx.ConnectError, httpx.ReadTimeout, httpx.TransportError) as exc:
sys.stderr.write(f"[network_error] {type(exc).__name__}: {exc}\n")
return 21
# v0.8.0: append local tier-3 entries not already in the remote list.
from ratatoskr.local_agents import load_local_agents
remote_ids = {a.agent_id for a in agents}
for entry in load_local_agents():
if entry.agent_id in remote_ids:
continue
agents.append(AgentInfo(
agent_id=entry.agent_id,
name=entry.agent_name,
description=entry.description,
version=None,
capabilities=[],
supported_models=[],
persona_traits={},
ui_hints={},
))
if not agents:
sys.stderr.write("[no_agents] server returned empty agent list\n")
return 13
@@ -980,12 +1094,23 @@ async def _cancel_via_sse(
session_id: str,
turn_id: int,
*,
log: RichLog,
transcript: VerticalScroll,
) -> None:
"""Fire-and-forget cancel; never raises (mirrors cli._cancel_and_log; #3 INV-009)."""
"""Fire-and-forget cancel; never raises (mirrors cli._cancel_and_log; #3 INV-009).
v0.9.0: mounts a `[cancel_failed]` Static into the transcript-scroll
container on failure (was log.write to RichLog).
"""
assert client is not None
assert isinstance(turn_id, int) and turn_id > 0
try:
await cancel_turn(client, session_id, turn_id)
except (CancelFailed, CancelTurnNotFound, CancelAlreadyCompleted, httpx.RequestError) as exc:
log.write(f"[cancel_failed] {type(exc).__name__}: {exc}")
try:
transcript.mount(Static(
f"[cancel_failed] {type(exc).__name__}: {exc}",
classes="error-label",
))
transcript.scroll_end(animate=False)
except Exception:
pass
+187
View File
@@ -0,0 +1,187 @@
"""Tests for ratatoskr.local_agents.
Use ``$RATATOSKR_LOCAL_AGENTS`` env-var override + pytest tmp_path to
isolate from the operator's real ``~/.config/ratatoskr/local_agents.json``.
"""
from __future__ import annotations
import json
from pathlib import Path
import pytest
from ratatoskr.local_agents import (
LocalAgentEntry,
_local_agents_path,
add_local_agent,
load_local_agents,
make_description,
remove_local_agent,
update_local_agent,
)
@pytest.fixture
def local_path(tmp_path: Path, monkeypatch: pytest.MonkeyPatch) -> Path:
"""Point $RATATOSKR_LOCAL_AGENTS at a fresh tmp file for the test."""
path = tmp_path / "local_agents.json"
monkeypatch.setenv("RATATOSKR_LOCAL_AGENTS", str(path))
return path
def _entry(
agent_id: str = "ratatoskr:wizard",
agent_name: str = "wizard",
model: str = "qwen3.6-35-a3b",
description: str = "(tier 3) test agent",
defined_at: str = "2026-05-25T00:00:00+00:00",
) -> LocalAgentEntry:
return LocalAgentEntry(
agent_id=agent_id,
agent_name=agent_name,
model=model,
description=description,
defined_at=defined_at,
)
class TestPathResolution:
def test_env_override(self, monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.setenv("RATATOSKR_LOCAL_AGENTS", "/tmp/custom-agents.json")
assert _local_agents_path() == Path("/tmp/custom-agents.json")
def test_xdg_config_home(self, monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.delenv("RATATOSKR_LOCAL_AGENTS", raising=False)
monkeypatch.setenv("XDG_CONFIG_HOME", "/tmp/xdg-config")
assert (
_local_agents_path()
== Path("/tmp/xdg-config/ratatoskr/local_agents.json")
)
def test_default_home(self, monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.delenv("RATATOSKR_LOCAL_AGENTS", raising=False)
monkeypatch.delenv("XDG_CONFIG_HOME", raising=False)
path = _local_agents_path()
assert path == Path.home() / ".config" / "ratatoskr" / "local_agents.json"
class TestLoadEmpty:
def test_missing_file_returns_empty(self, local_path: Path) -> None:
assert not local_path.exists()
assert load_local_agents() == []
def test_corrupt_json_returns_empty(self, local_path: Path) -> None:
local_path.parent.mkdir(parents=True, exist_ok=True)
local_path.write_text("not json at all")
assert load_local_agents() == []
def test_wrong_schema_version_returns_empty(self, local_path: Path) -> None:
local_path.parent.mkdir(parents=True, exist_ok=True)
local_path.write_text(json.dumps({"version": 999, "agents": []}))
assert load_local_agents() == []
def test_missing_version_key_returns_empty(self, local_path: Path) -> None:
local_path.parent.mkdir(parents=True, exist_ok=True)
local_path.write_text(json.dumps({"agents": []}))
assert load_local_agents() == []
def test_malformed_row_skipped(self, local_path: Path) -> None:
local_path.parent.mkdir(parents=True, exist_ok=True)
local_path.write_text(
json.dumps(
{
"version": 1,
"agents": [
{"agent_id": "incomplete"}, # missing required fields
{
"agent_id": "ratatoskr:good",
"agent_name": "good",
"model": "m",
"description": "d",
"defined_at": "t",
},
],
}
)
)
entries = load_local_agents()
assert len(entries) == 1
assert entries[0].agent_id == "ratatoskr:good"
class TestAdd:
def test_add_one(self, local_path: Path) -> None:
add_local_agent(_entry())
entries = load_local_agents()
assert len(entries) == 1
assert entries[0].agent_id == "ratatoskr:wizard"
def test_add_two_different(self, local_path: Path) -> None:
add_local_agent(_entry(agent_id="ratatoskr:a", agent_name="a"))
add_local_agent(_entry(agent_id="ratatoskr:b", agent_name="b"))
ids = {e.agent_id for e in load_local_agents()}
assert ids == {"ratatoskr:a", "ratatoskr:b"}
def test_add_replaces_same_id(self, local_path: Path) -> None:
add_local_agent(_entry(model="old-model"))
add_local_agent(_entry(model="new-model"))
entries = load_local_agents()
assert len(entries) == 1
assert entries[0].model == "new-model"
def test_creates_parent_dirs(
self, tmp_path: Path, monkeypatch: pytest.MonkeyPatch
) -> None:
nested = tmp_path / "deep" / "nested" / "path" / "agents.json"
monkeypatch.setenv("RATATOSKR_LOCAL_AGENTS", str(nested))
add_local_agent(_entry())
assert nested.exists()
class TestUpdate:
def test_update_changes_existing(self, local_path: Path) -> None:
add_local_agent(_entry(model="v1"))
update_local_agent(_entry(model="v2"))
entries = load_local_agents()
assert len(entries) == 1
assert entries[0].model == "v2"
class TestRemove:
def test_remove_existing(self, local_path: Path) -> None:
add_local_agent(_entry())
remove_local_agent("ratatoskr:wizard")
assert load_local_agents() == []
def test_remove_missing_is_noop(self, local_path: Path) -> None:
add_local_agent(_entry())
remove_local_agent("ratatoskr:doesnotexist")
assert len(load_local_agents()) == 1
class TestMakeDescription:
def test_first_nonempty_line(self) -> None:
prompt = "\n\n# IDENTITY\nYou are a test agent..."
desc = make_description(prompt)
assert desc.startswith("(tier 3) IDENTITY")
def test_strips_heading_markers(self) -> None:
prompt = "# A nice heading\nMore prompt..."
desc = make_description(prompt)
assert "(tier 3) A nice heading" == desc
def test_truncates_long(self) -> None:
prompt = "x" * 200
desc = make_description(prompt)
# 80 char cap including the prefix
assert len(desc) == 81 # 80 + ellipsis char
assert desc.endswith("")
def test_empty_prompt_fallback(self) -> None:
desc = make_description("")
assert desc == "(tier 3) custom system prompt"
def test_whitespace_only_fallback(self) -> None:
desc = make_description(" \n\n ")
assert desc == "(tier 3) custom system prompt"
+41
View File
@@ -711,6 +711,47 @@ def _sse_raw_chunk(sse_id: str, raw_data: str) -> bytes:
return f"id: {sse_id}\ndata: {raw_data}\n\n".encode()
def _sse_no_id_chunk(data: str) -> bytes:
"""SSE frame with NO id line + arbitrary data (v0.8.1: keepalive shape)."""
return f"data: {data}\n\n".encode()
class TestEmptyIdSkipped:
@respx.mock
async def test_empty_id_on_first_event_skipped(self) -> None:
"""empty_id_on_first_event_skipped [v0.8.1]: stream starts with an
event carrying NO `id:` line → httpx_sse exposes sse.id == ''
(no prior id to inherit). Pre-v0.8.1: MalformedSseId raw='' crashed
the turn. v0.8.1: treat same as empty-data keepalive — skip silently.
Observed 2026-05-25 on Worldtree's qwen3.6-35-a3b-heretic provider:
the first stream frame had no id line, every turn died with
`[malformed_sse_id] raw=''`.
"""
from ratatoskr.sse_client import Done as _Done
from ratatoskr.sse_client import Text as _Text
# First frame: no id line (httpx_sse → sse.id = ""). Skip it.
# Subsequent frames have ids; normal processing resumes.
stream = (
_sse_no_id_chunk('{"type":"keepalive"}') # ← skipped (sse.id == "")
+ _sse_chunk("42:1", {"type": "text", "content": "first"})
+ _sse_chunk("42:2", _DONE_42_6)
)
respx.post("https://w.example/sessions/s1/messages").mock(
return_value=httpx.Response(
200, headers={"content-type": "text/event-stream"}, content=stream
)
)
async with httpx.AsyncClient(base_url="https://w.example") as client:
events = [e async for e in stream_turn(client, "s1", "hi")]
# 2 events — the no-id frame is invisible (no MalformedSseId crash).
assert len(events) == 2
assert isinstance(events[0], _Text)
assert events[0].content == "first"
assert isinstance(events[1], _Done)
class TestEmptyDataSkipped:
@respx.mock
async def test_empty_data_skipped(self) -> None:
+60 -6
View File
@@ -1,5 +1,7 @@
"""Tests for ratatoskr.tier3 per docs/contracts/issues/15.contract.md."""
from pathlib import Path
import httpx
import pytest
import respx
@@ -315,12 +317,29 @@ class TestDeleteAgent:
assert exc.value.status == 500
@pytest.fixture
def _isolated_local_agents(
tmp_path: "Path", monkeypatch: pytest.MonkeyPatch
) -> "Path":
"""Isolate the v0.8.0 local-tier-3 index from the operator's real file."""
path = tmp_path / "local_agents.json"
monkeypatch.setenv("RATATOSKR_LOCAL_AGENTS", str(path))
return path
class TestCli:
@respx.mock
def test_cli_define_happy(
self, capsys: pytest.CaptureFixture[str], monkeypatch: pytest.MonkeyPatch
self,
capsys: pytest.CaptureFixture[str],
monkeypatch: pytest.MonkeyPatch,
_isolated_local_agents: "Path",
) -> None:
"""cli_define_happy [happy]: argv → 201 mock → stdout confirmation."""
"""cli_define_happy [happy]: argv → 201 mock → stdout confirmation;
local index updated with the new entry (v0.8.0 hook).
"""
from ratatoskr.local_agents import load_local_agents
monkeypatch.setenv("WORLDTREE_API_URL", "https://w.example")
monkeypatch.setenv("WORLDTREE_API_KEY", "k")
respx.post("https://w.example/agents/define").mock(
@@ -335,12 +354,24 @@ class TestCli:
out = capsys.readouterr()
assert rc == 0
assert out.out.strip() == "defined ratatoskr:wizard (qwen3.6-35-a3b)"
# v0.8.0: local index now has the new entry.
entries = load_local_agents()
assert len(entries) == 1
assert entries[0].agent_id == "ratatoskr:wizard"
assert entries[0].model == "qwen3.6-35-a3b"
@respx.mock
def test_cli_patch_happy(
self, capsys: pytest.CaptureFixture[str], monkeypatch: pytest.MonkeyPatch
self,
capsys: pytest.CaptureFixture[str],
monkeypatch: pytest.MonkeyPatch,
_isolated_local_agents: "Path",
) -> None:
"""cli_patch_happy [happy]: argv → 200 mock → stdout confirmation."""
"""cli_patch_happy [happy]: argv → 200 mock → stdout confirmation;
local index refreshed with the post-patch state.
"""
from ratatoskr.local_agents import load_local_agents
monkeypatch.setenv("WORLDTREE_API_URL", "https://w.example")
monkeypatch.setenv("WORLDTREE_API_KEY", "k")
respx.patch("https://w.example/agents/ratatoskr:wizard").mock(
@@ -350,12 +381,34 @@ class TestCli:
out = capsys.readouterr()
assert rc == 0
assert out.out.strip() == "patched ratatoskr:wizard"
entries = load_local_agents()
assert len(entries) == 1
assert entries[0].agent_id == "ratatoskr:wizard"
@respx.mock
def test_cli_delete_happy(
self, capsys: pytest.CaptureFixture[str], monkeypatch: pytest.MonkeyPatch
self,
capsys: pytest.CaptureFixture[str],
monkeypatch: pytest.MonkeyPatch,
_isolated_local_agents: "Path",
) -> None:
"""cli_delete_happy [happy]: argv → 204 mock → stdout confirmation."""
"""cli_delete_happy [happy]: argv → 204 mock → stdout confirmation;
local index entry removed (v0.8.0 hook).
"""
from ratatoskr.local_agents import (
LocalAgentEntry,
add_local_agent,
load_local_agents,
)
# Pre-populate so we can verify removal.
add_local_agent(LocalAgentEntry(
agent_id="ratatoskr:wizard",
agent_name="wizard",
model="m",
description="d",
defined_at="t",
))
monkeypatch.setenv("WORLDTREE_API_URL", "https://w.example")
monkeypatch.setenv("WORLDTREE_API_KEY", "k")
respx.delete("https://w.example/agents/ratatoskr:wizard").mock(
@@ -365,6 +418,7 @@ class TestCli:
out = capsys.readouterr()
assert rc == 0
assert out.out.strip() == "deleted ratatoskr:wizard"
assert load_local_agents() == []
def test_cli_missing_auth(
self, capsys: pytest.CaptureFixture[str], monkeypatch: pytest.MonkeyPatch
+330 -183
View File
@@ -1,5 +1,6 @@
"""Tests for ratatoskr.tui per docs/contracts/issues/4.contract.md."""
from pathlib import Path
from unittest.mock import MagicMock
import httpx
@@ -60,20 +61,43 @@ def _args_existing(session_id: str = "s-1existing", **overrides) -> ParsedArgs:
def _spy_writes(monkeypatch) -> list:
"""Patch RichLog.write to record every arg into a list (returned).
"""Patch RichLog.write AND VerticalScroll.mount to record every renderable
or mounted-widget content into a single list (returned).
Accepts *args/**kwargs so Textual's internal deferred-render path
(which calls write positionally with width/expand/shrink/scroll_end)
still works after a write-during-mount + Resize sequence.
v0.9.0: transcript content is mounted into a VerticalScroll, not written
to a RichLog. The spy captures both shapes — for each mounted Static, the
Static's `renderable` (Markdown / RichText / str) lands in the list,
indistinguishably from RichLog.write entries. Integration tests assert
on substrings or types in `writes` so the merged shape is the right
abstraction.
Accepts *args/**kwargs so Textual's internal deferred-render paths still
work after a write-during-mount + Resize sequence.
"""
from textual.containers import VerticalScroll
from textual.widgets import Static
writes: list = []
original = RichLog.write
def spy(self, content, *args, **kw):
original_write = RichLog.write
def spy_write(self, content, *args, **kw):
writes.append(content)
return original(self, content, *args, **kw)
return original_write(self, content, *args, **kw)
monkeypatch.setattr(RichLog, "write", spy)
monkeypatch.setattr(RichLog, "write", spy_write)
original_mount = VerticalScroll.mount
def spy_mount(self, *children, **kw):
for child in children:
if isinstance(child, Static):
writes.append(child.content)
else:
writes.append(child)
return original_mount(self, *children, **kw)
monkeypatch.setattr(VerticalScroll, "mount", spy_mount)
return writes
@@ -125,16 +149,15 @@ class TestTuiPresenterState:
from ratatoskr.tui import TuiPresenterState
log = MagicMock()
transcript = MagicMock()
thinking_log = MagicMock()
state = TuiPresenterState()
for chunk in ("Let", " me", " think"):
state.render(
Thinking(sse_id=SID, content=chunk),
log=log,
transcript=transcript,
tools_log=MagicMock(),
debug_log=MagicMock(),
current_text=MagicMock(),
thinking_log=thinking_log,
raw=False,
)
@@ -143,7 +166,7 @@ class TestTuiPresenterState:
assert len(writes) == 1
assert isinstance(writes[0], Rule)
assert state.thinking_chunk_buffer == "Let me think"
assert log.write.call_count == 0
assert transcript.mount.call_count == 0
def test_thinking_flushes_on_newline(self) -> None:
"""thinking_flushes_on_newline [happy, v0.7.1]:
@@ -156,10 +179,9 @@ class TestTuiPresenterState:
for chunk in ("Hello", " world", "\n"):
state.render(
Thinking(sse_id=SID, content=chunk),
log=MagicMock(),
transcript=MagicMock(),
tools_log=MagicMock(),
debug_log=MagicMock(),
current_text=MagicMock(),
thinking_log=thinking_log,
raw=False,
)
@@ -179,26 +201,24 @@ class TestTuiPresenterState:
from ratatoskr.tui import TuiPresenterState
log = MagicMock()
transcript = MagicMock()
debug_log = MagicMock()
thinking_log = MagicMock()
state = TuiPresenterState()
for content in ("a", "b"):
state.render(
Thinking(sse_id=SID, content=content),
log=log,
transcript=transcript,
tools_log=MagicMock(),
debug_log=debug_log,
current_text=MagicMock(),
thinking_log=thinking_log,
raw=False,
)
state.render(
WorkerPhase(sse_id=SID, phase="streaming", turn_id=42),
log=log,
transcript=transcript,
tools_log=MagicMock(),
debug_log=debug_log,
current_text=MagicMock(),
thinking_log=thinking_log,
raw=False,
)
@@ -210,24 +230,25 @@ class TestTuiPresenterState:
assert isinstance(thinking_writes[2], Rule)
# worker_phase still goes to debug_log; transcript untouched.
assert "· worker_phase:" in _text_of(debug_log.write.call_args_list[-1][0][0])
assert not log.write.called
assert not transcript.mount.called
# v0.6.5: thinking-current Static removed; test_thinking_widget_truncation
# and test_thinking_widget_visibility_lifecycle deleted (no longer apply).
def test_multiple_thinking_runs_each_get_thinking_log_section(self) -> None:
"""multiple_thinking_runs_each_get_section [scenario, v0.6.5]:
"""multiple_thinking_runs_each_get_section [scenario, v0.8.1]:
Thinking → Text → Thinking → Done → TWO start/end Rule pairs in
thinking_log, each wrapping their delta lines. Text goes to
current_text (buffered). Transcript: [done] + Markdown body.
thinking_log (deltas coalesced into tail-flushes per run).
Text deltas now stream into the transcript via coalesce-on-newline
(no current-text Static); "hi" with no `\\n` stays buffered until
Done's tail-flush.
"""
from rich.rule import Rule
from ratatoskr.tui import TuiPresenterState
log = MagicMock()
transcript = MagicMock()
thinking_log = MagicMock()
current_text = MagicMock()
state = TuiPresenterState()
for evt in (
Thinking(sse_id=SID, content="first"),
@@ -235,28 +256,27 @@ class TestTuiPresenterState:
Thinking(sse_id=SID, content="second"),
):
state.render(
evt, log=log,
evt, transcript=transcript,
tools_log=MagicMock(), debug_log=MagicMock(),
current_text=current_text, thinking_log=thinking_log, raw=False,
thinking_log=thinking_log, raw=False,
)
state.render(
_make_tui_done(),
log=log,
transcript=transcript,
tools_log=MagicMock(), debug_log=MagicMock(),
current_text=current_text, thinking_log=thinking_log, raw=False,
thinking_log=thinking_log, raw=False,
)
# v0.6.5: thinking_log holds 4 Rules (start + end per run) + 2 delta lines.
# thinking_log: 4 Rules (start+end per run) + 2 tail-flush strings.
thinking_writes = [c[0][0] for c in thinking_log.write.call_args_list]
rules = [w for w in thinking_writes if isinstance(w, Rule)]
delta_strs = [w for w in thinking_writes if isinstance(w, str)]
assert len(rules) == 4, f"expected 4 Rules (2 start + 2 end), got {len(rules)}"
assert "first" in delta_strs
assert "second" in delta_strs
# Text "hi" went to current_text (buffered), not the transcript directly.
current_text.update.assert_any_call("hi")
# Transcript: [done] label + Markdown(response) (raw=False).
log_writes = [_text_of(c[0][0]) for c in log.write.call_args_list]
assert any(w.startswith("[done]") for w in log_writes if isinstance(w, str))
# v0.8.1: Text "hi" flushes as a line in transcript on Done.
transcript_renderables = [_text_of(r) for r in _mounted_renderables(transcript)]
assert "hi" in transcript_renderables
assert any(w.startswith("[done]") for w in transcript_renderables if isinstance(w, str))
def test_render_exception_fallback(self) -> None:
"""render_exception_fallback [adversarial, v0.6.5]:
@@ -266,7 +286,7 @@ class TestTuiPresenterState:
"""
from ratatoskr.tui import TuiPresenterState
log = MagicMock()
transcript = MagicMock()
thinking_log = MagicMock()
# First call (Rule write) raises; subsequent calls succeed for fallback.
thinking_log.write.side_effect = [
@@ -277,10 +297,9 @@ class TestTuiPresenterState:
state = TuiPresenterState()
state.render(
Thinking(sse_id=SID, content="x"),
log=log,
transcript=transcript,
tools_log=MagicMock(),
debug_log=MagicMock(),
current_text=MagicMock(),
thinking_log=thinking_log,
raw=False,
)
@@ -288,7 +307,7 @@ class TestTuiPresenterState:
assert any(w.startswith("[thinking]") for w in writes), writes
assert any(w == "[render_error] AttributeError" for w in writes), writes
assert not any("rule write failed" in w for w in writes), writes
assert not log.write.called
assert not transcript.mount.called
def test_state_reset_per_worker(self) -> None:
"""state_reset_per_worker [trace]: fresh TuiPresenterState() starts no thinking open."""
@@ -297,10 +316,9 @@ class TestTuiPresenterState:
s1 = TuiPresenterState()
s1.render(
Thinking(sse_id=SID, content="x"),
log=MagicMock(),
transcript=MagicMock(),
tools_log=MagicMock(),
debug_log=MagicMock(),
current_text=MagicMock(),
thinking_log=MagicMock(),
raw=False,
)
@@ -315,98 +333,101 @@ class TestTuiPresenterState:
"""
from ratatoskr.tui import TuiPresenterState
log = MagicMock()
transcript = MagicMock()
thinking_log = MagicMock()
state = TuiPresenterState()
state.render(
Thinking(sse_id=SID, content="partial"),
log=log,
transcript=transcript,
tools_log=MagicMock(),
debug_log=MagicMock(),
current_text=MagicMock(), thinking_log=thinking_log, raw=False,
debug_log=MagicMock(), thinking_log=thinking_log, raw=False,
)
state.render(
Cancelled(
sse_id=SID, phase="cancelled", turn_id=42, reason="user", partial_message_id=None
),
log=log,
transcript=transcript,
tools_log=MagicMock(),
debug_log=MagicMock(),
current_text=MagicMock(), thinking_log=thinking_log, raw=False,
debug_log=MagicMock(), thinking_log=thinking_log, raw=False,
)
# v0.6.5: streamed thinking + Rule(end) in thinking_log; [cancelled] in transcript.
log_writes = [_text_of(c[0][0]) for c in log.write.call_args_list]
assert any(w.startswith("[cancelled]") for w in log_writes)
transcript_renderables = [_text_of(r) for r in _mounted_renderables(transcript)]
assert any(w.startswith("[cancelled]") for w in transcript_renderables)
# thinking_log got at least Rule(start) + "partial" delta + Rule(end)
assert thinking_log.write.call_count >= 3
def test_done_renders_markdown_after_label(self) -> None:
"""done_renders_markdown_after_label [happy, v0.6.0]:
Text("hi") accumulates into current_text Static (buffered streaming);
Done(response="hi") with raw=False → [done] label + Rule + Markdown
in transcript. current_text cleared on terminal.
def test_text_then_done_mounts_widget_and_finalizes(self) -> None:
"""text_then_done_mounts_widget_and_finalizes [happy, v0.9.0]:
First Text delta mounts a Static(Markdown(buffer)) into the transcript;
Done finalizes the widget reference and mounts a styled [done] label.
No duplicate content (v0.9.0 replaces v0.8.x's flush-on-Done with
live in-place Markdown updates).
"""
from rich.markdown import Markdown
from rich.rule import Rule
from ratatoskr.tui import TuiPresenterState
log = MagicMock()
current_text = MagicMock()
transcript = MagicMock()
state = TuiPresenterState()
state.render(
Text(sse_id=SID, content="hi"),
log=log,
transcript=transcript,
tools_log=MagicMock(),
debug_log=MagicMock(),
current_text=current_text,
thinking_log=MagicMock(),
raw=False,
)
# Text accumulated to current_text, NOT written to log.
current_text.update.assert_any_call("hi")
# v0.9.0: response widget mounted on first Text delta with Markdown wrapper.
assert transcript.mount.called
first_widget = transcript.mount.call_args_list[0][0][0]
assert isinstance(first_widget.content, Markdown)
assert first_widget.content.markup == "hi"
assert state.text_chunk_buffer == "hi"
# Done finalizes: text_chunk_buffer cleared, widget ref released, label mounted.
state.render(
_make_tui_done(),
log=log,
transcript=transcript,
tools_log=MagicMock(),
debug_log=MagicMock(),
current_text=current_text,
thinking_log=MagicMock(),
raw=False,
)
# Done cleared current_text and wrote [done] label + Rule + Markdown.
current_text.update.assert_any_call("")
writes = [c[0][0] for c in log.write.call_args_list]
writes = _mounted_renderables(transcript)
assert any(_text_of(w).startswith("[done]") for w in writes)
assert any(isinstance(w, Rule) for w in writes)
assert any(isinstance(w, Markdown) for w in writes)
# v0.9.0: response Markdown rendered live during stream — only ONE
# Markdown renderable lands in the transcript (no post-Done re-render).
markdowns = [w for w in writes if isinstance(w, Markdown)]
assert len(markdowns) == 1
assert state.text_chunk_buffer == ""
assert state.current_response_widget is None
def test_raw_flag_skips_markdown(self) -> None:
"""raw_flag_skips_markdown [trace]: raw=True → no Rule, no Markdown."""
"""raw_flag_skips_markdown [v0.9.0]: raw=True → response widget holds
plain str instead of Markdown. Live in-place update still happens;
only the wrapper differs.
"""
from rich.markdown import Markdown
from rich.rule import Rule
from ratatoskr.tui import TuiPresenterState
log = MagicMock()
transcript = MagicMock()
state = TuiPresenterState()
state.render(
Text(sse_id=SID, content="hi"),
log=log,
transcript=transcript,
tools_log=MagicMock(),
debug_log=MagicMock(),
current_text=MagicMock(), thinking_log=MagicMock(), raw=True,
debug_log=MagicMock(), thinking_log=MagicMock(), raw=True,
)
state.render(
_make_tui_done(),
log=log,
transcript=transcript,
tools_log=MagicMock(),
debug_log=MagicMock(),
current_text=MagicMock(), thinking_log=MagicMock(), raw=True,
debug_log=MagicMock(), thinking_log=MagicMock(), raw=True,
)
writes = [c[0][0] for c in log.write.call_args_list]
assert not any(isinstance(w, Rule) for w in writes)
writes = _mounted_renderables(transcript)
# Raw mode bypasses Markdown entirely — content lives as plain str.
assert not any(isinstance(w, Markdown) for w in writes)
assert "hi" in writes
def test_worker_phase_demoted_to_debug_log(self) -> None:
"""worker_phase_demoted_to_debug_log [trace, v0.5.0]: WorkerPhase → debug_log
@@ -417,18 +438,17 @@ class TestTuiPresenterState:
from ratatoskr.tui import TuiPresenterState
log = MagicMock()
transcript = MagicMock()
debug_log = MagicMock()
state = TuiPresenterState()
state.render(
WorkerPhase(sse_id=SID, phase="streaming", turn_id=42),
log=log,
transcript=transcript,
tools_log=MagicMock(),
debug_log=debug_log,
current_text=MagicMock(), thinking_log=MagicMock(), raw=False,
debug_log=debug_log, thinking_log=MagicMock(), raw=False,
)
# v0.5.0: WorkerPhase routes to debug_log, NOT transcript.
assert not log.write.called
assert not transcript.mount.called
renderable = debug_log.write.call_args[0][0]
# INV-003: must be a styled Rich Text renderable, not a plain str.
# v0.4.1 retheme: style is now Australis Sea dark-60 ("#86929d") instead
@@ -453,101 +473,108 @@ class TestTuiPresenterState:
"""
from ratatoskr.tui import TuiPresenterState
log = MagicMock()
transcript = MagicMock()
tools_log = MagicMock()
state = TuiPresenterState()
state.render(
ToolStart(sse_id=SID, name="read_file", arguments={"path": "/x"}),
log=log,
transcript=transcript,
tools_log=tools_log,
debug_log=MagicMock(),
current_text=MagicMock(), thinking_log=MagicMock(), raw=False,
debug_log=MagicMock(), thinking_log=MagicMock(), raw=False,
)
# INV-014: write went to tools_log
assert tools_log.write.called
assert _text_of(tools_log.write.call_args[0][0]).startswith("· tool_start:")
# INV-014: transcript was NOT written to
assert not log.write.called
assert not transcript.mount.called
def test_tool_result_routes_to_tools_log(self) -> None:
"""tool_result_routes_to_tools_log [INV-014]: ToolResult → tools_log, NOT transcript."""
from ratatoskr.tui import TuiPresenterState
log = MagicMock()
transcript = MagicMock()
tools_log = MagicMock()
state = TuiPresenterState()
state.render(
ToolResult(sse_id=SID, name="read_file", result="ok", duration_ms=12),
log=log,
transcript=transcript,
tools_log=tools_log,
debug_log=MagicMock(),
current_text=MagicMock(), thinking_log=MagicMock(), raw=False,
debug_log=MagicMock(), thinking_log=MagicMock(), raw=False,
)
assert tools_log.write.called
assert _text_of(tools_log.write.call_args[0][0]).startswith("· tool_result:")
assert not log.write.called
assert not transcript.mount.called
def test_text_event_buffers_into_current_text(self) -> None:
"""text_event_buffers_into_current_text [v0.6.0]: Text → current_text Static
(accumulated), NOT log or tools_log. Streaming UX fix — no per-token spam.
def test_text_first_delta_mounts_response_widget(self) -> None:
"""text_first_delta_mounts_response_widget [v0.9.0]: first Text delta
mounts a Static carrying Markdown(buffer) into the transcript. The
text_chunk_buffer holds the accumulated content for the next delta's
in-place update.
"""
from rich.markdown import Markdown
from ratatoskr.tui import TuiPresenterState
log = MagicMock()
transcript = MagicMock()
tools_log = MagicMock()
current_text = MagicMock()
state = TuiPresenterState()
state.render(
Text(sse_id=SID, content="hello"),
log=log,
transcript=transcript,
tools_log=tools_log,
debug_log=MagicMock(),
current_text=current_text,
thinking_log=MagicMock(),
raw=False,
)
current_text.update.assert_called_once_with("hello")
assert not log.write.called
assert state.text_chunk_buffer == "hello"
assert transcript.mount.call_count == 1
widget = transcript.mount.call_args[0][0]
assert isinstance(widget.content, Markdown)
assert widget.content.markup == "hello"
assert state.current_response_widget is widget
assert not tools_log.write.called
def test_text_deltas_accumulate(self) -> None:
"""text_deltas_accumulate [v0.6.0]: multiple Text deltas → current_text shows
concatenated content, NOT separate per-delta lines.
def test_text_subsequent_deltas_update_in_place(self) -> None:
"""text_subsequent_deltas_update_in_place [v0.9.0]: deltas after the
first do NOT mount a new widget — they update the existing widget's
Markdown content in place. The text_chunk_buffer accumulates.
"""
from ratatoskr.tui import TuiPresenterState
current_text = MagicMock()
transcript = MagicMock()
state = TuiPresenterState()
for tok in ("Hel", "lo", " ", "world"):
state.render(
Text(sse_id=SID, content=tok),
log=MagicMock(),
transcript=transcript,
tools_log=MagicMock(),
debug_log=MagicMock(),
current_text=current_text,
thinking_log=MagicMock(),
raw=False,
)
# Final update reflects the full concatenation.
assert current_text.update.call_args_list[-1][0][0] == "Hello world"
# Exactly ONE mount (the first delta); subsequent deltas update.
assert transcript.mount.call_count == 1
assert state.text_chunk_buffer == "Hello world"
# Widget reference held; buffer is the source of truth re-rendered
# into Markdown(...) for each Static.update call.
assert state.current_response_widget is not None
def test_duration_format_seconds(self) -> None:
"""duration_format_seconds [trace]: Done(duration_ms=5467) → label has "duration=5.5s"."""
from ratatoskr.tui import TuiPresenterState
log = MagicMock()
transcript = MagicMock()
state = TuiPresenterState()
state.render(
_make_tui_done(duration_ms=5467),
log=log,
transcript=transcript,
tools_log=MagicMock(),
debug_log=MagicMock(),
current_text=MagicMock(), thinking_log=MagicMock(), raw=True,
debug_log=MagicMock(), thinking_log=MagicMock(), raw=True,
)
done_line = next(
_text_of(c[0][0])
for c in log.write.call_args_list
if _text_of(c[0][0]).startswith("[done]")
_text_of(r)
for r in _mounted_renderables(transcript)
if _text_of(r).startswith("[done]")
)
assert "duration=5.5s" in done_line
assert "duration_ms=5467" not in done_line
@@ -556,7 +583,7 @@ class TestTuiPresenterState:
"""usage_format_unicode_arrow [trace]: TUI Done label uses → (Unicode), not -> (ASCII)."""
from ratatoskr.tui import TuiPresenterState
log = MagicMock()
transcript = MagicMock()
state = TuiPresenterState()
usage = {
"prompt_tokens": 6756,
@@ -566,15 +593,14 @@ class TestTuiPresenterState:
}
state.render(
_make_tui_done(usage=usage),
log=log,
transcript=transcript,
tools_log=MagicMock(),
debug_log=MagicMock(),
current_text=MagicMock(), thinking_log=MagicMock(), raw=True,
debug_log=MagicMock(), thinking_log=MagicMock(), raw=True,
)
done_line = next(
_text_of(c[0][0])
for c in log.write.call_args_list
if _text_of(c[0][0]).startswith("[done]")
_text_of(r)
for r in _mounted_renderables(transcript)
if _text_of(r).startswith("[done]")
)
assert "usage 6756 in → 126 out (6882 total, 0 cached)" in done_line
@@ -585,14 +611,38 @@ def _text_of(write_arg: object) -> str:
Issue #12 wraps demoted-telemetry entries in `rich.text.Text(..., style="dim")`
so the RichLog can apply dim styling; non-demoted writes stay as plain str.
Tests that want to assert against content need both shapes flattened.
v0.9.0: also extracts plain text from Markdown wrappers (the streaming-text
response path uses Markdown(buffer) now; tests assert against the source
markup, which lives in `Markdown.markup`).
"""
from rich.markdown import Markdown
from rich.text import Text as RichText
if isinstance(write_arg, RichText):
return write_arg.plain
if isinstance(write_arg, Markdown):
return write_arg.markup
if isinstance(write_arg, str):
return write_arg
return "" # Markdown / Rule / etc. — not text content
return "" # Rule / etc. — not text content
def _mounted_renderables(transcript_mock: MagicMock) -> list:
"""v0.9.0: TuiPresenterState now mounts Static widgets into the transcript
VerticalScroll instead of writing renderables to a RichLog. Tests using a
MagicMock transcript inspect `transcript.mount.call_args_list`; each call's
first positional arg is the Static child whose `.content` carries the
Markdown / RichText / str that pre-v0.9.0 would have been the write arg.
Returns those renderables in mount-call order so tests can assert on them
with the same shape they used for `log.write.call_args_list` previously.
"""
out: list = []
for call in transcript_mock.mount.call_args_list:
for child in call.args:
renderable = getattr(child, "content", child)
out.append(renderable)
return out
def _make_tui_done(*, duration_ms: int = 1, usage: dict[str, int] | None = None) -> Done:
@@ -616,15 +666,15 @@ def _make_tui_done(*, duration_ms: int = 1, usage: dict[str, int] | None = None)
class TestCancelViaSse:
@respx.mock
async def test_happy_cancel(self) -> None:
"""happy_cancel [happy,tracer]: 200 OK → returns None; log has no [cancel_failed]."""
"""happy_cancel [happy,tracer]: 200 OK → returns None; transcript has no [cancel_failed]."""
respx.post("https://w.example/sessions/s-1/turns/42/cancel").mock(
return_value=httpx.Response(200, json=_CANCEL_OK_RESP)
)
log = MagicMock()
transcript = MagicMock()
async with httpx.AsyncClient(base_url="https://w.example") as client:
result = await _cancel_via_sse(client, "s-1", 42, log=log)
result = await _cancel_via_sse(client, "s-1", 42, transcript=transcript)
assert result is None
log.write.assert_not_called()
transcript.mount.assert_not_called()
@respx.mock
async def test_cancel_failed_500(self) -> None:
@@ -632,10 +682,10 @@ class TestCancelViaSse:
respx.post("https://w.example/sessions/s-1/turns/42/cancel").mock(
return_value=httpx.Response(500, content=b"boom")
)
log = MagicMock()
transcript = MagicMock()
async with httpx.AsyncClient(base_url="https://w.example") as client:
await _cancel_via_sse(client, "s-1", 42, log=log)
line = log.write.call_args[0][0]
await _cancel_via_sse(client, "s-1", 42, transcript=transcript)
line = transcript.mount.call_args[0][0].content
assert "[cancel_failed]" in line
assert "CancelFailed" in line
@@ -645,10 +695,10 @@ class TestCancelViaSse:
respx.post("https://w.example/sessions/s-1/turns/42/cancel").mock(
return_value=httpx.Response(409)
)
log = MagicMock()
transcript = MagicMock()
async with httpx.AsyncClient(base_url="https://w.example") as client:
await _cancel_via_sse(client, "s-1", 42, log=log)
line = log.write.call_args[0][0]
await _cancel_via_sse(client, "s-1", 42, transcript=transcript)
line = transcript.mount.call_args[0][0].content
assert "[cancel_failed]" in line
assert "CancelAlreadyCompleted" in line
@@ -658,10 +708,10 @@ class TestCancelViaSse:
respx.post("https://w.example/sessions/s-1/turns/42/cancel").mock(
side_effect=httpx.ConnectError("network down")
)
log = MagicMock()
transcript = MagicMock()
async with httpx.AsyncClient(base_url="https://w.example") as client:
await _cancel_via_sse(client, "s-1", 42, log=log)
line = log.write.call_args[0][0]
await _cancel_via_sse(client, "s-1", 42, transcript=transcript)
line = transcript.mount.call_args[0][0].content
assert "[cancel_failed]" in line
assert "ConnectError" in line
@@ -728,18 +778,18 @@ class TestLayoutShape:
assert row is not None
async def test_left_column_content_only(self) -> None:
"""left_column_content_only [v0.6.5]: left column = transcript + prompt
+ current-text (streaming text Static). thinking-current Static
removed entirely as of v0.6.5.
"""left_column_content_only [v0.9.0]: left column = transcript-scroll
VerticalScroll + prompt Input. thinking-current Static removed in
v0.6.5; transcript RichLog replaced by VerticalScroll in v0.9.0.
"""
from textual.containers import Vertical
from textual.widgets import Input, RichLog
from textual.containers import Vertical, VerticalScroll
from textual.widgets import Input
app = _resolved_app(_args_new(), session_id="s-new12345", agent_id="mimir")
async with app.run_test() as pilot:
await pilot.pause()
left = app.query_one("#left-column", Vertical)
transcript = app.query_one("#transcript", RichLog)
transcript = app.query_one("#transcript-scroll", VerticalScroll)
prompt = app.query_one("#prompt", Input)
assert transcript in left.walk_children()
assert prompt in left.walk_children()
@@ -764,7 +814,7 @@ class TestLayoutShape:
assert tools_tab is not None
async def test_tools_log_inside_tools_tab(self) -> None:
"""tools_log_inside_tools_tab: tools-log RichLog is a descendant of tools-tab TabPane."""
"""tools_log_inside_tools_tab: tools-transcript RichLog is a descendant of tools-tab TabPane."""
from textual.widgets import RichLog, TabPane
app = _resolved_app(_args_new(), session_id="s-new12345", agent_id="mimir")
@@ -815,7 +865,7 @@ class TestLayoutShape:
)
async def test_debug_tab_exists(self) -> None:
"""debug_tab_exists [v0.5.0]: right column has Debug TabPane + #debug-log RichLog."""
"""debug_tab_exists [v0.5.0]: right column has Debug TabPane + #debug-transcript RichLog."""
from textual.widgets import RichLog, TabPane
app = _resolved_app(_args_new(), session_id="s-new12345", agent_id="mimir")
@@ -837,35 +887,45 @@ class TestLayoutShape:
assert app.query_one("#side-panes", TabbedContent).active == "debug-tab"
async def test_done_label_styled_success(self) -> None:
"""done_label_styled_success [v0.5.1]: [done] label renders in Aurora green."""
"""done_label_styled_success [v0.9.0]: [done] label mounts as Static
carrying a RichText with Aurora green style. Inspect the mounted
Static's `.content`.
"""
from rich.text import Text as RichText
from textual.containers import VerticalScroll
from textual.widgets import RichLog
app = _resolved_app(_args_new(), session_id="s-new12345", agent_id="mimir")
async with app.run_test() as pilot:
await pilot.pause()
# Probe the presenter directly — write a Done via state.render.
from ratatoskr.tui import TuiPresenterState
log = app.query_one("#transcript", RichLog)
transcript = app.query_one("#transcript-scroll", VerticalScroll)
state = TuiPresenterState()
seen: list = []
orig = log.write
log.write = lambda c, *a, **kw: (seen.append(c), orig(c, *a, **kw))[1]
mounted: list = []
orig_mount = transcript.mount
def spy_mount(*ch, **kw):
mounted.extend(ch)
return orig_mount(*ch, **kw)
transcript.mount = spy_mount # type: ignore[method-assign]
state.render(
_make_tui_done(),
log=log,
transcript=transcript,
tools_log=app.query_one("#tools-log", RichLog),
debug_log=app.query_one("#debug-log", RichLog),
current_text=MagicMock(), thinking_log=MagicMock(), raw=True,
thinking_log=MagicMock(),
raw=True,
)
done = next(
c for c in seen
if isinstance(c, RichText) and _text_of(c).startswith("[done]")
w.content for w in mounted
if isinstance(getattr(w, "content", None), RichText)
and _text_of(w.content).startswith("[done]")
)
assert done.style == "#16B866" # Aurora green
async def test_empty_state_placeholders_present(self) -> None:
"""empty_state_placeholders_present [v0.5.1]: tools-log + debug-log show
"""empty_state_placeholders_present [v0.5.1]: tools-transcript + debug-transcript show
placeholder lines before any turn fires."""
from textual.widgets import RichLog
@@ -1073,10 +1133,14 @@ async def _submit_and_wait(app: RatatoskrApp, pilot, content: str) -> None:
class TestStreamTurnWorker:
@respx.mock
async def test_happy_text_done_renders_markdown(self, monkeypatch: pytest.MonkeyPatch) -> None:
"""happy_text_done_renders_markdown [happy,tracer, v0.6.0]:
Text deltas go to current_text (not transcript); on Done, transcript
gets turn-header Rule, [done] label, post-Done Rule + Markdown body.
async def test_happy_text_done_no_double_print(self, monkeypatch: pytest.MonkeyPatch) -> None:
"""happy_text_done_no_double_print [happy,tracer, v0.9.0]:
Text("hello") mounts a Static(Markdown("hello")) into the transcript;
Done mounts a [done] label Static. The Markdown is rendered live (one
widget for the whole stream, updated in place), so there is NO
post-Done re-render — exactly ONE Markdown renderable lands in the
transcript for the response body. v0.9.0 supersedes v0.8.2's
drop-Markdown patch with proper live rendering.
"""
stream = _sse_chunk("42:1", {"type": "text", "content": "hello"}) + _sse_chunk(
"42:2", _DONE_BODY
@@ -1092,24 +1156,27 @@ class TestStreamTurnWorker:
await pilot.pause()
await _submit_and_wait(app, pilot, "hi")
assert app.state == "idle"
# v0.6.0: Text("hello") goes to current_text Static, NOT log.
# writes spy captures RichLog.write only, so "hello" SHOULD NOT appear.
from rich.markdown import Markdown
from rich.rule import Rule
assert not any(w == "hello" for w in writes)
assert any("[done]" in str(w) for w in writes)
# Post-Done: Markdown body + Rule + turn-header Rule all present.
assert any(isinstance(w, Markdown) for w in writes)
assert any(isinstance(w, Rule) for w in writes)
# The response body lives as ONE Markdown renderable mounted into
# the transcript; live updates happen via Static.update, not via
# re-mount, so there's exactly one Markdown in the spy stream.
markdowns = [w for w in writes if isinstance(w, Markdown)]
assert len(markdowns) == 1, (
f"v0.9.0: expected exactly ONE Markdown mounted, got {len(markdowns)}"
)
assert markdowns[0].markup == "hello"
# [done] label fires too.
assert any("[done]" in _text_of(w) for w in writes)
@respx.mock
async def test_raw_flag_skips_markdown_render(self, monkeypatch: pytest.MonkeyPatch) -> None:
"""raw_flag_skips_markdown_render [trace, v0.6.0]:
With --raw, no Markdown render. A turn-header Rule IS still written
(v0.6.0 INV — turn correlation lives in every pane). The post-Done
Rule(separator) is suppressed; accumulated streamed text is written
as a plain string instead.
"""raw_flag_skips_markdown_render [trace, v0.9.0]:
With --raw, the response widget holds plain str instead of Markdown.
Turn-header markers still appear in every pane: 3 RichLog panes
receive a Rule, the transcript-scroll receives a Static-wrapped
RichText (mounted, not written), giving 3 Rules in the captured
writes list.
"""
stream = _sse_chunk("42:1", {"type": "text", "content": "hi"}) + _sse_chunk(
"42:2", _DONE_BODY
@@ -1127,11 +1194,11 @@ class TestStreamTurnWorker:
# No Markdown in raw mode.
assert not any(isinstance(w, Markdown) for w in writes)
# Only turn-header Rules — one per pane (transcript + tools +
# debug + thinking = 4). No post-Done separator Rule.
# 3 Rules — one per RichLog pane (tools / debug / thinking).
# Transcript-scroll uses a Static turn-header Markdown alternative.
rules = [w for w in writes if isinstance(w, Rule)]
assert len(rules) == 4, f"expected 4 turn-header Rules, got {len(rules)}"
# Accumulated text "hi" written as plain string post-Done.
assert len(rules) == 3, f"expected 3 turn-header Rules, got {len(rules)}"
# Accumulated text "hi" mounted as plain str into transcript.
assert "hi" in writes
@respx.mock
@@ -1505,8 +1572,14 @@ class TestActionInterrupt:
await pilot.pause(0.02)
# Give _cancel_via_sse time to write the [cancel_failed] line
await pilot.pause(0.05)
log = app.query_one("#transcript", RichLog)
rendered = "\n".join(str(strip.text) for strip in log.lines)
from textual.containers import VerticalScroll
from textual.widgets import Static
transcript = app.query_one("#transcript-scroll", VerticalScroll)
rendered = "\n".join(
str(child.content)
for child in transcript.children
if isinstance(child, Static)
)
assert "[cancel_failed]" in rendered
assert app.state == "cancelling"
stream_gate.set() # let stream finish for teardown
@@ -2121,6 +2194,72 @@ class TestResolveThenRunWithPicker:
assert agents_route.call_count == 0
assert sessions_route.call_count == 1
@respx.mock
def test_picker_merges_local_tier3_agents(
self,
monkeypatch: pytest.MonkeyPatch,
tmp_path: Path,
) -> None:
"""picker_merges_local_tier3_agents [v0.8.0]: local index entries
appear in the picker's agent list alongside remote agents."""
from ratatoskr.local_agents import LocalAgentEntry, add_local_agent
# Isolate the local index in a tmp file.
monkeypatch.setenv(
"RATATOSKR_LOCAL_AGENTS", str(tmp_path / "local_agents.json")
)
add_local_agent(LocalAgentEntry(
agent_id="ratatoskr:wizard",
agent_name="wizard",
model="qwen3.6-35-a3b",
description="(tier 3) test wizard",
defined_at="2026-05-25T00:00:00+00:00",
))
respx.get("https://w.example/agents").mock(
return_value=httpx.Response(200, json=_AGENTS_RESP)
)
respx.post("https://w.example/sessions").mock(
return_value=httpx.Response(201, json=_CREATE_OK_RESP)
)
from ratatoskr.tui import AgentPickerApp
captured: list = []
async def capture_picker_init(self, *a, **kw):
captured.append(list(self.agents))
return "mimir" # auto-pick something so the rest succeeds
# Patch __init__ to capture the agent list passed to the picker.
orig_init = AgentPickerApp.__init__
def init_spy(self, agents):
captured.append(list(agents))
orig_init(self, agents)
monkeypatch.setattr(AgentPickerApp, "__init__", init_spy)
async def picker_returns_mimir(self):
return "mimir"
monkeypatch.setattr(AgentPickerApp, "run_async", picker_returns_mimir)
async def fake_main(self, *a, **kw):
return 0
monkeypatch.setattr(RatatoskrApp, "run_async", fake_main)
from ratatoskr.tui import run_tui
rc = run_tui(_args_new_no_agent())
assert rc == 0
# The local tier-3 agent should appear in the picker's agents list.
assert captured, "AgentPickerApp.__init__ was never called"
agent_ids = {a.agent_id for a in captured[0]}
assert "ratatoskr:wizard" in agent_ids
# Plus the remote agents.
assert "mimir" in agent_ids
assert "lofn" in agent_ids
@respx.mock
def test_picker_skipped_when_session_mode(self, monkeypatch: pytest.MonkeyPatch) -> None:
"""picker_skipped_when_session_mode: --session s-1 → no list_agents, no create_session."""
@@ -2187,8 +2326,16 @@ class TestResolveThenRunWithPicker:
self,
monkeypatch: pytest.MonkeyPatch,
capsys: pytest.CaptureFixture[str],
tmp_path: Path,
) -> None:
"""list_agents returns [] → stderr [no_agents]; exit 13; picker NOT opened."""
"""list_agents returns [] AND no local tier-3 entries → stderr
[no_agents]; exit 13; picker NOT opened. Isolate
$RATATOSKR_LOCAL_AGENTS so the operator's real local index
doesn't merge in and turn this into a non-empty list."""
# v0.8.0 isolation: point local agents at an empty tmp file.
monkeypatch.setenv(
"RATATOSKR_LOCAL_AGENTS", str(tmp_path / "empty_local_agents.json")
)
respx.get("https://w.example/agents").mock(return_value=httpx.Response(200, json=[]))
from ratatoskr.tui import AgentPickerApp
Generated
+1 -1
View File
@@ -968,7 +968,7 @@ wheels = [
[[package]]
name = "ratatoskr"
version = "0.7.1"
version = "0.9.0"
source = { editable = "." }
dependencies = [
{ name = "httpx" },