a358cc915014945d463c7aaf12abe70978f958de
56 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
a358cc9150 |
memory: snapshot — b1→b2 409/503 adaptation + bifrost 1.0.0 + combined-bind default + admin key + regard-dead-axis finding
Worldtree on v1.0.0b2 (both demo + personal); ratatoskr v0.18.4 all pushed. Session arc: web combined-bind default (v0.18.1), bifrost 1.0.0 repin (v0.18.2), b1/b2 eager 409/503 status mapping (v0.18.3/.4), readonly-admin key collected (#11 prereq cleared), and the regard-dead-axis finding (provider-side catch -> worldtree-dev escalating to Vuong). Next substantive effort = the v1 coverage-audit (folds in the deferred live-409 + b2 spec re-vendor). |
||
|
|
c5c8ecf9d5 |
memory: snapshot — #17 CLOSED + #18 composite final leg PROVEN end-to-end
The Worldtree-driven composite :8392 smoke ran and is proven + persisted: one bound session drove the full both-plane lifecycle through one endpoint (handshake both caps -> affect.fetch + memory.search -> affect.emit stored:true -> memory.upsert_many upserted:1), both writes verified in our SQLite stores. infra-ops allowlisted :8392 (01KVHWJGTT); #17 closed in the tracker. No open legs remain on the composite; repo at a converged checkpoint. |
||
|
|
4f16ba588d |
memory: snapshot — #18 CLOSED end-to-end + v0.18.0 (composite both-plane binding)
#18 D1 SHIPPED: build_combined_provider_app on :8392 wraps bifrost 0.10.0's public build_combined_app over both stores + the shared affect read route; one bound WT session drives memory.* AND affect.* through one endpoint; op-feed plane='combined' per-path. Shipped v0.17.15 (affect.fetch, the strong-or-absent prerequisite) -> v0.17.16 (composite) -> v0.17.17 (#17 op-feed field-name fix) -> v0.18.0 (publishing minor). Suite 503 green. Live-smoke PROVEN at wire+dispatch (real stores + bifrost 0.10.0 on a running :8392): handshake grants both caps, PAD read route serves real sindra PAD, both planes dispatch at one bound session_id. WT-driven turn gated on infra-ops adding :8392 to WT's BIFROST_CLIENT_ALLOWED_HOSTS (requested). New decisions: reference-impl-adopt-canonical (operator); v1-derived-from-WT-I/O-coverage (operator). New foot-guns: memory-store check_same_thread bug (same as affect D2, exposed by the contract-mandated search test via TestClient); :8392 infra-allowlist gate; heid-review test-fidelity nudge cascaded into 2 latent-bug fixes. |
||
|
|
a0c6c73ab9 |
memory: snapshot — #18 D2 SHIPPED+PUSHED (v0.17.14, 39eebd1): web pane renders live PAD/valence from our :8390 store, persona-telemetry gap closed; full #17+#18 arc now on origin. D1 (composite :8392) PARKED on bifrost build_combined_app (~v0.9.0, design locked, after WT #289). FR-1 RESOLVED — composite is bifrost-only, ZERO WT change (single-endpoint caps-routed, worldtree-dev code-verified). New decisions: #18 split + Option-C canonical-surface routing; D2 TDD + heid-code-review (1 INV-001 drift + 4 test-gaps fixed). Foot-guns: rationalized-away a known INV-001 deviation that only the post-impl cross-model review caught; latent sqlite check_same_thread bug exposed by the HTTP read route. FOOT-GUN: running :8390/:8765 are PRE-#18 code — restart with new code + RATATOSKR_AFFECT_READ_URL to see D2 live.
|
||
|
|
f3bac46238 | memory: snapshot — persona-telemetry diagnosis sharpened + #18 split; archived the 2026-05-* build-era cluster (59 entries: 41 decisions + 18 foot-guns) to archival-memory.md. New: wire-verified Tier-3 emits ZERO affect_update SSE (both WT persona sources dead → #18 PAD-display half is the only path); PAD confirmed in our :8390 store (vuong 8 turns, familiarity 0.18→0.59); affect.emit is POST-TURN ASYNC foot-gun. persistent-memory.md trimmed 331→~190. | ||
|
|
f15c8c6153 | memory: snapshot — #17 SHIPPED end-to-end (slices 1-3c, v0.17.8-.13, suite 470 green, live-smoke PROVEN: bound CLI->sindra->op-feed captured 2 recall searches @ exact bound session_id 2c0c7482 with #297/#298 union scopes; dispatch JWT carries session_id=sub, open-q resolved). Operator session UP: web :8765 bind-configured + plane selector, providers :8390/:8391 with op-feed, althing monitor armed. Persona-pane PAD gap diagnosed (affect persists to :8390 stored:true but pane reads Tier-3-404 persona_state) -> #18 filed (composite endpoint + PAD read-endpoint, operator approved 'A', contract-first next). | ||
|
|
f533464c54 | memory: snapshot — persona-pane reframe (worldtree-dev): persona_state GET is Tier-1-only by ADR-0009 (colon-404 correct-by-design, not a stub); Tier-3 affect is CLIENT-persisted — we already hold PAD/valence @ :8390 from affect.emit, so the pane is an OUR-side render via #17 affect-binding (→ affect.emit → :8390 → render), NOT a WT endpoint wait. WT #289 affect.fetch = optional mediated-read; #300 = WT client-impl guide. Expands #17 payoff: memory AND the persona pane. | ||
|
|
37cdef511f |
fix(web): de-ugly the Tier-3 persona pane — clear message instead of bare HTTP 404
persona_state hard-404s every Tier-3 (colon-id) agent by design upstream (WT api.py:1220, "Phase 2.0 has no Tier 3 persona") — so the Persona pane showed "persona not available (HTTP 404)" for consumer-defined characters. loadPersona now reads error_code + renders a clear Tier-3-aware message (she still responds in character; only the affect/OCEAN readout is gated), with distinct text for persona_not_configured / 403 / other. Also (snapshot): sindra switched to thoughtful-character role (mistral-small-4-reasoning); worldtree-dev pinged re Tier-3 persona_state roadmap (thread 01KVCR6P); #17 (bifrost-binding the chat client) teed up as the next-context target. v0.17.7 |
||
|
|
835375d22b | memory: snapshot — FULL COVERAGE proven (verbose persona too): sindra-probe theatrical turn promoted the user fact cleanly under Stage 2/v0.36.0 + cold-recalled @0.694; :8081 confirmed on v0.36.0; closes the verbose-persona caveat end-to-end. Operator session: :8391 wiped, ratatoskr-web up :8765 (consumer key, sindra in picker) — persona+debug only, web client does NOT bind :8391 (#17 unbuilt = no memory persistence in web chat) | ||
|
|
7666203722 | memory: snapshot — Tier-3 memory PROVEN end-to-end live (terse-probe cold recall @0.6994, fresh history-free session); #296 arc closed: Stage 1 (v0.35.19) recallability gate validated live + bisect localized residual to verbose-persona volume, Stage 2 (v0.36.0) MERGED at worldtree-codex (user-only per-turn extraction), live-validated eval fixture pair -> #305; root-cause chain v0.35.16 emit-2-meta -> v0.35.19 emit-then-reject -> v0.36.0 fix; foot-gun: :8391 store-wipe != WT promotion-dedup reset (clean promotion smoke needs a fresh agent+end_user) | ||
|
|
84d8c3f65f | memory: snapshot — cold-recall arc PROVEN live e2e (#297/#298 union recall; WT v0.35.16 emits scope_any into our v0.17.6 store); #296 extraction quality the isolated upstream gap (triage→worldtree-dev, both symptoms localized in-code: empty _EXTRACTOR_SYSTEM + both-roles prefilter); sindra restored (DELETE+redefine, role:character→mistral-small-4, memory:{}); learnings: Tier-3 owner-scoped, define-takes-role, promotion 4-trigger hybrid, DELETE≠drain | ||
|
|
4eee7c89b2 |
pin: bump Worldtree spec to f1b59f8 (v0.35.16) — cold recall closes end-to-end
Worldtree shipped its half of the union-recall fix: #297 (client-side per-scope-value union recall) + #298/#299 (adopt the bifrost v0.6 scope_any/scope_all wire, v0.35.16). It now emits scope_any on the recall path, pairing with our v0.17.6 provider — cold cross-session recall is closed end-to-end (pending a live re-smoke against a v0.35.16 instance). Re-vendored conversation-api-spec.md + conversation_api.contract.md; 285-commit catch-up (v0.29.0 -> v0.35.16). Diff-reviewed: no client-facing breaking changes for our consumer. - #211 agent-slug rename (saga->echo, actor->mask) — slugs only, we pass --agent - #245 end_user_id persistence + memory-scope resolver (additive) - #187/#188/#219 Tier-3 define/PATCH policy (additive); error codes stable - bifrost binding field + ephemeral_does_not_accept_bifrost 422 now documented (#17 surface) - docs: SPEC-PIN.md pin table + history; bifrost-self-test recall status; persistent-memory No package version bump (docs/pin-only, no ratatoskr code change). |
||
|
|
96d61a4bb1 |
feat(provider): split memory search scope_filter → scope_all + scope_any (bifrost 0.8.0/wire v0.6)
Repin bifrost 0.7.0→0.8.0 and reimplement the memory store's search scope filter to the v0.6 split (#11): scope_all (AND/intersection) + scope_any (OR/union over a list of conjunctive scopes), at parity with the v0.6 reference _matches_scope / _validate_scope. No-compat: scope_filter removed. scope_any is the union-visibility primitive that resolves the #295/#297 silent-zero AND foot-gun — a subset-scoped chunk now recalls via an OR member. End-to-end cold recall now gated only on Worldtree emitting scope_any on its recall path (#297, upstream). - store: search(scope_all, scope_any); _scope_subset + _matches_scope + _validate_scope - contract v1.2: search FN sig, INV-005 recomposed, PRE-003 both fields, scope_any_union test - tests: scope_any union, scope_all∧scope_any compose, both-empty match-all; parity vs real 0.8.0 dispatch (433 green) - #17 contract: sync stale scope_filter/_scope_matches-AND refs to scope_all/scope_any - runbook + persistent-memory updated; provider bounced onto 0.8.0 (fresh empty db) v0.17.6 |
||
|
|
43f2e148ad | memory: snapshot — observe brick + self-drive proven; #295 root-caused (upstream); agent_self canonical shipped both sides + 4-axis parity (v0.17.5); #17 contract reviewed, TDD next | ||
|
|
ca02c70b7c |
docs(#17): self-drive+observe contract, bifrost self-test runbook + snapshot
- docs/contracts/issues/17.contract.md — issue-scoped v2.1 contract for #17 (Bifrost-binding the chat client). v1 scope = single-plane bind + dispatch-layer op-feed (composite endpoint + turn-pane UI parked). Design consulted via /heid, paraphrase-gated via /heid-contract-review panel; two internal inconsistencies fixed (OpEvent turn_id reservation made literal; session_id-for-all-verbs correction). Validates OK, prd drift-clean. - docs/bifrost-self-test.md — reusable runbook for driving + observing the full Bifrost round-trip against our own provider (the manual form of #17; pins the consumer-key-as-bearer tripwire). - persistent-memory.md — snapshot: observe brick shipped, self-drive proven, #295 root-caused (upstream, scope-axis asymmetry) -> #296/#297, agent_self -> canonical decided. |
||
|
|
2b47dcff5a | memory: snapshot — memory provider live-proven (persist/dispatch/search); recall-injection upstream; #17 filed | ||
|
|
e57b054054 | memory: snapshot — memory plane shipped (v0.17.3), canonical sync + #3/#4 refresh | ||
|
|
5cdb69a6f3 | memory: snapshot — Bifrost consumer arc (affect live, memory contract v1.1) | ||
|
|
d96415806b | memory: snapshot — v0.17.0 operator-confirmed + issue-tracker cleanup | ||
|
|
489cfee1f0 |
fix(tui): drop post-Done Markdown body re-render (v0.8.2)
Operator: "first turn double prints agent's turn." Root cause: v0.8.1 wrote both the streamed Text lines AND the post- Done `Markdown(event.response)` body into the transcript. Same content rendered twice — once as plain streaming, once as a full markdown re-render. The v0.8.1 commit message documented this as "some duplication is acceptable" but the live UX read as a bug. ## Fix Drop the post-Done `Rule + Markdown(response)` writes in non-raw mode. The streamed text IS the response; whatever the model emitted flows into the transcript line-by-line via coalesce-on-newline. Markdown formatting (bold, lists, code blocks) renders as plain text — a known regression from v0.8.1's polished output but the right tradeoff vs the duplication bug. ## What this loses temporarily Pre-v0.8.2 (after Done): [done] turn_id=... ─── ─── (Rule separator) ─── **Bold text** rendered bold, `code` highlighted, lists as bullets, etc. v0.8.2 (after Done): [done] turn_id=... ─── **Bold text** as plain asterisks, `code` as backticks, lists as plain dashes ## v0.9.0 plan Restore markdown rendering via LIVE rendering during the stream (not post-Done re-render). Replace `RichLog#transcript` with a `VerticalScroll` container that mounts a fresh `Markdown` widget per turn; Text deltas update the widget; markdown renders as content arrives. No duplication, no snap, full formatting. Operator-confirmed direction (2026-05-25 AskUserQuestion). ## Tests 287/287 GREEN; ruff clean. Two tests updated for the new shape: - test_done_renders_markdown_after_label → renamed test_done_flushes_tail_and_writes_label; asserts NO Markdown, NO Rule (post-Done) in the writes. - test_happy_text_done_renders_markdown → renamed test_happy_text_done_no_double_print; asserts NO Markdown in the spy. Patch bump (v0.8.1 → v0.8.2): bug fix; no public API change. |
||
|
|
11ef6830ab |
fix(tui,sse): inline Text streaming + empty-id keepalive skip (v0.8.1)
Two related fixes for the same user-reported bug pattern from a
running session against ratatoskr:sindra (qwen3.6-35-a3b-heretic):
## 1. Streaming text overlapping the transcript
Operator: "new text comes at the bottom and overwrites the existing
pane information instead of pushing it up naturally."
Root cause: the v0.6.0 `#current-text` Static was `dock: bottom`
with `height: auto`, sitting between the transcript RichLog (1fr)
and the prompt Input (dock: bottom). As text streamed, the Static
grew UPWARD but Textual didn't dynamically resize the 1fr transcript
to accommodate — the growing Static visually OVERLAPPED the
transcript's bottom rows. On Done, `current_text.update("")` snapped
it to height 0 and the transcript re-laid-out — "boom, everything
updates."
Fix: remove `#current-text` Static entirely. Apply the same
coalesce-on-newline pattern v0.7.1 used for thinking — Text deltas
accumulate in `TuiPresenterState.text_chunk_buffer`, flushing whole
lines (each `\n` boundary) directly to `log` (transcript). On Done:
flush remaining tail, then [done] label + Rule + Markdown body.
Trade-off accepted: streamed lines + post-Done Markdown body are
both in the transcript (some content duplication). The Markdown
body re-renders the same content with proper formatting (lists,
bold, code blocks). Acceptable — operator gets both the live-progress
streaming AND the canonical rendered version.
## 2. MalformedSseId raw='' crashing every turn
Operator: "current session is erroring on every turn with
[malformed_sse_id] raw=''"
Worldtree's qwen3.6-35-a3b-heretic provider emits some events
without `id:` lines (observed 2026-05-25 mid-stream). When the FIRST
such event arrives before any prior id has been seen, httpx_sse's
`ServerSentEvent.id` is `""`. `_parse_sse_id('')` raised ValueError
→ MalformedSseId → turn worker bailed → operator saw the label
every turn.
Per SSE RFC, events without `id:` are legitimate (they just don't
update Last-Event-ID). Issue #7 already covered the empty-DATA
keepalive case with skip-silently semantics. Empty-id is the same
shape of wire weirdness; same fix shape:
if sse.id == "":
continue # treat as keepalive
Ordered AFTER the empty-data branch so an empty-data + empty-id
event still gets skipped on the data check.
## Tests + smoke
287/287 GREEN (was 286, +1 for empty-id skip; +1 net Text-flow test
adjustments). Ruff clean.
Verified Worldtree alive when the user hit the empty-id bug
(/healthz returned ok in 18ms) — not a server-down issue, just
wire-format mid-stream.
## Caveats
The fix doesn't recover content from the dropped empty-id event.
If the event happened to carry meaningful data (not a true
keepalive), we silently lose it. Acceptable trade-off: pre-v0.8.1
EVERY turn died on the offending agent; post-v0.8.1 the turn
continues and any single dropped frame is recoverable from logs if
debugging. Worldtree-side fix (always emit ids) is the right
upstream answer; ratatoskr just stops panicking on wire weirdness.
Patch bump (v0.8.0 → v0.8.1) — both fixes are bug fixes; no public
API change. The `TuiPresenterState.render` signature loses the
`current_text` parameter (was added v0.6.0), but presenter is an
internal contract; no external callers.
|
||
|
|
9fade55901 |
feat(local_agents): tier-3 index + picker merge (v0.8.0)
Worldtree's GET /agents doesn't return consumer-defined (tier-3)
agents — the public list excludes them by design. Confirmed live in
v0.7.0's smoke. Without server-side knowledge, ratatoskr's picker
couldn't show tier-3 agents the operator had defined; the workflow
was "remember the agent_id, pass --agent ratatoskr:<name>
explicitly." Friction grows with every tier-3 agent.
## Fix: client-side index, merged at picker time
New module `ratatoskr.local_agents` maintains a JSON-backed index at
$XDG_CONFIG_HOME/ratatoskr/local_agents.json (override via
$RATATOSKR_LOCAL_AGENTS). `tier3` CLI define / patch / delete update
the index as side-effects. `tui._resolve_then_run` loads the index
after `list_agents(client)` and appends entries not already in the
remote list (dedup by agent_id; remote wins on conflict).
Library-level `tier3.define_agent` / `patch_agent` / `delete_agent`
stay pure — local persistence lives in the CLI layer (`_run_define`
etc.), not in the library functions. Tests of the library don't
touch the filesystem.
## Public surface
ratatoskr.local_agents:
LocalAgentEntry (frozen dataclass)
load_local_agents() -> list[LocalAgentEntry]
add_local_agent(entry)
update_local_agent(entry) # same semantics as add (agent_id key)
remove_local_agent(agent_id)
make_description(system_prompt) -> str # synthetic picker label
Failure modes are lenient: missing file → empty index; corrupt JSON
or schema mismatch → empty index (no crash). The picker continues
to show foundational agents either way; tier-3 surface degrades to
the pre-v0.8.0 workflow.
## Picker integration
Local entries convert to ratatoskr.sessions.AgentInfo with synthetic
fields:
name = agent_name (from LocalAgentEntry)
description = "(tier 3) <first non-empty line of system prompt>"
version, capabilities, supported_models, persona_traits, ui_hints
= None / [] / [] / {} / {}
If Worldtree later starts returning tier-3 in GET /agents, this
module's role narrows to redundant local cache; can be removed
cleanly since the dedup-by-agent-id keeps remote-wins behavior.
## Tests
286/286 GREEN (was 265, +21: 20 local_agents + 1 picker-merge
integration). Ruff clean. Tests isolate the index via
$RATATOSKR_LOCAL_AGENTS pointed at pytest's tmp_path — no pollution
of operator's real ~/.config/ratatoskr/.
## Manual smoke
Sindra-like define against personal Worldtree:
python -m ratatoskr.tier3 define --name foo --system-prompt "..." --model X
cat ~/.config/ratatoskr/local_agents.json
# ratatoskr --new picker now shows ratatoskr:foo alongside mimir et al.
Cross-machine: the file is per-host. Operator can sync via dotfiles
if needed; out of scope for this commit.
Minor bump (v0.7.1 → v0.8.0) — new public module + new picker
behavior (more agents shown). No caller-side breaking changes.
|
||
|
|
9918c10acf |
fix(tui): coalesce thinking deltas on \n (v0.7.1)
Operator: "thinking tokens seem to be split by token — each on a
newline, is that correct? We don't want that."
Root cause: v0.6.5 wrote each Thinking SSE delta as its own
`thinking_log.write(event.content)` call. Worldtree emits Thinking
events at token granularity (per-token or per-few-tokens), so EACH
token became its own RichLog line — visually choppy, one short
fragment per visual row. Wrong UX.
## Fix: coalesce-on-newline
Thinking deltas accumulate in `TuiPresenterState.thinking_chunk_buffer`
(new str field). On each Thinking event:
1. Append delta content to buffer.
2. Flush every COMPLETE line (chars before each `\n`) as one
thinking_log.write(line) call.
3. Leave the post-final-`\n` tail in the buffer for the next delta.
On any non-thinking event (run close):
1. Flush remaining buffer tail (if any) as one final line.
2. Write Rule(end).
Empty lines (blank paragraph separators in the model's `\n\n` flow)
are skipped — they'd render as no-content RichLog entries which
just add vertical noise. Natural paragraph breaks become single
visible lines; multi-paragraph thinking renders top-to-bottom.
## Verified live (tier-3 smoke against personal Worldtree)
Defined a `thinky-smoke` agent via `python -m ratatoskr.tier3 define`,
asked "What is 12 times 13?". Thinking pane rendered with natural
paragraph chunks:
── turn N · thinking #1 start ──
Thinking Process:
1. **Analyze the Request:** The user wants to know the result of $12 \times 13$.
2. **Calculate:**
* Method 1: Standard multiplication.
$$12 \times 10 = 120$$
$$12 \times 3 = 36$$
$$120 + 36 = 156$$
* Method 2: $(10 + 2)(10 + 3) = 100 + 30 + 20 + 6 = 156$.
── turn N · thinking #1 end ──
Each line = one natural paragraph or list item. No per-token fragments.
## Edge cases noted
- Long-running thinking with NO `\n` at all stays buffered until run
close → operator sees nothing until close. Possible follow-up: add
a length-threshold flush (e.g., > 500 chars → flush at the last
space). For now this is acceptable; thinking content typically has
`\n` breaks every few sentences.
- Empty deltas (`""`) are ignored implicitly — no buffer growth, no
flush.
- `\n` at the very start of a delta flushes whatever was buffered
before, then leaves the empty post-`\n` tail (empty string) in the
buffer, which doesn't show up as an empty line because of the
`if line:` guard.
## Contract amendment
docs/contracts/issues/13.contract.md INV-022 amended for v0.7.1
coalesce semantics. Drift-check clean.
## Tests
265/265 GREEN; ruff clean. Two updated tests:
- `test_thinking_streams_into_thinking_log` → renamed
`test_thinking_coalesces_until_newline`: 3 token-shaped deltas
with no `\n` → only Rule(start) writes, buffer holds accumulated.
- NEW `test_thinking_flushes_on_newline`: delta carrying `\n` →
Rule(start) + accumulated line + clear buffer.
- `test_thinking_closes_to_thinking_log`: 2 deltas "a", "b" +
close → Rule(start) + tail-flush "ab" + Rule(end) = 3 writes
(was 4 with per-delta).
Patch bump (v0.7.0 → v0.7.1) — internal presenter routing change;
no public-API or layout change.
|
||
|
|
c086ae2b32 |
feat(tier3): ratatoskr.tier3 module + CLI (v0.7.0)
Issue #15. Worldtree Phase 2.0 ships Tier 3 (consumer-defined) agents at `<user_id>:<agent_name>`; ratatoskr now exposes their lifecycle via a dedicated module + CLI tool. The picker handles the colon-containing agent_id generically (per issue #8 out-of- scope clause); session creation works unchanged. What was missing was a way to DEFINE / PATCH / DELETE these agents from ratatoskr itself — operators previously had to curl the API directly. ## Public surface (ratatoskr.tier3) Tier3AgentInfo (frozen dataclass) define_agent (client, *, agent_name, system_prompt, model) → Info patch_agent (client, agent_id, *, system_prompt?, model?) → Info delete_agent (client, agent_id) → None Tier3QuotaExceeded — 429 agent_quota_exceeded (50-agent cap) Tier3UserIdUnsupported — 403 tier3_user_id_unsupported Tier3FieldNotMutable — 422 field_not_mutable (PATCH) Tier3LayerDeferred — 422 layer_deferred (define, defense-only) Tier3AgentNotFound — 404 SessionApiFailed (reused) — all other non-2xx Caller-owned httpx.AsyncClient posture (same as ratatoskr.sessions). Module is standalone — does NOT import sessions/sse_client/tui/cli beyond reusing the USER_AGENT constant from cli. ## CLI (python -m ratatoskr.tier3 <subcommand>) define --name <slug> --system-prompt <str> --model <id> patch <agent_id> [--system-prompt <str>] [--model <id>] delete <agent_id> Auth resolution mirrors ratatoskr.cli verbatim — --api-key flag > $WORLDTREE_API_KEY > exit 11. Server URL via --server > $WORLDTREE_API_URL > http://localhost:8000. Exit codes follow the cli.py matrix: 0 / 10 (usage) / 11 (auth) / 20 (api-failure) / 21 (network). ## Real-world finding from live smoke Tier-3 agents do NOT appear in `GET /agents` — the public list filters them out. The picker won't surface tier-3 agents; operators bypass it via `ratatoskr --send "..." --new --agent ratatoskr:<n>` directly. This contradicts the contract's acceptance assumption ("the new tier-3 agent should appear in the list") — caught at smoke time. The picker integration was hopeful; the real shape is "you know your tier-3 agent_id because you defined it." Adding a ratatoskr-side `tier3 list` subcommand would need a Worldtree endpoint that doesn't exist today; surfacing to worldtree-dev as a followup. ## Live lifecycle smoke (personal Worldtree v0.16.2) $ python -m ratatoskr.tier3 define --name smoke-tier3 \ --system-prompt "..." --model qwen3.6-35-a3b → defined ratatoskr:smoke-tier3 (qwen3.6-35-a3b) $ ratatoskr --send "hello via tier-3" --new --agent ratatoskr:smoke-tier3 → [done] turn_id=286 model=qwen3.6-35-a3b duration=14.2s usage 44 in → 390 out (434 total, 0 cached) $ python -m ratatoskr.tier3 delete ratatoskr:smoke-tier3 → deleted ratatoskr:smoke-tier3 $ python -m ratatoskr.tier3 delete ratatoskr:smoke-tier3 → [agent_not_found] ratatoskr:smoke-tier3 (exit 20) The colon-containing agent_id flowed transparently through ratatoskr.sessions.create_session, the SSE stream's text + worker_phase + done events all rendered correctly, and the ratatoskr.sessions module needed zero changes. ## Contract docs/contracts/issues/15.contract.md — new module spec; drift-check clean. Acceptance criterion about "appears in GET /agents" should be amended in a follow-up to reflect the empirical finding. ## Tests +26 tests (264 total GREEN, was 238). Covers all error paths via respx mocking — quota, user_id, layer_deferred, field_not_mutable, 404, 5xx — plus CLI happy + error paths. ruff clean. Minor bump (v0.6.5 → v0.7.0) per SemVer etiquette: new public module + CLI surface; new caller-visible behavior. |
||
|
|
d3569904bc |
refactor(tui): thinking streams into whole pane (v0.6.5)
Operator: "Why does the thinking scroll a little section at the
bottom of the thinking pane instead of scrolling the whole pane?"
Root cause: v0.6.1's thinking-current Static was docked to the
bottom of the Thinking pane and rendered the last 200 chars of
streaming content. As deltas arrived, the displayed 200-char tail
shifted — old text fell off the left, new text appeared on the
right — visually reading as "a little section scrolling at the
bottom" while the larger thinking-log RichLog above showed only
the previous run's closed content (or nothing on first turn).
## Fix: stream directly into thinking-log
The Static is gone. Thinking deltas now write straight to the
`thinking-log` RichLog (one delta = one line in the scrollable
log). The whole pane scrolls naturally as content arrives —
operator can switch to Ctrl+3 and see streaming content fill
the pane top-to-bottom.
Routing pattern:
First Thinking delta of run:
→ write Rule(title="turn N · thinking #K start") to thinking_log
→ write delta content as a line
→ set thinking_open = True
Subsequent Thinking deltas:
→ write delta content as a line
Non-thinking event (closes the run):
→ write Rule(title="turn N · thinking #K end") to thinking_log
→ reset thinking_open
The Rule(start) at the top of an in-progress run is now the
"thinking is happening" indicator. No more separate live-preview
widget required.
## Trade-off: no markdown re-render
Pre-v0.6.5 closed runs got a Markdown(full_content) render between
the start/end Rules. v0.6.5 drops that — the streamed deltas ARE
the content; re-rendering as Markdown would either need to wait
for run-end (no streaming) OR re-render incrementally per delta
(bad UX). Streaming wins for "live observability" framing.
The downside: if model thinking has Markdown structure (lists,
code), it renders as raw text. Acceptable per operator's "stream
in line" framing.
## Removed widgets
- `Static#thinking-current` (right column / Thinking pane bottom)
- `TuiPresenterState.render` no longer takes a `thinking_widget` param
- `TuiPresenterState.thinking_buffer` field dropped (no accumulation)
- `_stream_turn_worker` no longer queries `#thinking-current`
- `on_mount` no longer hides `#thinking-current`
- DEFAULT_CSS `#thinking-current` block removed
## Contract amendment
INV-022 amended: thinking now streams as raw delta lines, not
Markdown-rendered on close. INV-024 amended: thinking-current
Static removed entirely (was relocated v0.6.1, removed v0.6.5).
Drift-check clean.
## Tests
238/238 GREEN (was 241 — 3 obsolete widget tests deleted:
test_thinking_widget_truncation, test_thinking_widget_visibility_lifecycle,
test_terminal_events_belt_and_braces_widget_cleanup). 5 routing tests
rewritten for the new streaming shape (test_thinking_streams_into_thinking_log,
test_thinking_closes_to_thinking_log, test_multiple_thinking_runs_...,
test_render_exception_fallback, test_cancelled_mid_thinking_closes,
test_left_column_content_only).
ruff clean. Manual injection test confirms routing: Rule(start) +
delta lines write to thinking_log; transcript untouched.
Patch bump (v0.6.4 → v0.6.5) — internal restructure within Thinking
pane; presenter signature narrowed; no caller-visible public API
change (RatatoskrApp + AgentPickerApp surfaces identical).
|
||
|
|
82437bd4b9 |
style(tui): picker highlighted item → Aurora blue (v0.6.4)
Operator request: the agent picker's highlighted selection should
get the brand-color treatment — Aurora blue background — instead of
the v0.6.1 dark-30 muted bg.
## Two-fix landing
**The selector**: v0.6.1's `ListView > ListItem.--highlight` (double
dash) never actually matched. Textual's class is `-highlight` (single
dash). The v0.6.1 "fix" silently fell through to Textual's defaults,
which happened to be invisible because $block-cursor-background was
configured but the selector path didn't reach the rendered widget.
Probed live: `item.classes = frozenset({'-highlight'})`. Selector
corrected, plus dropped the `>` combinator since Textual's internal
DOM puts wrappers between `ListView` and `ListItem`.
**The background**: explicit `#agent-list:focus ListItem.-highlight
{ background: $primary }` — Aurora blue (#6388D8) for the focused-
list highlight band.
**The contrast**: bright-blue id-line text on Aurora-blue background
would be unreadable. Highlighted-state child overrides:
.agent-id-line → $au-bright-white (#cce7ec) + bold
.agent-desc → $au-bright-80 (#b3cbcf)
Non-highlighted items keep their default colors (bright-blue id +
bright-70 desc on App bg).
## Verified live
13 fill rects of `#6388d8` in the picker SVG export (was 0
before this commit). Other Australis brand colors intact:
chrome surface #373b46, dark-50 #6e7882, dark-30 #414751.
## Tests
241/241 GREEN; ruff clean. No test rewrites needed — picker tests
assert structure (widget tree, key bindings), not colors.
Patch bump (v0.6.3 → v0.6.4) — cosmetic; no public-API change.
|
||
|
|
ac690c11d5 |
style(tui): restore Australis, $background → pure black (v0.6.3)
Reverts v0.6.2's over-correction. The operator clarified: the complaint was specifically about the APP BACKGROUND going from black to a shade of blue, not about the cumulative cast across all Australis dark surfaces. v0.6.2 globally neutralized Ice + Sea darks → too far. ## v0.6.3 = v0.6.1 palette + $background override only Restored verbatim from v0.6.1: $foreground #a9bcc3 (Ice white) $surface #373b46 (Sea bright-black, chrome bg) $panel #414751 (Sea dark 30, borders) $au-dark-30..60 Australis Sea palette $au-bright-70/80 Australis Sea brights $au-bright-white #cce7ec (Ice highlight) Aurora accents bright-blue/cyan/green — verbatim Dawn accents red/yellow — verbatim _AU_DEMOTED #86929d (Sea dark 60) _AU_DEMOTED_FAINT #6e7882 (Sea dark 50) ONE deviation from Australis spec: $background #222531 (Ice black) → #000000 (pure black) Rationale: Ice black is RGB(34, 37, 49) — blue +44% over red. At App-wide scale (the dominant fill across the entire screen) the cumulative cast reads as "the app is blue" even though no single rect is in the conventional-blue range. Other dark surfaces are smaller chrome bands where the cool lean reads as character not background; only $background gets the override. ## What stayed Australis Every cosmetic element where the operator hasn't pushed back: identity widget (Aurora bright-blue), pane-name (Aurora bright-cyan), [done]/[error]/[cancelled] labels (Aurora green / Dawn red/yellow), focus borders (Aurora blue / accent cyan), demoted telemetry text (Sea dark-60), placeholder lines (Sea dark-50), Header/Footer chrome (Sea bright-black bg + Ice white-blue fg), separators (Sea dark-30). Brand fidelity preserved; only the dominant background surface neutralized. ## Tests + smoke 241/241 GREEN; ruff clean. Live screenshot export: - $background = #000000 (230 fill rects — dominant surface) - $surface = #373b46 (29 fill rects — Australis Sea bright-black) - Sea panels + dark-50 still present in chrome - Aurora #6388D8 still primary Patch bump (v0.6.2 → v0.6.3) — cosmetic refinement; no public-API change. |
||
|
|
d845b20efd |
style(tui): neutralize Australis dark palette (v0.6.2)
Operator-flagged third pass: "overall background for the whole app is blue." The previous "zero blue rects" investigations missed the structural cause — Australis's design principle "all colors are cooler than neutral" bakes a blue cast into every dark surface: Ice black #222531 = RGB(34, 37, 49) — blue +44% over red Sea bright #373b46 = RGB(55, 59, 70) — blue +27% over red Sea dark-30 #414751 = RGB(65, 71, 81) — blue +25% over red Sea dark-60 #86929d = RGB(134,146,157) — blue +17% over red Every chrome surface inherits the lean. The user reading "the whole app is blue" is correct — the SVG export just rendered hex values that aren't named "blue" but ARE measurably blue-tinted. ## Fix: keep accents, neutralize darks Australis brand signature lives in the ACCENTS — Aurora blue, cyan, green; Dawn red, yellow. Those are unchanged. The Ice/Sea dark palette is replaced with LAB-matched neutral grays (R=G=B) so the chrome reads truly neutral: $background #222531 → #1a1a1a (neutral near-black) $surface #373b46 → #2a2a2a (neutral dark gray) $panel #414751 → #3a3a3a (neutral mid gray) $foreground #a9bcc3 → #bdbdbd (neutral light gray) $au-dark-30 #414751 → #3a3a3a $au-dark-40 #565f69 → #4f4f4f $au-dark-50 #6e7882 → #6b6b6b $au-dark-60 #86929d → #878787 $au-bright-70 #9daeb6 → #9e9e9e $au-bright-80 #b3cbcf → #bdbdbd $au-bright-white #cce7ec → #e0e0e0 `_AU_DEMOTED` and `_AU_DEMOTED_FAINT` constants (Rich Text styling for demoted telemetry + placeholders) updated to the neutral equivalents. The Aurora bright variants (`$au-bright-blue`, `$au-bright-cyan`, `$au-bright-green`) stay verbatim — those are where the brand voice lives. ## What this preserves vs sacrifices **Preserved**: - Aurora accents: focus borders, active-tab indicator, pane-name widget, user-prompt echo all still render in cyan/blue/green. - Done/Error/Cancelled labels still tinted in Aurora green / Dawn red / Dawn yellow. - Identity widget still Aurora bright-blue. - The "Australis" theme name + variable slugs ($au-*) — downstream CSS rules don't have to change. **Sacrificed**: - The "all colors cooler than neutral" Australis design principle. Deliberate per-operator-feedback deviation; documented in the AUSTRALIS_THEME docstring as a v0.6.2 conscious break with spec. ## Tests + smoke 241/241 GREEN; ruff clean. Live screenshot exports: - Main app: chrome colors are #1a1a1a / #2a2a2a / #3a3a3a / #6b6b6b / #bdbdbd — all neutral grays. Aurora accents preserved as textual highlights. - Agent picker: same — neutral chrome, Aurora accents intact for highlighted item border + agent-id-line. Patch bump (v0.6.1 → v0.6.2): purely cosmetic palette adjustment; no public-API change. |
||
|
|
eb93e6d5f0 |
style(tui): kill remaining blue + thinking-current into pane (v0.6.1)
Three operator-flagged issues:
## 1. "Background is still blue" — Header sub-widgets + scrollbar
Two surviving blue sources after v0.6.0:
- **Header sub-widgets** (HeaderIcon, HeaderTitle, HeaderClock) each
carry their own `$primary` tint that the parent
`Header { background }` rule alone doesn't override. Sub-selectors
added: `Header, HeaderIcon, HeaderTitle, HeaderClock { background:
$surface; color: $au-bright-blue; }`.
- **Scrollbar gutter** uses Textual's `$primary-tint` (#32436a) by
default. Per-widget scrollbar overrides: `ListView` (picker) and
`RichLog` (every pane) get explicit Sea darks for gutter + thumb.
Live verification: both AgentPickerApp and RatatoskrApp now render
ZERO instances of `#6388d8` (Aurora blue) or `#32436a` (its dark
derivative) in the export-screenshot SVG.
## 2. "Picker is bright cyan with unreadable text" — ListView focus
Textual's default `ListView:focus > ListItem.--highlight { background:
$primary }` was overriding my v0.6.0 `#agent-list > ListItem.--highlight
{ background: $au-dark-30 }` because `:focus` carries higher
specificity. The highlighted item was rendering with Aurora-blue
background + bright-blue text = unreadable.
Fix: both selectors targeted explicitly with sufficient specificity:
`ListView > ListItem.--highlight, ListView:focus > ListItem.--highlight
{ background: $au-dark-30 }`. Description text bumped to Sea bright-70
for better contrast against the dark-30 highlight.
## 3. "Streaming everywhere, should just stream in line"
User flagged the disconnect: live thinking rendered above the
TabbedContent in the right-column header, then on closure the content
"moved" to thinking-log inside the Thinking pane. Read as jarring
discontinuity.
Fix: `thinking-current` Static moved INTO the Thinking TabPane (docked
bottom), below `thinking-log`. Both surfaces co-located now — live
streaming + closed runs share the same pane. Operator switches to
Ctrl+3 (Thinking) to see chronological closed runs ABOVE + live
streaming line BELOW. Same pattern as the transcript: closed history
+ inline streaming tail.
Trade-off: live thinking is now visible only when on the Thinking
tab. Pre-v0.6.1 it was always visible above the tabs. The user
explicitly prefers the co-located shape; this is the right call.
## Contract amendment
docs/contracts/issues/13.contract.md INV-024 amended: thinking-current
now docks bottom of the Thinking TabPane (was right-column header).
v0.6.0 layout-spec snapshot updated to reflect the new shape. Drift-
check clean.
## Tests + smoke
241/241 GREEN; ruff clean. Live smoke against personal Worldtree
confirmed:
- thinking-current AND thinking-log both inside thinking-tab.walk_children().
- Post-Done state: 23 closed-run lines in thinking-log, thinking-current
cleared to empty.
- Picker exports zero blue rects; main App exports zero blue rects.
Patch bump (v0.6.0 → v0.6.1): purely cosmetic + layout adjustment
within the existing pane structure; no public-API change.
|
||
|
|
cfee89ac1c |
refactor(tui): streaming + turn headers + Thinking pane + picker fix (v0.6.0)
Operator-driven big-batch polish + restructure:
## 1. Streaming text — no more per-token RichLog spam
Pre-v0.6.0, every Text SSE delta wrote its own RichLog line, so
"Let me read the..." became 4+ separate lines (a Worldtree-style
sentence-by-sentence reveal that read as broken). v0.6.0 adds a
`#current-text` Static docked above the prompt; TuiPresenterState
buffers Text deltas in `text_buffer` and updates the Static in
place. On terminal event the Static clears and the transcript
gets:
- raw=False: post-Done Markdown body + Rule separator
- raw=True: accumulated plain text
The Static collapses to height=0 when empty so the prompt sits at
the column bottom unchanged.
## 2. Turn-ID headers across every pane
`_stream_turn_worker` writes a `Rule(title="turn N")` to all four
log panes (transcript, tools, debug, thinking) on the first event
of each new turn. Operators can now visually correlate "what
happened in Tools during turn 42" by section markers in matching
positions across panes.
## 3. New Thinking TabPane (Ctrl+3)
Closed thinking runs now route to `#thinking-log` (a dedicated
TabPane) instead of `#debug-log`. Each closed run writes three
entries:
- Rule(title="turn N · thinking #K start")
- Markdown(thinking_content)
- Rule(title="turn N · thinking #K end")
Model reasoning often has lists/code/structure — rendering as
Markdown (instead of the previous "· thinking: ..." prefix line)
makes it scannable. The `thinking_run_index` counter scopes per
turn so multi-thinking-run turns get distinct markers.
`thinking-current` Static (live per-delta preview) stays in the
right column above TabbedContent (unchanged from v0.5.0) — live
visibility persists across tab switches.
## 4. Agent picker — multi-line items, full description visible
Pre-v0.6.0 the picker rendered each agent as a single Label with
"{id} · {name} — {description}", which truncated descriptions
visually. v0.6.0 uses two Static children per ListItem:
- bold Aurora bright-blue line: "{agent_id} · {name}"
- wrapped Sea dark-60 line(s): full description
ListItems are auto-height so long descriptions wrap as needed.
Highlighted (--highlight) row uses Sea dark-30 background instead
of Aurora blue (which the operator flagged as ugly).
## 5. Kill residual blue chrome
The user's "background is still blue" report traced to the prompt
Input's focused border, which I'd set to $primary (Aurora blue).
Switched to $au-bright-cyan (#42dcd1) — focus highlight is now
cyan, consistent with the operator's-voice accent throughout the
TUI. Also added explicit overrides for ContentTabs strip
background + active-tab underline color → Australis cyan.
## 6. Surfaced emotion-appraisal request to worldtree-dev
User asked for emotion-appraisal telemetry, but no SSE event for
this exists in the spec — persona/Vili affect lives in persona.log
(file-tail, blocked on remote-Worldtree topology) and per-character
state (poll endpoint, not per-turn). Posted an althing thread
proposing two shapes (worker_phase payload extension OR new
affect_update event type) and routing the decision to their team.
A 4th `Emotion` TabPane plugs in trivially when a wire event lands.
Low-priority / quality-of-life framing — not blocking ship.
## Contract amendment
docs/contracts/issues/13.contract.md amended in-place: INV-019
extended to 3 TabPanes; new INV-021 (Text → current_text Static),
INV-022 (thinking closed runs → thinking_log with Markdown +
start/end Rules), INV-023 (turn-ID headers across all panes),
INV-024 (thinking-current Static stays in right column with
"thinking… " prefix per v0.5.1 polish). INV-020 (render-exception
fallback routing) updated for Thinking → thinking_log. Drift-check
clean.
## Tests
241 GREEN (down from 244 in test count — 5 routing tests rewritten
for the new shape, replacing the v0.5.0 thinking-in-debug-log
assertions with the v0.6.0 thinking-log-as-Markdown shape; net
test coverage equivalent). ruff clean.
Live smoke against personal Worldtree's mimir confirmed:
- transcript: 27 lines (turn header + user echo + done +
markdown body, NO per-token spam)
- thinking_log: 19 lines (turn header + 2x thinking start/end
Rule sections with Markdown bodies)
- current_text cleared post-Done
Minor bump (v0.5.1 → v0.6.0) per SemVer etiquette: visible routing
+ new pane = operator-observable surface change.
|
||
|
|
7106af5c09 |
style(tui): UI polish pass (v0.5.1)
Cosmetic refinements on top of v0.5.0's content-only main pane. No
behavior change; ships as a patch bump.
## Color signal — terminal labels tinted per outcome
The transcript's [done]/[error]/[cancelled] labels were plain
foreground (Australis #a9bcc3 white), which made them slow to scan
against the surrounding assistant text. Now tinted per outcome:
- [done] → Aurora green (#16B866 / $success)
- [error] → Dawn red (#ff491a / $error)
- [cancelled] → Dawn yellow (#e1c631 / $warning)
The post-Done Rule() separator is also tinted to Australis dark-60
(#86929d) so the streamed-text → markdown-body boundary reads as
chrome, not a content artifact.
## Empty-state placeholders
Tools and Debug panes were stark-empty before any turn fired — easy
to misread as "the pane is broken." Now show placeholder lines on
mount in Sea dark-50 italic:
Tools tab: (no tool events yet — start a turn that uses tools)
Debug tab: (waiting for telemetry — start a turn)
The placeholders scroll off naturally as real events fill the panes.
## Live thinking widget self-explains
The thinking-current Static at the top of the right column used to
just display raw thinking content with no context — an operator
glancing at the screen mid-stream might not realize they were
looking at LLM chain-of-thought. Now prefixed with "thinking… " so
the widget self-identifies.
## Spacing + chrome
- Transcript / tools-log / debug-log: 1-cell horizontal padding so
content doesn't hug the column border.
- thinking-current: italic text-style on top of the dark-60 color,
so the live-preview band is visually distinct from solid-colored
log content.
- Active tab in TabbedContent: Aurora bright-cyan label + bold
text-style, so the eye lands on the currently selected pane.
- Input placeholder text: tinted to Sea dark-50 so it reads as
placeholder, not content.
## Test impact
3 new tests added (test_done_label_styled_success,
test_empty_state_placeholders_present, plus the polish hits
test_thinking_widget_truncation / test_thinking_coalesce updated for
the "thinking… " prefix). 4 existing tests that checked
`isinstance(w, str) and w.startswith("[done]")` updated to use the
_text_of helper (terminal labels are now RichText, not str).
_spy_writes helper widened to accept positional args after Textual's
internal deferred-render path started passing them positionally
post-Resize.
241/241 GREEN; ruff clean. Live smoke against personal Worldtree
confirmed: Done line renders in Aurora green #16B866 verbatim;
both placeholder lines appear in dark-50; thinking widget shows
"thinking… <content>" during a turn.
Patch bump (v0.5.0 → v0.5.1) per SemVer etiquette: purely cosmetic;
no signature change; no caller-visible behavioral shift.
|
||
|
|
ffd22fb587 |
refactor(tui): content-only main pane + Debug tab + dark chrome (v0.5.0)
Two operator-driven changes off v0.4.1: 1. **Main pane is content-only.** Pre-v0.5.0 the transcript mixed assistant text with telemetry (Thinking closed runs, WorkerPhase, TextBoundary) — only tool events were factored out per #13. The transcript now receives ONLY: user-prompt echo, assistant Text deltas, [done]/[error]/[cancelled] terminal labels, and the post-Done Markdown render. All telemetry routes to a new Debug tab in the right column. 2. **Chrome no longer blue.** Textual's default Header / Footer / active-tab styling tints with `$primary` (Aurora blue under Australis), which read as garish on dark terminals. Header, Footer, and the TabbedContent tab strip get explicit `background: $surface` (Sea bright-black #373b46) so the chrome sits cool and unobtrusive against the Ice black background. ## Layout reshape ``` LEFT COLUMN (content only): RIGHT COLUMN (telemetry): transcript (RichLog, 1fr) thinking-current (Static, dock top) prompt (Input, dock bottom) TabbedContent: Tools (tool_start, tool_result) Debug (thinking, worker_phase, text_boundary) ``` The thinking-current live-preview Static moves from left → right column so the left column is genuinely content-only. Live thinking visibility now persists across tab switches (it docks above the TabbedContent, not inside any tab). ## Presenter routing (TuiPresenterState.render) Signature widens with `debug_log: RichLog`. Routing matrix: Text → log (transcript) Done / Error / Cancelled → log (transcript) [terminal labels] ToolStart / ToolResult → tools_log (Tools tab) Thinking (closed run) → debug_log (Debug tab) WorkerPhase → debug_log (Debug tab) TextBoundary → debug_log (Debug tab) Thinking (per-delta) → thinking_widget (live preview) INV-009 render-exception fallback preserves routing per event class (new INV-020) — ToolStart/Result falls back to tools_log; Thinking/WorkerPhase/TextBoundary to debug_log; everything else to log. ## Keybindings - Ctrl+1 → Tools tab (existing, unchanged) - Ctrl+2 → Debug tab (NEW) `pane-name` footer widget updates dynamically as the operator switches tabs ("Tools" ↔ "Debug"). This was previously deferred to "the multi-tab issue" per the Volva contract-review amendment; multi-tab now exists, so the dynamic update lands here. ## Contract amendments docs/contracts/issues/13.contract.md amended in-place: - INV-015 amended: transcript is content-only; telemetry routes to debug_log. Old routing (telemetry in transcript) retired under the no-backwards-compat rule. - INV-017 amended: thinking-current docks to right column (was left). - INV-019 new: two TabPanes (Tools + Debug), Ctrl+1/Ctrl+2 bindings, dynamic pane-name update. - INV-020 new: render-exception fallback preserves per-event-class routing. - Layout-spec snapshot ASCII diagram updated. Drift-check clean. ## Tests 239/239 GREEN (+3 new: debug_tab_exists, ctrl_2_activates_debug_tab, pane_name_updates_on_tab_switch). 6 existing tests adjusted for the new routing (test_thinking_closes_one_debuglog_entry, test_multiple_thinking_runs_each_get_debuglog_entry, test_render_exception_fallback, test_cancelled_mid_thinking_closes, test_worker_phase_demoted_to_debug_log, test_left_column_content_only). ruff clean. Live smoke against personal Worldtree: mimir KB-search turn populated tools_log with 11 lines of tool events (search_library + read_note); debug_log with 20 lines of worker_phase + thinking content; transcript stayed content-only with `❯ user-prompt` (Aurora bright-cyan) + assistant text deltas. Routing matrix holds end-to-end. (Diagnostic note: RichLog.lines is the rendered- output buffer; inactive TabPane content shows lines=0 until the tab activates and renders. Internal write store is correct — this is a Textual rendering quirk, not a routing bug.) Minor bump (v0.4.1 → v0.5.0) per SemVer etiquette: visible routing surface change for operators; transcript and Debug tab contents look different from yesterday's v0.4.1. |
||
|
|
2756f5f1dd |
style(tui): apply Australis theme to TUI chrome + widgets (v0.4.1)
Retheme the Textual TUI under the Australis Dark colorscheme (github.com/lkraven/australis) — Aurora blue/cyan/green primary, Ice neutrals (#222531 bg, #a9bcc3 fg, #cce7ec highlight), Sea darks for chrome separators, Dawn accents reserved for terminal- event labels (red/yellow). 16-color cool-tone palette with medium contrast. Mechanism: Textual `Theme` API. AUSTRALIS_THEME defined at module scope mapped to semantic tokens (primary/secondary/accent/success/ warning/error/foreground/background/surface/panel) plus a Sea variables block (`au-dark-30..60`, `au-bright-70/80/white`, plus Aurora bright variants). Both `RatatoskrApp` and `AgentPickerApp` register the theme in `__init__` and set `self.theme = "australis"`. Per-widget styling via DEFAULT_CSS theme variables (no hex sprinkled in CSS): - `#identity` → `$au-bright-blue` (Aurora bright-blue). - `#pane-name` → `$au-bright-cyan` (current-pane indicator). - `#hint` → `$au-dark-60` (subtle Ctrl-C state line). - `#prompt` border → `$panel` unfocused / `$primary` focused (focus highlight in Aurora blue). - `#left-column` gets a `border-right: solid $panel` separator. - `#thinking-current` color → `$au-dark-60` (matches demoted style). - `#transcript` + `#tools-log` background → `$background`. - Agent picker's highlighted ListItem → `$primary` bg + `$au-bright-white` fg. Rich Text styling (where Textual's theme system doesn't apply): - `_dim()` helper for demoted telemetry now uses explicit `#86929d` (Sea dark-60) instead of the terminal-dim filter `"dim"`. Renders consistently across emulators and stays anchored to the brand palette. - User-prompt echo (`❯ <content>` in transcript) wrapped in `RichText` styled with `#42dcd1` (Aurora bright-cyan) — calls out the operator's voice in the primary palette. CLI mode (`--send`) is unaffected by design — raw stdout has no theme concept. The retheme is TUI-only. 236/236 tests GREEN; ruff clean. Live smoke against personal Worldtree: `app.theme == "australis"`, all widget colors resolve to expected Australis hex values (identity → `#a4c4ff`, pane-name → `#42dcd1`, hint → `#86929d`). One existing test adjusted: `test_worker_phase_demoted` previously asserted `style == "dim"`; now asserts non-empty style (the demoted style is now an explicit Australis hex, not the terminal "dim" sentinel). Behavior-equivalent — the demotion intent is preserved. Patch bump (v0.4.0 → v0.4.1) per SemVer etiquette: cosmetic refinement of just-shipped surface, no public-API change. |
||
|
|
24e4371ec7 |
feat(tui): issue #13 — §5 layout reshape + Tools pane (v0.4.0)
Reshape the TUI from vertical-stack single-pane to Horizontal two-column with TabbedContent on the right; v1 has a single Tools tab that consumes ToolStart/ToolResult SSE events previously rendered inline in the transcript. Foundation for the rest of design-brief §5; subsequent panes (Persona/AdminEvents/BifrostState/ ServerLog) plug in as sibling TabPanes when their substrate blockers resolve. Three coupled pieces, all in-place amendments to issues #4 + #12: - **Layout**: compose() yields Horizontal#main-row containing Vertical#left-column (transcript + thinking-current + prompt) and Vertical#right-column (TabbedContent#side-panes with TabPane#tools-tab → RichLog#tools-log). Width split 2fr:1fr. CSS dock rules narrow to per-container scope so thinking-current toggling doesn't reflow the right column. - **Tools pane**: TuiPresenterState.render() signature widens with tools_log: RichLog. ToolStart/ToolResult route there per INV-014; every other event keeps its issue-#12 routing. Plain-label fallback under render-exception preserves routing (INV-009). - **Ctrl+1 binding + pane-name widget**: BINDINGS gains Binding("ctrl+1", "focus_tools") which programmatically sets TabbedContent.active; Textual's default preserves Input focus per INV-016 (test asserts; regression path documented). Static#pane-name in the footer renders "Tools" v1 (static — no tab-switch handler wiring lands in #13 per amendment-2 from Volva paraphrase review). CLI mode (--send) is unaffected by design per INV-018 — non- interactive, no tabs concept; CLI keeps inline tool-event rendering. Contract: docs/contracts/issues/13.contract.md (drift-check clean, two amendments applied from Volva contract-paraphrase pass). Tests: +9 net (TestLayoutShape × 7 + TestTuiPresenterState routing × 3, minus 1 deprecated test_tool_start_demoted superseded by test_tool_start_routes_to_tools_log). 236 total GREEN; ruff clean. Live smoke against personal Worldtree's mimir: tool-using turn (KB search) populated tools_log with tool_start + tool_result for search_library + read_note; transcript stayed chat-only with worker_phase + thinking. Routing-not-duplication confirmed end-to-end. |
||
|
|
d30be12deb |
feat(sessions,cli,tui): issue #8 — startup agent picker (v0.3.0)
Adds GET /agents fetch + ListView picker for bare `--new` (TUI mode without --agent). Three in-place amendments: - ratatoskr.sessions: new `list_agents()` + `AgentInfo` frozen dataclass with omit-when-null/empty defaults mirroring SessionInfo's INV-001/INV-002 origin-conditional pattern. Non-200 responses raise the existing SessionApiFailed (no new exception). - ratatoskr.cli: `_parse_args` softens `--agent` from absolute to mode-conditional — required for `--send --new`, optional for bare `--new`, forbidden with `--session` (unchanged INV-004). - ratatoskr.tui: new `AgentPickerApp(App[str | None])` — separate Textual App (not Screen-within-RatatoskrApp) so list_agents errors land on real stderr before any alt-screen opens (preserves issue #6's INV-001). `_resolve_then_run` gains a pre-create branch: fetch agents → empty list → exit 13; non-200 → exit 20; network error → exit 21; picker dismissed → exit 0; otherwise thread chosen agent_id into create_session. Contract: docs/contracts/issues/8.contract.md (drift-check clean). Tests: +18 (227 total, was 209). Live smoke against personal Worldtree (:8081) returned 12 agents; programmatic picker drive auto-picked lofn and created a real session with `end_user_id="ratatoskr-tui"`. |
||
|
|
a77a872810 |
snapshot: persistent-memory — v0.2.1 layout fix + §5 sequencing decision + Worldtree-stall diagnostic
Captures three things accumulated since the v0.2.0 snapshot in |
||
|
|
3b9c610587 |
feat(cli,tui): issue #12 — presenter contract semantics amendment (v0.2.0)
Replaces the stateless _render_event / _render_event_to_log helpers with stateful per-turn presenters (CliPresenterState / TuiPresenterState). Coalesces thinking-event deltas into a single growing display per run; demotes telemetry events with editorial hierarchy; formats duration + usage for human reading. Headline behavior change: a 50-token thinking phase now renders as ONE coalesced growing line in CLI (or one closed RichLog entry + per-delta live Static widget in TUI), not 50 lines of [thinking] spam. Editorial promotion line (issue #12 INV-002): - Load-bearing (no demotion prefix): Text, Done, Error, Cancelled - Demoted telemetry (`. ` ASCII prefix in CLI; dim `· ` in TUI): WorkerPhase, Thinking, TextBoundary, ToolStart, ToolResult Stateful coalescing: - Thinking deltas accumulate into thinking_buffer; first non-thinking event closes the run with a single \n boundary in CLI / one closed dim RichLog entry in TUI. - TUI adds a dedicated Static(id="thinking-current") widget that shows the last ~200 chars of the active run, mirroring per-delta updates. Two-views-of-thinking decoupling per INV-004: chronological RichLog + always-visible widget. - CLI INV-005: when stdout text was streamed mid-line, text_written_since_newline triggers a stdout flush + \n before the next stderr terminal label — guarantees [done] / [error] / [cancelled] land on their own line in a TTY without breaking pipe-to-file scripted consumers. Formatting helpers (issue #12 INV-006 / INV-007): - _format_duration_ms — autoscale `347ms` / `5.5s` / `1.2m` - _format_usage — natural-language `6756 in -> 126 out (6882 total, 0 cached)` with arrow="->" CLI / "→" TUI Cross-frontier design pass (eitri-smithy-dev, althing 01KSBE52YZR5E3SPTKA672JE43) returned 16-of-16 confirmed decisions + 4 material divergences applied: - ASCII `. ` prefix in CLI (`·` is U+00B7, not ASCII) - RichLog one-closed-entry-per-run + Static per-delta updates (not inline-mirror as initially proposed) - presenter-state object instead of pure-function rendering - Framed as "contract semantics amendment", not "polish" Volva paraphrase round (5 prose-precision fixes applied to 12.contract.md): INV-001 "growing display" semantics; single hide mechanism for the Static widget (Textual reactive `display: bool`); [render_error] security clause (type-only, no exception message); text_written_since_newline `\n`-terminated text corner case; [create_session] integration path (bypasses state.render — not an SSE Event variant). Volva code-review round (5 findings applied): - F1 drift: render-exception fallback now writes BOTH a plain-label fallback line for the original event AND the `[render_error] <type>` line (was missing the fallback half). - F2 drift: dim Rich style applied to all demoted-telemetry RichLog writes via `rich.text.Text(..., style="dim")` (was plain str). - F3 drift: belt-and-braces widget clear+hide on EVERY terminal event (Done/Error/Cancelled), even when thinking_open was False. - F4 precision: _format_usage gains PRE-001 assertion on the four expected usage keys. - F5 precision: _run_turn signature amended in issue #3 contract to document the new `state: CliPresenterState | None = None` test- injection kwarg. [create_session] lifecycle line demoted to `. create_session:` (written directly by _amain; bypasses state.render since it's not a wire-level SSE Event variant). Pre-amendment _render_event / _render_event_to_log and their test classes removed under the no-backwards-compat rule. Issues #3 and #4 contracts amended in-place: #3 (CliPresenterState CLASS + FN block + helper FN blocks + _run_turn signature + _amain create_session demotion); #4 (TuiPresenterState CLASS + FN block + compose Static widget + _stream_turn_worker state construction). 209 tests GREEN; ruff clean. Bumps v0.1.0 → v0.2.0 (minor — output shape change breaks pre-amendment grep patterns like `[thinking] '`; no public API surface change beyond the rendering contract). Persistent-memory commit-along: captures the issue #12 decision, forward direction (require end_user_id for every access — declined worldtree-dev's requires_end_user_id offer because we'll send it universally), and the Heimdall scope-model foot-gun note (the "per-Tier-1-agent scope add" diagnosis was a phantom ask resolved by worldtree-dev's correction; agent.call:* baseline covers all Tier 1). |
||
|
|
82821561e6 |
snapshot: persistent-memory — Heimdall scope-model foot-gun note (post-v0.1.0)
Captures the lesson from today's lofn-scope chase: "per-Tier-1-agent scope add" is a phantom ask. The `agent.call:*` (singular) baseline rule in config/policies.yaml covers ALL Tier 1 foundational agents (mimir, lofn, all Asgardians) for every authenticated tier; there is no per-agent grant in this path. The plural `agents.call:<owner>:<agent>` namespace is Tier 3 only (consumer-defined agents via POST /agents/define). Worldtree-dev shipped a corresponding doc fix (dd6e091) — new "Authorization model — agent invocation" section at docs/conversation-api-spec.md lines 58-114 + a heimdall.contract.md fix removing a misleading agent.call:mimir example. Don't ping infra-ops for "per-Tier-1-agent scope adds" again. Real future infra-ops asks remain: admin-tier key for the AdminEvents pane (admin.events.read scope, different tier) and Tier 3 custom-agent registration (POST /agents/define flow, different from scope-add). |
||
|
|
804c2df6eb |
feat(sessions,cli,tui): issues #5 + #6 + worldtree-dev consumer-API follow-up
Issue #6 (TUI startup error visibility): restructure run_tui lifecycle so pre-App.run() failures land on real stderr instead of getting eaten by the alt-screen teardown. New _resolve_then_run async helper opens the AsyncClient via async-with, does pre-flight session resolution, routes AgentNotFound / SessionApiFailed / network errors to sys.stderr (verbatim same labels + exit codes as cli._amain), then constructs RatatoskrApp with pre-resolved state and awaits app.run_async(). RatatoskrApp.__init__ signature widens to (args, *, session_id, agent_id, client) — all three required. on_mount narrows to identity-widget population; on_unmount becomes a no-op (client lifetime owned by run_tui's async-with). Issue #5 (--end-user-id for per-end-user agents): sessions.create_session gains keyword-only end_user_id kwarg with PRE-003 non-empty assertion; ParsedArgs.end_user_id field added (default None); --end-user-id flag with non-empty validation; _amain + _resolve_then_run thread it to their create_session calls. RATATOSKR_END_USER_ID env-var fallback (flag > env > None) per the post-2026-05-23 amendment; env.sh (gitignored) ships "ratatoskr-tui" as project-stable partition default. Worldtree-dev consumer-API follow-up (althing 01KSBARG2B8M): User-Agent header added (ratatoskr/<version> (vh@phasefinal.com), version pulled via importlib.metadata) to both AsyncClient constructions so server logs can distinguish ratatoskr traffic from other consumers. Volva code-review (2 rounds on #6) found 8 test-precision gaps + 1 PRE assertion drift, all Category 1 fixed: missing PRE-001 at _resolve_then_run entry; Rule separator assertions on markdown render; RichLog-write spy on empty submit; input-cleared + no-new-worker on cancelling busy; worker.cancel observation on three force-exit paths; on_unmount-no-close focused test (the prior client-lifetime test patched run_async so on_unmount was never exercised); happy --new resolve test verifying POST count + identity propagation. Issues #2/#3/#4/#5 contracts amended in-place to reflect: - create_session widened (PRE-003, body construction step, body shape POST) - ParsedArgs description + _parse_args STEPS + _amain create_session call + new TESTS for end_user_id + env-var fallback - _resolve_then_run STEPS + new TEST entries; on_mount narrowed; INV-007 amended for new client ownership - Post-#6 adjustment note on issue #5 (_resolve_then_run replaces on_mount as the threading site since #6 moved session resolution out of the alt-screen) 188 tests GREEN; ruff clean. Bumps to v0.1.0 — first minor release, the load-bearing reason is RatatoskrApp.__init__'s breaking signature change (additive end_user_id alone wouldn't have triggered a minor pre-v1.x). Files Gitea issues #9 (spec-pin refresh v0.19.0 → v0.22.1), #10 (track Worldtree #196 subject:{type,id} migration), #11 (AdminEvents pane auth prerequisite admin.events.read). Infra-ops pinged via althing for agents.call:lofn scope add (broker pattern; they forwarded to worldtree-dev because personal Worldtree exposes no public scope-mutation endpoint). |
||
|
|
c713208585 |
feat(sse_client,cli,tui): implement issue #7 — empty-data skip + MalformedSseData
Bundles initial TDD impl + Volva-code-review F1/F3 amendments. sse_client.py: - New MalformedSseData(raw) exception; truncates raw to 200 chars at __init__ (mirrors MalformedSseId.raw[:64] precedent). - _iter_events gains `if sse.data == '': continue` BEFORE _parse_sse_id. Empty-data frames are silently skipped per issue #7 INV-001 (keepalive semantics). Empty-data + bad-id is still a keepalive; intentional ordering, don't reorder. - _iter_events json.loads(sse.data) now wrapped — JSONDecodeError → MalformedSseData(raw=sse.data). cli.py: - Imports MalformedSseData; _run_turn ERROR_ROUTING gains the case → stderr `[malformed_sse_data] raw={exc.raw!r}` + exit 22 (protocol- failure bucket, same as MalformedSseId/TurnIdFlip). tui.py: - Imports MalformedSseData; _stream_turn_worker ERROR_ROUTING gains the case → transcript label; finally block restores state→idle per INV-008 (mid-session errors don't exit the app). Tests (6 new): - test_sse_client.py: empty_data_skipped (tracer — 4 frames in, 3 events out), malformed_data_raises, whitespace_data_raises, malformed_data_truncation, AND empty_data_skip_preserves_last_seen_sse_id (F1 from Volva code-review — drop-after-empty probes internal last_sse_id non-advancement via SseConnectionDropped.last_seen_sse_id). - test_cli.py: malformed_sse_data (tightened to assert exact `[malformed_sse_data] raw='not-json'` shape per F3), malformed_sse_data_truncation (5000-char payload — verifies truncation carries through presenter rendering, F3). - test_tui.py: malformed_sse_data_returns_to_idle (state→idle per INV-008; app does NOT exit). Smoke validation (2026-05-22): the original crashing prompt ("what about system 1 and system 2 framing?") now completes cleanly end-to-end. mimir streamed 3193 tokens (50 seconds, 374980-token context), `[done] turn_id=96 duration_ms=50436`. Empty-data frames somewhere in the stream silently skipped; no crash. 172/172 tests GREEN; ruff clean; all 5 issue contracts (#1, #3, #4, #5, #7) drift-check clean. Persistent-memory updated per the commit-along rule: status reflects v0+#7 milestone; new dated decisions for #5/#6/#7 filing + #7 implementation; foot-gun entry for unguarded json.loads(sse.data). |
||
|
|
6f7192f8db |
snapshot: persistent-memory after v0 milestone (4 issues landed)
Tighten Current state — drop file-by-file inventory (redundant with ls src/ratatoskr/ + each module's docstring + CLAUDE.md's architecture-map pointer to docs/architecture.md). Keep high-level status + next moves. Add two dated decisions capturing session-level lessons: - Volva paraphrase + code-review calibration consistent across all 4 issues (hit rates + recurring gap classes the post-TDD review catches) - Manual smoke is load-bearing — found a real defect (httpx 5s read timeout killing SSE) the mock-only test layer couldn't surface Add two foot-gun entries to Tried and abandoned: - RichLog markup=True silently strips [xxx] labels - Worker query_one timing trap (widened signature dodge → reverted; real fix was test-side pilot.pause) Also fold per-issue TDD detail entries into a single git-log pointer — the structured commit messages (9703eb2..61c3941) carry the per-issue trail; persistent-memory shouldn't duplicate. Net: 114 → 106 lines. Well under 300 soft cap; no archival needed. |
||
|
|
61c3941ec3 |
fix(client): disable SSE read timeout — caught by personal Worldtree smoke
First manual smoke against personal Worldtree (10.250.50.152:8081) produced httpx.ReadTimeout mid-stream after the worker_phase BuildingPrompt event. Root cause: httpx's default 5s read timeout killed the connection during mimir's thinking phase (LLM streaming has multi-second idle gaps between SSE events). Fix at the caller layer (where the AsyncClient is owned): - cli._amain and tui.on_mount now construct AsyncClient with timeout=httpx.Timeout(connect=10.0, read=None, write=10.0, pool=10.0). read=None disables the SSE-killing timeout; connect/write/pool keep modest timeouts so true network failures still surface promptly. Defense in depth in sse_client.stream_turn: - ERROR_ROUTING now also catches httpx.ReadTimeout (was just ReadError | RemoteProtocolError) and surfaces it as SseConnectionDropped, so if a caller misconfigures their client the failure is at least a named exception the presenters handle. Contract amendments (in-place): - Issue #1: new [compatibility] constraint documents the read=None recommendation; ERROR_ROUTING for stream_turn lists ReadTimeout alongside ReadError/RemoteProtocolError. - Issues #3 + #4: AsyncClient construction step now spells out the timeout shape explicitly. Smoke after fix: SSE stream consumed cleanly, agent responded, [done] turn_id=88 model=qwen3.6-35-a3b duration_ms=2351. Stdout-only (2>/dev/null) returned clean agent text + exit 0 — INV-002 stdout/stderr split holds end-to-end against real wire. Wire-compat envelope (personal v0.16.2 vs ratatoskr's v0.19.0 pin) confirmed. 164/164 tests GREEN; ruff clean; all three drift checks clean. Note: TUI mode not smoke-tested from this CC session (needs a TTY; operator-side check via `source env.sh && uv run ratatoskr --new --agent mimir`). |
||
|
|
942e33898c |
fix(tui): address Volva code-vs-contract drift (issue #4)
Volva code-review surfaced 8 findings against the TDD-passing
TUI shell. All 8 addressed.
Drift fixes (code):
- Primary: INV-002 + INV-003 require visible Footer-area rendering
of session-identity + Ctrl-C state hint. Implementation stored
the strings in `self.sub_title` (which lands in the Header, not
Footer) and `self.hint` (a plain attribute, never rendered). Fixed
by adding two `Static` widgets (id="identity" and id="hint") in
compose; the `_set_hint()` helper mirrors state into the widget on
every state transition. Same-model TDD missed this because tests
asserted internal state, not visible widget content.
- Reverted `_stream_turn_worker(content, log)` to single-param
`(content)` per the contract FN signature. The widened signature
was a TDD-time workaround for a NoMatches-during-worker
execution; root cause was test timing (added `await pilot.pause()`
before the polling loop in `_submit_and_wait`).
- Restored `exclusive=True` on `self.run_worker(...)` per the
contract STEP 6 spec.
- Added missing `isinstance(args, ParsedArgs)` PRE assertion to
`run_tui`. Required hoisting `from ratatoskr.cli import
ParsedArgs` out of TYPE_CHECKING — runtime import is fine (no
circular dependency: cli lazy-imports tui inside main; tui
imports cli unconditionally at module load).
- Added missing union-type PRE assertion to `_render_event_to_log`.
Contract amendments (precision):
- COMPOSE shape: RichLog `markup=False, highlight=False` (was True,
True). Explanatory comment in-line: bracketed labels like
[cancel_failed] would otherwise be interpreted+stripped as Rich
style spans; the post-Done Markdown rendering still works via
Markdown() Renderable.
- INV-002 reworded: identity rendered via dedicated
Static(id="identity") widget composed adjacent to Footer (Textual's
built-in Footer renders BINDINGS descriptions; a sibling Static
carries custom content in the same visual region).
- on_mount POST-003 amended to allow `agent_id is None` when
--session is used without --agent (matches INV-002 carve-out;
GET /sessions/{id} agent lookup is out of scope for this shell).
- run_tui happy_returns_zero_on_quit test description clarified:
App.run() is sync and can't be driven by Pilot, so run_tui's
wrapping behavior is tested via monkeypatch; the piloted Ctrl-D
exit path is covered separately by TestActionQuit.
Test fixes:
- footer_identity_visible_first_frame, footer_hint_flips_to_cancel,
streaming_first_ctrl_c_cancels: now query the Static(#identity) /
Static(#hint) widgets via `widget.render()` instead of asserting
on `app.sub_title` / `app.hint` internal state. The internal
state still exists (mirror), but the load-bearing assertion is
on visible widget content.
Meta-note from Volva: "TDD pass caught most stream/session/error
mechanics, but tested internal state where the contract required
visible Footer behavior, so same-model TDD would plausibly miss the
primary drift." Calibration shape continues across all four issues:
the post-TDD cross-model review consistently catches assert-boundary
+ observability-shape gaps the test-author's hypotheses don't cover
(#1: 4 findings, #2: 3, #3: 5, #4: 8).
164/164 tests GREEN; ruff clean; both contract drift checks clean.
|
||
|
|
dd89239c34 |
feat(tui): implement issue #4 contract via TDD; amend cli for TUI dispatch
47 contract-listed tests authored + GREEN (43 tui + 4 issue-#3 amendments). 164/164 tests GREEN suite-wide; ruff clean. Vertical-slice ordering: _render_event_to_log → _cancel_via_sse → CLI amendments → RatatoskrApp class + on_mount + on_unmount → on_input_submitted → _stream_turn_worker → action_interrupt + action_quit → run_tui. Two in-flight contract amendments caught during TDD: - PRE-002 of run_tui was `(args.session_id is None) != args.new` — backwards (fails when --session is set + new=False). Corrected to `bool(args.session_id) != bool(args.new)`. - RichLog created with markup=False (contract drafted markup=True). Rich interprets `[xxx]` as style markup and strips it, which would break every labeled stderr-style line ([cancel_failed], [done], [error], etc.). The post-Done Markdown rendering still works because rich.markdown.Markdown is a Renderable and doesn't need widget-level markup. Implementation notes: - _stream_turn_worker takes the log widget as a parameter passed from on_input_submitted. Querying #transcript from inside a Textual worker context fails with NoMatches; capturing the reference once at handler-time and threading it through the worker sidesteps the issue. - _spy_writes(monkeypatch) test helper records every RichLog.write call. RichLog's `.lines` Strip buffer isn't populated synchronously after .write() returns, which makes post-app-shutdown inspection unreliable; a write-spy gives deterministic verification. - SIGINT-mid-stream tests use custom httpx.AsyncByteStream subclasses with asyncio.Event gates to make timing deterministic without sleep-based polling — the cancel-respx-mock sets the gate event when its endpoint is observed, releasing the next SSE chunk. - _submit_and_wait test helper needs `await pilot.pause()` BEFORE the polling loop so the Input.Submitted message has a chance to dispatch. Discovered via debug-print trace; tracked in the test helper. CLI amendments (per issue #4 in-place amendment of #3 contract): - ParsedArgs.send_content: str | None (was str) - ParsedArgs.raw: bool added - _parse_args: --send default=None; empty-string still rejected; --raw added - main: branches on args.send_content — None → lazy `from ratatoskr.tui import run_tui` + run_tui(args); else asyncio.run(_amain(args)). Lazy import preserves issue #3 INV-001. Persistent-memory updated per the commit-along rule: tui module landed, recent-decisions entries for #4 (contract + Volva + TDD), next natural moves rotated to Volva code-review + manual smoke against the personal Worldtree (key landed in env.sh per infra-ops's earlier delivery). |
||
|
|
9717fb80e2 |
fix(cli): address Volva code-vs-contract drift (issue #3)
Volva code-review surfaced 5 findings against the TDD-passing implementation; all 5 addressed. Drift fixes (code): - Add `assert argv is None or all(isinstance(a, str) for a in argv)` at both `main` and `_parse_args` entry points (PRE-001 was unenforced). - `main` now catches `SystemExit` and returns `exc.code` verbatim — argparse's --help (SystemExit(0)) was escaping through main as an unhandled exception. Contract amended in-place to spell out the SystemExit-from-argparse-clean-exits passthrough in both `main` and `_parse_args` ERROR_ROUTING. New `help_exits_cleanly` test added per the contract amendment. - Add the PRE-001 union-type assert at `_render_event` entry — unmatched Event variants would have silently no-op'd. - `_run_turn` now awaits `cancel_task` in the `finally` block before returning. Under fast-stream + slow-cancel scenarios the `[cancel_failed]` line could miss being written before _run_turn returns, AND _amain could close the AsyncClient while the cancel POST was still in flight. `_cancel_and_log` swallows all errors per INV-009 so the await is safe. Test gap fix: - New `_FlushCountingIO` subclass counts flush() calls; `test_text_to_stdout_only` and `test_done_writes_newline_and_label` now assert `flush_count == 1` to verify INV-010 (per-chunk flush). Previously the tests would have passed even with flush removed. Meta-note carried in persistent-memory: TDD caught central behavior (stdout/stderr routing, exit-code mapping, create-session ordering, SIGINT idempotence); the cross-model code review consistently catches assert-boundary + observability-shape gaps across all three issues (#1: 4 findings, #2: 3 findings, #3: 5 findings). 118/118 tests GREEN; ruff clean; drift check clean. |
||
|
|
db27774c51 |
feat(cli): implement issue #3 contract via TDD
54 contract-listed tests authored + GREEN per the vertical-slice ordering (_parse_args → _render_event → _cancel_and_log → _run_turn → _amain → main). 117/117 tests GREEN suite-wide; ruff clean. The _run_turn race-loop is the load-bearing piece. Per iteration, the await on the next event is raced against sigint_event.wait() when NOT cancelling. Once SIGINT fires (with last_turn_id known), _cancel_and_log is spawned, cancelling=True flips, and subsequent iterations skip wait()-task creation entirely — the bug Volva flagged in contract review would otherwise busy-wake on the already-set event each iteration. Implementation notes: - _UsageErrorParser subclasses argparse.ArgumentParser and overrides error() to raise _ArgparseError instead of calling sys.exit; _parse_args catches and re-raises as UsageError per the contract's ERROR_ROUTING. - _GatedStream test helper (custom httpx.AsyncByteStream that pauses on asyncio.Event entries) makes SIGINT-mid-stream tests deterministic without sleep-based timing — gates release via side-channels (the cancel-mock sets an event when its endpoint is observed). - _sse_resp test helper wraps respx Response with the text/event-stream content-type, dedupes the boilerplate across the 13 _run_turn tests. - Strong-ref cancel_task local in _run_turn holds the fire-and-forget cancel task to suppress RUF006 / asyncio GC warning. One in-flight contract amendment during TDD: no_busy_loop_after_cancel test description originally said "exactly ONE wait()-shaped task" but the natural race-loop shape produces 2 (iter 1 raced w/ text, iter 2 raced w/ sigint → flipped cancelling; iter 3+ skipped). Amended to "TWO total wait() coroutines" with rationale; the busy-loop check is preserved (iter 3+ MUST skip). Persistent-memory updated per the commit-along rule: new module landed, recent-decisions log entries for #3 (contract + Volva paraphrase + TDD), next natural moves rotated to /volva-code-review on the implementation. |
||
|
|
d6f9327ec1 |
fix(sessions): address Volva code-vs-contract drift (issue #2)
Volva's code-spec review (thread 01KS4EKVKKGF) surfaced three findings
on the TDD-passing sessions module. All three addressed; one carries
a collateral contract amendment to keep INV-002 truthful.
1) drift: archived=item.get("archived", False) returned None for an
explicit "archived": null in the response. dict.get(k, default) only
fires the default when the key is absent — it does NOT default for
explicit-null values. The dataclass type is `bool` (not `bool | None`)
and INV-002 says explicit-null → False; the .get() form silently
violated both. Fixed: archived=item.get("archived") or False
(handles absent, null, false, and true cleanly).
INV-002 wording was the source of the bug — I introduced the
mis-spelled form during the Volva amendment round. Updated to spell
out the .get(default) foot-gun explicitly so future readers (and
future paraphrase rounds) don't fall back to the broken pattern.
2) test-gap: no test exercised explicit-null archived/tags. The
_list_item() helper had its own defaulting layer (tags=None →
["work"]) so a happy path test couldn't catch the underlying drift.
Added test_explicit_null_list_defaults using a raw dict to bypass
the helper. Catches the drift directly.
3) precision: message_count=body.get("message_count") could silently
default to None while POST-003 required it non-None. INV-001 prose
literally said "body['message_count']" (bracket access) so the
STEP 5 .get() was the contract's own internal inconsistency.
Aligned the code to bracket access (matches sibling required
fields like session_id) and amended STEP 5 + INV-001 to spell out
the strict semantics explicitly.
Volva's meta-note: "modest weight" — TDD caught the main surface;
this round caught a narrow Python .get() semantics edge that no
human reading would have spotted without explicit-null priors.
Still pulls real weight: that's the kind of bug that ships and
shows up months later when a server starts emitting null where
it used to omit a field.
63 tests GREEN (42 sse_client + 20 sessions + 1 boundary).
Ruff clean. Drift check still GREEN against the pinned issue body.
|
||
|
|
4ba143c563 |
feat(sessions): implement issue #2 contract via TDD
Implements docs/contracts/issues/2.contract.md. Two functions (create_session, list_sessions), two frozen dataclasses (SessionInfo, SessionPage), three exception types (AgentNotFound, InvalidCursor, SessionApiFailed). 19 contract-listed tests cover every TESTS: entry verbatim per the tracer-bullet vertical-slice ordering. SessionInfo uses one shape across both endpoints with origin- conditional defaults per INV-001 (create) and INV-002 (list). create- origin always sets list-only fields to (name=None, archived=False, tags=[]); list-origin reads them from the response item with absent/null treated as those same defaults — keeps the dataclass uniform without forcing callers to handle two types. Spotted an internal-inconsistency in the contract at TDD start — POST-003 and happy_create's test description still said "archived is None, tags is None" while the freshly-applied Volva amendment had moved INV-001 to (archived=False, tags=[]). Fixed in-place before writing any tests so the spec stayed coherent. SessionApiFailed.body truncates to <= 1024 bytes at construction, matching the SseConnectFailed / CancelFailed precedent from issue #1. No code shared with sse_client.py (convention-dependency only per issue #2's dependencies: block). 62 tests GREEN total (42 sse_client + 19 sessions + 1 boundary smoke). Ruff clean. No refactor pass — the two functions are ~25 LOC each with distinct error-routing branches that don't naturally share more than they already do. |
||
|
|
a6e6c1bbd8 |
contract(issue#2): amend per Volva paraphrase — defaults, query, metadata
Volva's contract paraphrase round (thread 01KS4DTCW8CV) surfaced five
ambiguities; three are real contract-text gaps and addressed here.
1) tags/archived/name defaulting was inconsistent across prose, INV-002,
and STEP 6. open_questions said "defaulting to sensible None/empty",
INV-002 said "populated from the response item shape", STEP 6 said
`item.get("tags", []) if "tags" in item else None` (which collapses
absent and explicit-null into the same None branch while letting an
explicit [] pass through). Tightened to: tags is always list[str]
defaulting to [] for absent/null/empty in list items; archived is
always bool defaulting to False; name remains str | None (the only
field where None is a meaningful value). create_session always sets
list-only fields to their fixed defaults (name=None, archived=False,
tags=[]) instead of None to keep the dataclass shape uniform.
2) The include_archived_query test said "URL has no include_archived
param OR explicit false". STEP 2 prescribes "ADD include_archived='true'
iff include_archived" — the OR-clause weakened the test against the
prescribed behavior. Tightened to: default (include_archived=False)
asserts NO include_archived param at all, not an explicit false.
5) metadata's populated semantics: INV-001 said "populated from the
201 response", STEP 5 said body.get("metadata", {}) — two valid
readings (trust the spec vs defensive default). Aligned to the
defensive shape: INV-001 + INV-002 now explicitly state "defaults
to {} when absent" as spec-drift tolerance.
Volva flags #3 (exception .body sensitivity — truncation reduces
size not sensitivity) and #4 (assert for runtime validation — Python
-O disables) reviewed and kept as-is. Both are intentional carryovers
from issue #1's precedent: exception .body is for caller debugging
bound to 1024 bytes (caller's responsibility to not log raw); assert
chosen for fast-path validation, trading -O robustness for normal-mode
speed.
Drift check still clean — amendments don't touch the pinned issue
body, so prd: hashes remain valid.
|
||
|
|
9df9bb8757 |
contract(issue#2): scaffold ratatoskr.sessions — create + list
Issue #2: ratatoskr.sessions covers the session-lifecycle endpoints needed by --send --new (POST /sessions) and the eventual TUI startup picker (GET /sessions). Two FN blocks (create_session, list_sessions) plus two shared frozen dataclasses (SessionInfo, SessionPage). complexity=low; estimated 150 LOC. Bundles both endpoints in one contract because they share the response- envelope shape — SessionInfo carries the union of POST-response fields (message_count) and list-item fields (name, archived, tags), with the origin-conditional fields defaulting to None. INV-001/002 spell out which fields come from which source so callers can rely on the discriminator. INV-006 refuses out-of-range limit (< 1 or > 200) client-side: spec §GET /sessions says the server returns 422; the client checks first so a 422 from this endpoint indicates server-side spec drift, not a client bug. INV-003 codifies the opaque-cursor discipline (spec §Pagination: "Cursors are opaque to clients — do not parse or construct them."). list_sessions threads next_cursor verbatim; never base64-decodes. Exception .body truncation to [:1024] inherited from issue #1's SseConnectFailed/CancelFailed precedent. First contract in this repo to carry a ## Out of scope H2. Future /volva-code-review consults will auto-resolve that section instead of needing --out-of-scope overrides. Six explicit exclusions: Bifrost binding (Worldtree #160), ephemeral/Saga sessions, single- session fetch, PATCH/DELETE mutation, history pagination, transparent multi-page iteration. Server retry/backoff is caller's policy. dependencies: lists issue #1 as a convention-dependency only — no code import; same API-consumption posture (caller-owned httpx client, async-native, frozen dataclasses, no Worldtree-source imports). prd: pinned to issue #2 body SHA-256 01fbbd52b6d90eb0 at 2026-05-21T04:45:06+00:00; scripts/contract_drift_check.py returns clean. |