19b499ab500faf0ad57b21b43c69f6884020ab2f
8
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
19b499ab50 |
feat(tts): migrate off Zonos to chatterbox-fast; drop affect, hold English
Repoint the TTS client from the Zonos gateway (:8890 /v1/audio/speech) to
chatterbox-fast (:8197 /tts — bespoke non-OpenAI {text,voice,format,stream}
schema, no auth, 24kHz, infra-ops-verified). tts.py stays the single swap seam.
Dropped, no backward-compat (pre-v1):
- Affect (DEC-7): the Turbo checkpoint has no emotion knob, so PadState,
EmotionDials, pad_to_dials, the /api/tts p/a fields, and the browser pad
argument are deleted. Voice is now flat.
- Client-side chunking (DEC-10): chatterbox has no per-synth cap and chunks
internally, so chunk_text/tts_stream_long/_pcm_after_header are deleted; a
single tts_stream call voices a whole turn, the mid-stream yielded_any degrade
folded into it.
- Language pin (DEC-9): no language field; re-purposed to sampling curbs (below).
Fixed / added:
- Browser Web Audio sample rate 44100 -> 24000 (the chatterbox rate).
- Default voice Cora -> glados_25s; donut registered lowercase at /refs/donut.wav.
- English-drift curb: Turbo is multilingual-leaky and wanders off English on a
long generation (the gateway scheduler ratchets chunk size unbounded). Tighten
sampling in gateway_body: top_k 1000->80, top_p 0.95->0.85, temperature
0.8->0.5. These reduce drift probability; the guaranteed fix is a server-side
max-chunk cap (infra-ops, greenlit).
- OOM guard (DEC-9a): a long generation can OOM the shared 3090, returning 200
with a 0-byte body; /api/tts surfaces an empty 200 as 503 rather than
committing silent audio.
Contract donut_voiced_interview.contract.md amended: migration banner, DEC-1/3/8
amended, DEC-7/9/10 retired with historical notes, DEC-9a added.
Tests rewritten to the new wire; 520 green. Live-smoked against the gateway
(24kHz synth + endpoint proxy + web console). persistent-memory.md committed
alongside (commit-along).
|
||
|
|
d59f907962 |
feat(tts): pin English, stream long turns via chunking, dialogue-only Donut
TTS fixes + hardening for the Donut voiced interview. Feature: - gibberish -> pin `language: "en-us"` on every gateway call (DEC-9); the multilingual model drifted into other-language phonemes without it. - truncation -> the Zonos model hard-caps one synthesis at 6144 tokens / 71.2s (infra-ops). Chunk client-side (paragraph-first, greedy to ~75% of cap for prosody; sentence/clause fallback) and concatenate the int16 PCM behind ONE WAV header (DEC-10). /api/tts becomes POST so a long turn rides the body, not a length-capped URL (DEC-10a). - persona -> dialogue-only rewrite (no asterisk RP beats -- they were being voiced as gibberish) + always consult the native `reference_knowledge` tool before answering (retires the stale kb_bridge references). Pushed live to ratatoskr:donut. Heid code-review + bug-hunt hardening (4-arm panels, triaged): - untrusted /api/tts body fields degrade, never 500: huge-int PAD (OverflowError), non-str agent_id (unhashable .get), lone surrogates (utf-8 encode), whitespace-only text. - serialize lock + client released on every peek escape (cancel / InvalidURL) -- previously a permanent deadlock. - a mid-stream drop after a committed 200 degrades (keeps what played), never raises into the response; a non-WAV 200 body is rejected (RIFF sniff + bounded header scan) instead of decoded as garbage. 546 tests green; long-form live-verified (106.6s, one header). Contract brought canonical (DEC-9/10, FN chunk_text/tts_stream_long, POST endpoint, INV-TTS-4 logging scope, FN pad_to_dials domain). reference_knowledge empty-recall root-caused to a Worldtree wing-misfile (escalated to worldtree-dev; not ratatoskr code). |
||
|
|
9041f1f402 |
fix: Web Audio streaming playback — fixes Safari NotSupportedError
Operator confirmed the "TTS blocked" was NotSupportedError on Safari — WebKit refuses a streaming 0xFFFFFFFF-length WAV via <audio src> (can't compute duration/seek), exactly as infra-ops warned. Replaced the <audio src> playback with a Web Audio path that works in all engines: - speakOnDone: fetch the chunked /api/tts stream, skip the WAV header to the data chunk, decode int16 LE PCM -> Float32, and schedule the samples GAPLESSLY into an AudioContext as they arrive (BufferSource per chunk, playAt += buf.duration). Progressive, TTFA ~0.5s. Decoding the raw PCM ourselves sidesteps every WAV-container quirk. - unlock: an AudioContext starts suspended; Safari + Chrome need resume() from a user gesture. _unlockTtsAudio() now resumes the ctx on the first interaction anywhere + toggle-on + submit, so it's running before the ~15s-delayed speak-on-done. - cancelTts: aborts the fetch + stops all scheduled BufferSource nodes. Validated in Chromium (Playwright, strict autoplay): 43 nodes scheduled, 5.1s of PCM decoded, ctx "running" 6.5s post-gesture, zero errors. Headless WebKit can't launch here (missing system libs — an infra-ops install), so the operator's live Safari is the final check; the code is standard Web Audio Safari has supported for years. Contract FN client:speakOnDone updated (Web Audio; the Safari NotSupportedError reason). |
||
|
|
7856ec5438 |
feat: stream Donut TTS play-as-it-arrives + autoplay unlock (supersedes buffered)
Operator: play-as-it-arrives, don't wait for the whole clip. infra-ops confirmed the
Zonos gateway ALREADY streams (chunked int16 WAV, TTFB ~0.44s vs ~7s total; placeholder
0xFFFFFFFF sizes are DESIGNED for progressive <audio src>). The buffering was entirely
in our proxy, and the _finalize_wav_header rewrite (
|
||
|
|
09e425787b |
refactor: retire the KB-recall bridge — WT #383 native reference_knowledge (b167)
Worldtree #383 shipped native Tier-3 reference_knowledge (v1.0.0b167, live on :8081 + demo): every Tier-3 agent context now carries the tool automatically, with evidence packets (note_id + path provenance, confidence bucket) and a server-side grounding rule. That supersedes the interim consumer-side memory_context pinning bridge (slice 3), so it is deleted per its INV-KB-1 retire seam. Removed: - src/ratatoskr/kb_bridge.py + tests/test_kb_bridge.py (the whole module). - server.py: the pin_kb_context import + the single turn-path call-site (reverted to the pre-bridge wt.stream_turn call), the SSE keepalive that only covered the consult delay, and the bridge-only agent_id plumbing (TurnHandle.agent_id + the submit read). - index.html: agent_id dropped from the turn POST body. - test_web_server.py: TestKbBridgeWiring (tested the removed call-site). Kept: - wt.stream_turn's memory_context param (inert SDK-parity passthrough; worldtree-dev concurred it stays) + its forwarding tests. - the non-str content 400 guard (general input hygiene, not bridge-specific). Retirement LIVE-VERIFIED before deletion: a Donut session on :8081/b167 carries builtin_tools=['reference_knowledge']; she called it and grounded in the DCC Collapse content fully in-voice, degrading gracefully on absent content. 524 green. Contract marks slice-3 RETIRED (historical record retained). #383 closed. |
||
|
|
eb0767e96d |
fix: heid-code-review fixups — donut voiced-interview slices 2+3
Triaged the heid-code-review panel (3 arms; reconciled against
|
||
|
|
71689142bc |
feat: Donut voiced-interview slice-3 — retire-ready KB-recall bridge
Grounds the interview character in the ingested corpus while she stays in-voice. Tier-3 agents are tool-less by design in v1, so this is the consumer-side workaround (DEC-6, worldtree-dev ruling): per opted-in interview turn, ratatoskr consults Mimir out-of-band, extracts the passages, and pins them as memory_context on the character's turn. She frames the pinned corpus as her own memory. - src/ratatoskr/kb_bridge.py (new, RETIRE-READY): pin_kb_context — THE single seam (INV-KB-1). Allowlist-gated (INV-KB-4: ratatoskr:donut only), hard-timeout-bounded, degrades to [] on any failure/timeout/empty (INV-KB-3, never raises; CancelledError propagates). Imports nothing from the SDK-adapter / TTS core. aclosing() closes the SDK stream deterministically on the DoneEvent break. - wt.stream_turn: memory_context passthrough (defaults None — inert for every other caller and for the bridge's own retirement). Seam-review catch: the contract's original touch list undercounted wt.py by one file (recorded in the contract). - web/server.py: TurnHandle.agent_id + the single pin_kb_context call-site on the turn path; the browser now sends agent_id so the allowlist can gate. - web/static/index.html: the turn POST carries agent_id. Consult prompt tuned live: "search_library EXACTLY ONCE, no read_note" converges Mimir in ~3-15s (the softer "do one search" phrasing looped past 25s on conversational questions). TDD: 12 kb_bridge unit tests + wt memory_context forwarding + 2 server wiring tests (531 green). Live-smoked on :8081/b128: pin_kb_context grounds in the DCC corpus (real excerpts, <20s) and Donut answers in-voice; degrades cleanly on a slow consult. KNOWN LIMIT surfaced (not a bridge defect): DCC's fiction index is weak (failed backfill, a worldtree-dev item), so grounding is opportunistic — the bridge's real payoff is a corpus the model does not already know. Per docs/contracts/donut_voiced_interview.contract.md (slice 3 of 3). |
||
|
|
3e912b13b3 |
feat: Donut voiced-interview slice-1 (contract + persona + define) + /snapshot
Slice 1 of the auto-TTS/voiced-KB-character build (operator ask "add auto-tts to the web gui"): the donut_voiced_interview contract (validated), the Princess Donut persona (corpus-grounded from a Mimir DCC pull), and ratatoskr:donut defined on :8081 (server-side; in the picker). Slices 2 (Zonos auto-TTS) + 3 (retire-ready KB-bridge) are TO BUILD. Snapshot captures the full build state + design (Zonos gateway :8890, voice "donut" registered, affect-driven emotion dials; the worldtree-dev-ruled consumer-side retrieval + memory_context pinning bridge, retire-ready) for the post-clear resume, plus the arcs since v0.22.0 (SDK 1.1.2 repin, bifrost 1.1.5, canonical sync, release-only versioning, the Sindra saga + local-index schema-burial foot-gun, the Mimir #382 reference-consumer finding). Handoff at /tmp/ratatoskr-dev-handoff.md. Release-only cadence: no tag. |