feat(tts): migrate RP-surface TTS chatterbox-fast → dots-tts
Swap the voice synthesis backend from chatterbox-fast (:8197 bespoke /tts)
to dots-tts (rednote-hilab dots.tts-soar, :8198 OpenAI-shaped
/v1/audio/speech), operator-directed after an A/B win. tts.py stays the
single swap seam.
- gateway body OpenAI-shaped: {input, voice, response_format, stream}
(was chatterbox {text, voice, format, stream})
- sample rate 24000 -> 48000 Hz (browser Web Audio SR)
- default voice glados_25s -> glados; donut voice carries over
- serialized single-consumer (satisfied by the existing DEC-5 lock)
- affect stays dropped (dots has no emotion knob, same as chatterbox)
DOTS_TTS_URL replaces CHATTERBOX_TTS_URL; RATATOSKR_TTS_URL override
unchanged. chatterbox-fast :8197 kept up as rollback. Contract amended
(donut_voiced_interview.contract.md). Live-verified end-to-end on :8765
(RIFF/WAVE 48kHz mono s16le through /api/tts). 520 tests green.
This commit is contained in:
@@ -4,8 +4,9 @@ module: "ratatoskr.web.tts_kb"
|
||||
purpose: >
|
||||
A voiced, corpus-grounded Tier-3 interview character in the ratatoskr web
|
||||
console. Two capabilities plus one character: (a) auto-TTS via the
|
||||
chatterbox-fast gateway, spoken on SSE `done` (migrated off Zonos 2026-08-07;
|
||||
no affect modulation — chatterbox Turbo has no emotion knob); (b) a
|
||||
dots-tts gateway, spoken on SSE `done` (migrated Zonos→chatterbox-fast
|
||||
2026-08-07, then chatterbox-fast→dots-tts 2026-08-10; no affect modulation —
|
||||
dots has no emotion knob); (b) a
|
||||
consumer-side KB-retrieval + `memory_context` pinning
|
||||
BRIDGE that grounds the character's recall in the ingested corpus while she
|
||||
stays in-voice; (c) Princess Donut (Dungeon Crawler Carl) as the first
|
||||
@@ -20,14 +21,14 @@ scope: >
|
||||
consult (an existing agent turn).
|
||||
touches:
|
||||
- src/ratatoskr/web/server.py # /api/tts route + the retrieval-pinning seam on the turn path
|
||||
- src/ratatoskr/web/static/index.html # speak-on-done playback, 🔊 toggle, <audio> sink; turn POST carries agent_id
|
||||
- src/ratatoskr/web/static/index.html # speak-on-done playback (SR 48000), 🔊 toggle, <audio> sink; turn POST carries agent_id
|
||||
- src/ratatoskr/web/entrypoint.py # RATATOSKR_TTS_URL override (the tts swap seam)
|
||||
- src/ratatoskr/tts.py # chatterbox-fast gateway client (was Zonos + PAD->dial; migrated 2026-08-07)
|
||||
- src/ratatoskr/tts.py # dots-tts gateway client (Zonos→chatterbox 2026-08-07→dots 2026-08-10; OpenAI-shaped)
|
||||
- src/ratatoskr/kb_bridge.py # NEW, RETIRE-READY — consumer-side retrieval + memory_context pinning
|
||||
- src/ratatoskr/wt.py # stream_turn gains a memory_context passthrough (seam-review: the contract's original touch list undercounted this by one file; the param defaults None so the bridge's RETIREMENT stays inert — deleting kb_bridge.py + the one call-site leaves wt.stream_turn's SDK-parity param harmless)
|
||||
- docs/characters/donut.md # NEW — Princess Donut persona (content; the tier3 define source)
|
||||
depends_on:
|
||||
- "chatterbox-fast gateway: POST http://10.100.79.3:8197/tts (infra-ops; WG-internal, no auth; bespoke non-OpenAI schema {text,voice,format,stream}; streaming placeholder-header wav @ 24000 Hz; no per-synth cap; NO affect controls; verified 2026-08-07 against image local/chatterbox-fast:v1)"
|
||||
- "dots-tts gateway: POST http://10.100.79.3:8198/v1/audio/speech (infra-ops; WG-internal, no auth; OpenAI-shaped schema {input,voice,response_format,stream}; streaming placeholder-header wav @ 48000 Hz mono s16le; dots streams a whole turn from one call; SERIALIZED single-consumer; zero-shot voice cloning, voices donut/glados/emmie/miranda; NO affect controls; verified 2026-08-10 against dots-studio/dots.tts-soar). chatterbox-fast :8197 kept up as rollback."
|
||||
- "Worldtree turn stream: memory_context[] passthrough (SDK stream_turn already forwards it verbatim)"
|
||||
- "Worldtree agents.define (Tier-3) for Donut; Mimir (search_kb) for the out-of-band retrieval consult"
|
||||
used_by:
|
||||
@@ -41,6 +42,28 @@ confidence: 0.8
|
||||
|
||||
# Contract: Donut voiced interview (auto-TTS + KB-recall bridge)
|
||||
|
||||
> **⚠ TTS MIGRATED chatterbox-fast → dots-tts 2026-08-10 (operator-directed, after an
|
||||
> A/B win).** The synthesis backend moved from chatterbox-fast (:8197 bespoke `/tts`)
|
||||
> to dots-tts (rednote-hilab `dots.tts-soar`, :8198 OpenAI-shaped `/v1/audio/speech`),
|
||||
> verified live. Four deltas; everything else (the streaming placeholder-header WAV
|
||||
> shape, the browser Web-Audio PCM decode path, POST `/api/tts`, the serialize lock,
|
||||
> INV-TTS-1..4) is UNCHANGED:
|
||||
> - **Gateway body OpenAI-shaped.** `{input, voice, response_format:"wav", stream:true}`
|
||||
> — `input` (not chatterbox's `text`), `response_format` (not `format`). Closer to the
|
||||
> Zonos-era client. `tts.py` stays the single swap seam (DEC-1), now translating the
|
||||
> OpenAI schema; `DOTS_TTS_URL` replaces `CHATTERBOX_TTS_URL`.
|
||||
> - **Sample rate 24000 → 48000 Hz.** The browser Web Audio decode MUST use 48000 or the
|
||||
> voice plays ~2× too fast (`index.html` `SR = 48000`).
|
||||
> - **Default voice `glados_25s` → `glados`.** dots voices are donut/glados/emmie/miranda
|
||||
> (GET /v1/voices); `donut` carries over. Non-interview agents fall to `glados`.
|
||||
> - **Serialized single-consumer.** dots renders one generation at a time — satisfied by
|
||||
> the existing DEC-5 lock (no code change). If concurrent streams are ever needed,
|
||||
> infra-ops escalates the backend behind the same API (client unchanged).
|
||||
> Affect stays dropped (DEC-7): dots has no emotion knob, same as chatterbox — NOT a fresh
|
||||
> regression. chatterbox-fast :8197 is kept up as the rollback until dots is confirmed
|
||||
> solid. The 2026-08-07 chatterbox banner + DEC-7/9/9a/10 below are retained as historical
|
||||
> record.
|
||||
|
||||
> **⚠ TTS MIGRATED OFF ZONOS → chatterbox-fast 2026-08-07 (operator-directed).**
|
||||
> Slice 2's synthesis backend moved from the Zonos gateway (:8890
|
||||
> `/v1/audio/speech`) to chatterbox-fast (:8197 `/tts`). Three architecture deltas,
|
||||
@@ -147,13 +170,14 @@ each independently shippable. Slice order is chosen for fastest visible result.
|
||||
record of the Zonos build. (Original: map live PAD from the `affect_update` SSE →
|
||||
Zonos `emotion_valence`/`emotion_arousal`, reframing the feature as voice
|
||||
OBSERVABILITY. The observability framing dies with the knob.)
|
||||
- **DEC-8 — voice: custom "donut" is REGISTERED (amended 2026-08-07 for chatterbox).**
|
||||
chatterbox voices are `*.wav` reference clips in `/refs` (GET /voices lists the
|
||||
stems). infra-ops registered `/refs/donut.wav` (the same reference clip behind the
|
||||
Zonos Donut voice) at operator direction, so `_TTS_VOICE_MAP` maps
|
||||
`ratatoskr:donut → "donut"` directly. NOTE the case: chatterbox wants lowercase
|
||||
`"donut"` (Zonos used `"Donut"`). Non-interview agents fall to the chatterbox
|
||||
default `"glados_25s"` (was Zonos `"Cora"`, which does not exist on chatterbox).
|
||||
- **DEC-8 — voice: custom "donut" is REGISTERED (amended 2026-08-10 for dots).**
|
||||
dots clones a voice server-side from a reference clip + transcript; the client just
|
||||
passes a voice NAME (GET /v1/voices lists them: donut/glados/emmie/miranda). The
|
||||
`donut` voice carries over from chatterbox, so `_TTS_VOICE_MAP` maps
|
||||
`ratatoskr:donut → "donut"` unchanged. NOTE the case: lowercase `"donut"` (Zonos used
|
||||
`"Donut"`). Non-interview agents fall to the dots default `"glados"` (was chatterbox
|
||||
`"glados_25s"` / Zonos `"Cora"`, neither of which exists on dots). New voices are a
|
||||
one-line request to infra-ops (derived from the canonical voice corpus).
|
||||
- **DEC-9 — hold English: RESOLVED SERVER-SIDE 2026-08-07 (client sends full text, default
|
||||
sampling).** The Zonos `language:"en-us"` pin is dropped — chatterbox has no `language` field.
|
||||
The long-turn garble ("swaps to German halfway through") went through two WRONG hypotheses
|
||||
@@ -241,19 +265,19 @@ each independently shippable. Slice order is chosen for fastest visible result.
|
||||
|
||||
## FN blocks
|
||||
|
||||
### FN tts_stream (the sole synthesis primitive — DEC-2 streaming; amended 2026-08-07)
|
||||
### FN tts_stream (the sole synthesis primitive — DEC-2 streaming; amended 2026-08-10 dots)
|
||||
```
|
||||
tts_stream(text, *, voice, client: httpx.AsyncClient, url=CHATTERBOX_TTS_URL) -> AsyncIterator[bytes]
|
||||
tts_stream(text, *, voice, client: httpx.AsyncClient, url=DOTS_TTS_URL) -> AsyncIterator[bytes]
|
||||
# Open the gateway's CHUNKED stream (client.stream("POST", url, json=gateway_body(text, voice))) and
|
||||
# YIELD wav chunks as they synthesize. Pass through verbatim — never buffer, never rewrite the placeholder
|
||||
# header. chatterbox chunks arbitrary-length text INTERNALLY (no per-synth cap, DEC-10 RETIRED), so this
|
||||
# SINGLE call voices a whole turn — no client-side chunk-and-concatenate wrapper.
|
||||
# gateway_body(text, voice) = {text, voice, format:"wav", stream:true}. Full text, DEFAULT sampling — the
|
||||
# :v2 server caps each generation at 250 chars, which fixes the long-turn tail garble (DEC-9). NO dials,
|
||||
# NO language, NO client sampling curbs (a curb was counterproductive — it pulled the garble onset earlier).
|
||||
precondition: text non-empty. Voice membership in GET /voices is GATEWAY-enforced, not client-asserted.
|
||||
# header. dots streams a whole turn from this SINGLE call (DEC-10 RETIRED) — no client-side
|
||||
# chunk-and-concatenate wrapper.
|
||||
# gateway_body(text, voice) = {input, voice, response_format:"wav", stream:true} (OpenAI-shaped: `input`
|
||||
# not `text`, `response_format` not `format`). Full text, DEFAULT sampling. NO dials, NO language,
|
||||
# NO client sampling curbs.
|
||||
precondition: text non-empty. Voice membership in GET /v1/voices is GATEWAY-enforced, not client-asserted.
|
||||
postcondition: yields the gateway's chunked int16 streaming WAV bytes unmodified (0xFFFFFFFF placeholder
|
||||
sizes intact), one leading header then s16le PCM @ 24000 Hz to EOF.
|
||||
sizes intact), one leading header then mono s16le PCM @ 48000 Hz to EOF.
|
||||
error (the yielded_any pivot, folded in from the retired tts_stream_long):
|
||||
- a non-200 OPEN or a connect/transport failure BEFORE the first byte -> TtsUnavailable (so the endpoint
|
||||
peek can still return 503; nothing committed yet).
|
||||
@@ -365,8 +389,8 @@ on SSE `done`:
|
||||
cancelTts() # INV-TTS-3: abort fetch + stop scheduled nodes
|
||||
POST /api/tts {text (sliced to the 8000 cap), agent_id?} -> reader # DEC-10a: POST body. NO p/a (DEC-7 retired).
|
||||
loop: read chunk -> skip ONE WAV header up to the data chunk (bounded 64KiB) -> int16 LE PCM -> Float32 ->
|
||||
AudioBuffer(sampleRate=24000) -> BufferSource.start(playAt) GAPLESSLY -> playAt += buf.duration
|
||||
# SR = 24000 (chatterbox; was 44100 for Zonos — MUST match or the voice plays ~1.8x too fast). TTFA ~0.5s.
|
||||
AudioBuffer(sampleRate=48000) -> BufferSource.start(playAt) GAPLESSLY -> playAt += buf.duration
|
||||
# SR = 48000 (dots; was 24000 for chatterbox — MUST match or the voice plays ~2x too fast). TTFA ~0.5s.
|
||||
first scheduled node -> "▶ voiced". HARD failure (non-OK HTTP, or 64KiB with no WAV header) -> ticker + skip;
|
||||
ABORT/cancel (INV-TTS-3 new-turn) + bare network error -> SILENT skip (INV-TTS-4, cancel is not a failure)
|
||||
|
||||
|
||||
Reference in New Issue
Block a user