catalog(chatterbox): route to /tts, expose emotion levers + sane defaults
Chatterbox was producing poor output because the catalog pointed at the thin OpenAI /v1/audio/speech endpoint, which exposes none of Resemble's emotion/ pacing knobs — and the devnen server's shipped default exaggeration is 1.3 (tuned for its theatrical demo presets), which over-acts. Re-point to the wrapper's richer /tts and expose the real control surface (exaggeration, cfg_weight, temperature, speed_factor, seed, voice_mode), mirroring the sibling dia stack (same devnen author). Defaults sourced live 2026-06-01: exaggeration + cfg_weight = 0.5 (Resemble README 'works well for most prompts'), temperature 0.8 / speed 1.0 / seed 0 (server generation_ defaults). The shipped 1.3 exaggeration is deliberately NOT adopted. Voices: expose the 28 built-in predefined voices via /get_predefined_voices (default Emily.wav, the server default_voice_id) + clone via /get_reference_ files — replacing the wrong 'OpenAI aliases only' claim. Corrected seedable: false -> true (/tts has seed) and image_tag_mutable -> true (:latest). Bumped service version 1 -> 2 (breaking field-shape change); status down -> ready (live + healthy). catalog_version unchanged (no new field types).
This commit is contained in:
+150
-33
@@ -199,68 +199,185 @@ services:
|
|||||||
renderer ships. Once present, flip to status: ready.
|
renderer ships. Once present, flip to status: ready.
|
||||||
|
|
||||||
- id: chatterbox
|
- id: chatterbox
|
||||||
name: Chatterbox Turbo TTS
|
name: Chatterbox TTS
|
||||||
description: >
|
description: >
|
||||||
Resemble AI's low-latency English TTS (350M, ~75ms TTFB, 6× realtime).
|
Resemble AI's low-latency English TTS (Chatterbox-Turbo, 350M, ~75ms TTFB,
|
||||||
Zero-shot voice cloning from ~5s reference. 9 paralinguistic tags.
|
6× realtime). 28 built-in predefined voices + zero-shot cloning from a
|
||||||
|
5–30s reference. Inline paralinguistic tags, plus Resemble's signature
|
||||||
|
exaggeration / cfg_weight emotion + pacing control.
|
||||||
category: tts
|
category: tts
|
||||||
version: 1
|
version: 2
|
||||||
status: down
|
status: ready
|
||||||
host: irv-ml1
|
host: irv-ml1
|
||||||
lifecycle:
|
lifecycle:
|
||||||
stack: chatterbox
|
stack: chatterbox
|
||||||
vram_gb: 4
|
vram_gb: 4
|
||||||
gpu_device_id: 0
|
gpu_device_id: 0
|
||||||
endpoint: http://10.100.79.3:8196/v1/audio/speech
|
endpoint: http://10.100.79.3:8196/tts
|
||||||
method: POST
|
method: POST
|
||||||
content_type: application/json
|
content_type: application/json
|
||||||
model:
|
model:
|
||||||
id: ResembleAI/chatterbox-turbo
|
id: ResembleAI/chatterbox-turbo
|
||||||
revision: null
|
revision: null
|
||||||
image: devnen/Chatterbox-TTS-Server:latest
|
image: devnen/Chatterbox-TTS-Server:latest
|
||||||
|
section_groups:
|
||||||
|
- id: basic
|
||||||
|
label: Text & voice
|
||||||
|
- id: sampling
|
||||||
|
label: Expression & sampling
|
||||||
|
hint: Resemble's neutral defaults (exaggeration 0.5 / cfg_weight 0.5). Raise exaggeration or lower cfg_weight for drama.
|
||||||
|
- id: advanced
|
||||||
|
label: Advanced
|
||||||
fields:
|
fields:
|
||||||
- name: input
|
- name: text
|
||||||
type: textarea
|
type: textarea
|
||||||
label: Text (with optional [tags])
|
label: Text (with optional [tags])
|
||||||
|
section: basic
|
||||||
required: true
|
required: true
|
||||||
max_length: 5000
|
max_length: 5000
|
||||||
description: >
|
description: >
|
||||||
Inline tags: [laugh] [chuckle] [sigh] [gasp] [cough] [clear throat]
|
Inline paralinguistic tags honored by Turbo: [laugh] [chuckle] [sigh]
|
||||||
[sniff] [groan] [shush]. Turbo loses base-Chatterbox's exaggeration knob.
|
[gasp] [cough] [clear throat] [sniff] [groan] [shush]. Best results
|
||||||
- name: model
|
when a physical tag is paired with surrounding emotional context.
|
||||||
|
- name: voice_mode
|
||||||
type: select
|
type: select
|
||||||
options: [chatterbox-turbo]
|
label: Voice mode
|
||||||
default: chatterbox-turbo
|
section: basic
|
||||||
- name: voice
|
options: [predefined, clone]
|
||||||
type: select
|
default: predefined
|
||||||
label: Voice
|
|
||||||
default: alloy
|
|
||||||
description: >
|
description: >
|
||||||
Built-in OpenAI-compat aliases (alloy, echo, fable, onyx, nova, shimmer).
|
`predefined` -> a built-in voice (predefined_voice_id below).
|
||||||
Cloned: 5–15s WAV files in /worktank/chatterbox/reference_audio/.
|
`clone` -> a reference clip (reference_audio_filename). predefined is
|
||||||
- name: response_format
|
the out-of-box default; the empty/"undefined" case is avoided by
|
||||||
|
defaulting the voice below.
|
||||||
|
- name: predefined_voice_id
|
||||||
type: select
|
type: select
|
||||||
options: [wav, opus, aac, flac, pcm_s16]
|
label: Voice (built-in)
|
||||||
|
section: basic
|
||||||
|
optional: true
|
||||||
|
default: "Emily.wav"
|
||||||
|
source_url: http://10.100.79.3:8196/get_predefined_voices
|
||||||
|
source_jsonpath: $[*].filename
|
||||||
|
description: >
|
||||||
|
Required when voice_mode=predefined. 28 built-in voices staged in the
|
||||||
|
devnen image (Abigail, Adrian, Alexander, Alice, Austin, Axel, Connor,
|
||||||
|
Cora, Elena, Eli, Emily, Everett, Gabriel, Gianna, Henry, Ian, Jade,
|
||||||
|
Jeremiah, Jordan, Julian, Layla, Leonardo, Michael, Miles, Olivia,
|
||||||
|
Ryan, Taylor, Thomas — each <name>.wav). Default Emily.wav is the
|
||||||
|
server's own default_voice_id. Verified live via /get_predefined_voices.
|
||||||
|
- name: reference_audio_filename
|
||||||
|
type: select
|
||||||
|
label: Voice (clone reference)
|
||||||
|
section: basic
|
||||||
|
optional: true
|
||||||
|
source_url: http://10.100.79.3:8196/get_reference_files
|
||||||
|
source_jsonpath: $[*]
|
||||||
|
description: >
|
||||||
|
Required when voice_mode=clone. 5–30s clean WAV (16 kHz+ mono) under
|
||||||
|
/worktank/chatterbox/reference_audio/; upload via the server's
|
||||||
|
/upload_reference. Match the clip's language to `language` to avoid
|
||||||
|
accent transfer (or set cfg_weight=0).
|
||||||
|
- name: exaggeration
|
||||||
|
type: slider
|
||||||
|
section: sampling
|
||||||
|
min: 0.25
|
||||||
|
max: 2.0
|
||||||
|
step: 0.05
|
||||||
|
default: 0.5
|
||||||
|
description: >
|
||||||
|
Emotional intensity. Resemble's docs: 0.5 "works well for most prompts
|
||||||
|
across all languages"; ~0.7+ for dramatic delivery (which also speeds
|
||||||
|
speech up). NOTE: the devnen server *ships* 1.3 (tuned for its
|
||||||
|
theatrical demo presets) — 0.5 is the general-use value and the catalog
|
||||||
|
default; the shipped 1.3 is the likely cause of over-acted/unstable output.
|
||||||
|
- name: cfg_weight
|
||||||
|
type: slider
|
||||||
|
section: sampling
|
||||||
|
min: 0.0
|
||||||
|
max: 1.0
|
||||||
|
step: 0.05
|
||||||
|
default: 0.5
|
||||||
|
description: >
|
||||||
|
Pacing / prompt adherence (Resemble default 0.5). Lower to ~0.3 to
|
||||||
|
slow delivery, for fast/intense reference speakers, or alongside a
|
||||||
|
raised exaggeration for drama; 0 effectively disables guidance (useful
|
||||||
|
to reduce reference-accent transfer).
|
||||||
|
- name: temperature
|
||||||
|
type: slider
|
||||||
|
section: sampling
|
||||||
|
min: 0.05
|
||||||
|
max: 2.0
|
||||||
|
step: 0.05
|
||||||
|
default: 0.8
|
||||||
|
description: Sampling temperature; lower = steadier. Server + Resemble default 0.8.
|
||||||
|
- name: speed_factor
|
||||||
|
type: slider
|
||||||
|
section: sampling
|
||||||
|
min: 0.5
|
||||||
|
max: 2.0
|
||||||
|
step: 0.05
|
||||||
|
default: 1.0
|
||||||
|
description: Post-hoc playback speed. Server default 1.0.
|
||||||
|
- name: seed
|
||||||
|
type: number
|
||||||
|
section: sampling
|
||||||
|
required: false
|
||||||
|
default: 0
|
||||||
|
description: 0 = random; a fixed integer repeats the same take.
|
||||||
|
- name: output_format
|
||||||
|
type: select
|
||||||
|
section: basic
|
||||||
|
options: [wav, opus, mp3]
|
||||||
default: wav
|
default: wav
|
||||||
- name: stream
|
description: 24 kHz. Live-verified enum (wav/opus/mp3).
|
||||||
|
- name: language
|
||||||
|
type: text
|
||||||
|
section: advanced
|
||||||
|
required: false
|
||||||
|
default: en
|
||||||
|
description: >
|
||||||
|
Language override. Base Turbo is English; the multilingual variant
|
||||||
|
(23 languages, via the stack .env) honors other codes. Leave `en`.
|
||||||
|
- name: split_text
|
||||||
type: bool
|
type: bool
|
||||||
default: false
|
section: advanced
|
||||||
|
default: true
|
||||||
|
description: Auto-split long text into chunks.
|
||||||
|
- name: chunk_size
|
||||||
|
type: slider
|
||||||
|
section: advanced
|
||||||
|
min: 100
|
||||||
|
max: 1000
|
||||||
|
step: 10
|
||||||
|
default: 120
|
||||||
|
description: Target chunk length in chars when splitting (server default 120).
|
||||||
response:
|
response:
|
||||||
type: audio
|
type: audio
|
||||||
mime_from_field: response_format
|
mime_from_field: output_format
|
||||||
reproducibility:
|
reproducibility:
|
||||||
seedable: false
|
seedable: true
|
||||||
deterministic: true
|
deterministic: false
|
||||||
|
seed_field: seed
|
||||||
notes: >
|
notes: >
|
||||||
No seed. Wrapper repo updates ~weekly; pin SHA in .env. PerTh watermark
|
/tts exposes `seed` (0=random); a fixed seed + identical params repeats a
|
||||||
unconditionally applied (Resemble policy).
|
take. Temperature-sampled → not guaranteed byte-exact, and Resemble's
|
||||||
|
PerTh watermark is applied unconditionally. (Prior catalog claimed no
|
||||||
|
seed support — corrected against the live OpenAPI 2026-06-01.)
|
||||||
estimated_latency:
|
estimated_latency:
|
||||||
cold_start_s: 3
|
cold_start_s: 3
|
||||||
warm_per_unit: "~75ms TTFB, 6× realtime"
|
warm_per_unit: "~75ms TTFB, 6× realtime"
|
||||||
license: MIT
|
license: MIT
|
||||||
notes: |
|
notes: |
|
||||||
Python 3.10 only (wrapper hardcoding).
|
Routes to the devnen wrapper's richer /tts (full control surface:
|
||||||
Multilingual variant (23 languages) also available via .env.
|
exaggeration / cfg_weight / temperature / speed_factor / seed / voice_mode)
|
||||||
|
instead of the thin OpenAI /v1/audio/speech, which exposes NONE of the
|
||||||
|
emotion knobs — that omission was why prior output was poor. Same wrapper
|
||||||
|
author as the `dia` stack; identical predefined/clone voice model.
|
||||||
|
Defaults sourced from Resemble's README (exaggeration + cfg_weight = 0.5)
|
||||||
|
and the server's generation_defaults (temperature 0.8, speed 1.0, seed 0),
|
||||||
|
read live 2026-06-01; the server's shipped exaggeration 1.3 is demo-tuned
|
||||||
|
and deliberately NOT adopted. Python 3.10 only (wrapper hardcoding);
|
||||||
|
multilingual (23-language) variant available via the stack .env.
|
||||||
|
|
||||||
- id: index-tts
|
- id: index-tts
|
||||||
name: IndexTTS-2
|
name: IndexTTS-2
|
||||||
@@ -1724,10 +1841,10 @@ reproducibility_audit:
|
|||||||
image_tag_mutable: true
|
image_tag_mutable: true
|
||||||
notes: "Same image as kokoro proper; same mutability story. Response carries timestamps."
|
notes: "Same image as kokoro proper; same mutability story. Response carries timestamps."
|
||||||
- service: chatterbox
|
- service: chatterbox
|
||||||
seedable: false
|
seedable: true
|
||||||
model_deterministic: true
|
model_deterministic: false
|
||||||
image_tag_mutable: false
|
image_tag_mutable: true
|
||||||
notes: "PerTh watermark unconditional (Resemble policy)."
|
notes: "/tts exposes seed (0=random); temperature-sampled, not byte-exact. PerTh watermark unconditional (Resemble policy). image :latest is mutable — pin a digest/SHA for true repro."
|
||||||
- service: index-tts
|
- service: index-tts
|
||||||
seedable: false
|
seedable: false
|
||||||
model_deterministic: true
|
model_deterministic: true
|
||||||
|
|||||||
Reference in New Issue
Block a user