Files
esh-pfi-infrastructure/docs/asset-engine/services.yaml
T
vh c5bbb90980 fix(fish-s2): reference_id was a silent no-op — populate per-voice dirs + guard the regression
reference_id=<name> resolves against the DIRECTORY references/<name>/
(audio + same-basename .lab), not a flat references/<name>.wav. Voices
were staged flat with the per-name dirs left empty, so every
reference_id resolved to nothing and Fish fell back to its default
speaker — every dropdown voice produced byte-identical audio (proven:
Abigail == Imogen == no-ref, same text+seed). This was the real "no
accent" root cause, independent of the asset-engine "undefined" select
bug.

Server fix (applied to irv-ml1): populated references/<name>/<name>.wav
+ <name>.lab for all 32 voices; re-test confirms Imogen/Eleanor/
Beatrice/Abigail/no-ref now all distinct.

Durable hardening + record correction:
- playbook: normalize-layout step (flat <name>.wav -> nested dir, cp -u
  idempotent, when-gated on count mismatch) + an A/B verify gate that
  hard-fails the deploy if two reference_ids yield identical output.
- services.yaml: correct the reference_id resolution doc (dir + .lab,
  not flat wav).
- README + persistent-memory: correct the "reference_id-by-name is THE
  working path, verified" claim — it was a no-op until this fix; the
  prior ECAPA 0.79 result came through the inline base64 path.
2026-06-01 16:42:30 -07:00

2069 lines
72 KiB
YAML
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# services.yaml — canonical catalog of inference services on irv-ml1
#
# THIS IS A FIRST-CLASS CONTRACT. External consumers depend on the
# shape and contents of this file. See CATALOG-CONTRACT.md alongside
# for: schema reference, change-management policy, sync workflow for
# downstream consumers, and known consumers list.
#
# Drives form generators that auto-render UIs against the inference
# services. Each entry produces one form (from `fields:`) and one
# response renderer (dispatched on `response.type`). The field-type
# vocabulary and response-type vocabulary are closed sets — see the
# JSON Schema at services.schema.json or the contract doc.
#
# `catalog_version`: bump when the SCHEMA changes (a field type added,
# a required field removed, etc). Per-service `version:` bumps when a
# specific service's parameter shape changes. Both let downstream
# consumers detect drift.
catalog_version: 1
services:
- id: kokoro
name: Kokoro 82M TTS
description: >
Lowest-latency English TTS (82M params, fp16). ~300ms TTFA, 35100x realtime,
60+ built-in voices in 8 languages. Voice mixing via inline weights.
OpenAI-compatible /v1/audio/speech.
category: tts
version: 3
host: irv-ml1
lifecycle:
stack: kokoro
vram_gb: 2
gpu_device_id: null
endpoint: http://10.100.79.3:8193/v1/audio/speech
method: POST
content_type: application/json
model:
id: hexgrad/Kokoro-82M
revision: null
image: ghcr.io/remsky/kokoro-fastapi-gpu:v0.2.4-master
fields:
- name: input
type: textarea
label: Text
required: true
max_length: 5000
- name: voice
type: select
label: Voice
source_url: http://10.100.79.3:8193/v1/audio/voices
source_jsonpath: $.voices[*]
default: af_heart
description: >
60+ built-in voices. Default `af_heart` is the only A-rated voice per
upstream VOICES.md (also the kokoro-fastapi container's own default).
Custom blends: af_bella(2)+af_aoede(1) syntax for weighted mixing.
Persistent custom voices via playbooks/blend-kokoro-voice.yaml.
- name: response_format
type: select
options: [wav, mp3, opus, flac, pcm]
default: wav
- name: speed
type: slider
label: Speed
required: false
min: 0.25
max: 4.0
step: 0.05
default: 1.0
description: Playback speed multiplier. 1.0 is natural.
- name: volume_multiplier
type: slider
label: Volume
required: false
min: 0.5
max: 2.0
step: 0.05
default: 1.0
description: >
Output gain. Useful for level-matching against music beds.
Upstream is unbounded; UI bounds (0.52.0) are a sensible default —
override the slider client-side if you need extremes.
- name: lang_code
type: text
label: Language code
required: false
description: >
Optional language hint (e.g. "a" for American English). When unset,
kokoro infers from the first letter of the voice name. Set when text
contains code-switching kokoro guesses wrong on.
- name: stream
type: bool
default: false
description: Phrase-by-phrase streaming via chunked HTTP.
response:
type: audio
mime_from_field: response_format
reproducibility:
seedable: false
deterministic: true
notes: >
No seed parameter. Model fully deterministic (no sampling); identical
params always produce identical bytes. Image tag is mutable; pin to
digest for v3 reproducibility audit.
estimated_latency:
cold_start_s: 2
warm_per_unit: "~300ms TTFA, 35100x realtime"
license: Apache-2.0
notes: |
Voice mixing: voice="name1(w1)+name2(w2)+..." normalizes weights.
Custom voices persist at /worktank/kokoro/user_voices (bind-mounted).
v2 (2026-05-11): added speed, volume_multiplier, lang_code per
kokoro's OpenAPI surface.
- id: kokoro-captioned
name: Kokoro Captioned Speech
description: >
Kokoro TTS with word-level timestamps returned alongside the audio.
For subtitle generation and video sync. Same model as `kokoro`;
separate catalog entry because the wire shape is a JSON envelope
carrying base64-encoded audio plus a structured timestamps array.
category: tts
version: 2
status: ready
host: irv-ml1
endpoint: http://10.100.79.3:8193/dev/captioned_speech
method: POST
content_type: application/json
model:
id: hexgrad/Kokoro-82M
revision: null
image: ghcr.io/remsky/kokoro-fastapi-gpu:v0.2.4-master
fields:
- name: input
type: textarea
label: Text
required: true
max_length: 5000
- name: voice
type: select
label: Voice
source_url: http://10.100.79.3:8193/v1/audio/voices
source_jsonpath: $.voices[*]
default: af_heart
- name: response_format
type: select
options: [wav, mp3, opus, flac, pcm]
default: wav
- name: speed
type: slider
label: Speed
required: false
min: 0.25
max: 4.0
step: 0.05
default: 1.0
- name: lang_code
type: text
label: Language code
required: false
response:
# Stays in the closed type vocabulary: from a renderer-dispatch
# standpoint this IS audio. The audio_field/audio_format_field/
# timestamps_field decomposition tells consumers how to extract
# those parts from the JSON envelope wire shape — added to the
# catalog schema in 2026-05 specifically to support response
# shapes like this one without extending the type vocab.
type: audio
audio_field: audio
audio_format_field: audio_format
timestamps_field: timestamps
reproducibility:
seedable: false
deterministic: true
notes: >
Same determinism story as kokoro proper. Verified wire shape
(2026-05-11 against live API):
{
"audio": "<base64-encoded bytes in response_format>",
"audio_format": "audio/wav" (or matching response_format),
"timestamps": [{"word": str, "start_time": float, "end_time": float}, ...]
}
Consumer base64-decodes `audio` to play; `timestamps` drives
subtitle/karaoke UI. The response decomposition fields above
encode this so the renderer doesn't need per-service-id branches.
estimated_latency:
cold_start_s: 2
warm_per_unit: "~same as kokoro proper, plus minor overhead for timestamp emission"
license: Apache-2.0
notes: |
`return_timestamps` and `stream` upstream params deliberately
omitted from the catalog: timestamps must be on for this endpoint
to be meaningful, and streaming + JSON-with-base64 don't compose.
`download_format` / `return_download_link` skipped — same as kokoro
proper.
status: experimental until the consumer's audio-with-timestamps
renderer ships. Once present, flip to status: ready.
- id: chatterbox
name: Chatterbox TTS
description: >
Resemble AI's low-latency English TTS (Chatterbox-Turbo, 350M, ~75ms TTFB,
6× realtime). 28 built-in predefined voices + zero-shot cloning from a
530s reference. Inline paralinguistic tags, plus Resemble's signature
exaggeration / cfg_weight emotion + pacing control.
category: tts
version: 2
status: ready
host: irv-ml1
lifecycle:
stack: chatterbox
vram_gb: 4
gpu_device_id: 0
endpoint: http://10.100.79.3:8196/tts
method: POST
content_type: application/json
model:
id: ResembleAI/chatterbox-turbo
revision: null
image: devnen/Chatterbox-TTS-Server:latest
section_groups:
- id: basic
label: Text & voice
- id: sampling
label: Expression & sampling
hint: Resemble's neutral defaults (exaggeration 0.5 / cfg_weight 0.5). Raise exaggeration or lower cfg_weight for drama.
- id: advanced
label: Advanced
fields:
- name: text
type: textarea
label: Text (with optional [tags])
section: basic
required: true
max_length: 5000
description: >
Inline paralinguistic tags honored by Turbo: [laugh] [chuckle] [sigh]
[gasp] [cough] [clear throat] [sniff] [groan] [shush]. Best results
when a physical tag is paired with surrounding emotional context.
- name: voice_mode
type: select
label: Voice mode
section: basic
options: [predefined, clone]
default: predefined
description: >
`predefined` -> a built-in voice (predefined_voice_id below).
`clone` -> a reference clip (reference_audio_filename). predefined is
the out-of-box default; the empty/"undefined" case is avoided by
defaulting the voice below.
- name: predefined_voice_id
type: select
label: Voice (built-in)
section: basic
optional: true
default: "Emily.wav"
source_url: http://10.100.79.3:8196/get_predefined_voices
source_jsonpath: $[*].filename
description: >
Required when voice_mode=predefined. 28 built-in voices staged in the
devnen image (Abigail, Adrian, Alexander, Alice, Austin, Axel, Connor,
Cora, Elena, Eli, Emily, Everett, Gabriel, Gianna, Henry, Ian, Jade,
Jeremiah, Jordan, Julian, Layla, Leonardo, Michael, Miles, Olivia,
Ryan, Taylor, Thomas — each <name>.wav). Default Emily.wav is the
server's own default_voice_id. Verified live via /get_predefined_voices.
- name: reference_audio_filename
type: select
label: Voice (clone reference)
section: basic
optional: true
source_url: http://10.100.79.3:8196/get_reference_files
source_jsonpath: $[*]
description: >
Required when voice_mode=clone. 530s clean WAV (16 kHz+ mono) under
/worktank/chatterbox/reference_audio/; upload via the server's
/upload_reference. Match the clip's language to `language` to avoid
accent transfer (or set cfg_weight=0).
- name: exaggeration
type: slider
section: sampling
min: 0.25
max: 2.0
step: 0.05
default: 0.5
description: >
Emotional intensity. Resemble's docs: 0.5 "works well for most prompts
across all languages"; ~0.7+ for dramatic delivery (which also speeds
speech up). NOTE: the devnen server *ships* 1.3 (tuned for its
theatrical demo presets) — 0.5 is the general-use value and the catalog
default; the shipped 1.3 is the likely cause of over-acted/unstable output.
- name: cfg_weight
type: slider
section: sampling
min: 0.0
max: 1.0
step: 0.05
default: 0.5
description: >
Pacing / prompt adherence (Resemble default 0.5). Lower to ~0.3 to
slow delivery, for fast/intense reference speakers, or alongside a
raised exaggeration for drama; 0 effectively disables guidance (useful
to reduce reference-accent transfer).
- name: temperature
type: slider
section: sampling
min: 0.05
max: 2.0
step: 0.05
default: 0.8
description: Sampling temperature; lower = steadier. Server + Resemble default 0.8.
- name: speed_factor
type: slider
section: sampling
min: 0.5
max: 2.0
step: 0.05
default: 1.0
description: Post-hoc playback speed. Server default 1.0.
- name: seed
type: number
section: sampling
required: false
default: 0
description: 0 = random; a fixed integer repeats the same take.
- name: output_format
type: select
section: basic
options: [wav, opus, mp3]
default: wav
description: 24 kHz. Live-verified enum (wav/opus/mp3).
- name: language
type: text
section: advanced
required: false
default: en
description: >
Language override. Base Turbo is English; the multilingual variant
(23 languages, via the stack .env) honors other codes. Leave `en`.
- name: split_text
type: bool
section: advanced
default: true
description: Auto-split long text into chunks.
- name: chunk_size
type: slider
section: advanced
min: 100
max: 1000
step: 10
default: 120
description: Target chunk length in chars when splitting (server default 120).
response:
type: audio
mime_from_field: output_format
reproducibility:
seedable: true
deterministic: false
seed_field: seed
notes: >
/tts exposes `seed` (0=random); a fixed seed + identical params repeats a
take. Temperature-sampled → not guaranteed byte-exact, and Resemble's
PerTh watermark is applied unconditionally. (Prior catalog claimed no
seed support — corrected against the live OpenAPI 2026-06-01.)
estimated_latency:
cold_start_s: 3
warm_per_unit: "~75ms TTFB, 6× realtime"
license: MIT
notes: |
Routes to the devnen wrapper's richer /tts (full control surface:
exaggeration / cfg_weight / temperature / speed_factor / seed / voice_mode)
instead of the thin OpenAI /v1/audio/speech, which exposes NONE of the
emotion knobs — that omission was why prior output was poor. Same wrapper
author as the `dia` stack; identical predefined/clone voice model.
Defaults sourced from Resemble's README (exaggeration + cfg_weight = 0.5)
and the server's generation_defaults (temperature 0.8, speed 1.0, seed 0),
read live 2026-06-01; the server's shipped exaggeration 1.3 is demo-tuned
and deliberately NOT adopted. Python 3.10 only (wrapper hardcoding);
multilingual (23-language) variant available via the stack .env.
- id: index-tts
name: IndexTTS-2
description: >
Bilibili's emotion-controllable zero-shot TTS. Disentangled emotion control
(timbre from one reference, emotion from another). 22050 Hz fixed output.
category: tts
version: 1
status: down
host: irv-ml1
lifecycle:
stack: index-tts
vram_gb: 6
gpu_device_id: null
endpoint: http://10.100.79.3:8192/v1/audio/speech
method: POST
content_type: application/json
model:
id: IndexTeam/IndexTTS-2
revision: null
image: local/index-tts:v2
fields:
- name: input
type: textarea
label: Text
required: true
max_length: 5000
- name: voice
type: select
label: Speaker Voice
description: <name>.wav in /worktank/index-tts/voices/. 530s clean clips.
- name: response_format
type: select
options: [wav]
default: wav
description: 22050 Hz PCM_16 mono only; no negotiation.
- name: stream
type: bool
default: false
description: >
Segment-level streaming (~120 tokens). Streaming WAV uses placeholder
data-length (0xFFFFFFFF); browsers fine, strict parsers may complain.
- name: emotion_voice
type: select
label: Emotion Reference Voice
optional: true
description: <name>.wav in /worktank/index-tts/emotions/.
- name: emotion_vector
type: json
label: Emotion Vector
optional: true
description: >
8 floats [happy, angry, sad, afraid, disgusted, melancholic, surprised, calm],
each 0.01.0.
- name: emotion_text
type: textarea
label: Emotion Description (free text)
optional: true
- name: emotion_alpha
type: slider
min: 0.0
max: 1.0
default: 1.0
label: Emotion Strength
response:
type: audio
mime: audio/wav
reproducibility:
seedable: false
deterministic: true
notes: >
No seed. 22050 Hz hardcoded — resample in caller if 24/48 kHz needed.
Precedence if multiple emotion sources: emotion_voice > vector > text.
estimated_latency:
cold_start_s: 5
warm_per_unit: "segment-latency streaming"
license: "Custom Bilibili (free at small scale; commercial tier 100M MAU)"
notes: |
Three-way mutual-exclusion among emotion_voice / emotion_vector / emotion_text;
precedence as above. UI should expose this as a single picker.
- id: qwen3-tts
name: Qwen3-TTS 1.7B
description: >
Alibaba's open English-first TTS (Apache 2.0). 10 languages, 97ms TTFB,
instruction-driven emotion, voice cloning.
category: tts
version: 2
status: down
host: irv-ml1
lifecycle:
stack: qwen3-tts
vram_gb: 16
gpu_device_id: null
endpoint: http://10.100.79.3:8191/v1/audio/speech
method: POST
content_type: application/json
model:
id: Qwen/Qwen3-TTS-12Hz-1.7B
revision: null
image: local/qwen3-tts:v2
fields:
- name: model
type: select
options: [Qwen/Qwen3-TTS-12Hz-1.7B, Qwen/Qwen3-TTS-12Hz-0.6B-Base]
default: Qwen/Qwen3-TTS-12Hz-1.7B
- name: input
type: textarea
label: Text
required: true
max_length: 5000
- name: voice
type: select
label: Voice
source_url: http://10.100.79.3:8191/v1/voices
source_jsonpath: $.voices[*].id
default: vivian
options:
- vivian
- serena
- uncle_fu
- ryan
- aiden
- ono_anna
- sohee
- eric
- dylan
- alloy
- echo
- fable
- nova
- onyx
- shimmer
description: >
Wrapper ships 9 Qwen presets (vivian/serena/uncle_fu/ryan/aiden/
ono_anna/sohee/eric/dylan) + 6 OpenAI-compat aliases (alloy/echo/
fable/nova/onyx/shimmer) — verified live 2026-05-12 against
/v1/voices. Default `vivian` matches the wrapper's OpenAPI default
(case-normalized; the Qwen3-TTS blog cites "Vivian" with caps).
For cloning: clone:<name> where <name> is a profile dir under
/worktank/qwen3-tts/voices/profiles/. Create via /voice-studio web
UI or manually (meta.json + reference.wav). The 17-preset list
cited in the Qwen3-TTS upstream blog (Cherry/Ethan/Chelsie/...)
differs from what this wrapper exposes — the wrapper is the
source of truth.
- name: instructions
type: textarea
label: Emotion/Style Instructions
optional: true
description: >
Natural-language directive (e.g. "speak with cold contempt").
English instructions verified working ~2026-04.
- name: response_format
type: select
options: [wav, mp3, pcm]
default: wav
response:
type: audio
mime_from_field: response_format
reproducibility:
seedable: false
deterministic: true
estimated_latency:
cold_start_s: 5
warm_per_unit: "~97ms TTFB"
license: Apache-2.0
notes: |
Voice cloning shape differs from CosyVoice: profile-based, not voice-id-based.
- id: cosyvoice
name: CosyVoice 3 (Multilingual)
description: >
FunAudioLLM's multilingual expressive TTS. 18+ Chinese dialects + 8 other langs.
English prosody not ElevenLabs-grade — use Qwen3-TTS for English.
category: tts
version: 1
status: down
host: irv-ml1
lifecycle:
stack: cosyvoice
vram_gb: 4
gpu_device_id: null
endpoint: http://10.100.79.3:8190/v1/audio/speech
method: POST
content_type: application/json
model:
id: FunAudioLLM/Fun-CosyVoice3-0.5B-2512
revision: null
image: neosun/cosyvoice:v1.3.2
fields:
- name: model
type: select
options: [cosyvoice-v3, cosyvoice-v2]
default: cosyvoice-v3
- name: input
type: textarea
label: Text (with optional XML emotion tags)
required: true
max_length: 5000
description: >
For English: use XML tags <angry>, <sad>, <surprised>, <fast>, <whisper>, etc.
NOT the instruct field — English instruct values get vocalized literally
(upstream bug).
- name: voice
type: select
label: Voice (cloned only — no presets)
description: >
Create via POST /v1/voices/create (multipart with reference audio ≤30s).
Reference must be 16kHz mono ≤30s; longer = AssertionError.
- name: response_format
type: select
options: [wav]
default: wav
- name: speed
type: slider
min: 0.5
max: 2.0
default: 1.0
response:
type: audio
mime: audio/wav
reproducibility:
seedable: false
deterministic: true
notes: >
Reference audio MUST be ≤30s (16kHz mono). instruct field broken for English.
estimated_latency:
cold_start_s: 5
warm_per_unit: "~150ms TTFB streaming"
license: Apache-2.0
notes: |
GOTCHA: instruct field is Chinese-context only — UI should hide it for English
or surface a strong warning.
- id: fish-s2
name: Fish Audio S2-Pro
description: >
Fishaudio's richest-paralinguistic English TTS (15,000+ inline tags).
Trained 10M+ hours, dual-AR. Released March 2026. Heavy: ~240s compile
warmup on cold start, ~realtime throughput once warm.
category: tts
version: 6
host: irv-ml1
lifecycle:
stack: fish-s2
vram_gb: 24
gpu_device_id: 1
endpoint: http://10.100.79.3:8195/v1/tts
method: POST
content_type: application/json
model:
id: fishaudio/s2-pro
revision: null
image: local/fish-s2:v1
section_groups:
- id: basic
label: Text & output
- id: sampling
label: Sampling & latency
hint: fish-speech ServeTTSRequest defaults. Lower temperature/top_p for steadier reads; latency=balanced for lower latency.
- id: advanced
label: Advanced & voice cloning
fields:
- name: text
type: textarea
label: Text (with optional [tags])
section: basic
required: true
max_length: 5000
description: >
Inline tags: [laugh] [whispers] [super happy] [sigh] [excited]
[heavy breathing] [angry] [sleepy] [crying] [surprise] ... (15,000+).
Upstream best-practice: physical tags like [whispers] feel flat
without emotional context — pair with an emotion tag. Example:
`[excited] Welcome back! [whisper] I have a secret. [laugh]
Just kidding.`
- name: format
type: select
section: basic
options: [wav, mp3, opus]
default: wav
description: >
Output container (upstream also accepts raw `pcm`, omitted here).
wav + mp3 verified live 2026-06-01.
- name: temperature
type: slider
section: sampling
min: 0.1
max: 1.0
step: 0.05
default: 0.8
description: Sampling temperature; lower = steadier. fish-speech default 0.8.
- name: top_p
type: slider
section: sampling
min: 0.1
max: 1.0
step: 0.05
default: 0.8
description: Nucleus sampling cutoff. fish-speech default 0.8.
- name: repetition_penalty
type: slider
section: sampling
min: 0.9
max: 2.0
step: 0.05
default: 1.1
description: >
Penalizes repeats; raise if you hear stutters/loops. fish-speech
default 1.1.
- name: latency
type: select
section: sampling
options: [normal, balanced]
default: normal
description: >
`balanced` trades a little quality for lower latency — useful given
S2-Pro's ~realtime throughput. fish-speech default `normal`.
- name: seed
type: number
section: sampling
required: false
description: Omit for a random seed; a fixed integer repeats a take.
- name: reference_id
type: select
label: Voice
section: basic
optional: true
default: "Emily"
options:
- Abigail
- Adrian
- Alexander
- Alice
- Austin
- Axel
- Beatrice
- Eleanor
- Imogen
- Connor
- Cora
- Elena
- Eli
- Emily
- Everett
- Gabriel
- Gianna
- glados
- Henry
- Ian
- Jade
- Jeremiah
- Jordan
- Julian
- Layla
- Leonardo
- Michael
- Miles
- Olivia
- Ryan
- Taylor
- Thomas
description: >
Voice = a staged clone reference picked by name (THE working voice
path on this build; verified live 2026-06-01). 32 voices staged in
/worktank/fish-s2/references/: 28 from the dia library + glados + 3
British-female VCTK voices (Beatrice/Eleanor/Imogen). Default Emily.
Subtle British (Southern-England) female accents: Imogen (VCTK p225),
Eleanor (p228), Beatrice (p229) — consenting VCTK volunteers (CC BY
4.0), NOT modeled on or representing any public figure. Other female:
Abigail, Alice, Cora, Elena, Emily, Gianna, Jade, Layla, Olivia,
glados. Resolves against the DIRECTORY references/<name>/ (audio +
a same-basename <name>.lab transcript inside it) — NOT a flat
references/<name>.wav, which fish-speech ignores (empty/flat-only
dirs silently collapse every voice to the model default; root-caused
+ fixed 2026-06-01). Blank = model default/random speaker. To add:
create references/<name>/ with <name>.wav (clean 515s) + <name>.lab
(transcript) and add the name here; the deploy playbook's normalize
step also mirrors any flat <name>.wav into this layout. (No /voices
API → static list; a list-endpoint is the durable fix — see notes.)
- name: references
type: json
label: Custom clone (inline base64)
section: advanced
optional: true
description: >
Power-use: zero-shot clone from a reference NOT in the staged library
— array of {audio: <base64-encoded WAV bytes>, text: "transcript"}.
IMPORTANT: `audio` must be base64 bytes, NOT a filesystem path (a path
returns HTTP 500 on this build). For staged voices use `reference_id`.
- name: max_new_tokens
type: number
section: advanced
required: false
default: 1024
description: Max audio tokens per chunk. fish-speech default 1024.
- name: chunk_length
type: slider
section: advanced
min: 100
max: 1000
step: 10
default: 200
description: Target chunk length in chars when splitting. fish-speech default 200.
- name: normalize
type: bool
section: advanced
default: true
description: Text normalization (numbers, punctuation). fish-speech default on.
- name: use_memory_cache
type: select
section: advanced
options: ["off", "on"]
default: "off"
description: >
`on` caches the reference encoding to speed up repeated calls with the
same voice. fish-speech default off.
response:
type: audio
mime_from_field: format
reproducibility:
seedable: true
deterministic: false
seed_field: seed
notes: >
/v1/tts exposes `seed` (omit = random); a fixed seed repeats a take.
Temperature/top_p sampled → not guaranteed byte-exact. (Prior catalog
claimed no seed support — corrected against the upstream ServeTTSRequest
schema + a live generation 2026-06-01.)
estimated_latency:
cold_start_s: 240
warm_per_unit: "~realtime throughput on the A6000 (heavy). latency=balanced trades quality for lower latency."
license: "Research-only (Fishaudio terms; non-commercial)"
license_warning: |
Research/internal only. Not clear for commercial use. Same flag as
Voxtral and SAO — UI must surface when output destined for products.
notes: |
NOT OpenAI-compatible: only /v1/tts (no /v1/audio/voices, /v1/models, or
/openapi.json). Voice discovery is manual (host filesystem). Full control
surface (temperature/top_p/repetition_penalty/latency/seed/format + the
cloning fields) sourced from fish-speech's ServeTTSRequest schema and
verified live 2026-06-01; prior catalog exposed only text + references.
HEAVY: ~240s torch.compile warmup measured on cold start; budget the wait.
- id: kyutai-tts
name: Kyutai TTS 1.6B EN/FR
description: >
Kyutai's bilingual streaming TTS (1.6B, 2.5M hours). Heritage from Mimi codec
+ Moshi dialogue framework. OpenAI-compat HTTP wrapper over Rust streaming core.
category: tts
version: 2
status: down
host: irv-ml1
lifecycle:
stack: kyutai-tts
vram_gb: 6
gpu_device_id: 0
endpoint: http://10.100.79.3:8198/v1/audio/speech
method: POST
content_type: application/json
model:
id: kyutai/tts-1.6b-en_fr
revision: null
image: local/kyutai-tts:v1
fields:
- name: model
type: select
options: [tts-1.6b-en_fr]
default: tts-1.6b-en_fr
- name: input
type: textarea
label: Text
required: true
max_length: 5000
- name: voice
type: text
label: Voice (HF-relative path)
default: unmute-prod-website/default_voice.wav
description: >
Path relative to the `kyutai/tts-voices` HF repo, cached on
irv-ml1 at /worktank/kyutai-tts/hf_cache. NillPointer wrapper
does NOT expose /v1/audio/voices (404; only /health + POST
/v1/audio/speech are wired) — voices are specified by path.
Default `unmute-prod-website/default_voice.wav` is upstream's
named-default. Categories available locally (count):
vctk/* (212, CC BY 4.0 — commercial-safe)
ears/* (306)
voice-donations/* (291, mixed)
expresso/* (103, CC BY-NC — research-only)
cml-tts/* (70)
unmute-prod-website/* (6, incl. default)
alba-mackenna/* (4)
For commercial paths, prefer a vctk/* voice (e.g.,
`vctk/p225_023_mic1.wav`). Full repo:
https://huggingface.co/kyutai/tts-voices
- name: response_format
type: select
options: [wav, mp3, pcm]
default: wav
- name: stream
type: bool
default: false
response:
type: audio
mime_from_field: response_format
reproducibility:
seedable: false
deterministic: true
notes: >
Wrapper adds Python overhead vs Kyutai's bare 220ms claim.
Voice list is filesystem-discovered, not API-exposed — the
NillPointer wrapper only wires /health + /v1/audio/speech.
estimated_latency:
cold_start_s: 4
warm_per_unit: "~220400ms TTFB (with wrapper)"
license: "TBD (Kyutai — verify)"
- id: vibevoice
name: VibeVoice 1.5B (Long-form)
description: >
Microsoft's diffusion-based long-form multi-speaker TTS. Multi-minute scripts
with speaker switching. Not for low-latency single-line use.
category: tts
version: 3
host: irv-ml1
lifecycle:
stack: vibevoice
vram_gb: 20
gpu_device_id: 1
endpoint: http://10.100.79.3:8194/v1/audio/speech
method: POST
content_type: application/json
model:
id: microsoft/VibeVoice-1.5B
revision: null
image: local/vibevoice:v1
fields:
- name: model
type: select
options: [tts-1, tts-1-hd, vibevoice]
default: tts-1
description: >
OpenAI-compatibility aliases all map to VibeVoice internally.
Wire default is tts-1 (verified against /openapi.json).
- name: input
type: textarea
label: "Text (or Speaker N: ... script)"
required: true
description: >
Single-speaker: plain text. Multi-speaker: "Speaker 0: ...\nSpeaker 1: ..."
via /v1/vibevoice/generate (extended endpoint).
- name: voice
type: select
label: Voice
default: en-Alice_woman
description: >
Default `en-Alice_woman` matches the upstream model card's stated
default voice. Foot-gun: Alice has been observed to insert
intro-BGM for openers like "Welcome to ..." — flip to
en-Maya_woman if BGM artifacts are unwanted.
OpenAI presets (compat aliases): alloy, echo, fable, onyx, nova, shimmer.
VibeVoice presets: en-Alice_woman, en-Carter_man, en-Frank_man,
en-Mary_woman_bgm, en-Maya_woman, in-Samuel_man, zh-Anchen_man_bgm,
zh-Bowen_man, zh-Xinran_woman. The `_bgm`-suffixed voices were
trained with background music intentionally — avoid for clean speech.
Custom: drop WAV/MP3/FLAC/M4A into /worktank/vibevoice/voices/,
restart container. Voice cloning training NOT released by Microsoft.
options:
- alloy
- echo
- fable
- onyx
- nova
- shimmer
- en-Alice_woman
- en-Carter_man
- en-Frank_man
- en-Mary_woman_bgm
- en-Maya_woman
- in-Samuel_man
- zh-Anchen_man_bgm
- zh-Bowen_man
- zh-Xinran_woman
- name: response_format
type: select
options: [wav, mp3, opus, flac, pcm]
default: mp3
description: >
Wire default is mp3 (verified /openapi.json). All five values
confirmed by exercising the endpoint 2026-05-11; aac and m4a
fail with 500 despite OpenAPI not enforcing an enum, so excluded.
- name: stream
type: bool
default: false
description: Single-shot endpoint doesn't stream; segment-level on multi-speaker.
response:
type: audio
mime_from_field: response_format
reproducibility:
seedable: false
deterministic: true
estimated_latency:
cold_start_s: 8
warm_per_unit: "3060s first generation; segment-based after"
license: MIT
notes: |
flash_attention_2 default; sdpa fallback.
7B variant (rsxdalv/VibeVoice-Large, ~18GB) needs int8 for <10GB VRAM.
- id: voxtral
name: Voxtral 4B TTS
description: >
Mistral AI's 4B multilingual streaming TTS (CC BY-NC — research/internal only).
8 languages, 70ms model latency, 9.7× realtime. Served via vLLM-Omni.
category: tts
version: 2
status: down
host: irv-ml1
lifecycle:
stack: voxtral
vram_gb: 12
gpu_device_id: 1
endpoint: http://10.100.79.3:8197/v1/audio/speech
method: POST
content_type: application/json
model:
id: mistralai/Voxtral-4B-TTS-2603
revision: null
image: vllm/vllm-omni:v0.18.0
fields:
- name: model
type: select
options: [mistralai/Voxtral-4B-TTS-2603]
default: mistralai/Voxtral-4B-TTS-2603
- name: input
type: textarea
label: Text
required: true
max_length: 5000
- name: voice
type: select
label: Voice
source_url: http://10.100.79.3:8197/v1/audio/voices
default: neutral_female
options:
- neutral_female
- neutral_male
- casual_female
- casual_male
- cheerful_female
- ar_male
- de_female
- de_male
- es_female
- es_male
- fr_female
- fr_male
- hi_female
- hi_male
- it_female
- it_male
- nl_female
- nl_male
- pt_female
- pt_male
description: >
Native Voxtral presets — vLLM-Omni does NOT translate OpenAI
aliases (verified 2026-05-12 live against /v1/audio/voices). Naming
shape is `<register>_<gender>` for English (neutral_*, casual_*,
cheerful_*) and `<lang>_<gender>` for non-English. Default
`neutral_female` matches Mistral docs' stated default. The full
list above was captured live from the running container; uploaded
user voices appear in /v1/audio/voices under `uploaded_voices`.
- name: response_format
type: select
options: [wav, mp3, pcm]
default: wav
- name: stream
type: bool
default: false
response:
type: audio
mime_from_field: response_format
reproducibility:
seedable: false
deterministic: true
notes: vLLM-Omni v0.18.0+ required (first Voxtral-aware release, 2026-03-29).
estimated_latency:
cold_start_s: 6
warm_per_unit: "~70ms model latency, 9.7× realtime"
license: "CC BY-NC (Mistral) — non-commercial only"
license_warning: |
NOT clear for commercial use. UI must surface this when output destined
for products. Use Kokoro/Chatterbox/Fish/IndexTTS/Qwen3/CosyVoice for
commercial paths.
- id: parakeet
name: Parakeet TDT (ASR)
description: >
NVIDIA Parakeet-TDT 0.6B (int8 ONNX, ~400MB) via sherpa-onnx.
Transcription only.
category: asr
version: 1
host: irv-ml1
lifecycle:
stack: parakeet
vram_gb: 2
gpu_device_id: null
endpoint: http://10.100.79.3:8765/transcribe
method: POST
content_type: multipart/form-data
model:
id: nvidia/parakeet-tdt-v2-en-int8
revision: null
image: local/parakeet:sherpa-onnx-v2
fields:
- name: file
type: file
label: Audio File
required: true
accepted_types: [audio/wav, audio/mp3, audio/flac, audio/ogg]
response:
type: text
output_field: text
reproducibility:
seedable: false
deterministic: true
estimated_latency:
cold_start_s: 1
warm_per_unit: "~realtime"
license: "Proprietary (NVIDIA model) + Apache-2.0 (sherpa-onnx)"
notes: |
Aliased at /v1/audio/transcriptions for OpenAI compat.
v3 model (25 languages) available via env-only change.
- id: stable-audio-open
name: Stable Audio Open 1.0 (SFX)
description: >
Stability AI's text-to-audio diffusion. SFX/foley/ambience only — NOT music.
Max 47s clips. Seeded, fully deterministic.
category: sfx
version: 2
host: irv-ml1
lifecycle:
stack: stable-audio-open
vram_gb: 6
gpu_device_id: 1
endpoint: http://10.100.79.3:8211/v1/audio/sfx
method: POST
content_type: application/json
model:
id: stabilityai/stable-audio-open-1.0
revision: null
image: local/stable-audio-open:v1
fields:
- name: prompt
type: textarea
label: Prompt
required: true
description: >
Description of the audio to generate. Per diffusers Tips section,
descriptive multi-element prompts produce better results than terse
ones — `"melodic techno with a fast beat and synths, 128 BPM"` is
stronger than `"techno"`. Other canonical examples from the model
card: `"128 BPM tech house drum loop"`, `"the sound of a hammer
hitting a wooden surface"`.
- name: negative_prompt
type: textarea
label: Negative Prompt
default: "low quality, average quality"
description: >
Official diffusers-docs recommendation. `"low quality, average
quality"` (lowercase, comma-separated) per the Stable Audio
pipeline's Tips section — was `"Low quality."` in v0, fixed in v1.
- name: duration
type: slider
min: 0.5
max: 47.0
default: 10.0
label: Duration (seconds)
- name: steps
type: slider
min: 10
max: 300
default: 100
label: Diffusion Steps
- name: cfg_scale
type: slider
min: 0.0
max: 20.0
default: 7.0
label: CFG Scale
- name: seed
type: number
label: Seed
optional: true
response:
type: audio
mime: audio/wav
reproducibility:
seedable: true
seed_field: seed
deterministic: true
notes: >
Identical seed+prompt+steps+cfg = bit-identical bytes. Wrapper serializes
concurrent requests via asyncio.Lock (StableAudioPipeline not reentrant).
estimated_latency:
cold_start_s: 3
warm_per_unit: "~3060s per 10s clip"
license: "Stability AI Community (non-commercial)"
license_warning: |
Non-commercial only. Same flag as Voxtral for commercial paths.
- id: ace-step
name: ACE-Step 1.5 (Music)
description: >
Apache-2.0 hybrid diffusion+LLM music generation. Multi-minute lyric-aware
songs with vocals + instrumentation.
category: music
version: 5
host: irv-ml1
lifecycle:
stack: ace-step
vram_gb: 16
gpu_device_id: 1
endpoint: http://10.100.79.3:8210/generate
method: POST
content_type: application/json
model:
id: ace-step/ACE-Step
revision: main
image: local/ace-step:v1
section_groups:
- id: basic
label: Basic
hint: "Prompt + lyrics + duration. The 80% case."
- id: generation
label: Generation
hint: "Sampler steps, guidance, scheduler, seed."
- id: conditioning
label: Conditioning
hint: "ERG flags, OSS schedule, guidance shape."
- id: a2a
label: "Audio → Audio"
hint: "Generate from a reference clip."
- id: lora
label: LoRA
hint: "Adapter weights."
- id: output
label: Output
hint: "Format and device."
fields:
- name: prompt
type: textarea
label: Musical Prompt
required: true
section: basic
description: >
Style/mood/instrumentation as a comma-separated tag string.
The upstream Gradio UI ships pre-filled with: `"funk, pop, soul,
rock, melodic, guitar, drums, bass, keyboard, percussion, 105 BPM,
energetic, upbeat, groovy, vibrant, dynamic"` — same shape works
well here. Shorter examples: `"uplifting pop with synth leads"`,
`"lo-fi hip-hop beat with vinyl crackle"`.
- name: lyrics
type: textarea
label: Lyrics
required: true
default: ""
section: basic
description: >
Required by the upstream Pydantic model (`lyrics: str`) but empty
string is a valid value (instrumental). Default '' so the inference
client always sends a string.
- name: audio_duration
type: slider
min: 5.0
max: 240.0
default: 30.0
label: Duration (seconds)
section: basic
description: >
Upstream caps at 240s (the model's training horizon). Lower bound
5s is our choice — upstream uses -1 as a "random duration" sentinel
which is hostile UX for a slider. Default 30s also kept (upstream
uses -1; explicit 30 is the better first-time-user experience).
- name: infer_step
type: number
default: 60
label: Inference Steps
section: generation
description: >
Upstream Gradio default is 60 (matches benchmark numbers in the
README). Lower values (20-30) are useful for "preview" passes;
higher (80-100) marginal returns.
- name: guidance_scale
type: slider
min: 0.0
max: 30.0
default: 15.0
section: generation
- name: scheduler_type
type: select
options: [euler, heun, pingpong]
default: euler
section: generation
description: >
Flow-matching scheduler. `euler` is the canonical default;
`heun` trades 2x compute for higher-order accuracy; `pingpong`
is a specialized noise-bouncing variant.
- name: cfg_type
type: select
options: [apg, cfg, cfg_star]
default: apg
section: generation
description: >
Classifier-free guidance variant. Upstream Gradio default is `apg`
(adaptive projected guidance); `cfg` is the standard SD-style
formulation; `cfg_star` is an advanced alternative. See the
upstream pipeline source for trade-offs.
- name: omega_scale
type: slider
min: -10.0
max: 30.0
default: 10.0
section: generation
description: >
Upstream technically allows [-100, 100] but values that wide are
unusable as a slider. UI-bounded to [-10, 30] which covers the
typical zone with headroom. Hit the API directly for extremes.
- name: actual_seeds
type: json
label: Seeds (empty = random)
default: []
section: generation
description: >
Empty list = wrapper sends empty string to pipeline = pipeline
picks a random seed per batch element. Explicit seeds (e.g. [42]
or [42, 137, 9999]) for reproducibility.
REPRODUCIBILITY GAP (queued for follow-up): the pipeline returns
the chosen seed in its result dict, but our wrapper currently
throws it away. Assets generated with the default `[]` cannot
currently be regenerated. Workaround: set actual_seeds explicitly
when reproducibility matters. Wrapper enhancement to surface
random-resolved seeds via X-Actual-Seeds header + a catalog
schema field for header→accessory capture is the planned fix.
- name: guidance_interval
type: slider
min: 0.0
max: 1.0
default: 0.5
section: conditioning
- name: guidance_interval_decay
type: slider
min: 0.0
max: 1.0
default: 0.0
section: conditioning
description: >
Upstream Gradio default is 0.0 (no decay). Catalog v4 had this at
1.0 (full decay) — wrong; produced under-conditioned outputs.
- name: min_guidance_scale
type: slider
min: 0.0
max: 20.0
default: 3.0
section: conditioning
description: >
Upstream technically allows up to 200; UI-bounded to 20 (covers
the typical zone). Hit the API directly for extremes.
- name: use_erg_tag
type: bool
default: true
section: conditioning
- name: use_erg_lyric
type: bool
default: false
section: conditioning
- name: use_erg_diffusion
type: bool
default: true
section: conditioning
- name: oss_steps
type: json
default: []
section: conditioning
- name: guidance_scale_text
type: slider
min: 0.0
max: 10.0
default: 0.0
section: conditioning
- name: guidance_scale_lyric
type: slider
min: 0.0
max: 10.0
default: 0.0
section: conditioning
- name: audio2audio_enable
type: bool
default: false
section: a2a
- name: ref_audio_strength
type: slider
min: 0.0
max: 1.0
default: 0.5
section: a2a
- name: ref_audio_input
type: text
label: Reference Audio Path
optional: true
section: a2a
- name: lora_name_or_path
type: text
label: LoRA Repo/Path
optional: true
section: lora
- name: lora_weight
type: slider
min: -3.0
max: 3.0
default: 1.0
section: lora
description: >
Negative weights are legitimate (apply the LoRA in inverse).
Upstream Gradio range adopted verbatim.
- name: audio_format
type: select
options: [wav, mp3, flac]
default: wav
section: output
- name: bf16
type: bool
default: true
section: output
- name: torch_compile
type: bool
default: false
section: output
- name: device_id
type: number
default: 0
label: GPU Device Index
section: output
- name: checkpoint_path
type: text
label: Checkpoint Path
required: true
default: /app/checkpoints
section: output
description: >
Required by the upstream wrapper (`infer-api.py:29`, no default in
the Pydantic model). The container always mounts checkpoints at
/app/checkpoints, so the catalog default works out-of-the-box.
Only override if you've mounted an alternate checkpoint path.
Wrapper-side cleanup queued — once the upstream model defaults this,
the catalog field will become optional or be dropped entirely.
response:
# As of wrapper version that ships with image local/ace-step:v1
# post 2026-05-11, /generate streams audio bytes inline with
# Content-Type set from the audio_format request field. The
# in-container output_path is exposed via X-Output-Path header
# for log correlation but is no longer load-bearing.
type: audio
mime_from_field: audio_format
reproducibility:
seedable: true
deterministic: true
notes: >
actual_seeds parameter exposed; identical seeds + params = identical audio.
Local infer-api.py patches upstream's broken 24-arg pipeline signature
(v2: was 18 in upstream — caused crashes with audio_duration in `format`
slot) AND inline-streams the generated audio bytes (v4: was returning a
JSON path reference to a file inside the container, which was
unreachable from outside).
REPRODUCIBILITY GAP (v5): default `actual_seeds: []` triggers random
seed selection inside the pipeline. The chosen seed IS available in
the pipeline's return dict (`actual_seeds` key) but our wrapper
doesn't capture or surface it — so default-defaulted assets cannot be
regenerated bit-exact. Set actual_seeds explicitly when reproducibility
is required. Wrapper enhancement to surface chosen seeds via response
header + a catalog schema for header→accessory capture is queued.
estimated_latency:
cold_start_s: 30
warm_per_unit: "~1060s depending on audio_duration + infer_step"
license: Apache-2.0
notes: |
27-field surface — UI must do progressive disclosure (basic/advanced).
Gradio UI fallback: docker exec -it ace-step python3 acestep/gui.py.
- id: comfyui
name: ComfyUI (workflow engine — catalog-deferred)
description: >
Node-based diffusion workflow engine. Native API is workflow JSON DAG, not
form fields. Catalog-deferred until per-asset-type workflow templates are
defined and a wrapper maps form inputs to template node parameters.
category: image
status: catalog-deferred
version: 1
host: irv-ml1
endpoint: http://10.100.79.3:8188
method: WebSocket + POST
content_type: application/json
model:
id: "various (SD, SDXL, Flux, ControlNet, LoRA, upscalers)"
revision: null
image: mmartial/comfyui-nvidia-docker:ubuntu24_cuda12.8-20260312
response:
type: image
mime: image/png
reproducibility:
seedable: true
deterministic: true
notes: >
Seed is a workflow node parameter. Reproducibility requires persisting
full workflow JSON + seed alongside the asset.
license: GPL-3.0
notes: |
To bring into the catalog: define workflow templates per asset-type
(portrait, landscape, sfx-thumbnail, ...), build a wrapper that does
form-field-to-node-parameter substitution, submit via POST /prompt, poll
/history. Until then, expose ComfyUI as an external link in the UI.
User state at /worktank/comfyui/basedir/.
# Dia2 runs as TWO fixed-model instances: the devnen wrapper is
# single-model and IGNORES the OpenAI `model` field (serves whatever
# checkpoint is loaded), so the only way to offer both Dia2 models to
# asset-engine as real per-request choices is one endpoint per model.
# Both live in the `dia` stack (one compose, two services); `compose up`
# starts both. The legacy Dia 1.6B entry was retired 2026-05-31 in favor
# of these — drop the old `dia` tile when re-vendoring.
- id: dia2-2b
name: Dia2 2B (High Quality)
description: >
Nari Labs' Dia2 dialogue TTS, 2B high-quality checkpoint — multi-speaker
turn-taking in one pass with [S1]/[S2] speaker tags and nonverbals
(laughs)/(coughs)/(sighs). 24 kHz. Served via devnen/Dia-TTS-Server.
category: tts
version: 2
status: ready
host: irv-ml1
lifecycle:
stack: dia
vram_gb: 8
gpu_device_id: 0
endpoint: http://10.100.79.3:8200/tts
method: POST
content_type: application/json
model:
id: nari-labs/Dia2-2B
revision: null
image: local/dia:v2
section_groups:
- id: basic
label: Text & voice
- id: sampling
label: Sampling
hint: Wrapper-canonical Dia2 defaults; lower temperature/cfg for steadier output.
- id: advanced
label: Advanced
fields:
- name: text
type: textarea
label: Text ([S1]/[S2] dialogue + nonverbals)
section: basic
required: true
max_length: 5000
default: "[S1] Hello Dia2!" # canonical: nari-labs/dia2 README usage example
description: >
[S1]/[S2] tags mark speaker turns; nonverbals like (laughs),
(coughs), (sighs), (clears throat) go inline.
- name: voice_mode
type: select
label: Voice mode
section: basic
options: [clone, single_s1, single_s2, dialogue]
default: clone
description: >
Defaults to `clone` -> a stable predefined voice (clone_reference_filename
below, default Abigail.wav). The unconditioned modes (single_s1/s2,
dialogue) sample a RANDOM speaker each call (gender varies) — pick one
only if you want that. NOTE: voice_mode=clone REQUIRES a real
clone_reference_filename; an empty/"undefined" value 404s.
- name: clone_reference_filename
type: select
label: Voice (clone reference)
section: basic
optional: true
default: "Abigail.wav"
source_url: http://10.100.79.3:8200/get_reference_files
source_jsonpath: $[*]
description: >
Required when voice_mode=clone. Lists the reference_audio library —
43 curated voices (singles + [S1]/[S2] dialogue pairs like
Abigail_Taylor.wav) staged there from the image's /app/voices, plus
any you upload. Pins voice/gender deterministically.
- name: transcript
type: textarea
label: Reference transcript
section: advanced
optional: true
description: >
Optional transcript of a custom clone reference (overrides its local
.txt sidecar). Leave blank for predefined voices.
- name: output_format
type: select
section: basic
options: [opus, wav]
default: opus
description: 24 kHz (Dia2).
- name: cfg_scale
type: slider
section: sampling
min: 1.0
max: 5.0
step: 0.1
default: 3.0
description: Classifier-free guidance scale (wrapper default 3.0).
- name: temperature
type: slider
section: sampling
min: 0.1
max: 1.5
step: 0.05
default: 1.3
description: Sampling temperature; lower = steadier (wrapper default 1.3).
- name: top_p
type: slider
section: sampling
min: 0.1
max: 1.0
step: 0.01
default: 0.95
- name: cfg_filter_top_k
type: number
section: sampling
required: false
default: 35
description: Top-k filter for CFG guidance (1100; wrapper default 35).
- name: speed_factor
type: slider
section: sampling
min: 0.5
max: 2.0
step: 0.01
default: 0.94
description: Playback speed (wrapper canonical default 0.94).
- name: seed
type: number
section: sampling
required: false
default: -1
description: -1 = random; a fixed integer repeats the same take.
- name: split_text
type: bool
section: advanced
default: true
description: Auto-split long text into chunks.
- name: chunk_size
type: slider
section: advanced
min: 100
max: 1000
step: 10
default: 300
description: Target chunk length in chars when splitting (wrapper default 300).
- name: max_tokens
type: number
section: advanced
required: false
description: Max audio tokens per chunk; omit to use the model's internal default.
response:
type: audio
mime_from_field: output_format
reproducibility:
seedable: true
deterministic: false
notes: >
Per the dia2 README, voices vary per generation unless anchored — set
voice_mode=clone + a clone_reference_filename for a stable voice/gender.
seed (-1=random) repeats a take but doesn't choose the voice;
temperature-sampled, not byte-exact; raw bytes -> no seed_field. All
field defaults are the wrapper's CustomTTSRequest Pydantic blessed values.
estimated_latency:
cold_start_s: 30
warm_per_unit: "dialogue one-pass; ~realtime on the 3090"
license: "Apache-2.0 (Dia2 weights); MIT (devnen wrapper)"
notes: |
Routes to the wrapper's richer /tts endpoint (full control surface:
cfg_scale/temperature/top_p/cfg_filter_top_k/voice_mode/clone) rather than
the OpenAI /v1/audio/speech. One of two fixed-model instances in the `dia`
stack; sibling dia2-1b (:8202). `docker compose up` starts BOTH (~8 GB
here; budget both if gating the whole stack). Image local/dia:v2 (see
stacks/dia/dia2-image/). Verified end-to-end 2026-05-31.
- id: dia2-1b
name: Dia2 1B (Streaming)
description: >
Nari Labs' Dia2 dialogue TTS, 1B streaming checkpoint — same [S1]/[S2]
dialogue + nonverbals as dia2-2b, lower latency / lighter weight. 24 kHz.
Served via devnen/Dia-TTS-Server.
category: tts
version: 2
status: ready
host: irv-ml1
lifecycle:
stack: dia
vram_gb: 5
gpu_device_id: 0
endpoint: http://10.100.79.3:8202/tts
method: POST
content_type: application/json
model:
id: nari-labs/Dia2-1B
revision: null
image: local/dia:v2
section_groups:
- id: basic
label: Text & voice
- id: sampling
label: Sampling
hint: Wrapper-canonical Dia2 defaults; lower temperature/cfg for steadier output.
- id: advanced
label: Advanced
fields:
- name: text
type: textarea
label: Text ([S1]/[S2] dialogue + nonverbals)
section: basic
required: true
max_length: 5000
default: "[S1] Hello Dia2!" # canonical: nari-labs/dia2 README usage example
description: >
[S1]/[S2] tags mark speaker turns; nonverbals like (laughs),
(coughs), (sighs), (clears throat) go inline.
- name: voice_mode
type: select
label: Voice mode
section: basic
options: [clone, single_s1, single_s2, dialogue]
default: clone
description: >
Defaults to `clone` -> a stable predefined voice (clone_reference_filename
below, default Abigail.wav). The unconditioned modes (single_s1/s2,
dialogue) sample a RANDOM speaker each call (gender varies) — pick one
only if you want that. NOTE: voice_mode=clone REQUIRES a real
clone_reference_filename; an empty/"undefined" value 404s.
- name: clone_reference_filename
type: select
label: Voice (clone reference)
section: basic
optional: true
default: "Abigail.wav"
source_url: http://10.100.79.3:8202/get_reference_files
source_jsonpath: $[*]
description: >
Required when voice_mode=clone. Lists the reference_audio library —
43 curated voices (singles + [S1]/[S2] dialogue pairs like
Abigail_Taylor.wav) staged there from the image's /app/voices, plus
any you upload. Pins voice/gender deterministically.
- name: transcript
type: textarea
label: Reference transcript
section: advanced
optional: true
description: >
Optional transcript of a custom clone reference (overrides its local
.txt sidecar). Leave blank for predefined voices.
- name: output_format
type: select
section: basic
options: [opus, wav]
default: opus
description: 24 kHz (Dia2).
- name: cfg_scale
type: slider
section: sampling
min: 1.0
max: 5.0
step: 0.1
default: 3.0
description: Classifier-free guidance scale (wrapper default 3.0).
- name: temperature
type: slider
section: sampling
min: 0.1
max: 1.5
step: 0.05
default: 1.3
description: Sampling temperature; lower = steadier (wrapper default 1.3).
- name: top_p
type: slider
section: sampling
min: 0.1
max: 1.0
step: 0.01
default: 0.95
- name: cfg_filter_top_k
type: number
section: sampling
required: false
default: 35
description: Top-k filter for CFG guidance (1100; wrapper default 35).
- name: speed_factor
type: slider
section: sampling
min: 0.5
max: 2.0
step: 0.01
default: 0.94
description: Playback speed (wrapper canonical default 0.94).
- name: seed
type: number
section: sampling
required: false
default: -1
description: -1 = random; a fixed integer repeats the same take.
- name: split_text
type: bool
section: advanced
default: true
description: Auto-split long text into chunks.
- name: chunk_size
type: slider
section: advanced
min: 100
max: 1000
step: 10
default: 300
description: Target chunk length in chars when splitting (wrapper default 300).
- name: max_tokens
type: number
section: advanced
required: false
description: Max audio tokens per chunk; omit to use the model's internal default.
response:
type: audio
mime_from_field: output_format
reproducibility:
seedable: true
deterministic: false
notes: >
Same story as dia2-2b: voices vary per generation unless anchored
(voice_mode=clone + clone_reference_filename); temperature-sampled,
not byte-exact; raw bytes -> no seed_field. Defaults are the wrapper's
CustomTTSRequest Pydantic blessed values.
estimated_latency:
cold_start_s: 30
warm_per_unit: "streaming; lower latency than dia2-2b"
license: "Apache-2.0 (Dia2 weights); MIT (devnen wrapper)"
notes: |
Routes to the wrapper's richer /tts endpoint (same control surface as
dia2-2b). Sibling of dia2-2b in the same `dia` stack (see its notes for
the single-model-per-endpoint rationale + both-start-together caveat).
Image local/dia:v2. Verified end-to-end 2026-05-31.
- id: zonos
name: Zonos-v0.1 TTS
description: >
Zyphra's expressive multilingual open-weight TTS (Apache-2.0) — 44 kHz,
zero-shot voice cloning from a 1030s clip, and an 8-axis emotion
vector plus pitch/rate/fmax conditioning. First genuinely seedable TTS
in the fleet: the adapter returns the seed so regenerate/fork can pin it.
category: tts
version: 1
status: down
host: irv-ml1
lifecycle:
stack: zonos
# Adapter alone (~6 GB). NOTE: the zonos stack currently ships TWO
# services (the Gradio eval UI + this adapter); if orchestration does a
# whole-stack `compose up` it loads the model twice (~12 GB). Target the
# zonos-api service, or drop the Gradio service at promotion.
vram_gb: 6
gpu_device_id: 0
endpoint: http://10.100.79.3:8203/v1/audio/speech
method: POST
content_type: application/json
model:
id: Zyphra/Zonos-v0.1-transformer
revision: null
image: local/zonos-api:v1
section_groups:
- id: text
label: Text
- id: voice
label: Voice
- id: expression
label: Expression
hint: Conditioning knobs; leave emotion blank to keep it unconditional.
- id: generation
label: Generation
fields:
- name: input
type: textarea
label: Text to synthesize
section: text
required: true
max_length: 500
default: "Zonos uses eSpeak for text to phoneme conversion!" # canonical: upstream gradio_interface.py default
description: eSpeak-phonemized. Upstream Gradio caps the box at ~500 chars.
- name: language
type: select
label: Language code
section: text
default: en-us
# eSpeak code list is large; these are the documented multilingual
# set. The adapter accepts any code in zonos.supported_language_codes.
options: [en-us, ja, cmn, fr-fr, de]
description: Any eSpeak code Zonos supports; 400 on an unknown code.
- name: voice
type: select
label: Clone voice (optional)
section: voice
optional: true
source_url: http://10.100.79.3:8203/v1/audio/voices
source_jsonpath: $.voices[*]
description: >
Reference clip filename under /worktank/zonos/voices/ (1030s clean
WAV). Omit for Zonos's default unconditioned voice.
- name: response_format
type: select
section: generation
options: [wav, flac]
default: wav
description: 44.1 kHz. Both lossless (libsndfile); no mp3/opus in v1.
- name: emotion
type: json
label: Emotion vector (optional)
section: expression
optional: true
description: >
8 floats 0.01.0 [happy, sad, disgust, fear, surprise, anger, other,
neutral]. Omit to leave emotion unconditional (the Gradio default;
upstream warns the sliders are finicky and can destabilize output).
- name: speaking_rate
type: slider
section: expression
min: 5.0
max: 30.0
step: 0.5
default: 15.0
- name: pitch_std
type: slider
label: Pitch std
section: expression
min: 0.0
max: 300.0
step: 1
default: 45.0
- name: fmax
type: slider
label: Fmax (Hz)
section: expression
min: 0
max: 24000
step: 1
default: 24000
- name: cfg_scale
type: slider
label: CFG scale
section: generation
min: 1.0
max: 5.0
step: 0.1
default: 2.0
- name: seed
type: number
section: generation
optional: true
description: Omit for a random seed; the seed used is returned in the envelope.
response:
# JSON envelope (like kokoro-captioned) so the seed rides back with the
# audio. Adapter returns: {audio: <base64 in response_format>,
# audio_format: "audio/wav"|"audio/flac", seed: <int>}.
type: audio
audio_field: audio
audio_format_field: audio_format
reproducibility:
seedable: true
deterministic: true
seed_field: seed
notes: >
torch.manual_seed(seed); the adapter echoes the seed it used in the
envelope's `seed` key. Same seed + params + GPU → byte-stable in
practice; cross-GPU (3090 vs A6000) may differ at bf16. Wire shape
is defined by the adapter (stacks/zonos/adapter/server.py) — verify
against the live endpoint at deploy per CATALOG-CONTRACT.
estimated_latency:
cold_start_s: 30
warm_per_unit: "~2x realtime on the 3090 (transformer variant)"
license: Apache-2.0
notes: |
Served via a thin OpenAI-ish FastAPI adapter (stacks/zonos/adapter/) in
front of the Zonos Python SDK — upstream ships only Gradio + SDK, no REST.
Transformer variant only (mamba-ssm absent → no hybrid). Fields read from
upstream gradio_interface.py (2026-05-31), the CATALOG-CONTRACT blessed
source for defaults/ranges. Adapter not yet deployed/verified — flip to
ready (or experimental) after the first successful generation through 8203.
# Reproducibility audit — answers per service: (a) seedable, (b) model
# deterministic without seed, (c) image tag mutable (security/reproducibility risk).
reproducibility_audit:
- service: kokoro
seedable: false
model_deterministic: true
image_tag_mutable: true
notes: ":v0.2.4-master is mutable; pin to digest for v3."
- service: kokoro-captioned
seedable: false
model_deterministic: true
image_tag_mutable: true
notes: "Same image as kokoro proper; same mutability story. Response carries timestamps."
- service: chatterbox
seedable: true
model_deterministic: false
image_tag_mutable: true
notes: "/tts exposes seed (0=random); temperature-sampled, not byte-exact. PerTh watermark unconditional (Resemble policy). image :latest is mutable — pin a digest/SHA for true repro."
- service: index-tts
seedable: false
model_deterministic: true
image_tag_mutable: false
notes: "22050 Hz hardcoded — caller must resample."
- service: qwen3-tts
seedable: false
model_deterministic: true
image_tag_mutable: false
- service: cosyvoice
seedable: false
model_deterministic: true
image_tag_mutable: false
notes: "instruct field broken for English; XML tags only."
- service: fish-s2
seedable: true
model_deterministic: false
image_tag_mutable: false
notes: "/v1/tts exposes seed (omit=random); temperature/top_p sampled, not byte-exact. Research-only license — non-commercial."
- service: kyutai-tts
seedable: false
model_deterministic: true
image_tag_mutable: false
- service: vibevoice
seedable: false
model_deterministic: true
image_tag_mutable: false
notes: "Voice cloning training not released."
- service: voxtral
seedable: false
model_deterministic: true
image_tag_mutable: false
notes: "CC BY-NC — non-commercial."
- service: parakeet
seedable: false
model_deterministic: true
image_tag_mutable: false
- service: stable-audio-open
seedable: true
model_deterministic: true
image_tag_mutable: false
notes: "Wrapper serializes concurrent requests (StableAudioPipeline not reentrant)."
- service: ace-step
seedable: true
model_deterministic: true
image_tag_mutable: false
notes: "Local infer-api.py patches upstream's broken pipeline signature."
- service: comfyui
seedable: true
model_deterministic: true
image_tag_mutable: false
notes: "Reproducibility requires persisting full workflow JSON + seed."
- service: dia2-2b
seedable: true
model_deterministic: false
image_tag_mutable: false
notes: "Dia2 2B, temperature-sampled (seed gives consistent prosody, not byte-exact). local/dia:v2 (derived; upstream wrapper SHA not pinned)."
- service: dia2-1b
seedable: true
model_deterministic: false
image_tag_mutable: false
notes: "Dia2 1B streaming; same determinism story as dia2-2b. local/dia:v2."
- service: zonos
seedable: true
model_deterministic: true
image_tag_mutable: false
notes: "Adapter echoes the seed used (reproducibility.seed_field=seed). Byte-stable same-GPU; bf16 may drift cross-GPU. local/zonos-api:v1 built FROM local/zonos (pin ZONOS_SHA for true repro)."