catalog: defaults audit — Kokoro/SAO/VibeVoice/Fish-s2/ace-step

Six default + description tightenings from upstream-source research:

- **Kokoro voice**: af_bella → af_heart. Per upstream VOICES.md, af_heart
  is the only A-rated voice; also the kokoro-fastapi container's own
  default. Applied to both kokoro (v2 → v3) and kokoro-captioned (v1 → v2).
- **SAO negative_prompt**: "Low quality." → "low quality, average quality".
  Per diffusers official docs Tips section — the lowercase comma-separated
  shape is the explicit recommendation; the period form was a code-example
  string, never a documented default. SAO bumped v1 → v2.
- **SAO prompt description**: added the model card's canonical examples
  ("128 BPM tech house drum loop", "the sound of a hammer hitting a wooden
  surface") plus the Tips advice on descriptive prompts.
- **VibeVoice voice**: en-Carter_man → en-Alice_woman. Per upstream model
  card, Alice is the documented default. Description now flags the
  Alice-injects-BGM-for-intros foot-gun + the `_bgm`-suffix meaning.
  VibeVoice bumped v2 → v3.
- **Fish-s2 text description**: added multi-tag placeholder example
  combining emotion + physical tags, per upstream best-practice docs
  (physical tags "feel flat without emotional context").
- **ace-step prompt description**: added the upstream Gradio UI's
  pre-filled tag string as the canonical prompt-shape example.

Surfaced separately to lkraven (NOT applied here, need decisions):
- Voxtral voice "alloy" may need to become a native preset like
  "neutral_female" — depends on whether vLLM-Omni translates OpenAI
  aliases.
- Qwen3-TTS — research found 17 presets (Cherry, Ethan, ...) that the
  catalog currently says don't exist. Catalog says clone-only; needs
  wrapper-level verification before adding.
- Kyutai-TTS — service down, can't probe /v1/audio/voices. Default
  remains undefaulted.
This commit is contained in:
2026-05-12 22:29:22 -07:00
parent 079c7b15e3
commit 51c77370b3
+41 -13
View File
@@ -26,7 +26,7 @@ services:
60+ built-in voices in 8 languages. Voice mixing via inline weights.
OpenAI-compatible /v1/audio/speech.
category: tts
version: 2
version: 3
host: irv-ml1
lifecycle:
stack: kokoro
@@ -50,10 +50,12 @@ services:
label: Voice
source_url: http://10.100.79.3:8193/v1/audio/voices
source_jsonpath: $.voices[*]
default: af_bella
default: af_heart
description: >
60+ built-in voices. Custom blends: af_bella(2)+af_aoede(1) syntax for
weighted mixing. Persistent custom voices via playbooks/blend-kokoro-voice.yaml.
60+ built-in voices. Default `af_heart` is the only A-rated voice per
upstream VOICES.md (also the kokoro-fastapi container's own default).
Custom blends: af_bella(2)+af_aoede(1) syntax for weighted mixing.
Persistent custom voices via playbooks/blend-kokoro-voice.yaml.
- name: response_format
type: select
options: [wav, mp3, opus, flac, pcm]
@@ -119,7 +121,7 @@ services:
separate catalog entry because the wire shape is a JSON envelope
carrying base64-encoded audio plus a structured timestamps array.
category: tts
version: 1
version: 2
status: ready
host: irv-ml1
endpoint: http://10.100.79.3:8193/dev/captioned_speech
@@ -140,7 +142,7 @@ services:
label: Voice
source_url: http://10.100.79.3:8193/v1/audio/voices
source_jsonpath: $.voices[*]
default: af_bella
default: af_heart
- name: response_format
type: select
options: [wav, mp3, opus, flac, pcm]
@@ -494,6 +496,10 @@ services:
description: >
Inline tags: [laugh] [whispers] [super happy] [sigh] [excited]
[heavy breathing] [angry] [sleepy] [crying] [surprise] ... (15,000+).
Upstream best-practice: physical tags like [whispers] feel flat
without emotional context — pair with an emotion tag. Example:
`[excited] Welcome back! [whisper] I have a secret. [laugh]
Just kidding.`
- name: references
type: json
label: Voice References
@@ -578,7 +584,7 @@ services:
Microsoft's diffusion-based long-form multi-speaker TTS. Multi-minute scripts
with speaker switching. Not for low-latency single-line use.
category: tts
version: 2
version: 3
host: irv-ml1
lifecycle:
stack: vibevoice
@@ -609,12 +615,17 @@ services:
- name: voice
type: select
label: Voice
default: en-Carter_man
default: en-Alice_woman
description: >
OpenAI presets: alloy, echo, fable, onyx, nova, shimmer.
Default `en-Alice_woman` matches the upstream model card's stated
default voice. Foot-gun: Alice has been observed to insert
intro-BGM for openers like "Welcome to ..." — flip to
en-Maya_woman if BGM artifacts are unwanted.
OpenAI presets (compat aliases): alloy, echo, fable, onyx, nova, shimmer.
VibeVoice presets: en-Alice_woman, en-Carter_man, en-Frank_man,
en-Mary_woman_bgm, en-Maya_woman, in-Samuel_man, zh-Anchen_man_bgm,
zh-Bowen_man, zh-Xinran_woman.
zh-Bowen_man, zh-Xinran_woman. The `_bgm`-suffixed voices were
trained with background music intentionally — avoid for clean speech.
Custom: drop WAV/MP3/FLAC/M4A into /worktank/vibevoice/voices/,
restart container. Voice cloning training NOT released by Microsoft.
options:
@@ -762,7 +773,7 @@ services:
Stability AI's text-to-audio diffusion. SFX/foley/ambience only — NOT music.
Max 47s clips. Seeded, fully deterministic.
category: sfx
version: 1
version: 2
host: irv-ml1
lifecycle:
stack: stable-audio-open
@@ -780,10 +791,21 @@ services:
type: textarea
label: Prompt
required: true
description: >
Description of the audio to generate. Per diffusers Tips section,
descriptive multi-element prompts produce better results than terse
ones — `"melodic techno with a fast beat and synths, 128 BPM"` is
stronger than `"techno"`. Other canonical examples from the model
card: `"128 BPM tech house drum loop"`, `"the sound of a hammer
hitting a wooden surface"`.
- name: negative_prompt
type: textarea
label: Negative Prompt
default: "Low quality."
default: "low quality, average quality"
description: >
Official diffusers-docs recommendation. `"low quality, average
quality"` (lowercase, comma-separated) per the Stable Audio
pipeline's Tips section — was `"Low quality."` in v0, fixed in v1.
- name: duration
type: slider
min: 0.5
@@ -867,7 +889,13 @@ services:
label: Musical Prompt
required: true
section: basic
description: Style/mood/instrumentation, e.g. "uplifting pop with synth leads".
description: >
Style/mood/instrumentation as a comma-separated tag string.
The upstream Gradio UI ships pre-filled with: `"funk, pop, soul,
rock, melodic, guitar, drums, bass, keyboard, percussion, 105 BPM,
energetic, upbeat, groovy, vibrant, dynamic"` — same shape works
well here. Shorter examples: `"uplifting pop with synth leads"`,
`"lo-fi hip-hop beat with vinyl crackle"`.
- name: lyrics
type: textarea
label: Lyrics