catalog(dia2): expose full /tts control surface + stable predefined voices

Repoint both dia2 entries from /v1/audio/speech to the wrapper's richer /tts
endpoint (CustomTTSRequest), exposing the levers that fix the random-voice
problem: voice_mode, clone_reference_filename, cfg_scale, temperature, top_p,
cfg_filter_top_k, speed_factor, seed, split_text, chunk_size, transcript,
max_tokens. All defaults are the wrapper's Pydantic blessed values (cfg 3.0 /
temp 1.3 / top_p 0.95 / top_k 35 / speed_factor 0.94 / chunk 300). Fields
grouped (basic/sampling/advanced). dia2 -> version 2 (field-shape change).

Voice stability: Dia2 samples a random speaker per call unless anchored. The
43 curated voices baked at /app/voices aren't reachable from /tts's clone path
(reference_audio dir only), so they're staged into reference_audio; the
clone_reference_filename picker now sources /get_reference_files. voice_mode=
clone + a reference filename pins voice/gender. Verified /tts clone end-to-end
(HTTP 200, Ogg/Opus 24 kHz). README documents the staging + two-instance shape.
This commit is contained in:
vh
2026-05-31 15:18:53 -07:00
parent e97d80cb8e
commit 55602b7251
2 changed files with 260 additions and 60 deletions
+45 -10
View File
@@ -54,15 +54,50 @@ scripts/deploy-stack.sh irv-ml1 dia
First boot pulls the checkpoint (~6-10 GB) into `DIA_CACHE_DIR` and can
take several minutes; the healthcheck's 600 s `start_period` covers it.
## Deployment shape (as of 2026-05-31)
This stack now runs **two fixed-model instances** from `local/dia:v2`
(the dia2-capable image — see [`dia2-image/Dockerfile`](dia2-image/Dockerfile)):
| service | model | port | notes |
|---|---|---|---|
| `dia2-2b` | nari-labs/Dia2-2B | 8200 | highest quality |
| `dia2-1b` | nari-labs/Dia2-1B | 8202 | streaming / lower latency |
The wrapper is single-model and ignores per-request model selection, so
one fixed instance per model is the only way to offer both as real
asset-engine choices. Legacy Dia 1.6B was retired. `local/dia:v2` is built
in two stages: upstream wrapper → `local/dia:v1`, then `dia2-image/` layers
in the dia2 package + its deps. Each instance pins its model via a mounted
`/opt/docker/conf/dia2-*/config.yaml`.
## Voices — stabilizing the random-voice behavior
Dia2 samples a **random speaker (random gender) per generation** unless
anchored (per the dia2 README: "voices vary per generation … use with
prefix … for stable output"). To pin a voice, use the richer **`/tts`**
endpoint with `voice_mode: clone` + a `clone_reference_filename`.
The image bakes **43 curated voices** at `/app/voices` (singles +
`[S1]/[S2]` dialogue pairs like `Abigail_Taylor.wav`), but `/tts`'s clone
path only reads the **reference_audio** dir — so they're **staged** into it:
```bash
# one-time per host (writes through the shared /worktank/dia/reference_audio mount):
docker exec dia2-2b sh -c 'cp -n /app/voices/* /app/reference_audio/'
```
After staging, `GET /get_reference_files` lists them and asset-engine's
`clone_reference_filename` picker (sourced from that endpoint) offers a
stable, known voice. Restic-included, so it survives once staged.
## Notes
- **Pin `DIA_SHA`** to a full 40-char commit before relying on this —
`.env.example` ships `main` for convenience, but `main` is not
reproducible (chatterbox learned this when an upstream restructure
broke its `main` build).
- **Model switching** is config.yaml-driven in the wrapper, or live from
the Web UI at `http://10.100.79.3:8200/`. To pin a non-default model
declaratively, mount a host `config.yaml` (see the commented volume in
`compose.yaml`).
- Endpoints: `/v1/audio/speech` (OpenAI-compat), `/health` (liveness),
`/api/model-status` (download/load progress), `/api/model-info`.
- **Model switching** within an instance is config.yaml-driven (the mounted
`config.yaml` pins `model.repo_id`); the Web UI can hot-swap live but only
the mounted config survives recreate.
- Endpoints: `/tts` (rich: cfg_scale/temperature/top_p/cfg_filter_top_k/
voice_mode/clone — what asset-engine targets), `/v1/audio/speech`
(OpenAI-compat; its `voice` param also resolves predefined voices by name),
`/get_reference_files`, `/get_predefined_voices`, `/health`,
`/api/model-status`, `/api/model-info`.