c5bbb90980
reference_id=<name> resolves against the DIRECTORY references/<name>/ (audio + same-basename .lab), not a flat references/<name>.wav. Voices were staged flat with the per-name dirs left empty, so every reference_id resolved to nothing and Fish fell back to its default speaker — every dropdown voice produced byte-identical audio (proven: Abigail == Imogen == no-ref, same text+seed). This was the real "no accent" root cause, independent of the asset-engine "undefined" select bug. Server fix (applied to irv-ml1): populated references/<name>/<name>.wav + <name>.lab for all 32 voices; re-test confirms Imogen/Eleanor/ Beatrice/Abigail/no-ref now all distinct. Durable hardening + record correction: - playbook: normalize-layout step (flat <name>.wav -> nested dir, cp -u idempotent, when-gated on count mismatch) + an A/B verify gate that hard-fails the deploy if two reference_ids yield identical output. - services.yaml: correct the reference_id resolution doc (dir + .lab, not flat wav). - README + persistent-memory: correct the "reference_id-by-name is THE working path, verified" claim — it was a no-op until this fix; the prior ECAPA 0.79 result came through the inline base64 path.
154 lines
6.0 KiB
Markdown
154 lines
6.0 KiB
Markdown
# Fish Audio S2-Pro
|
|
|
|
[fishaudio/s2-pro](https://huggingface.co/fishaudio/s2-pro) — the most
|
|
expressive open-source TTS model as of 2026-04, served via the
|
|
official [fishaudio/fish-speech](https://github.com/fishaudio/fish-speech)
|
|
inference engine.
|
|
|
|
## Why this stack exists
|
|
|
|
Three of the existing TTS already cover the basics — Kokoro for raw
|
|
speed, Chatterbox for speed-with-cloning, IndexTTS-2 for precision
|
|
emotion control. Fish Audio S2-Pro fills a different slot:
|
|
**dramatically richer paralinguistic control via natural-language
|
|
tags** (15,000+ vs Chatterbox Turbo's 9 fixed tags), with comparable
|
|
latency (~150 ms streaming) and voice cloning.
|
|
|
|
Released March 9, 2026; we missed it during the original irv-ml1
|
|
build-out in early April.
|
|
|
|
| | use case |
|
|
|---|---|
|
|
| **Fish Audio S2-Pro** | richest emotive / paralinguistic English TTS — 15k+ tags |
|
|
| Kokoro | low-latency English, fixed voice library |
|
|
| Chatterbox Turbo | low-latency English w/ cloning + 9 paralinguistic tags |
|
|
| IndexTTS-2 | English voice cloning + emotion vector / text control |
|
|
| Qwen3-TTS-1.7B | English voice cloning (slow on official backend) |
|
|
| CosyVoice 3 | multilingual (Chinese-leaning) |
|
|
| VibeVoice 1.5B | long-form / multi-speaker dialogue |
|
|
|
|
## Architecture
|
|
|
|
Dual-AR (Slow + Fast):
|
|
- **Slow AR** operates along the time axis, predicts the primary
|
|
semantic codebook.
|
|
- **Fast AR** generates the remaining 9 residual codebooks per time
|
|
step, reconstructing fine-grained acoustic detail.
|
|
|
|
Trained on 10M+ hours of audio across 80+ languages with
|
|
reinforcement-learning alignment. Win rates per upstream:
|
|
|
|
| benchmark | S2-Pro |
|
|
|---|---|
|
|
| EmergentTTS-Eval paralinguistics | 91.61% |
|
|
| Blind A/B vs ElevenLabs Flash v2.5 (multilingual) | strong |
|
|
|
|
## Headline features
|
|
|
|
- **15,000+ paralinguistic / emotion tags** via natural language:
|
|
```
|
|
[laugh] [whispers] [super happy] [sigh] [excited] [heavy breathing]
|
|
[angry] [sleepy] [crying] [surprise] ...
|
|
```
|
|
Drop them inline in the input text. Different shape from
|
|
IndexTTS-2's 8-vector emotion control — this is "say it like this"
|
|
markup directly in the prompt, with a far larger vocabulary.
|
|
- **Voice cloning** from ~5-15 s reference WAV.
|
|
- **Multi-speaker / multi-turn** generation natively supported.
|
|
- **80+ languages** (English-strong, not Chinese-leaning like
|
|
CosyVoice).
|
|
|
|
## API
|
|
|
|
**Fish ships a custom API, NOT OpenAI-compatible.** The wrapper has
|
|
exactly one TTS endpoint (`POST /v1/tts`) plus liveness probes — no
|
|
`/v1/audio/speech`, no `/v1/audio/voices`, no `/v1/models`. Voice
|
|
cloning happens via the `references` field in the request body
|
|
(pointing at files under `/app/references`).
|
|
|
|
```bash
|
|
# Minimal POST — text only, default voice.
|
|
curl -fsS -X POST http://10.100.79.3:8195/v1/tts \
|
|
-H 'Content-Type: application/json' \
|
|
-d '{"text":"Oh wow [super happy] I cannot believe it. [laugh] What a day."}' \
|
|
> out.wav
|
|
|
|
# With voice cloning — point at a reference file (drop the .wav into
|
|
# /worktank/fish-s2/references/ on the host first).
|
|
curl -fsS -X POST http://10.100.79.3:8195/v1/tts \
|
|
-H 'Content-Type: application/json' \
|
|
-d '{"text":"...", "references":[{"audio":"/app/references/glados.wav","text":"transcript of the reference"}]}' \
|
|
> out.wav
|
|
```
|
|
|
|
Other endpoints:
|
|
|
|
| path | purpose |
|
|
|---|---|
|
|
| `GET /v1/health` | liveness probe (used by our Docker healthcheck) |
|
|
| `GET /heartbeat` | alternate liveness signal |
|
|
| `GET /` | Swagger Editor UI for the OpenAPI spec |
|
|
|
|
The 200-line OpenAPI spec is rendered through Swagger Editor at the
|
|
root path; there's no `/openapi.json` endpoint exposed directly.
|
|
|
|
## Voice library
|
|
|
|
Named voices are selected via the `reference_id` field, which Fish
|
|
resolves against the **directory** `references/<name>/` — NOT a flat
|
|
`references/<name>.wav`. Stage each voice as:
|
|
|
|
```
|
|
/worktank/fish-s2/references/<name>/<name>.wav # clean ~5-15 s, single speaker
|
|
/worktank/fish-s2/references/<name>/<name>.lab # transcript of that clip
|
|
```
|
|
|
|
A flat `references/<name>.wav` (or an empty `references/<name>/` dir)
|
|
is **silently ignored** — `reference_id` resolves to nothing and Fish
|
|
falls back to its default speaker, so every voice sounds identical.
|
|
This was the 2026-06-01 "no accent" root cause; the deploy playbook now
|
|
carries a normalize-layout step + an A/B smoke gate (`reference_id`
|
|
MUST change the output) to keep it from regressing. The wrapper scans
|
|
on request — no restart needed after adding a voice.
|
|
|
|
## Deploy
|
|
|
|
```bash
|
|
scripts/elway irv-ml1 --playbook playbooks/deploy-fish-s2.yaml
|
|
```
|
|
|
|
First boot pulls the s2-pro checkpoint (~9 GB BF16) into the HF
|
|
cache + warms torch.compile (adds ~60 s). Both are cached afterwards.
|
|
|
|
## Hardware footprint
|
|
|
|
- **VRAM**: ~17 GB practical, 24 GB recommended. Pinned to GPU 1
|
|
(RTX A6000) by default — plenty of headroom for long contexts and
|
|
large mmproj if a future checkpoint adds vision.
|
|
- **Disk**: ~11 GB for the s2-pro checkpoint
|
|
(codec.pth 1.9 GB + 2 safetensors shards 9 GB + tokenizer/config).
|
|
|
|
## Lessons learned during deploy (2026-04-27)
|
|
|
|
Took 5 iterations to land. Recording for next time:
|
|
|
|
1. **`dockerfile`** (lowercase, root) — doesn't exist. Fish doesn't
|
|
ship a plain Dockerfile.
|
|
2. **`dockerfile.dev`** — exists at root, but it's a 2-line wrapper
|
|
(`FROM ghcr.io/fishaudio/fish-speech:${VERSION}`) over a private
|
|
GHCR base image. Anonymous pulls 403.
|
|
3. **`docker/Dockerfile`** — the real build path (referenced by
|
|
upstream's `compose.base.yml`).
|
|
4. **Multi-stage default builds the wrong target.** Without
|
|
`target: server`, docker builds the last stage which is `webui`
|
|
(gradio only — no `start_server.sh`, container exits silently
|
|
rc=0 because the API entrypoint is missing).
|
|
5. **Fish doesn't auto-download checkpoints.** `start_server.sh`
|
|
validates `/app/checkpoints/s2-pro/` exists and exits clean if
|
|
not. The playbook now pre-pulls `fishaudio/s2-pro` (~11 GB) via a
|
|
one-shot `huggingface_hub.snapshot_download` container before
|
|
starting the service.
|
|
6. **API is NOT OpenAI-compatible.** Endpoint is `POST /v1/tts`, not
|
|
`/v1/audio/speech`. No `/v1/audio/voices` or `/v1/models`. Voice
|
|
cloning is via `references` field in the POST body.
|