Move the ~22-service flat "AI Systems" group off the Main tab into a new four-tab layout (Main / AI / Infrastructure / Toolchain). The AI tab sorts the inference fleet by function into seven groups: AI - Inference gen, char-rp, char-rp-reasoning, Granite summarizer AI - Eval & Retrieval Selene, Skywork Reward, Qwen3 rerank/embed, image-bench AI - Gateways & Chat LiteLLM, Asset Engine, Gateway Chat, Open WebUI, ... AI - Speech (TTS) Chatterbox Fast, Kokoro, mOrpheus AI - Audio Tools Parakeet ASR, YT Voice Clipper AI - Image & Media ComfyUI, Arbo AI - Dormant stopped rollback seats + retired auditions Relabel each stack's homepage.group so canonical stacks/ matches the live containers on ana-ml2, ana-docker, and irv-ml1. Dormant stacks were refreshed with `docker compose up --no-start` so they carry the new label while staying stopped (compose-start rollback preserved). settings.yaml drives tab/order/ columns; services.yaml and README updated to the new scheme.
CosyVoice
Multilingual expressive TTS with zero-shot voice cloning, served by
the neosun/cosyvoice wrapper around FunAudioLLM's
Fun-CosyVoice3-0.5B-2512.
Server: irv-ml1 (Irvine, WireGuard-only)
Port: 8190 (configurable via .env; container listens on 8188
internally but host port moved to avoid ComfyUI's 8188)
GPUs: both exposed (NVIDIA_VISIBLE_DEVICES=all); image reads
CUDA_VISIBLE_DEVICES for pinning
Upstream wrapper: neosun100/cosyvoice-docker
Upstream model: FunAudioLLM/Fun-CosyVoice3-0.5B-2512
Why CosyVoice 3 (and not Kokoro / v2)
Emotive content was the deal-breaker for Kokoro. CosyVoice 3 extended
the v2 instruction-following dataset from 1,500 → 5,000 hours
specifically covering emotions, speed, tones, dialects, accents, and
role-playing. Streaming is available on /api/tts at ~150 ms TTFB
(not on the OpenAI-compat /v1/audio/speech path — see API table
below).
Language center of gravity is Mandarin + Cantonese (18+ Chinese dialects) and then 8 other languages (English, Japanese, Korean, German, Spanish, French, Italian, Russian). English works fine but don't expect ElevenLabs-grade English prosody polish — ear-test with your actual content.
API endpoints
| Method + path | Purpose | Streaming? |
|---|---|---|
POST /v1/voices/create |
Clone a speaker from reference audio | n/a |
GET /v1/voices |
List cloned voices | n/a |
GET /v1/voices/{voice_id} |
Inspect one voice | n/a |
DELETE /v1/voices/{voice_id} |
Remove a cloned voice | n/a |
GET /v1/models |
Lists cosyvoice-v3 + cosyvoice-v2 (both bundled) |
n/a |
POST /v1/audio/speech |
OpenAI drop-in TTS | no — single response body |
POST /api/tts |
Native TTS with finer control | yes when stream=true |
POST /api/tts/async |
Job-based TTS; poll /api/task/{task_id} |
no |
GET /health |
Health probe | n/a |
Web UI at / |
Gradio-style voice-cloning + inference UI | — |
Notably not exposed:
- No
preset_voices/spk_id—/api/speakersalways returns[]in this image. CosyVoice's built-in SFT speakers (中文男/女 etc.) are not surfaced. Every synthesis needs a voice that came from/v1/voices/create. - No WebSocket / SSE endpoints — the upstream CosyVoice 3 model
supports bi-directional streaming at ~150 ms TTFB, but this wrapper
doesn't surface WS. Use
/api/ttswithstream=truefor HTTP chunked streaming (still ~150 ms TTFB).
Path layout
| Host path | Container path | Purpose | Restic? |
|---|---|---|---|
/worktank/cosyvoice/voices/ |
/data/voices |
Cloned speaker profiles | included (precious — reproducing a clone needs the original reference audio) |
/worktank/cosyvoice/input/ |
/data/input |
Scratch for uploaded source audio | excluded |
/worktank/cosyvoice/output/ |
/data/output |
Synthesized clips | excluded (regenerable) |
Model weights (~2–3 GB) are NOT bind-mounted. The image places
them at pretrained_models/Fun-CosyVoice3-0.5B/ inside the container
on first run. Docker's image-layer cache keeps them across normal
compose up -d recreates; a docker image rm or tag bump re-downloads.
First-time deploy on irv-ml1
# 1. Push compose + env template
scripts/deploy-stack.sh irv-ml1 cosyvoice
# 2. Create the host dirs. One-time sudo — /worktank is root-owned.
ssh -t irv-ml1 'sudo mkdir -p /worktank/cosyvoice/{voices,input,output} && \
sudo chown -R lkraven:lkraven /worktank/cosyvoice'
# 3. Pull + up. First `up` downloads ~2–3 GB of model weights inside
# the container; allow a few minutes before the health probe passes.
ssh irv-ml1 '
cd /opt/docker/compose/cosyvoice && \
cp -n .env.example .env && \
docker compose config >/dev/null && \
docker compose pull && \
docker compose up -d && \
docker compose logs -f --tail=30
'
Gotchas (read before first use)
- No default voice ships with the image. You must clone one via
/v1/voices/createbefore/v1/audio/speechcan return audio. Calling/v1/audio/speechwithvoice="default"(or any unregistered id) returns HTTP 400; curl writing the error JSON into a.wavfile is what the 124-byte phantom output was in an early smoke test. - Reference audio must be ≤ 30 seconds. The frontend asserts
speech.shape[1] / 16000 <= 30— longer clips are accepted at upload time but synthesis raisesAssertionError: do not support extract speech token for audio longer than 30s, producing a 500. Trim with ffmpeg before/v1/voices/create:ffmpeg -i src.wav -ss 0 -t 25 -ac 1 -ar 16000 ref.wav - Provide an accurate transcript with the clone. The
textfield of/v1/voices/createdefaults to empty, in which case the wrapper auto-transcribes via Fun-ASR-Nano. That often misses names / technical terms. Supplying your own transcript (e.g., from the Parakeet stack at:8765) produces better clones. - Voice id vs. name. The response's
voice_idis what you pass to subsequent/v1/audio/speechcalls. Thenameis purely for humans — reusing a name with a new upload creates a new voice id, it doesn't overwrite. - Model choice.
cosyvoice-v3is default and best;cosyvoice-v2is available as a fallback if you have voices tuned against the v2 model's idiosyncrasies. - The
instructfield is Chinese-context-only — DON'T use it for English emotion control. Verified empirically 2026-04-24: any value ininstruct(English, Chinese, or empty hints inside it) routes through CosyVoice'sinference_instruct2which expects a directive in the upstream Chinese template (You are a helpful assistant. <directive>请用…<|endofprompt|>). English values likeinstruct: "speak with cold contempt"cause the model to vocalize the instruction text itself — you'll hear Chinese-flavored phonemes reading the English directive instead of yourinput. Workaround: use the XML inline tags (<angry>,<sad>,<surprised>,<fearful>,<fast>,<slow>,<whisper>,<shout>,<peppa>,<robot>) for English emotion control. They're parsed as control tokens, not content. See upstream issue [reference TBD].
Smoke test (full clone → synthesize)
# 1. Trim reference audio to ≤30s (16 kHz mono is cleanest input)
ffmpeg -i glados.wav -ss 0 -t 25 -ac 1 -ar 16000 glados_25s.wav
# 2. Transcribe the trimmed clip via Parakeet for an accurate text field
curl -s -F "file=@./glados_25s.wav" http://10.100.79.3:8765/transcribe \
| tee glados_25s.txt
# 3. Clone the voice
curl -X POST http://10.100.79.3:8190/v1/voices/create \
-F "audio=@./glados_25s.wav" \
-F "name=glados" \
-F "text=<paste from glados_25s.txt>"
# → {"voice_id": "<id>", "name": "glados", ...}
# 4. Synthesize via OpenAI-compat endpoint (single-shot).
# For English emotion control, use the XML inline tag — NOT the
# `instruct` field (that's Chinese-context only; see Gotchas).
curl -X POST http://10.100.79.3:8190/v1/audio/speech \
-H 'Content-Type: application/json' \
-d '{
"model": "cosyvoice-v3",
"voice": "<voice_id>",
"input": "<angry>I did not ask for visitors today.</angry>",
"response_format": "wav",
"speed": 1.0
}' \
-o out.wav
# 5. Or synthesize via native endpoint with streaming — curl writes
# audio chunks to the file as they arrive (~150 ms TTFB).
# `instruct_text` here has the same Chinese-only caveat as the
# OpenAI path's `instruct` — leave it empty for English work.
curl -N -X POST http://10.100.79.3:8190/api/tts \
-F "text=<angry>I did not ask for visitors today.</angry>" \
-F "voice=<voice_id>" \
-F "mode=zero_shot" \
-F "speed=1.0" \
-F "stream=true" \
-o out_stream.pcm
Field naming differs between the two synthesis endpoints:
/v1/audio/speech (JSON) |
/api/tts (multipart) |
|---|---|
input |
text |
instruct |
instruct_text |
voice |
voice (same) |
response_format |
— (always wav/pcm chunks) |
| — | stream (bool) |
| — | mode (zero_shot / instruct / sft) |
| — | prompt_wav + prompt_text (inline ref audio; skip /v1/voices/create) |
Emotion / style control
For English content, use XML inline tags. Do not use the instruct
field — see the Gotcha above; instruct routes through CosyVoice's
Chinese-context inference path and English values get spoken aloud
instead of treated as direction.
| Syntax | Where | Example | English-safe? |
|---|---|---|---|
| XML tags in the text | inline in input / text |
<angry>You're late.</angry> |
✅ |
| Natural-language directive | instruct / instruct_text |
"用愤怒的语气说" (Chinese) |
⚠️ Chinese only |
Tag vocabulary includes <angry>, <sad>, <surprised>,
<fearful>, <fast>, <slow>, <whisper>, <shout>, plus
character/style tags like <peppa>, <robot>. Tags can be nested
or applied at sentence granularity (the model accepts paragraph-
length input with multiple tagged segments).
Version bump
# Pick a new tag from https://hub.docker.com/r/neosun/cosyvoice/tags
ssh irv-ml1 '
cd /opt/docker/compose/cosyvoice && \
sed -i "s/^COSYVOICE_VERSION=.*/COSYVOICE_VERSION=<new-tag>/" .env && \
docker compose pull && \
docker compose up -d
'
Voices / input / output persist across bumps. Model weights inside the image re-download on first run of the new tag.