parakeet + cosyvoice: add stacks + deploy to irv-ml1
Two new speech stacks on irv-ml1, both on the /worktank/<stack>/
pattern, no tnet (irv-ml1 is local-endpoints-only for now).
parakeet — ASR via Shadowfita/parakeet-tdt-0.6b-v2-fastapi:
- docker buildx git context pinned to SHA 31c5652; no source
vendored. Rebuild on SHA bump.
- GPU-capable FastAPI + Silero VAD + WS streaming.
- API: POST /transcribe, WS /ws/transcribe, GET /healthz. Not the
literal OpenAI `/v1/audio/transcriptions` path — note in README.
- HF cache at /worktank/parakeet/models/ (excluded from restic).
- Build ~158s first time; steady-state start ~40s.
cosyvoice — TTS via neosun/cosyvoice:v1.3.2 shipping
Fun-CosyVoice3-0.5B-2512 (CosyVoice 3, chosen over v2 for the
expanded 5,000-hour instruction-following data covering emotions,
speed, tones, dialects, accents, role-playing; ~150ms streaming
TTFB matches v2). API: /v1/audio/speech (OpenAI drop-in),
/v1/voices/create (cloning), /health.
- Host port 8190 (container 8188; host 8188 already taken by comfyui).
- /worktank/cosyvoice/{voices,input,output}/; voices include in
restic (precious — reproducing a clone needs the original ref
audio), input+output excluded (scratch).
- Model weights (~2-3 GB) live inside image layer; re-download on
tag bump, persist across `compose up -d`.
Both healthy on first deploy.
This commit is contained in:
@@ -0,0 +1,104 @@
|
||||
# CosyVoice
|
||||
|
||||
Multilingual expressive TTS with zero-shot voice cloning, served by
|
||||
the `neosun/cosyvoice` wrapper around FunAudioLLM's
|
||||
Fun-CosyVoice3-0.5B-2512.
|
||||
|
||||
**Server:** irv-ml1 (Irvine, WireGuard-only)
|
||||
**Port:** 8190 (configurable via `.env`; container listens on 8188
|
||||
internally but host port moved to avoid ComfyUI's 8188)
|
||||
**GPUs:** both exposed (`NVIDIA_VISIBLE_DEVICES=all`); image reads
|
||||
`CUDA_VISIBLE_DEVICES` for pinning
|
||||
**Upstream wrapper:** [neosun100/cosyvoice-docker](https://github.com/neosun100/cosyvoice-docker)
|
||||
**Upstream model:** [FunAudioLLM/Fun-CosyVoice3-0.5B-2512](https://huggingface.co/FunAudioLLM/Fun-CosyVoice3-0.5B-2512)
|
||||
|
||||
## Why CosyVoice 3 (and not Kokoro / v2)
|
||||
|
||||
Emotive content was the deal-breaker for Kokoro. CosyVoice 3 extended
|
||||
the v2 instruction-following dataset from 1,500 → 5,000 hours
|
||||
specifically covering emotions, speed, tones, dialects, accents, and
|
||||
role-playing. Streaming TTFB stays at ~150 ms.
|
||||
|
||||
Two ways to request emotion / style:
|
||||
|
||||
- **XML tags** — `<angry>That's my line!</angry>`, `<sad>…</sad>`,
|
||||
`<fast>…</fast>`, `<slow>…</slow>`, `<peppa>…</peppa>`,
|
||||
`<robot>…</robot>`
|
||||
- **Instruction prompts** — `You are a helpful assistant. 请用尽可能快地语速说一句话。<|endofprompt|>`
|
||||
gives finer control via natural-language directives in the
|
||||
instruction channel
|
||||
|
||||
Language center of gravity is Mandarin + Cantonese (18+ Chinese
|
||||
dialects) and then 8 other languages (English, Japanese, Korean,
|
||||
German, Spanish, French, Italian, Russian). English works fine but
|
||||
don't expect ElevenLabs-grade English prosody polish — ear-test with
|
||||
your actual content.
|
||||
|
||||
## API endpoints
|
||||
|
||||
| Method + path | Purpose |
|
||||
|---|---|
|
||||
| `POST /v1/audio/speech` | OpenAI drop-in for TTS |
|
||||
| `POST /v1/voices/create` | Clone a speaker from reference audio (auto-transcription via built-in ASR) |
|
||||
| `GET /v1/voices` | List cloned voices by `voice_id` |
|
||||
| `GET /health` | Health probe (used by docker healthcheck) |
|
||||
|
||||
## Path layout
|
||||
|
||||
| Host path | Container path | Purpose | Restic? |
|
||||
|---|---|---|---|
|
||||
| `/worktank/cosyvoice/voices/` | `/data/voices` | Cloned speaker profiles | **included** (precious — reproducing a clone needs the original reference audio) |
|
||||
| `/worktank/cosyvoice/input/` | `/data/input` | Scratch for uploaded source audio | excluded |
|
||||
| `/worktank/cosyvoice/output/` | `/data/output` | Synthesized clips | excluded (regenerable) |
|
||||
|
||||
**Model weights (~2–3 GB) are NOT bind-mounted.** The image places
|
||||
them at `pretrained_models/Fun-CosyVoice3-0.5B/` inside the container
|
||||
on first run. Docker's image-layer cache keeps them across normal
|
||||
`compose up -d` recreates; a `docker image rm` or tag bump re-downloads.
|
||||
|
||||
## First-time deploy on irv-ml1
|
||||
|
||||
```bash
|
||||
# 1. Push compose + env template
|
||||
scripts/deploy-stack.sh irv-ml1 cosyvoice
|
||||
|
||||
# 2. Create the host dirs. One-time sudo — /worktank is root-owned.
|
||||
ssh -t irv-ml1 'sudo mkdir -p /worktank/cosyvoice/{voices,input,output} && \
|
||||
sudo chown -R lkraven:lkraven /worktank/cosyvoice'
|
||||
|
||||
# 3. Pull + up. First `up` downloads ~2–3 GB of model weights inside
|
||||
# the container; allow a few minutes before the health probe passes.
|
||||
ssh irv-ml1 '
|
||||
cd /opt/docker/compose/cosyvoice && \
|
||||
cp -n .env.example .env && \
|
||||
docker compose config >/dev/null && \
|
||||
docker compose pull && \
|
||||
docker compose up -d && \
|
||||
docker compose logs -f --tail=30
|
||||
'
|
||||
```
|
||||
|
||||
## Smoke test
|
||||
|
||||
```bash
|
||||
# From the workstation over WG — OpenAI-compatible request
|
||||
curl -X POST http://10.100.79.3:8190/v1/audio/speech \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"model":"cosyvoice","voice":"default","input":"<angry>Hello there.</angry>","response_format":"wav"}' \
|
||||
-o /tmp/out.wav
|
||||
```
|
||||
|
||||
## Version bump
|
||||
|
||||
```bash
|
||||
# Pick a new tag from https://hub.docker.com/r/neosun/cosyvoice/tags
|
||||
ssh irv-ml1 '
|
||||
cd /opt/docker/compose/cosyvoice && \
|
||||
sed -i "s/^COSYVOICE_VERSION=.*/COSYVOICE_VERSION=<new-tag>/" .env && \
|
||||
docker compose pull && \
|
||||
docker compose up -d
|
||||
'
|
||||
```
|
||||
|
||||
Voices / input / output persist across bumps. Model weights inside
|
||||
the image re-download on first run of the new tag.
|
||||
Reference in New Issue
Block a user