parakeet + cosyvoice: add stacks + deploy to irv-ml1

Two new speech stacks on irv-ml1, both on the /worktank/<stack>/
pattern, no tnet (irv-ml1 is local-endpoints-only for now).

parakeet — ASR via Shadowfita/parakeet-tdt-0.6b-v2-fastapi:
  - docker buildx git context pinned to SHA 31c5652; no source
    vendored. Rebuild on SHA bump.
  - GPU-capable FastAPI + Silero VAD + WS streaming.
  - API: POST /transcribe, WS /ws/transcribe, GET /healthz. Not the
    literal OpenAI `/v1/audio/transcriptions` path — note in README.
  - HF cache at /worktank/parakeet/models/ (excluded from restic).
  - Build ~158s first time; steady-state start ~40s.

cosyvoice — TTS via neosun/cosyvoice:v1.3.2 shipping
Fun-CosyVoice3-0.5B-2512 (CosyVoice 3, chosen over v2 for the
expanded 5,000-hour instruction-following data covering emotions,
speed, tones, dialects, accents, role-playing; ~150ms streaming
TTFB matches v2). API: /v1/audio/speech (OpenAI drop-in),
/v1/voices/create (cloning), /health.
  - Host port 8190 (container 8188; host 8188 already taken by comfyui).
  - /worktank/cosyvoice/{voices,input,output}/; voices include in
    restic (precious — reproducing a clone needs the original ref
    audio), input+output excluded (scratch).
  - Model weights (~2-3 GB) live inside image layer; re-download on
    tag bump, persist across `compose up -d`.

Both healthy on first deploy.
This commit is contained in:
vh
2026-04-23 23:40:36 -07:00
parent 06476745f1
commit 1a67370138
6 changed files with 410 additions and 0 deletions
+104
View File
@@ -0,0 +1,104 @@
# CosyVoice
Multilingual expressive TTS with zero-shot voice cloning, served by
the `neosun/cosyvoice` wrapper around FunAudioLLM's
Fun-CosyVoice3-0.5B-2512.
**Server:** irv-ml1 (Irvine, WireGuard-only)
**Port:** 8190 (configurable via `.env`; container listens on 8188
internally but host port moved to avoid ComfyUI's 8188)
**GPUs:** both exposed (`NVIDIA_VISIBLE_DEVICES=all`); image reads
`CUDA_VISIBLE_DEVICES` for pinning
**Upstream wrapper:** [neosun100/cosyvoice-docker](https://github.com/neosun100/cosyvoice-docker)
**Upstream model:** [FunAudioLLM/Fun-CosyVoice3-0.5B-2512](https://huggingface.co/FunAudioLLM/Fun-CosyVoice3-0.5B-2512)
## Why CosyVoice 3 (and not Kokoro / v2)
Emotive content was the deal-breaker for Kokoro. CosyVoice 3 extended
the v2 instruction-following dataset from 1,500 → 5,000 hours
specifically covering emotions, speed, tones, dialects, accents, and
role-playing. Streaming TTFB stays at ~150 ms.
Two ways to request emotion / style:
- **XML tags** — `<angry>That's my line!</angry>`, `<sad>…</sad>`,
`<fast>…</fast>`, `<slow>…</slow>`, `<peppa>…</peppa>`,
`<robot>…</robot>`
- **Instruction prompts** — `You are a helpful assistant. 请用尽可能快地语速说一句话。<|endofprompt|>`
gives finer control via natural-language directives in the
instruction channel
Language center of gravity is Mandarin + Cantonese (18+ Chinese
dialects) and then 8 other languages (English, Japanese, Korean,
German, Spanish, French, Italian, Russian). English works fine but
don't expect ElevenLabs-grade English prosody polish — ear-test with
your actual content.
## API endpoints
| Method + path | Purpose |
|---|---|
| `POST /v1/audio/speech` | OpenAI drop-in for TTS |
| `POST /v1/voices/create` | Clone a speaker from reference audio (auto-transcription via built-in ASR) |
| `GET /v1/voices` | List cloned voices by `voice_id` |
| `GET /health` | Health probe (used by docker healthcheck) |
## Path layout
| Host path | Container path | Purpose | Restic? |
|---|---|---|---|
| `/worktank/cosyvoice/voices/` | `/data/voices` | Cloned speaker profiles | **included** (precious — reproducing a clone needs the original reference audio) |
| `/worktank/cosyvoice/input/` | `/data/input` | Scratch for uploaded source audio | excluded |
| `/worktank/cosyvoice/output/` | `/data/output` | Synthesized clips | excluded (regenerable) |
**Model weights (~2–3 GB) are NOT bind-mounted.** The image places
them at `pretrained_models/Fun-CosyVoice3-0.5B/` inside the container
on first run. Docker's image-layer cache keeps them across normal
`compose up -d` recreates; a `docker image rm` or tag bump re-downloads.
## First-time deploy on irv-ml1
```bash
# 1. Push compose + env template
scripts/deploy-stack.sh irv-ml1 cosyvoice
# 2. Create the host dirs. One-time sudo — /worktank is root-owned.
ssh -t irv-ml1 'sudo mkdir -p /worktank/cosyvoice/{voices,input,output} && \
sudo chown -R lkraven:lkraven /worktank/cosyvoice'
# 3. Pull + up. First `up` downloads ~2–3 GB of model weights inside
# the container; allow a few minutes before the health probe passes.
ssh irv-ml1 '
cd /opt/docker/compose/cosyvoice && \
cp -n .env.example .env && \
docker compose config >/dev/null && \
docker compose pull && \
docker compose up -d && \
docker compose logs -f --tail=30
'
```
## Smoke test
```bash
# From the workstation over WG — OpenAI-compatible request
curl -X POST http://10.100.79.3:8190/v1/audio/speech \
-H 'Content-Type: application/json' \
-d '{"model":"cosyvoice","voice":"default","input":"<angry>Hello there.</angry>","response_format":"wav"}' \
-o /tmp/out.wav
```
## Version bump
```bash
# Pick a new tag from https://hub.docker.com/r/neosun/cosyvoice/tags
ssh irv-ml1 '
cd /opt/docker/compose/cosyvoice && \
sed -i "s/^COSYVOICE_VERSION=.*/COSYVOICE_VERSION=<new-tag>/" .env && \
docker compose pull && \
docker compose up -d
'
```
Voices / input / output persist across bumps. Model weights inside
the image re-download on first run of the new tag.