diff --git a/playbooks/deploy-fish-s2.yaml b/playbooks/deploy-fish-s2.yaml index a6f13f2..b4c4746 100644 --- a/playbooks/deploy-fish-s2.yaml +++ b/playbooks/deploy-fish-s2.yaml @@ -103,10 +103,20 @@ verify: shell: curl -sf -o /dev/null http://localhost:{{ host_port }}/v1/health changed_when: "false" - - name: /v1/audio/voices returns valid JSON + - name: /v1/tts returns a real WAV (POST with text body) + # Fish's API is NOT OpenAI-compatible — there's no /v1/audio/speech + # and no /v1/audio/voices. The single TTS endpoint is POST /v1/tts + # with at minimum {"text":"..."} returning audio/wav. Voice cloning + # is via reference= field in the body (paths under /app/references). + # Verify by POST + asserting the response is a real RIFF WAV. shell: | - curl -sf http://localhost:{{ host_port }}/v1/audio/voices \ - | python3 -c "import json,sys; json.load(sys.stdin)" + out=$(mktemp --suffix=.wav) + curl -sf -X POST http://localhost:{{ host_port }}/v1/tts \ + -H 'Content-Type: application/json' \ + -d '{"text":"Verify."}' \ + -o "$out" --max-time 30 + file -b "$out" | grep -q '^RIFF.*WAVE' + rm -f "$out" changed_when: "false" - name: Container is running diff --git a/stacks/fish-s2/README.md b/stacks/fish-s2/README.md index 7965a65..3ef74bb 100644 --- a/stacks/fish-s2/README.md +++ b/stacks/fish-s2/README.md @@ -60,26 +60,37 @@ reinforcement-learning alignment. Win rates per upstream: ## API -OpenAI-compat at `http://10.100.79.3:8195`: +**Fish ships a custom API, NOT OpenAI-compatible.** The wrapper has +exactly one TTS endpoint (`POST /v1/tts`) plus liveness probes — no +`/v1/audio/speech`, no `/v1/audio/voices`, no `/v1/models`. Voice +cloning happens via the `references` field in the request body +(pointing at files under `/app/references`). ```bash -# Single-shot synthesis with paralinguistic tags inline. -curl -fsS -X POST http://10.100.79.3:8195/v1/audio/speech \ +# Minimal POST — text only, default voice. +curl -fsS -X POST http://10.100.79.3:8195/v1/tts \ -H 'Content-Type: application/json' \ - -d '{"model":"fish-s2","input":"Oh wow [super happy] I cannot believe it. [laugh] What a day.","voice":"glados","response_format":"wav"}' \ + -d '{"text":"Oh wow [super happy] I cannot believe it. [laugh] What a day."}' \ > out.wav -# Built-in voices. -curl http://10.100.79.3:8195/v1/audio/voices - -# Streaming. -curl -fsS -X POST http://10.100.79.3:8195/v1/audio/speech \ +# With voice cloning — point at a reference file (drop the .wav into +# /worktank/fish-s2/references/ on the host first). +curl -fsS -X POST http://10.100.79.3:8195/v1/tts \ -H 'Content-Type: application/json' \ - -d '{"model":"fish-s2","input":"long passage…","voice":"glados","stream":true}' \ - | mpv --no-cache - + -d '{"text":"...", "references":[{"audio":"/app/references/glados.wav","text":"transcript of the reference"}]}' \ + > out.wav ``` -WebUI at `/`. OpenAPI / docs at `/docs`. Healthcheck at `/v1/health`. +Other endpoints: + +| path | purpose | +|---|---| +| `GET /v1/health` | liveness probe (used by our Docker healthcheck) | +| `GET /heartbeat` | alternate liveness signal | +| `GET /` | Swagger Editor UI for the OpenAPI spec | + +The 200-line OpenAPI spec is rendered through Swagger Editor at the +root path; there's no `/openapi.json` endpoint exposed directly. ## Voice library @@ -102,4 +113,29 @@ cache + warms torch.compile (adds ~60 s). Both are cached afterwards. - **VRAM**: ~17 GB practical, 24 GB recommended. Pinned to GPU 1 (RTX A6000) by default — plenty of headroom for long contexts and large mmproj if a future checkpoint adds vision. -- **Disk**: ~9 GB for the model checkpoint + HF cache. +- **Disk**: ~11 GB for the s2-pro checkpoint + (codec.pth 1.9 GB + 2 safetensors shards 9 GB + tokenizer/config). + +## Lessons learned during deploy (2026-04-27) + +Took 5 iterations to land. Recording for next time: + +1. **`dockerfile`** (lowercase, root) — doesn't exist. Fish doesn't + ship a plain Dockerfile. +2. **`dockerfile.dev`** — exists at root, but it's a 2-line wrapper + (`FROM ghcr.io/fishaudio/fish-speech:${VERSION}`) over a private + GHCR base image. Anonymous pulls 403. +3. **`docker/Dockerfile`** — the real build path (referenced by + upstream's `compose.base.yml`). +4. **Multi-stage default builds the wrong target.** Without + `target: server`, docker builds the last stage which is `webui` + (gradio only — no `start_server.sh`, container exits silently + rc=0 because the API entrypoint is missing). +5. **Fish doesn't auto-download checkpoints.** `start_server.sh` + validates `/app/checkpoints/s2-pro/` exists and exits clean if + not. The playbook now pre-pulls `fishaudio/s2-pro` (~11 GB) via a + one-shot `huggingface_hub.snapshot_download` container before + starting the service. +6. **API is NOT OpenAI-compatible.** Endpoint is `POST /v1/tts`, not + `/v1/audio/speech`. No `/v1/audio/voices` or `/v1/models`. Voice + cloning is via `references` field in the POST body.