s2.cpp's README example uses `-v 0` which is `--vulkan 0` (Vulkan device 0), easy to misread as "voice 0". The shim copied that verbatim, so even after fixing the libcuda.so build problem AND the libgomp.so runtime dep, every synthesis ran on CPU because the wrong backend was selected. Direct verification: `[Model] NPU not compiled, falling back to CPU` in stderr; nvidia-smi showed no s2 process; bench timed out at 60s on phrases that fish-s2 (HF, GPU) does in 7s. s2.cpp's CLI: -v <id> = --vulkan <device> -c <id> = --cuda <device> -M = --metal (Apple Silicon) Switched the shim to `-c 0`. The CUDA backend IS in the build (-DS2_CUDA=ON worked, libggml-cuda.so links fine per ldd, libcuda.so.1 mounts at runtime via NVIDIA container runtime) — just wasn't being told to use it.
fish-cpp
Fish s2-pro served via s2.cpp
— a pure C++/GGML inference engine for Fish s2-pro, with weights from
rodrigomt/s2-pro-gguf.
Wrapped by a tiny FastAPI shim exposing Fish's /v1/tts HTTP contract
so it slots into the same bench harness + client patterns as fish-s2.
Why this stack alongside fish-s2
fish-s2 (HF transformers wrapper) measured at 0.78× realtime on
the A6000 — phenomenal quality but sub-realtime, meaning streaming
clients hit buffer underruns on phrases longer than ~3-4 seconds of
audio. fish-cpp targets the same s2-pro architecture but runs it
through s2.cpp's C++/GGML/CUDA inference path with q6_k quantization
— typically 2-5× faster than HF transformers for equivalent precision
(GGML is what makes llama.cpp fast).
Goal: hit ≥ 1× realtime on the A6000 so streaming actually flows without stutters, while keeping Fish's quality near-equivalent to BF16.
Status: alpha
s2.cpp is alpha software per its upstream README. Pin the SHA in
.env, don't track main blindly — community alpha projects break
weekly.
What works (and what doesn't) vs fish-s2
| feature | fish-s2 (HF) | fish-cpp (this) |
|---|---|---|
/v1/tts POST endpoint |
✓ | ✓ |
text body field |
✓ | ✓ |
references body field (cloning) |
✓ | ✓ (single ref only) |
streaming: true |
✓ (TTFB → 26 ms) | ✗ accepted but ignored |
| Paralinguistic tags | ✓ | should work (same weights) |
| Quantization | bf16 | q4_k_m / q5_k_m / q6_k (default) / q8_0 |
| Realtime factor | 0.78× | targeting 1-1.5× |
The streaming gap matters: fish-s2 with streaming: true returns the
first audio chunk in 26 ms (perceived latency feels instant). fish-cpp
returns nothing until generation completes. So fish-cpp's appeal is
RAW THROUGHPUT, not perceived latency. Combined with realtime+
generation, total wall time stays low enough that polling clients
don't notice.
API
OpenAPI shape mirrors fish-s2:
# Basic — text only, default voice
curl -fsS -X POST http://10.100.79.3:8199/v1/tts \
-H 'Content-Type: application/json' \
-d '{"text":"Hello there."}' \
> out.wav
# With voice cloning — base64-encoded reference audio inline
B64=$(base64 -w 0 /worktank/fish-cpp/references/glados.wav) # on irv-ml1
echo "{\"text\":\"Welcome.\",\"references\":[{\"audio\":\"$B64\",\"text\":\"transcript\"}]}" \
| curl -fsS -X POST http://10.100.79.3:8199/v1/tts \
-H 'Content-Type: application/json' --data-binary @- \
> out.wav
Health probe at GET /v1/health.
Deploy
scripts/elway irv-ml1 --playbook playbooks/deploy-fish-cpp.yaml
Cold deploy ~30-45 min: ~10 min image build (CUDA dev toolchain + CMake + s2.cpp compile), ~3 min weights pull (~5 GB for q6_k + 12 MB tokenizer), ~1 min container boot.
Hardware footprint
- VRAM: ~8 GB practical for q6_k (5 GB weights + 3 GB runtime).
Pinned to GPU 1 (A6000) by default to share with
fish-s2for direct A/B comparison. Could also run on GPU 0 (3090) with room to spare. - Disk: ~5 GB for q6_k checkpoint + tokenizer.
Bench plan
Same 3-phrase suite as the other TTS:
P1 = "Hello, this is a test of the voice synthesis system. The quick brown fox jumps over the lazy dog."
P2 = "Oh my god, I cannot believe what just happened. That was absolutely incredible!"
P3 = "What the hell is going on. This is some bullshit and I am not putting up with it."
Compare:
- TTFB / total wall-clock per phrase
- audio-seconds / wall-seconds (realtime factor)
- Quality (ear test) vs
fish-s2BF16 baseline
If fish-cpp lands ≥ 1× realtime AND the q6_k quality holds up under ear test, this stack becomes the default Fish path. fish-s2 stays deployed for paralinguistic tag fidelity reference + streaming (if that turns out to matter for any specific use case).
Lessons learned
(Populate after deploy iteration.)