14f052461e
New stack scaffolding for the Fish quantized-realtime experiment. Not deployed yet — this commit lands the canonical files; deploy follows. Architecture decisions made in Phase 1: * CUDA backend, NOT Vulkan. s2.cpp's CMakeLists exposes both -DS2_VULKAN and -DS2_CUDA; the most recent upstream commit (2026-04-12) was specifically about CUDA improvements, and CUDA on the A6000 will be substantially faster than Vulkan for ML matmul. -DS2_CUDA=ON in the Dockerfile build args. * Pinned to s2.cpp commit e48ce8e02d8335bd9a0ba94679f605724b31d12 (2026-04-12 HEAD of main). Repo is alpha software per README; pin tightly so future churn doesn't break our build. Bump deliberately when wanting upstream improvements. * Multi-stage Dockerfile: nvidia/cuda:12.6.0-devel for build (needs CMake + ninja + git + the CUDA toolchain) → nvidia/cuda:12.6.0-runtime for serve (slimmer; just the s2 binary + GGML libs + a small Python shim). Cuts image size by ~50% vs single-stage devel. * FastAPI shim (server.py) wraps s2.cpp CLI in Fish's `/v1/tts` contract so the same bench harness + clients work against fish-cpp with no changes. Per-request flow: decode optional reference WAV from base64 → write to temp → subprocess.run the s2 binary → stream resulting WAV back. Adds ~50-100ms per-request fork+exec overhead; negligible vs the multi-second generation cost. * `streaming: true` accepted in request body but IGNORED — s2.cpp writes a complete WAV before returning, so chunked output isn't available. Unlike fish-s2 (HF wrapper) where streaming drops TTFB to 26ms, fish-cpp's TTFB ≈ total wall time. Speed depends entirely on raw generation throughput. * q6_k as default quant — sweet spot per typical GGUF guidance: near-bf16 quality at ~5GB. Other variants (q4_k_m, q5_k_m, q8_0, f16) selectable via FISH_CPP_MODEL env. * Pinned to GPU 1 (A6000) by default to share with fish-s2 for direct A/B benching. q6_k weights ~5GB + runtime ~3GB ≈ 8GB — comfortable on either GPU. * Port 8199 (next free in the irv-ml1 TTS slate). Phase 2 (next) is the actual deploy + first build. Reserved 30-45 min for cold-cache build + weights pull.
56 lines
2.3 KiB
YAML
56 lines
2.3 KiB
YAML
# fish-cpp — Fish s2-pro served via s2.cpp (pure C++/GGML inference,
|
||
# CUDA backend) with a tiny FastAPI shim exposing Fish's /v1/tts
|
||
# contract. Built locally from rodrigomatta/s2.cpp + a pinned SHA.
|
||
#
|
||
# Why this stack alongside fish-s2:
|
||
# * fish-s2 (HF transformers wrapper): ~7-8 s TTFB, 0.78× realtime.
|
||
# Phenomenal quality but sub-realtime, buffer-underruns on long
|
||
# phrases even with streaming.
|
||
# * fish-cpp (this): targets ~2-5× speedup from the GGML inference
|
||
# path + q6_k quantization. Goal: hit realtime for streaming use.
|
||
#
|
||
# CAVEATS:
|
||
# * s2.cpp is alpha software (per upstream README). Pin the SHA;
|
||
# don't track main blindly.
|
||
# * Per-request subprocess spawn — each /v1/tts call forks the s2
|
||
# binary. Adds ~50-100 ms over a long-running daemon. Negligible
|
||
# vs the multi-second generation cost.
|
||
# * No streaming — s2.cpp writes a complete WAV before returning,
|
||
# so the shim's `streaming: true` field is accepted but ignored.
|
||
# TTFB ≈ total wall time (fast generation is the only path to
|
||
# low-latency here, not chunked output).
|
||
|
||
services:
|
||
fish-cpp:
|
||
image: local/fish-cpp:${FISH_CPP_TAG}
|
||
build:
|
||
context: .
|
||
dockerfile: Dockerfile
|
||
args:
|
||
S2_CPP_SHA: ${FISH_CPP_S2_SHA:-e48ce8e02d8335bd9a0ba94679f605724b31d123}
|
||
container_name: fish-cpp
|
||
restart: unless-stopped
|
||
runtime: nvidia
|
||
ports:
|
||
- "${FISH_CPP_BIND:-0.0.0.0}:${FISH_CPP_PORT}:8000"
|
||
environment:
|
||
- NVIDIA_VISIBLE_DEVICES=${FISH_CPP_GPU_DEVICES:-1}
|
||
- FISH_CPP_MODEL=${FISH_CPP_MODEL:-s2-pro-q6_k.gguf}
|
||
- FISH_CPP_TOKENIZER=tokenizer.json
|
||
- FISH_CPP_DEVICE=0
|
||
volumes:
|
||
- ${FISH_CPP_WEIGHTS_DIR}:/weights:ro
|
||
- ${FISH_CPP_REFERENCE_DIR}:/references:ro
|
||
healthcheck:
|
||
test: ["CMD-SHELL", "python3 -c \"import urllib.request,sys; sys.exit(0 if urllib.request.urlopen('http://127.0.0.1:8000/v1/health', timeout=5).status==200 else 1)\""]
|
||
interval: 30s
|
||
timeout: 10s
|
||
retries: 3
|
||
start_period: 60s
|
||
labels:
|
||
- homepage.group=AI Systems
|
||
- homepage.name=Fish (s2.cpp)
|
||
- homepage.icon=mdi-fish
|
||
- homepage.description=Fish s2-pro via s2.cpp/GGML — quantized for realtime (irv-ml1)
|
||
- homepage.href=http://10.100.79.3:${FISH_CPP_PORT}
|