Files
esh-pfi-infrastructure/stacks/fish-cpp/compose.yaml
T
vh 14f052461e stacks/fish-cpp: Phase 1 — s2.cpp + GGML CUDA backend image, FastAPI shim, deploy playbook
New stack scaffolding for the Fish quantized-realtime experiment. Not
deployed yet — this commit lands the canonical files; deploy follows.

Architecture decisions made in Phase 1:
* CUDA backend, NOT Vulkan. s2.cpp's CMakeLists exposes both
  -DS2_VULKAN and -DS2_CUDA; the most recent upstream commit
  (2026-04-12) was specifically about CUDA improvements, and CUDA
  on the A6000 will be substantially faster than Vulkan for ML
  matmul. -DS2_CUDA=ON in the Dockerfile build args.

* Pinned to s2.cpp commit e48ce8e02d8335bd9a0ba94679f605724b31d12
  (2026-04-12 HEAD of main). Repo is alpha software per README;
  pin tightly so future churn doesn't break our build. Bump
  deliberately when wanting upstream improvements.

* Multi-stage Dockerfile: nvidia/cuda:12.6.0-devel for build (needs
  CMake + ninja + git + the CUDA toolchain) → nvidia/cuda:12.6.0-runtime
  for serve (slimmer; just the s2 binary + GGML libs + a small Python
  shim). Cuts image size by ~50% vs single-stage devel.

* FastAPI shim (server.py) wraps s2.cpp CLI in Fish's `/v1/tts`
  contract so the same bench harness + clients work against fish-cpp
  with no changes. Per-request flow: decode optional reference WAV
  from base64 → write to temp → subprocess.run the s2 binary → stream
  resulting WAV back. Adds ~50-100ms per-request fork+exec overhead;
  negligible vs the multi-second generation cost.

* `streaming: true` accepted in request body but IGNORED — s2.cpp
  writes a complete WAV before returning, so chunked output isn't
  available. Unlike fish-s2 (HF wrapper) where streaming drops TTFB
  to 26ms, fish-cpp's TTFB ≈ total wall time. Speed depends entirely
  on raw generation throughput.

* q6_k as default quant — sweet spot per typical GGUF guidance:
  near-bf16 quality at ~5GB. Other variants (q4_k_m, q5_k_m, q8_0,
  f16) selectable via FISH_CPP_MODEL env.

* Pinned to GPU 1 (A6000) by default to share with fish-s2 for
  direct A/B benching. q6_k weights ~5GB + runtime ~3GB ≈ 8GB —
  comfortable on either GPU.

* Port 8199 (next free in the irv-ml1 TTS slate).

Phase 2 (next) is the actual deploy + first build. Reserved 30-45 min
for cold-cache build + weights pull.
2026-04-28 01:06:14 -07:00

56 lines
2.3 KiB
YAML
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# fish-cpp — Fish s2-pro served via s2.cpp (pure C++/GGML inference,
# CUDA backend) with a tiny FastAPI shim exposing Fish's /v1/tts
# contract. Built locally from rodrigomatta/s2.cpp + a pinned SHA.
#
# Why this stack alongside fish-s2:
# * fish-s2 (HF transformers wrapper): ~7-8 s TTFB, 0.78× realtime.
# Phenomenal quality but sub-realtime, buffer-underruns on long
# phrases even with streaming.
# * fish-cpp (this): targets ~2-5× speedup from the GGML inference
# path + q6_k quantization. Goal: hit realtime for streaming use.
#
# CAVEATS:
# * s2.cpp is alpha software (per upstream README). Pin the SHA;
# don't track main blindly.
# * Per-request subprocess spawn — each /v1/tts call forks the s2
# binary. Adds ~50-100 ms over a long-running daemon. Negligible
# vs the multi-second generation cost.
# * No streaming — s2.cpp writes a complete WAV before returning,
# so the shim's `streaming: true` field is accepted but ignored.
# TTFB ≈ total wall time (fast generation is the only path to
# low-latency here, not chunked output).
services:
fish-cpp:
image: local/fish-cpp:${FISH_CPP_TAG}
build:
context: .
dockerfile: Dockerfile
args:
S2_CPP_SHA: ${FISH_CPP_S2_SHA:-e48ce8e02d8335bd9a0ba94679f605724b31d123}
container_name: fish-cpp
restart: unless-stopped
runtime: nvidia
ports:
- "${FISH_CPP_BIND:-0.0.0.0}:${FISH_CPP_PORT}:8000"
environment:
- NVIDIA_VISIBLE_DEVICES=${FISH_CPP_GPU_DEVICES:-1}
- FISH_CPP_MODEL=${FISH_CPP_MODEL:-s2-pro-q6_k.gguf}
- FISH_CPP_TOKENIZER=tokenizer.json
- FISH_CPP_DEVICE=0
volumes:
- ${FISH_CPP_WEIGHTS_DIR}:/weights:ro
- ${FISH_CPP_REFERENCE_DIR}:/references:ro
healthcheck:
test: ["CMD-SHELL", "python3 -c \"import urllib.request,sys; sys.exit(0 if urllib.request.urlopen('http://127.0.0.1:8000/v1/health', timeout=5).status==200 else 1)\""]
interval: 30s
timeout: 10s
retries: 3
start_period: 60s
labels:
- homepage.group=AI Systems
- homepage.name=Fish (s2.cpp)
- homepage.icon=mdi-fish
- homepage.description=Fish s2-pro via s2.cpp/GGML — quantized for realtime (irv-ml1)
- homepage.href=http://10.100.79.3:${FISH_CPP_PORT}