Two issues from the first deploy attempt:
1) Build failure (real): linker errors on s2.cpp's CUDA build —
undefined references to cuMemSetAccess, cuDeviceGet, etc. These
are CUDA Driver API symbols (in libcuda.so), not Runtime API
(libcudart.so). The driver lib is provided by NVIDIA's container
runtime at RUN time, not BUILD time.
Fix: nvidia/cuda:devel images ship a stubs library at
/usr/local/cuda/lib64/stubs/libcuda.so that provides the symbols
for linking but is non-runnable. Adding that path via
LIBRARY_PATH + CMAKE_LIBRARY_PATH lets the linker resolve while
leaving runtime unchanged (real libcuda.so comes from the
driver mount).
2) Verify false positive: the /v1/tts verify step's last command was
`rm -f "$out"` — which always exits 0. This made the shell's
final exit code 0 regardless of whether curl/file/grep succeeded,
so verify reported OK even when nothing was running on host_port.
Fix: `set -e` at top + trap-based cleanup. Failures now propagate;
the rm still runs on either path via EXIT trap.
New stack scaffolding for the Fish quantized-realtime experiment. Not
deployed yet — this commit lands the canonical files; deploy follows.
Architecture decisions made in Phase 1:
* CUDA backend, NOT Vulkan. s2.cpp's CMakeLists exposes both
-DS2_VULKAN and -DS2_CUDA; the most recent upstream commit
(2026-04-12) was specifically about CUDA improvements, and CUDA
on the A6000 will be substantially faster than Vulkan for ML
matmul. -DS2_CUDA=ON in the Dockerfile build args.
* Pinned to s2.cpp commit e48ce8e02d8335bd9a0ba94679f605724b31d12
(2026-04-12 HEAD of main). Repo is alpha software per README;
pin tightly so future churn doesn't break our build. Bump
deliberately when wanting upstream improvements.
* Multi-stage Dockerfile: nvidia/cuda:12.6.0-devel for build (needs
CMake + ninja + git + the CUDA toolchain) → nvidia/cuda:12.6.0-runtime
for serve (slimmer; just the s2 binary + GGML libs + a small Python
shim). Cuts image size by ~50% vs single-stage devel.
* FastAPI shim (server.py) wraps s2.cpp CLI in Fish's `/v1/tts`
contract so the same bench harness + clients work against fish-cpp
with no changes. Per-request flow: decode optional reference WAV
from base64 → write to temp → subprocess.run the s2 binary → stream
resulting WAV back. Adds ~50-100ms per-request fork+exec overhead;
negligible vs the multi-second generation cost.
* `streaming: true` accepted in request body but IGNORED — s2.cpp
writes a complete WAV before returning, so chunked output isn't
available. Unlike fish-s2 (HF wrapper) where streaming drops TTFB
to 26ms, fish-cpp's TTFB ≈ total wall time. Speed depends entirely
on raw generation throughput.
* q6_k as default quant — sweet spot per typical GGUF guidance:
near-bf16 quality at ~5GB. Other variants (q4_k_m, q5_k_m, q8_0,
f16) selectable via FISH_CPP_MODEL env.
* Pinned to GPU 1 (A6000) by default to share with fish-s2 for
direct A/B benching. q6_k weights ~5GB + runtime ~3GB ≈ 8GB —
comfortable on either GPU.
* Port 8199 (next free in the irv-ml1 TTS slate).
Phase 2 (next) is the actual deploy + first build. Reserved 30-45 min
for cold-cache build + weights pull.