fish-cpp: delete the stack — s2.cpp is too alpha to use today

Three deploy iterations + four backend attempts (subprocess CUDA,
resident-server CUDA, Vulkan rebuild) all failed to deliver speedup
over fish-s2:

* CUDA path: ggml_cuda_init succeeded, weights loaded onto GPU per
  s2's logs, but nvidia-smi showed 0% utilization during synthesis.
  Wall time 20s/long phrase vs fish-s2's 7.5s. The "CUDA get_rows
  unsupported for type q6_K" warning hints at incomplete op coverage
  in s2.cpp's alpha CUDA backend for fish-speech architecture.

* Vulkan path: vk::IncompatibleDriverError on container init. NVIDIA
  Vulkan ICD not accessible inside the container despite
  NVIDIA_DRIVER_CAPABILITIES=compute,utility,graphics. Would need
  host-side nvidia-utils-vulkan installation or manual ICD bind
  mount. Didn't pursue.

Both are fixable — CUDA needs op coverage upstream (author actively
working on it; "selective embedding dequant" commit landed 16 days
ago), Vulkan needs host-side ICD setup. Neither is a config-flip,
both are real work for marginal-or-zero return. Better to delete the
stack and revisit when s2.cpp matures or when we tackle FP8
quantization on ana-ml2's RTX 6000 Ada (sm_89, native FP8 hardware).

Local image rmi'd, /opt/docker/compose/fish-cpp removed on irv-ml1.
/worktank/fish-cpp left for user-side sudo cleanup.

Future Fish acceleration paths (in order of decreasing certainty):
1. Wait for s2.cpp CUDA op coverage to mature (track upstream commits).
2. Quantize Fish BF16 → FP8 via TransformerEngine, deploy on
   ana-ml2's RTX 6000 Ada (Ada has native FP8 tensor cores, A6000
   doesn't). ~2x speedup if it works.
3. vLLM port of Fish (no upstream support today).
This commit is contained in:
vh
2026-04-28 01:52:57 -07:00
parent 67813bbef4
commit 0ba41e02ea
7 changed files with 0 additions and 624 deletions
-142
View File
@@ -1,142 +0,0 @@
# Deploy fish-cpp (Fish s2-pro via s2.cpp + GGML CUDA inference) to irv-ml1.
#
# Builds the image locally — multi-stage CUDA devel base (CMake + s2.cpp
# compile, ~10 min cold) → CUDA runtime base + binary + python shim.
# Pre-pulls rodrigomt/s2-pro-gguf weights (q6_k default, ~5 GB) into
# the bind-mounted weights dir.
#
# Usage:
# scripts/elway irv-ml1 --playbook playbooks/deploy-fish-cpp.yaml
#
# Idempotent — every step is creates-/when-gated; rerun is safe.
vars:
compose_dir: /opt/docker/compose/fish-cpp
references_dir: /worktank/fish-cpp/references
weights_dir: /worktank/fish-cpp/weights
host_port: "8199"
weights_repo: rodrigomt/s2-pro-gguf
default_quant: s2-pro-q6_k.gguf
steps:
# ── host-side dirs ──────────────────────────────────────────────────
- name: Ensure /worktank/fish-cpp root exists (one-time, sudo)
shell: mkdir -p /worktank/fish-cpp
sudo: true
creates: /worktank/fish-cpp
- name: Chown /worktank/fish-cpp to lkraven
shell: chown -R lkraven:lkraven /worktank/fish-cpp
sudo: true
when: "[ \"$(stat -c %U /worktank/fish-cpp)\" != \"lkraven\" ]"
- name: Ensure references dir exists
shell: mkdir -p {{ references_dir }}
creates: "{{ references_dir }}"
- name: Ensure weights dir exists
shell: mkdir -p {{ weights_dir }}
creates: "{{ weights_dir }}"
- name: Ensure compose dir exists
shell: mkdir -p {{ compose_dir }}
creates: "{{ compose_dir }}"
# ── deploy build context ────────────────────────────────────────────
# s2.cpp is built INSIDE the docker image, but the Dockerfile + shim
# need to be present in the compose dir so `docker compose build`
# can find them.
- name: Upload compose.yaml
upload:
src: stacks/fish-cpp/compose.yaml
dest: "{{ compose_dir }}/compose.yaml"
mode: "0644"
- name: Upload Dockerfile
upload:
src: stacks/fish-cpp/Dockerfile
dest: "{{ compose_dir }}/Dockerfile"
mode: "0644"
- name: Upload server.py (FastAPI shim)
upload:
src: stacks/fish-cpp/server.py
dest: "{{ compose_dir }}/server.py"
mode: "0644"
- name: Upload entrypoint.sh (starts s2 server + uvicorn shim)
upload:
src: stacks/fish-cpp/entrypoint.sh
dest: "{{ compose_dir }}/entrypoint.sh"
mode: "0755"
- name: Seed .env from template (only if absent)
upload:
src: stacks/fish-cpp/.env.example
dest: "{{ compose_dir }}/.env"
mode: "0644"
when: "[ ! -f {{ compose_dir }}/.env ]"
# ── pre-pull weights ────────────────────────────────────────────────
# q6_k + tokenizer.json (~5 GB total). Same one-shot
# python:3.12-slim + huggingface_hub.snapshot_download + hf_transfer
# pattern we've used for fish-s2, voxtral, etc. Idempotent on rerun
# via `creates:` on the model file.
- name: Pre-pull rodrigomt/s2-pro-gguf weights (q6_k + tokenizer, ~5 GB)
shell: |
docker run --rm --user 1000:1000 \
-e HOME=/tmp/h -e HF_HUB_ENABLE_HF_TRANSFER=1 \
-v {{ weights_dir }}:/dest \
python:3.12-slim sh -c 'set -e; mkdir -p /tmp/h /tmp/pip /tmp/site; PIP_CACHE_DIR=/tmp/pip pip install --quiet --target /tmp/site huggingface_hub hf_transfer; PYTHONPATH=/tmp/site python -c "from huggingface_hub import snapshot_download; snapshot_download(repo_id=\"{{ weights_repo }}\", local_dir=\"/dest\", allow_patterns=[\"{{ default_quant }}\",\"tokenizer.json\"])"'
creates: "{{ weights_dir }}/{{ default_quant }}"
# ── build + bring up ────────────────────────────────────────────────
- name: docker compose build (~10 min first time; CUDA toolchain + s2.cpp compile)
shell: |
set -o pipefail
cd {{ compose_dir }} && docker compose build 2>&1 \
| grep -vE '^#[0-9]+ |^ => |^=> |Collecting|Downloading|Requirement|Using cached|Installing collected|Successfully (installed|built)|━'
- name: docker compose up -d
shell: cd {{ compose_dir }} && docker compose up -d
- name: Wait for /v1/health to respond
shell: |
for i in $(seq 1 60); do
curl -sf -o /dev/null --max-time 3 http://localhost:{{ host_port }}/v1/health && exit 0
sleep 5
done
exit 1
changed_when: "false"
verify:
- name: /v1/health returns 200 + reports model loaded
shell: |
curl -sf http://localhost:{{ host_port }}/v1/health \
| python3 -c "import json,sys; d=json.load(sys.stdin); assert d.get('status')=='ok' and d.get('model')"
changed_when: "false"
- name: /v1/tts returns a real WAV (POST with text body)
# `set -e` so curl/file/grep failures actually propagate. The
# previous version put `rm -f` as the last command, which always
# exits 0 — masking real failures (verify reported OK even when
# nothing was running on host_port). Trap-based cleanup runs the
# rm even on failure.
shell: |
set -e
out=$(mktemp --suffix=.wav)
trap 'rm -f "$out"' EXIT
curl -sf -X POST http://localhost:{{ host_port }}/v1/tts \
-H 'Content-Type: application/json' \
-d '{"text":"Verify."}' \
-o "$out" --max-time 60
file -b "$out" | grep -q '^RIFF.*WAVE'
changed_when: "false"
- name: Container is running
shell: docker inspect fish-cpp --format '{{.State.Status}}' | grep -q running
changed_when: "false"