Two new dense headline rails above the existing reddit/tech cards.
Designed for high-volume "what happened" coverage where the title
is the deliverable — no LLM summarization, ~15 items per section,
6-column-collapsing grid (title / source / time).
Digest pipeline:
* fetch_miniflux_headlines(category) — flat list per category, dedup
by lowercased title (different feeds syndicate the same wire stories)
* 8h look-back window (vs 12h for tech/reddit) since headlines move
faster
* cap of 15 per section (DIGEST_MINIFLUX_HEADLINES_MAX)
Frontend:
* .headline element parallels .item for the hide-button machinery
(both have data-id, both honored by app.js)
* dense 3-col layout collapses to 1-col on narrow screens
* jumpnav now numbers world=01, local=02, reddit=03, tech=04
Setup:
* seed-headlines.py — one-shot script (lives in the image at
/app/seed-headlines.py). Creates the World + Local categories in
miniflux, subscribes a curated feed list, and renames each feed
to a short display title (BBC vs "BBC News", "LA Times" vs "California").
Idempotent — reruns only add new feeds.
* Default world: BBC, NPR, Al Jazeera. Default local: LA Times Local,
LA Times CA, Voice of OC. (OC Register blocks miniflux; left out.)
* entrypoint.sh now syncs templates/{style.css,app.js,favicon.svg}
to /output on container start so frontend asset updates land
without a manual copy after rebuild.
Three upstream gaps surfaced once /generate was actually exercised:
1. infer-api.py builds an 18-arg positional tuple but the pipeline
expects 24 — first missing arg is `format`, so audio_duration
shifts into format's slot and the pipeline calls len() on an
int. Ship a patched copy of infer-api.py and COPY over upstream's
in the Dockerfile. Also handle empty lora_name_or_path -> "none"
(empty string trips HF Hub's repo-id validator).
2. torchcodec + ffmpeg are required by the WAV save path but neither
is in upstream requirements.txt. Without them every /generate
runs to completion and then 500s at write-time.
3. ACE-Step caches checkpoints at /root/.cache/ace-step/checkpoints
(HARDCODED, not honored by HF_HOME). Mount our persistent dir
there so the ~7 GB model survives container recreates.
Bench on A6000 (cached model, lo-fi hip hop, 60-step euler/apg):
10s @ 27 steps -> 9.4s (0.94x)
30s @ 60 steps -> 11.2s (0.37x, ~2.7x realtime)
60s @ 60 steps -> 14.8s (0.24x, ~4x realtime)
Two new audio-generation stacks alongside the TTS slate:
ace-step :8210 — Apache 2.0 music generation foundation model
(hybrid diffusion + LLM). Lyric-aware multi-minute songs. ~10-12 GB
VRAM during inference, A6000-pinned. Custom Dockerfile patches
upstream's torch/cu126 resolution bug (--extra-index-url cu126 was
falling back to pypi-default cu13 wheels, mismatching torchvision).
stable-audio-open :8211 — Stability AI 1.21B latent-diffusion SFX +
ambience. Up to 47s clips at 44.1 kHz. ~6 GB VRAM in fp16,
A6000-pinned. Custom FastAPI shim around diffusers' StableAudioPipeline
(no upstream HTTP server). Dockerfile pins torchsde explicitly —
diffusers doesn't pull it as a hard dep but
CosineDPMSolverMultistepScheduler needs it.
Three deploy iterations + four backend attempts (subprocess CUDA,
resident-server CUDA, Vulkan rebuild) all failed to deliver speedup
over fish-s2:
* CUDA path: ggml_cuda_init succeeded, weights loaded onto GPU per
s2's logs, but nvidia-smi showed 0% utilization during synthesis.
Wall time 20s/long phrase vs fish-s2's 7.5s. The "CUDA get_rows
unsupported for type q6_K" warning hints at incomplete op coverage
in s2.cpp's alpha CUDA backend for fish-speech architecture.
* Vulkan path: vk::IncompatibleDriverError on container init. NVIDIA
Vulkan ICD not accessible inside the container despite
NVIDIA_DRIVER_CAPABILITIES=compute,utility,graphics. Would need
host-side nvidia-utils-vulkan installation or manual ICD bind
mount. Didn't pursue.
Both are fixable — CUDA needs op coverage upstream (author actively
working on it; "selective embedding dequant" commit landed 16 days
ago), Vulkan needs host-side ICD setup. Neither is a config-flip,
both are real work for marginal-or-zero return. Better to delete the
stack and revisit when s2.cpp matures or when we tackle FP8
quantization on ana-ml2's RTX 6000 Ada (sm_89, native FP8 hardware).
Local image rmi'd, /opt/docker/compose/fish-cpp removed on irv-ml1.
/worktank/fish-cpp left for user-side sudo cleanup.
Future Fish acceleration paths (in order of decreasing certainty):
1. Wait for s2.cpp CUDA op coverage to mature (track upstream commits).
2. Quantize Fish BF16 → FP8 via TransformerEngine, deploy on
ana-ml2's RTX 6000 Ada (Ada has native FP8 tensor cores, A6000
doesn't). ~2x speedup if it works.
3. vLLM port of Fish (no upstream support today).
CUDA backend confirmed broken for fish-speech ops on s2.cpp v0.x — alpha,
incomplete op coverage, GPU stays at 0% during generation despite
ggml_cuda_init succeeding. Vulkan was the original README example
(`-v 0`), so likely the more battle-tested path.
Build the image with BOTH backends so we can flip via env without
rebuilding:
* libvulkan-dev + glslc in the build stage (GGML's Vulkan backend
compiles its shaders with glslc at build time; without it the
cmake configure silently disables Vulkan).
* libvulkan1 + the libggml-vulkan.so copy in the runtime stage.
* compose env NVIDIA_DRIVER_CAPABILITIES=compute,utility,graphics —
default nvidia-container-toolkit only mounts compute libs; Vulkan
needs the graphics ICD (libGLX_nvidia + nvidia_icd.json) too.
* entrypoint reads FISH_CPP_BACKEND (cuda/vulkan/cpu) and selects
the appropriate -c/-v/no-flag invocation.
* Default backend = vulkan.
Subprocess-per-request architecture forced CUDA + model load on every
/v1/tts call (~10-20s init, then 5-15s generation). Even though CUDA
is now actually being used (`-c 0` fix landed), 32s for "Verify."
proved per-request init was the bottleneck.
s2.cpp ships a built-in HTTP server (`--server -H -P`) that keeps the
model resident on the GPU. Refactor:
* entrypoint.sh — backgrounds `s2 --server -P 3030 -c 0 -m ... -t ...`,
waits for it to bind 3030, then foregrounds uvicorn. tini supervises
via `wait -n` so either child dying takes down the container.
* server.py — drops subprocess.run; instead httpx-POSTs Fish-shaped
/v1/tts JSON to s2's localhost:3030/generate (multipart form: text
+ optional prompt_text/prompt_audio for cloning). Model load + CUDA
init now happen once at container start, not per-request.
* Dockerfile — added httpx (shim dep), curl (entrypoint readiness
probe), and the entrypoint.sh COPY+chmod. CMD now invokes
entrypoint.sh instead of uvicorn directly.
* deploy-fish-cpp.yaml — uploads entrypoint.sh alongside server.py.
s2.cpp's README example uses `-v 0` which is `--vulkan 0` (Vulkan
device 0), easy to misread as "voice 0". The shim copied that
verbatim, so even after fixing the libcuda.so build problem AND the
libgomp.so runtime dep, every synthesis ran on CPU because the wrong
backend was selected.
Direct verification: `[Model] NPU not compiled, falling back to CPU`
in stderr; nvidia-smi showed no s2 process; bench timed out at 60s
on phrases that fish-s2 (HF, GPU) does in 7s.
s2.cpp's CLI:
-v <id> = --vulkan <device>
-c <id> = --cuda <device>
-M = --metal (Apple Silicon)
Switched the shim to `-c 0`. The CUDA backend IS in the build (-DS2_CUDA=ON
worked, libggml-cuda.so links fine per ldd, libcuda.so.1 mounts at
runtime via NVIDIA container runtime) — just wasn't being told to use it.
Build succeeded after the libcuda.so symlink fix, but the first
/v1/tts request returned HTTP 500 with:
s2 binary failed (rc=127): /usr/local/bin/s2: error while loading
shared libraries: libgomp.so.1: cannot open shared object file
CMake auto-enabled OpenMP during the build (gcc's -fopenmp flag), so
the s2 binary dynamically links libgomp.so.1. The build-stage devel
image had it; the slim cuda:runtime base doesn't ship it by default.
Adding libgomp1 to the runtime image's apt install resolves it.
Second attempt's CMAKE_LIBRARY_PATH + LIBRARY_PATH didn't get picked
up by ggml's nested CMake — same linker errors as the first run.
Robust fix: symlink the stub at /usr/local/cuda/lib64/stubs/libcuda.so
into /usr/local/lib (which ld searches unconditionally) and provide
both libcuda.so AND libcuda.so.1 (the SONAME ggml-cuda's
libggml-cuda.so links against). ldconfig refreshes the cache.
The symlinks live only in the build stage. The runtime image inherits
the real driver-provided libcuda.so.1 via NVIDIA's container runtime
mount, so the stubs never get used at execution time.
Two issues from the first deploy attempt:
1) Build failure (real): linker errors on s2.cpp's CUDA build —
undefined references to cuMemSetAccess, cuDeviceGet, etc. These
are CUDA Driver API symbols (in libcuda.so), not Runtime API
(libcudart.so). The driver lib is provided by NVIDIA's container
runtime at RUN time, not BUILD time.
Fix: nvidia/cuda:devel images ship a stubs library at
/usr/local/cuda/lib64/stubs/libcuda.so that provides the symbols
for linking but is non-runnable. Adding that path via
LIBRARY_PATH + CMAKE_LIBRARY_PATH lets the linker resolve while
leaving runtime unchanged (real libcuda.so comes from the
driver mount).
2) Verify false positive: the /v1/tts verify step's last command was
`rm -f "$out"` — which always exits 0. This made the shell's
final exit code 0 regardless of whether curl/file/grep succeeded,
so verify reported OK even when nothing was running on host_port.
Fix: `set -e` at top + trap-based cleanup. Failures now propagate;
the rm still runs on either path via EXIT trap.
New stack scaffolding for the Fish quantized-realtime experiment. Not
deployed yet — this commit lands the canonical files; deploy follows.
Architecture decisions made in Phase 1:
* CUDA backend, NOT Vulkan. s2.cpp's CMakeLists exposes both
-DS2_VULKAN and -DS2_CUDA; the most recent upstream commit
(2026-04-12) was specifically about CUDA improvements, and CUDA
on the A6000 will be substantially faster than Vulkan for ML
matmul. -DS2_CUDA=ON in the Dockerfile build args.
* Pinned to s2.cpp commit e48ce8e02d8335bd9a0ba94679f605724b31d12
(2026-04-12 HEAD of main). Repo is alpha software per README;
pin tightly so future churn doesn't break our build. Bump
deliberately when wanting upstream improvements.
* Multi-stage Dockerfile: nvidia/cuda:12.6.0-devel for build (needs
CMake + ninja + git + the CUDA toolchain) → nvidia/cuda:12.6.0-runtime
for serve (slimmer; just the s2 binary + GGML libs + a small Python
shim). Cuts image size by ~50% vs single-stage devel.
* FastAPI shim (server.py) wraps s2.cpp CLI in Fish's `/v1/tts`
contract so the same bench harness + clients work against fish-cpp
with no changes. Per-request flow: decode optional reference WAV
from base64 → write to temp → subprocess.run the s2 binary → stream
resulting WAV back. Adds ~50-100ms per-request fork+exec overhead;
negligible vs the multi-second generation cost.
* `streaming: true` accepted in request body but IGNORED — s2.cpp
writes a complete WAV before returning, so chunked output isn't
available. Unlike fish-s2 (HF wrapper) where streaming drops TTFB
to 26ms, fish-cpp's TTFB ≈ total wall time. Speed depends entirely
on raw generation throughput.
* q6_k as default quant — sweet spot per typical GGUF guidance:
near-bf16 quality at ~5GB. Other variants (q4_k_m, q5_k_m, q8_0,
f16) selectable via FISH_CPP_MODEL env.
* Pinned to GPU 1 (A6000) by default to share with fish-s2 for
direct A/B benching. q6_k weights ~5GB + runtime ~3GB ≈ 8GB —
comfortable on either GPU.
* Port 8199 (next free in the irv-ml1 TTS slate).
Phase 2 (next) is the actual deploy + first build. Reserved 30-45 min
for cold-cache build + weights pull.
Voxtral final fix (8th iteration):
* The bundled voxtral_tts.yaml hardcodes gpu_memory_utilization: 0.8
on the language_model stage — overrides the CLI flag. Mounted a
patched copy (0.4) at /etc/voxtral/voxtral_tts.yaml and pointed
--stage-configs-path there.
* With Kyutai stopped to free 5 GB on the 3090, both stages fit
(target 9.4 + 2.4 GB ≈ 11.8 GB; 17 GB free post-kyutai-stop).
* Voxtral now healthy on GPU 0 — bench: 1.9-2.7 s TTFB, real WAV.
Fish s2-pro optimization (per-request sweep, no model swap):
* `streaming: true` in request body drops TTFB from 7.7 s → 0.026 s
(300×). Total time goes up ~1 s (chunked HTTP overhead) but
perceived latency = TTFB. Use stream:true for any interactive use.
* `latency: "balanced"` actually slower than default — bad name; skip.
* `use_memory_cache: "on"` no measurable benefit.
* `chunk_length: 100` (default 200) no TTFB benefit non-streaming.
* Server-side `--half` (fp16 inference) added via compose `command`
override — passes through start_server.sh's $@ unchanged into
api_server.py. Should reduce total time too. Validation pending
the post-restart bench.
Kyutai stopped to free GPU 0 budget — the bench numbers earlier
(3.4 s avg) were unimpressive vs Voxtral's 2.3 s in the same
multilingual slot. Kept the stack files for future re-deploy if
needed; just the running container is gone.
Fourth attempt finally found the right invocation. Voxtral is a
two-stage TTS pipeline (language_model → acoustic_transformer →
audio output), not a flat MistralForCausalLM. Standard `vllm serve`
errored with "no module named 'acoustic_transformer'" because it
loads the model as a vanilla Mistral causal LM.
Pattern from /workspace/vllm-omni/examples/online_serving/
qwen3_tts/run_server.sh (closest in-image analog):
vllm-omni serve <MODEL> \
--stage-configs-path vllm_omni/model_executor/stage_configs/voxtral_tts.yaml \
--host 0.0.0.0 --port 8000 \
--gpu-memory-utilization 0.45 \
--trust-remote-code --omni
Key differences from previous attempt:
* `vllm-omni` binary, not `vllm`
* `--omni` flag activates multi-stage pipeline
* `--stage-configs-path` points at the bundled YAML that maps
stages to GPU + scheduler + worker classes
* Dropped --load-format/--tokenizer-mode/--config-format=mistral
flags — the stage config handles tokenizer_mode internally
* --trust-remote-code is required for the acoustic_transformer
custom code path
Default .env.example now: GPU 0 (3090) with util 0.45 (~10.6 GB
target on 24 GB GPU). The A6000 is fully booked by Fish s2-pro.
Third voxtral attempt: image pulled clean (3 min, v0.18.0), entrypoint
parsed correctly, vLLM started, but engine init failed two ways:
1. HF rate-limited the irv-ml1 IP (38.120.94.3) during the metadata
fetch — 429 Too Many Requests from too many large unauthenticated
pulls today (heretic, 27b, fish-s2, fish-s1-mini, voxtral). Added
HF_TOKEN env passthrough; user generates a token at
https://huggingface.co/settings/tokens and sets VOXTRAL_HF_TOKEN
in .env.
2. Voxtral uses Mistral's native model format (params.json +
tekken.json tokenizer + consolidated.safetensors single file),
NOT HF transformers format (config.json + tokenizer.json + sharded
.safetensors). vLLM errored with "ensure presence of params.json
for Mistral models." Fix: pass --load-format=mistral
--tokenizer-mode=mistral --config-format=mistral to vllm serve.
Confirmed by inspecting the Voxtral-4B-TTS-2603 HF tree:
25 files, ships params.json + tekken.json + consolidated.safetensors.
Both fixes baked into compose. User needs to drop their HF_TOKEN into
.env once and recreate.
Side note discovered while debugging: fish-s2 s1-mini variant uses
the tiktoken tokenizer format; the wrapper can't load it (errors with
"NoneType has no attribute encode" on warmup). So s1-mini isn't a
drop-in optimization for s2-pro — different code path needed. Fish
back on s2-pro for now.
Second voxtral attempt got past the image pull (v0.18.0 published,
~3 min download) but container init failed:
unable to start container process: error during container init:
exec: "--model=mistralai/Voxtral-4B-TTS-2603": stat ...: no such file
vllm/vllm-omni:v0.18.0 has Entrypoint=null AND Cmd=null — there's no
default executable. The compose's `command:` array becomes the full
exec invocation, with --model=... interpreted as the binary name.
Standard vLLM serving CLI is `vllm serve <model> [flags]`. The
binary's at /usr/local/bin/vllm. Set entrypoint: ["vllm", "serve"]
and pass the model as a positional arg.
While we're here: HF cache was empty too (Voxtral 4B BF16 ~8 GB
download on first start) — vLLM auto-downloads from HF on model
load, so no separate pre-pull step needed.
Three fixes from the second-wave deploy attempts:
* voxtral: vllm/vllm-omni doesn't publish a `latest` tag — pull
failed with "manifest unknown". Pinned VOXTRAL_VLLM_TAG to v0.18.0
(released 2026-03-29, the day after the Voxtral 4B TTS release —
first cut with Voxtral support).
* kyutai-tts: NillPointer wrapper exposes ONLY /health (root) and
POST /v1/audio/speech. No /v1/models, no /v1/audio/voices —
those return 404. Verified by /openapi.json against the live
container. Compose healthcheck + playbook wait + verify steps
all repointed at the actual paths. POST /v1/audio/speech is now
smoke-tested with a RIFF WAV assertion (same pattern as fish-s2).
* fish-s2: added FISH_S2_MODEL env var so the model variant is
swappable via .env without rebuilding. Both s2-pro (default) and
s1-mini are pre-pulled into the bind-mount; LLAMA_CHECKPOINT_PATH
+ DECODER_CHECKPOINT_PATH now use ${FISH_S2_MODEL:-s2-pro}.
s1-mini was originally gated on fishaudio's HF org (401), but
niobures/OpenAudio-S1 mirrors the same files openly — pulled
from there via a one-shot snapshot_download.
After getting fish-s2 finally healthy on attempt #5, the playbook's
verify still failed because /v1/audio/voices doesn't exist. Discovery:
the Fish wrapper has a custom API surface, not OpenAI-compatible.
Real endpoints:
POST /v1/tts — synthesis (text body, optional `references`
field for voice cloning, returns audio/wav)
GET /v1/health — liveness (used by Docker healthcheck)
GET /heartbeat — alternate liveness signal
GET / — Swagger Editor UI for the OpenAPI spec
No /v1/audio/speech, /v1/audio/voices, /v1/models — those return 404.
Updated:
* Playbook verify — replaced the JSON-shape /v1/audio/voices check
with a POST /v1/tts smoke that asserts a real RIFF WAV comes back.
* README API section — replaced the OpenAI-compat examples with
Fish's actual {"text":"...","references":[...]} body shape.
* README disk footprint — corrected ~9 GB → ~11 GB (codec.pth was
larger than I estimated; 1.9 GB + 9 GB safetensors).
* README Lessons learned section — recorded the 5-iteration deploy
story so the next time we touch a Fish-style upstream we don't
re-walk the dockerfile / target / pre-pull / API-shape traps.
Fourth fish-s2 attempt got past build + checkpoints, then container
crashlooped silently again. Diagnosis: the upstream docker/Dockerfile
is multi-stage with `webui` and `server` targets; without specifying
a target, docker builds the LAST stage (webui — gradio-only, no
start_server.sh, no API server). start_server.sh is the entrypoint
script that lives only in the `server` stage.
Confirmed by `cat /app/start_server.sh` inside the built image:
"No such file or directory."
Upstream's compose.yml uses target: server on its server service —
doing the same here.
Third deploy attempt got past the build but crashlooped at container
start: Fish's start_server.sh validates checkpoints/s2-pro/ exists
and exits cleanly (rc=0) if missing — no auto-download, no helpful
message. /worktank/fish-s2/checkpoints/ was empty, so the container
exited every ~52s under restart policy.
Added an idempotent pre-pull step using the same one-shot
python:3.12-slim + huggingface_hub.snapshot_download + hf_transfer
pattern we used for the Qwen 3.6 GGUFs earlier today. Pulls the 9
relevant files (~11 GB total: codec.pth + 2 safetensors shards +
config + tokenizer/template) directly into the bind-mount at
/worktank/fish-s2/checkpoints/s2-pro/ — gated by `creates:` on
codec.pth so the pre-pull step is a no-op on reruns.
~83 s wall-clock for the 11 GB pull on first deploy.
Wrapped .masthead-brand in <a href="index.html"> in both digest.html.j2
and archive.html.j2 so the hero is a clickable shortcut to the latest
edition. Useful when reading an archived edition and you want to jump
back to the freshest one without going through the archive list.
CSS: color: inherit + text-decoration: none keeps the visual
identical; hover drops opacity to 0.85 for affordance; focus-visible
gets an accent outline so keyboard nav is discoverable.
Second deploy attempt failed at build time:
failed to fetch anonymous token: ... ghcr.io/fishaudio/fish-speech ... 403 Forbidden
Root cause: dockerfile.dev is a thin two-line wrapper around
`FROM ghcr.io/fishaudio/fish-speech:${VERSION}`, which is a private
GHCR base image. Anonymous pulls 403, and we'd need GHCR auth to use
that path. The dev variant is meant for upstream's CI / fish-speech
contributors, not external consumers.
The REAL production path (from upstream's compose.base.yml) is to
build from `docker/Dockerfile` with build args BACKEND=cuda,
CUDA_VER=12.9.0, UV_EXTRA=cu129, UV_VERSION=0.8.15. That builds
everything from source — slower (15-20 min cold), but fully self-
contained.
irv-ml1's driver (595.58.03, CUDA 13.2 capable) is forward-compatible
with the 12.9 PyTorch wheels.
Took three iterations to find the right Dockerfile because:
1. First try: dockerfile (lowercase) — doesn't exist
2. Second try: dockerfile.dev — exists but pulls a private base
3. Third try: docker/Dockerfile — actual production path
First fish-s2 deploy attempt failed in step 9/11:
failed to read dockerfile: open dockerfile: no such file or directory
Upstream fishaudio/fish-speech ships:
* dockerfile.dev (lowercase, dev/test image)
* compose.yml + compose.base.yml (intended deploy path:
`docker compose --profile server up`)
There is no standalone production Dockerfile. The dockerfile.dev
image is what their own compose.yml builds from anyway, so building
against it directly is functionally equivalent to using their compose
profile — we just keep our own restart-policy / labels / bind-mount
conventions on the outer compose.
Comment in the build block now documents this so future-Claude doesn't
re-walk the path.
Adds the three premier 2026 TTS releases we missed during the original
fleet build-out (early April), all licensed for self-host:
* Fish Audio S2-Pro (port 8195, GPU 1 / A6000) — released 2026-03-09.
4B dual-AR (Slow + Fast) trained on 10M+ hours / 80+ languages.
Headline: 15,000+ paralinguistic / emotion tags via natural language
([laugh] [whispers] [super happy] etc.) — a step-function over
Chatterbox Turbo's 9 fixed tags. 91.61% paralinguistic win rate on
EmergentTTS-Eval. ~150 ms streaming TTFB, voice cloning, MIT-style
open. ~17 GB VRAM.
* Voxtral TTS (port 8197, GPU 1 / A6000) — Mistral, released 2026-03-28.
4B open-weight, 70 ms model latency, 9.7× realtime. 68.4% blind A/B
win rate vs ElevenLabs Flash v2.5 in cloning. 8 languages
(EN/FR/DE/ES/IT/PT/NL/HI). Served via vLLM-Omni (Mistral's partner
serving stack) — published Docker image, no local build. ~16 GB VRAM.
CC BY-NC license — personal/research use only; flagged in README.
* Kyutai TTS (port 8198, GPU 0 / 3090) — kyutai/tts-1.6b-en_fr.
Trained on 2.5M hours from the Moshi/Mimi team. Claimed 220 ms in
solo setup, 32 simultaneous streams under 350 ms on L40. Kyutai's
official deploy is Rust + websockets only; using NillPointer's
community OpenAI-compat wrapper to bridge to /v1/audio/speech so
it slots into the same bench harness. ~4-6 GB VRAM.
Each stack: compose.yaml (build context, env, volumes, healthcheck,
homepage label), .env.example (all tunables documented), README.md
(why it exists, headline numbers, API, deploy + hardware notes).
Playbooks at playbooks/deploy-{fish-s2,voxtral,kyutai-tts}.yaml are
idempotent in the same shape as the existing deploy-vibevoice /
deploy-chatterbox playbooks.
Port allocations on irv-ml1 after this lands: 8188 ComfyUI, 8190
CosyVoice, 8191 Qwen3-TTS, 8192 IndexTTS-2, 8193 Kokoro, 8194
VibeVoice, 8195 Fish, 8196 Chatterbox, 8197 Voxtral, 8198 Kyutai,
8765 Parakeet ASR.
Investigation of the slow (8-12s) qwen3-tts TTFB found the upstream
wrapper has 5 backend options. The advertised path to fast TTFB is
TTS_BACKEND=optimized (torch.compile + CUDA graphs + real-time
streaming). It loads cleanly but crashes the container during its
hardcoded warmup phase — silent exit (ExitCode 0, no traceback,
no OOM kill), repeats every ~22s under restart policy.
TTS_WARMUP_ON_START=false suppresses the factory-level warmup but
the optimized backend has its own internal warmup that fires
regardless and triggers the crash.
Updated the .env.example block to enumerate all 5 backend options
with their actual current behavior so future-Claude doesn't re-walk
this path. official is staying as the default.
The wrapper's `optimized` backend (torch.compile + CUDA graphs +
real-time streaming) reads its model registry from a YAML config:
default path is ~/qwen3-tts/config.yaml inside the container, which
doesn't exist. Without TTS_CONFIG set, the backend boots with an
empty registry and every synthesis request fails with
"Unknown model key: '<name>'. Available: []".
The repo ships /app/config.yaml with all 4 model variants defined.
Pointing TTS_CONFIG at it lets the optimized backend load cleanly.
This is a prerequisite for benching the optimized backend properly
— it's the path to the upstream's claimed 97 ms streaming TTFB. The
default `official` backend uses naive HF transformers autoregressive
generation that pegged GPU at only 27% utilization and gave us 8-12 s
TTFB on bench (no recompile theory needed — same phrase repeated 4x
plateaued at 8.5 s, ruling out shape-specific recompilation).
qwen3-tts: deploy was using the -Base checkpoint, which sounds like
the right one ("supports voice cloning") but the upstream wrapper's
only synthesis path goes through generate_custom_voice. The -Base
variant doesn't expose that, so every request — including ones with
the wrapper's listed built-in voices like Ryan/Vivian — errored with
"does not support generate_custom_voice". The -CustomVoice variant
exposes both the cloning machinery and the preset voices, and is
what the wrapper actually needs.
The .env.example comments had the variant labels backward; fixed in
this commit. Live host already updated to -CustomVoice via direct
.env edit (model downloaded on container restart).
chatterbox README listed [whisper] and [breath] as supported tags —
those are in the base Chatterbox tag set but NOT in the Turbo set
that's actually loaded. Replaced with the canonical 9-tag list
verified against /api/model-info: laugh, chuckle, sigh, gasp, cough,
clear throat, sniff, groan, shush.
New runbook captures the three-phase process:
Phase 1 — Drop --append-only via DSM Container Manager web UI
Phase 2 — sudo resticprofile forget --prune --verbose on each of
nh3-docker, nh3-dev, irv-ml1 (interactive sudo per host)
Phase 3 — Restore --append-only via DSM
Why each phase looks the way it does, what to expect (largely no-op
runs for the first 6 months while no snapshots have aged out of the
keep window), how to verify each phase non-destructively (curl 401
on the rest-server root proves the container's up + serving), what
to do if Phase 2 fails with `repository is configured as append-only`
(skipped Phase 1 / DSM didn't apply), and the path to future
automation (find docker bin path on DSM, NOPASSWD-lock syncuser to
the specific recreate command).
Includes a "last run history" table seeded with today's first
post-pipeline run (no-op, irv-ml1 only had 3 snapshots due to the
04-25→27 CUDA stall).
Cross-referenced from docs/README.md (runbook tree), docs/
orientation.md (where-to-look table), and STATUS.md item 9 (which
now points at the runbook + records the next-round date 2026-07-27).
Both reported (unhealthy) in docker ps. Two distinct root causes:
* news-digest-web: switched from nginx:alpine to python:3.12-alpine
(uvicorn) but kept the wget healthcheck against `localhost`. Alpine's
/etc/hosts maps localhost to BOTH ::1 and 127.0.0.1; busybox wget
tries IPv6 first, hits "connection refused" because uvicorn binds
IPv4-only, and doesn't fall back. Pinned to 127.0.0.1.
* chatterbox: devnen's image is built from a python:3.10 base and
doesn't ship curl, so `curl -fsS http://localhost:8004/api/model-info`
failed with `/bin/sh: 1: curl: not found`. Replaced with a python
urllib one-liner that fetches + asserts `b'"loaded":true' in body`,
also pinned to 127.0.0.1 to dodge the same IPv4/IPv6 race.
Both YAML extractions tested directly inside the running containers
(via `sh < script`) — chatterbox python check returns 0 when the model
is loaded.
Big update for 2026-04-27. Sections added:
* Marked the "🟥 Blocked — irv-ml1 stalled" header as RECOVERED with
resolution notes (driver 595.58.03 / CUDA 13.2 IS working, both GPUs
detected; the original "stall" must have been a one-shot
post-install hiccup that resolved on a later boot).
* New "Session milestones — 2026-04-27" section covering:
- irv-ml1 unstall + 5 pre-existing GPU stacks restored
- Kokoro GPU variant deployed (irv-ml1:8193) with the .env.example
default flipped to gpu now that the driver works
- VibeVoice 1.5B deployed (irv-ml1:8194) after fixing two bugs:
full 40-char SHA required by buildx + verify regex didn't match
the OpenAI list-format response shape
- Chatterbox Turbo deployed (irv-ml1:8196) after fixing three:
upstream moved Dockerfile path (docker/Dockerfile.gpu →
Dockerfile.cu128 at root), pinned to current SHA instead of `main`,
/health doesn't exist (switched all probes to /api/model-info
which is the wrapper's own ready-after-loaded signal)
- llama-swap qwen3.6 ttl removal across non-pinned variants;
qwen3.6-35-a3b unpinned (was OOM'ing other loads via the pinned
group's persistent: true flag); granite-4-small added to the
pinned group to stop it swapping with qwen3.6-27b
- Backup verification: all three layers green (per-host restic,
PBS-ANA, PBS-NH3 mirror — 2026-04-27 snapshots everywhere). Noted
that backrest's empty dashboard is expected (no plans configured;
the actual orchestration is the per-host resticprofile timers).
Symptom: granite-4-small and qwen3.6-27b were evicting each other
when called in alternation. granite is the news-digest curator (fires
twice daily on cron) — being evicted means a cold reload (~5s) on
every digest tick, plus visible churn whenever the user uses 27b
concurrently.
Added granite-4-small to the `pinned` group as a persistent member.
~5-6 GB at Q4_K_M + 120K KV ≈ comfortable inside the existing pin
budget (qwen3.5-9b ~6 GB → ~12 GB total persistent). Single RTX 6000
Ada is 48 GB, leaves ~36 GB headroom for whichever non-pinned model
the user invokes (qwen3.6-27b at ~30 GB fits cleanly).
Updated the pinned group's docstring to capture the current member set
+ VRAM math + the historical context (qwen3.6-35-a3b was here, was
too heavy, got removed yesterday). Marked the granite ttl: 0 with the
matching "pinned — never unloads" comment as the other group members.
Symptom: qwen3.6-35-a3b refused to deload when other models needed
the VRAM, even with the model itself at ttl: 0. The pinning came from
the `pinned` group's `persistent: true` flag, which exempts members
from eviction by the scheduler regardless of memory pressure. The
model's ttl: 0 only governs idle-timeout, NOT scheduler eviction —
those are separate concerns.
Removed qwen3.6-35-a3b from the group's members. Kept ttl: 0 on the
model itself: still no idle-unload, but the scheduler CAN now evict
it when another non-coexistent model is requested. qwen3.5-9b stays
pinned (~6 GB at Q4 — cheap to hold).
Updated the inline comment + the group-header docstring to reflect
the new semantics so future-Claude doesn't undo this.
The base qwen3.6-35-a3b is already ttl: 0 via the `pinned` group.
The three other Qwen 3.6 variants (abliterated, heretic, 27b) had
ttl: 600 → llama-swap auto-unloaded them after 10 min idle, costing
the next request a full reload (~5-15s). Removed so they stay loaded
once warm. Still get evicted by the normal swap when another
non-pinned model is requested — these aren't joining the pinned group,
just losing their idle-unload timer.
devnen/Chatterbox-TTS-Server doesn't expose /health — neither in code
nor OpenAPI. The deploy hung on the playbook's `Wait for /health to
respond` loop indefinitely (each curl -> 404, retry forever) even
though the container was up and the model loaded clean to CUDA at
22:52:21 (~42s after start).
/api/model-info returns `{"loaded":true,...}` only after the model
finishes loading, so it doubles as liveness + readiness. Updated:
* compose.yaml healthcheck — grep for `"loaded":true` from
/api/model-info.
* playbook wait step — same probe instead of /health.
* verify /health → verify /api/model-info reports loaded.
* verify /v1/audio/voices — switched from greping for `voice|alloy|echo`
literals to parsing JSON and asserting the actual response shape:
`{"status":"ok","voices":[...]}` (devnen's shape — note this is NOT
the OpenAI list-format vibevoice uses).
Build + container + /health all came up clean on the re-run; only the
voices-endpoint verify failed. The check greped the response body for
"voices"/"voice"/alloy/Carter — but VibeVoice's actual response shape
is OpenAI list-format `{"object":"list","data":[...]}`, which contains
none of those substrings. On a fresh install the data array is also
empty (voices live at /worktank/vibevoice/voices/ and the user seeds
them).
Switched the check to parse the JSON and assert the shape (object="list",
data is a list). Robust against empty voices, robust against future
schema additions.
Both deploys failed against irv-ml1 today with upstream-changed-on-us
errors:
* vibevoice: VIBEVOICE_SHA=7614c469a145 (12-char short) made docker
buildx report "repository does not contain ref 7614c469a145" — same
commit IS still HEAD of main, but buildx's git source resolver
doesn't accept short hashes even when unambiguous. Now full 40-char.
* chatterbox: dockerfile: docker/Dockerfile.gpu — devnen restructured
the repo to put Dockerfiles at root, renamed by CUDA version
(Dockerfile.cu128, .cpu, .rocm). Switched to Dockerfile.cu128 (GPU
build for CUDA 12.8 toolkit; works on irv-ml1's 595.58.03 driver).
Also pinned CHATTERBOX_SHA to a full 40-char SHA instead of `main`
so future upstream churn doesn't break the deploy without warning.
Live host .env files patched directly (the playbook only seeds .env
when absent, so canonical edits don't propagate to existing installs).
irv-ml1's driver upgrade to 595.58.03 (kernel 6.1.0-37, CUDA 13.2) is
working — both GPUs detected, modules loaded. The gpu variant of the
Kokoro-FastAPI image (which requires CUDA >= 12.9) is now the right
default for new deploys. Flipping KOKORO_VARIANT=gpu, KOKORO_USE_GPU=true,
KOKORO_GPU_DEVICES=0 (pins to the RTX 3090 — Kokoro is ~1 GB VRAM and
doesn't need the A6000).
Driver bump survived after all (595.58.03, kernel 6.1.0-37, both GPUs
detected and modules loaded). 5 GPU stacks back up clean (comfyui,
cosyvoice, qwen3-tts, index-tts, parakeet — all healthy). Homepage
discovery can resume polling 10.100.79.3:2375 over the WG tunnel.
Two new sections:
* "llama-swap — added two vision-capable Qwen 3.6 entries" documents
the heretic + 27b additions, their pre-pull into HF_HOME=/hfcache
via the one-shot python:3.12-slim + hf_transfer recipe (4:10 and
3:46 wall-clock for 29 GB and 26.5 GB respectively), and the fact
that llama-server's -hf flag auto-loads mmproj when present.
* "Stack tree convention (canonical vs mirror) — clarified" captures
the deploy-stack.sh-was-reading-from-the-wrong-tree bug and the
resolution: stacks/<stack>/ is canonical/intent (deploy source),
stacks-mirror/<host>/<stack>/ is gitignored snapshot for drift
detection only. CLAUDE.md and memory updated separately in the
prior commit.
Decision recorded in CLAUDE.md ("Stack tree convention") and memory
(convention_stacks_vs_mirror.md):
stacks/<stack>/ canonical / intent. git-tracked.
deploy-stack.sh reads from here.
stacks-mirror/<host>/<stack>/ snapshot / reality. gitignored.
sync-stacks.sh writes here. Used
for drift inspection only — never
a deploy source.
Bug this fixes: deploy-stack.sh was reading from the mirror, so edits
to stacks/llama-swap/config.yaml never reached ana-ml2. Today's
two new model entries (qwen3.6-35-a3b-heretic + qwen3.6-27b) lived
in the canonical for hours but the deploy reported "in sync" because
the script only diffed mirror vs server.
Changes:
* deploy-stack.sh: source switched from MIRROR_DIR/$HOST/$STACK to
STACKS_DIR/$STACK. Header comment + error message updated.
* sync-stacks.sh: header explicitly identifies its role as drift
detection; documents the diff command for comparing canonical vs
mirror.
* stacks/llama-swap/{config.yaml → conf/config.yaml}: matches the
deploy mapping (conf/ in canonical → /opt/docker/conf/ on host).
* CLAUDE.md: "Stack mirror (pull / push)" section rewritten as
"Stack tree convention (canonical vs mirror)" with the role table
+ workflow rules + diff recipe. Layout diagram updated.
Both models pre-pulled into /tank/aimodels/huggingface (HF_HOME=/hfcache
inside the container) via huggingface_hub.snapshot_download with
hf_transfer for parallel chunked download — heretic's 29 GB landed in
~4 min, unsloth's 26.5 GB in ~3:46 (~118 MB/s each).
heretic: llmfan46/Qwen3.6-35B-A3B-uncensored-heretic-GGUF:Q6_K
27b: unsloth/Qwen3.6-27B-GGUF:UD-Q6_K_XL
Both repos include mmproj-BF16.gguf alongside the main GGUF, and
llama-server's -hf flag auto-loads the mmproj when present in the same
repo (-hf docs: "mmproj is also downloaded automatically if available").
So both entries get vision (image-text-to-text) without needing an
explicit --mmproj path. ttl: 600 (10-min idle unload), matching the
existing abliterated entry's style.
The server-rendered .source-count / .desk-count badges were correct
at render time but went stale the moment the user hid anything —
"r/HOMELAB (4)" stayed at 4 even after all 4 items were hidden.
Worse, the entire source header still rendered with a (0) badge
once every item underneath was gone.
app.js gains a refreshCounts() pass that walks every .source and
.desk, recomputes the visible (non-.is-hidden) child count, updates
the badge text, and toggles an .is-empty class. CSS rule for
.source.is-empty and .desk.is-empty sets display:none so empty
groups collapse out entirely. Hooked into hideItem, restoreItem,
and the initial-paint hidden-set application.
New section documenting the architecture change (news-digest-web
moved from nginx:alpine to a FastAPI app on uvicorn built from the
worker's same Dockerfile), the three new endpoints
(GET /api/hidden, POST /api/hide, POST /api/restore), the item-id
scheme (12-char sha1 of reddit:<post_id> or miniflux:<entry_id>
so hide-once = hide-forever-for-that-article), and the playbook
changes (dropped DOCKER_BUILDKIT=0 now that ana-docker is on
docker-ce 29, added round-trip API verify steps).
Adds a small × on each item that hides it from the page. State is
server-side at /output/hidden.json so the same hidden set follows
the user across devices (home, ipad, laptop, work). A "Hidden (N)"
tray at the bottom shows what's hidden on the current page with a
restore button per row; older hidden ids that aren't on this page
sit silently and continue to filter future editions that include
the same article.
Architecture change: news-digest-web swaps from nginx:alpine to a
FastAPI app on uvicorn, built from the same Dockerfile as the
worker. Same image, different command (`uvicorn web:app` overrides
the worker's cron entrypoint via compose). Drops one image dependency,
adds /api/{hidden,hide,restore}.
Item ids are stable 12-char sha1 prefixes (`reddit:<post_id>` /
`miniflux:<entry_id>`) computed in digest.py at render time and
emitted as `data-id` on each .item. The frontend reads /api/hidden
once on load, applies `is-hidden` to matching items, and POSTs
hide/restore on user interaction (optimistic, with rollback on
network error).
Storage: single JSON array at /output/hidden.json, atomic writes
via tempfile + rename, threading.Lock around the read-modify-write
inside the single uvicorn worker. No auth — the digest itself is
unauthenticated on LAN; same trust boundary applies.
Playbook also drops the DOCKER_BUILDKIT=0 fallback now that
ana-docker is on docker-ce 29, and adds three verify steps
(/api/hidden returns a JSON array, app.js is reachable, full
hide/restore round-trip with a synthetic id).
autorestic removal completed on both esh-docker-vm and esh-vm-db
after two playbook fixes (YAML plain-scalar folding ate a backslash
continuation; YAML tag indicator stripped a leading `!`). Both
documented inline.
seafile seahub race resolved by adding a healthcheck to mariadb
(bundled healthcheck.sh --connect --innodb_initialized) and
converting seafile's depends_on to long-form with
condition: service_healthy on db. Compose now waits for InnoDB
to initialize before starting seahub, so the daemon-restart race
that wedged the python frontend can't recur. Verified: seahub log
clean post-recreate, traefik 502 rate dropped to zero on
seafile@docker. Compose change lives on the server (the mirror is
gitignored by design).
YAML treats a leading `!` as a tag indicator, so the unquoted
`shell: ! command -v autorestic >/dev/null` was parsed as a tagged
scalar with the `!` stripped. The verify ended up running just
`command -v autorestic >/dev/null` — which exits non-zero when
autorestic is absent, the OPPOSITE of what the assertion needed.
Quoted version `"! command -v autorestic >/dev/null"` survives
parsing and gives the intended bash negation.
The previous version listed four unit paths separated by `\` + newline.
That looks fine in source but YAML plain-scalar folding collapses the
sequence to a literal `\ ` — the backslash + space no longer functions
as a shell line continuation, and only the first path actually gets
passed to rm. End result on esh-docker-vm's first run: backup.service
removed; backup.timer + prune.service + prune.timer survived; verify
correctly caught the partial state.
Switched to `rm -f /etc/systemd/system/autorestic-*.{service,timer}`
form — single string, no folding hazard, and idempotent on hosts where
some or all of the files are already gone. Re-running on esh-docker-vm
will mop up the leftovers cleanly.
Migration complete:
* ana-docker on docker-ce 29.4.1, all 29 containers back up. Traefik
routing live (verified 200s on matrix.phasefinal.com presence +
seafile.phasefinal.com syncs).
* traefik-postboot.service installed + enabled on both traefik hosts
(esh-docker-vm, ana-docker) — one-shot systemd unit that restarts
traefik 60s after every boot, fixing the long-standing routing-races-
after-reboot symptom.
New playbook: remove-autorestic. Triggered by a typo (`D:escription`
in autorestic-backup.timer line 2) flagged by systemd-analyze during
the traefik-postboot install on esh-docker-vm. Rather than fix it,
remove autorestic — it's redundant with the PBS + structured-restic
two-layer pipeline that's been operational since 2026-04-22. Detected
on two ESH-side hosts: esh-docker-vm and esh-vm-db. Playbook removes
the four unit files + the /usr/local/bin/autorestic binary; leaves
/srv/backups/autorestic/.autorestic.yml (archival) and
/mnt/backup/restic/repo/esh (historical snapshots) for separate
disposition.
Sub-finding from ana-docker upgrade: seafile's seahub (the Python
frontend at port 8000 inside the container) failed to start because
mysql wasn't ready when seafile booted, and a single restart didn't
recover it. Traefik routes return 502 on seafile dynamic endpoints
until seahub is up. Needs separate triage of seafile's depends_on
wiring or seahub's retry behavior — not a docker-ce regression.
Traefik often misses backends after a reboot or daemon swap because
(a) its docker provider debounces / drops events when 30+ containers
start in a burst, and (b) backends can be `Created` on the docker
socket but not yet attached to traefik-net when traefik scans. The
empirical workaround is `docker restart traefik` once the topology
settles — this unit bakes that in.
Type=oneshot, After=docker.service, ExecStartPre=/bin/sleep 60,
ExecStart=docker restart traefik. Runs once per boot. delay_seconds
and container name are tunable via --var.
Verify phase: file mode, enabled state, ExecStart references the
right container, container actually exists on the host, and
systemd-analyze parses the unit cleanly (lint without executing —
avoids needlessly bouncing traefik on healthy hosts).
In scope: esh-docker-vm, ana-docker (the two hosts that run traefik).
docker-ce 29 ships docker-compose-plugin renumbered to v5.x (was v2.x
with docker-ce 26-28). Same Compose v2 codebase under the hood —
Docker just realigned the major number. The verify regex was hardcoded
to `v2\.[0-9]+\.[0-9]+`, so a successful migration on esh-docker-vm
(29.4.1, 16/16 stacks back up clean) reported FAILED on the verify
phase. Switched to `docker compose version --short` parsed for major,
gated `>= 2` — works across future plugin renumbers too.
STATUS.md: mark esh-docker-vm done. ana-docker is the last host.
After nh3-docker's swap, two systemd unit gotchas surfaced that the
playbook now handles automatically:
* The docker.io-era /etc/systemd/system/docker.service.d/override.conf
hardcoded ExecStart=/usr/sbin/dockerd; docker-ce installs at
/usr/bin/dockerd → daemon failed status=203/EXEC.
* The shipped docker-ce unit's ExecStart=dockerd -H fd:// conflicts
with daemon.json hosts: (defined for the 0.0.0.0:2375 homepage
discovery binding) → "conflicting host options".
The "Rewrite docker.service drop-in" step now backs up any existing
override, probes daemon.json for a hosts: setting, and installs an
override that strips -H from ExecStart when needed. Also added an
explicit systemctl reset-failed step to clear the start-rate-limit
state that 3 failed install-time starts leave behind.
configs/homepage/docker.yaml: comment out irv-ml1-docker provider —
20s-per-poll ETIMEDOUTs from the stalled host were drowning homepage's
logs and apparently blocking ana-pfi-docker discovery (the Miniflux
card in the News group wouldn't render until removal). Re-enable when
irv-ml1 is back.
STATUS.md: new "Active migration" section tracking the docker-ce
rollout — nh3-docker done; esh-docker-vm + ana-docker queued.