Temporary diagnostic for the class of bug story 83ff386d47c6 hit
2026-05-23: POST /generation/start returned 202, then total silence
— no log, no DB state update, py-spy showed event loop idle with no
GenerationRunner frame anywhere. Strongly suggests a created_task()
result not held → GC'd → silent destroy.
PYTHONASYNCIODEBUG=1 emits "Task was destroyed but it is pending"
and "Task exception was never retrieved" warnings to stderr; that
should distinguish lost-task from cancelled-task on the next attempt.
Per skaldsong-dev's note, remove once they wire proper task-exception
capture upstream.
Diagnosis thread: althing 01KSBGKQBXA756JWW1KD4MPGXM
The compose set SKALDSONG_DB_PATH + SKALDSONG_RUNS_DIR, but skaldsong's
app reads SKALDSONG_HOST_SQLITE_PATH + SKALDSONG_HOST_RUNS_ROOT (per
its Dockerfile ENV defaults). Our values were orthogonal — the app
fell back to Dockerfile defaults pointing at /app/data/... which is
NOT bind-mounted, so every --force-recreate wiped the SQLite DB +
runs/ tree along with the ephemeral container layer.
Surfaced by skaldsong-dev (althing thread 01KS4DPF6SXTBP4Q360JZVWPNT)
after the operator noticed stories vanishing on every deploy.
Confirmed on ana-docker: container had a 40KB skaldsong-ui.db sitting
in /app/data/, while /opt/docker/conf/skaldsong/db/ on the host was
empty. Rescued the live DB to the bind-mount target before recreate.
Fix: rename env vars to match what the app reads. Bind targets stay
at /app/state/{db,runs} (parent-dir mount for SQLite WAL+SHM).
Two corrections surfaced by the first end-to-end deploy that didn't
land in the pre-flight align:
- SPA static assets are at /app/spa, not /app/web/dist (Dockerfile
COPYs the SvelteKit build output flat into /app/spa, not into
/app/spa/dist). Mismatch caused /health to 500 with
"RuntimeError: File at path /app/web/dist/index.html does not
exist."
- SKALDSONG_HOST_CORS_ORIGINS must be a JSON array literal in .env.
Pydantic-settings parses complex-typed env vars via json.loads();
bare URL string fails first-boot with SettingsError.
Container now reports Up (healthy) on ana-docker; /health 200.
skaldsong-dev surfaced three contract corrections before the first
deploy:
- WORLDTREE_TOKEN (outbound HTTP Bearer) was missing — separate code
path from SKALDSONG_BIFROST_JWT_KEY (inbound HS256 verify) but
same secret value.
- WORLDTREE_BASE_URL replaces SKALDSONG_WORLDTREE_API_URL (the
former is what the app actually reads).
- SKALDSONG_HOST_WIZARD_AGENT_ID was missing entirely — must pin to
skaldsong:wizard-v2 to inherit the existing Worldtree agent slot;
blank would burn another slot of the 50-per-key Heimdall quota.
Registry-pull pattern matching Worldtree: CI on vh/skaldsong builds and
pushes gitea.phasefinal.com/vh/skaldsong:<sha>, this playbook pulls +
recreates. SHA-pin only per current preference; no :latest moving-tag
advance yet (revisit once /health exercises Worldtree + Kokoro reach).
Host port 8300 (host) → 8000 (container). Persistent state under
/opt/docker/conf/skaldsong/{db,runs}.
Bifrost endpoint URL 10.250.50.70:8300 will need a paired
BIFROST_CLIENT_ALLOWED_HOSTS update on corviduo-dev Worldtree at first
deploy.
Phase 3.1 closes the cross-process gap the Phase 3 smoke surfaced —
streaming events (msg_start/thinking/delta/complete/curated) flow
from agent-runner → chamber via valkey pub/sub rather than the
SQLite bridge (too high-volume + ephemeral for the DB).
New service: `althing-valkey` (stock `valkey/valkey:8-alpine`).
Internal-only — no exposed port, no volume. chamber + agent-runner
reach via docker DNS at `valkey:6379` on the compose default
network. healthcheck via `valkey-cli ping` (5s interval). chamber
+ agent-runner gain `depends_on: valkey: service_healthy` so the
bridge is up before either side starts publishing or subscribing.
Forseti unchanged — never publishes Phase 3 events.
Operational properties (per forseti's deployment notes):
- Mixed-state safe at every step. Missing valkey.url config key
→ chamber + runner stay on v3.0 / Phase 2 equivalent paths.
- Backward path is single config-key delete + restart.
- streaming_enabled: true (set on agent-runner 2026-05-17) is
unaffected by this change.
README's services table + playbook header + verify section all
extended to reflect the four-service shape. Forseti's contract
at vh/althing:docs/contracts/phase3_1_valkey_bridge.contract.md
carries the wire-protocol spec.
Phase 2 daemon added to the althing-chamber stack per forseti's request
(vh/althing@5cd088a..ad1d025). Polls floor_grants WHERE consumed_at IS
NULL AND agents.driver='worldtree', claims via atomic UPDATE, calls
Worldtree's conversation API, posts the response back through the bus
as a broadcast.
Shape matches the existing forseti daemon:
- Same ${ALTHING_IMAGE} (the binary is already in [project.scripts]
as of ad1d025)
- command: ["althing-agent-runner"]
- Same shared SQLite bind-mount at /app/data
- No port, no healthcheck (CLI doesn't expose one; same liveness
story as forseti)
Safe to enable preemptively per forseti — when no driver=worldtree
handles are declared in config, the runner sleeps at
poll_interval_seconds. Multi-instance safe via the atomic claim
primitive (no flock needed).
Compose top comment, README "Services in this stack" table, playbook
header + verify steps all extended to reflect the three-service
shape. Will land on ana-docker on vh/althing's next push (compose
deployed via the elway playbook's upload step; image already carries
the binary).
Two-service compose (chamber + forseti sidecar daemon) sharing a single
SQLite store via bind-mount under /opt/docker/conf/althing-chamber/data.
eventbus.bridge_from_db is the cross-process glue — forseti's commits
reach chamber's SSE subscribers via the bridge.
Pattern matches task-board's build-on-host deploy:
- elway playbook clones vh/althing into /opt/docker/build/
- docker build -t althing-chamber:local . (no registry)
- playbook uploads compose + seeds .env one-time, brings both
services up, polls /health
- Gitea Actions workflow lives in vh/althing; reference copy here.
Internal tooling — host port 7881 (chamber's default of 7878 collides
with task-board). LAN-direct, no Traefik. Container always listens on
8000 internally.
Scaffold will fail to bring the chamber container up healthy until
galdrabok-side commits land:
- Dockerfile at vh/althing repo root (two-stage: uv-bookworm-slim
build → python:3.12-slim runtime, locked per open_questions §2
of the v1 contract).
- GET /health endpoint on the chamber app (200, no DB read).
- ALTHING_BIND / ALTHING_PORT env-var support in
core.cli.chamber_serve / core.chamber.cli (env > config.yaml >
defaults precedence).
Coordinated via althing thread 01KRMAK7RD7TP6C8DF4KXV31RT.
Stack was retired and replaced by the vllm stack (originally vllm-qwen3,
renamed 2026-05-13). Its README still framed it as a current solution
while ana-ml2's README + vllm's README both documented the retirement.
stacks/vllm/README.md "Migrating off Infinity" step 3 explicitly said
"Delete stacks/infinity/ from this workspace" — actioning that now.
No backwards-compat shims (PRACTICES §4): contract of a deleted system
has no historical value the next contributor needs; the replacement
path is documented in stacks/vllm/README.md.
Surfaced by /tend-docs audit 2026-05-14.
Two related changes shipped together. The stack rename is independent
but adding `vllm-reward` to the existing `vllm-qwen3` would have made
that name actively misleading.
**Rename:** `stacks/vllm-qwen3/ → stacks/vllm/`. Updated all in-repo
references (README.md root, servers/ana-ml2/, stacks/llama-swap/,
configs/restic/ana-ml2/, docs/runbooks/disaster-recovery.md). Two
intentional history mentions retained (servers/ana-ml2 + stacks/vllm
README).
**Add `vllm-reward` service:** serves Skywork-Reward-V2-Llama-3.1-8B-AWQ
on port 8003. The AWQ output is a locally-quantized model (not from HF),
so bind-mounts `/tank/aimodels/llm:/local-models:ro` rather than the
shared HF cache. Model config.json declares LlamaForSequenceClassification
which vLLM's pooling runner picks up automatically — produces a single
reward score per input via /classify.
**Flag note:** the user's spec listed `--task classify`, but vLLM 0.19.1
deprecated --task in favor of --runner pooling (model architecture in
config.json drives the classification head). Compose uses --runner
pooling with a comment explaining the substitution.
**GPU memory:** no rebalance needed — production had already tuned
EMBED/RERANK down from 0.40 to 0.20 each (canonical .env.example now
matches reality). Adding REWARD at 0.30 totals 0.70, leaving ~14 GB
headroom on the 48 GB Ada.
**Server-side:** brought existing vllm-qwen3 down, mv'd
/opt/docker/compose/vllm-qwen3 → /opt/docker/compose/vllm, appended
REWARD_* lines to existing .env (preserving API_KEY/HF_TOKEN), deployed
new compose via scripts/deploy-stack.sh, brought all 3 services up.
**Smoke tests:**
- /health on 8001/8002/8003 → 200
- /v1/models on 8003 → lists Skywork/Skywork-Reward-V2-Llama-3.1-8B-AWQ
with max_model_len 16384
- /classify with a sample conversation → returns LABEL_0 with prob 0.9999
(single-output regression-style reward score, expected shape for a
reward model)
Playbook handles models, datasets, and spaces (via --var repo_type=...)
since 3025d49 — the "-model" suffix was misleading. Renaming to match
actual scope.
Updates the single in-repo reference (changelog comment in
stacks/llama-swap/conf/config.yaml). config.yaml was scp'd to ana-ml2;
no docker compose restart needed (comment-only).
AtlaAI's Selene-1-Mini judge model for evaluation/scoring tasks.
Llama 3.1 8B base, mradermacher imatrix-quantized Q6_K (~6.5GB,
quality-leaning quant). Apache-2.0. Per Atla cookbook these defaults
hit 84% on RAGTruth hallucination eval.
New 'JUDGE / EVAL MODELS' section between the dense chat models and
the embedding models — separate category from chat/reasoning since
the run-params shape is different (deterministic-leaning: temp 0.01,
top-p 1.0, no repeat penalty).
q8_0 KV cache to fit 32K ctx cleanly on the 3090 with headroom.
Pre-pulled into the shared HF cache via the new
playbooks/pull-hf-model.yaml playbook (canonical replacement for
ad-hoc huggingface_hub.snapshot_download calls; see CHANGELOG).
Smoke-tested 2026-05-13: GET /v1/models lists selene-1-mini-8b,
POST /v1/chat/completions returns expected output cleanly.
Three changes prepping infra for asset_engine's orchestration feature
(SSH-driven bring-up / bring-down of irv-ml1 inference services with
per-device VRAM gating, contract in vh/asset-engine commit 5a36f8c):
1. asset-engine compose + .env.example + playbook gain a read-only
bind-mount for /app/runtime/ssh — the dedicated ed25519 keypair
(generated on ana-docker, not in the repo) plus a pinned known_hosts
for irv-ml1's host fingerprint. Env vars SSH_KEY_PATH and
SSH_KNOWN_HOSTS are exposed for the app to consume.
2. docs/asset-engine/services.yaml gains a `lifecycle: { stack, vram_gb,
gpu_device_id }` block on each of 12 orchestratable irv-ml1 services
(kokoro, chatterbox, index-tts, qwen3-tts, cosyvoice, fish-s2,
kyutai-tts, vibevoice, voxtral, parakeet, stable-audio-open, ace-step).
VRAM numbers are estimates from model footprint at fp16 — tune from
real nvidia-smi measurements once the gate is live. comfyui and
kokoro-captioned are deliberately excluded (variable-VRAM and
shared-container respectively).
3. servers/irv-ml1/README.md docker-stacks table now lists all 13
inference stacks (was only dockge + agents + comfyui) with port +
GPU pinning columns.
Pubkey deployed to ~lkraven/.ssh/authorized_keys on irv-ml1;
end-to-end SSH from ana-docker → irv-ml1 verified with strict
host-key checking.
Adds VOR_WORLDTREE_KEY + VOR_WORLDTREE_BASE + VOR_WORLDTREE_MODEL to vor's
compose environment with sane defaults. Empty key falls back to the
in-process MockWorldtree (the /mockup/ surface returns canned fixtures);
a real key issued by architect routes LLM calls at the demo Saga instance.
Key itself lives in ana-docker:/opt/docker/compose/vor/.env (not in the
repo).
Internal tooling — accessed at http://10.250.50.70:8200, not through
Traefik. Removes the unused traefik labels (router rule, TLS, crowdsec
middleware, loadbalancer port) and the traefik-net network membership;
homepage.href now points at host:port for direct discovery, matching
task-board's pattern. Playbook verify drops the traefik-net membership
check.
Mirrors task-board's build-on-host pattern: elway playbook clones
vh/asset-engine into /opt/docker/build/, docker build, install compose +
seed .env, up -d, verify /health. No registry.
Internal-only tool — LAN port 8200 (bind 0.0.0.0) is primary; Traefik
labels additionally route asset-engine.phasefinal.com with TLS via the
anaprod cert resolver. DB and outputs are separate bind-mounts under
/opt/docker/conf/asset-engine/ so outputs/ can move volumes later
without touching DB state. INFERENCE_HOST defaults to 10.100.79.3
(irv-ml1 over WG). OIDC env seam is pre-allocated empty for v2.
The pre-fix wrapper at stacks/ace-step/infer-api.py returned a JSON
{output_path: "..."} reference to a file written inside the
container at /app/outputs/. That path was unreachable from outside
the container — every consumer got 134 bytes of JSON-pretending-to-
be-WAV instead of audio. Surfaced by the asset_engine consumer's
end-to-end smoke (althing thread 01KRCJF7NGMXYE9F62Q1A6KFD4 msg 5);
my own earlier smoke missed it because I checked HTTP=200 and stopped
reading instead of inspecting the response body.
Wrapper now reads back the file the pipeline writes and streams the
bytes via fastapi.responses.Response with media_type set from the
audio_format request field (audio/wav | audio/mpeg | audio/flac).
The in-container path is exposed via X-Output-Path header for log
correlation but is no longer load-bearing.
Verified end-to-end against live ace-step on irv-ml1:
POST /generate -> HTTP 200 in 80s
content-type: audio/wav
content-length: 945226
x-output-path: /app/outputs/output_cfe87d1d....wav
$ file response.wav
RIFF (little-endian) data, WAVE audio, Microsoft PCM, 16 bit,
stereo 48000 Hz
Catalog: ace-step bumped version 3 -> 4. Dropped
response.output_field (no longer applicable). reproducibility.notes
expanded to record both the v2 18-arg-tuple fix and this v4
inline-streaming change so the history is auditable from the
catalog itself.
Stale ACEStepOutput Pydantic model left in infer-api.py for now —
unused but small; future cleanup.
StableAudioPipeline isn't reentrant — concurrent requests share the
scheduler's step_index counter and corrupt each other mid-run
(observed: IndexError in cosine_dpmsolver_multistep when two requests
overlap). Wrap the pipeline call + audio decode in a single
asyncio.Lock created at startup, and run the (sync, GPU-bound)
pipeline call via asyncio.to_thread so the event loop stays
responsive. Concurrent requests now queue cleanly instead of racing.
Verified: 5 parallel POSTs at steps=50 all return 200, clear ~4s
serialization spacing (4, 8, 12, 16, 20s wall time), distinct
output hashes per seed.
server.py accepted cfg_scale in the request schema and the README
documented its 0–20 range, but the pipeline call never received it
— so changing cfg_scale between requests silently produced identical
output (the pipeline ran at its own default of 7.0 every time). Add
guidance_scale=req.cfg_scale to the pipe(...) call.
Verified: (prompt, seed, steps) held constant, cfg_scale=3.0 vs 15.0
now produce different SHA256s; same triple at cfg_scale=7.0 is
deterministic across repeated calls.
Wrapper only enumerates one voice directory (settings.voices_dir,
default /app/api/src/voices/v1_0 — inside the container's writable
layer, not bind-mounted). Override via VOICES_DIR=/app/user_voices
(host bind mount) and add a command shim that cp -r's built-ins from
the in-image v1_0 into user_voices on every start. Built-ins re-seed
fresh from the image (so upgrades that add voices propagate); custom
.pt files in user_voices are preserved (cp -r is additive).
Also adds scripts/blend_kokoro_voice.py + a playbook around it that
mirrors the wrapper's request-time voice="a(w)+b(w)" math but writes
the result as a named .pt to user_voices, making it discoverable via
GET /v1/audio/voices and persistent across recreate. Defaults to
athena = af_bella(2)+af_aoede(1) normalized.
Same shape as task-board: build-on-host from vh/vor, bind-mounted
persistence for sessions/ and responses/ (the user-published markdown
files), exposed at port 7879 (adjacent to task-board's 7878 since both
are claude-tooling sidecars).
Workflow template assumes the same DEPLOY_SSH_KEY + MGMT_REPO_TOKEN
secrets at user scope; nothing new to provision. Playbook accepts SHA
or branch refs (same fix as deploy-task-board.yaml) so manual runs
and CI runs share the same code path.
Centralized vs upstream-local: README documents the trade. Claude
fetches response markdown via /api/sessions/{id} JSON instead of a
local file read — the only API-flow change from the upstream README.
Runner is now re-registered with `:docker://node:20-bookworm-slim`
schema in its labels, so workflows targeting `pfi-fleet` get that
image automatically. Saves a few lines per workflow and gives us one
place (the runner config) to bump the default image when a new
node/debian release lands.
Runner's .runner registration cached :host mode at first start; env-var
label updates aren't sticky once the runner is registered. Until we
re-register with docker-schema labels, workflows must declare their
own container. node:20-bookworm-slim has node (for actions/checkout)
and apt (for python3-yaml + openssh-client install).
debian:bookworm-slim lacks node, so actions/checkout@v4 (a JS action
running dist/index.js) fails with `exec: "node": executable file not
found in $PATH`. Dropping the explicit `container:` directive lets
the runner use its label-default — node:20-bookworm-slim has node +
git out of the box. Install step shrinks to python3 + pyyaml +
openssh-client.
Central runner on ana-docker (gitea is local; existing fleet tooling
already SSHes from there). Playbook is parameterized so future
site-local runners (nh3-docker, esh-docker-vm) drop in via --var
overrides instead of copy-paste.
Includes a workflow template for vh/task-board that calls the existing
deploy-task-board.yaml playbook — keeps the playbook as the single
source of truth for "how task-board is deployed", manual or automated.
Labels embed `:docker://node:20-bookworm-slim` schema; without it,
act_runner v0.6+ silently falls back to host-mode and runs job steps
inside the Alpine runner container (no apt/python/node), breaking any
real workflow. node:20-bookworm-slim is small + has git + node so
actions/checkout works out of the box.
The applet outgrew "stack alongside the infra-management workspace" —
it has its own pyproject, multi-tenant deploy story, separate
release cadence, and isn't actually about managing infrastructure.
Lives at https://gitea.phasefinal.com/vh/nevermore now, with
provenance noted in its initial commit.
This commit removes:
stacks/news-digest/ (full stack tree)
playbooks/deploy-news-digest.yaml
scripts/add-digest-user.sh
The existing ana-docker deployment continues running on its baked
local/news-digest:v5 image — nothing changes for the live install
until you choose to redeploy from the new repo. Migration steps
(rename data dir, redeploy, retire old compose dir) are in
nevermore's README.
Updated:
README.md — Current stacks listing now points at the new repo
STATUS.md — milestones entry for the extraction
The compose-side default was still pinning qwen3.5-35-a3b — broken
on launch since its GGUF stopped working months ago. Real .env on
ana-docker overrides to granite-4-small so live deploys are unaffected,
but the default was misleading for anyone forking the stack. Found
via /tend-docs.
The original default model in .env.example was changed to
granite-4-small months ago when qwen3.5-35-a3b's GGUF file started
exiting on launch, but the README still named the old one as
"current". Also bumped the summarization-style description from
"one sentence" to "2-3 sentences" to match the post-trafilatura
prompt rewrite. Found via /tend-docs.
Stock neosmemo/memos:stable, port 5230, SQLite at
/opt/docker/conf/memos/data/. Joins traefik-net and ships homepage
labels (group=Notes) so it auto-appears on the dashboard via docker
discovery — no edit to configs/homepage/services.yaml needed.
First-run bootstrap is via the UI: visit http://10.250.50.70:5230
and create the Host account through the sign-up form.
Playbook idiom note: docker compose pull lines need the literal
block scalar (|) when the grep pattern contains colons — bare-string
shell value made YAML parse the colon as a mapping separator and
elway choked on first try.
Adds a "Customizing the run schedule" section (DIGEST_CRON_AM/PM env
vars, edit-and-recreate flow) and a "Multi-tenant: one instance per
teammate" section covering scripts/add-digest-user.sh end to end:
what it does, the per-user file layout on ana-docker, idempotent
schedule/password updates, and the teardown path.
Updated the stale "two editions per day" intro line to note the
schedule is now configurable.
Hardcoded crontab → render at container start from
DIGEST_CRON_AM + DIGEST_CRON_PM. Defaults match the original
0800 / 2000 so existing deploys are no-ops.
scripts/add-digest-user.sh learns --am and --pm flags so each
teammate's stack can fire on their hours:
scripts/add-digest-user.sh bob --am "0 6 * * *" --pm "0 17 * * *"
scripts/add-digest-user.sh carol --pm "30 18 * * 1-5" # weekdays only
Standard 5-field cron syntax; busybox crond honors the container's
\$TZ. Removed the now-unused stacks/news-digest/crontab file and
the matching COPY in the Dockerfile.
Two pieces:
1) Multi-tenant onboarding via scripts/add-digest-user.sh
Shared miniflux + per-user digest stack. Onboarding a teammate
takes one command (plus a one-time sudo for dir creation):
scripts/add-digest-user.sh <username>
What the script does:
- Reads miniflux admin creds from ana-docker
- Allocates next free port (scans existing digest-*/.env)
- Generates a random password (or accepts one as 2nd arg)
- Creates the miniflux user via the admin API
- Materializes a per-user .env at /opt/docker/compose/digest-<user>/
(inherits NEWS_DIGEST_TAG from the canonical stack so all
tenants run the same image)
- Brings up `docker compose -p digest-<user> up -d`
- Seeds default world/local feeds in the new user's miniflux
- Triggers a first digest run
compose.yaml now uses ${DIGEST_PROJECT:-news-digest} to namespace
container_name + homepage labels. Default keeps backward-compat
for the singleton install — existing stacks unaffected.
2) Masthead overlap on phone widths
Desktop CSS pinned .masthead-edition to grid-row 1, which collided
with .masthead-brand once the mobile media query collapsed both
to grid-column 1. Result: "MORNING EDITION" badge stacked on top
of the "DAILY DIGEST" hero. Reset grid-row to `auto` for all
three masthead children in the ≤720 px breakpoint so they
auto-flow vertically.
Three things were broken on phones:
1. The collapse button I added to .desk-head had no grid placement,
so it auto-flowed into the desk-sub row and looked like a floating
chevron. Made the desk-head grid 4 columns explicit (num | title |
count | collapse) and pinned the button to col 4 row 1.
2. The 720px breakpoint was the only one — everything inherited
tablet rules at iPhone widths. Added a true-phone tier at
≤480 px that hides the section number badge and the rail
gutter, floats chips inline above the title, makes the jumpnav
horizontally scrollable for narrow widths, drops the edition
number, and bumps touch targets.
3. Long URLs / unbroken tokens could push horizontal overflow.
Added overflow-wrap: anywhere on titles + tldrs and overflow-x:
hidden on body as a belt-and-suspenders catch.
Two upgrades to make the digest actually readable:
1) Article-grounded 2-3 sentence summaries (everywhere)
The old prompt got just the title + miniflux's content excerpt,
which for HN/Lobsters/wire feeds is barely more than the title
itself — so summaries paraphrased the title and added nothing.
Now every URL gets fetched and main-content-extracted via
trafilatura on a parallel pre-pass (10 workers, ~15s for ~50
URLs). Extracted text caches to /output/.article-cache.json with
a 7-day TTL so repeat runs in the same window don't re-pull.
Headlines also get summarized now — one batched LLM call per
category (world / local). Rendered as a paragraph below the
title with source + time on the right rail.
Prompt rewrites tell the model to pull names/numbers/places
from the body and explicitly forbid restating the title.
Result: real specifics ("71% saw no pay increase globally",
"third time in less than two weeks", "Islamabad and Moscow
intermediaries") instead of title paraphrase.
2) Per-desk collapse buttons
Chevron next to .desk-count toggles a .is-collapsed class.
Collapsed state is per-device (localStorage by section id) since
collapse is a viewing preference, not content state.
Browsers were serving stale frontend assets after rebuilds, which hid
the new world/local headline desks: the OLD app.js's refreshCounts()
only counted .item children (not .headline), so the new headline desks
came up with visibleItems=0 and got the .is-empty class which is
display:none. Hard refresh fixed it but only for the user who knew
to do that.
Append ?v=<generated_at strftime> to both link/script tags in
digest.html.j2 and archive.html.j2 so every digest run produces a new
asset URL. Works with the existing entrypoint.sh static-asset sync —
no other infra needed.
Two new dense headline rails above the existing reddit/tech cards.
Designed for high-volume "what happened" coverage where the title
is the deliverable — no LLM summarization, ~15 items per section,
6-column-collapsing grid (title / source / time).
Digest pipeline:
* fetch_miniflux_headlines(category) — flat list per category, dedup
by lowercased title (different feeds syndicate the same wire stories)
* 8h look-back window (vs 12h for tech/reddit) since headlines move
faster
* cap of 15 per section (DIGEST_MINIFLUX_HEADLINES_MAX)
Frontend:
* .headline element parallels .item for the hide-button machinery
(both have data-id, both honored by app.js)
* dense 3-col layout collapses to 1-col on narrow screens
* jumpnav now numbers world=01, local=02, reddit=03, tech=04
Setup:
* seed-headlines.py — one-shot script (lives in the image at
/app/seed-headlines.py). Creates the World + Local categories in
miniflux, subscribes a curated feed list, and renames each feed
to a short display title (BBC vs "BBC News", "LA Times" vs "California").
Idempotent — reruns only add new feeds.
* Default world: BBC, NPR, Al Jazeera. Default local: LA Times Local,
LA Times CA, Voice of OC. (OC Register blocks miniflux; left out.)
* entrypoint.sh now syncs templates/{style.css,app.js,favicon.svg}
to /output on container start so frontend asset updates land
without a manual copy after rebuild.
Three upstream gaps surfaced once /generate was actually exercised:
1. infer-api.py builds an 18-arg positional tuple but the pipeline
expects 24 — first missing arg is `format`, so audio_duration
shifts into format's slot and the pipeline calls len() on an
int. Ship a patched copy of infer-api.py and COPY over upstream's
in the Dockerfile. Also handle empty lora_name_or_path -> "none"
(empty string trips HF Hub's repo-id validator).
2. torchcodec + ffmpeg are required by the WAV save path but neither
is in upstream requirements.txt. Without them every /generate
runs to completion and then 500s at write-time.
3. ACE-Step caches checkpoints at /root/.cache/ace-step/checkpoints
(HARDCODED, not honored by HF_HOME). Mount our persistent dir
there so the ~7 GB model survives container recreates.
Bench on A6000 (cached model, lo-fi hip hop, 60-step euler/apg):
10s @ 27 steps -> 9.4s (0.94x)
30s @ 60 steps -> 11.2s (0.37x, ~2.7x realtime)
60s @ 60 steps -> 14.8s (0.24x, ~4x realtime)
Two new audio-generation stacks alongside the TTS slate:
ace-step :8210 — Apache 2.0 music generation foundation model
(hybrid diffusion + LLM). Lyric-aware multi-minute songs. ~10-12 GB
VRAM during inference, A6000-pinned. Custom Dockerfile patches
upstream's torch/cu126 resolution bug (--extra-index-url cu126 was
falling back to pypi-default cu13 wheels, mismatching torchvision).
stable-audio-open :8211 — Stability AI 1.21B latent-diffusion SFX +
ambience. Up to 47s clips at 44.1 kHz. ~6 GB VRAM in fp16,
A6000-pinned. Custom FastAPI shim around diffusers' StableAudioPipeline
(no upstream HTTP server). Dockerfile pins torchsde explicitly —
diffusers doesn't pull it as a hard dep but
CosineDPMSolverMultistepScheduler needs it.
Three deploy iterations + four backend attempts (subprocess CUDA,
resident-server CUDA, Vulkan rebuild) all failed to deliver speedup
over fish-s2:
* CUDA path: ggml_cuda_init succeeded, weights loaded onto GPU per
s2's logs, but nvidia-smi showed 0% utilization during synthesis.
Wall time 20s/long phrase vs fish-s2's 7.5s. The "CUDA get_rows
unsupported for type q6_K" warning hints at incomplete op coverage
in s2.cpp's alpha CUDA backend for fish-speech architecture.
* Vulkan path: vk::IncompatibleDriverError on container init. NVIDIA
Vulkan ICD not accessible inside the container despite
NVIDIA_DRIVER_CAPABILITIES=compute,utility,graphics. Would need
host-side nvidia-utils-vulkan installation or manual ICD bind
mount. Didn't pursue.
Both are fixable — CUDA needs op coverage upstream (author actively
working on it; "selective embedding dequant" commit landed 16 days
ago), Vulkan needs host-side ICD setup. Neither is a config-flip,
both are real work for marginal-or-zero return. Better to delete the
stack and revisit when s2.cpp matures or when we tackle FP8
quantization on ana-ml2's RTX 6000 Ada (sm_89, native FP8 hardware).
Local image rmi'd, /opt/docker/compose/fish-cpp removed on irv-ml1.
/worktank/fish-cpp left for user-side sudo cleanup.
Future Fish acceleration paths (in order of decreasing certainty):
1. Wait for s2.cpp CUDA op coverage to mature (track upstream commits).
2. Quantize Fish BF16 → FP8 via TransformerEngine, deploy on
ana-ml2's RTX 6000 Ada (Ada has native FP8 tensor cores, A6000
doesn't). ~2x speedup if it works.
3. vLLM port of Fish (no upstream support today).
CUDA backend confirmed broken for fish-speech ops on s2.cpp v0.x — alpha,
incomplete op coverage, GPU stays at 0% during generation despite
ggml_cuda_init succeeding. Vulkan was the original README example
(`-v 0`), so likely the more battle-tested path.
Build the image with BOTH backends so we can flip via env without
rebuilding:
* libvulkan-dev + glslc in the build stage (GGML's Vulkan backend
compiles its shaders with glslc at build time; without it the
cmake configure silently disables Vulkan).
* libvulkan1 + the libggml-vulkan.so copy in the runtime stage.
* compose env NVIDIA_DRIVER_CAPABILITIES=compute,utility,graphics —
default nvidia-container-toolkit only mounts compute libs; Vulkan
needs the graphics ICD (libGLX_nvidia + nvidia_icd.json) too.
* entrypoint reads FISH_CPP_BACKEND (cuda/vulkan/cpu) and selects
the appropriate -c/-v/no-flag invocation.
* Default backend = vulkan.
Subprocess-per-request architecture forced CUDA + model load on every
/v1/tts call (~10-20s init, then 5-15s generation). Even though CUDA
is now actually being used (`-c 0` fix landed), 32s for "Verify."
proved per-request init was the bottleneck.
s2.cpp ships a built-in HTTP server (`--server -H -P`) that keeps the
model resident on the GPU. Refactor:
* entrypoint.sh — backgrounds `s2 --server -P 3030 -c 0 -m ... -t ...`,
waits for it to bind 3030, then foregrounds uvicorn. tini supervises
via `wait -n` so either child dying takes down the container.
* server.py — drops subprocess.run; instead httpx-POSTs Fish-shaped
/v1/tts JSON to s2's localhost:3030/generate (multipart form: text
+ optional prompt_text/prompt_audio for cloning). Model load + CUDA
init now happen once at container start, not per-request.
* Dockerfile — added httpx (shim dep), curl (entrypoint readiness
probe), and the entrypoint.sh COPY+chmod. CMD now invokes
entrypoint.sh instead of uvicorn directly.
* deploy-fish-cpp.yaml — uploads entrypoint.sh alongside server.py.
s2.cpp's README example uses `-v 0` which is `--vulkan 0` (Vulkan
device 0), easy to misread as "voice 0". The shim copied that
verbatim, so even after fixing the libcuda.so build problem AND the
libgomp.so runtime dep, every synthesis ran on CPU because the wrong
backend was selected.
Direct verification: `[Model] NPU not compiled, falling back to CPU`
in stderr; nvidia-smi showed no s2 process; bench timed out at 60s
on phrases that fish-s2 (HF, GPU) does in 7s.
s2.cpp's CLI:
-v <id> = --vulkan <device>
-c <id> = --cuda <device>
-M = --metal (Apple Silicon)
Switched the shim to `-c 0`. The CUDA backend IS in the build (-DS2_CUDA=ON
worked, libggml-cuda.so links fine per ldd, libcuda.so.1 mounts at
runtime via NVIDIA container runtime) — just wasn't being told to use it.
Build succeeded after the libcuda.so symlink fix, but the first
/v1/tts request returned HTTP 500 with:
s2 binary failed (rc=127): /usr/local/bin/s2: error while loading
shared libraries: libgomp.so.1: cannot open shared object file
CMake auto-enabled OpenMP during the build (gcc's -fopenmp flag), so
the s2 binary dynamically links libgomp.so.1. The build-stage devel
image had it; the slim cuda:runtime base doesn't ship it by default.
Adding libgomp1 to the runtime image's apt install resolves it.
Second attempt's CMAKE_LIBRARY_PATH + LIBRARY_PATH didn't get picked
up by ggml's nested CMake — same linker errors as the first run.
Robust fix: symlink the stub at /usr/local/cuda/lib64/stubs/libcuda.so
into /usr/local/lib (which ld searches unconditionally) and provide
both libcuda.so AND libcuda.so.1 (the SONAME ggml-cuda's
libggml-cuda.so links against). ldconfig refreshes the cache.
The symlinks live only in the build stage. The runtime image inherits
the real driver-provided libcuda.so.1 via NVIDIA's container runtime
mount, so the stubs never get used at execution time.
Two issues from the first deploy attempt:
1) Build failure (real): linker errors on s2.cpp's CUDA build —
undefined references to cuMemSetAccess, cuDeviceGet, etc. These
are CUDA Driver API symbols (in libcuda.so), not Runtime API
(libcudart.so). The driver lib is provided by NVIDIA's container
runtime at RUN time, not BUILD time.
Fix: nvidia/cuda:devel images ship a stubs library at
/usr/local/cuda/lib64/stubs/libcuda.so that provides the symbols
for linking but is non-runnable. Adding that path via
LIBRARY_PATH + CMAKE_LIBRARY_PATH lets the linker resolve while
leaving runtime unchanged (real libcuda.so comes from the
driver mount).
2) Verify false positive: the /v1/tts verify step's last command was
`rm -f "$out"` — which always exits 0. This made the shell's
final exit code 0 regardless of whether curl/file/grep succeeded,
so verify reported OK even when nothing was running on host_port.
Fix: `set -e` at top + trap-based cleanup. Failures now propagate;
the rm still runs on either path via EXIT trap.
New stack scaffolding for the Fish quantized-realtime experiment. Not
deployed yet — this commit lands the canonical files; deploy follows.
Architecture decisions made in Phase 1:
* CUDA backend, NOT Vulkan. s2.cpp's CMakeLists exposes both
-DS2_VULKAN and -DS2_CUDA; the most recent upstream commit
(2026-04-12) was specifically about CUDA improvements, and CUDA
on the A6000 will be substantially faster than Vulkan for ML
matmul. -DS2_CUDA=ON in the Dockerfile build args.
* Pinned to s2.cpp commit e48ce8e02d8335bd9a0ba94679f605724b31d12
(2026-04-12 HEAD of main). Repo is alpha software per README;
pin tightly so future churn doesn't break our build. Bump
deliberately when wanting upstream improvements.
* Multi-stage Dockerfile: nvidia/cuda:12.6.0-devel for build (needs
CMake + ninja + git + the CUDA toolchain) → nvidia/cuda:12.6.0-runtime
for serve (slimmer; just the s2 binary + GGML libs + a small Python
shim). Cuts image size by ~50% vs single-stage devel.
* FastAPI shim (server.py) wraps s2.cpp CLI in Fish's `/v1/tts`
contract so the same bench harness + clients work against fish-cpp
with no changes. Per-request flow: decode optional reference WAV
from base64 → write to temp → subprocess.run the s2 binary → stream
resulting WAV back. Adds ~50-100ms per-request fork+exec overhead;
negligible vs the multi-second generation cost.
* `streaming: true` accepted in request body but IGNORED — s2.cpp
writes a complete WAV before returning, so chunked output isn't
available. Unlike fish-s2 (HF wrapper) where streaming drops TTFB
to 26ms, fish-cpp's TTFB ≈ total wall time. Speed depends entirely
on raw generation throughput.
* q6_k as default quant — sweet spot per typical GGUF guidance:
near-bf16 quality at ~5GB. Other variants (q4_k_m, q5_k_m, q8_0,
f16) selectable via FISH_CPP_MODEL env.
* Pinned to GPU 1 (A6000) by default to share with fish-s2 for
direct A/B benching. q6_k weights ~5GB + runtime ~3GB ≈ 8GB —
comfortable on either GPU.
* Port 8199 (next free in the irv-ml1 TTS slate).
Phase 2 (next) is the actual deploy + first build. Reserved 30-45 min
for cold-cache build + weights pull.
Voxtral final fix (8th iteration):
* The bundled voxtral_tts.yaml hardcodes gpu_memory_utilization: 0.8
on the language_model stage — overrides the CLI flag. Mounted a
patched copy (0.4) at /etc/voxtral/voxtral_tts.yaml and pointed
--stage-configs-path there.
* With Kyutai stopped to free 5 GB on the 3090, both stages fit
(target 9.4 + 2.4 GB ≈ 11.8 GB; 17 GB free post-kyutai-stop).
* Voxtral now healthy on GPU 0 — bench: 1.9-2.7 s TTFB, real WAV.
Fish s2-pro optimization (per-request sweep, no model swap):
* `streaming: true` in request body drops TTFB from 7.7 s → 0.026 s
(300×). Total time goes up ~1 s (chunked HTTP overhead) but
perceived latency = TTFB. Use stream:true for any interactive use.
* `latency: "balanced"` actually slower than default — bad name; skip.
* `use_memory_cache: "on"` no measurable benefit.
* `chunk_length: 100` (default 200) no TTFB benefit non-streaming.
* Server-side `--half` (fp16 inference) added via compose `command`
override — passes through start_server.sh's $@ unchanged into
api_server.py. Should reduce total time too. Validation pending
the post-restart bench.
Kyutai stopped to free GPU 0 budget — the bench numbers earlier
(3.4 s avg) were unimpressive vs Voxtral's 2.3 s in the same
multilingual slot. Kept the stack files for future re-deploy if
needed; just the running container is gone.
Fourth attempt finally found the right invocation. Voxtral is a
two-stage TTS pipeline (language_model → acoustic_transformer →
audio output), not a flat MistralForCausalLM. Standard `vllm serve`
errored with "no module named 'acoustic_transformer'" because it
loads the model as a vanilla Mistral causal LM.
Pattern from /workspace/vllm-omni/examples/online_serving/
qwen3_tts/run_server.sh (closest in-image analog):
vllm-omni serve <MODEL> \
--stage-configs-path vllm_omni/model_executor/stage_configs/voxtral_tts.yaml \
--host 0.0.0.0 --port 8000 \
--gpu-memory-utilization 0.45 \
--trust-remote-code --omni
Key differences from previous attempt:
* `vllm-omni` binary, not `vllm`
* `--omni` flag activates multi-stage pipeline
* `--stage-configs-path` points at the bundled YAML that maps
stages to GPU + scheduler + worker classes
* Dropped --load-format/--tokenizer-mode/--config-format=mistral
flags — the stage config handles tokenizer_mode internally
* --trust-remote-code is required for the acoustic_transformer
custom code path
Default .env.example now: GPU 0 (3090) with util 0.45 (~10.6 GB
target on 24 GB GPU). The A6000 is fully booked by Fish s2-pro.
Third voxtral attempt: image pulled clean (3 min, v0.18.0), entrypoint
parsed correctly, vLLM started, but engine init failed two ways:
1. HF rate-limited the irv-ml1 IP (38.120.94.3) during the metadata
fetch — 429 Too Many Requests from too many large unauthenticated
pulls today (heretic, 27b, fish-s2, fish-s1-mini, voxtral). Added
HF_TOKEN env passthrough; user generates a token at
https://huggingface.co/settings/tokens and sets VOXTRAL_HF_TOKEN
in .env.
2. Voxtral uses Mistral's native model format (params.json +
tekken.json tokenizer + consolidated.safetensors single file),
NOT HF transformers format (config.json + tokenizer.json + sharded
.safetensors). vLLM errored with "ensure presence of params.json
for Mistral models." Fix: pass --load-format=mistral
--tokenizer-mode=mistral --config-format=mistral to vllm serve.
Confirmed by inspecting the Voxtral-4B-TTS-2603 HF tree:
25 files, ships params.json + tekken.json + consolidated.safetensors.
Both fixes baked into compose. User needs to drop their HF_TOKEN into
.env once and recreate.
Side note discovered while debugging: fish-s2 s1-mini variant uses
the tiktoken tokenizer format; the wrapper can't load it (errors with
"NoneType has no attribute encode" on warmup). So s1-mini isn't a
drop-in optimization for s2-pro — different code path needed. Fish
back on s2-pro for now.
Second voxtral attempt got past the image pull (v0.18.0 published,
~3 min download) but container init failed:
unable to start container process: error during container init:
exec: "--model=mistralai/Voxtral-4B-TTS-2603": stat ...: no such file
vllm/vllm-omni:v0.18.0 has Entrypoint=null AND Cmd=null — there's no
default executable. The compose's `command:` array becomes the full
exec invocation, with --model=... interpreted as the binary name.
Standard vLLM serving CLI is `vllm serve <model> [flags]`. The
binary's at /usr/local/bin/vllm. Set entrypoint: ["vllm", "serve"]
and pass the model as a positional arg.
While we're here: HF cache was empty too (Voxtral 4B BF16 ~8 GB
download on first start) — vLLM auto-downloads from HF on model
load, so no separate pre-pull step needed.