Files
esh-pfi-infrastructure/stacks/scriberr/compose.yaml
T
vh 39da1d4a97 feat(homepage): recategorise on "do I open this?", collapse the API groups
The board mixed tools with endpoints. A vLLM seat whose href is a /docs page
sat in the same band as ComfyUI; the MQTT broker and the RustDesk relay, which
have no page at all, sat in Apps; and `Service Networking` was thirteen members
spanning three AdGuards, five Dockges, two Traefiks and four headless agents.

Every group is now one of two kinds and they never mix. TOOLS are expanded and
sit at the top of their tab. ENDPOINTS — an API, a broker, a background agent,
an href that is /docs or /ping or nothing — carry `initiallyCollapsed: true`
and sit at the bottom. Collapsed is not hidden: the eyebrow and its rule still
render, so the tab still says the thing exists and one click expands it.

A second rule fell out of the same pass and now shapes the group boundaries: a
group's members should all carry a widget or none should. A stat strip makes a
card ~50px taller, so one widget card in a row of plain ones opens a void under
the plain ones. That is why AdGuard and Traefik get their own groups rather
than sharing one with Dockge, and it is most of why the old Service Networking
band looked broken. AdGuard (ANA) was the last short card in its row and now
carries the same query/blocked/latency strip as its two siblings — one
infra-ops AdGuard login authenticates against all three instances, verified
against each; it lives in that stack's .env on the host and is vaulted.

The sixteen GPU-backed model seats were deliberately NOT relabelled.
`homepage.group` is read at container creation, so clearer names for
`AI - Inference` and friends would have cost a recreate on six vLLM seats, four
eval seats and four TTS engines — multi-minute model reloads on endpoints peers
reach through the gateway. Order plus `initiallyCollapsed` buys the same
separation for nothing, so those names stay as they are on purpose.

28 containers that ARE cheap to bounce were relabelled, across five hosts, via
rerunnable elway playbooks. Their label steps are gated on the old value still
being present, so a second run reports skipped rather than churning. Two verify
steps were wrong on first contact and are fixed with the reason recorded: the
traefik check raced its own recreate, and asserting a model seat is "running"
cannot answer "did I bounce it" when a seat may be legitimately stopped —
container age can, and now does.

The canonical stacks/ tree was synced to the deployed labels afterwards, so
intent and reality agree again on all fourteen tracked stacks.

Also documents the real nature of the post-recreate blank dashboard, which cost
~25 minutes here and an hour on 2026-08-19. `initialSettings":{}` in the served
HTML is the catch branch of the page's data loader, not a warm-up and not a
cache — and the error can vanish entirely, because the logger is assigned inside
the same try and the catch only logs if the logger exists. Ruled out by
measurement this time: all four API routes return 200 with correct content while
the page serves {}, and the previous known-good settings.yaml reproduces it
identically. The README now carries the one-command test and the next lead.

Before/after, all four tabs: http://10.100.10.50:8090/b/homepage-relayout/
2026-08-24 08:54:06 -07:00

112 lines
5.9 KiB
YAML

# Scriberr — self-hosted audio/video transcription with speaker diarization.
# Upstream: https://github.com/rishikanthc/Scriberr (Go + SvelteKit, SQLite).
#
# Transcription runs locally via WhisperX (Whisper + pyannote diarization);
# NVIDIA Parakeet / Canary models are also selectable in the UI. Optional
# summarisation / transcript chat talks to any OpenAI-compatible endpoint —
# point it at the LiteLLM gateway rather than a paid API (see README).
#
# ── IMAGE: BUILT LOCALLY, ON PURPOSE ──────────────────────────────────────
# ana-ml2's RTX PRO 6000 Blackwell cards are **sm_120**. Upstream publishes
# `scriberr-cuda` (built for sm_61…sm_89 — no sm_120 kernels) and documents a
# `scriberr-cuda-blackwell` image that **has never actually been published**
# (GHCR returns no tags for it, checked 2026-08-23). The sm_120 path upstream
# ships is `Dockerfile.cuda.12.9` (CUDA 12.9.1 + cu128 torch), built locally.
# Do NOT "simplify" this to the published `scriberr-cuda` image — it will
# fail on these cards or silently fall back to CPU.
# Rebuild: see README "Rebuilding" — checkout lives at
# /tank/scriberr/src/Scriberr on ana-ml2.
#
# ── GPU PINNING ───────────────────────────────────────────────────────────
# Pinned to **GPU1** via explicit device_ids, per the house convention and
# because GPU0 is fully committed to the `gen` seat. GPU1 shares space with
# the `sec` seat, so this stack is a guest there — keep an eye on VRAM.
# NOTE: do NOT add `NVIDIA_VISIBLE_DEVICES=all` (as upstream's compose does).
# It overrides the device_ids reservation and exposes both cards.
#
# All tunables live in .env — edit that, not this file.
services:
scriberr:
image: ${SCRIBERR_IMAGE:-scriberr:local-blackwell}
container_name: scriberr
restart: unless-stopped
ports:
- "${SCRIBERR_BIND:-0.0.0.0}:${SCRIBERR_PORT}:8080"
volumes:
# Bind mounts rather than named volumes: /var/lib/docker on ana-ml2
# lives on zroot with limited headroom, while /tank has terabytes.
# Model weights (Whisper, pyannote, NeMo) land in whisperx-env and are
# multi-GB — they must not go anywhere near the root pool.
- ${SCRIBERR_DATA_DIR}:/app/data
- ${SCRIBERR_ENV_DIR}:/app/whisperx-env
environment:
# ⚠ 10001, NOT the fleet-usual 1000 — this is load-bearing.
# Dockerfile.cuda.12.9 creates `appuser` at uid 10001 (Ubuntu 24.04's
# base image already owns uid 1000 as `ubuntu`, so upstream moved it) and
# chowns /app to 10001. The entrypoint's PUID remapping only chowns
# /app/data + /app/whisperx-env, not /app itself, so running as 1000
# leaves the app unable to open its SQLite DB and it crash-loops with
# `unable to open database file: out of memory (14)` — which is
# SQLITE_CANTOPEN wearing a misleading message, not a real OOM.
# The host bind-mount dirs are therefore chowned to 10001:10001 too.
# Verified 2026-08-23: PUID=1000 crash-loops, PUID=10001 starts clean.
- PUID=${SCRIBERR_PUID:-10001}
- PGID=${SCRIBERR_PGID:-10001}
- APP_ENV=production
# Served over plain HTTP on the LAN. Left at the production default of
# `true`, the session cookie is marked Secure and the browser silently
# drops it — you log in, get bounced back to the login page, and the
# logs show nothing wrong. This must stay false while access is HTTP.
- SECURE_COOKIES=${SCRIBERR_SECURE_COOKIES:-false}
# Upstream defaults to localhost origins only, which fails CORS when
# reached by host IP. Keep this in sync with how the app is reached.
- ALLOWED_ORIGINS=${SCRIBERR_ALLOWED_ORIGINS}
- NVIDIA_DRIVER_CAPABILITIES=compute,utility
# Scriberr builds each model backend's Python env with `uv` at runtime.
# uv's default link mode reflink/hardlinks out of its cache, which fails
# on this overlayfs+ZFS combination with a misleading
# "Failed to clone ... Resource temporarily unavailable (os error 11)"
# and takes out the Parakeet + Sortformer backends (WhisperX survives).
# `copy` trades a little disk and time for it actually working.
- UV_LINK_MODE=${SCRIBERR_UV_LINK_MODE:-copy}
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids: ["${SCRIBERR_GPU_ID:-1}"]
capabilities: [gpu]
healthcheck:
# 127.0.0.1 rather than localhost — the IPv6-first resolution trap has
# bitten news-digest and chatterbox in this fleet before.
# start_period is generous: first boot builds a Python env and pulls
# several GB of model weights before the port answers.
test: ["CMD-SHELL", "curl -fsS http://127.0.0.1:8080/ >/dev/null || exit 1"]
interval: 30s
timeout: 5s
retries: 3
start_period: 600s
networks:
- tnet
labels:
# ⚠ The group name MUST match a key in the dashboard's settings.yaml
# `layout:` block. A group that appears nowhere in that block gets no
# `tab:` assignment, and Homepage renders an untabbed group on EVERY tab.
# This label read `AI Systems` — a group that existed nowhere — from
# 2026-08-23 until it was caught on 2026-08-24.
# `AI - Studios` and not one of the ASR groups because Scriberr is a
# transcription UI you open and work in, which is what Studios collects;
# the bare ASR endpoints (Parakeet, Speaches) live in the collapsed
# `AI - Audio Tools` group instead.
- homepage.group=AI - Studios
- homepage.name=Scriberr
- homepage.icon=mdi-microphone-message
- homepage.description=Audio/video transcription + diarization (ana-ml2, GPU1)
- homepage.href=http://10.250.50.54:${SCRIBERR_PORT}
networks:
tnet:
name: traefik-net
external: true