Scriberr transcribes audio and video locally with WhisperX and speaker diarization, and it lands on ana-ml2 rather than ana-docker because the work is GPU-shaped: ana-docker offers eight cores already shared with fifty containers and thirty-seven gigabytes of disk, against ninety-six cores, terabytes on /tank and idle capacity on GPU1. The reservation names device 1 explicitly, since GPU0 is fully committed to the gen seat, and the container is confirmed to see that card alone. The image is built from source, which is not a preference. These are Blackwell cards at sm_120; the published CUDA image covers Pascal through Ada only, and the blackwell image the upstream README documents has never been published at all. The path upstream actually ships for sm_120 is Dockerfile.cuda.12.9, carrying CUDA 12.9 and cu128 torch, so that is what gets built. The compose header says so, because the obvious cleanup is to swap in the published image and that would silently drop the deployment to CPU. Two configuration details are load-bearing and documented where someone would go to change them. The application runs as uid 10001 rather than the usual 1000: that Dockerfile moves its user aside for Ubuntu 24.04's own uid-1000 account and chowns /app accordingly, while the entrypoint's remapping covers only the data directories, so at 1000 the process cannot open its database and restarts forever behind a SQLite error that reads as though the machine were out of memory. Secure cookies stay off while the service is reached over plain HTTP, or sessions are dropped by the browser and login appears to loop for no visible reason. Storage is bind-mounted onto /tank because model weights run to several gigabytes and the root pool on that host is nearly full. Also adds the scriberr service alias to internal DNS, following the existing alias convention so consumers name the service rather than the box.
5.5 KiB
scriberr — self-hosted transcription + diarization (ana-ml2, GPU1)
Web UI for transcribing audio/video locally. WhisperX (Whisper + pyannote speaker diarization) with NVIDIA Parakeet/Canary also selectable; SQLite for state; optional summarisation and transcript chat against any OpenAI-compatible endpoint.
- Host:
ana-ml2(10.250.50.54) — GPU1 - URL: http://10.250.50.54:8080
- Upstream: https://github.com/rishikanthc/Scriberr
The image is built locally, and that is not incidental
ana-ml2's RTX PRO 6000 Blackwell cards are sm_120. Upstream's published images do not cover that:
| image | built for | usable here |
|---|---|---|
ghcr.io/rishikanthc/scriberr |
CPU | yes, but no GPU |
ghcr.io/rishikanthc/scriberr-cuda |
sm_61 … sm_89 (Pascal→Ada) | no — no sm_120 kernels |
ghcr.io/rishikanthc/scriberr-cuda-blackwell |
sm_120 | does not exist — documented in the upstream README but never published; GHCR returns no tags (checked 2026-08-23) |
The sm_120 path upstream actually ships is Dockerfile.cuda.12.9
(CUDA 12.9.1 + cuDNN, PYTORCH_CUDA_VERSION=cu128), built from source. So we
build it. Do not "simplify" the compose back to the published scriberr-cuda
image — it will fail on these cards or quietly fall back to CPU.
Rebuilding
ssh ana-ml2
cd /tank/scriberr/src/Scriberr
git pull
docker build -f Dockerfile.cuda.12.9 -t scriberr:local-blackwell .
cd /opt/docker/compose/scriberr && docker compose up -d
Source checkout lives on /tank, not the root pool — see storage below.
Deploy
# from this workstation
scripts/deploy-stack.sh ana-ml2 scriberr
Then on the host, the usual:
cd /opt/docker/compose/scriberr
docker compose config # dry parse first
docker compose up -d scriberr # target the service, not the whole stack
Storage — deliberately on /tank
/var/lib/docker on ana-ml2 sits on zroot at ~87% used. Whisper, pyannote
and NeMo weights are multi-GB and land in the whisperx-env volume, so both
mounts are bind-mounted onto /tank (4+ TB) instead of named volumes:
| host path | container path | holds |
|---|---|---|
/tank/scriberr/data |
/app/data |
SQLite DB, uploads, transcripts |
/tank/scriberr/whisperx-env |
/app/whisperx-env |
Python env + model weights |
/tank/scriberr/src/Scriberr |
— | build checkout |
Both are owned by uid/gid 1000 to match PUID/PGID.
First run takes a while
On first start the container builds a Python environment and downloads several
GB of model weights before the port answers — upstream says "several minutes".
The healthcheck therefore has a 600 s start_period; the container will
show starting, not unhealthy, during that window. Watch it with:
docker logs -f scriberr
Subsequent starts are fast because the env volume persists.
The PUID trap — read this before "fixing" the uid
This stack runs as uid/gid 10001, not the fleet-usual 1000, and the
/tank/scriberr dirs are chowned to match. That is deliberate.
Dockerfile.cuda.12.9 creates appuser at uid 10001 — Ubuntu 24.04's base
image already owns uid 1000 as ubuntu, so upstream moved their app user out of
the way. It then chowns /app to 10001. But the entrypoint's PUID remapping
only chowns /app/data and /app/whisperx-env — not /app itself. So
running with PUID=1000 leaves the app unable to open its SQLite database and
it crash-loops with:
Failed to connect to database: unable to open database file: out of memory (14)
That message is a red herring twice over: error 14 is SQLITE_CANTOPEN, not an
OOM, and the machine has 566 GB of RAM. Diagnosis notes from 2026-08-23:
- SQLite itself writes fine to
/tankas uid 1000 — the mount is not at fault. - The app fails on a plain Docker named volume too — storage is not at fault.
- The published CPU image runs fine at
PUID=1000, because inDockerfile(the non-CUDA one)appuseris uid 1000. Only the CUDA 12.9 variant moved it. - Same image at
PUID=10001starts clean. That is the whole difference.
If you ever want host files owned by 1000 instead, the fix is to patch
Dockerfile.cuda.12.9 to userdel ubuntu and recreate appuser at 1000, then
rebuild — a local patch to carry, which is why it was not done.
Gotchas
SECURE_COOKIESmust stayfalsewhile served over plain HTTP. At the production default oftruethe session cookie is markedSecure, the browser drops it, and login appears to succeed then bounces you straight back to the login page with nothing useful in the logs.ALLOWED_ORIGINSmust list the real origin. Upstream defaults tolocalhostonly; reaching the UI by host IP fails CORS until it is set.- Never add
NVIDIA_VISIBLE_DEVICES=all. Upstream's compose sets it, but here it would override thedevice_idsreservation and expose both cards — GPU0 belongs to thegenseat. - This stack is a guest on GPU1, which it shares with the
secseat. If VRAM gets tight, this is the thing that should yield.
Optional: summarisation via the LiteLLM gateway
Scriberr speaks the OpenAI API, so point it at the fleet gateway instead of a paid vendor. In the UI under the AI provider settings:
- base URL:
http://10.250.50.70:4000/v1 - model:
summarizer(orgen/gen-reasoning) - key: the shared all-agents gateway key
⚠ That key also reaches paid passthrough models (GLM, Kimi) on a shared tab. Keep the configured model on a free local seat.