167 lines
7.0 KiB
Markdown
167 lines
7.0 KiB
Markdown
# scriberr — self-hosted transcription + diarization (fv-ml1, GPU1)
|
|
|
|
Web UI for transcribing audio/video locally. WhisperX (Whisper + pyannote
|
|
speaker diarization) with NVIDIA Parakeet/Canary also selectable; SQLite for
|
|
state; optional summarisation and transcript chat against any OpenAI-compatible
|
|
endpoint.
|
|
|
|
- **Host:** `fv-ml1` (10.251.50.54) — GPU1
|
|
- **URL:** http://10.251.50.54:8080
|
|
- **Upstream:** https://github.com/rishikanthc/Scriberr
|
|
|
|
## The image is built locally, and that is not incidental
|
|
|
|
fv-ml1's RTX PRO 6000 Blackwell cards are **sm_120**. Upstream's published
|
|
images do not cover that:
|
|
|
|
| image | built for | usable here |
|
|
|---|---|---|
|
|
| `ghcr.io/rishikanthc/scriberr` | CPU | yes, but no GPU |
|
|
| `ghcr.io/rishikanthc/scriberr-cuda` | sm_61 … sm_89 (Pascal→Ada) | **no** — no sm_120 kernels |
|
|
| `ghcr.io/rishikanthc/scriberr-cuda-blackwell` | sm_120 | **does not exist** — documented in the upstream README but never published; GHCR returns no tags (checked 2026-08-23) |
|
|
|
|
The sm_120 path upstream actually ships is `Dockerfile.cuda.12.9`
|
|
(CUDA 12.9.1 + cuDNN, `PYTORCH_CUDA_VERSION=cu128`), built from source. So we
|
|
build it. **Do not "simplify" the compose back to the published `scriberr-cuda`
|
|
image** — it will fail on these cards or quietly fall back to CPU.
|
|
|
|
### Rebuilding
|
|
|
|
```bash
|
|
ssh fv-ml1
|
|
cd /tank/scriberr/src/Scriberr
|
|
git pull
|
|
docker build -f Dockerfile.cuda.12.9 -t scriberr:local-blackwell .
|
|
cd /opt/docker/compose/scriberr && docker compose up -d
|
|
```
|
|
|
|
Source checkout lives on `/tank`, not the root pool — see storage below.
|
|
|
|
## Deploy
|
|
|
|
```bash
|
|
# from this workstation
|
|
scripts/deploy-stack.sh fv-ml1 scriberr
|
|
```
|
|
|
|
Then on the host, the usual:
|
|
|
|
```bash
|
|
cd /opt/docker/compose/scriberr
|
|
docker compose config # dry parse first
|
|
docker compose up -d scriberr # target the service, not the whole stack
|
|
```
|
|
|
|
## Storage — deliberately on /tank
|
|
|
|
`/var/lib/docker` on fv-ml1 sits on `zroot` at ~87% used. Whisper, pyannote
|
|
and NeMo weights are multi-GB and land in the `whisperx-env` volume, so both
|
|
mounts are bind-mounted onto `/tank` (4+ TB) instead of named volumes:
|
|
|
|
| host path | container path | holds |
|
|
|---|---|---|
|
|
| `/tank/scriberr/data` | `/app/data` | SQLite DB, uploads, transcripts |
|
|
| `/tank/scriberr/whisperx-env` | `/app/whisperx-env` | Python env + model weights |
|
|
| `/tank/scriberr/src/Scriberr` | — | build checkout |
|
|
|
|
Both are owned by uid/gid 1000 to match `PUID`/`PGID`.
|
|
|
|
## First run takes a while
|
|
|
|
On first start the container builds a Python environment and downloads several
|
|
GB of model weights before the port answers — upstream says "several minutes".
|
|
The healthcheck therefore has a **600 s `start_period`**; the container will
|
|
show `starting`, not `unhealthy`, during that window. Watch it with:
|
|
|
|
```bash
|
|
docker logs -f scriberr
|
|
```
|
|
|
|
Subsequent starts are fast because the env volume persists.
|
|
|
|
## The PUID trap — read this before "fixing" the uid
|
|
|
|
This stack runs as **uid/gid 10001**, not the fleet-usual 1000, and the
|
|
`/tank/scriberr` dirs are chowned to match. That is deliberate.
|
|
|
|
`Dockerfile.cuda.12.9` creates `appuser` at **uid 10001** — Ubuntu 24.04's base
|
|
image already owns uid 1000 as `ubuntu`, so upstream moved their app user out of
|
|
the way. It then `chown`s `/app` to 10001. But the entrypoint's `PUID` remapping
|
|
only chowns `/app/data` and `/app/whisperx-env` — **not `/app` itself**. So
|
|
running with `PUID=1000` leaves the app unable to open its SQLite database and
|
|
it crash-loops with:
|
|
|
|
```
|
|
Failed to connect to database: unable to open database file: out of memory (14)
|
|
```
|
|
|
|
That message is a red herring twice over: error 14 is `SQLITE_CANTOPEN`, not an
|
|
OOM, and the machine has 566 GB of RAM. Diagnosis notes from 2026-08-23:
|
|
|
|
- SQLite itself writes fine to `/tank` as uid 1000 — the mount is not at fault.
|
|
- The app fails on a plain Docker **named volume** too — storage is not at fault.
|
|
- The **published CPU image runs fine at `PUID=1000`**, because in `Dockerfile`
|
|
(the non-CUDA one) `appuser` *is* uid 1000. Only the CUDA 12.9 variant moved it.
|
|
- Same image at `PUID=10001` starts clean. That is the whole difference.
|
|
|
|
If you ever want host files owned by 1000 instead, the fix is to patch
|
|
`Dockerfile.cuda.12.9` to `userdel ubuntu` and recreate `appuser` at 1000, then
|
|
rebuild — a local patch to carry, which is why it was not done.
|
|
|
|
## Gotchas
|
|
|
|
- **`SECURE_COOKIES` must stay `false` while served over plain HTTP.** At the
|
|
production default of `true` the session cookie is marked `Secure`, the
|
|
browser drops it, and login appears to succeed then bounces you straight back
|
|
to the login page with nothing useful in the logs.
|
|
- **`ALLOWED_ORIGINS` must list the real origin.** Upstream defaults to
|
|
`localhost` only; reaching the UI by host IP fails CORS until it is set.
|
|
- **Never add `NVIDIA_VISIBLE_DEVICES=all`.** Upstream's compose sets it, but
|
|
here it would override the `device_ids` reservation and expose both cards —
|
|
GPU0 belongs to the `gen` seat.
|
|
- **This stack is a guest on GPU1**, which it shares with the `sec` seat. If
|
|
VRAM gets tight, this is the thing that should yield.
|
|
|
|
## Optional: summarisation via the LiteLLM gateway
|
|
|
|
Scriberr speaks the OpenAI API, so point it at the fleet gateway instead of a
|
|
paid vendor. In the UI under the AI provider settings:
|
|
|
|
- base URL: `http://10.250.50.70:4000/v1`
|
|
- model: `summarizer` (or `gen` / `gen-reasoning`)
|
|
- key: the shared all-agents gateway key
|
|
|
|
⚠ That key also reaches **paid** passthrough models (GLM, Kimi) on a shared
|
|
tab. Keep the configured model on a free local seat.
|
|
|
|
## Parakeet memory and slicing (measured 2026-09-30)
|
|
|
|
Scriberr shares fv-ml1 GPU 1 with intern-decision (9.0 GiB cap, 9,876 MiB card peak).
|
|
GPU 1's nvidia-smi Free is 15,442 MiB, so the budget left for Scriberr is about 5.5 GB
|
|
(70 MiB spare at both peaks), so the compose file sets
|
|
`PARAKEET_CHUNK_THRESHOLD_SECS=120` and
|
|
`PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True`. The numbers behind those
|
|
settings are in the compose comments.
|
|
|
|
Peak GPU memory on a 35-minute file:
|
|
|
|
| setting | peak |
|
|
|---|---|
|
|
| 300 s slices | 9,384 MiB |
|
|
| 300 s slices + expandable_segments | 6,962 MiB |
|
|
| 120 s slices + expandable_segments | 5,496 MiB (live setting) |
|
|
|
|
- **Whole file in one pass with local attention** (`--context-left/right 255`,
|
|
standard script): **CUDA OOM at >16 GB**. It asked for another 6.46 GiB at
|
|
9.8 GiB in use. It is not viable on this card budget.
|
|
- ⚠ **Hand-editing the Parakeet scripts does not persist.** Scriberr's
|
|
`PrepareEnvironment` rewrites `parakeet_transcribe.py` and
|
|
`parakeet_transcribe_buffered.py` into `whisperx-env/parakeet/` from the copies
|
|
embedded in the Go binary every time it prepares the environment. A change to the
|
|
slicer has to go into the source checkout
|
|
(`internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py`,
|
|
embedded at build) and be rebuilt. Or it goes upstream (MIT).
|
|
- The two env knobs are read by upstream's Go code (`parakeet_adapter.go`), so
|
|
they survive image upgrades for as long as upstream keeps them. Re-measure the
|
|
peak after any upgrade.
|