Files
esh-pfi-infrastructure/stacks/scriberr/README.md
T
vh 6b201e1d4a scriberr to fv-ml1 GPU 3 (on-demand, steps aside to irv-ml1 A6000); intern-decision 32k-token calls (cap 14.4 GiB)
Prime 2026-09-30: move scriberr to GPU 3 and extend the Jev endpoint to 32k tokens.
Scriberr holds 0 VRAM idle; verified a 20-min job on GPU 3 at 5,496 MiB. With GPU 1
freed, intern-decision's measured card peak at MAX_TOKENS=32768 is 15,220 MiB against
a 15,437 MiB budget (n=3, 1 and 16 questions); 32,769 tokens is refused 422 up front.
JevBench v1.2.16 via /v1/systemone unchanged: 202/231, 0 diffs vs the bench.
2026-09-30 13:35:32 -07:00

185 lines
8.4 KiB
Markdown

# scriberr — self-hosted transcription + diarization (fv-ml1, GPU1)
Web UI for transcribing audio/video locally. WhisperX (Whisper + pyannote
speaker diarization) with NVIDIA Parakeet/Canary also selectable; SQLite for
state; optional summarisation and transcript chat against any OpenAI-compatible
endpoint.
- **Host:** `fv-ml1` (10.251.50.54) — GPU1
- **URL:** http://10.251.50.54:8080
- **Upstream:** https://github.com/rishikanthc/Scriberr
## The image is built locally, and that is not incidental
fv-ml1's RTX PRO 6000 Blackwell cards are **sm_120**. Upstream's published
images do not cover that:
| image | built for | usable here |
|---|---|---|
| `ghcr.io/rishikanthc/scriberr` | CPU | yes, but no GPU |
| `ghcr.io/rishikanthc/scriberr-cuda` | sm_61 … sm_89 (Pascal→Ada) | **no** — no sm_120 kernels |
| `ghcr.io/rishikanthc/scriberr-cuda-blackwell` | sm_120 | **does not exist** — documented in the upstream README but never published; GHCR returns no tags (checked 2026-08-23) |
The sm_120 path upstream actually ships is `Dockerfile.cuda.12.9`
(CUDA 12.9.1 + cuDNN, `PYTORCH_CUDA_VERSION=cu128`), built from source. So we
build it. **Do not "simplify" the compose back to the published `scriberr-cuda`
image** — it will fail on these cards or quietly fall back to CPU.
### Rebuilding
We carry local patches (`patches/`, currently the pause-aware Parakeet
slicer), so a rebuild is one command from nh3-dev, pinned to an upstream sha:
```bash
scripts/scriberr-rebuild --sha <full upstream sha> --suffix slicer1
```
It makes a clean clone in `/opt/docker/src/scriberr-<sha7>-<suffix>` on
fv-ml1, `git apply --check`s the patches (a conflict stops it), builds
`scriberr:local-blackwell-<sha7>-<suffix>` beside the old images, and checks
the embed, the unit tests, the Go↔Python JSON seam, and the GPU memory budget.
Deploying it is a separate manual step: `patches/README.md` § Deploy.
The old checkout at `/tank/scriberr/src/Scriberr` (lkraven-owned, shallow)
built the original `scriberr:local-blackwell` and is left as it was.
## Deploy
```bash
# from this workstation
scripts/deploy-stack.sh fv-ml1 scriberr
```
Then on the host, the usual:
```bash
cd /opt/docker/compose/scriberr
docker compose config # dry parse first
docker compose up -d scriberr # target the service, not the whole stack
```
## Storage — deliberately on /tank
`/var/lib/docker` on fv-ml1 sits on `zroot` at ~87% used. Whisper, pyannote
and NeMo weights are multi-GB and land in the `whisperx-env` volume, so both
mounts are bind-mounted onto `/tank` (4+ TB) instead of named volumes:
| host path | container path | holds |
|---|---|---|
| `/tank/scriberr/data` | `/app/data` | SQLite DB, uploads, transcripts |
| `/tank/scriberr/whisperx-env` | `/app/whisperx-env` | Python env + model weights |
| `/tank/scriberr/src/Scriberr` | — | original build checkout (patched builds: `/opt/docker/src/scriberr-<sha7>-<suffix>`) |
Both are owned by uid/gid 1000 to match `PUID`/`PGID`.
## First run takes a while
On first start the container builds a Python environment and downloads several
GB of model weights before the port answers — upstream says "several minutes".
The healthcheck therefore has a **600 s `start_period`**; the container will
show `starting`, not `unhealthy`, during that window. Watch it with:
```bash
docker logs -f scriberr
```
Subsequent starts are fast because the env volume persists.
## The PUID trap — read this before "fixing" the uid
This stack runs as **uid/gid 10001**, not the fleet-usual 1000, and the
`/tank/scriberr` dirs are chowned to match. That is deliberate.
`Dockerfile.cuda.12.9` creates `appuser` at **uid 10001** — Ubuntu 24.04's base
image already owns uid 1000 as `ubuntu`, so upstream moved their app user out of
the way. It then `chown`s `/app` to 10001. But the entrypoint's `PUID` remapping
only chowns `/app/data` and `/app/whisperx-env` — **not `/app` itself**. So
running with `PUID=1000` leaves the app unable to open its SQLite database and
it crash-loops with:
```
Failed to connect to database: unable to open database file: out of memory (14)
```
That message is a red herring twice over: error 14 is `SQLITE_CANTOPEN`, not an
OOM, and the machine has 566 GB of RAM. Diagnosis notes from 2026-08-23:
- SQLite itself writes fine to `/tank` as uid 1000 — the mount is not at fault.
- The app fails on a plain Docker **named volume** too — storage is not at fault.
- The **published CPU image runs fine at `PUID=1000`**, because in `Dockerfile`
(the non-CUDA one) `appuser` *is* uid 1000. Only the CUDA 12.9 variant moved it.
- Same image at `PUID=10001` starts clean. That is the whole difference.
If you ever want host files owned by 1000 instead, the fix is to patch
`Dockerfile.cuda.12.9` to `userdel ubuntu` and recreate `appuser` at 1000, then
rebuild — a local patch to carry, which is why it was not done.
## Gotchas
- **`SECURE_COOKIES` must stay `false` while served over plain HTTP.** At the
production default of `true` the session cookie is marked `Secure`, the
browser drops it, and login appears to succeed then bounces you straight back
to the login page with nothing useful in the logs.
- **`ALLOWED_ORIGINS` must list the real origin.** Upstream defaults to
`localhost` only; reaching the UI by host IP fails CORS until it is set.
- **Never add `NVIDIA_VISIBLE_DEVICES=all`.** Upstream's compose sets it, but
here it would override the `device_ids` reservation and expose both cards —
GPU0 belongs to the `gen` seat.
- **This stack is a guest on GPU1**, which it shares with the `sec` seat. If
VRAM gets tight, this is the thing that should yield.
## Optional: summarisation via the LiteLLM gateway
Scriberr speaks the OpenAI API, so point it at the fleet gateway instead of a
paid vendor. In the UI under the AI provider settings:
- base URL: `http://10.250.50.70:4000/v1`
- model: `summarizer` (or `gen` / `gen-reasoning`)
- key: the shared all-agents gateway key
⚠ That key also reaches **paid** passthrough models (GLM, Kimi) on a shared
tab. Keep the configured model on a free local seat.
## Parakeet memory and slicing (measured 2026-09-30)
> **Moved to fv-ml1 GPU 3 at 1322 on 2026-09-30 (Prime).** The GPU 1 budget below no longer binds. Scriberr is
> an on-demand tenant of GPU 3's reserve: it steps aside to irv-ml1's A6000 when a full-size seat claims GPU 3.
> A 20-min file was verified on GPU 3 at a 5,496 MiB peak. The settings below were not changed by the move.
Scriberr shares fv-ml1 GPU 1 with intern-decision (9.0 GiB cap, 9,876 MiB card peak).
GPU 1's nvidia-smi Free is 15,442 MiB, so the budget left for Scriberr is about 5.5 GB
(70 MiB spare at both peaks), so the compose file sets
`PARAKEET_CHUNK_THRESHOLD_SECS=120` and
`PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True`. The numbers behind those
settings are in the compose comments.
Peak GPU memory on a 35-minute file:
| setting | peak |
|---|---|
| 300 s slices | 9,384 MiB |
| 300 s slices + expandable_segments | 6,962 MiB |
| 120 s slices + expandable_segments | 5,496 MiB (live setting) |
- **Whole file in one pass with local attention** (`--context-left/right 255`,
standard script): **CUDA OOM at >16 GB**. It asked for another 6.46 GiB at
9.8 GiB in use. It is not viable on this card budget.
- ⚠ **Hand-editing the Parakeet scripts does not persist.** Scriberr's
`PrepareEnvironment` rewrites `parakeet_transcribe.py` and
`parakeet_transcribe_buffered.py` into `whisperx-env/parakeet/` from the copies
embedded in the Go binary every time it prepares the environment. A change to the
slicer has to go into the source checkout
(`internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py`,
embedded at build) and be rebuilt. Or it goes upstream (MIT).
- The two env knobs are read by upstream's Go code (`parakeet_adapter.go`), so
they survive image upgrades for as long as upstream keeps them. Re-measure the
peak after any upgrade.
- **The slicer itself is patched** (`patches/0001-parakeet-pause-aware-slicer.patch`,
2026-09-30). Adjacent 120 s slices now overlap by 4 s and are stitched at a word
both transcribed, which cut the share of cuts with an error nearby from 52 % to
22 % against a 19 % background. The overlap sits *inside* the 120 s, so the
peak is unchanged (5,496 MiB, n=3). See `patches/README.md` and
`docs/pfi/scriberr-slicer-bench-2026-09-30.md`. That bench also found that
Parakeet sometimes skips stretches of ≥10 words mid-slice, with or without the
patch; that is still open.