Files
esh-pfi-infrastructure/stacks/scriberr/README.md
T
vh 38015a1977 docs(scriberr): Parakeet dropout investigation; proposed 0002 (gap retry + model path)
Prime's ask (via the coordinator): investigate the "Parakeet skips
stretches of speech" finding, including other Parakeet weights.
Investigation only; nothing deployed.

Against ground truth (official SCOTUS transcript, Gutenberg #38916) the
drops are real: production v3 loses 140 / 66 clean words per transcript on
the two public files and ~50 on each private one (Whisper-referenced,
Canary-confirmed; adjudicator 129/129 correct on the calibration). Cause:
the v2/v3 0.6B weights collapse deep inside long full-attention windows;
the encoder output is degraded, the audio alone transcribes fine, and
1.1B TDT/RNNT/CTC and CTC-0.6B never do it. Decoding (CUDA graphs, greedy
variants, max_symbols, beam), slice length, local attention, loudness,
resampling and a noise floor do not fix it. Controls: A-vs-A, silence
positive control (>=15 words 36/36), null control, bootstrap floor.

Proposed patch 0002 re-transcribes >=3 s stretches where the audio holds
speech but no word came out (-80 to -90 % lost words on all four
recordings, lower WER, no invented text, +10 MiB) and adds an explicit
PARAKEET_MODEL_PATH with the loaded model recorded in JSON and ModelUsed.
Reviewed at high effort, all findings fixed; built and tested as
scriberr:local-blackwell-a353078-dropout2, not deployed.

scriberr-rebuild: --patches takes DIR[:DIR...]; embeds and seam-checks
both Parakeet scripts (seam-check --standard for the short-audio one).
2026-09-30 15:48:02 -07:00

188 lines
8.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# scriberr — self-hosted transcription + diarization (fv-ml1, GPU1)
Web UI for transcribing audio/video locally. WhisperX (Whisper + pyannote
speaker diarization) with NVIDIA Parakeet/Canary also selectable; SQLite for
state; optional summarisation and transcript chat against any OpenAI-compatible
endpoint.
- **Host:** `fv-ml1` (10.251.50.54) — GPU1
- **URL:** http://10.251.50.54:8080
- **Upstream:** https://github.com/rishikanthc/Scriberr
## The image is built locally, and that is not incidental
fv-ml1's RTX PRO 6000 Blackwell cards are **sm_120**. Upstream's published
images do not cover that:
| image | built for | usable here |
|---|---|---|
| `ghcr.io/rishikanthc/scriberr` | CPU | yes, but no GPU |
| `ghcr.io/rishikanthc/scriberr-cuda` | sm_61 … sm_89 (Pascal→Ada) | **no** — no sm_120 kernels |
| `ghcr.io/rishikanthc/scriberr-cuda-blackwell` | sm_120 | **does not exist** — documented in the upstream README but never published; GHCR returns no tags (checked 2026-08-23) |
The sm_120 path upstream actually ships is `Dockerfile.cuda.12.9`
(CUDA 12.9.1 + cuDNN, `PYTORCH_CUDA_VERSION=cu128`), built from source. So we
build it. **Do not "simplify" the compose back to the published `scriberr-cuda`
image** — it will fail on these cards or quietly fall back to CPU.
### Rebuilding
We carry local patches (`patches/`, currently the pause-aware Parakeet
slicer), so a rebuild is one command from nh3-dev, pinned to an upstream sha:
```bash
scripts/scriberr-rebuild --sha <full upstream sha> --suffix slicer1
```
It makes a clean clone in `/opt/docker/src/scriberr-<sha7>-<suffix>` on
fv-ml1, `git apply --check`s the patches (a conflict stops it), builds
`scriberr:local-blackwell-<sha7>-<suffix>` beside the old images, and checks
the embed, the unit tests, the Go↔Python JSON seam, and the GPU memory budget.
Deploying it is a separate manual step: `patches/README.md` § Deploy.
The old checkout at `/tank/scriberr/src/Scriberr` (lkraven-owned, shallow)
built the original `scriberr:local-blackwell` and is left as it was.
## Deploy
```bash
# from this workstation
scripts/deploy-stack.sh fv-ml1 scriberr
```
Then on the host, the usual:
```bash
cd /opt/docker/compose/scriberr
docker compose config # dry parse first
docker compose up -d scriberr # target the service, not the whole stack
```
## Storage — deliberately on /tank
`/var/lib/docker` on fv-ml1 sits on `zroot` at ~87% used. Whisper, pyannote
and NeMo weights are multi-GB and land in the `whisperx-env` volume, so both
mounts are bind-mounted onto `/tank` (4+ TB) instead of named volumes:
| host path | container path | holds |
|---|---|---|
| `/tank/scriberr/data` | `/app/data` | SQLite DB, uploads, transcripts |
| `/tank/scriberr/whisperx-env` | `/app/whisperx-env` | Python env + model weights |
| `/tank/scriberr/src/Scriberr` | — | original build checkout (patched builds: `/opt/docker/src/scriberr-<sha7>-<suffix>`) |
Both are owned by uid/gid 1000 to match `PUID`/`PGID`.
## First run takes a while
On first start the container builds a Python environment and downloads several
GB of model weights before the port answers — upstream says "several minutes".
The healthcheck therefore has a **600 s `start_period`**; the container will
show `starting`, not `unhealthy`, during that window. Watch it with:
```bash
docker logs -f scriberr
```
Subsequent starts are fast because the env volume persists.
## The PUID trap — read this before "fixing" the uid
This stack runs as **uid/gid 10001**, not the fleet-usual 1000, and the
`/tank/scriberr` dirs are chowned to match. That is deliberate.
`Dockerfile.cuda.12.9` creates `appuser` at **uid 10001** — Ubuntu 24.04's base
image already owns uid 1000 as `ubuntu`, so upstream moved their app user out of
the way. It then `chown`s `/app` to 10001. But the entrypoint's `PUID` remapping
only chowns `/app/data` and `/app/whisperx-env` — **not `/app` itself**. So
running with `PUID=1000` leaves the app unable to open its SQLite database and
it crash-loops with:
```
Failed to connect to database: unable to open database file: out of memory (14)
```
That message is a red herring twice over: error 14 is `SQLITE_CANTOPEN`, not an
OOM, and the machine has 566 GB of RAM. Diagnosis notes from 2026-08-23:
- SQLite itself writes fine to `/tank` as uid 1000 — the mount is not at fault.
- The app fails on a plain Docker **named volume** too — storage is not at fault.
- The **published CPU image runs fine at `PUID=1000`**, because in `Dockerfile`
(the non-CUDA one) `appuser` *is* uid 1000. Only the CUDA 12.9 variant moved it.
- Same image at `PUID=10001` starts clean. That is the whole difference.
If you ever want host files owned by 1000 instead, the fix is to patch
`Dockerfile.cuda.12.9` to `userdel ubuntu` and recreate `appuser` at 1000, then
rebuild — a local patch to carry, which is why it was not done.
## Gotchas
- **`SECURE_COOKIES` must stay `false` while served over plain HTTP.** At the
production default of `true` the session cookie is marked `Secure`, the
browser drops it, and login appears to succeed then bounces you straight back
to the login page with nothing useful in the logs.
- **`ALLOWED_ORIGINS` must list the real origin.** Upstream defaults to
`localhost` only; reaching the UI by host IP fails CORS until it is set.
- **Never add `NVIDIA_VISIBLE_DEVICES=all`.** Upstream's compose sets it, but
here it would override the `device_ids` reservation and expose both cards —
GPU0 belongs to the `gen` seat.
- **This stack is a guest on GPU1**, which it shares with the `sec` seat. If
VRAM gets tight, this is the thing that should yield.
## Optional: summarisation via the LiteLLM gateway
Scriberr speaks the OpenAI API, so point it at the fleet gateway instead of a
paid vendor. In the UI under the AI provider settings:
- base URL: `http://10.250.50.70:4000/v1`
- model: `summarizer` (or `gen` / `gen-reasoning`)
- key: the shared all-agents gateway key
⚠ That key also reaches **paid** passthrough models (GLM, Kimi) on a shared
tab. Keep the configured model on a free local seat.
## Parakeet memory and slicing (measured 2026-09-30)
> **Moved to fv-ml1 GPU 3 at 1322 on 2026-09-30 (Prime).** The GPU 1 budget below no longer binds. Scriberr is
> an on-demand tenant of GPU 3's reserve: it steps aside to irv-ml1's A6000 when a full-size seat claims GPU 3.
> A 20-min file was verified on GPU 3 at a 5,496 MiB peak. The settings below were not changed by the move.
Scriberr shares fv-ml1 GPU 1 with intern-decision (9.0 GiB cap, 9,876 MiB card peak).
GPU 1's nvidia-smi Free is 15,442 MiB, so the budget left for Scriberr is about 5.5 GB
(70 MiB spare at both peaks), so the compose file sets
`PARAKEET_CHUNK_THRESHOLD_SECS=120` and
`PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True`. The numbers behind those
settings are in the compose comments.
Peak GPU memory on a 35-minute file:
| setting | peak |
|---|---|
| 300 s slices | 9,384 MiB |
| 300 s slices + expandable_segments | 6,962 MiB |
| 120 s slices + expandable_segments | 5,496 MiB (live setting) |
- **Whole file in one pass with local attention** (`--context-left/right 255`,
standard script): **CUDA OOM at >16 GB**. It asked for another 6.46 GiB at
9.8 GiB in use. It is not viable on this card budget.
- ⚠ **Hand-editing the Parakeet scripts does not persist.** Scriberr's
`PrepareEnvironment` rewrites `parakeet_transcribe.py` and
`parakeet_transcribe_buffered.py` into `whisperx-env/parakeet/` from the copies
embedded in the Go binary every time it prepares the environment. A change to the
slicer has to go into the source checkout
(`internal/transcription/adapters/py/nvidia/parakeet_transcribe_buffered.py`,
embedded at build) and be rebuilt. Or it goes upstream (MIT).
- The two env knobs are read by upstream's Go code (`parakeet_adapter.go`), so
they survive image upgrades for as long as upstream keeps them. Re-measure the
peak after any upgrade.
- **The slicer itself is patched** (`patches/0001-parakeet-pause-aware-slicer.patch`,
2026-09-30). Adjacent 120 s slices now overlap by 4 s and are stitched at a word
both transcribed, which cut the share of cuts with an error nearby from 52 % to
22 % against a 19 % background. The overlap sits *inside* the 120 s, so the
peak is unchanged (5,496 MiB, n=3). See `patches/README.md` and
`docs/pfi/scriberr-slicer-bench-2026-09-30.md`. That bench also found that
Parakeet sometimes skips stretches of ≥10 words mid-slice, with or without the
patch. **Investigated 2026-09-30** (`docs/pfi/parakeet-dropout-investigation-2026-09-30.md`):
the losses are real (against ground truth) and belong to the v2/v3 weights over
long windows; a proposed patch, `patches/proposed/0002`, re-transcribes speech that
got no words and cuts them 80–90 %. It is not deployed; that is Prime's call.