scriberr: fit GPU 1 beside intern-decision — 120 s Parakeet slices + expandable_segments

Measured Parakeet peak on a 35-min file (n=3 each, deterministic): 300 s
9,384 MiB, 120 s 6,510, 60 s 5,976, 10 s 5,634 (fixed floor); with
expandable_segments 120 s 5,496 and 60 s 5,502. Verified 5,496 under the
recreated container's own env. Transcripts: 95.8% word-sequence similarity
vs 300 s, diffs mostly casing/punctuation.
This commit is contained in:
vh
2026-09-30 09:02:45 -07:00
parent d9bbaa07b2
commit 0176ec0a5b
2 changed files with 23 additions and 0 deletions
+6
View File
@@ -42,3 +42,9 @@ SCRIBERR_SECURE_COOKIES=false
# ⚠ That key also reaches PAID passthrough models (GLM, Kimi) on a shared
# tab — keep the configured model on a free local seat.
# SCRIBERR_OPENAI_API_KEY=
# GPU 1 memory budget (2026-09-30). Parakeet slice length in seconds and the
# torch allocator mode; see compose.yaml for the measurements. Defaults apply
# when unset; override only with a re-measured peak.
# SCRIBERR_PARAKEET_CHUNK_SECS=120
# SCRIBERR_PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
+17
View File
@@ -70,6 +70,23 @@ services:
# and takes out the Parakeet + Sortformer backends (WhisperX survives).
# `copy` trades a little disk and time for it actually working.
- UV_LINK_MODE=${SCRIBERR_UV_LINK_MODE:-copy}
# ── GPU 1 memory budget (2026-09-30, Prime: Scriberr shares GPU 1 with
# intern-decision, which holds ~9.7 GB resting / 10.3 GB peak). ──────────
# Parakeet's buffered path cuts audio into slices of this many seconds
# (Scriberr reads it in parakeet_adapter.go for BOTH the "is this long
# audio" threshold and --chunk-len; upstream default 300). Measured peak
# GPU memory on a 35-min file, n=3 each, deterministic:
# 300 s 9,384 MiB · 120 s 6,510 · 60 s 5,976 · 10 s 5,634 (fixed floor)
# 120 s + expandable_segments 5,496 · 60 s + expandable_segments 5,502
# The floor, not the slice, dominates below ~120 s; expandable_segments is
# what removes the fragmentation on top of it. 120 s + expandable fits
# beside intern-decision even at both peaks (295 MiB spare). Transcripts
# change slightly: 95.8% word-sequence similarity vs 300 s (7,599 vs
# 7,645 words); the diffs are mostly casing/punctuation spread through
# the file, ~40 words at the 17 cuts. A-vs-A at 300 s: identical.
# Scriberr passes os.Environ() to the uv subprocess, so both reach NeMo.
- PARAKEET_CHUNK_THRESHOLD_SECS=${SCRIBERR_PARAKEET_CHUNK_SECS:-120}
- PYTORCH_CUDA_ALLOC_CONF=${SCRIBERR_PYTORCH_CUDA_ALLOC_CONF:-expandable_segments:True}
deploy:
resources:
reservations: