From 0176ec0a5b227122ec680bdb228660439e0f6c48 Mon Sep 17 00:00:00 2001 From: Vuong Hoang Date: Wed, 30 Sep 2026 09:02:45 -0700 Subject: [PATCH] =?UTF-8?q?scriberr:=20fit=20GPU=201=20beside=20intern-dec?= =?UTF-8?q?ision=20=E2=80=94=20120=20s=20Parakeet=20slices=20+=20expandabl?= =?UTF-8?q?e=5Fsegments?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Measured Parakeet peak on a 35-min file (n=3 each, deterministic): 300 s 9,384 MiB, 120 s 6,510, 60 s 5,976, 10 s 5,634 (fixed floor); with expandable_segments 120 s 5,496 and 60 s 5,502. Verified 5,496 under the recreated container's own env. Transcripts: 95.8% word-sequence similarity vs 300 s, diffs mostly casing/punctuation. --- stacks/scriberr/.env.example | 6 ++++++ stacks/scriberr/compose.yaml | 17 +++++++++++++++++ 2 files changed, 23 insertions(+) diff --git a/stacks/scriberr/.env.example b/stacks/scriberr/.env.example index 43490d1..5e24d73 100644 --- a/stacks/scriberr/.env.example +++ b/stacks/scriberr/.env.example @@ -42,3 +42,9 @@ SCRIBERR_SECURE_COOKIES=false # ⚠ That key also reaches PAID passthrough models (GLM, Kimi) on a shared # tab — keep the configured model on a free local seat. # SCRIBERR_OPENAI_API_KEY= + +# GPU 1 memory budget (2026-09-30). Parakeet slice length in seconds and the +# torch allocator mode; see compose.yaml for the measurements. Defaults apply +# when unset; override only with a re-measured peak. +# SCRIBERR_PARAKEET_CHUNK_SECS=120 +# SCRIBERR_PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True diff --git a/stacks/scriberr/compose.yaml b/stacks/scriberr/compose.yaml index e46809b..1b85341 100644 --- a/stacks/scriberr/compose.yaml +++ b/stacks/scriberr/compose.yaml @@ -70,6 +70,23 @@ services: # and takes out the Parakeet + Sortformer backends (WhisperX survives). # `copy` trades a little disk and time for it actually working. - UV_LINK_MODE=${SCRIBERR_UV_LINK_MODE:-copy} + # ── GPU 1 memory budget (2026-09-30, Prime: Scriberr shares GPU 1 with + # intern-decision, which holds ~9.7 GB resting / 10.3 GB peak). ────────── + # Parakeet's buffered path cuts audio into slices of this many seconds + # (Scriberr reads it in parakeet_adapter.go for BOTH the "is this long + # audio" threshold and --chunk-len; upstream default 300). Measured peak + # GPU memory on a 35-min file, n=3 each, deterministic: + # 300 s 9,384 MiB · 120 s 6,510 · 60 s 5,976 · 10 s 5,634 (fixed floor) + # 120 s + expandable_segments 5,496 · 60 s + expandable_segments 5,502 + # The floor, not the slice, dominates below ~120 s; expandable_segments is + # what removes the fragmentation on top of it. 120 s + expandable fits + # beside intern-decision even at both peaks (295 MiB spare). Transcripts + # change slightly: 95.8% word-sequence similarity vs 300 s (7,599 vs + # 7,645 words); the diffs are mostly casing/punctuation spread through + # the file, ~40 words at the 17 cuts. A-vs-A at 300 s: identical. + # Scriberr passes os.Environ() to the uv subprocess, so both reach NeMo. + - PARAKEET_CHUNK_THRESHOLD_SECS=${SCRIBERR_PARAKEET_CHUNK_SECS:-120} + - PYTORCH_CUDA_ALLOC_CONF=${SCRIBERR_PYTORCH_CUDA_ALLOC_CONF:-expandable_segments:True} deploy: resources: reservations: