scriberr to fv-ml1 GPU 3 (on-demand, steps aside to irv-ml1 A6000); intern-decision 32k-token calls (cap 14.4 GiB)

Prime 2026-09-30: move scriberr to GPU 3 and extend the Jev endpoint to 32k tokens.
Scriberr holds 0 VRAM idle; verified a 20-min job on GPU 3 at 5,496 MiB. With GPU 1
freed, intern-decision's measured card peak at MAX_TOKENS=32768 is 15,220 MiB against
a 15,437 MiB budget (n=3, 1 and 16 questions); 32,769 tokens is refused 422 up front.
JevBench v1.2.16 via /v1/systemone unchanged: 202/231, 0 diffs vs the bench.
This commit is contained in:
vh
2026-09-30 13:35:32 -07:00
parent 92501a29c1
commit 6b201e1d4a
8 changed files with 79 additions and 29 deletions
+4 -3
View File
@@ -19,8 +19,9 @@ SCRIBERR_BIND=0.0.0.0
SCRIBERR_ALLOWED_ORIGINS=http://10.251.50.54:8080,http://scriberr.fv.internal:8080
# ── GPU ──────────────────────────────────────────────────────────────────
# GPU0 is fully committed to the `gen` seat; GPU1 is the one with headroom.
SCRIBERR_GPU_ID=1
# GPU 3 since 2026-09-30 (Prime): an on-demand tenant of the full-size-seat reserve; it steps aside
# (to irv-ml1's A6000) when a full-size seat claims GPU 3. See compose.yaml.
SCRIBERR_GPU_ID=3
# ── Storage (on /tank — NOT the root pool, weights are multi-GB) ─────────
SCRIBERR_DATA_DIR=/tank/scriberr/data
@@ -48,7 +49,7 @@ SCRIBERR_SECURE_COOKIES=false
# tab — keep the configured model on a free local seat.
# SCRIBERR_OPENAI_API_KEY=
# GPU 1 memory budget (2026-09-30). Parakeet slice length in seconds and the
# Memory settings measured on GPU 1 (2026-09-30; still in force on GPU 3). Parakeet slice length in seconds and the
# torch allocator mode; see compose.yaml for the measurements. Defaults apply
# when unset; override only with a re-measured peak.
# SCRIBERR_PARAKEET_CHUNK_SECS=120
+4
View File
@@ -142,6 +142,10 @@ tab. Keep the configured model on a free local seat.
## Parakeet memory and slicing (measured 2026-09-30)
> **Moved to fv-ml1 GPU 3 at 1322 on 2026-09-30 (Prime).** The GPU 1 budget below no longer binds. Scriberr is
> an on-demand tenant of GPU 3's reserve: it steps aside to irv-ml1's A6000 when a full-size seat claims GPU 3.
> A 20-min file was verified on GPU 3 at a 5,496 MiB peak. The settings below were not changed by the move.
Scriberr shares fv-ml1 GPU 1 with intern-decision (9.0 GiB cap, 9,876 MiB card peak).
GPU 1's nvidia-smi Free is 15,442 MiB, so the budget left for Scriberr is about 5.5 GB
(70 MiB spare at both peaks), so the compose file sets
+10 -7
View File
@@ -18,11 +18,13 @@
# /tank/scriberr/src/Scriberr on fv-ml1.
#
# ── GPU PINNING ───────────────────────────────────────────────────────────
# Pinned to **GPU1** via explicit device_ids, per the house convention and
# because GPU0 is fully committed to the `gen` seat. GPU1 shares space with
# the `sec` seat, so this stack is a guest there — keep an eye on VRAM.
# Pinned to **GPU 3** (Prime, 2026-09-30 1322) via explicit device_ids. Scriberr holds 0 VRAM
# when idle, so it is an ON-DEMAND tenant of GPU 3's full-size-seat reserve, like Blender: when a
# full-size seat (Flash-Next) claims GPU 3, Scriberr STEPS ASIDE. Its planned landing spot then is
# irv-ml1's A6000, not GPU 1 (GPU 1's headroom now funds intern-decision's 32k-token calls).
# (It was on GPU 1 until 2026-09-30, beside the vLLM seats; history in stacks/scriberr/README.md.)
# NOTE: do NOT add `NVIDIA_VISIBLE_DEVICES=all` (as upstream's compose does).
# It overrides the device_ids reservation and exposes both cards.
# It overrides the device_ids reservation and exposes every card.
#
# All tunables live in .env — edit that, not this file.
@@ -70,7 +72,8 @@ services:
# and takes out the Parakeet + Sortformer backends (WhisperX survives).
# `copy` trades a little disk and time for it actually working.
- UV_LINK_MODE=${SCRIBERR_UV_LINK_MODE:-copy}
# ── GPU 1 memory budget (2026-09-30, Prime: Scriberr shares GPU 1 with
# ── Memory settings, measured while Scriberr shared GPU 1 (2026-09-30; it moved to
# GPU 3 at 1322 the same day, so the 5.5 GB budget no longer binds, but the values stand). Prime: Scriberr shared GPU 1 with
# intern-decision, which holds ~9.7 GB resting / 10.3 GB peak). ──────────
# Parakeet's buffered path cuts audio into slices of this many seconds
# (Scriberr reads it in parakeet_adapter.go for BOTH the "is this long
@@ -94,7 +97,7 @@ services:
reservations:
devices:
- driver: nvidia
device_ids: ["${SCRIBERR_GPU_ID:-1}"]
device_ids: ["${SCRIBERR_GPU_ID:-3}"]
capabilities: [gpu]
healthcheck:
# 127.0.0.1 rather than localhost — the IPv6-first resolution trap has
@@ -121,7 +124,7 @@ services:
- homepage.group=AI - Studios
- homepage.name=Scriberr
- homepage.icon=mdi-microphone-message
- homepage.description=Audio/video transcription + diarization (fv-ml1, GPU1)
- homepage.description=Audio/video transcription + diarization (fv-ml1, GPU3)
- homepage.href=http://10.251.50.54:${SCRIBERR_PORT}
networks: