scriberr to fv-ml1 GPU 3 (on-demand, steps aside to irv-ml1 A6000); intern-decision 32k-token calls (cap 14.4 GiB)
Prime 2026-09-30: move scriberr to GPU 3 and extend the Jev endpoint to 32k tokens. Scriberr holds 0 VRAM idle; verified a 20-min job on GPU 3 at 5,496 MiB. With GPU 1 freed, intern-decision's measured card peak at MAX_TOKENS=32768 is 15,220 MiB against a 15,437 MiB budget (n=3, 1 and 16 questions); 32,769 tokens is refused 422 up front. JevBench v1.2.16 via /v1/systemone unchanged: 202/231, 0 diffs vs the bench.
This commit is contained in:
@@ -19,8 +19,9 @@ SCRIBERR_BIND=0.0.0.0
|
||||
SCRIBERR_ALLOWED_ORIGINS=http://10.251.50.54:8080,http://scriberr.fv.internal:8080
|
||||
|
||||
# ── GPU ──────────────────────────────────────────────────────────────────
|
||||
# GPU0 is fully committed to the `gen` seat; GPU1 is the one with headroom.
|
||||
SCRIBERR_GPU_ID=1
|
||||
# GPU 3 since 2026-09-30 (Prime): an on-demand tenant of the full-size-seat reserve; it steps aside
|
||||
# (to irv-ml1's A6000) when a full-size seat claims GPU 3. See compose.yaml.
|
||||
SCRIBERR_GPU_ID=3
|
||||
|
||||
# ── Storage (on /tank — NOT the root pool, weights are multi-GB) ─────────
|
||||
SCRIBERR_DATA_DIR=/tank/scriberr/data
|
||||
@@ -48,7 +49,7 @@ SCRIBERR_SECURE_COOKIES=false
|
||||
# tab — keep the configured model on a free local seat.
|
||||
# SCRIBERR_OPENAI_API_KEY=
|
||||
|
||||
# GPU 1 memory budget (2026-09-30). Parakeet slice length in seconds and the
|
||||
# Memory settings measured on GPU 1 (2026-09-30; still in force on GPU 3). Parakeet slice length in seconds and the
|
||||
# torch allocator mode; see compose.yaml for the measurements. Defaults apply
|
||||
# when unset; override only with a re-measured peak.
|
||||
# SCRIBERR_PARAKEET_CHUNK_SECS=120
|
||||
|
||||
Reference in New Issue
Block a user