scriberr to fv-ml1 GPU 3 (on-demand, steps aside to irv-ml1 A6000); intern-decision 32k-token calls (cap 14.4 GiB)

Prime 2026-09-30: move scriberr to GPU 3 and extend the Jev endpoint to 32k tokens.
Scriberr holds 0 VRAM idle; verified a 20-min job on GPU 3 at 5,496 MiB. With GPU 1
freed, intern-decision's measured card peak at MAX_TOKENS=32768 is 15,220 MiB against
a 15,437 MiB budget (n=3, 1 and 16 questions); 32,769 tokens is refused 422 up front.
JevBench v1.2.16 via /v1/systemone unchanged: 202/231, 0 diffs vs the bench.
This commit is contained in:
vh
2026-09-30 13:35:32 -07:00
parent 92501a29c1
commit 6b201e1d4a
8 changed files with 79 additions and 29 deletions
+4
View File
@@ -183,6 +183,10 @@ _As of 2026-09-30 ~0120 PT._
- **Upstream PR prepared, NOT opened; it needs Prime's yes** (`stacks/scriberr/patches/upstream-pr/PR.md`). - **Upstream PR prepared, NOT opened; it needs Prime's yes** (`stacks/scriberr/patches/upstream-pr/PR.md`).
- **OPEN, not fixed:** Parakeet skips runs of ≥10 words mid-chunk with ANY slicer, upstream's included (12–17 runs, 500–720 words per 12 transcripts; p2 lost 85 words at today's old setting). It is chaotic with cut placement. The investigation (decoder, chunk length, model) is Prime's call. - **OPEN, not fixed:** Parakeet skips runs of ≥10 words mid-chunk with ANY slicer, upstream's included (12–17 runs, 500–720 words per 12 transcripts; p2 lost 85 words at today's old setting). It is chaotic with cut placement. The investigation (decoder, chunk length, model) is Prime's call.
- Private bench data (copies of Prime's two uploads + transcripts) sits in fv-ml1 `/tank/spikes/scriberr-slicer/private/` (mode 700), kept pending Prime; the public audio and metrics are beside it. - Private bench data (copies of Prime's two uploads + transcripts) sits in fv-ml1 `/tank/spikes/scriberr-slicer/private/` (mode 700), kept pending Prime; the public audio and metrics are beside it.
- **2026-09-30 1322–1335, Prime: "Go GPU 3 now and extend the jev endpoint to hit 32k tokens".** DONE.
- **Scriberr is on fv-ml1 GPU 3** (`SCRIBERR_GPU_ID=3`; a 20-min file verified at 5,496 MiB). It is an on-demand tenant of the reserve, like Blender: it STEPS ASIDE when a full-size seat claims GPU 3, and it goes to **irv-ml1's A6000**, NOT back to GPU 1.
- **intern-decision: `VRAM_CAP_GIB=14.4`, `MAX_TOKENS=32768`** (Jev's 32k). The measured card peak at the limit is 15,220 MiB (1 and 16 questions, n=3) against a 15,437 budget; 32,769 tokens → 422; latency 2.1 s at 32k. JevBench is still 202/231 with 0 diffs.
- ⚠ The first call in a new length bucket after a restart costs ~6.5 s (kernel autotune per shape bucket). A startup warm-up across the buckets would fix it; not done.
- **intern-decision LIVE on fv-ml1 GPU 1 since 0941 2026-09-30, REPLACING SemIf (Prime: "replace semif with intern-decision now", with Scriberr fixed alongside).** - **intern-decision LIVE on fv-ml1 GPU 1 since 0941 2026-09-30, REPLACING SemIf (Prime: "replace semif with intern-decision now", with Scriberr fixed alongside).**
- Where: `http://intern-decision.fv.internal:8033`, image `intern-decision-serve:0.1.0`, token `intern-decision/api-token`. Code and contract are in `services/intern-decision-serve/`, the stack in `stacks/intern-decision`. - Where: `http://intern-decision.fv.internal:8033`, image `intern-decision-serve:0.1.0`, token `intern-decision/api-token`. Code and contract are in `services/intern-decision-serve/`, the stack in `stacks/intern-decision`.
- Surface: semif-compatible `/decide`, `/decide/shared`, `/health`. It has 12 documented deltas; the main one is that the questions in one call share a prompt, in calls of at most 16. - Surface: semif-compatible `/decide`, `/decide/shared`, `/health`. It has 12 documented deltas; the main one is that the questions in one call share a prompt, in calls of at most 16.
+13 -6
View File
@@ -181,13 +181,14 @@ embed/rerank/reward trio. GPUs are pinned per container via
**GPU 1 — light / eval / retrieval + char-RP GGUF (~91/98 GB, on-demand):** **GPU 1 — light / eval / retrieval + char-RP GGUF (~91/98 GB, on-demand):**
> ⚠ **This table is stale (checked 2026-09-27).** Live GPU 1 residents are `scriberr`, > ⚠ **This table is stale (checked 2026-09-27).** Live GPU 1 residents are `vllm-coder`,
> `vllm-coder`, `vllm-erp-seat` and `vllm-meromero-rp`. > `vllm-erp-seat` and `vllm-meromero-rp` (`scriberr` moved to GPU 3 on 2026-09-30 1322).
> **`intern-decision`** joined them on 2026-09-30, 0941 PT: :8033, 8,812 MiB at rest, 9,866 MiB > **`intern-decision`** joined them on 2026-09-30, 0941 PT: :8033, 8,812 MiB at rest, 9,866 MiB
> peak, hard-capped at 9.0 GiB with `MAX_TOKENS` 7,168; see `stacks/intern-decision`. It replaced > peak at first; since 1330 it is capped at **14.4 GiB with `MAX_TOKENS` 32,768** (card peak 15,220 MiB at the
> **`semif`** (:8032), which is stopped and kept as the rollback, per Prime's ruling of 2026-09-30. > limit), because scriberr left this card. See `stacks/intern-decision`. It replaced **`semif`** (:8032),
> **GPU 1 budget:** nvidia-smi `Free` read 15,442 MiB before intern-decision and 6,581 MiB after, > whose container was removed and is kept as the rollback, per Prime's ruling of 2026-09-30.
> at rest. That covers scriberr's 5,496 MiB peak even while intern-decision is at its own peak. > **GPU 1 budget:** nvidia-smi `Free` reads 6,625 MiB with intern-decision at rest. All of it is intern-decision's
> headroom for 32k calls (217 MiB spare at its peak). Nothing else fits on this card now.
> Measure nvidia-smi `Free` (the driver reserves 640 MiB per card) before adding anything to > Measure nvidia-smi `Free` (the driver reserves 640 MiB per card) before adding anything to
> this card. Read the host > this card. Read the host
> (`docker inspect … DeviceRequests`), not this table. (A fixtures-only `augaman` instance ran > (`docker inspect … DeviceRequests`), not this table. (A fixtures-only `augaman` instance ran
@@ -236,6 +237,12 @@ GPU 2 at 0.96).
while in use (`restart: "no"`, about 270 MiB when idle with the desktop running, 0 when down). while in use (`restart: "no"`, about 270 MiB when idle with the desktop running, 0 when down).
The reserve still stands: whenever a full-size seat takes GPU 3, Blender stays down. The reserve still stands: whenever a full-size seat takes GPU 3, Blender stays down.
**GPU 3 on-demand tenant (Prime, 2026-09-30 1322): `scriberr`** (`stacks/scriberr/`). It holds 0 VRAM when
idle and peaks at ~5.5 GB per job. Same rule as Blender: when a full-size seat claims GPU 3, Scriberr steps aside,
and its planned landing spot is **irv-ml1's A6000** (~32 GB free on 2026-09-30, but shared with bursty ComfyUI
work; check the peaks first). It does NOT go back to GPU 1, because GPU 1's headroom now funds intern-decision's
32k-token calls (cap 14.4 GiB).
**Retired:** **Retired:**
- `llama-swap` (former GGUF multiplexer on :9292) — replaced by dedicated - `llama-swap` (former GGUF multiplexer on :9292) — replaced by dedicated
per-model seats (e.g. `llama-charrp`); no longer running. per-model seats (e.g. `llama-charrp`); no longer running.
+10 -9
View File
@@ -1,18 +1,19 @@
# intern-decision — copy to /opt/docker/compose/intern-decision/.env on fv-ml1 (mode 0600). # intern-decision — copy to /opt/docker/compose/intern-decision/.env on fv-ml1 (mode 0600).
# Built on fv-ml1 from services/intern-decision-serve (see README "Building"). # Built on fv-ml1 from services/intern-decision-serve (see README "Building").
IMAGE=intern-decision-serve:0.1.0 IMAGE=intern-decision-serve:0.1.2
PORT=8033 PORT=8033
HOST_IP=10.251.50.54 HOST_IP=10.251.50.54
# fv-ml1 GPU 1 = the utility card (vllm-coder, erp, meromero, scriberr). # fv-ml1 GPU 1 = the utility card (vllm-coder, erp, meromero). scriberr moved to GPU 3 (2026-09-30).
GPU_ID=1 GPU_ID=1
# HARD torch-allocator cap: the single knob that holds the container's WHOLE nvidia-smi footprint # HARD torch-allocator cap: the single knob that holds the container's WHOLE nvidia-smi footprint
# (CUDA context included) inside GPU 1's budget next to scriberr (infra-ops, 2026-09-30): # (CUDA context included, ~660 MiB outside the allocator) inside GPU 1's free memory (infra-ops,
# GPU 1 nvidia-smi Free >= 15,400 MiB = our card peak 9,876 (cap 9.0 GiB + 660 MiB outside the # 2026-09-30 1330): with the vLLM seats static, this container may use its rest (8,812) + GPU 1
# allocator, measured) + scriberr's peak 5,496, rounded up. README "VRAM". # nvidia-smi Free (6,625) = 15,437 MiB. 14.4 GiB cap -> card ceiling ~15,408. README "VRAM".
VRAM_CAP_GIB=9.0 VRAM_CAP_GIB=14.4
# Tokens per CALL (state + up to 16 questions), checked BEFORE the forward pass: a longer call is a # Tokens per CALL (state + up to 16 questions), checked BEFORE the forward pass: a longer call is a
# clear 422. 7,168 is the largest call measured to fit under VRAM_CAP_GIB=9.0. Change the two TOGETHER, # clear 422. 32,768 = Jev's "32k for state plus the longest question"; measured to fit under 14.4 GiB
# and re-measure (README "VRAM"): a larger value would let a call reach the cap and return 503. # at a card peak of 15,220 MiB (1 question and 16 questions, n=3 each). Change the two TOGETHER, and
MAX_TOKENS=7168 # re-measure (README "VRAM"): a larger value would let a call reach the cap and return 503.
MAX_TOKENS=32768
# >= 32 characters; source of truth: secret get intern-decision/api-token # >= 32 characters; source of truth: secret get intern-decision/api-token
INTERN_DECISION_API_TOKEN= INTERN_DECISION_API_TOKEN=
+30 -1
View File
@@ -80,7 +80,36 @@ curl -s -H "Authorization: Bearer $T" http://intern-decision.fv.internal:8033/de
never truncated. never truncated.
- A request with more decisions is split into more calls, and each call must fit. - A request with more decisions is split into more calls, and each call must fit.
## VRAM: fits beside scriberr's peak, whatever the request ## VRAM: 32k-token calls since 2026-09-30 1330 (scriberr moved to GPU 3)
**Current setting: `VRAM_CAP_GIB=14.4`, `MAX_TOKENS=32768`** (Prime: move scriberr to GPU 3 and "extend the jev
endpoint to hit 32k tokens if possible"). With scriberr gone, GPU 1 holds only the static vLLM seats and this
service. This container may therefore use its rest (8,812 MiB) plus GPU 1's nvidia-smi `Free` (6,625 MiB), which is
**15,437 MiB**. The 14.4 GiB cap plus the ~660 MiB outside the allocator puts the card ceiling at ~15,408 MiB.
Measured on the live service (per-process nvidia-smi every 0.1 s; single `noul` question unless noted; n=3, deterministic):
| call tokens | card peak MiB | wall (warm) |
|---|---|---|
| 3,187 | 9,306 | 0.15 s |
| 12,187 | 11,206 | 0.68 s |
| 24,187 | 13,552 | 1.49 s |
| 29,987 | 14,692 | 1.92 s |
| **32,768 (limit)** | **15,220** | 2.12 s |
| 32,765, 16 questions | 15,220 | 2.15 s |
| 32,769 | refused 422 before the forward | 0.30 s |
Spare at the limit: 217 MiB. JevBench v1.2.16 through `/v1/systemone` after the change: 202/231, with 0 answer and 0
probability diffs against the bench's r1..r4. `/decide` is unchanged.
⚠ **Cold-shape latency:** the first call in a new length bucket after a (re)start costs ~6.5 s extra
(3,001 and 4,000 words were slow; 3,002–3,500 and 4,097–5,000 were not). The fast kernels autotune per shape
bucket, and the result is cached in-process. Warm calls are as tabled. A startup warm-up across the buckets would
remove it; that is not done yet.
The table below is the ORIGINAL budget (9.0 GiB / 7,168 tokens, beside scriberr) and is kept for history.
## VRAM (history): fits beside scriberr's peak, whatever the request
**Budget (infra-ops, 2026-09-30):** GPU 1 needs nvidia-smi `Free` ≥ **15,400 MiB** before this **Budget (infra-ops, 2026-09-30):** GPU 1 needs nvidia-smi `Free` ≥ **15,400 MiB** before this
service starts. That is our card peak of 9,876 MiB plus scriberr's peak of 5,496 MiB (with its service starts. That is our card peak of 9,876 MiB plus scriberr's peak of 5,496 MiB (with its
+4 -3
View File
@@ -1,5 +1,6 @@
# intern-decision: Intern-Decision-4B (internlm, Apache-2.0) behind intern-decision-serve, on # intern-decision: Intern-Decision-4B (internlm, Apache-2.0) behind intern-decision-serve, on
# fv-ml1 GPU 1 (the utility card, beside vllm-coder, the erp/meromero seats and scriberr). # fv-ml1 GPU 1 (the utility card, beside vllm-coder and the erp/meromero seats; scriberr moved to
# GPU 3 on 2026-09-30 1322, Prime, to free this card's headroom for 32k-token calls).
# Replaces semif (Prime, 2026-09-30: "replace semif with intern-decision now"). # Replaces semif (Prime, 2026-09-30: "replace semif with intern-decision now").
# #
# One forward pass per call, scored by the checkpoint's OWN inference.py (sha256-pinned); the # One forward pass per call, scored by the checkpoint's OWN inference.py (sha256-pinned); the
@@ -8,8 +9,8 @@
# fv-ml1 from that dir. # fv-ml1 from that dir.
# #
# ⚠ VRAM_CAP_GIB is a HARD cap on torch's allocator (per-process memory fraction), set so the # ⚠ VRAM_CAP_GIB is a HARD cap on torch's allocator (per-process memory fraction), set so the
# container's WHOLE nvidia-smi footprint, CUDA context included, fits beside scriberr's peak # container's WHOLE nvidia-smi footprint, CUDA context included, fits GPU 1's free memory beside
# (infra-ops budget, 2026-09-30); MAX_TOKENS keeps every accepted call under the cap. A request that needs more # the static vLLM seats (infra-ops budget, 2026-09-30); MAX_TOKENS keeps every accepted call under the cap. A request that needs more
# gets 503 out_of_memory and the service stays up. See the README before changing it. # gets 503 out_of_memory and the service stays up. See the README before changing it.
# #
# .env (tunables): IMAGE, PORT, GPU_ID, VRAM_CAP_GIB, MAX_TOKENS, HOST_IP, INTERN_DECISION_API_TOKEN # .env (tunables): IMAGE, PORT, GPU_ID, VRAM_CAP_GIB, MAX_TOKENS, HOST_IP, INTERN_DECISION_API_TOKEN
+4 -3
View File
@@ -19,8 +19,9 @@ SCRIBERR_BIND=0.0.0.0
SCRIBERR_ALLOWED_ORIGINS=http://10.251.50.54:8080,http://scriberr.fv.internal:8080 SCRIBERR_ALLOWED_ORIGINS=http://10.251.50.54:8080,http://scriberr.fv.internal:8080
# ── GPU ────────────────────────────────────────────────────────────────── # ── GPU ──────────────────────────────────────────────────────────────────
# GPU0 is fully committed to the `gen` seat; GPU1 is the one with headroom. # GPU 3 since 2026-09-30 (Prime): an on-demand tenant of the full-size-seat reserve; it steps aside
SCRIBERR_GPU_ID=1 # (to irv-ml1's A6000) when a full-size seat claims GPU 3. See compose.yaml.
SCRIBERR_GPU_ID=3
# ── Storage (on /tank — NOT the root pool, weights are multi-GB) ───────── # ── Storage (on /tank — NOT the root pool, weights are multi-GB) ─────────
SCRIBERR_DATA_DIR=/tank/scriberr/data SCRIBERR_DATA_DIR=/tank/scriberr/data
@@ -48,7 +49,7 @@ SCRIBERR_SECURE_COOKIES=false
# tab — keep the configured model on a free local seat. # tab — keep the configured model on a free local seat.
# SCRIBERR_OPENAI_API_KEY= # SCRIBERR_OPENAI_API_KEY=
# GPU 1 memory budget (2026-09-30). Parakeet slice length in seconds and the # Memory settings measured on GPU 1 (2026-09-30; still in force on GPU 3). Parakeet slice length in seconds and the
# torch allocator mode; see compose.yaml for the measurements. Defaults apply # torch allocator mode; see compose.yaml for the measurements. Defaults apply
# when unset; override only with a re-measured peak. # when unset; override only with a re-measured peak.
# SCRIBERR_PARAKEET_CHUNK_SECS=120 # SCRIBERR_PARAKEET_CHUNK_SECS=120
+4
View File
@@ -142,6 +142,10 @@ tab. Keep the configured model on a free local seat.
## Parakeet memory and slicing (measured 2026-09-30) ## Parakeet memory and slicing (measured 2026-09-30)
> **Moved to fv-ml1 GPU 3 at 1322 on 2026-09-30 (Prime).** The GPU 1 budget below no longer binds. Scriberr is
> an on-demand tenant of GPU 3's reserve: it steps aside to irv-ml1's A6000 when a full-size seat claims GPU 3.
> A 20-min file was verified on GPU 3 at a 5,496 MiB peak. The settings below were not changed by the move.
Scriberr shares fv-ml1 GPU 1 with intern-decision (9.0 GiB cap, 9,876 MiB card peak). Scriberr shares fv-ml1 GPU 1 with intern-decision (9.0 GiB cap, 9,876 MiB card peak).
GPU 1's nvidia-smi Free is 15,442 MiB, so the budget left for Scriberr is about 5.5 GB GPU 1's nvidia-smi Free is 15,442 MiB, so the budget left for Scriberr is about 5.5 GB
(70 MiB spare at both peaks), so the compose file sets (70 MiB spare at both peaks), so the compose file sets
+10 -7
View File
@@ -18,11 +18,13 @@
# /tank/scriberr/src/Scriberr on fv-ml1. # /tank/scriberr/src/Scriberr on fv-ml1.
# #
# ── GPU PINNING ─────────────────────────────────────────────────────────── # ── GPU PINNING ───────────────────────────────────────────────────────────
# Pinned to **GPU1** via explicit device_ids, per the house convention and # Pinned to **GPU 3** (Prime, 2026-09-30 1322) via explicit device_ids. Scriberr holds 0 VRAM
# because GPU0 is fully committed to the `gen` seat. GPU1 shares space with # when idle, so it is an ON-DEMAND tenant of GPU 3's full-size-seat reserve, like Blender: when a
# the `sec` seat, so this stack is a guest there — keep an eye on VRAM. # full-size seat (Flash-Next) claims GPU 3, Scriberr STEPS ASIDE. Its planned landing spot then is
# irv-ml1's A6000, not GPU 1 (GPU 1's headroom now funds intern-decision's 32k-token calls).
# (It was on GPU 1 until 2026-09-30, beside the vLLM seats; history in stacks/scriberr/README.md.)
# NOTE: do NOT add `NVIDIA_VISIBLE_DEVICES=all` (as upstream's compose does). # NOTE: do NOT add `NVIDIA_VISIBLE_DEVICES=all` (as upstream's compose does).
# It overrides the device_ids reservation and exposes both cards. # It overrides the device_ids reservation and exposes every card.
# #
# All tunables live in .env — edit that, not this file. # All tunables live in .env — edit that, not this file.
@@ -70,7 +72,8 @@ services:
# and takes out the Parakeet + Sortformer backends (WhisperX survives). # and takes out the Parakeet + Sortformer backends (WhisperX survives).
# `copy` trades a little disk and time for it actually working. # `copy` trades a little disk and time for it actually working.
- UV_LINK_MODE=${SCRIBERR_UV_LINK_MODE:-copy} - UV_LINK_MODE=${SCRIBERR_UV_LINK_MODE:-copy}
# ── GPU 1 memory budget (2026-09-30, Prime: Scriberr shares GPU 1 with # ── Memory settings, measured while Scriberr shared GPU 1 (2026-09-30; it moved to
# GPU 3 at 1322 the same day, so the 5.5 GB budget no longer binds, but the values stand). Prime: Scriberr shared GPU 1 with
# intern-decision, which holds ~9.7 GB resting / 10.3 GB peak). ────────── # intern-decision, which holds ~9.7 GB resting / 10.3 GB peak). ──────────
# Parakeet's buffered path cuts audio into slices of this many seconds # Parakeet's buffered path cuts audio into slices of this many seconds
# (Scriberr reads it in parakeet_adapter.go for BOTH the "is this long # (Scriberr reads it in parakeet_adapter.go for BOTH the "is this long
@@ -94,7 +97,7 @@ services:
reservations: reservations:
devices: devices:
- driver: nvidia - driver: nvidia
device_ids: ["${SCRIBERR_GPU_ID:-1}"] device_ids: ["${SCRIBERR_GPU_ID:-3}"]
capabilities: [gpu] capabilities: [gpu]
healthcheck: healthcheck:
# 127.0.0.1 rather than localhost — the IPv6-first resolution trap has # 127.0.0.1 rather than localhost — the IPv6-first resolution trap has
@@ -121,7 +124,7 @@ services:
- homepage.group=AI - Studios - homepage.group=AI - Studios
- homepage.name=Scriberr - homepage.name=Scriberr
- homepage.icon=mdi-microphone-message - homepage.icon=mdi-microphone-message
- homepage.description=Audio/video transcription + diarization (fv-ml1, GPU1) - homepage.description=Audio/video transcription + diarization (fv-ml1, GPU3)
- homepage.href=http://10.251.50.54:${SCRIBERR_PORT} - homepage.href=http://10.251.50.54:${SCRIBERR_PORT}
networks: networks: