scriberr to fv-ml1 GPU 3 (on-demand, steps aside to irv-ml1 A6000); intern-decision 32k-token calls (cap 14.4 GiB)

Prime 2026-09-30: move scriberr to GPU 3 and extend the Jev endpoint to 32k tokens.
Scriberr holds 0 VRAM idle; verified a 20-min job on GPU 3 at 5,496 MiB. With GPU 1
freed, intern-decision's measured card peak at MAX_TOKENS=32768 is 15,220 MiB against
a 15,437 MiB budget (n=3, 1 and 16 questions); 32,769 tokens is refused 422 up front.
JevBench v1.2.16 via /v1/systemone unchanged: 202/231, 0 diffs vs the bench.
This commit is contained in:
vh
2026-09-30 13:35:32 -07:00
parent 92501a29c1
commit 6b201e1d4a
8 changed files with 79 additions and 29 deletions
+13 -6
View File
@@ -181,13 +181,14 @@ embed/rerank/reward trio. GPUs are pinned per container via
**GPU 1 — light / eval / retrieval + char-RP GGUF (~91/98 GB, on-demand):**
> ⚠ **This table is stale (checked 2026-09-27).** Live GPU 1 residents are `scriberr`,
> `vllm-coder`, `vllm-erp-seat` and `vllm-meromero-rp`.
> ⚠ **This table is stale (checked 2026-09-27).** Live GPU 1 residents are `vllm-coder`,
> `vllm-erp-seat` and `vllm-meromero-rp` (`scriberr` moved to GPU 3 on 2026-09-30 1322).
> **`intern-decision`** joined them on 2026-09-30, 0941 PT: :8033, 8,812 MiB at rest, 9,866 MiB
> peak, hard-capped at 9.0 GiB with `MAX_TOKENS` 7,168; see `stacks/intern-decision`. It replaced
> **`semif`** (:8032), which is stopped and kept as the rollback, per Prime's ruling of 2026-09-30.
> **GPU 1 budget:** nvidia-smi `Free` read 15,442 MiB before intern-decision and 6,581 MiB after,
> at rest. That covers scriberr's 5,496 MiB peak even while intern-decision is at its own peak.
> peak at first; since 1330 it is capped at **14.4 GiB with `MAX_TOKENS` 32,768** (card peak 15,220 MiB at the
> limit), because scriberr left this card. See `stacks/intern-decision`. It replaced **`semif`** (:8032),
> whose container was removed and is kept as the rollback, per Prime's ruling of 2026-09-30.
> **GPU 1 budget:** nvidia-smi `Free` reads 6,625 MiB with intern-decision at rest. All of it is intern-decision's
> headroom for 32k calls (217 MiB spare at its peak). Nothing else fits on this card now.
> Measure nvidia-smi `Free` (the driver reserves 640 MiB per card) before adding anything to
> this card. Read the host
> (`docker inspect … DeviceRequests`), not this table. (A fixtures-only `augaman` instance ran
@@ -236,6 +237,12 @@ GPU 2 at 0.96).
while in use (`restart: "no"`, about 270 MiB when idle with the desktop running, 0 when down).
The reserve still stands: whenever a full-size seat takes GPU 3, Blender stays down.
**GPU 3 on-demand tenant (Prime, 2026-09-30 1322): `scriberr`** (`stacks/scriberr/`). It holds 0 VRAM when
idle and peaks at ~5.5 GB per job. Same rule as Blender: when a full-size seat claims GPU 3, Scriberr steps aside,
and its planned landing spot is **irv-ml1's A6000** (~32 GB free on 2026-09-30, but shared with bursty ComfyUI
work; check the peaks first). It does NOT go back to GPU 1, because GPU 1's headroom now funds intern-decision's
32k-token calls (cap 14.4 GiB).
**Retired:**
- `llama-swap` (former GGUF multiplexer on :9292) — replaced by dedicated
per-model seats (e.g. `llama-charrp`); no longer running.