scriberr to fv-ml1 GPU 3 (on-demand, steps aside to irv-ml1 A6000); intern-decision 32k-token calls (cap 14.4 GiB)
Prime 2026-09-30: move scriberr to GPU 3 and extend the Jev endpoint to 32k tokens. Scriberr holds 0 VRAM idle; verified a 20-min job on GPU 3 at 5,496 MiB. With GPU 1 freed, intern-decision's measured card peak at MAX_TOKENS=32768 is 15,220 MiB against a 15,437 MiB budget (n=3, 1 and 16 questions); 32,769 tokens is refused 422 up front. JevBench v1.2.16 via /v1/systemone unchanged: 202/231, 0 diffs vs the bench.
This commit is contained in:
@@ -181,13 +181,14 @@ embed/rerank/reward trio. GPUs are pinned per container via
|
||||
|
||||
**GPU 1 — light / eval / retrieval + char-RP GGUF (~91/98 GB, on-demand):**
|
||||
|
||||
> ⚠ **This table is stale (checked 2026-09-27).** Live GPU 1 residents are `scriberr`,
|
||||
> `vllm-coder`, `vllm-erp-seat` and `vllm-meromero-rp`.
|
||||
> ⚠ **This table is stale (checked 2026-09-27).** Live GPU 1 residents are `vllm-coder`,
|
||||
> `vllm-erp-seat` and `vllm-meromero-rp` (`scriberr` moved to GPU 3 on 2026-09-30 1322).
|
||||
> **`intern-decision`** joined them on 2026-09-30, 0941 PT: :8033, 8,812 MiB at rest, 9,866 MiB
|
||||
> peak, hard-capped at 9.0 GiB with `MAX_TOKENS` 7,168; see `stacks/intern-decision`. It replaced
|
||||
> **`semif`** (:8032), which is stopped and kept as the rollback, per Prime's ruling of 2026-09-30.
|
||||
> **GPU 1 budget:** nvidia-smi `Free` read 15,442 MiB before intern-decision and 6,581 MiB after,
|
||||
> at rest. That covers scriberr's 5,496 MiB peak even while intern-decision is at its own peak.
|
||||
> peak at first; since 1330 it is capped at **14.4 GiB with `MAX_TOKENS` 32,768** (card peak 15,220 MiB at the
|
||||
> limit), because scriberr left this card. See `stacks/intern-decision`. It replaced **`semif`** (:8032),
|
||||
> whose container was removed and is kept as the rollback, per Prime's ruling of 2026-09-30.
|
||||
> **GPU 1 budget:** nvidia-smi `Free` reads 6,625 MiB with intern-decision at rest. All of it is intern-decision's
|
||||
> headroom for 32k calls (217 MiB spare at its peak). Nothing else fits on this card now.
|
||||
> Measure nvidia-smi `Free` (the driver reserves 640 MiB per card) before adding anything to
|
||||
> this card. Read the host
|
||||
> (`docker inspect … DeviceRequests`), not this table. (A fixtures-only `augaman` instance ran
|
||||
@@ -236,6 +237,12 @@ GPU 2 at 0.96).
|
||||
while in use (`restart: "no"`, about 270 MiB when idle with the desktop running, 0 when down).
|
||||
The reserve still stands: whenever a full-size seat takes GPU 3, Blender stays down.
|
||||
|
||||
**GPU 3 on-demand tenant (Prime, 2026-09-30 1322): `scriberr`** (`stacks/scriberr/`). It holds 0 VRAM when
|
||||
idle and peaks at ~5.5 GB per job. Same rule as Blender: when a full-size seat claims GPU 3, Scriberr steps aside,
|
||||
and its planned landing spot is **irv-ml1's A6000** (~32 GB free on 2026-09-30, but shared with bursty ComfyUI
|
||||
work; check the peaks first). It does NOT go back to GPU 1, because GPU 1's headroom now funds intern-decision's
|
||||
32k-token calls (cap 14.4 GiB).
|
||||
|
||||
**Retired:**
|
||||
- `llama-swap` (former GGUF multiplexer on :9292) — replaced by dedicated
|
||||
per-model seats (e.g. `llama-charrp`); no longer running.
|
||||
|
||||
Reference in New Issue
Block a user