scriberr to fv-ml1 GPU 3 (on-demand, steps aside to irv-ml1 A6000); intern-decision 32k-token calls (cap 14.4 GiB)

Prime 2026-09-30: move scriberr to GPU 3 and extend the Jev endpoint to 32k tokens.
Scriberr holds 0 VRAM idle; verified a 20-min job on GPU 3 at 5,496 MiB. With GPU 1
freed, intern-decision's measured card peak at MAX_TOKENS=32768 is 15,220 MiB against
a 15,437 MiB budget (n=3, 1 and 16 questions); 32,769 tokens is refused 422 up front.
JevBench v1.2.16 via /v1/systemone unchanged: 202/231, 0 diffs vs the bench.
This commit is contained in:
vh
2026-09-30 13:35:32 -07:00
parent 92501a29c1
commit 6b201e1d4a
8 changed files with 79 additions and 29 deletions
+30 -1
View File
@@ -80,7 +80,36 @@ curl -s -H "Authorization: Bearer $T" http://intern-decision.fv.internal:8033/de
never truncated.
- A request with more decisions is split into more calls, and each call must fit.
## VRAM: fits beside scriberr's peak, whatever the request
## VRAM: 32k-token calls since 2026-09-30 1330 (scriberr moved to GPU 3)
**Current setting: `VRAM_CAP_GIB=14.4`, `MAX_TOKENS=32768`** (Prime: move scriberr to GPU 3 and "extend the jev
endpoint to hit 32k tokens if possible"). With scriberr gone, GPU 1 holds only the static vLLM seats and this
service. This container may therefore use its rest (8,812 MiB) plus GPU 1's nvidia-smi `Free` (6,625 MiB), which is
**15,437 MiB**. The 14.4 GiB cap plus the ~660 MiB outside the allocator puts the card ceiling at ~15,408 MiB.
Measured on the live service (per-process nvidia-smi every 0.1 s; single `noul` question unless noted; n=3, deterministic):
| call tokens | card peak MiB | wall (warm) |
|---|---|---|
| 3,187 | 9,306 | 0.15 s |
| 12,187 | 11,206 | 0.68 s |
| 24,187 | 13,552 | 1.49 s |
| 29,987 | 14,692 | 1.92 s |
| **32,768 (limit)** | **15,220** | 2.12 s |
| 32,765, 16 questions | 15,220 | 2.15 s |
| 32,769 | refused 422 before the forward | 0.30 s |
Spare at the limit: 217 MiB. JevBench v1.2.16 through `/v1/systemone` after the change: 202/231, with 0 answer and 0
probability diffs against the bench's r1..r4. `/decide` is unchanged.
⚠ **Cold-shape latency:** the first call in a new length bucket after a (re)start costs ~6.5 s extra
(3,001 and 4,000 words were slow; 3,002–3,500 and 4,097–5,000 were not). The fast kernels autotune per shape
bucket, and the result is cached in-process. Warm calls are as tabled. A startup warm-up across the buckets would
remove it; that is not done yet.
The table below is the ORIGINAL budget (9.0 GiB / 7,168 tokens, beside scriberr) and is kept for history.
## VRAM (history): fits beside scriberr's peak, whatever the request
**Budget (infra-ops, 2026-09-30):** GPU 1 needs nvidia-smi `Free` ≥ **15,400 MiB** before this
service starts. That is our card peak of 9,876 MiB plus scriberr's peak of 5,496 MiB (with its