scriberr to fv-ml1 GPU 3 (on-demand, steps aside to irv-ml1 A6000); intern-decision 32k-token calls (cap 14.4 GiB)
Prime 2026-09-30: move scriberr to GPU 3 and extend the Jev endpoint to 32k tokens. Scriberr holds 0 VRAM idle; verified a 20-min job on GPU 3 at 5,496 MiB. With GPU 1 freed, intern-decision's measured card peak at MAX_TOKENS=32768 is 15,220 MiB against a 15,437 MiB budget (n=3, 1 and 16 questions); 32,769 tokens is refused 422 up front. JevBench v1.2.16 via /v1/systemone unchanged: 202/231, 0 diffs vs the bench.
This commit is contained in:
@@ -1,18 +1,19 @@
|
||||
# intern-decision — copy to /opt/docker/compose/intern-decision/.env on fv-ml1 (mode 0600).
|
||||
# Built on fv-ml1 from services/intern-decision-serve (see README "Building").
|
||||
IMAGE=intern-decision-serve:0.1.0
|
||||
IMAGE=intern-decision-serve:0.1.2
|
||||
PORT=8033
|
||||
HOST_IP=10.251.50.54
|
||||
# fv-ml1 GPU 1 = the utility card (vllm-coder, erp, meromero, scriberr).
|
||||
# fv-ml1 GPU 1 = the utility card (vllm-coder, erp, meromero). scriberr moved to GPU 3 (2026-09-30).
|
||||
GPU_ID=1
|
||||
# HARD torch-allocator cap: the single knob that holds the container's WHOLE nvidia-smi footprint
|
||||
# (CUDA context included) inside GPU 1's budget next to scriberr (infra-ops, 2026-09-30):
|
||||
# GPU 1 nvidia-smi Free >= 15,400 MiB = our card peak 9,876 (cap 9.0 GiB + 660 MiB outside the
|
||||
# allocator, measured) + scriberr's peak 5,496, rounded up. README "VRAM".
|
||||
VRAM_CAP_GIB=9.0
|
||||
# (CUDA context included, ~660 MiB outside the allocator) inside GPU 1's free memory (infra-ops,
|
||||
# 2026-09-30 1330): with the vLLM seats static, this container may use its rest (8,812) + GPU 1
|
||||
# nvidia-smi Free (6,625) = 15,437 MiB. 14.4 GiB cap -> card ceiling ~15,408. README "VRAM".
|
||||
VRAM_CAP_GIB=14.4
|
||||
# Tokens per CALL (state + up to 16 questions), checked BEFORE the forward pass: a longer call is a
|
||||
# clear 422. 7,168 is the largest call measured to fit under VRAM_CAP_GIB=9.0. Change the two TOGETHER,
|
||||
# and re-measure (README "VRAM"): a larger value would let a call reach the cap and return 503.
|
||||
MAX_TOKENS=7168
|
||||
# clear 422. 32,768 = Jev's "32k for state plus the longest question"; measured to fit under 14.4 GiB
|
||||
# at a card peak of 15,220 MiB (1 question and 16 questions, n=3 each). Change the two TOGETHER, and
|
||||
# re-measure (README "VRAM"): a larger value would let a call reach the cap and return 503.
|
||||
MAX_TOKENS=32768
|
||||
# >= 32 characters; source of truth: secret get intern-decision/api-token
|
||||
INTERN_DECISION_API_TOKEN=
|
||||
|
||||
@@ -80,7 +80,36 @@ curl -s -H "Authorization: Bearer $T" http://intern-decision.fv.internal:8033/de
|
||||
never truncated.
|
||||
- A request with more decisions is split into more calls, and each call must fit.
|
||||
|
||||
## VRAM: fits beside scriberr's peak, whatever the request
|
||||
## VRAM: 32k-token calls since 2026-09-30 1330 (scriberr moved to GPU 3)
|
||||
|
||||
**Current setting: `VRAM_CAP_GIB=14.4`, `MAX_TOKENS=32768`** (Prime: move scriberr to GPU 3 and "extend the jev
|
||||
endpoint to hit 32k tokens if possible"). With scriberr gone, GPU 1 holds only the static vLLM seats and this
|
||||
service. This container may therefore use its rest (8,812 MiB) plus GPU 1's nvidia-smi `Free` (6,625 MiB), which is
|
||||
**15,437 MiB**. The 14.4 GiB cap plus the ~660 MiB outside the allocator puts the card ceiling at ~15,408 MiB.
|
||||
|
||||
Measured on the live service (per-process nvidia-smi every 0.1 s; single `noul` question unless noted; n=3, deterministic):
|
||||
|
||||
| call tokens | card peak MiB | wall (warm) |
|
||||
|---|---|---|
|
||||
| 3,187 | 9,306 | 0.15 s |
|
||||
| 12,187 | 11,206 | 0.68 s |
|
||||
| 24,187 | 13,552 | 1.49 s |
|
||||
| 29,987 | 14,692 | 1.92 s |
|
||||
| **32,768 (limit)** | **15,220** | 2.12 s |
|
||||
| 32,765, 16 questions | 15,220 | 2.15 s |
|
||||
| 32,769 | refused 422 before the forward | 0.30 s |
|
||||
|
||||
Spare at the limit: 217 MiB. JevBench v1.2.16 through `/v1/systemone` after the change: 202/231, with 0 answer and 0
|
||||
probability diffs against the bench's r1..r4. `/decide` is unchanged.
|
||||
|
||||
⚠ **Cold-shape latency:** the first call in a new length bucket after a (re)start costs ~6.5 s extra
|
||||
(3,001 and 4,000 words were slow; 3,002–3,500 and 4,097–5,000 were not). The fast kernels autotune per shape
|
||||
bucket, and the result is cached in-process. Warm calls are as tabled. A startup warm-up across the buckets would
|
||||
remove it; that is not done yet.
|
||||
|
||||
The table below is the ORIGINAL budget (9.0 GiB / 7,168 tokens, beside scriberr) and is kept for history.
|
||||
|
||||
## VRAM (history): fits beside scriberr's peak, whatever the request
|
||||
|
||||
**Budget (infra-ops, 2026-09-30):** GPU 1 needs nvidia-smi `Free` ≥ **15,400 MiB** before this
|
||||
service starts. That is our card peak of 9,876 MiB plus scriberr's peak of 5,496 MiB (with its
|
||||
|
||||
@@ -1,5 +1,6 @@
|
||||
# intern-decision: Intern-Decision-4B (internlm, Apache-2.0) behind intern-decision-serve, on
|
||||
# fv-ml1 GPU 1 (the utility card, beside vllm-coder, the erp/meromero seats and scriberr).
|
||||
# fv-ml1 GPU 1 (the utility card, beside vllm-coder and the erp/meromero seats; scriberr moved to
|
||||
# GPU 3 on 2026-09-30 1322, Prime, to free this card's headroom for 32k-token calls).
|
||||
# Replaces semif (Prime, 2026-09-30: "replace semif with intern-decision now").
|
||||
#
|
||||
# One forward pass per call, scored by the checkpoint's OWN inference.py (sha256-pinned); the
|
||||
@@ -8,8 +9,8 @@
|
||||
# fv-ml1 from that dir.
|
||||
#
|
||||
# ⚠ VRAM_CAP_GIB is a HARD cap on torch's allocator (per-process memory fraction), set so the
|
||||
# container's WHOLE nvidia-smi footprint, CUDA context included, fits beside scriberr's peak
|
||||
# (infra-ops budget, 2026-09-30); MAX_TOKENS keeps every accepted call under the cap. A request that needs more
|
||||
# container's WHOLE nvidia-smi footprint, CUDA context included, fits GPU 1's free memory beside
|
||||
# the static vLLM seats (infra-ops budget, 2026-09-30); MAX_TOKENS keeps every accepted call under the cap. A request that needs more
|
||||
# gets 503 out_of_memory and the service stays up. See the README before changing it.
|
||||
#
|
||||
# .env (tunables): IMAGE, PORT, GPU_ID, VRAM_CAP_GIB, MAX_TOKENS, HOST_IP, INTERN_DECISION_API_TOKEN
|
||||
|
||||
Reference in New Issue
Block a user