scriberr to fv-ml1 GPU 3 (on-demand, steps aside to irv-ml1 A6000); intern-decision 32k-token calls (cap 14.4 GiB)
Prime 2026-09-30: move scriberr to GPU 3 and extend the Jev endpoint to 32k tokens. Scriberr holds 0 VRAM idle; verified a 20-min job on GPU 3 at 5,496 MiB. With GPU 1 freed, intern-decision's measured card peak at MAX_TOKENS=32768 is 15,220 MiB against a 15,437 MiB budget (n=3, 1 and 16 questions); 32,769 tokens is refused 422 up front. JevBench v1.2.16 via /v1/systemone unchanged: 202/231, 0 diffs vs the bench.
This commit is contained in:
@@ -1,18 +1,19 @@
|
||||
# intern-decision — copy to /opt/docker/compose/intern-decision/.env on fv-ml1 (mode 0600).
|
||||
# Built on fv-ml1 from services/intern-decision-serve (see README "Building").
|
||||
IMAGE=intern-decision-serve:0.1.0
|
||||
IMAGE=intern-decision-serve:0.1.2
|
||||
PORT=8033
|
||||
HOST_IP=10.251.50.54
|
||||
# fv-ml1 GPU 1 = the utility card (vllm-coder, erp, meromero, scriberr).
|
||||
# fv-ml1 GPU 1 = the utility card (vllm-coder, erp, meromero). scriberr moved to GPU 3 (2026-09-30).
|
||||
GPU_ID=1
|
||||
# HARD torch-allocator cap: the single knob that holds the container's WHOLE nvidia-smi footprint
|
||||
# (CUDA context included) inside GPU 1's budget next to scriberr (infra-ops, 2026-09-30):
|
||||
# GPU 1 nvidia-smi Free >= 15,400 MiB = our card peak 9,876 (cap 9.0 GiB + 660 MiB outside the
|
||||
# allocator, measured) + scriberr's peak 5,496, rounded up. README "VRAM".
|
||||
VRAM_CAP_GIB=9.0
|
||||
# (CUDA context included, ~660 MiB outside the allocator) inside GPU 1's free memory (infra-ops,
|
||||
# 2026-09-30 1330): with the vLLM seats static, this container may use its rest (8,812) + GPU 1
|
||||
# nvidia-smi Free (6,625) = 15,437 MiB. 14.4 GiB cap -> card ceiling ~15,408. README "VRAM".
|
||||
VRAM_CAP_GIB=14.4
|
||||
# Tokens per CALL (state + up to 16 questions), checked BEFORE the forward pass: a longer call is a
|
||||
# clear 422. 7,168 is the largest call measured to fit under VRAM_CAP_GIB=9.0. Change the two TOGETHER,
|
||||
# and re-measure (README "VRAM"): a larger value would let a call reach the cap and return 503.
|
||||
MAX_TOKENS=7168
|
||||
# clear 422. 32,768 = Jev's "32k for state plus the longest question"; measured to fit under 14.4 GiB
|
||||
# at a card peak of 15,220 MiB (1 question and 16 questions, n=3 each). Change the two TOGETHER, and
|
||||
# re-measure (README "VRAM"): a larger value would let a call reach the cap and return 503.
|
||||
MAX_TOKENS=32768
|
||||
# >= 32 characters; source of truth: secret get intern-decision/api-token
|
||||
INTERN_DECISION_API_TOKEN=
|
||||
|
||||
Reference in New Issue
Block a user