feat(intern-decision): cap 9.0 GiB with MAX_TOKENS 7168, the largest call measured to fit

Both are required in compose because they are coupled: MAX_TOKENS is checked before the forward
pass, so an oversized call is a clear 422 instead of reaching the cap as a 503. Pre-deploy floor
is nvidia-smi Free >= 15,400 MiB on GPU 1 (card peak 9,876 + scriberr 5,496).
This commit is contained in:
vh
2026-09-30 09:40:08 -07:00
parent 66034cc69e
commit 750675e391
4 changed files with 38 additions and 13 deletions
+8 -5
View File
@@ -6,10 +6,13 @@ HOST_IP=10.251.50.54
# fv-ml1 GPU 1 = the utility card (vllm-coder, erp, meromero, scriberr).
GPU_ID=1
# HARD torch-allocator cap: the single knob that holds the container's WHOLE nvidia-smi footprint
# (CUDA context included) <= 10,300 MiB, the GPU 1 budget next to scriberr (2026-09-30).
# footprint <= cap + non-allocator overhead = 9,472 MiB + 662 MiB = 10,134 MiB
# (overhead measured flat at 660-662 MiB with the single inference thread; README "VRAM").
# 9.25 still answers the largest request the API accepts (4 calls x 8,191 tokens) with a 200.
VRAM_CAP_GIB=9.25
# (CUDA context included) inside GPU 1's budget next to scriberr (infra-ops, 2026-09-30):
# GPU 1 nvidia-smi Free >= 15,400 MiB = our card peak 9,876 (cap 9.0 GiB + 660 MiB outside the
# allocator, measured) + scriberr's peak 5,496, rounded up. README "VRAM".
VRAM_CAP_GIB=9.0
# Tokens per CALL (state + up to 16 questions), checked BEFORE the forward pass: a longer call is a
# clear 422. 7,168 is the largest call measured to fit under VRAM_CAP_GIB=9.0. Change the two TOGETHER,
# and re-measure (README "VRAM"): a larger value would let a call reach the cap and return 503.
MAX_TOKENS=7168
# >= 32 characters; source of truth: secret get intern-decision/api-token
INTERN_DECISION_API_TOKEN=
+7 -2
View File
@@ -190,9 +190,14 @@ cd /opt/docker/compose/intern-decision && new=$(sed 's/^IMAGE=.*/IMAGE=intern-de
&& printf '%s\n' "$new" > .env && docker compose config -q && docker compose up -d
```
**Before any deploy onto GPU 1:** GPU 1 must have at least 15,800 MiB free
(`nvidia-smi -i 1 --query-gpu=memory.free --format=csv`). If it has less, stop. Do not squeeze
**Before any deploy onto GPU 1**, check two things. If either fails, stop; do not squeeze
scriberr.
- nvidia-smi's own `Free` on GPU 1 must be at least **15,400 MiB**:
`nvidia-smi -i 1 --query-gpu=memory.free --format=csv`. That is our card peak of 9,876 MiB plus
scriberr's 5,496, rounded up. Do not use total − used, which misses the driver's 640 MiB
reserve.
- Scriberr must not be running a job. This command must print 0:
`docker logs --since 2m scriberr | grep -c "Processing single-track job"`.
Startup fails closed. A container that never reaches healthy did not pass its own checks: the
`inference.py` hash, the pinned snapshot, the warm-up, the text-only swap and the prompt hash.
+5 -4
View File
@@ -8,11 +8,11 @@
# fv-ml1 from that dir.
#
# ⚠ VRAM_CAP_GIB is a HARD cap on torch's allocator (per-process memory fraction), set so the
# container's WHOLE nvidia-smi footprint, CUDA context included, stays <= 10,300 MiB whatever
# the request (Prime/infra-ops budget with scriberr, 2026-09-30). A request that needs more
# container's WHOLE nvidia-smi footprint, CUDA context included, fits beside scriberr's peak
# (infra-ops budget, 2026-09-30); MAX_TOKENS keeps every accepted call under the cap. A request that needs more
# gets 503 out_of_memory and the service stays up. See the README before changing it.
#
# .env (tunables): IMAGE, PORT, GPU_ID, VRAM_CAP_GIB, HOST_IP, INTERN_DECISION_API_TOKEN
# .env (tunables): IMAGE, PORT, GPU_ID, VRAM_CAP_GIB, MAX_TOKENS, HOST_IP, INTERN_DECISION_API_TOKEN
# (vault intern-decision/api-token, >= 32 chars).
name: intern-decision
@@ -28,7 +28,8 @@ services:
INTERN_DECISION_API_TOKEN: ${INTERN_DECISION_API_TOKEN:?set INTERN_DECISION_API_TOKEN}
INTERN_DECISION_DEVICE: cuda
INTERN_DECISION_VRAM_CAP_GIB: ${VRAM_CAP_GIB:?set VRAM_CAP_GIB}
INTERN_DECISION_MAX_TOKENS: ${MAX_TOKENS:-8192}
# Coupled to VRAM_CAP_GIB: the largest call measured to fit under the cap (README "VRAM").
INTERN_DECISION_MAX_TOKENS: ${MAX_TOKENS:?set MAX_TOKENS}
INTERN_DECISION_MAX_DECISIONS: ${MAX_DECISIONS:-64}
# POSTs in progress (queued + scoring) before new ones get 429 busy.
INTERN_DECISION_MAX_QUEUE: ${MAX_QUEUE:-32}