Both are required in compose because they are coupled: MAX_TOKENS is checked before the forward pass, so an oversized call is a clear 422 instead of reaching the cap as a 503. Pre-deploy floor is nvidia-smi Free >= 15,400 MiB on GPU 1 (card peak 9,876 + scriberr 5,496).
19 lines
1.1 KiB
Bash
19 lines
1.1 KiB
Bash
# intern-decision — copy to /opt/docker/compose/intern-decision/.env on fv-ml1 (mode 0600).
|
|
# Built on fv-ml1 from services/intern-decision-serve (see README "Building").
|
|
IMAGE=intern-decision-serve:0.1.0
|
|
PORT=8033
|
|
HOST_IP=10.251.50.54
|
|
# fv-ml1 GPU 1 = the utility card (vllm-coder, erp, meromero, scriberr).
|
|
GPU_ID=1
|
|
# HARD torch-allocator cap: the single knob that holds the container's WHOLE nvidia-smi footprint
|
|
# (CUDA context included) inside GPU 1's budget next to scriberr (infra-ops, 2026-09-30):
|
|
# GPU 1 nvidia-smi Free >= 15,400 MiB = our card peak 9,876 (cap 9.0 GiB + 660 MiB outside the
|
|
# allocator, measured) + scriberr's peak 5,496, rounded up. README "VRAM".
|
|
VRAM_CAP_GIB=9.0
|
|
# Tokens per CALL (state + up to 16 questions), checked BEFORE the forward pass: a longer call is a
|
|
# clear 422. 7,168 is the largest call measured to fit under VRAM_CAP_GIB=9.0. Change the two TOGETHER,
|
|
# and re-measure (README "VRAM"): a larger value would let a call reach the cap and return 503.
|
|
MAX_TOKENS=7168
|
|
# >= 32 characters; source of truth: secret get intern-decision/api-token
|
|
INTERN_DECISION_API_TOKEN=
|