# intern-decision — copy to /opt/docker/compose/intern-decision/.env on fv-ml1 (mode 0600). # Built on fv-ml1 from services/intern-decision-serve (see README "Building"). IMAGE=intern-decision-serve:0.1.2 PORT=8033 HOST_IP=10.251.50.54 # fv-ml1 GPU 1 = the utility card (vllm-coder, erp, meromero). scriberr moved to GPU 3 (2026-09-30). GPU_ID=1 # HARD torch-allocator cap: the single knob that holds the container's WHOLE nvidia-smi footprint # (CUDA context included, ~660 MiB outside the allocator) inside GPU 1's free memory (infra-ops, # 2026-09-30 1330): with the vLLM seats static, this container may use its rest (8,812) + GPU 1 # nvidia-smi Free (6,625) = 15,437 MiB. 14.4 GiB cap -> card ceiling ~15,408. README "VRAM". VRAM_CAP_GIB=14.4 # Tokens per CALL (state + up to 16 questions), checked BEFORE the forward pass: a longer call is a # clear 422. 32,768 = Jev's "32k for state plus the longest question"; measured to fit under 14.4 GiB # at a card peak of 15,220 MiB (1 question and 16 questions, n=3 each). Change the two TOGETHER, and # re-measure (README "VRAM"): a larger value would let a call reach the cap and return 503. MAX_TOKENS=32768 # >= 32 characters; source of truth: secret get intern-decision/api-token INTERN_DECISION_API_TOKEN=