feat(intern-decision): cap 9.0 GiB with MAX_TOKENS 7168, the largest call measured to fit
Both are required in compose because they are coupled: MAX_TOKENS is checked before the forward pass, so an oversized call is a clear 422 instead of reaching the cap as a 503. Pre-deploy floor is nvidia-smi Free >= 15,400 MiB on GPU 1 (card peak 9,876 + scriberr 5,496).
This commit is contained in:
@@ -8,11 +8,11 @@
|
||||
# fv-ml1 from that dir.
|
||||
#
|
||||
# ⚠ VRAM_CAP_GIB is a HARD cap on torch's allocator (per-process memory fraction), set so the
|
||||
# container's WHOLE nvidia-smi footprint, CUDA context included, stays <= 10,300 MiB whatever
|
||||
# the request (Prime/infra-ops budget with scriberr, 2026-09-30). A request that needs more
|
||||
# container's WHOLE nvidia-smi footprint, CUDA context included, fits beside scriberr's peak
|
||||
# (infra-ops budget, 2026-09-30); MAX_TOKENS keeps every accepted call under the cap. A request that needs more
|
||||
# gets 503 out_of_memory and the service stays up. See the README before changing it.
|
||||
#
|
||||
# .env (tunables): IMAGE, PORT, GPU_ID, VRAM_CAP_GIB, HOST_IP, INTERN_DECISION_API_TOKEN
|
||||
# .env (tunables): IMAGE, PORT, GPU_ID, VRAM_CAP_GIB, MAX_TOKENS, HOST_IP, INTERN_DECISION_API_TOKEN
|
||||
# (vault intern-decision/api-token, >= 32 chars).
|
||||
|
||||
name: intern-decision
|
||||
@@ -28,7 +28,8 @@ services:
|
||||
INTERN_DECISION_API_TOKEN: ${INTERN_DECISION_API_TOKEN:?set INTERN_DECISION_API_TOKEN}
|
||||
INTERN_DECISION_DEVICE: cuda
|
||||
INTERN_DECISION_VRAM_CAP_GIB: ${VRAM_CAP_GIB:?set VRAM_CAP_GIB}
|
||||
INTERN_DECISION_MAX_TOKENS: ${MAX_TOKENS:-8192}
|
||||
# Coupled to VRAM_CAP_GIB: the largest call measured to fit under the cap (README "VRAM").
|
||||
INTERN_DECISION_MAX_TOKENS: ${MAX_TOKENS:?set MAX_TOKENS}
|
||||
INTERN_DECISION_MAX_DECISIONS: ${MAX_DECISIONS:-64}
|
||||
# POSTs in progress (queued + scoring) before new ones get 429 busy.
|
||||
INTERN_DECISION_MAX_QUEUE: ${MAX_QUEUE:-32}
|
||||
|
||||
Reference in New Issue
Block a user