The ~6.5-9 s first-call-per-bucket autotune lived in the container's writable layer and died on every recreate. 0.1.3 creates /tmp/triton-cache in the image owned by 10001 so the named volume intern-decision_triton-cache inherits a writable mount point, and compose mounts it. Bucket model PROVEN, not inferred: 2,048-token buckets, 16 up to 32,768. After one warmed call per bucket, 12 random sizes across 8k-32k were all warm (worst 2.09 s); cold entries cost 6.5-9 s. Full cold warm-up 109 s; warm re-run 17 s. scripts/intern-decision-warmup: one noul call per bucket, MAX_TOKENS from /health, two-point live calibration of the tokenizer's linear token model (a single probe overcorrects and the aim oscillates around the bucket edge), per-bucket wall times, non-zero exit on a missed bucket. Run it after an IMAGE CHANGE only; the volume carries ordinary recreates (measured: force-recreate, then a warmed 32k call answered in 2.11 s). Acceptance on 0.1.3: JevBench 202/231, hard 83/111, 0 diffs / 924; warm 32k GPU 1 peak 15,218 MiB (budget 15,220; a COLD autotune touched 15,224 once, README caveat); /decide answers. Artifacts in the acceptance dir.
75 lines
3.4 KiB
YAML
75 lines
3.4 KiB
YAML
# intern-decision: Intern-Decision-4B (internlm, Apache-2.0) behind intern-decision-serve, on
|
|
# fv-ml1 GPU 1 (the utility card, beside vllm-coder and the erp/meromero seats; scriberr moved to
|
|
# GPU 3 on 2026-09-30 1322, Prime, to free this card's headroom for 32k-token calls).
|
|
# Replaces semif (Prime, 2026-09-30: "replace semif with intern-decision now").
|
|
#
|
|
# One forward pass per call, scored by the checkpoint's OWN inference.py (sha256-pinned); the
|
|
# service keeps semif-serve's HTTP surface (/decide, /decide/shared, /health). Service code +
|
|
# contract: services/intern-decision-serve/ (intern-decision-serve.contract.md). Image built on
|
|
# fv-ml1 from that dir.
|
|
#
|
|
# ⚠ VRAM_CAP_GIB is a HARD cap on torch's allocator (per-process memory fraction), set so the
|
|
# container's WHOLE nvidia-smi footprint, CUDA context included, fits GPU 1's free memory beside
|
|
# the static vLLM seats (infra-ops budget, 2026-09-30); MAX_TOKENS keeps every accepted call under the cap. A request that needs more
|
|
# gets 503 out_of_memory and the service stays up. See the README before changing it.
|
|
#
|
|
# .env (tunables): IMAGE, PORT, GPU_ID, VRAM_CAP_GIB, MAX_TOKENS, HOST_IP, INTERN_DECISION_API_TOKEN
|
|
# (vault intern-decision/api-token, >= 32 chars).
|
|
|
|
name: intern-decision
|
|
|
|
services:
|
|
intern-decision:
|
|
image: ${IMAGE:?set IMAGE}
|
|
container_name: intern-decision
|
|
restart: unless-stopped
|
|
ports:
|
|
- "${PORT:-8033}:8000"
|
|
environment:
|
|
INTERN_DECISION_API_TOKEN: ${INTERN_DECISION_API_TOKEN:?set INTERN_DECISION_API_TOKEN}
|
|
INTERN_DECISION_DEVICE: cuda
|
|
INTERN_DECISION_VRAM_CAP_GIB: ${VRAM_CAP_GIB:?set VRAM_CAP_GIB}
|
|
# Coupled to VRAM_CAP_GIB: the largest call measured to fit under the cap (README "VRAM").
|
|
INTERN_DECISION_MAX_TOKENS: ${MAX_TOKENS:?set MAX_TOKENS}
|
|
INTERN_DECISION_MAX_DECISIONS: ${MAX_DECISIONS:-64}
|
|
# POSTs in progress (queued + scoring) before new ones get 429 busy.
|
|
INTERN_DECISION_MAX_QUEUE: ${MAX_QUEUE:-32}
|
|
volumes:
|
|
# Pinned weights AND the checkpoint's inference.py, read offline. Never downloads.
|
|
- /tank/aimodels/huggingface:/hf:ro
|
|
# Triton/fla autotune cache (README "Cold-shape latency"): the ~6.5 s first-call-per-bucket
|
|
# autotune results survive RECREATES (deploys, image upgrades, .env edits), not just restarts.
|
|
# The mount point exists in the image owned by 10001, so the fresh volume is intern-writable.
|
|
- triton-cache:/tmp/triton-cache
|
|
deploy:
|
|
resources:
|
|
reservations:
|
|
devices:
|
|
- driver: nvidia
|
|
device_ids: ["${GPU_ID:-1}"]
|
|
capabilities: [gpu]
|
|
healthcheck:
|
|
test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=5)"]
|
|
interval: 30s
|
|
timeout: 10s
|
|
retries: 3
|
|
# Startup loads ~9 GB of weights and scores the warm-up three times before it serves.
|
|
start_period: 300s
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=AI - Eval & Retrieval
|
|
- homepage.name=Intern-Decision — typed decisions
|
|
- homepage.icon=mdi-scale-balance
|
|
- homepage.description=Typed decisions from one forward pass (Intern-Decision-4B, fv-ml1 GPU1)
|
|
- homepage.href=http://${HOST_IP:-10.251.50.54}:${PORT:-8033}/health
|
|
|
|
networks:
|
|
tnet:
|
|
name: traefik-net
|
|
external: true
|
|
|
|
volumes:
|
|
# -> intern-decision_triton-cache on the host
|
|
triton-cache:
|