feat(intern-decision): persist the triton autotune cache across recreates

The ~6.5-9 s first-call-per-bucket autotune lived in the container's writable
layer and died on every recreate. 0.1.3 creates /tmp/triton-cache in the image
owned by 10001 so the named volume intern-decision_triton-cache inherits a
writable mount point, and compose mounts it.

Bucket model PROVEN, not inferred: 2,048-token buckets, 16 up to 32,768. After
one warmed call per bucket, 12 random sizes across 8k-32k were all warm (worst
2.09 s); cold entries cost 6.5-9 s. Full cold warm-up 109 s; warm re-run 17 s.

scripts/intern-decision-warmup: one noul call per bucket, MAX_TOKENS from
/health, two-point live calibration of the tokenizer's linear token model (a
single probe overcorrects and the aim oscillates around the bucket edge),
per-bucket wall times, non-zero exit on a missed bucket. Run it after an IMAGE
CHANGE only; the volume carries ordinary recreates (measured: force-recreate,
then a warmed 32k call answered in 2.11 s).

Acceptance on 0.1.3: JevBench 202/231, hard 83/111, 0 diffs / 924; warm 32k GPU
1 peak 15,218 MiB (budget 15,220; a COLD autotune touched 15,224 once, README
caveat); /decide answers. Artifacts in the acceptance dir.
This commit is contained in:
vh
2026-09-30 16:02:16 -07:00
parent 38015a1977
commit f65e27b08f
15 changed files with 4683 additions and 14 deletions
+8
View File
@@ -37,6 +37,10 @@ services:
volumes:
# Pinned weights AND the checkpoint's inference.py, read offline. Never downloads.
- /tank/aimodels/huggingface:/hf:ro
# Triton/fla autotune cache (README "Cold-shape latency"): the ~6.5 s first-call-per-bucket
# autotune results survive RECREATES (deploys, image upgrades, .env edits), not just restarts.
# The mount point exists in the image owned by 10001, so the fresh volume is intern-writable.
- triton-cache:/tmp/triton-cache
deploy:
resources:
reservations:
@@ -64,3 +68,7 @@ networks:
tnet:
name: traefik-net
external: true
volumes:
# -> intern-decision_triton-cache on the host
triton-cache: