feat(intern-decision): persist the triton autotune cache across recreates

The ~6.5-9 s first-call-per-bucket autotune lived in the container's writable
layer and died on every recreate. 0.1.3 creates /tmp/triton-cache in the image
owned by 10001 so the named volume intern-decision_triton-cache inherits a
writable mount point, and compose mounts it.

Bucket model PROVEN, not inferred: 2,048-token buckets, 16 up to 32,768. After
one warmed call per bucket, 12 random sizes across 8k-32k were all warm (worst
2.09 s); cold entries cost 6.5-9 s. Full cold warm-up 109 s; warm re-run 17 s.

scripts/intern-decision-warmup: one noul call per bucket, MAX_TOKENS from
/health, two-point live calibration of the tokenizer's linear token model (a
single probe overcorrects and the aim oscillates around the bucket edge),
per-bucket wall times, non-zero exit on a missed bucket. Run it after an IMAGE
CHANGE only; the volume carries ordinary recreates (measured: force-recreate,
then a warmed 32k call answered in 2.11 s).

Acceptance on 0.1.3: JevBench 202/231, hard 83/111, 0 diffs / 924; warm 32k GPU
1 peak 15,218 MiB (budget 15,220; a COLD autotune touched 15,224 once, README
caveat); /decide answers. Artifacts in the acceptance dir.
This commit is contained in:
vh
2026-09-30 16:02:16 -07:00
parent 38015a1977
commit f65e27b08f
15 changed files with 4683 additions and 14 deletions
@@ -0,0 +1,14 @@
tokens~ tokens wall_s
8285 8280 0.45
11080 11075 0.62
14286 14285 0.84
16164 16160 0.97
17447 17445 1.05
21661 21660 1.35
21966 21965 1.37
23112 23110 1.45
24358 24355 1.54
24567 24565 1.56
24911 24910 1.58
31822 31820 2.09
worst wall 2.09 s — every call warm