feat(intern-decision): persist the triton autotune cache across recreates

The ~6.5-9 s first-call-per-bucket autotune lived in the container's writable
layer and died on every recreate. 0.1.3 creates /tmp/triton-cache in the image
owned by 10001 so the named volume intern-decision_triton-cache inherits a
writable mount point, and compose mounts it.

Bucket model PROVEN, not inferred: 2,048-token buckets, 16 up to 32,768. After
one warmed call per bucket, 12 random sizes across 8k-32k were all warm (worst
2.09 s); cold entries cost 6.5-9 s. Full cold warm-up 109 s; warm re-run 17 s.

scripts/intern-decision-warmup: one noul call per bucket, MAX_TOKENS from
/health, two-point live calibration of the tokenizer's linear token model (a
single probe overcorrects and the aim oscillates around the bucket edge),
per-bucket wall times, non-zero exit on a missed bucket. Run it after an IMAGE
CHANGE only; the volume carries ordinary recreates (measured: force-recreate,
then a warmed 32k call answered in 2.11 s).

Acceptance on 0.1.3: JevBench 202/231, hard 83/111, 0 diffs / 924; warm 32k GPU
1 peak 15,218 MiB (budget 15,220; a COLD autotune touched 15,224 once, README
caveat); /decide answers. Artifacts in the acceptance dir.
This commit is contained in:
vh
2026-09-30 16:02:16 -07:00
parent 38015a1977
commit f65e27b08f
15 changed files with 4683 additions and 14 deletions
@@ -41,3 +41,17 @@ infra-ops audit PASS with two low findings, both fixed in 0.1.2:
Re-run on 0.1.2 (jb2-* files, the 0.1.1 artifacts kept with their version suffix): all **202/231**,
hard **83/111**, **0 changed rows** across 924. 124 tests green. Rollback: 0.1.1.
## 0.1.3 — triton-cache volume + warm-up script (same day, Prime task)
Image change only (Dockerfile creates /tmp/triton-cache owned 10001; compose mounts the named
volume there). Measured facts in the stack README "Cold-shape latency": bucket width 2,048 PROVEN,
16 buckets, cold full warm-up 109 s, warm re-run 17 s. Acceptance here:
- warmup-cold.txt / warmup-warm.txt: the two runs; every bucket in-bucket, cold 109.3 s vs 17.2 s.
- random-validation.txt: 12 random sizes 8.3k–31.8k after warm-up, worst wall 2.09 s — no cold call.
- Recreate survival: force-recreate, then a warmed 32k call answered in 2.11 s; volume holds
3,200+ files owned by uid 10001 (new files newer than the image build = the volume is in use).
- Memory: warm 32k peak 15,218 MiB (budget 15,220); a COLD autotune touched 15,224 once, during
the pre-deploy 0.1.2-era measurement — see README caveat. gpu1-cold-run.csv, gpu1-32k-warm.csv.
- JevBench on 0.1.3: all 202/231, hard 83/111, 0 diffs / 924 (jb-results-0.1.3.jsonl).
- /decide positive control: answered (yes, 0.969).