Files
esh-pfi-infrastructure/services/intern-decision-serve/acceptance/systemone-2026-09-30/warmup-warm.txt
T
vh f65e27b08f feat(intern-decision): persist the triton autotune cache across recreates
The ~6.5-9 s first-call-per-bucket autotune lived in the container's writable
layer and died on every recreate. 0.1.3 creates /tmp/triton-cache in the image
owned by 10001 so the named volume intern-decision_triton-cache inherits a
writable mount point, and compose mounts it.

Bucket model PROVEN, not inferred: 2,048-token buckets, 16 up to 32,768. After
one warmed call per bucket, 12 random sizes across 8k-32k were all warm (worst
2.09 s); cold entries cost 6.5-9 s. Full cold warm-up 109 s; warm re-run 17 s.

scripts/intern-decision-warmup: one noul call per bucket, MAX_TOKENS from
/health, two-point live calibration of the tokenizer's linear token model (a
single probe overcorrects and the aim oscillates around the bucket edge),
per-bucket wall times, non-zero exit on a missed bucket. Run it after an IMAGE
CHANGE only; the volume carries ordinary recreates (measured: force-recreate,
then a warmed 32k call answered in 2.11 s).

Acceptance on 0.1.3: JevBench 202/231, hard 83/111, 0 diffs / 924; warm 32k GPU
1 peak 15,218 MiB (budget 15,220; a COLD autotune touched 15,224 once, README
caveat); /decide answers. Artifacts in the acceptance dir.
2026-09-30 16:02:16 -07:00

21 lines
685 B
Plaintext

endpoint=http://intern-decision.fv.internal:8033 max_tokens=32768 bucket_width=2048 buckets=16
calibrated: input_tokens = 5.0000*reps + 183.0
bucket<= tokens wall_s note
2048 1988 0.11
4096 4038 0.20
6144 6083 0.31
8192 8133 0.43
10240 10178 0.56
12288 12228 0.70
14336 14278 0.84
16384 16323 0.97
18432 18373 1.12
20480 20418 1.26
22528 22468 1.40
24576 24518 1.55
26624 26563 1.71
28672 28613 1.85
30720 30658 2.01
32768 32708 2.17
total 17.2 s over 16 buckets