The ~6.5-9 s first-call-per-bucket autotune lived in the container's writable layer and died on every recreate. 0.1.3 creates /tmp/triton-cache in the image owned by 10001 so the named volume intern-decision_triton-cache inherits a writable mount point, and compose mounts it. Bucket model PROVEN, not inferred: 2,048-token buckets, 16 up to 32,768. After one warmed call per bucket, 12 random sizes across 8k-32k were all warm (worst 2.09 s); cold entries cost 6.5-9 s. Full cold warm-up 109 s; warm re-run 17 s. scripts/intern-decision-warmup: one noul call per bucket, MAX_TOKENS from /health, two-point live calibration of the tokenizer's linear token model (a single probe overcorrects and the aim oscillates around the bucket edge), per-bucket wall times, non-zero exit on a missed bucket. Run it after an IMAGE CHANGE only; the volume carries ordinary recreates (measured: force-recreate, then a warmed 32k call answered in 2.11 s). Acceptance on 0.1.3: JevBench 202/231, hard 83/111, 0 diffs / 924; warm 32k GPU 1 peak 15,218 MiB (budget 15,220; a COLD autotune touched 15,224 once, README caveat); /decide answers. Artifacts in the acceptance dir.
15 lines
363 B
Plaintext
15 lines
363 B
Plaintext
tokens~ tokens wall_s
|
|
8285 8280 0.45
|
|
11080 11075 0.62
|
|
14286 14285 0.84
|
|
16164 16160 0.97
|
|
17447 17445 1.05
|
|
21661 21660 1.35
|
|
21966 21965 1.37
|
|
23112 23110 1.45
|
|
24358 24355 1.54
|
|
24567 24565 1.56
|
|
24911 24910 1.58
|
|
31822 31820 2.09
|
|
worst wall 2.09 s — every call warm
|