The ~6.5-9 s first-call-per-bucket autotune lived in the container's writable layer and died on every recreate. 0.1.3 creates /tmp/triton-cache in the image owned by 10001 so the named volume intern-decision_triton-cache inherits a writable mount point, and compose mounts it. Bucket model PROVEN, not inferred: 2,048-token buckets, 16 up to 32,768. After one warmed call per bucket, 12 random sizes across 8k-32k were all warm (worst 2.09 s); cold entries cost 6.5-9 s. Full cold warm-up 109 s; warm re-run 17 s. scripts/intern-decision-warmup: one noul call per bucket, MAX_TOKENS from /health, two-point live calibration of the tokenizer's linear token model (a single probe overcorrects and the aim oscillates around the bucket edge), per-bucket wall times, non-zero exit on a missed bucket. Run it after an IMAGE CHANGE only; the volume carries ordinary recreates (measured: force-recreate, then a warmed 32k call answered in 2.11 s). Acceptance on 0.1.3: JevBench 202/231, hard 83/111, 0 diffs / 924; warm 32k GPU 1 peak 15,218 MiB (budget 15,220; a COLD autotune touched 15,224 once, README caveat); /decide answers. Artifacts in the acceptance dir.
21 lines
686 B
Plaintext
21 lines
686 B
Plaintext
endpoint=http://intern-decision.fv.internal:8033 max_tokens=32768 bucket_width=2048 buckets=16
|
|
calibrated: input_tokens = 5.0000*reps + 183.0
|
|
bucket<= tokens wall_s note
|
|
2048 1988 0.11
|
|
4096 4038 6.53
|
|
6144 6083 0.31
|
|
8192 8133 6.75
|
|
10240 10178 7.00
|
|
12288 12228 7.16
|
|
14336 14278 7.28
|
|
16384 16323 7.52
|
|
18432 18373 7.72
|
|
20480 20418 7.92
|
|
22528 22468 8.08
|
|
24576 24518 8.24
|
|
26624 26563 8.40
|
|
28672 28613 8.60
|
|
30720 30658 8.77
|
|
32768 32708 8.90
|
|
total 109.3 s over 16 buckets
|