docs(intern-decision): Triton warm state survives restart, lost on recreate; ~16 x 2048-token buckets, ~2 min full warm-up

This commit is contained in:
vh
2026-09-30 15:41:55 -07:00
parent 6b201e1d4a
commit af450f4ef7
2 changed files with 13 additions and 4 deletions
+10 -2
View File
@@ -104,8 +104,16 @@ probability diffs against the bench's r1..r4. `/decide` is unchanged.
⚠ **Cold-shape latency:** the first call in a new length bucket after a (re)start costs ~6.5 s extra
(3,001 and 4,000 words were slow; 3,002–3,500 and 4,097–5,000 were not). The fast kernels autotune per shape
bucket, and the result is cached in-process. Warm calls are as tabled. A startup warm-up across the buckets would
remove it; that is not done yet.
bucket. Warm calls are as tabled.
- **How long the warm state lasts (measured 1541):** it lives in the Triton cache at `/tmp/triton-cache`, including
fla's `*.autotune.json`, in the container's writable layer. It SURVIVES `docker restart` (and so host reboots):
after a restart, a seen 32k bucket answered in 2.0 s. It is LOST on recreate (`compose up -d` after a change, a
deploy, an image upgrade). An unseen bucket after the restart still cost +6.4 s (12,187 tokens), which is the
positive control.
- **Buckets:** every observation so far fits buckets of 2,048 tokens (inferred from ~20 calls, not proven), so ~16
buckets up to 32,768. A full warm-up is estimated at ~2 min, once per container.
- **Durable fix (not done):** put `/tmp/triton-cache` on a named volume so that recreates keep it, and run one
~2 min warm-up across the buckets after an image change (a new triton/fla version keys new entries).
The table below is the ORIGINAL budget (9.0 GiB / 7,168 tokens, beside scriberr) and is kept for history.