docs(intern-decision): Triton warm state survives restart, lost on recreate; ~16 x 2048-token buckets, ~2 min full warm-up
This commit is contained in:
@@ -104,8 +104,16 @@ probability diffs against the bench's r1..r4. `/decide` is unchanged.
|
||||
|
||||
⚠ **Cold-shape latency:** the first call in a new length bucket after a (re)start costs ~6.5 s extra
|
||||
(3,001 and 4,000 words were slow; 3,002–3,500 and 4,097–5,000 were not). The fast kernels autotune per shape
|
||||
bucket, and the result is cached in-process. Warm calls are as tabled. A startup warm-up across the buckets would
|
||||
remove it; that is not done yet.
|
||||
bucket. Warm calls are as tabled.
|
||||
- **How long the warm state lasts (measured 1541):** it lives in the Triton cache at `/tmp/triton-cache`, including
|
||||
fla's `*.autotune.json`, in the container's writable layer. It SURVIVES `docker restart` (and so host reboots):
|
||||
after a restart, a seen 32k bucket answered in 2.0 s. It is LOST on recreate (`compose up -d` after a change, a
|
||||
deploy, an image upgrade). An unseen bucket after the restart still cost +6.4 s (12,187 tokens), which is the
|
||||
positive control.
|
||||
- **Buckets:** every observation so far fits buckets of 2,048 tokens (inferred from ~20 calls, not proven), so ~16
|
||||
buckets up to 32,768. A full warm-up is estimated at ~2 min, once per container.
|
||||
- **Durable fix (not done):** put `/tmp/triton-cache` on a named volume so that recreates keep it, and run one
|
||||
~2 min warm-up across the buckets after an image change (a new triton/fla version keys new entries).
|
||||
|
||||
The table below is the ORIGINAL budget (9.0 GiB / 7,168 tokens, beside scriberr) and is kept for history.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user