feat(intern-decision): persist the triton autotune cache across recreates
The ~6.5-9 s first-call-per-bucket autotune lived in the container's writable layer and died on every recreate. 0.1.3 creates /tmp/triton-cache in the image owned by 10001 so the named volume intern-decision_triton-cache inherits a writable mount point, and compose mounts it. Bucket model PROVEN, not inferred: 2,048-token buckets, 16 up to 32,768. After one warmed call per bucket, 12 random sizes across 8k-32k were all warm (worst 2.09 s); cold entries cost 6.5-9 s. Full cold warm-up 109 s; warm re-run 17 s. scripts/intern-decision-warmup: one noul call per bucket, MAX_TOKENS from /health, two-point live calibration of the tokenizer's linear token model (a single probe overcorrects and the aim oscillates around the bucket edge), per-bucket wall times, non-zero exit on a missed bucket. Run it after an IMAGE CHANGE only; the volume carries ordinary recreates (measured: force-recreate, then a warmed 32k call answered in 2.11 s). Acceptance on 0.1.3: JevBench 202/231, hard 83/111, 0 diffs / 924; warm 32k GPU 1 peak 15,218 MiB (budget 15,220; a COLD autotune touched 15,224 once, README caveat); /decide answers. Artifacts in the acceptance dir.
This commit is contained in:
@@ -102,18 +102,27 @@ Measured on the live service (per-process nvidia-smi every 0.1 s; single `noul`
|
||||
Spare at the limit: 217 MiB. JevBench v1.2.16 through `/v1/systemone` after the change: 202/231, with 0 answer and 0
|
||||
probability diffs against the bench's r1..r4. `/decide` is unchanged.
|
||||
|
||||
⚠ **Cold-shape latency:** the first call in a new length bucket after a (re)start costs ~6.5 s extra
|
||||
(3,001 and 4,000 words were slow; 3,002–3,500 and 4,097–5,000 were not). The fast kernels autotune per shape
|
||||
bucket. Warm calls are as tabled.
|
||||
- **How long the warm state lasts (measured 1541):** it lives in the Triton cache at `/tmp/triton-cache`, including
|
||||
fla's `*.autotune.json`, in the container's writable layer. It SURVIVES `docker restart` (and so host reboots):
|
||||
after a restart, a seen 32k bucket answered in 2.0 s. It is LOST on recreate (`compose up -d` after a change, a
|
||||
deploy, an image upgrade). An unseen bucket after the restart still cost +6.4 s (12,187 tokens), which is the
|
||||
positive control.
|
||||
- **Buckets:** every observation so far fits buckets of 2,048 tokens (inferred from ~20 calls, not proven), so ~16
|
||||
buckets up to 32,768. A full warm-up is estimated at ~2 min, once per container.
|
||||
- **Durable fix (not done):** put `/tmp/triton-cache` on a named volume so that recreates keep it, and run one
|
||||
~2 min warm-up across the buckets after an image change (a new triton/fla version keys new entries).
|
||||
⚠ **Cold-shape latency (MEASURED 2026-09-30, was inferred):** the fast kernels (triton + fla) autotune
|
||||
per shape bucket. **Buckets are 2,048 tokens wide** — proven: after warming one call per bucket, 12
|
||||
calls at random sizes spread 8k–32k were all warm (worst 2.09 s, the same curve as the warmed sizes;
|
||||
no +6 s), and every cold call that crossed into its own bucket cost 6.5–9 s while same-bucket repeats
|
||||
cost <0.3 s. **16 buckets** to MAX_TOKENS=32,768. A full cold warm-up took **109 s** (16 buckets; the
|
||||
service's own startup warm-up and the script's two calibration probes warm 3 of them along the way;
|
||||
a second run totals 17 s). Per-bucket cold cost grows with size: 6.5 s at 4k → 8.9 s at 32k.
|
||||
- **Durable now:** the autotune cache (TRITON_CACHE_DIR=/tmp/triton-cache, incl. fla's
|
||||
*.autotune.json) lives on the named volume `intern-decision_triton-cache`, mounted at a mount
|
||||
point baked into the image owned by 10001 (a fresh named volume inherits it; without that, the
|
||||
volume is root-owned and Triton cannot write). Proven: 3,200+ files owned by uid 10001 in the
|
||||
volume; after `up -d --force-recreate` a previously warmed 32k call answered in 2.11 s.
|
||||
- **When to warm:** run `scripts/intern-decision-warmup` (on nh3-dev; token from the vault) after
|
||||
an IMAGE CHANGE — a new triton/fla version keys new cache entries. Not needed after ordinary
|
||||
recreates. The script reads MAX_TOKENS from /health, aims each bucket just under its edge with a
|
||||
two-point live calibration (single-probe "calibration" oscillates ±500 tokens around the edge),
|
||||
prints per-bucket wall times, and exits non-zero if any bucket was missed.
|
||||
- Cold autotune also costs a few MiB of transient scratch: GPU 1 per-process peak measured
|
||||
**15,218 MiB warm** at the 32k limit (budget 15,220), but **15,224** during a COLD 32k autotune.
|
||||
The volume means the cold path now happens only after an image change — worth knowing when
|
||||
GPU 1 headroom is being read during a deploy.
|
||||
|
||||
The table below is the ORIGINAL budget (9.0 GiB / 7,168 tokens, beside scriberr) and is kept for history.
|
||||
|
||||
@@ -253,6 +262,9 @@ scripts/deploy-stack.sh fv-ml1 intern-decision
|
||||
# on fv-ml1: .env is infra-ops's, mode 0600. Rewrite it IN PLACE to keep the owner and mode.
|
||||
cd /opt/docker/compose/intern-decision && new=$(sed 's/^IMAGE=.*/IMAGE=intern-decision-serve:X.Y.Z/' .env) \
|
||||
&& printf '%s\n' "$new" > .env && docker compose config -q && docker compose up -d
|
||||
# IMAGE CHANGE ONLY (not ordinary recreates — the triton-cache volume carries those):
|
||||
# warm every length bucket once the container is healthy (~2 min, prints per-bucket times):
|
||||
scripts/intern-decision-warmup # token from vault; MAX_TOKENS read from /health
|
||||
```
|
||||
|
||||
**Before any deploy onto GPU 1**, check two things. If either fails, stop; do not squeeze
|
||||
|
||||
Reference in New Issue
Block a user