feat(intern-decision): persist the triton autotune cache across recreates

The ~6.5-9 s first-call-per-bucket autotune lived in the container's writable
layer and died on every recreate. 0.1.3 creates /tmp/triton-cache in the image
owned by 10001 so the named volume intern-decision_triton-cache inherits a
writable mount point, and compose mounts it.

Bucket model PROVEN, not inferred: 2,048-token buckets, 16 up to 32,768. After
one warmed call per bucket, 12 random sizes across 8k-32k were all warm (worst
2.09 s); cold entries cost 6.5-9 s. Full cold warm-up 109 s; warm re-run 17 s.

scripts/intern-decision-warmup: one noul call per bucket, MAX_TOKENS from
/health, two-point live calibration of the tokenizer's linear token model (a
single probe overcorrects and the aim oscillates around the bucket edge),
per-bucket wall times, non-zero exit on a missed bucket. Run it after an IMAGE
CHANGE only; the volume carries ordinary recreates (measured: force-recreate,
then a warmed 32k call answered in 2.11 s).

Acceptance on 0.1.3: JevBench 202/231, hard 83/111, 0 diffs / 924; warm 32k GPU
1 peak 15,218 MiB (budget 15,220; a COLD autotune touched 15,224 once, README
caveat); /decide answers. Artifacts in the acceptance dir.
This commit is contained in:
vh
2026-09-30 16:02:16 -07:00
parent 38015a1977
commit f65e27b08f
15 changed files with 4683 additions and 14 deletions
+24 -12
View File
@@ -102,18 +102,27 @@ Measured on the live service (per-process nvidia-smi every 0.1 s; single `noul`
Spare at the limit: 217 MiB. JevBench v1.2.16 through `/v1/systemone` after the change: 202/231, with 0 answer and 0
probability diffs against the bench's r1..r4. `/decide` is unchanged.
⚠ **Cold-shape latency:** the first call in a new length bucket after a (re)start costs ~6.5 s extra
(3,001 and 4,000 words were slow; 3,002–3,500 and 4,097–5,000 were not). The fast kernels autotune per shape
bucket. Warm calls are as tabled.
- **How long the warm state lasts (measured 1541):** it lives in the Triton cache at `/tmp/triton-cache`, including
fla's `*.autotune.json`, in the container's writable layer. It SURVIVES `docker restart` (and so host reboots):
after a restart, a seen 32k bucket answered in 2.0 s. It is LOST on recreate (`compose up -d` after a change, a
deploy, an image upgrade). An unseen bucket after the restart still cost +6.4 s (12,187 tokens), which is the
positive control.
- **Buckets:** every observation so far fits buckets of 2,048 tokens (inferred from ~20 calls, not proven), so ~16
buckets up to 32,768. A full warm-up is estimated at ~2 min, once per container.
- **Durable fix (not done):** put `/tmp/triton-cache` on a named volume so that recreates keep it, and run one
~2 min warm-up across the buckets after an image change (a new triton/fla version keys new entries).
⚠ **Cold-shape latency (MEASURED 2026-09-30, was inferred):** the fast kernels (triton + fla) autotune
per shape bucket. **Buckets are 2,048 tokens wide** — proven: after warming one call per bucket, 12
calls at random sizes spread 8k–32k were all warm (worst 2.09 s, the same curve as the warmed sizes;
no +6 s), and every cold call that crossed into its own bucket cost 6.5–9 s while same-bucket repeats
cost <0.3 s. **16 buckets** to MAX_TOKENS=32,768. A full cold warm-up took **109 s** (16 buckets; the
service's own startup warm-up and the script's two calibration probes warm 3 of them along the way;
a second run totals 17 s). Per-bucket cold cost grows with size: 6.5 s at 4k → 8.9 s at 32k.
- **Durable now:** the autotune cache (TRITON_CACHE_DIR=/tmp/triton-cache, incl. fla's
*.autotune.json) lives on the named volume `intern-decision_triton-cache`, mounted at a mount
point baked into the image owned by 10001 (a fresh named volume inherits it; without that, the
volume is root-owned and Triton cannot write). Proven: 3,200+ files owned by uid 10001 in the
volume; after `up -d --force-recreate` a previously warmed 32k call answered in 2.11 s.
- **When to warm:** run `scripts/intern-decision-warmup` (on nh3-dev; token from the vault) after
an IMAGE CHANGE — a new triton/fla version keys new cache entries. Not needed after ordinary
recreates. The script reads MAX_TOKENS from /health, aims each bucket just under its edge with a
two-point live calibration (single-probe "calibration" oscillates ±500 tokens around the edge),
prints per-bucket wall times, and exits non-zero if any bucket was missed.
- Cold autotune also costs a few MiB of transient scratch: GPU 1 per-process peak measured
**15,218 MiB warm** at the 32k limit (budget 15,220), but **15,224** during a COLD 32k autotune.
The volume means the cold path now happens only after an image change — worth knowing when
GPU 1 headroom is being read during a deploy.
The table below is the ORIGINAL budget (9.0 GiB / 7,168 tokens, beside scriberr) and is kept for history.
@@ -253,6 +262,9 @@ scripts/deploy-stack.sh fv-ml1 intern-decision
# on fv-ml1: .env is infra-ops's, mode 0600. Rewrite it IN PLACE to keep the owner and mode.
cd /opt/docker/compose/intern-decision && new=$(sed 's/^IMAGE=.*/IMAGE=intern-decision-serve:X.Y.Z/' .env) \
&& printf '%s\n' "$new" > .env && docker compose config -q && docker compose up -d
# IMAGE CHANGE ONLY (not ordinary recreates — the triton-cache volume carries those):
# warm every length bucket once the container is healthy (~2 min, prints per-bucket times):
scripts/intern-decision-warmup # token from vault; MAX_TOKENS read from /health
```
**Before any deploy onto GPU 1**, check two things. If either fails, stop; do not squeeze