feat(intern-decision): persist the triton autotune cache across recreates
The ~6.5-9 s first-call-per-bucket autotune lived in the container's writable layer and died on every recreate. 0.1.3 creates /tmp/triton-cache in the image owned by 10001 so the named volume intern-decision_triton-cache inherits a writable mount point, and compose mounts it. Bucket model PROVEN, not inferred: 2,048-token buckets, 16 up to 32,768. After one warmed call per bucket, 12 random sizes across 8k-32k were all warm (worst 2.09 s); cold entries cost 6.5-9 s. Full cold warm-up 109 s; warm re-run 17 s. scripts/intern-decision-warmup: one noul call per bucket, MAX_TOKENS from /health, two-point live calibration of the tokenizer's linear token model (a single probe overcorrects and the aim oscillates around the bucket edge), per-bucket wall times, non-zero exit on a missed bucket. Run it after an IMAGE CHANGE only; the volume carries ordinary recreates (measured: force-recreate, then a warmed 32k call answered in 2.11 s). Acceptance on 0.1.3: JevBench 202/231, hard 83/111, 0 diffs / 924; warm 32k GPU 1 peak 15,218 MiB (budget 15,220; a COLD autotune touched 15,224 once, README caveat); /decide answers. Artifacts in the acceptance dir.
This commit is contained in:
@@ -1,6 +1,6 @@
|
||||
[project]
|
||||
name = "intern-decision-serve"
|
||||
version = "0.1.2"
|
||||
version = "0.1.3"
|
||||
description = "Intern-Decision-4B (its own inference.py) behind semif-serve's HTTP surface"
|
||||
requires-python = ">=3.12"
|
||||
dependencies = [
|
||||
|
||||
Reference in New Issue
Block a user