Commit Graph
6 Commits
Author SHA1 Message Date
vh f65e27b08f feat(intern-decision): persist the triton autotune cache across recreates
The ~6.5-9 s first-call-per-bucket autotune lived in the container's writable
layer and died on every recreate. 0.1.3 creates /tmp/triton-cache in the image
owned by 10001 so the named volume intern-decision_triton-cache inherits a
writable mount point, and compose mounts it.

Bucket model PROVEN, not inferred: 2,048-token buckets, 16 up to 32,768. After
one warmed call per bucket, 12 random sizes across 8k-32k were all warm (worst
2.09 s); cold entries cost 6.5-9 s. Full cold warm-up 109 s; warm re-run 17 s.

scripts/intern-decision-warmup: one noul call per bucket, MAX_TOKENS from
/health, two-point live calibration of the tokenizer's linear token model (a
single probe overcorrects and the aim oscillates around the bucket edge),
per-bucket wall times, non-zero exit on a missed bucket. Run it after an IMAGE
CHANGE only; the volume carries ordinary recreates (measured: force-recreate,
then a warmed 32k call answered in 2.11 s).

Acceptance on 0.1.3: JevBench 202/231, hard 83/111, 0 diffs / 924; warm 32k GPU
1 peak 15,218 MiB (budget 15,220; a COLD autotune touched 15,224 once, README
caveat); /decide answers. Artifacts in the acceptance dir.
2026-09-30 16:02:16 -07:00
vh af450f4ef7 docs(intern-decision): Triton warm state survives restart, lost on recreate; ~16 x 2048-token buckets, ~2 min full warm-up 2026-09-30 15:41:55 -07:00
vh 6b201e1d4a scriberr to fv-ml1 GPU 3 (on-demand, steps aside to irv-ml1 A6000); intern-decision 32k-token calls (cap 14.4 GiB)
Prime 2026-09-30: move scriberr to GPU 3 and extend the Jev endpoint to 32k tokens.
Scriberr holds 0 VRAM idle; verified a 20-min job on GPU 3 at 5,496 MiB. With GPU 1
freed, intern-decision's measured card peak at MAX_TOKENS=32768 is 15,220 MiB against
a 15,437 MiB budget (n=3, 1 and 16 questions); 32,769 tokens is refused 422 up front.
JevBench v1.2.16 via /v1/systemone unchanged: 202/231, 0 diffs vs the bench.
2026-09-30 13:35:32 -07:00
vh 1cf763a7b1 docs(intern-decision): live on fv-ml1 GPU 1; semif marked REPLACED
intern-decision deployed 0941 PT (cap 9.0 GiB, MAX_TOKENS 7168). Live acceptance: positive
control 240/259 and Wyrd 79/84, bit-identical to the bench (0/560 rows, Δp 0); negative control
10/122/14; largest accepted requests 200 with no 503; per-process 8,812 MiB at rest and 9,866 peak;
GPU 1 Free 15,442 before and 6,581 after (lowest 5,569 under load). Latency from nh3-dev:
21 criteria 114 ms, 16 over ~3,900 tokens 238 ms. semif README banner now REPLACED with the
rollback; fv-ml1 GPU 1 note updated.
2026-09-30 09:48:30 -07:00
vh 750675e391 feat(intern-decision): cap 9.0 GiB with MAX_TOKENS 7168, the largest call measured to fit
Both are required in compose because they are coupled: MAX_TOKENS is checked before the forward
pass, so an oversized call is a clear 422 instead of reaching the cap as a 503. Pre-deploy floor
is nvidia-smi Free >= 15,400 MiB on GPU 1 (card peak 9,876 + scriberr 5,496).
2026-09-30 09:40:08 -07:00
vh a262477a61 feat(intern-decision): stack, DNS and GPU 3 acceptance for the SemIf replacement
stacks/intern-decision: compose (GPU 1, :8033, hard VRAM cap as the single .env knob,
healthcheck, Homepage group 'AI - Eval & Retrieval'), .env.example and README.
dns: intern-decision.fv.internal -> fv-ml1 (synced to ana/esh/nh3).
acceptance on fv-ml1 GPU 3, 3 fresh processes: bit-identical to the Jev bench's native rows
(pooled 240/259, Wyrd 79/84, 0/560 flips, Δp 0), negative control 10/122/14, 0 flips across
restarts; largest accepted request 200 at a 10,134 MiB card peak under a 9.25 GiB cap; 503 and
recovery proven at a tight cap. GPU 1 deploy held: nvidia-smi Free on GPU 1 is 15,442 MiB.
2026-09-30 09:38:00 -07:00