The ~6.5-9 s first-call-per-bucket autotune lived in the container's writable layer and died on every recreate. 0.1.3 creates /tmp/triton-cache in the image owned by 10001 so the named volume intern-decision_triton-cache inherits a writable mount point, and compose mounts it. Bucket model PROVEN, not inferred: 2,048-token buckets, 16 up to 32,768. After one warmed call per bucket, 12 random sizes across 8k-32k were all warm (worst 2.09 s); cold entries cost 6.5-9 s. Full cold warm-up 109 s; warm re-run 17 s. scripts/intern-decision-warmup: one noul call per bucket, MAX_TOKENS from /health, two-point live calibration of the tokenizer's linear token model (a single probe overcorrects and the aim oscillates around the bucket edge), per-bucket wall times, non-zero exit on a missed bucket. Run it after an IMAGE CHANGE only; the volume carries ordinary recreates (measured: force-recreate, then a warmed 32k call answered in 2.11 s). Acceptance on 0.1.3: JevBench 202/231, hard 83/111, 0 diffs / 924; warm 32k GPU 1 peak 15,218 MiB (budget 15,220; a COLD autotune touched 15,224 once, README caveat); /decide answers. Artifacts in the acceptance dir.
58 lines
3.4 KiB
Markdown
58 lines
3.4 KiB
Markdown
# /v1/systemone acceptance — 0.1.1 on fv-ml1, 2026-09-30
|
||
|
||
Live service: `http://intern-decision.fv.internal:8033`, image `intern-decision-serve:0.1.1`
|
||
(built in `/opt/docker/src/intern-decision-serve-0.1.1`; 0.1.0 kept for rollback).
|
||
|
||
## JevBench (the pass/fail line)
|
||
|
||
JevBench v1.2.16 (5e95f23), its own `typesafe` adapter, over the 231 public items:
|
||
|
||
TYPESAFE_ENDPOINT=http://intern-decision.fv.internal:8033 TYPESAFE_API_KEY=$(secret get intern-decision/api-token) \
|
||
PYTHONPATH=/tmp/jevbench python3 -m jevbench.cli run --tasks datasets/public/{easy,original,hard}.jsonl \
|
||
--adapter typesafe --results ... --ledger ... --raw-dir ... --price-in-per-m 0 --price-out-per-m 0
|
||
|
||
- **all 202/231, hard 83/111** — matches the expected numbers exactly.
|
||
- Row-by-row vs the bench's own runs r1..r4 (`bench-jev-2026-09-30/raw/out/intern-decision-4b-native/`):
|
||
**0 changed answers/probabilities across 924 rows.**
|
||
- Artifacts here: jb-results.jsonl, jb-summary.json, jb-ledger.jsonl, jb-manifest.json.
|
||
- The first-run finding: the response `model` must be a string — the runner hashes it into
|
||
its manifest and a dict crashed the manifest step. Fixed before the run above.
|
||
|
||
## Controls (live)
|
||
|
||
- wrong token → 401; 17 questions → 422 ("never chunked"); `images` → 422 "images not supported".
|
||
- token limit at the boundary: 7,167 → 200, **7,169 → 422** (binary search; "one token over").
|
||
- `/decide/shared` positive control: 2 decisions, 1 call, fields q1/q2 — unchanged.
|
||
|
||
## Memory (GPU 1 is packed beside scriberr)
|
||
|
||
`gpu1-peak.csv`: per-process GPU 1 footprint, nvidia-smi at 0.2 s, container PID resolved via
|
||
docker inspect. Largest accepted request (16 noul questions, 7,167-token prompt) found by binary
|
||
search. **Peak 9,866 MiB ≤ 9,876 budget** (10 MiB inside).
|
||
|
||
## 0.1.2 (audit fixes, same day)
|
||
|
||
infra-ops audit PASS with two low findings, both fixed in 0.1.2:
|
||
1. doc-only: the contract's response example showed `model` as an object; the wire returns the
|
||
`"<name>@<revision>"` string. Example and prose aligned.
|
||
2. `images: []` / `null` now count as ABSENT (only a non-empty value is 422) — a client that
|
||
always sends the field must not be rejected for nothing. Live: `[]`→200, `null`→200,
|
||
`["a.png"]`→422.
|
||
|
||
Re-run on 0.1.2 (jb2-* files, the 0.1.1 artifacts kept with their version suffix): all **202/231**,
|
||
hard **83/111**, **0 changed rows** across 924. 124 tests green. Rollback: 0.1.1.
|
||
|
||
## 0.1.3 — triton-cache volume + warm-up script (same day, Prime task)
|
||
|
||
Image change only (Dockerfile creates /tmp/triton-cache owned 10001; compose mounts the named
|
||
volume there). Measured facts in the stack README "Cold-shape latency": bucket width 2,048 PROVEN,
|
||
16 buckets, cold full warm-up 109 s, warm re-run 17 s. Acceptance here:
|
||
- warmup-cold.txt / warmup-warm.txt: the two runs; every bucket in-bucket, cold 109.3 s vs 17.2 s.
|
||
- random-validation.txt: 12 random sizes 8.3k–31.8k after warm-up, worst wall 2.09 s — no cold call.
|
||
- Recreate survival: force-recreate, then a warmed 32k call answered in 2.11 s; volume holds
|
||
3,200+ files owned by uid 10001 (new files newer than the image build = the volume is in use).
|
||
- Memory: warm 32k peak 15,218 MiB (budget 15,220); a COLD autotune touched 15,224 once, during
|
||
the pre-deploy 0.1.2-era measurement — see README caveat. gpu1-cold-run.csv, gpu1-32k-warm.csv.
|
||
- JevBench on 0.1.3: all 202/231, hard 83/111, 0 diffs / 924 (jb-results-0.1.3.jsonl).
|
||
- /decide positive control: answered (yes, 0.969).
|