Files
esh-pfi-infrastructure/services/intern-decision-serve/acceptance/systemone-2026-09-30/README.md
T
vh ff552abf1f feat(intern-decision-serve): 0.1.1 adds POST /v1/systemone (Jev wire shape)
Straight passthrough to the checkpoint's own DecisionEngine.predict — never the
semif mapping, whose different prompt would change the answers. Reuses the one
inference thread, bearer auth, MAX_QUEUE, VRAM cap and error envelope; no new
concurrency. 1..16 questions in ONE call (never chunked: Jev questions share a
prompt); images 422; over MAX_TOKENS 422 before the forward. /health advertises
the surface. Response 'model' is a string name@revision (JevBench's runner
hashes it; a dict broke its manifest step).

Acceptance on the live service (see acceptance/systemone-2026-09-30/): JevBench
v1.2.16 typesafe adapter over the 231 public items scores all 202/231, hard
83/111, with 0 changed answers across all 924 rows of the bench's own r1..r4;
controls 401/422x3 (token boundary proven at 7168 pass / 7169 refuse); GPU 1
per-process peak 9,866 MiB under the largest accepted request (budget 9,876);
/decide/shared positive control unchanged. 120 tests green.
2026-09-30 12:58:19 -07:00

1.8 KiB

/v1/systemone acceptance — 0.1.1 on fv-ml1, 2026-09-30

Live service: http://intern-decision.fv.internal:8033, image intern-decision-serve:0.1.1 (built in /opt/docker/src/intern-decision-serve-0.1.1; 0.1.0 kept for rollback).

JevBench (the pass/fail line)

JevBench v1.2.16 (5e95f23), its own typesafe adapter, over the 231 public items:

TYPESAFE_ENDPOINT=http://intern-decision.fv.internal:8033 TYPESAFE_API_KEY=$(secret get intern-decision/api-token) \
PYTHONPATH=/tmp/jevbench python3 -m jevbench.cli run --tasks datasets/public/{easy,original,hard}.jsonl \
  --adapter typesafe --results ... --ledger ... --raw-dir ... --price-in-per-m 0 --price-out-per-m 0
  • all 202/231, hard 83/111 — matches the expected numbers exactly.
  • Row-by-row vs the bench's own runs r1..r4 (bench-jev-2026-09-30/raw/out/intern-decision-4b-native/): 0 changed answers/probabilities across 924 rows.
  • Artifacts here: jb-results.jsonl, jb-summary.json, jb-ledger.jsonl, jb-manifest.json.
  • The first-run finding: the response model must be a string — the runner hashes it into its manifest and a dict crashed the manifest step. Fixed before the run above.

Controls (live)

  • wrong token → 401; 17 questions → 422 ("never chunked"); images → 422 "images not supported".
  • token limit at the boundary: 7,167 → 200, 7,169 → 422 (binary search; "one token over").
  • /decide/shared positive control: 2 decisions, 1 call, fields q1/q2 — unchanged.

Memory (GPU 1 is packed beside scriberr)

gpu1-peak.csv: per-process GPU 1 footprint, nvidia-smi at 0.2 s, container PID resolved via docker inspect. Largest accepted request (16 noul questions, 7,167-token prompt) found by binary search. Peak 9,866 MiB ≤ 9,876 budget (10 MiB inside).