Files
esh-pfi-infrastructure/services/intern-decision-serve/acceptance/systemone-2026-09-30/README.md
T
vh 1866c003e8 fix(intern-decision-serve): 0.1.2 treats empty/null images as absent
Audit finding on 0.1.1: a Jev client that always sends an images array with no
images was rejected for nothing. Only a non-empty value is 422 now. Also aligns
the contract's response example with the wire (model is the name@revision
string, not an object).

Re-accepted live: images []/null -> 200, ["a.png"] -> 422; JevBench all 202/231,
hard 83/111, 0 changed rows across the bench's r1..r4 (924).
2026-09-30 13:04:31 -07:00

44 lines
2.4 KiB
Markdown

# /v1/systemone acceptance — 0.1.1 on fv-ml1, 2026-09-30
Live service: `http://intern-decision.fv.internal:8033`, image `intern-decision-serve:0.1.1`
(built in `/opt/docker/src/intern-decision-serve-0.1.1`; 0.1.0 kept for rollback).
## JevBench (the pass/fail line)
JevBench v1.2.16 (5e95f23), its own `typesafe` adapter, over the 231 public items:
TYPESAFE_ENDPOINT=http://intern-decision.fv.internal:8033 TYPESAFE_API_KEY=$(secret get intern-decision/api-token) \
PYTHONPATH=/tmp/jevbench python3 -m jevbench.cli run --tasks datasets/public/{easy,original,hard}.jsonl \
--adapter typesafe --results ... --ledger ... --raw-dir ... --price-in-per-m 0 --price-out-per-m 0
- **all 202/231, hard 83/111** — matches the expected numbers exactly.
- Row-by-row vs the bench's own runs r1..r4 (`bench-jev-2026-09-30/raw/out/intern-decision-4b-native/`):
**0 changed answers/probabilities across 924 rows.**
- Artifacts here: jb-results.jsonl, jb-summary.json, jb-ledger.jsonl, jb-manifest.json.
- The first-run finding: the response `model` must be a string — the runner hashes it into
its manifest and a dict crashed the manifest step. Fixed before the run above.
## Controls (live)
- wrong token → 401; 17 questions → 422 ("never chunked"); `images` → 422 "images not supported".
- token limit at the boundary: 7,167 → 200, **7,169 → 422** (binary search; "one token over").
- `/decide/shared` positive control: 2 decisions, 1 call, fields q1/q2 — unchanged.
## Memory (GPU 1 is packed beside scriberr)
`gpu1-peak.csv`: per-process GPU 1 footprint, nvidia-smi at 0.2 s, container PID resolved via
docker inspect. Largest accepted request (16 noul questions, 7,167-token prompt) found by binary
search. **Peak 9,866 MiB ≤ 9,876 budget** (10 MiB inside).
## 0.1.2 (audit fixes, same day)
infra-ops audit PASS with two low findings, both fixed in 0.1.2:
1. doc-only: the contract's response example showed `model` as an object; the wire returns the
`"<name>@<revision>"` string. Example and prose aligned.
2. `images: []` / `null` now count as ABSENT (only a non-empty value is 422) — a client that
always sends the field must not be rejected for nothing. Live: `[]`→200, `null`→200,
`["a.png"]`→422.
Re-run on 0.1.2 (jb2-* files, the 0.1.1 artifacts kept with their version suffix): all **202/231**,
hard **83/111**, **0 changed rows** across 924. 124 tests green. Rollback: 0.1.1.