feat(intern-decision-serve): 0.1.1 adds POST /v1/systemone (Jev wire shape)

Straight passthrough to the checkpoint's own DecisionEngine.predict — never the
semif mapping, whose different prompt would change the answers. Reuses the one
inference thread, bearer auth, MAX_QUEUE, VRAM cap and error envelope; no new
concurrency. 1..16 questions in ONE call (never chunked: Jev questions share a
prompt); images 422; over MAX_TOKENS 422 before the forward. /health advertises
the surface. Response 'model' is a string name@revision (JevBench's runner
hashes it; a dict broke its manifest step).

Acceptance on the live service (see acceptance/systemone-2026-09-30/): JevBench
v1.2.16 typesafe adapter over the 231 public items scores all 202/231, hard
83/111, with 0 changed answers across all 924 rows of the bench's own r1..r4;
controls 401/422x3 (token boundary proven at 7168 pass / 7169 refuse); GPU 1
per-process peak 9,866 MiB under the largest accepted request (budget 9,876);
/decide/shared positive control unchanged. 120 tests green.
This commit is contained in:
vh
2026-09-30 12:58:19 -07:00
parent 579b1f5d42
commit ff552abf1f
13 changed files with 3449 additions and 13 deletions
@@ -32,6 +32,7 @@ Every POST takes and returns JSON and needs `Authorization: Bearer <token>`. `GE
| GET | `/health` | none | 200 `{status: "ok", model, vram_cap_gib, max_tokens, max_decisions, max_questions_per_call: 16, chunking, workloads: []}` |
| POST | `/decide` | `{id, state, question, options[2..16], orderings?, workload?}` | 200 one decision result |
| POST | `/decide/shared` | `{state, decisions: [{id, question, options, orderings?}], workload?}` | 200 `{results: [...], timing: {...}}`, results in request order |
| POST | `/v1/systemone` | `{state, model?, questions: {"<id>": {type, instructions, criteria?}}, images? (rejected)}` | 200 the engine's own response (`answers`, `usage`, `model`) — see § POST /v1/systemone |
`options` items are `{id, description}`. Validation is SemIf's, re-stated here because SemIf is
gone. `id` and `question` are nonempty strings. `state` is a nonempty string, object or array,
@@ -39,6 +40,49 @@ and must be finite JSON. There are 2..16 options, each with a string `id` and a
`description`, and option ids are unique within a decision. Decision ids are unique within a
request. A violation is a 422.
## POST /v1/systemone (Jev compatibility, added 0.1.1)
The TypeSafe/Jev wire shape, served natively: the request goes **straight to the engine's own
`DecisionEngine.predict`** — the checkpoint's documented Jev format — and never through the semif
mapping below. The two surfaces use different prompts; routing one through the other would change
the answers.
```
request {"state": ..., "model": "<ignored>",
"questions": {"<id>": {"type": "noul"|"choice"|"score", "instructions": ..., "criteria": ...}},
"images": <rejected if present>}
response {"answers": {"<id>": {"type": "noul", "noul": p, ...}
| {"type": "choice", "choice": label, "probabilities": {...}, ...}
| {"type": "score", "score": v, "probabilities": {...}, ...}},
"usage": {...}, "model": {"name": ..., "revision": ...}}
```
- `model` is accepted with **any** value (Jev clients send `"jev-latest"`) and ignored. The
response `model` is a **string** naming the model actually serving and its pinned revision:
`"<name>@<revision>"` (JevBench's runner hashes the value into its manifest; a dict is not
hashable and broke its manifest step in the 0.1.1 first run).
- The `answers` block is the engine's own response, unmodified: probabilities are the same
temperature-scaled numbers `/decide`'s `native` carries. `usage` is whatever the engine
reports (token counts); if it reports nothing, `{}`. Extra engine fields (`timing`,
`calibration`, `backend`) ride along — additive fields Jev clients ignore.
- **Limits, all enforced before the forward pass:** 1..16 questions per request (NO chunking —
in Jev semantics the questions of one call share one prompt and chunking would silently
change that); 422 with a clear message above 16. `images` present → 422 `images not
supported` (the vision tower is dropped, INV-7). Over `MAX_TOKENS` → 422, the model's own
pre-forward check. The same bearer token (401), the same single inference thread, the same
`MAX_QUEUE` (429), `MAX_BODY_BYTES` (413), VRAM cap with 503 `out_of_memory` + recovery.
- Per-question shape validation (type must be noul/choice/score; choice criteria an object;
score criteria a list/object; 1..16 options; nonempty state; unique field names) is the
engine's own `validate_request`; its `ValueError`s map to 422 `invalid_request` like anywhere
else. A `noul` question needs no `criteria`.
- **What is NOT verified:** this covers the subset of TypeSafe's API that JevBench's `typesafe`
adapter exercises (231 public items, 4 repeats, bench 2026-09-30). Other Jev request features
(batching semantics beyond one prompt, streaming, price headers, `images` serving) are not
implemented and not verified against the closed API.
- **Discovery:** `/health` lists the surface: `endpoints` names all three paths, and
`systemone` reports `{max_questions: 16, chunking: "none", images: "not supported",
max_tokens: <MAX_TOKENS>}`.
## Mapping onto the model (the seam)
- **One call** is one `predict()` request: