feat(intern-decision-serve): 0.1.1 adds POST /v1/systemone (Jev wire shape)
Straight passthrough to the checkpoint's own DecisionEngine.predict — never the semif mapping, whose different prompt would change the answers. Reuses the one inference thread, bearer auth, MAX_QUEUE, VRAM cap and error envelope; no new concurrency. 1..16 questions in ONE call (never chunked: Jev questions share a prompt); images 422; over MAX_TOKENS 422 before the forward. /health advertises the surface. Response 'model' is a string name@revision (JevBench's runner hashes it; a dict broke its manifest step). Acceptance on the live service (see acceptance/systemone-2026-09-30/): JevBench v1.2.16 typesafe adapter over the 231 public items scores all 202/231, hard 83/111, with 0 changed answers across all 924 rows of the bench's own r1..r4; controls 401/422x3 (token boundary proven at 7168 pass / 7169 refuse); GPU 1 per-process peak 9,866 MiB under the largest accepted request (budget 9,876); /decide/shared positive control unchanged. 120 tests green.
This commit is contained in:
@@ -32,6 +32,7 @@ Every POST takes and returns JSON and needs `Authorization: Bearer <token>`. `GE
|
||||
| GET | `/health` | none | 200 `{status: "ok", model, vram_cap_gib, max_tokens, max_decisions, max_questions_per_call: 16, chunking, workloads: []}` |
|
||||
| POST | `/decide` | `{id, state, question, options[2..16], orderings?, workload?}` | 200 one decision result |
|
||||
| POST | `/decide/shared` | `{state, decisions: [{id, question, options, orderings?}], workload?}` | 200 `{results: [...], timing: {...}}`, results in request order |
|
||||
| POST | `/v1/systemone` | `{state, model?, questions: {"<id>": {type, instructions, criteria?}}, images? (rejected)}` | 200 the engine's own response (`answers`, `usage`, `model`) — see § POST /v1/systemone |
|
||||
|
||||
`options` items are `{id, description}`. Validation is SemIf's, re-stated here because SemIf is
|
||||
gone. `id` and `question` are nonempty strings. `state` is a nonempty string, object or array,
|
||||
@@ -39,6 +40,49 @@ and must be finite JSON. There are 2..16 options, each with a string `id` and a
|
||||
`description`, and option ids are unique within a decision. Decision ids are unique within a
|
||||
request. A violation is a 422.
|
||||
|
||||
## POST /v1/systemone (Jev compatibility, added 0.1.1)
|
||||
|
||||
The TypeSafe/Jev wire shape, served natively: the request goes **straight to the engine's own
|
||||
`DecisionEngine.predict`** — the checkpoint's documented Jev format — and never through the semif
|
||||
mapping below. The two surfaces use different prompts; routing one through the other would change
|
||||
the answers.
|
||||
|
||||
```
|
||||
request {"state": ..., "model": "<ignored>",
|
||||
"questions": {"<id>": {"type": "noul"|"choice"|"score", "instructions": ..., "criteria": ...}},
|
||||
"images": <rejected if present>}
|
||||
response {"answers": {"<id>": {"type": "noul", "noul": p, ...}
|
||||
| {"type": "choice", "choice": label, "probabilities": {...}, ...}
|
||||
| {"type": "score", "score": v, "probabilities": {...}, ...}},
|
||||
"usage": {...}, "model": {"name": ..., "revision": ...}}
|
||||
```
|
||||
|
||||
- `model` is accepted with **any** value (Jev clients send `"jev-latest"`) and ignored. The
|
||||
response `model` is a **string** naming the model actually serving and its pinned revision:
|
||||
`"<name>@<revision>"` (JevBench's runner hashes the value into its manifest; a dict is not
|
||||
hashable and broke its manifest step in the 0.1.1 first run).
|
||||
- The `answers` block is the engine's own response, unmodified: probabilities are the same
|
||||
temperature-scaled numbers `/decide`'s `native` carries. `usage` is whatever the engine
|
||||
reports (token counts); if it reports nothing, `{}`. Extra engine fields (`timing`,
|
||||
`calibration`, `backend`) ride along — additive fields Jev clients ignore.
|
||||
- **Limits, all enforced before the forward pass:** 1..16 questions per request (NO chunking —
|
||||
in Jev semantics the questions of one call share one prompt and chunking would silently
|
||||
change that); 422 with a clear message above 16. `images` present → 422 `images not
|
||||
supported` (the vision tower is dropped, INV-7). Over `MAX_TOKENS` → 422, the model's own
|
||||
pre-forward check. The same bearer token (401), the same single inference thread, the same
|
||||
`MAX_QUEUE` (429), `MAX_BODY_BYTES` (413), VRAM cap with 503 `out_of_memory` + recovery.
|
||||
- Per-question shape validation (type must be noul/choice/score; choice criteria an object;
|
||||
score criteria a list/object; 1..16 options; nonempty state; unique field names) is the
|
||||
engine's own `validate_request`; its `ValueError`s map to 422 `invalid_request` like anywhere
|
||||
else. A `noul` question needs no `criteria`.
|
||||
- **What is NOT verified:** this covers the subset of TypeSafe's API that JevBench's `typesafe`
|
||||
adapter exercises (231 public items, 4 repeats, bench 2026-09-30). Other Jev request features
|
||||
(batching semantics beyond one prompt, streaming, price headers, `images` serving) are not
|
||||
implemented and not verified against the closed API.
|
||||
- **Discovery:** `/health` lists the surface: `endpoints` names all three paths, and
|
||||
`systemone` reports `{max_questions: 16, chunking: "none", images: "not supported",
|
||||
max_tokens: <MAX_TOKENS>}`.
|
||||
|
||||
## Mapping onto the model (the seam)
|
||||
|
||||
- **One call** is one `predict()` request:
|
||||
|
||||
Reference in New Issue
Block a user