Files
esh-pfi-infrastructure/services/intern-decision-serve/intern-decision-serve.contract.md
T
vh ff552abf1f feat(intern-decision-serve): 0.1.1 adds POST /v1/systemone (Jev wire shape)
Straight passthrough to the checkpoint's own DecisionEngine.predict — never the
semif mapping, whose different prompt would change the answers. Reuses the one
inference thread, bearer auth, MAX_QUEUE, VRAM cap and error envelope; no new
concurrency. 1..16 questions in ONE call (never chunked: Jev questions share a
prompt); images 422; over MAX_TOKENS 422 before the forward. /health advertises
the surface. Response 'model' is a string name@revision (JevBench's runner
hashes it; a dict broke its manifest step).

Acceptance on the live service (see acceptance/systemone-2026-09-30/): JevBench
v1.2.16 typesafe adapter over the 231 public items scores all 202/231, hard
83/111, with 0 changed answers across all 924 rows of the bench's own r1..r4;
controls 401/422x3 (token boundary proven at 7168 pass / 7169 refuse); GPU 1
per-process peak 9,866 MiB under the largest accepted request (budget 9,876);
/decide/shared positive control unchanged. 120 tests green.
2026-09-30 12:58:19 -07:00

333 lines
21 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
title: intern-decision-serve
kind: module-contract
status: draft
owner: infra-ops
created: 2026-09-30
replaces: semif-serve 0.1.4 (services/semif-serve/semif-serve.contract.md), external surface kept
depends_on:
- internlm/Intern-Decision-4B at revision 0e5e6aa7d6d750e2b1504ba11a8136cb58aeb3cd, BF16 (Apache-2.0)
- that snapshot's own inference.py (DecisionEngine.predict), sha256 c904e2c67ca0775621a22375ee373d2ba30b52117cda870c6c9ef74143b29863
- torch 2.10.0+cu128, transformers 5.17.0, flash-linear-attention 0.5.2, causal-conv1d 1.7.0 (the 2026-09-30 bench stack)
---
# intern-decision-serve: Intern-Decision-4B behind semif-serve's HTTP surface
## Purpose
Prime, 2026-09-30: "replace semif with intern-decision now". The bench
(`docs/pfi/jev-candidates-bench-2026-09-30.md`) picked Intern-Decision-4B on its own runtime.
This service loads the model once and scores every request with the checkpoint's **own**
`inference.py` (`DecisionEngine.predict`). It re-implements neither the prompt nor the readout.
It keeps semif-serve's external surface, so a caller written for semif-serve works unchanged.
It only maps semif-shaped requests onto the model's Jev request schema and maps the answers back.
Every deliberate difference is listed under **Deltas from semif-serve**.
## Endpoints (same as semif-serve)
Every POST takes and returns JSON and needs `Authorization: Bearer <token>`. `GET /health` is open.
| method | path | body | success |
|---|---|---|---|
| GET | `/health` | none | 200 `{status: "ok", model, vram_cap_gib, max_tokens, max_decisions, max_questions_per_call: 16, chunking, workloads: []}` |
| POST | `/decide` | `{id, state, question, options[2..16], orderings?, workload?}` | 200 one decision result |
| POST | `/decide/shared` | `{state, decisions: [{id, question, options, orderings?}], workload?}` | 200 `{results: [...], timing: {...}}`, results in request order |
| POST | `/v1/systemone` | `{state, model?, questions: {"<id>": {type, instructions, criteria?}}, images? (rejected)}` | 200 the engine's own response (`answers`, `usage`, `model`) — see § POST /v1/systemone |
`options` items are `{id, description}`. Validation is SemIf's, re-stated here because SemIf is
gone. `id` and `question` are nonempty strings. `state` is a nonempty string, object or array,
and must be finite JSON. There are 2..16 options, each with a string `id` and a string
`description`, and option ids are unique within a decision. Decision ids are unique within a
request. A violation is a 422.
## POST /v1/systemone (Jev compatibility, added 0.1.1)
The TypeSafe/Jev wire shape, served natively: the request goes **straight to the engine's own
`DecisionEngine.predict`** — the checkpoint's documented Jev format — and never through the semif
mapping below. The two surfaces use different prompts; routing one through the other would change
the answers.
```
request {"state": ..., "model": "<ignored>",
"questions": {"<id>": {"type": "noul"|"choice"|"score", "instructions": ..., "criteria": ...}},
"images": <rejected if present>}
response {"answers": {"<id>": {"type": "noul", "noul": p, ...}
| {"type": "choice", "choice": label, "probabilities": {...}, ...}
| {"type": "score", "score": v, "probabilities": {...}, ...}},
"usage": {...}, "model": {"name": ..., "revision": ...}}
```
- `model` is accepted with **any** value (Jev clients send `"jev-latest"`) and ignored. The
response `model` is a **string** naming the model actually serving and its pinned revision:
`"<name>@<revision>"` (JevBench's runner hashes the value into its manifest; a dict is not
hashable and broke its manifest step in the 0.1.1 first run).
- The `answers` block is the engine's own response, unmodified: probabilities are the same
temperature-scaled numbers `/decide`'s `native` carries. `usage` is whatever the engine
reports (token counts); if it reports nothing, `{}`. Extra engine fields (`timing`,
`calibration`, `backend`) ride along — additive fields Jev clients ignore.
- **Limits, all enforced before the forward pass:** 1..16 questions per request (NO chunking —
in Jev semantics the questions of one call share one prompt and chunking would silently
change that); 422 with a clear message above 16. `images` present → 422 `images not
supported` (the vision tower is dropped, INV-7). Over `MAX_TOKENS` → 422, the model's own
pre-forward check. The same bearer token (401), the same single inference thread, the same
`MAX_QUEUE` (429), `MAX_BODY_BYTES` (413), VRAM cap with 503 `out_of_memory` + recovery.
- Per-question shape validation (type must be noul/choice/score; choice criteria an object;
score criteria a list/object; 1..16 options; nonempty state; unique field names) is the
engine's own `validate_request`; its `ValueError`s map to 422 `invalid_request` like anywhere
else. A `noul` question needs no `criteria`.
- **What is NOT verified:** this covers the subset of TypeSafe's API that JevBench's `typesafe`
adapter exercises (231 public items, 4 repeats, bench 2026-09-30). Other Jev request features
(batching semantics beyond one prompt, streaming, price headers, `images` serving) are not
implemented and not verified against the closed API.
- **Discovery:** `/health` lists the surface: `endpoints` names all three paths, and
`systemone` reports `{max_questions: 16, chunking: "none", images: "not supported",
max_tokens: <MAX_TOKENS>}`.
## Mapping onto the model (the seam)
- **One call** is one `predict()` request:
`{"state": state, "questions": {<field>: {"type": "choice", "instructions": question, "criteria": {option.id: option.description, ...}}}}`.
The criteria keep the caller's option order.
- **A one-question call is exactly the bench's native request.** Acceptance showed it
bit-identical, row for row.
- **A multi-question call differs from the bench's multifield run only in the field names.**
Those are positional here and were the decision names in the bench. Measured: 1 of 84 Wyrd
rows differs (78/84 against 77/84).
- **Field names are positional:** `q` when a call carries one question, and `q1`..`qN` in
request order when it carries several. Decision ids never reach the prompt. **Option ids do:**
the model prints `A = <id>: <description>`.
- **`/decide`** is one call with one question.
- **`/decide/shared`** is packed into calls of **at most 16 questions**, the model's own limit
(`validate_request`). The packing is greedy and in request order: questions 1–16, then 17–32,
and so on. Every call of a request runs back to back under the inference lock. `/health` reports
`max_questions_per_call: 16` and the rule in `chunking`.
- **Orderings.** A decision may set `orderings: "none" | "rotations" | "all"`, with semif's
meaning. `rotations` gives the n cyclic shifts, the caller's order first. `all` gives the n!
permutations, the caller's order first, and is 422 above 4 options. Ordering k of every decision
that has more than k orderings forms **wave k**. Each wave is packed as above. So wave 0 is the
request exactly as written, and no prompt ever holds two orderings of the same decision. Every
ordering counts toward `MAX_DECISIONS`.
- **The result for one ordering** (and for every plain decision):
```
{id, option_ids, probabilities, top, confidence, calibration, native,
input_tokens, prompt_sha256, prompt_version, model, readout, probability_status,
call: {index, field, questions}}
```
- `native` is the answer object `predict()` returned for that field, unchanged.
- `probabilities` is `native.probabilities` listed in `option_ids` order.
- `top` is `native.choice` and `confidence` is `native.confidence`.
- `calibration` is the `calibration` object from `predict()`.
- `input_tokens` and `prompt_sha256` describe the whole call.
- `/decide` results (plain) also carry `total_seconds` and `forward_seconds`. `forward_seconds`
is `predict()`'s `timing.inference_ms` / 1000.
- **Averaged result:** semif's shape, `{id, option_ids, combined: {method, orderings, probabilities, top, agreement, spread}, orderings: [...]}`.
- The per-ordering ids are `<id>#o<k>`.
- `combined.probabilities` renormalises the per-option mean of `log p`. A probability of
exactly 0 is floored at 1e-300 before the log.
- `combined.top` is the first maximum in the caller's order.
- `agreement` is the share of orderings whose `top` equals `combined.top`.
- `spread` holds each option's min and max `probabilities` across orderings.
- Temperature scaling is one monotone transform per ordering, so `combined.top` and
`agreement` are what combining T=1 scores would give.
- **`/decide/shared` timing:** `{total_seconds, batch_size (orderings scored), calls, questions_per_call: [...], input_tokens: [...], inference_seconds}`.
## Invariants
- **INV-1 pass-through.** Every number in `native`, `probabilities`, `top`, `confidence`,
`calibration` and `input_tokens` is what `predict()` returned. The wrapper only re-keys it.
If an answer lacks one of the decision's option ids, that is a 500 `scoring_failed`, never a
guess. So is any failure while building the response, a non-finite number included. That stays
inside the error envelope and is never a 422 or a bare 500.
- **INV-2 one model, one inference thread.** The model loads at startup on a dedicated
single-thread executor. The warm-ups and every later call run on that **same host thread**,
never on the event loop's threadpool. A lock also serialises each request's calls, so all of a
request's chunks run inside one hold. The app runs one worker, and `/health` answers during a
call.
- `/health` reads only allocator counters and a device name cached at load, so it adds no
per-thread CUDA state.
- **Why one thread:** torch keeps CUDA state per host thread (cuBLAS handles and workspaces),
and part of it sits outside the VRAM cap.
- **Measured on 2026-09-30, fv-ml1 GPU 3:** anyio's 40 worker threads added 252 MiB outside the
cap and 326 MiB inside it. That pushed the footprint past the GPU 1 budget.
- **INV-3 fail-closed startup.** Before the service serves, all of these must hold:
- `inference.py` in the checkpoint hashes to the pinned sha256 (it is executed code, loaded
from a data mount);
- the checkpoint path ends in `snapshots/<pinned revision>`;
- the model sits on CUDA, and torch's arch list has the card's `sm_XY`;
- one warm-up decision scores.
`device=cpu` is allowed only when set explicitly.
- **INV-4 VRAM cap.** `VRAM_CAP_GIB`, when set, becomes `torch.cuda.set_per_process_memory_fraction`
**before** the weights load.
- An OOM in any call makes the whole request a 503 `out_of_memory`. So does a RuntimeError
whose first line says "out of memory". The engine then frees the failed call's frames,
runs `gc.collect()` and `empty_cache()`, and raises unchained. The process stays up.
- A `ValueError` (the token limit, or one raised inside a forward) keeps its message and maps
to 422. It is released and raised unchained the same way.
- After every call, if reserved memory exceeds the post-warm-up baseline by more than
`RELEASE_SLACK_MIB` (default 512), the engine runs `empty_cache()`.
- Any other failure is logged with its traceback and released the same way, then raised
unchained as `ScoringFailed`.
- **INV-5 no network.** The entry point sets `HF_HUB_OFFLINE=1` and `TRANSFORMERS_OFFLINE=1`
before torch or transformers load. The weights are read from the mounted, read-only HF cache.
- **INV-6 constant-time auth.** The token is compared with `hmac.compare_digest`. It must be at
least 32 visible ASCII characters (33–126), or startup refuses it.
- **INV-7 the text-only model is the same model.**
- The service takes no images, and neither did semif. So after the first warm-up the engine
replaces the vision tower (`model.model.visual`, 0.62 GiB) with a stub that raises if it is
ever called.
- It then scores the warm-up again. It refuses to start unless the answer is bit-identical to
the first one.
- `KEEP_VISION=1` keeps the tower.
- **INV-8 honest prompt hash.** `prompt_sha256` is the sha256 of the chat-template text rendered
from `inference.compile_row(...)` with the same arguments `HFBackend.encode` uses. At startup,
that text must tokenise to exactly the `input_tokens` `predict()` reports for the warm-up.
Otherwise the service refuses to start rather than hash a prompt the model never saw.
## Limits and errors
- `MAX_TOKENS` (default 8192, the model's own default) is `DecisionEngine(max_length=...)`, and
applies to a **whole call**: the state plus all its questions. A longer call is a 422. It is
never truncated; the model raises.
- `MAX_DECISIONS` (default 64) caps the orderings scored per request, counted after expansion.
A request must score 1..max of them, else 422.
- The body may be at most `MAX_BODY_BYTES` (default 1 MiB), else 413. This is checked before
each chunk is kept. A declared `Content-Length` is trusted only as ASCII digits.
- **Admission:** at most `MAX_QUEUE` (default 32) POSTs may be queued or scoring at once. The
next one gets 429 `busy` before its body is read.
- `workload`: there is no per-workload table, and the model's own calibration always applies.
So any non-null `workload` is a 422, exactly as semif-serve behaved with its deployed empty
table.
| status | code | when |
|---|---|---|
| 401 | `unauthorized` | missing or wrong bearer |
| 413 | `request_too_large` | body over the limit |
| 422 | `invalid_request` | bad JSON or shape; a SemIf-rule violation; a model `ValueError` (token limit, reserved `<decision>` marker in the input, ...); a `workload`; `all` over 4 options; a row count outside 1..`MAX_DECISIONS` |
| 429 | `busy` | `MAX_QUEUE` requests in progress |
| 503 | `out_of_memory` | CUDA OOM in any call of the request |
| 500 | `scoring_failed` | any other model failure, including one while building the response |
The error body is `{error: {code, message}}`.
## Configuration (env, prefix `INTERN_DECISION_`)
- `API_TOKEN` is required.
- `CHECKPOINT` defaults to
`/hf/hub/models--internlm--Intern-Decision-4B/snapshots/<revision>`.
- `DEVICE` defaults to `cuda`.
- `VRAM_CAP_GIB` has no default: unset means uncapped, and when set it must be finite and > 0.
- `MAX_TOKENS`, `MAX_DECISIONS`, `MAX_BODY_BYTES`, `MAX_QUEUE` and `RELEASE_SLACK_MIB` are
integers; each must be ≥ 1, except the slack, which must be ≥ 0.
- `KEEP_VISION` is `0` or `1`.
A bad value is refused at startup with a `ValueError` naming the variable. In the stack's `.env`
on the host, the cap is the single knob `VRAM_CAP_GIB`.
## Deltas from semif-serve (deliberate)
1. **Prompt and model.** The prompt is Intern-Decision's own Jev prompt.
- **Option ids are shown to the model** (`A = <id>: <description>`); SemIf showed only the
descriptions. An option id is therefore part of the question, so give options meaningful
or neutral ids.
- In the bench negative control, the ids-in-prompt cue made the top stay on 10/144 rows
after the descriptions moved, against SemIf's 14/144.
2. **Shared requests are one prompt, not independent rows.** The questions in a call are asked
together.
- A decision's answer can depend on the other questions in its call and on their order. In
the bench, Wyrd scored 79/84 asked one decision at a time and 77/84 with a turn's 4
decisions in one prompt.
- SemIf only shared a KV prefix, so each of its rows was independent.
- Calls hold at most 16 questions, and the chunk boundaries follow request order.
3. **`probabilities` are temperature-scaled** by the checkpoint's shipped calibration
(T = 1.99241824).
- `probability_status` says so, and `calibration` carries the method and T.
- SemIf's were raw softmax, labelled uncalibrated. The argmax is the same either way.
- This is the vendor's calibration on the vendor's data, not ours.
4. **`option_logits` does not exist.** `predict()` does not expose logits. T=1 scores could only
be derived up to a constant, which would be a derived number, not logits.
5. **New fields:** `top`, `confidence`, `calibration`, `native`, `call`.
6. **`prompt_sha256` and `input_tokens` describe the whole call**, shared by every decision in
it. SemIf's were per row.
7. **`prompt_version`** names the pinned `inference.py`. **`readout`** and **`model`** describe
Intern-Decision.
8. **`workload` is always a 422, and `/health.workloads` is `[]`.** There is no per-workload
temperature table. This matches the deployed semif, whose table was empty.
9. **`/health`**: `semif_commit` is gone; the model's pins live in `model`. It adds
`max_questions_per_call` and `chunking`.
10. **`/decide/shared` timing:** SemIf's prefix-cache fields cannot exist, because there is no
prefix cache: `prefix_tokens`, `prefill_seconds`, `replicate_seconds`,
`suffix_forward_seconds`, `true_suffix_tokens`, `padded_suffix_tokens` and `encode_seconds`.
The timing adds `calls`, `questions_per_call`, `input_tokens` and `inference_seconds`.
`total_seconds` and `batch_size` keep their meaning.
11. **`MAX_TOKENS` is per call** (state plus up to 16 questions) and defaults to 8192. SemIf's
was per row and defaulted to 4096.
12. **Orderings run in waves**, one forward pass per wave per chunk. SemIf batched every
ordering in one shared forward.
13. **Ties.** Per-ordering `top` uses the model's argmax, whose exact tie goes to the smaller
option id string. SemIf used the first maximum in the ordering. `combined.top` keeps semif's
rule.
14. **Images.** The vision tower is dropped (INV-7). The surface never took images.
## Tests (TDD, fake engine: no torch, no model)
- **Auth.** A POST without the right bearer is 401 and never reaches the engine. `/health`
needs no auth. A short or non-visible-ASCII token is refused at startup.
- **Mapping.**
- `/decide` sends one call with field `q` and the criteria in the caller's order.
- `/decide/shared` sends `q1..qN` over the shared state.
- 17 questions become 2 calls (16 + 1), and 40 become 3 (16 + 16 + 8). Results come back in
request order, and `call.index` and `call.field` are right.
- Decision ids never appear in a call.
- **Result.** `probabilities` follows `option_ids`. `top`, `confidence`, `calibration`, `native`
and `input_tokens` are passed through. An answer missing an option id is a 500.
- **Orderings.**
- `rotations` puts each option in each position once, and the waves never repeat a decision
within a call.
- `all` is n!, and 422 above 4 options.
- A position bias cancels exactly.
- `agreement` and `spread` are computed from the orderings.
- A mixed request keeps plain results unchanged, and wave 0 equals the plain request.
- Orderings count toward the cap.
- **Validation (422).** Fewer than 2 or more than 16 options, duplicate option ids, duplicate
decision ids, an empty id or question, an empty or non-finite state, a `workload`, a row count
outside 1..max, malformed JSON, and a model `ValueError`.
- **Limits.** A body over the limit is 413. A queue past `MAX_QUEUE` is 429 before the body is
read.
- **Engine failures.** An engine `OutOfMemory` is 503, and any other failure is 500.
- **Concurrency.** Requests are serialised: two never overlap inside the engine, and the
chunks of one request are not interleaved with another's. `/health` answers while a call is
blocked. Every call runs on one thread, which is the thread the engine was loaded on.
- **Engine against a fake torch.**
- An OOM is re-raised unchained, and `empty_cache` runs only after the failed call's tensors
are freed.
- A RuntimeError saying "out of memory" becomes `OutOfMemory`.
- Another failure becomes `ScoringFailed`, unchained.
- A `ValueError` keeps its message but is released and raised unchained.
- A request the prompt builder rejects never reaches the model.
- A burst over baseline plus slack is released, and one at or under it is left alone.
- **Engine load (fake).**
- A wrong `inference.py` hash, or a checkpoint path that is not the pinned snapshot, refuses
to start.
- The cap is applied before the engine is constructed.
- The vision swap refuses to start when the warm-up changes.
- The prompt-hash check refuses to start on a token-count mismatch.
- **Config.** Every value is validated.
## Acceptance (fv-ml1, real model; not unit tests)
1. **Positive control:** through the service, the bench's native numbers on the pooled 259 rows
and on Wyrd, single ordering (bench: 240/259, Wyrd 79/84; floor: the pooled set resolves
±4 pts, and 0 labels moved across 4 restarts). Also a row-by-row comparison against the
bench's own rows.
2. **Negative control:** descriptions rotated one place. The top follows the moved description
(bench: 122/144 follow, 10/144 same top).
3. A-vs-A repeat stability, within the process and across restarts.
4. Latency at our shape: 21 binary criteria, and 16 criteria over the ~3,900-token state, both
server-side and from nh3-dev.
5. Resident and peak VRAM (nvidia-smi and torch), and the cap chosen from them.
6. An over-cap request is a 503, memory returns to baseline, and the service keeps answering.
7. A `/decide/shared` with more than 16 questions is chunked, and its answers equal the same
questions asked chunk by chunk.
8. 401 without the token, and 429 past the queue.