--- title: intern-decision-serve kind: module-contract status: draft owner: infra-ops created: 2026-09-30 replaces: semif-serve 0.1.4 (services/semif-serve/semif-serve.contract.md), external surface kept depends_on: - internlm/Intern-Decision-4B at revision 0e5e6aa7d6d750e2b1504ba11a8136cb58aeb3cd, BF16 (Apache-2.0) - that snapshot's own inference.py (DecisionEngine.predict), sha256 c904e2c67ca0775621a22375ee373d2ba30b52117cda870c6c9ef74143b29863 - torch 2.10.0+cu128, transformers 5.17.0, flash-linear-attention 0.5.2, causal-conv1d 1.7.0 (the 2026-09-30 bench stack) --- # intern-decision-serve: Intern-Decision-4B behind semif-serve's HTTP surface ## Purpose Prime, 2026-09-30: "replace semif with intern-decision now". The bench (`docs/pfi/jev-candidates-bench-2026-09-30.md`) picked Intern-Decision-4B on its own runtime. This service loads the model once and scores every request with the checkpoint's **own** `inference.py` (`DecisionEngine.predict`). It re-implements neither the prompt nor the readout. It keeps semif-serve's external surface, so a caller written for semif-serve works unchanged. It only maps semif-shaped requests onto the model's Jev request schema and maps the answers back. Every deliberate difference is listed under **Deltas from semif-serve**. ## Endpoints (same as semif-serve) Every POST takes and returns JSON and needs `Authorization: Bearer `. `GET /health` is open. | method | path | body | success | |---|---|---|---| | GET | `/health` | none | 200 `{status: "ok", model, vram_cap_gib, max_tokens, max_decisions, max_questions_per_call: 16, chunking, workloads: []}` | | POST | `/decide` | `{id, state, question, options[2..16], orderings?, workload?}` | 200 one decision result | | POST | `/decide/shared` | `{state, decisions: [{id, question, options, orderings?}], workload?}` | 200 `{results: [...], timing: {...}}`, results in request order | | POST | `/v1/systemone` | `{state, model?, questions: {"": {type, instructions, criteria?}}, images? (rejected)}` | 200 the engine's own response (`answers`, `usage`, `model`) — see § POST /v1/systemone | `options` items are `{id, description}`. Validation is SemIf's, re-stated here because SemIf is gone. `id` and `question` are nonempty strings. `state` is a nonempty string, object or array, and must be finite JSON. There are 2..16 options, each with a string `id` and a string `description`, and option ids are unique within a decision. Decision ids are unique within a request. A violation is a 422. ## POST /v1/systemone (Jev compatibility, added 0.1.1) The TypeSafe/Jev wire shape, served natively: the request goes **straight to the engine's own `DecisionEngine.predict`** — the checkpoint's documented Jev format — and never through the semif mapping below. The two surfaces use different prompts; routing one through the other would change the answers. ``` request {"state": ..., "model": "", "questions": {"": {"type": "noul"|"choice"|"score", "instructions": ..., "criteria": ...}}, "images": } response {"answers": {"": {"type": "noul", "noul": p, ...} | {"type": "choice", "choice": label, "probabilities": {...}, ...} | {"type": "score", "score": v, "probabilities": {...}, ...}}, "usage": {...}, "model": {"name": ..., "revision": ...}} ``` - `model` is accepted with **any** value (Jev clients send `"jev-latest"`) and ignored. The response `model` is a **string** naming the model actually serving and its pinned revision: `"@"` (JevBench's runner hashes the value into its manifest; a dict is not hashable and broke its manifest step in the 0.1.1 first run). - The `answers` block is the engine's own response, unmodified: probabilities are the same temperature-scaled numbers `/decide`'s `native` carries. `usage` is whatever the engine reports (token counts); if it reports nothing, `{}`. Extra engine fields (`timing`, `calibration`, `backend`) ride along — additive fields Jev clients ignore. - **Limits, all enforced before the forward pass:** 1..16 questions per request (NO chunking — in Jev semantics the questions of one call share one prompt and chunking would silently change that); 422 with a clear message above 16. `images` present → 422 `images not supported` (the vision tower is dropped, INV-7). Over `MAX_TOKENS` → 422, the model's own pre-forward check. The same bearer token (401), the same single inference thread, the same `MAX_QUEUE` (429), `MAX_BODY_BYTES` (413), VRAM cap with 503 `out_of_memory` + recovery. - Per-question shape validation (type must be noul/choice/score; choice criteria an object; score criteria a list/object; 1..16 options; nonempty state; unique field names) is the engine's own `validate_request`; its `ValueError`s map to 422 `invalid_request` like anywhere else. A `noul` question needs no `criteria`. - **What is NOT verified:** this covers the subset of TypeSafe's API that JevBench's `typesafe` adapter exercises (231 public items, 4 repeats, bench 2026-09-30). Other Jev request features (batching semantics beyond one prompt, streaming, price headers, `images` serving) are not implemented and not verified against the closed API. - **Discovery:** `/health` lists the surface: `endpoints` names all three paths, and `systemone` reports `{max_questions: 16, chunking: "none", images: "not supported", max_tokens: }`. ## Mapping onto the model (the seam) - **One call** is one `predict()` request: `{"state": state, "questions": {: {"type": "choice", "instructions": question, "criteria": {option.id: option.description, ...}}}}`. The criteria keep the caller's option order. - **A one-question call is exactly the bench's native request.** Acceptance showed it bit-identical, row for row. - **A multi-question call differs from the bench's multifield run only in the field names.** Those are positional here and were the decision names in the bench. Measured: 1 of 84 Wyrd rows differs (78/84 against 77/84). - **Field names are positional:** `q` when a call carries one question, and `q1`..`qN` in request order when it carries several. Decision ids never reach the prompt. **Option ids do:** the model prints `A = : `. - **`/decide`** is one call with one question. - **`/decide/shared`** is packed into calls of **at most 16 questions**, the model's own limit (`validate_request`). The packing is greedy and in request order: questions 1–16, then 17–32, and so on. Every call of a request runs back to back under the inference lock. `/health` reports `max_questions_per_call: 16` and the rule in `chunking`. - **Orderings.** A decision may set `orderings: "none" | "rotations" | "all"`, with semif's meaning. `rotations` gives the n cyclic shifts, the caller's order first. `all` gives the n! permutations, the caller's order first, and is 422 above 4 options. Ordering k of every decision that has more than k orderings forms **wave k**. Each wave is packed as above. So wave 0 is the request exactly as written, and no prompt ever holds two orderings of the same decision. Every ordering counts toward `MAX_DECISIONS`. - **The result for one ordering** (and for every plain decision): ``` {id, option_ids, probabilities, top, confidence, calibration, native, input_tokens, prompt_sha256, prompt_version, model, readout, probability_status, call: {index, field, questions}} ``` - `native` is the answer object `predict()` returned for that field, unchanged. - `probabilities` is `native.probabilities` listed in `option_ids` order. - `top` is `native.choice` and `confidence` is `native.confidence`. - `calibration` is the `calibration` object from `predict()`. - `input_tokens` and `prompt_sha256` describe the whole call. - `/decide` results (plain) also carry `total_seconds` and `forward_seconds`. `forward_seconds` is `predict()`'s `timing.inference_ms` / 1000. - **Averaged result:** semif's shape, `{id, option_ids, combined: {method, orderings, probabilities, top, agreement, spread}, orderings: [...]}`. - The per-ordering ids are `#o`. - `combined.probabilities` renormalises the per-option mean of `log p`. A probability of exactly 0 is floored at 1e-300 before the log. - `combined.top` is the first maximum in the caller's order. - `agreement` is the share of orderings whose `top` equals `combined.top`. - `spread` holds each option's min and max `probabilities` across orderings. - Temperature scaling is one monotone transform per ordering, so `combined.top` and `agreement` are what combining T=1 scores would give. - **`/decide/shared` timing:** `{total_seconds, batch_size (orderings scored), calls, questions_per_call: [...], input_tokens: [...], inference_seconds}`. ## Invariants - **INV-1 pass-through.** Every number in `native`, `probabilities`, `top`, `confidence`, `calibration` and `input_tokens` is what `predict()` returned. The wrapper only re-keys it. If an answer lacks one of the decision's option ids, that is a 500 `scoring_failed`, never a guess. So is any failure while building the response, a non-finite number included. That stays inside the error envelope and is never a 422 or a bare 500. - **INV-2 one model, one inference thread.** The model loads at startup on a dedicated single-thread executor. The warm-ups and every later call run on that **same host thread**, never on the event loop's threadpool. A lock also serialises each request's calls, so all of a request's chunks run inside one hold. The app runs one worker, and `/health` answers during a call. - `/health` reads only allocator counters and a device name cached at load, so it adds no per-thread CUDA state. - **Why one thread:** torch keeps CUDA state per host thread (cuBLAS handles and workspaces), and part of it sits outside the VRAM cap. - **Measured on 2026-09-30, fv-ml1 GPU 3:** anyio's 40 worker threads added 252 MiB outside the cap and 326 MiB inside it. That pushed the footprint past the GPU 1 budget. - **INV-3 fail-closed startup.** Before the service serves, all of these must hold: - `inference.py` in the checkpoint hashes to the pinned sha256 (it is executed code, loaded from a data mount); - the checkpoint path ends in `snapshots/`; - the model sits on CUDA, and torch's arch list has the card's `sm_XY`; - one warm-up decision scores. `device=cpu` is allowed only when set explicitly. - **INV-4 VRAM cap.** `VRAM_CAP_GIB`, when set, becomes `torch.cuda.set_per_process_memory_fraction` **before** the weights load. - An OOM in any call makes the whole request a 503 `out_of_memory`. So does a RuntimeError whose first line says "out of memory". The engine then frees the failed call's frames, runs `gc.collect()` and `empty_cache()`, and raises unchained. The process stays up. - A `ValueError` (the token limit, or one raised inside a forward) keeps its message and maps to 422. It is released and raised unchained the same way. - After every call, if reserved memory exceeds the post-warm-up baseline by more than `RELEASE_SLACK_MIB` (default 512), the engine runs `empty_cache()`. - Any other failure is logged with its traceback and released the same way, then raised unchained as `ScoringFailed`. - **INV-5 no network.** The entry point sets `HF_HUB_OFFLINE=1` and `TRANSFORMERS_OFFLINE=1` before torch or transformers load. The weights are read from the mounted, read-only HF cache. - **INV-6 constant-time auth.** The token is compared with `hmac.compare_digest`. It must be at least 32 visible ASCII characters (33–126), or startup refuses it. - **INV-7 the text-only model is the same model.** - The service takes no images, and neither did semif. So after the first warm-up the engine replaces the vision tower (`model.model.visual`, 0.62 GiB) with a stub that raises if it is ever called. - It then scores the warm-up again. It refuses to start unless the answer is bit-identical to the first one. - `KEEP_VISION=1` keeps the tower. - **INV-8 honest prompt hash.** `prompt_sha256` is the sha256 of the chat-template text rendered from `inference.compile_row(...)` with the same arguments `HFBackend.encode` uses. At startup, that text must tokenise to exactly the `input_tokens` `predict()` reports for the warm-up. Otherwise the service refuses to start rather than hash a prompt the model never saw. ## Limits and errors - `MAX_TOKENS` (default 8192, the model's own default) is `DecisionEngine(max_length=...)`, and applies to a **whole call**: the state plus all its questions. A longer call is a 422. It is never truncated; the model raises. - `MAX_DECISIONS` (default 64) caps the orderings scored per request, counted after expansion. A request must score 1..max of them, else 422. - The body may be at most `MAX_BODY_BYTES` (default 1 MiB), else 413. This is checked before each chunk is kept. A declared `Content-Length` is trusted only as ASCII digits. - **Admission:** at most `MAX_QUEUE` (default 32) POSTs may be queued or scoring at once. The next one gets 429 `busy` before its body is read. - `workload`: there is no per-workload table, and the model's own calibration always applies. So any non-null `workload` is a 422, exactly as semif-serve behaved with its deployed empty table. | status | code | when | |---|---|---| | 401 | `unauthorized` | missing or wrong bearer | | 413 | `request_too_large` | body over the limit | | 422 | `invalid_request` | bad JSON or shape; a SemIf-rule violation; a model `ValueError` (token limit, reserved `` marker in the input, ...); a `workload`; `all` over 4 options; a row count outside 1..`MAX_DECISIONS` | | 429 | `busy` | `MAX_QUEUE` requests in progress | | 503 | `out_of_memory` | CUDA OOM in any call of the request | | 500 | `scoring_failed` | any other model failure, including one while building the response | The error body is `{error: {code, message}}`. ## Configuration (env, prefix `INTERN_DECISION_`) - `API_TOKEN` is required. - `CHECKPOINT` defaults to `/hf/hub/models--internlm--Intern-Decision-4B/snapshots/`. - `DEVICE` defaults to `cuda`. - `VRAM_CAP_GIB` has no default: unset means uncapped, and when set it must be finite and > 0. - `MAX_TOKENS`, `MAX_DECISIONS`, `MAX_BODY_BYTES`, `MAX_QUEUE` and `RELEASE_SLACK_MIB` are integers; each must be ≥ 1, except the slack, which must be ≥ 0. - `KEEP_VISION` is `0` or `1`. A bad value is refused at startup with a `ValueError` naming the variable. In the stack's `.env` on the host, the cap is the single knob `VRAM_CAP_GIB`. ## Deltas from semif-serve (deliberate) 1. **Prompt and model.** The prompt is Intern-Decision's own Jev prompt. - **Option ids are shown to the model** (`A = : `); SemIf showed only the descriptions. An option id is therefore part of the question, so give options meaningful or neutral ids. - In the bench negative control, the ids-in-prompt cue made the top stay on 10/144 rows after the descriptions moved, against SemIf's 14/144. 2. **Shared requests are one prompt, not independent rows.** The questions in a call are asked together. - A decision's answer can depend on the other questions in its call and on their order. In the bench, Wyrd scored 79/84 asked one decision at a time and 77/84 with a turn's 4 decisions in one prompt. - SemIf only shared a KV prefix, so each of its rows was independent. - Calls hold at most 16 questions, and the chunk boundaries follow request order. 3. **`probabilities` are temperature-scaled** by the checkpoint's shipped calibration (T = 1.99241824). - `probability_status` says so, and `calibration` carries the method and T. - SemIf's were raw softmax, labelled uncalibrated. The argmax is the same either way. - This is the vendor's calibration on the vendor's data, not ours. 4. **`option_logits` does not exist.** `predict()` does not expose logits. T=1 scores could only be derived up to a constant, which would be a derived number, not logits. 5. **New fields:** `top`, `confidence`, `calibration`, `native`, `call`. 6. **`prompt_sha256` and `input_tokens` describe the whole call**, shared by every decision in it. SemIf's were per row. 7. **`prompt_version`** names the pinned `inference.py`. **`readout`** and **`model`** describe Intern-Decision. 8. **`workload` is always a 422, and `/health.workloads` is `[]`.** There is no per-workload temperature table. This matches the deployed semif, whose table was empty. 9. **`/health`**: `semif_commit` is gone; the model's pins live in `model`. It adds `max_questions_per_call` and `chunking`. 10. **`/decide/shared` timing:** SemIf's prefix-cache fields cannot exist, because there is no prefix cache: `prefix_tokens`, `prefill_seconds`, `replicate_seconds`, `suffix_forward_seconds`, `true_suffix_tokens`, `padded_suffix_tokens` and `encode_seconds`. The timing adds `calls`, `questions_per_call`, `input_tokens` and `inference_seconds`. `total_seconds` and `batch_size` keep their meaning. 11. **`MAX_TOKENS` is per call** (state plus up to 16 questions) and defaults to 8192. SemIf's was per row and defaulted to 4096. 12. **Orderings run in waves**, one forward pass per wave per chunk. SemIf batched every ordering in one shared forward. 13. **Ties.** Per-ordering `top` uses the model's argmax, whose exact tie goes to the smaller option id string. SemIf used the first maximum in the ordering. `combined.top` keeps semif's rule. 14. **Images.** The vision tower is dropped (INV-7). The surface never took images. ## Tests (TDD, fake engine: no torch, no model) - **Auth.** A POST without the right bearer is 401 and never reaches the engine. `/health` needs no auth. A short or non-visible-ASCII token is refused at startup. - **Mapping.** - `/decide` sends one call with field `q` and the criteria in the caller's order. - `/decide/shared` sends `q1..qN` over the shared state. - 17 questions become 2 calls (16 + 1), and 40 become 3 (16 + 16 + 8). Results come back in request order, and `call.index` and `call.field` are right. - Decision ids never appear in a call. - **Result.** `probabilities` follows `option_ids`. `top`, `confidence`, `calibration`, `native` and `input_tokens` are passed through. An answer missing an option id is a 500. - **Orderings.** - `rotations` puts each option in each position once, and the waves never repeat a decision within a call. - `all` is n!, and 422 above 4 options. - A position bias cancels exactly. - `agreement` and `spread` are computed from the orderings. - A mixed request keeps plain results unchanged, and wave 0 equals the plain request. - Orderings count toward the cap. - **Validation (422).** Fewer than 2 or more than 16 options, duplicate option ids, duplicate decision ids, an empty id or question, an empty or non-finite state, a `workload`, a row count outside 1..max, malformed JSON, and a model `ValueError`. - **Limits.** A body over the limit is 413. A queue past `MAX_QUEUE` is 429 before the body is read. - **Engine failures.** An engine `OutOfMemory` is 503, and any other failure is 500. - **Concurrency.** Requests are serialised: two never overlap inside the engine, and the chunks of one request are not interleaved with another's. `/health` answers while a call is blocked. Every call runs on one thread, which is the thread the engine was loaded on. - **Engine against a fake torch.** - An OOM is re-raised unchained, and `empty_cache` runs only after the failed call's tensors are freed. - A RuntimeError saying "out of memory" becomes `OutOfMemory`. - Another failure becomes `ScoringFailed`, unchained. - A `ValueError` keeps its message but is released and raised unchained. - A request the prompt builder rejects never reaches the model. - A burst over baseline plus slack is released, and one at or under it is left alone. - **Engine load (fake).** - A wrong `inference.py` hash, or a checkpoint path that is not the pinned snapshot, refuses to start. - The cap is applied before the engine is constructed. - The vision swap refuses to start when the warm-up changes. - The prompt-hash check refuses to start on a token-count mismatch. - **Config.** Every value is validated. ## Acceptance (fv-ml1, real model; not unit tests) 1. **Positive control:** through the service, the bench's native numbers on the pooled 259 rows and on Wyrd, single ordering (bench: 240/259, Wyrd 79/84; floor: the pooled set resolves ±4 pts, and 0 labels moved across 4 restarts). Also a row-by-row comparison against the bench's own rows. 2. **Negative control:** descriptions rotated one place. The top follows the moved description (bench: 122/144 follow, 10/144 same top). 3. A-vs-A repeat stability, within the process and across restarts. 4. Latency at our shape: 21 binary criteria, and 16 criteria over the ~3,900-token state, both server-side and from nh3-dev. 5. Resident and peak VRAM (nvidia-smi and torch), and the cap chosen from them. 6. An over-cap request is a 503, memory returns to baseline, and the service keeps answering. 7. A `/decide/shared` with more than 16 questions is chunked, and its answers equal the same questions asked chunk by chunk. 8. 401 without the token, and 429 past the queue.