feat(intern-decision-serve): Intern-Decision-4B behind semif-serve's HTTP surface

Contract, service and tests (fake engine, no GPU). Scores through the checkpoint's own
inference.py (DecisionEngine.predict, sha256-pinned); maps semif decisions onto Jev choice
questions, packs /decide/shared into calls of at most 16, runs orderings in waves, and keeps
semif's error mapping, admission, body limit and hard VRAM cap. Deltas from semif-serve are
listed in the contract.
This commit is contained in:
vh
2026-09-30 09:04:39 -07:00
parent bb806e3596
commit 5bbf0aaeba
18 changed files with 3196 additions and 0 deletions
@@ -0,0 +1,271 @@
---
title: intern-decision-serve
kind: module-contract
status: draft
owner: infra-ops
created: 2026-09-30
replaces: semif-serve 0.1.4 (services/semif-serve/semif-serve.contract.md), external surface kept
depends_on:
- internlm/Intern-Decision-4B at revision 0e5e6aa7d6d750e2b1504ba11a8136cb58aeb3cd, BF16 (Apache-2.0)
- that snapshot's own inference.py (DecisionEngine.predict), sha256 c904e2c67ca0775621a22375ee373d2ba30b52117cda870c6c9ef74143b29863
- torch 2.10.0+cu128, transformers 5.17.0, flash-linear-attention 0.5.2, causal-conv1d 1.7.0 (the 2026-09-30 bench stack)
---
# intern-decision-serve: Intern-Decision-4B behind semif-serve's HTTP surface
## Purpose
Prime, 2026-09-30: "replace semif with intern-decision now". The bench
(`docs/pfi/jev-candidates-bench-2026-09-30.md`) picked Intern-Decision-4B on its own runtime.
This service loads the model once and scores every request with the checkpoint's **own**
`inference.py` (`DecisionEngine.predict`). It re-implements neither the prompt nor the readout.
It keeps semif-serve's external surface, so a caller written for semif-serve works unchanged.
It only maps semif-shaped requests onto the model's Jev request schema and maps the answers back.
Every deliberate difference is listed under **Deltas from semif-serve**.
## Endpoints (same as semif-serve)
Every POST takes and returns JSON and needs `Authorization: Bearer <token>`. `GET /health` is open.
| method | path | body | success |
|---|---|---|---|
| GET | `/health` | none | 200 `{status: "ok", model, vram_cap_gib, max_tokens, max_decisions, max_questions_per_call: 16, chunking, workloads: []}` |
| POST | `/decide` | `{id, state, question, options[2..16], orderings?, workload?}` | 200 one decision result |
| POST | `/decide/shared` | `{state, decisions: [{id, question, options, orderings?}], workload?}` | 200 `{results: [...], timing: {...}}`, results in request order |
`options` items are `{id, description}`. Validation is SemIf's, re-stated here because SemIf is
gone. `id` and `question` are nonempty strings. `state` is a nonempty string, object or array,
and must be finite JSON. There are 2..16 options, each with a string `id` and a string
`description`, and option ids are unique within a decision. Decision ids are unique within a
request. A violation is a 422.
## Mapping onto the model (the seam)
- **One call** is one `predict()` request:
`{"state": state, "questions": {<field>: {"type": "choice", "instructions": question, "criteria": {option.id: option.description, ...}}}}`.
The criteria keep the caller's option order. This is the format the bench measured.
- **Field names are positional:** `q` when a call carries one question, and `q1`..`qN` in
request order when it carries several. Decision ids never reach the prompt. **Option ids do:**
the model prints `A = <id>: <description>`.
- **`/decide`** is one call with one question.
- **`/decide/shared`** is packed into calls of **at most 16 questions**, the model's own limit
(`validate_request`). The packing is greedy and in request order: questions 1–16, then 17–32,
and so on. Every call of a request runs back to back under the inference lock. `/health` reports
`max_questions_per_call: 16` and the rule in `chunking`.
- **Orderings.** A decision may set `orderings: "none" | "rotations" | "all"`, with semif's
meaning. `rotations` gives the n cyclic shifts, the caller's order first. `all` gives the n!
permutations, the caller's order first, and is 422 above 4 options. Ordering k of every decision
that has more than k orderings forms **wave k**. Each wave is packed as above. So wave 0 is the
request exactly as written, and no prompt ever holds two orderings of the same decision. Every
ordering counts toward `MAX_DECISIONS`.
- **The result for one ordering** (and for every plain decision):
```
{id, option_ids, probabilities, top, confidence, calibration, native,
input_tokens, prompt_sha256, prompt_version, model, readout, probability_status,
call: {index, field, questions}}
```
- `native` is the answer object `predict()` returned for that field, unchanged.
- `probabilities` is `native.probabilities` listed in `option_ids` order.
- `top` is `native.choice` and `confidence` is `native.confidence`.
- `calibration` is the `calibration` object from `predict()`.
- `input_tokens` and `prompt_sha256` describe the whole call.
- `/decide` results (plain) also carry `total_seconds` and `forward_seconds`. `forward_seconds`
is `predict()`'s `timing.inference_ms` / 1000.
- **Averaged result:** semif's shape, `{id, option_ids, combined: {method, orderings, probabilities, top, agreement, spread}, orderings: [...]}`.
- The per-ordering ids are `<id>#o<k>`.
- `combined.probabilities` renormalises the per-option mean of `log p`. A probability of
exactly 0 is floored at 1e-300 before the log.
- `combined.top` is the first maximum in the caller's order.
- `agreement` is the share of orderings whose `top` equals `combined.top`.
- `spread` holds each option's min and max `probabilities` across orderings.
- Temperature scaling is one monotone transform per ordering, so `combined.top` and
`agreement` are what combining T=1 scores would give.
- **`/decide/shared` timing:** `{total_seconds, batch_size (orderings scored), calls, questions_per_call: [...], input_tokens: [...], inference_seconds}`.
## Invariants
- **INV-1 pass-through.** Every number in `native`, `probabilities`, `top`, `confidence`,
`calibration` and `input_tokens` is what `predict()` returned. The wrapper only re-keys it.
If an answer lacks one of the decision's option ids, that is a 500 `scoring_failed`, never a
guess.
- **INV-2 one model, one inference at a time.** The model loads at startup, and a process-wide
lock serialises every request's calls (all of a request's chunks run inside one hold). The app
runs one worker. Calls run off the event loop, so `/health` answers during one.
- **INV-3 fail-closed startup.** Before the service serves, all of these must hold:
- `inference.py` in the checkpoint hashes to the pinned sha256 (it is executed code, loaded
from a data mount);
- the checkpoint path ends in `snapshots/<pinned revision>`;
- the model sits on CUDA, and torch's arch list has the card's `sm_XY`;
- one warm-up decision scores.
`device=cpu` is allowed only when set explicitly.
- **INV-4 VRAM cap.** `VRAM_CAP_GIB`, when set, becomes `torch.cuda.set_per_process_memory_fraction`
**before** the weights load.
- An OOM in any call makes the whole request a 503 `out_of_memory`. So does a RuntimeError
whose first line says "out of memory". The engine then frees the failed call's frames,
runs `gc.collect()` and `empty_cache()`, and raises unchained. The process stays up.
- After every call, if reserved memory exceeds the post-warm-up baseline by more than
`RELEASE_SLACK_MIB` (default 512), the engine runs `empty_cache()`.
- Any other failure except `ValueError` is logged with its traceback and released the same
way, then raised unchained as `ScoringFailed`.
- **INV-5 no network.** The entry point sets `HF_HUB_OFFLINE=1` and `TRANSFORMERS_OFFLINE=1`
before torch or transformers load. The weights are read from the mounted, read-only HF cache.
- **INV-6 constant-time auth.** The token is compared with `hmac.compare_digest`. It must be at
least 32 visible ASCII characters (33–126), or startup refuses it.
- **INV-7 the text-only model is the same model.**
- The service takes no images, and neither did semif. So after the first warm-up the engine
replaces the vision tower (`model.model.visual`, 0.62 GiB) with a stub that raises if it is
ever called.
- It then scores the warm-up again. It refuses to start unless the answer is bit-identical to
the first one.
- `KEEP_VISION=1` keeps the tower.
- **INV-8 honest prompt hash.** `prompt_sha256` is the sha256 of the chat-template text rendered
from `inference.compile_row(...)` with the same arguments `HFBackend.encode` uses. At startup,
that text must tokenise to exactly the `input_tokens` `predict()` reports for the warm-up.
Otherwise the service refuses to start rather than hash a prompt the model never saw.
## Limits and errors
- `MAX_TOKENS` (default 8192, the model's own default) is `DecisionEngine(max_length=...)`, and
applies to a **whole call**: the state plus all its questions. A longer call is a 422. It is
never truncated; the model raises.
- `MAX_DECISIONS` (default 64) caps the orderings scored per request, counted after expansion.
A request must score 1..max of them, else 422.
- The body may be at most `MAX_BODY_BYTES` (default 1 MiB), else 413. This is checked before
each chunk is kept. A declared `Content-Length` is trusted only as ASCII digits.
- **Admission:** at most `MAX_QUEUE` (default 32) POSTs may be queued or scoring at once. The
next one gets 429 `busy` before its body is read.
- `workload`: there is no per-workload table, and the model's own calibration always applies.
So any non-null `workload` is a 422, exactly as semif-serve behaved with its deployed empty
table.
| status | code | when |
|---|---|---|
| 401 | `unauthorized` | missing or wrong bearer |
| 413 | `request_too_large` | body over the limit |
| 422 | `invalid_request` | bad JSON or shape; a SemIf-rule violation; a model `ValueError` (token limit, reserved `<decision>` marker in the input, ...); a `workload`; `all` over 4 options; a row count outside 1..`MAX_DECISIONS` |
| 429 | `busy` | `MAX_QUEUE` requests in progress |
| 503 | `out_of_memory` | CUDA OOM in any call of the request |
| 500 | `scoring_failed` | any other model failure, including one while building the response |
The error body is `{error: {code, message}}`.
## Configuration (env, prefix `INTERN_DECISION_`)
- `API_TOKEN` is required.
- `CHECKPOINT` defaults to
`/hf/hub/models--internlm--Intern-Decision-4B/snapshots/<revision>`.
- `DEVICE` defaults to `cuda`.
- `VRAM_CAP_GIB` has no default: unset means uncapped, and when set it must be finite and > 0.
- `MAX_TOKENS`, `MAX_DECISIONS`, `MAX_BODY_BYTES`, `MAX_QUEUE` and `RELEASE_SLACK_MIB` are
integers; each must be ≥ 1, except the slack, which must be ≥ 0.
- `KEEP_VISION` is `0` or `1`.
A bad value is refused at startup with a `ValueError` naming the variable. In the stack's `.env`
on the host, the cap is the single knob `VRAM_CAP_GIB`.
## Deltas from semif-serve (deliberate)
1. **Prompt and model.** The prompt is Intern-Decision's own Jev prompt.
- **Option ids are shown to the model** (`A = <id>: <description>`); SemIf showed only the
descriptions. An option id is therefore part of the question, so give options meaningful
or neutral ids.
- In the bench negative control, the ids-in-prompt cue made the top stay on 10/144 rows
after the descriptions moved, against SemIf's 14/144.
2. **Shared requests are one prompt, not independent rows.** The questions in a call are asked
together.
- A decision's answer can depend on the other questions in its call and on their order. In
the bench, Wyrd scored 79/84 asked one decision at a time and 77/84 with a turn's 4
decisions in one prompt.
- SemIf only shared a KV prefix, so each of its rows was independent.
- Calls hold at most 16 questions, and the chunk boundaries follow request order.
3. **`probabilities` are temperature-scaled** by the checkpoint's shipped calibration
(T = 1.99241824).
- `probability_status` says so, and `calibration` carries the method and T.
- SemIf's were raw softmax, labelled uncalibrated. The argmax is the same either way.
- This is the vendor's calibration on the vendor's data, not ours.
4. **`option_logits` does not exist.** `predict()` does not expose logits. T=1 scores could only
be derived up to a constant, which would be a derived number, not logits.
5. **New fields:** `top`, `confidence`, `calibration`, `native`, `call`.
6. **`prompt_sha256` and `input_tokens` describe the whole call**, shared by every decision in
it. SemIf's were per row.
7. **`prompt_version`** names the pinned `inference.py`. **`readout`** and **`model`** describe
Intern-Decision.
8. **`workload` is always a 422, and `/health.workloads` is `[]`.** There is no per-workload
temperature table. This matches the deployed semif, whose table was empty.
9. **`/health`**: `semif_commit` is gone; the model's pins live in `model`. It adds
`max_questions_per_call` and `chunking`.
10. **`/decide/shared` timing:** SemIf's prefix-cache fields cannot exist, because there is no
prefix cache: `prefix_tokens`, `prefill_seconds`, `replicate_seconds`,
`suffix_forward_seconds`, `true_suffix_tokens`, `padded_suffix_tokens` and `encode_seconds`.
The timing adds `calls`, `questions_per_call`, `input_tokens` and `inference_seconds`.
`total_seconds` and `batch_size` keep their meaning.
11. **`MAX_TOKENS` is per call** (state plus up to 16 questions) and defaults to 8192. SemIf's
was per row and defaulted to 4096.
12. **Orderings run in waves**, one forward pass per wave per chunk. SemIf batched every
ordering in one shared forward.
13. **Ties.** Per-ordering `top` uses the model's argmax, whose exact tie goes to the smaller
option id string. SemIf used the first maximum in the ordering. `combined.top` keeps semif's
rule.
14. **Images.** The vision tower is dropped (INV-7). The surface never took images.
## Tests (TDD, fake engine: no torch, no model)
- **Auth.** A POST without the right bearer is 401 and never reaches the engine. `/health`
needs no auth. A short or non-visible-ASCII token is refused at startup.
- **Mapping.**
- `/decide` sends one call with field `q` and the criteria in the caller's order.
- `/decide/shared` sends `q1..qN` over the shared state.
- 17 questions become 2 calls (16 + 1), and 40 become 3 (16 + 16 + 8). Results come back in
request order, and `call.index` and `call.field` are right.
- Decision ids never appear in a call.
- **Result.** `probabilities` follows `option_ids`. `top`, `confidence`, `calibration`, `native`
and `input_tokens` are passed through. An answer missing an option id is a 500.
- **Orderings.**
- `rotations` puts each option in each position once, and the waves never repeat a decision
within a call.
- `all` is n!, and 422 above 4 options.
- A position bias cancels exactly.
- `agreement` and `spread` are computed from the orderings.
- A mixed request keeps plain results unchanged, and wave 0 equals the plain request.
- Orderings count toward the cap.
- **Validation (422).** Fewer than 2 or more than 16 options, duplicate option ids, duplicate
decision ids, an empty id or question, an empty or non-finite state, a `workload`, a row count
outside 1..max, malformed JSON, and a model `ValueError`.
- **Limits.** A body over the limit is 413. A queue past `MAX_QUEUE` is 429 before the body is
read.
- **Engine failures.** An engine `OutOfMemory` is 503, and any other failure is 500.
- **Concurrency.** Requests are serialised: two never overlap inside the engine, and the
chunks of one request are not interleaved with another's. `/health` answers while a call is
blocked.
- **Engine against a fake torch.**
- An OOM is re-raised unchained, and `empty_cache` runs only after the failed call's tensors
are freed.
- A RuntimeError saying "out of memory" becomes `OutOfMemory`.
- Another failure becomes `ScoringFailed`, unchained.
- `ValueError` passes through.
- A burst over baseline plus slack is released, and one at or under it is left alone.
- **Engine load (fake).**
- A wrong `inference.py` hash, or a checkpoint path that is not the pinned snapshot, refuses
to start.
- The cap is applied before the engine is constructed.
- The vision swap refuses to start when the warm-up changes.
- The prompt-hash check refuses to start on a token-count mismatch.
- **Config.** Every value is validated.
## Acceptance (fv-ml1, real model; not unit tests)
1. **Positive control:** through the service, the bench's native numbers on the pooled 259 rows
and on Wyrd, single ordering (bench: 240/259, Wyrd 79/84; floor: the pooled set resolves
±4 pts, and 0 labels moved across 4 restarts). Also a row-by-row comparison against the
bench's own rows.
2. **Negative control:** descriptions rotated one place. The top follows the moved description
(bench: 122/144 follow, 10/144 same top).
3. A-vs-A repeat stability, within the process and across restarts.
4. Latency at our shape: 21 binary criteria, and 16 criteria over the ~3,900-token state, both
server-side and from nh3-dev.
5. Resident and peak VRAM (nvidia-smi and torch), and the cap chosen from them.
6. An over-cap request is a 503, memory returns to baseline, and the service keeps answering.
7. A `/decide/shared` with more than 16 questions is chunked, and its answers equal the same
questions asked chunk by chunk.
8. 401 without the token, and 429 past the queue.