Files
esh-pfi-infrastructure/services/intern-decision-serve/intern-decision-serve.contract.md
T
vh 5bbf0aaeba feat(intern-decision-serve): Intern-Decision-4B behind semif-serve's HTTP surface
Contract, service and tests (fake engine, no GPU). Scores through the checkpoint's own
inference.py (DecisionEngine.predict, sha256-pinned); maps semif decisions onto Jev choice
questions, packs /decide/shared into calls of at most 16, runs orderings in waves, and keeps
semif's error mapping, admission, body limit and hard VRAM cap. Deltas from semif-serve are
listed in the contract.
2026-09-30 09:04:39 -07:00

272 lines
17 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
title: intern-decision-serve
kind: module-contract
status: draft
owner: infra-ops
created: 2026-09-30
replaces: semif-serve 0.1.4 (services/semif-serve/semif-serve.contract.md), external surface kept
depends_on:
- internlm/Intern-Decision-4B at revision 0e5e6aa7d6d750e2b1504ba11a8136cb58aeb3cd, BF16 (Apache-2.0)
- that snapshot's own inference.py (DecisionEngine.predict), sha256 c904e2c67ca0775621a22375ee373d2ba30b52117cda870c6c9ef74143b29863
- torch 2.10.0+cu128, transformers 5.17.0, flash-linear-attention 0.5.2, causal-conv1d 1.7.0 (the 2026-09-30 bench stack)
---
# intern-decision-serve: Intern-Decision-4B behind semif-serve's HTTP surface
## Purpose
Prime, 2026-09-30: "replace semif with intern-decision now". The bench
(`docs/pfi/jev-candidates-bench-2026-09-30.md`) picked Intern-Decision-4B on its own runtime.
This service loads the model once and scores every request with the checkpoint's **own**
`inference.py` (`DecisionEngine.predict`). It re-implements neither the prompt nor the readout.
It keeps semif-serve's external surface, so a caller written for semif-serve works unchanged.
It only maps semif-shaped requests onto the model's Jev request schema and maps the answers back.
Every deliberate difference is listed under **Deltas from semif-serve**.
## Endpoints (same as semif-serve)
Every POST takes and returns JSON and needs `Authorization: Bearer <token>`. `GET /health` is open.
| method | path | body | success |
|---|---|---|---|
| GET | `/health` | none | 200 `{status: "ok", model, vram_cap_gib, max_tokens, max_decisions, max_questions_per_call: 16, chunking, workloads: []}` |
| POST | `/decide` | `{id, state, question, options[2..16], orderings?, workload?}` | 200 one decision result |
| POST | `/decide/shared` | `{state, decisions: [{id, question, options, orderings?}], workload?}` | 200 `{results: [...], timing: {...}}`, results in request order |
`options` items are `{id, description}`. Validation is SemIf's, re-stated here because SemIf is
gone. `id` and `question` are nonempty strings. `state` is a nonempty string, object or array,
and must be finite JSON. There are 2..16 options, each with a string `id` and a string
`description`, and option ids are unique within a decision. Decision ids are unique within a
request. A violation is a 422.
## Mapping onto the model (the seam)
- **One call** is one `predict()` request:
`{"state": state, "questions": {<field>: {"type": "choice", "instructions": question, "criteria": {option.id: option.description, ...}}}}`.
The criteria keep the caller's option order. This is the format the bench measured.
- **Field names are positional:** `q` when a call carries one question, and `q1`..`qN` in
request order when it carries several. Decision ids never reach the prompt. **Option ids do:**
the model prints `A = <id>: <description>`.
- **`/decide`** is one call with one question.
- **`/decide/shared`** is packed into calls of **at most 16 questions**, the model's own limit
(`validate_request`). The packing is greedy and in request order: questions 1–16, then 17–32,
and so on. Every call of a request runs back to back under the inference lock. `/health` reports
`max_questions_per_call: 16` and the rule in `chunking`.
- **Orderings.** A decision may set `orderings: "none" | "rotations" | "all"`, with semif's
meaning. `rotations` gives the n cyclic shifts, the caller's order first. `all` gives the n!
permutations, the caller's order first, and is 422 above 4 options. Ordering k of every decision
that has more than k orderings forms **wave k**. Each wave is packed as above. So wave 0 is the
request exactly as written, and no prompt ever holds two orderings of the same decision. Every
ordering counts toward `MAX_DECISIONS`.
- **The result for one ordering** (and for every plain decision):
```
{id, option_ids, probabilities, top, confidence, calibration, native,
input_tokens, prompt_sha256, prompt_version, model, readout, probability_status,
call: {index, field, questions}}
```
- `native` is the answer object `predict()` returned for that field, unchanged.
- `probabilities` is `native.probabilities` listed in `option_ids` order.
- `top` is `native.choice` and `confidence` is `native.confidence`.
- `calibration` is the `calibration` object from `predict()`.
- `input_tokens` and `prompt_sha256` describe the whole call.
- `/decide` results (plain) also carry `total_seconds` and `forward_seconds`. `forward_seconds`
is `predict()`'s `timing.inference_ms` / 1000.
- **Averaged result:** semif's shape, `{id, option_ids, combined: {method, orderings, probabilities, top, agreement, spread}, orderings: [...]}`.
- The per-ordering ids are `<id>#o<k>`.
- `combined.probabilities` renormalises the per-option mean of `log p`. A probability of
exactly 0 is floored at 1e-300 before the log.
- `combined.top` is the first maximum in the caller's order.
- `agreement` is the share of orderings whose `top` equals `combined.top`.
- `spread` holds each option's min and max `probabilities` across orderings.
- Temperature scaling is one monotone transform per ordering, so `combined.top` and
`agreement` are what combining T=1 scores would give.
- **`/decide/shared` timing:** `{total_seconds, batch_size (orderings scored), calls, questions_per_call: [...], input_tokens: [...], inference_seconds}`.
## Invariants
- **INV-1 pass-through.** Every number in `native`, `probabilities`, `top`, `confidence`,
`calibration` and `input_tokens` is what `predict()` returned. The wrapper only re-keys it.
If an answer lacks one of the decision's option ids, that is a 500 `scoring_failed`, never a
guess.
- **INV-2 one model, one inference at a time.** The model loads at startup, and a process-wide
lock serialises every request's calls (all of a request's chunks run inside one hold). The app
runs one worker. Calls run off the event loop, so `/health` answers during one.
- **INV-3 fail-closed startup.** Before the service serves, all of these must hold:
- `inference.py` in the checkpoint hashes to the pinned sha256 (it is executed code, loaded
from a data mount);
- the checkpoint path ends in `snapshots/<pinned revision>`;
- the model sits on CUDA, and torch's arch list has the card's `sm_XY`;
- one warm-up decision scores.
`device=cpu` is allowed only when set explicitly.
- **INV-4 VRAM cap.** `VRAM_CAP_GIB`, when set, becomes `torch.cuda.set_per_process_memory_fraction`
**before** the weights load.
- An OOM in any call makes the whole request a 503 `out_of_memory`. So does a RuntimeError
whose first line says "out of memory". The engine then frees the failed call's frames,
runs `gc.collect()` and `empty_cache()`, and raises unchained. The process stays up.
- After every call, if reserved memory exceeds the post-warm-up baseline by more than
`RELEASE_SLACK_MIB` (default 512), the engine runs `empty_cache()`.
- Any other failure except `ValueError` is logged with its traceback and released the same
way, then raised unchained as `ScoringFailed`.
- **INV-5 no network.** The entry point sets `HF_HUB_OFFLINE=1` and `TRANSFORMERS_OFFLINE=1`
before torch or transformers load. The weights are read from the mounted, read-only HF cache.
- **INV-6 constant-time auth.** The token is compared with `hmac.compare_digest`. It must be at
least 32 visible ASCII characters (33–126), or startup refuses it.
- **INV-7 the text-only model is the same model.**
- The service takes no images, and neither did semif. So after the first warm-up the engine
replaces the vision tower (`model.model.visual`, 0.62 GiB) with a stub that raises if it is
ever called.
- It then scores the warm-up again. It refuses to start unless the answer is bit-identical to
the first one.
- `KEEP_VISION=1` keeps the tower.
- **INV-8 honest prompt hash.** `prompt_sha256` is the sha256 of the chat-template text rendered
from `inference.compile_row(...)` with the same arguments `HFBackend.encode` uses. At startup,
that text must tokenise to exactly the `input_tokens` `predict()` reports for the warm-up.
Otherwise the service refuses to start rather than hash a prompt the model never saw.
## Limits and errors
- `MAX_TOKENS` (default 8192, the model's own default) is `DecisionEngine(max_length=...)`, and
applies to a **whole call**: the state plus all its questions. A longer call is a 422. It is
never truncated; the model raises.
- `MAX_DECISIONS` (default 64) caps the orderings scored per request, counted after expansion.
A request must score 1..max of them, else 422.
- The body may be at most `MAX_BODY_BYTES` (default 1 MiB), else 413. This is checked before
each chunk is kept. A declared `Content-Length` is trusted only as ASCII digits.
- **Admission:** at most `MAX_QUEUE` (default 32) POSTs may be queued or scoring at once. The
next one gets 429 `busy` before its body is read.
- `workload`: there is no per-workload table, and the model's own calibration always applies.
So any non-null `workload` is a 422, exactly as semif-serve behaved with its deployed empty
table.
| status | code | when |
|---|---|---|
| 401 | `unauthorized` | missing or wrong bearer |
| 413 | `request_too_large` | body over the limit |
| 422 | `invalid_request` | bad JSON or shape; a SemIf-rule violation; a model `ValueError` (token limit, reserved `<decision>` marker in the input, ...); a `workload`; `all` over 4 options; a row count outside 1..`MAX_DECISIONS` |
| 429 | `busy` | `MAX_QUEUE` requests in progress |
| 503 | `out_of_memory` | CUDA OOM in any call of the request |
| 500 | `scoring_failed` | any other model failure, including one while building the response |
The error body is `{error: {code, message}}`.
## Configuration (env, prefix `INTERN_DECISION_`)
- `API_TOKEN` is required.
- `CHECKPOINT` defaults to
`/hf/hub/models--internlm--Intern-Decision-4B/snapshots/<revision>`.
- `DEVICE` defaults to `cuda`.
- `VRAM_CAP_GIB` has no default: unset means uncapped, and when set it must be finite and > 0.
- `MAX_TOKENS`, `MAX_DECISIONS`, `MAX_BODY_BYTES`, `MAX_QUEUE` and `RELEASE_SLACK_MIB` are
integers; each must be ≥ 1, except the slack, which must be ≥ 0.
- `KEEP_VISION` is `0` or `1`.
A bad value is refused at startup with a `ValueError` naming the variable. In the stack's `.env`
on the host, the cap is the single knob `VRAM_CAP_GIB`.
## Deltas from semif-serve (deliberate)
1. **Prompt and model.** The prompt is Intern-Decision's own Jev prompt.
- **Option ids are shown to the model** (`A = <id>: <description>`); SemIf showed only the
descriptions. An option id is therefore part of the question, so give options meaningful
or neutral ids.
- In the bench negative control, the ids-in-prompt cue made the top stay on 10/144 rows
after the descriptions moved, against SemIf's 14/144.
2. **Shared requests are one prompt, not independent rows.** The questions in a call are asked
together.
- A decision's answer can depend on the other questions in its call and on their order. In
the bench, Wyrd scored 79/84 asked one decision at a time and 77/84 with a turn's 4
decisions in one prompt.
- SemIf only shared a KV prefix, so each of its rows was independent.
- Calls hold at most 16 questions, and the chunk boundaries follow request order.
3. **`probabilities` are temperature-scaled** by the checkpoint's shipped calibration
(T = 1.99241824).
- `probability_status` says so, and `calibration` carries the method and T.
- SemIf's were raw softmax, labelled uncalibrated. The argmax is the same either way.
- This is the vendor's calibration on the vendor's data, not ours.
4. **`option_logits` does not exist.** `predict()` does not expose logits. T=1 scores could only
be derived up to a constant, which would be a derived number, not logits.
5. **New fields:** `top`, `confidence`, `calibration`, `native`, `call`.
6. **`prompt_sha256` and `input_tokens` describe the whole call**, shared by every decision in
it. SemIf's were per row.
7. **`prompt_version`** names the pinned `inference.py`. **`readout`** and **`model`** describe
Intern-Decision.
8. **`workload` is always a 422, and `/health.workloads` is `[]`.** There is no per-workload
temperature table. This matches the deployed semif, whose table was empty.
9. **`/health`**: `semif_commit` is gone; the model's pins live in `model`. It adds
`max_questions_per_call` and `chunking`.
10. **`/decide/shared` timing:** SemIf's prefix-cache fields cannot exist, because there is no
prefix cache: `prefix_tokens`, `prefill_seconds`, `replicate_seconds`,
`suffix_forward_seconds`, `true_suffix_tokens`, `padded_suffix_tokens` and `encode_seconds`.
The timing adds `calls`, `questions_per_call`, `input_tokens` and `inference_seconds`.
`total_seconds` and `batch_size` keep their meaning.
11. **`MAX_TOKENS` is per call** (state plus up to 16 questions) and defaults to 8192. SemIf's
was per row and defaulted to 4096.
12. **Orderings run in waves**, one forward pass per wave per chunk. SemIf batched every
ordering in one shared forward.
13. **Ties.** Per-ordering `top` uses the model's argmax, whose exact tie goes to the smaller
option id string. SemIf used the first maximum in the ordering. `combined.top` keeps semif's
rule.
14. **Images.** The vision tower is dropped (INV-7). The surface never took images.
## Tests (TDD, fake engine: no torch, no model)
- **Auth.** A POST without the right bearer is 401 and never reaches the engine. `/health`
needs no auth. A short or non-visible-ASCII token is refused at startup.
- **Mapping.**
- `/decide` sends one call with field `q` and the criteria in the caller's order.
- `/decide/shared` sends `q1..qN` over the shared state.
- 17 questions become 2 calls (16 + 1), and 40 become 3 (16 + 16 + 8). Results come back in
request order, and `call.index` and `call.field` are right.
- Decision ids never appear in a call.
- **Result.** `probabilities` follows `option_ids`. `top`, `confidence`, `calibration`, `native`
and `input_tokens` are passed through. An answer missing an option id is a 500.
- **Orderings.**
- `rotations` puts each option in each position once, and the waves never repeat a decision
within a call.
- `all` is n!, and 422 above 4 options.
- A position bias cancels exactly.
- `agreement` and `spread` are computed from the orderings.
- A mixed request keeps plain results unchanged, and wave 0 equals the plain request.
- Orderings count toward the cap.
- **Validation (422).** Fewer than 2 or more than 16 options, duplicate option ids, duplicate
decision ids, an empty id or question, an empty or non-finite state, a `workload`, a row count
outside 1..max, malformed JSON, and a model `ValueError`.
- **Limits.** A body over the limit is 413. A queue past `MAX_QUEUE` is 429 before the body is
read.
- **Engine failures.** An engine `OutOfMemory` is 503, and any other failure is 500.
- **Concurrency.** Requests are serialised: two never overlap inside the engine, and the
chunks of one request are not interleaved with another's. `/health` answers while a call is
blocked.
- **Engine against a fake torch.**
- An OOM is re-raised unchained, and `empty_cache` runs only after the failed call's tensors
are freed.
- A RuntimeError saying "out of memory" becomes `OutOfMemory`.
- Another failure becomes `ScoringFailed`, unchained.
- `ValueError` passes through.
- A burst over baseline plus slack is released, and one at or under it is left alone.
- **Engine load (fake).**
- A wrong `inference.py` hash, or a checkpoint path that is not the pinned snapshot, refuses
to start.
- The cap is applied before the engine is constructed.
- The vision swap refuses to start when the warm-up changes.
- The prompt-hash check refuses to start on a token-count mismatch.
- **Config.** Every value is validated.
## Acceptance (fv-ml1, real model; not unit tests)
1. **Positive control:** through the service, the bench's native numbers on the pooled 259 rows
and on Wyrd, single ordering (bench: 240/259, Wyrd 79/84; floor: the pooled set resolves
±4 pts, and 0 labels moved across 4 restarts). Also a row-by-row comparison against the
bench's own rows.
2. **Negative control:** descriptions rotated one place. The top follows the moved description
(bench: 122/144 follow, 10/144 same top).
3. A-vs-A repeat stability, within the process and across restarts.
4. Latency at our shape: 21 binary criteria, and 16 criteria over the ~3,900-token state, both
server-side and from nh3-dev.
5. Resident and peak VRAM (nvidia-smi and torch), and the cap chosen from them.
6. An over-cap request is a 503, memory returns to baseline, and the service keeps answering.
7. A `/decide/shared` with more than 16 questions is chunked, and its answers equal the same
questions asked chunk by chunk.
8. 401 without the token, and 429 past the queue.