- a failure while building the response (a non-finite number included) is a 500 inside the envelope, never a 422 or a render crash outside it - an engine ValueError keeps its message but is released and raised unchained, like an OOM - the prompt is built (and the model's own validation run) before the forward - the row cap is counted before any ordering is built - /health reads a device name cached at load, so it makes no driver call off the inference thread
289 lines
18 KiB
Markdown
289 lines
18 KiB
Markdown
---
|
||
title: intern-decision-serve
|
||
kind: module-contract
|
||
status: draft
|
||
owner: infra-ops
|
||
created: 2026-09-30
|
||
replaces: semif-serve 0.1.4 (services/semif-serve/semif-serve.contract.md), external surface kept
|
||
depends_on:
|
||
- internlm/Intern-Decision-4B at revision 0e5e6aa7d6d750e2b1504ba11a8136cb58aeb3cd, BF16 (Apache-2.0)
|
||
- that snapshot's own inference.py (DecisionEngine.predict), sha256 c904e2c67ca0775621a22375ee373d2ba30b52117cda870c6c9ef74143b29863
|
||
- torch 2.10.0+cu128, transformers 5.17.0, flash-linear-attention 0.5.2, causal-conv1d 1.7.0 (the 2026-09-30 bench stack)
|
||
---
|
||
|
||
# intern-decision-serve: Intern-Decision-4B behind semif-serve's HTTP surface
|
||
|
||
## Purpose
|
||
|
||
Prime, 2026-09-30: "replace semif with intern-decision now". The bench
|
||
(`docs/pfi/jev-candidates-bench-2026-09-30.md`) picked Intern-Decision-4B on its own runtime.
|
||
This service loads the model once and scores every request with the checkpoint's **own**
|
||
`inference.py` (`DecisionEngine.predict`). It re-implements neither the prompt nor the readout.
|
||
It keeps semif-serve's external surface, so a caller written for semif-serve works unchanged.
|
||
It only maps semif-shaped requests onto the model's Jev request schema and maps the answers back.
|
||
Every deliberate difference is listed under **Deltas from semif-serve**.
|
||
|
||
## Endpoints (same as semif-serve)
|
||
|
||
Every POST takes and returns JSON and needs `Authorization: Bearer <token>`. `GET /health` is open.
|
||
|
||
| method | path | body | success |
|
||
|---|---|---|---|
|
||
| GET | `/health` | none | 200 `{status: "ok", model, vram_cap_gib, max_tokens, max_decisions, max_questions_per_call: 16, chunking, workloads: []}` |
|
||
| POST | `/decide` | `{id, state, question, options[2..16], orderings?, workload?}` | 200 one decision result |
|
||
| POST | `/decide/shared` | `{state, decisions: [{id, question, options, orderings?}], workload?}` | 200 `{results: [...], timing: {...}}`, results in request order |
|
||
|
||
`options` items are `{id, description}`. Validation is SemIf's, re-stated here because SemIf is
|
||
gone. `id` and `question` are nonempty strings. `state` is a nonempty string, object or array,
|
||
and must be finite JSON. There are 2..16 options, each with a string `id` and a string
|
||
`description`, and option ids are unique within a decision. Decision ids are unique within a
|
||
request. A violation is a 422.
|
||
|
||
## Mapping onto the model (the seam)
|
||
|
||
- **One call** is one `predict()` request:
|
||
`{"state": state, "questions": {<field>: {"type": "choice", "instructions": question, "criteria": {option.id: option.description, ...}}}}`.
|
||
The criteria keep the caller's option order.
|
||
- **A one-question call is exactly the bench's native request.** Acceptance showed it
|
||
bit-identical, row for row.
|
||
- **A multi-question call differs from the bench's multifield run only in the field names.**
|
||
Those are positional here and were the decision names in the bench. Measured: 1 of 84 Wyrd
|
||
rows differs (78/84 against 77/84).
|
||
- **Field names are positional:** `q` when a call carries one question, and `q1`..`qN` in
|
||
request order when it carries several. Decision ids never reach the prompt. **Option ids do:**
|
||
the model prints `A = <id>: <description>`.
|
||
- **`/decide`** is one call with one question.
|
||
- **`/decide/shared`** is packed into calls of **at most 16 questions**, the model's own limit
|
||
(`validate_request`). The packing is greedy and in request order: questions 1–16, then 17–32,
|
||
and so on. Every call of a request runs back to back under the inference lock. `/health` reports
|
||
`max_questions_per_call: 16` and the rule in `chunking`.
|
||
- **Orderings.** A decision may set `orderings: "none" | "rotations" | "all"`, with semif's
|
||
meaning. `rotations` gives the n cyclic shifts, the caller's order first. `all` gives the n!
|
||
permutations, the caller's order first, and is 422 above 4 options. Ordering k of every decision
|
||
that has more than k orderings forms **wave k**. Each wave is packed as above. So wave 0 is the
|
||
request exactly as written, and no prompt ever holds two orderings of the same decision. Every
|
||
ordering counts toward `MAX_DECISIONS`.
|
||
- **The result for one ordering** (and for every plain decision):
|
||
```
|
||
{id, option_ids, probabilities, top, confidence, calibration, native,
|
||
input_tokens, prompt_sha256, prompt_version, model, readout, probability_status,
|
||
call: {index, field, questions}}
|
||
```
|
||
- `native` is the answer object `predict()` returned for that field, unchanged.
|
||
- `probabilities` is `native.probabilities` listed in `option_ids` order.
|
||
- `top` is `native.choice` and `confidence` is `native.confidence`.
|
||
- `calibration` is the `calibration` object from `predict()`.
|
||
- `input_tokens` and `prompt_sha256` describe the whole call.
|
||
- `/decide` results (plain) also carry `total_seconds` and `forward_seconds`. `forward_seconds`
|
||
is `predict()`'s `timing.inference_ms` / 1000.
|
||
- **Averaged result:** semif's shape, `{id, option_ids, combined: {method, orderings, probabilities, top, agreement, spread}, orderings: [...]}`.
|
||
- The per-ordering ids are `<id>#o<k>`.
|
||
- `combined.probabilities` renormalises the per-option mean of `log p`. A probability of
|
||
exactly 0 is floored at 1e-300 before the log.
|
||
- `combined.top` is the first maximum in the caller's order.
|
||
- `agreement` is the share of orderings whose `top` equals `combined.top`.
|
||
- `spread` holds each option's min and max `probabilities` across orderings.
|
||
- Temperature scaling is one monotone transform per ordering, so `combined.top` and
|
||
`agreement` are what combining T=1 scores would give.
|
||
- **`/decide/shared` timing:** `{total_seconds, batch_size (orderings scored), calls, questions_per_call: [...], input_tokens: [...], inference_seconds}`.
|
||
|
||
## Invariants
|
||
|
||
- **INV-1 pass-through.** Every number in `native`, `probabilities`, `top`, `confidence`,
|
||
`calibration` and `input_tokens` is what `predict()` returned. The wrapper only re-keys it.
|
||
If an answer lacks one of the decision's option ids, that is a 500 `scoring_failed`, never a
|
||
guess. So is any failure while building the response, a non-finite number included. That stays
|
||
inside the error envelope and is never a 422 or a bare 500.
|
||
- **INV-2 one model, one inference thread.** The model loads at startup on a dedicated
|
||
single-thread executor. The warm-ups and every later call run on that **same host thread**,
|
||
never on the event loop's threadpool. A lock also serialises each request's calls, so all of a
|
||
request's chunks run inside one hold. The app runs one worker, and `/health` answers during a
|
||
call.
|
||
- `/health` reads only allocator counters and a device name cached at load, so it adds no
|
||
per-thread CUDA state.
|
||
- **Why one thread:** torch keeps CUDA state per host thread (cuBLAS handles and workspaces),
|
||
and part of it sits outside the VRAM cap.
|
||
- **Measured on 2026-09-30, fv-ml1 GPU 3:** anyio's 40 worker threads added 252 MiB outside the
|
||
cap and 326 MiB inside it. That pushed the footprint past the GPU 1 budget.
|
||
- **INV-3 fail-closed startup.** Before the service serves, all of these must hold:
|
||
- `inference.py` in the checkpoint hashes to the pinned sha256 (it is executed code, loaded
|
||
from a data mount);
|
||
- the checkpoint path ends in `snapshots/<pinned revision>`;
|
||
- the model sits on CUDA, and torch's arch list has the card's `sm_XY`;
|
||
- one warm-up decision scores.
|
||
|
||
`device=cpu` is allowed only when set explicitly.
|
||
- **INV-4 VRAM cap.** `VRAM_CAP_GIB`, when set, becomes `torch.cuda.set_per_process_memory_fraction`
|
||
**before** the weights load.
|
||
- An OOM in any call makes the whole request a 503 `out_of_memory`. So does a RuntimeError
|
||
whose first line says "out of memory". The engine then frees the failed call's frames,
|
||
runs `gc.collect()` and `empty_cache()`, and raises unchained. The process stays up.
|
||
- A `ValueError` (the token limit, or one raised inside a forward) keeps its message and maps
|
||
to 422. It is released and raised unchained the same way.
|
||
- After every call, if reserved memory exceeds the post-warm-up baseline by more than
|
||
`RELEASE_SLACK_MIB` (default 512), the engine runs `empty_cache()`.
|
||
- Any other failure is logged with its traceback and released the same way, then raised
|
||
unchained as `ScoringFailed`.
|
||
- **INV-5 no network.** The entry point sets `HF_HUB_OFFLINE=1` and `TRANSFORMERS_OFFLINE=1`
|
||
before torch or transformers load. The weights are read from the mounted, read-only HF cache.
|
||
- **INV-6 constant-time auth.** The token is compared with `hmac.compare_digest`. It must be at
|
||
least 32 visible ASCII characters (33–126), or startup refuses it.
|
||
- **INV-7 the text-only model is the same model.**
|
||
- The service takes no images, and neither did semif. So after the first warm-up the engine
|
||
replaces the vision tower (`model.model.visual`, 0.62 GiB) with a stub that raises if it is
|
||
ever called.
|
||
- It then scores the warm-up again. It refuses to start unless the answer is bit-identical to
|
||
the first one.
|
||
- `KEEP_VISION=1` keeps the tower.
|
||
- **INV-8 honest prompt hash.** `prompt_sha256` is the sha256 of the chat-template text rendered
|
||
from `inference.compile_row(...)` with the same arguments `HFBackend.encode` uses. At startup,
|
||
that text must tokenise to exactly the `input_tokens` `predict()` reports for the warm-up.
|
||
Otherwise the service refuses to start rather than hash a prompt the model never saw.
|
||
|
||
## Limits and errors
|
||
|
||
- `MAX_TOKENS` (default 8192, the model's own default) is `DecisionEngine(max_length=...)`, and
|
||
applies to a **whole call**: the state plus all its questions. A longer call is a 422. It is
|
||
never truncated; the model raises.
|
||
- `MAX_DECISIONS` (default 64) caps the orderings scored per request, counted after expansion.
|
||
A request must score 1..max of them, else 422.
|
||
- The body may be at most `MAX_BODY_BYTES` (default 1 MiB), else 413. This is checked before
|
||
each chunk is kept. A declared `Content-Length` is trusted only as ASCII digits.
|
||
- **Admission:** at most `MAX_QUEUE` (default 32) POSTs may be queued or scoring at once. The
|
||
next one gets 429 `busy` before its body is read.
|
||
- `workload`: there is no per-workload table, and the model's own calibration always applies.
|
||
So any non-null `workload` is a 422, exactly as semif-serve behaved with its deployed empty
|
||
table.
|
||
|
||
| status | code | when |
|
||
|---|---|---|
|
||
| 401 | `unauthorized` | missing or wrong bearer |
|
||
| 413 | `request_too_large` | body over the limit |
|
||
| 422 | `invalid_request` | bad JSON or shape; a SemIf-rule violation; a model `ValueError` (token limit, reserved `<decision>` marker in the input, ...); a `workload`; `all` over 4 options; a row count outside 1..`MAX_DECISIONS` |
|
||
| 429 | `busy` | `MAX_QUEUE` requests in progress |
|
||
| 503 | `out_of_memory` | CUDA OOM in any call of the request |
|
||
| 500 | `scoring_failed` | any other model failure, including one while building the response |
|
||
|
||
The error body is `{error: {code, message}}`.
|
||
|
||
## Configuration (env, prefix `INTERN_DECISION_`)
|
||
|
||
- `API_TOKEN` is required.
|
||
- `CHECKPOINT` defaults to
|
||
`/hf/hub/models--internlm--Intern-Decision-4B/snapshots/<revision>`.
|
||
- `DEVICE` defaults to `cuda`.
|
||
- `VRAM_CAP_GIB` has no default: unset means uncapped, and when set it must be finite and > 0.
|
||
- `MAX_TOKENS`, `MAX_DECISIONS`, `MAX_BODY_BYTES`, `MAX_QUEUE` and `RELEASE_SLACK_MIB` are
|
||
integers; each must be ≥ 1, except the slack, which must be ≥ 0.
|
||
- `KEEP_VISION` is `0` or `1`.
|
||
|
||
A bad value is refused at startup with a `ValueError` naming the variable. In the stack's `.env`
|
||
on the host, the cap is the single knob `VRAM_CAP_GIB`.
|
||
|
||
## Deltas from semif-serve (deliberate)
|
||
|
||
1. **Prompt and model.** The prompt is Intern-Decision's own Jev prompt.
|
||
- **Option ids are shown to the model** (`A = <id>: <description>`); SemIf showed only the
|
||
descriptions. An option id is therefore part of the question, so give options meaningful
|
||
or neutral ids.
|
||
- In the bench negative control, the ids-in-prompt cue made the top stay on 10/144 rows
|
||
after the descriptions moved, against SemIf's 14/144.
|
||
2. **Shared requests are one prompt, not independent rows.** The questions in a call are asked
|
||
together.
|
||
- A decision's answer can depend on the other questions in its call and on their order. In
|
||
the bench, Wyrd scored 79/84 asked one decision at a time and 77/84 with a turn's 4
|
||
decisions in one prompt.
|
||
- SemIf only shared a KV prefix, so each of its rows was independent.
|
||
- Calls hold at most 16 questions, and the chunk boundaries follow request order.
|
||
3. **`probabilities` are temperature-scaled** by the checkpoint's shipped calibration
|
||
(T = 1.99241824).
|
||
- `probability_status` says so, and `calibration` carries the method and T.
|
||
- SemIf's were raw softmax, labelled uncalibrated. The argmax is the same either way.
|
||
- This is the vendor's calibration on the vendor's data, not ours.
|
||
4. **`option_logits` does not exist.** `predict()` does not expose logits. T=1 scores could only
|
||
be derived up to a constant, which would be a derived number, not logits.
|
||
5. **New fields:** `top`, `confidence`, `calibration`, `native`, `call`.
|
||
6. **`prompt_sha256` and `input_tokens` describe the whole call**, shared by every decision in
|
||
it. SemIf's were per row.
|
||
7. **`prompt_version`** names the pinned `inference.py`. **`readout`** and **`model`** describe
|
||
Intern-Decision.
|
||
8. **`workload` is always a 422, and `/health.workloads` is `[]`.** There is no per-workload
|
||
temperature table. This matches the deployed semif, whose table was empty.
|
||
9. **`/health`**: `semif_commit` is gone; the model's pins live in `model`. It adds
|
||
`max_questions_per_call` and `chunking`.
|
||
10. **`/decide/shared` timing:** SemIf's prefix-cache fields cannot exist, because there is no
|
||
prefix cache: `prefix_tokens`, `prefill_seconds`, `replicate_seconds`,
|
||
`suffix_forward_seconds`, `true_suffix_tokens`, `padded_suffix_tokens` and `encode_seconds`.
|
||
The timing adds `calls`, `questions_per_call`, `input_tokens` and `inference_seconds`.
|
||
`total_seconds` and `batch_size` keep their meaning.
|
||
11. **`MAX_TOKENS` is per call** (state plus up to 16 questions) and defaults to 8192. SemIf's
|
||
was per row and defaulted to 4096.
|
||
12. **Orderings run in waves**, one forward pass per wave per chunk. SemIf batched every
|
||
ordering in one shared forward.
|
||
13. **Ties.** Per-ordering `top` uses the model's argmax, whose exact tie goes to the smaller
|
||
option id string. SemIf used the first maximum in the ordering. `combined.top` keeps semif's
|
||
rule.
|
||
14. **Images.** The vision tower is dropped (INV-7). The surface never took images.
|
||
|
||
## Tests (TDD, fake engine: no torch, no model)
|
||
|
||
- **Auth.** A POST without the right bearer is 401 and never reaches the engine. `/health`
|
||
needs no auth. A short or non-visible-ASCII token is refused at startup.
|
||
- **Mapping.**
|
||
- `/decide` sends one call with field `q` and the criteria in the caller's order.
|
||
- `/decide/shared` sends `q1..qN` over the shared state.
|
||
- 17 questions become 2 calls (16 + 1), and 40 become 3 (16 + 16 + 8). Results come back in
|
||
request order, and `call.index` and `call.field` are right.
|
||
- Decision ids never appear in a call.
|
||
- **Result.** `probabilities` follows `option_ids`. `top`, `confidence`, `calibration`, `native`
|
||
and `input_tokens` are passed through. An answer missing an option id is a 500.
|
||
- **Orderings.**
|
||
- `rotations` puts each option in each position once, and the waves never repeat a decision
|
||
within a call.
|
||
- `all` is n!, and 422 above 4 options.
|
||
- A position bias cancels exactly.
|
||
- `agreement` and `spread` are computed from the orderings.
|
||
- A mixed request keeps plain results unchanged, and wave 0 equals the plain request.
|
||
- Orderings count toward the cap.
|
||
- **Validation (422).** Fewer than 2 or more than 16 options, duplicate option ids, duplicate
|
||
decision ids, an empty id or question, an empty or non-finite state, a `workload`, a row count
|
||
outside 1..max, malformed JSON, and a model `ValueError`.
|
||
- **Limits.** A body over the limit is 413. A queue past `MAX_QUEUE` is 429 before the body is
|
||
read.
|
||
- **Engine failures.** An engine `OutOfMemory` is 503, and any other failure is 500.
|
||
- **Concurrency.** Requests are serialised: two never overlap inside the engine, and the
|
||
chunks of one request are not interleaved with another's. `/health` answers while a call is
|
||
blocked. Every call runs on one thread, which is the thread the engine was loaded on.
|
||
- **Engine against a fake torch.**
|
||
- An OOM is re-raised unchained, and `empty_cache` runs only after the failed call's tensors
|
||
are freed.
|
||
- A RuntimeError saying "out of memory" becomes `OutOfMemory`.
|
||
- Another failure becomes `ScoringFailed`, unchained.
|
||
- A `ValueError` keeps its message but is released and raised unchained.
|
||
- A request the prompt builder rejects never reaches the model.
|
||
- A burst over baseline plus slack is released, and one at or under it is left alone.
|
||
- **Engine load (fake).**
|
||
- A wrong `inference.py` hash, or a checkpoint path that is not the pinned snapshot, refuses
|
||
to start.
|
||
- The cap is applied before the engine is constructed.
|
||
- The vision swap refuses to start when the warm-up changes.
|
||
- The prompt-hash check refuses to start on a token-count mismatch.
|
||
- **Config.** Every value is validated.
|
||
|
||
## Acceptance (fv-ml1, real model; not unit tests)
|
||
|
||
1. **Positive control:** through the service, the bench's native numbers on the pooled 259 rows
|
||
and on Wyrd, single ordering (bench: 240/259, Wyrd 79/84; floor: the pooled set resolves
|
||
±4 pts, and 0 labels moved across 4 restarts). Also a row-by-row comparison against the
|
||
bench's own rows.
|
||
2. **Negative control:** descriptions rotated one place. The top follows the moved description
|
||
(bench: 122/144 follow, 10/144 same top).
|
||
3. A-vs-A repeat stability, within the process and across restarts.
|
||
4. Latency at our shape: 21 binary criteria, and 16 criteria over the ~3,900-token state, both
|
||
server-side and from nh3-dev.
|
||
5. Resident and peak VRAM (nvidia-smi and torch), and the cap chosen from them.
|
||
6. An over-cap request is a 503, memory returns to baseline, and the service keeps answering.
|
||
7. A `/decide/shared` with more than 16 questions is chunked, and its answers equal the same
|
||
questions asked chunk by chunk.
|
||
8. 401 without the token, and 429 past the queue.
|