Contract, service and tests (fake engine, no GPU). Scores through the checkpoint's own inference.py (DecisionEngine.predict, sha256-pinned); maps semif decisions onto Jev choice questions, packs /decide/shared into calls of at most 16, runs orderings in waves, and keeps semif's error mapping, admission, body limit and hard VRAM cap. Deltas from semif-serve are listed in the contract.
17 KiB
title, kind, status, owner, created, replaces, depends_on
| title | kind | status | owner | created | replaces | depends_on | |||
|---|---|---|---|---|---|---|---|---|---|
| intern-decision-serve | module-contract | draft | infra-ops | 2026-09-30 | semif-serve 0.1.4 (services/semif-serve/semif-serve.contract.md), external surface kept |
|
intern-decision-serve: Intern-Decision-4B behind semif-serve's HTTP surface
Purpose
Prime, 2026-09-30: "replace semif with intern-decision now". The bench
(docs/pfi/jev-candidates-bench-2026-09-30.md) picked Intern-Decision-4B on its own runtime.
This service loads the model once and scores every request with the checkpoint's own
inference.py (DecisionEngine.predict). It re-implements neither the prompt nor the readout.
It keeps semif-serve's external surface, so a caller written for semif-serve works unchanged.
It only maps semif-shaped requests onto the model's Jev request schema and maps the answers back.
Every deliberate difference is listed under Deltas from semif-serve.
Endpoints (same as semif-serve)
Every POST takes and returns JSON and needs Authorization: Bearer <token>. GET /health is open.
| method | path | body | success |
|---|---|---|---|
| GET | /health |
none | 200 {status: "ok", model, vram_cap_gib, max_tokens, max_decisions, max_questions_per_call: 16, chunking, workloads: []} |
| POST | /decide |
{id, state, question, options[2..16], orderings?, workload?} |
200 one decision result |
| POST | /decide/shared |
{state, decisions: [{id, question, options, orderings?}], workload?} |
200 {results: [...], timing: {...}}, results in request order |
options items are {id, description}. Validation is SemIf's, re-stated here because SemIf is
gone. id and question are nonempty strings. state is a nonempty string, object or array,
and must be finite JSON. There are 2..16 options, each with a string id and a string
description, and option ids are unique within a decision. Decision ids are unique within a
request. A violation is a 422.
Mapping onto the model (the seam)
- One call is one
predict()request:{"state": state, "questions": {<field>: {"type": "choice", "instructions": question, "criteria": {option.id: option.description, ...}}}}. The criteria keep the caller's option order. This is the format the bench measured. - Field names are positional:
qwhen a call carries one question, andq1..qNin request order when it carries several. Decision ids never reach the prompt. Option ids do: the model printsA = <id>: <description>. /decideis one call with one question./decide/sharedis packed into calls of at most 16 questions, the model's own limit (validate_request). The packing is greedy and in request order: questions 1–16, then 17–32, and so on. Every call of a request runs back to back under the inference lock./healthreportsmax_questions_per_call: 16and the rule inchunking.- Orderings. A decision may set
orderings: "none" | "rotations" | "all", with semif's meaning.rotationsgives the n cyclic shifts, the caller's order first.allgives the n! permutations, the caller's order first, and is 422 above 4 options. Ordering k of every decision that has more than k orderings forms wave k. Each wave is packed as above. So wave 0 is the request exactly as written, and no prompt ever holds two orderings of the same decision. Every ordering counts towardMAX_DECISIONS. - The result for one ordering (and for every plain decision):
{id, option_ids, probabilities, top, confidence, calibration, native, input_tokens, prompt_sha256, prompt_version, model, readout, probability_status, call: {index, field, questions}}nativeis the answer objectpredict()returned for that field, unchanged.probabilitiesisnative.probabilitieslisted inoption_idsorder.topisnative.choiceandconfidenceisnative.confidence.calibrationis thecalibrationobject frompredict().input_tokensandprompt_sha256describe the whole call./decideresults (plain) also carrytotal_secondsandforward_seconds.forward_secondsispredict()'stiming.inference_ms/ 1000.
- Averaged result: semif's shape,
{id, option_ids, combined: {method, orderings, probabilities, top, agreement, spread}, orderings: [...]}.- The per-ordering ids are
<id>#o<k>. combined.probabilitiesrenormalises the per-option mean oflog p. A probability of exactly 0 is floored at 1e-300 before the log.combined.topis the first maximum in the caller's order.agreementis the share of orderings whosetopequalscombined.top.spreadholds each option's min and maxprobabilitiesacross orderings.- Temperature scaling is one monotone transform per ordering, so
combined.topandagreementare what combining T=1 scores would give.
- The per-ordering ids are
/decide/sharedtiming:{total_seconds, batch_size (orderings scored), calls, questions_per_call: [...], input_tokens: [...], inference_seconds}.
Invariants
-
INV-1 pass-through. Every number in
native,probabilities,top,confidence,calibrationandinput_tokensis whatpredict()returned. The wrapper only re-keys it. If an answer lacks one of the decision's option ids, that is a 500scoring_failed, never a guess. -
INV-2 one model, one inference at a time. The model loads at startup, and a process-wide lock serialises every request's calls (all of a request's chunks run inside one hold). The app runs one worker. Calls run off the event loop, so
/healthanswers during one. -
INV-3 fail-closed startup. Before the service serves, all of these must hold:
inference.pyin the checkpoint hashes to the pinned sha256 (it is executed code, loaded from a data mount);- the checkpoint path ends in
snapshots/<pinned revision>; - the model sits on CUDA, and torch's arch list has the card's
sm_XY; - one warm-up decision scores.
device=cpuis allowed only when set explicitly. -
INV-4 VRAM cap.
VRAM_CAP_GIB, when set, becomestorch.cuda.set_per_process_memory_fractionbefore the weights load.- An OOM in any call makes the whole request a 503
out_of_memory. So does a RuntimeError whose first line says "out of memory". The engine then frees the failed call's frames, runsgc.collect()andempty_cache(), and raises unchained. The process stays up. - After every call, if reserved memory exceeds the post-warm-up baseline by more than
RELEASE_SLACK_MIB(default 512), the engine runsempty_cache(). - Any other failure except
ValueErroris logged with its traceback and released the same way, then raised unchained asScoringFailed.
- An OOM in any call makes the whole request a 503
-
INV-5 no network. The entry point sets
HF_HUB_OFFLINE=1andTRANSFORMERS_OFFLINE=1before torch or transformers load. The weights are read from the mounted, read-only HF cache. -
INV-6 constant-time auth. The token is compared with
hmac.compare_digest. It must be at least 32 visible ASCII characters (33–126), or startup refuses it. -
INV-7 the text-only model is the same model.
- The service takes no images, and neither did semif. So after the first warm-up the engine
replaces the vision tower (
model.model.visual, 0.62 GiB) with a stub that raises if it is ever called. - It then scores the warm-up again. It refuses to start unless the answer is bit-identical to the first one.
KEEP_VISION=1keeps the tower.
- The service takes no images, and neither did semif. So after the first warm-up the engine
replaces the vision tower (
-
INV-8 honest prompt hash.
prompt_sha256is the sha256 of the chat-template text rendered frominference.compile_row(...)with the same argumentsHFBackend.encodeuses. At startup, that text must tokenise to exactly theinput_tokenspredict()reports for the warm-up. Otherwise the service refuses to start rather than hash a prompt the model never saw.
Limits and errors
MAX_TOKENS(default 8192, the model's own default) isDecisionEngine(max_length=...), and applies to a whole call: the state plus all its questions. A longer call is a 422. It is never truncated; the model raises.MAX_DECISIONS(default 64) caps the orderings scored per request, counted after expansion. A request must score 1..max of them, else 422.- The body may be at most
MAX_BODY_BYTES(default 1 MiB), else 413. This is checked before each chunk is kept. A declaredContent-Lengthis trusted only as ASCII digits. - Admission: at most
MAX_QUEUE(default 32) POSTs may be queued or scoring at once. The next one gets 429busybefore its body is read. workload: there is no per-workload table, and the model's own calibration always applies. So any non-nullworkloadis a 422, exactly as semif-serve behaved with its deployed empty table.
| status | code | when |
|---|---|---|
| 401 | unauthorized |
missing or wrong bearer |
| 413 | request_too_large |
body over the limit |
| 422 | invalid_request |
bad JSON or shape; a SemIf-rule violation; a model ValueError (token limit, reserved <decision> marker in the input, ...); a workload; all over 4 options; a row count outside 1..MAX_DECISIONS |
| 429 | busy |
MAX_QUEUE requests in progress |
| 503 | out_of_memory |
CUDA OOM in any call of the request |
| 500 | scoring_failed |
any other model failure, including one while building the response |
The error body is {error: {code, message}}.
Configuration (env, prefix INTERN_DECISION_)
API_TOKENis required.CHECKPOINTdefaults to/hf/hub/models--internlm--Intern-Decision-4B/snapshots/<revision>.DEVICEdefaults tocuda.VRAM_CAP_GIBhas no default: unset means uncapped, and when set it must be finite and > 0.MAX_TOKENS,MAX_DECISIONS,MAX_BODY_BYTES,MAX_QUEUEandRELEASE_SLACK_MIBare integers; each must be ≥ 1, except the slack, which must be ≥ 0.KEEP_VISIONis0or1.
A bad value is refused at startup with a ValueError naming the variable. In the stack's .env
on the host, the cap is the single knob VRAM_CAP_GIB.
Deltas from semif-serve (deliberate)
- Prompt and model. The prompt is Intern-Decision's own Jev prompt.
- Option ids are shown to the model (
A = <id>: <description>); SemIf showed only the descriptions. An option id is therefore part of the question, so give options meaningful or neutral ids. - In the bench negative control, the ids-in-prompt cue made the top stay on 10/144 rows after the descriptions moved, against SemIf's 14/144.
- Option ids are shown to the model (
- Shared requests are one prompt, not independent rows. The questions in a call are asked
together.
- A decision's answer can depend on the other questions in its call and on their order. In the bench, Wyrd scored 79/84 asked one decision at a time and 77/84 with a turn's 4 decisions in one prompt.
- SemIf only shared a KV prefix, so each of its rows was independent.
- Calls hold at most 16 questions, and the chunk boundaries follow request order.
probabilitiesare temperature-scaled by the checkpoint's shipped calibration (T = 1.99241824).probability_statussays so, andcalibrationcarries the method and T.- SemIf's were raw softmax, labelled uncalibrated. The argmax is the same either way.
- This is the vendor's calibration on the vendor's data, not ours.
option_logitsdoes not exist.predict()does not expose logits. T=1 scores could only be derived up to a constant, which would be a derived number, not logits.- New fields:
top,confidence,calibration,native,call. prompt_sha256andinput_tokensdescribe the whole call, shared by every decision in it. SemIf's were per row.prompt_versionnames the pinnedinference.py.readoutandmodeldescribe Intern-Decision.workloadis always a 422, and/health.workloadsis[]. There is no per-workload temperature table. This matches the deployed semif, whose table was empty./health:semif_commitis gone; the model's pins live inmodel. It addsmax_questions_per_callandchunking./decide/sharedtiming: SemIf's prefix-cache fields cannot exist, because there is no prefix cache:prefix_tokens,prefill_seconds,replicate_seconds,suffix_forward_seconds,true_suffix_tokens,padded_suffix_tokensandencode_seconds. The timing addscalls,questions_per_call,input_tokensandinference_seconds.total_secondsandbatch_sizekeep their meaning.MAX_TOKENSis per call (state plus up to 16 questions) and defaults to 8192. SemIf's was per row and defaulted to 4096.- Orderings run in waves, one forward pass per wave per chunk. SemIf batched every ordering in one shared forward.
- Ties. Per-ordering
topuses the model's argmax, whose exact tie goes to the smaller option id string. SemIf used the first maximum in the ordering.combined.topkeeps semif's rule. - Images. The vision tower is dropped (INV-7). The surface never took images.
Tests (TDD, fake engine: no torch, no model)
- Auth. A POST without the right bearer is 401 and never reaches the engine.
/healthneeds no auth. A short or non-visible-ASCII token is refused at startup. - Mapping.
/decidesends one call with fieldqand the criteria in the caller's order./decide/sharedsendsq1..qNover the shared state.- 17 questions become 2 calls (16 + 1), and 40 become 3 (16 + 16 + 8). Results come back in
request order, and
call.indexandcall.fieldare right. - Decision ids never appear in a call.
- Result.
probabilitiesfollowsoption_ids.top,confidence,calibration,nativeandinput_tokensare passed through. An answer missing an option id is a 500. - Orderings.
rotationsputs each option in each position once, and the waves never repeat a decision within a call.allis n!, and 422 above 4 options.- A position bias cancels exactly.
agreementandspreadare computed from the orderings.- A mixed request keeps plain results unchanged, and wave 0 equals the plain request.
- Orderings count toward the cap.
- Validation (422). Fewer than 2 or more than 16 options, duplicate option ids, duplicate
decision ids, an empty id or question, an empty or non-finite state, a
workload, a row count outside 1..max, malformed JSON, and a modelValueError. - Limits. A body over the limit is 413. A queue past
MAX_QUEUEis 429 before the body is read. - Engine failures. An engine
OutOfMemoryis 503, and any other failure is 500. - Concurrency. Requests are serialised: two never overlap inside the engine, and the
chunks of one request are not interleaved with another's.
/healthanswers while a call is blocked. - Engine against a fake torch.
- An OOM is re-raised unchained, and
empty_cacheruns only after the failed call's tensors are freed. - A RuntimeError saying "out of memory" becomes
OutOfMemory. - Another failure becomes
ScoringFailed, unchained. ValueErrorpasses through.- A burst over baseline plus slack is released, and one at or under it is left alone.
- An OOM is re-raised unchained, and
- Engine load (fake).
- A wrong
inference.pyhash, or a checkpoint path that is not the pinned snapshot, refuses to start. - The cap is applied before the engine is constructed.
- The vision swap refuses to start when the warm-up changes.
- The prompt-hash check refuses to start on a token-count mismatch.
- A wrong
- Config. Every value is validated.
Acceptance (fv-ml1, real model; not unit tests)
- Positive control: through the service, the bench's native numbers on the pooled 259 rows and on Wyrd, single ordering (bench: 240/259, Wyrd 79/84; floor: the pooled set resolves ±4 pts, and 0 labels moved across 4 restarts). Also a row-by-row comparison against the bench's own rows.
- Negative control: descriptions rotated one place. The top follows the moved description (bench: 122/144 follow, 10/144 same top).
- A-vs-A repeat stability, within the process and across restarts.
- Latency at our shape: 21 binary criteria, and 16 criteria over the ~3,900-token state, both server-side and from nh3-dev.
- Resident and peak VRAM (nvidia-smi and torch), and the cap chosen from them.
- An over-cap request is a 503, memory returns to baseline, and the service keeps answering.
- A
/decide/sharedwith more than 16 questions is chunked, and its answers equal the same questions asked chunk by chunk. - 401 without the token, and 429 past the queue.