Files
esh-pfi-infrastructure/services/intern-decision-serve/intern-decision-serve.contract.md
T
vh 5bbf0aaeba feat(intern-decision-serve): Intern-Decision-4B behind semif-serve's HTTP surface
Contract, service and tests (fake engine, no GPU). Scores through the checkpoint's own
inference.py (DecisionEngine.predict, sha256-pinned); maps semif decisions onto Jev choice
questions, packs /decide/shared into calls of at most 16, runs orderings in waves, and keeps
semif's error mapping, admission, body limit and hard VRAM cap. Deltas from semif-serve are
listed in the contract.
2026-09-30 09:04:39 -07:00

17 KiB
Raw Blame History

title, kind, status, owner, created, replaces, depends_on
title kind status owner created replaces depends_on
intern-decision-serve module-contract draft infra-ops 2026-09-30 semif-serve 0.1.4 (services/semif-serve/semif-serve.contract.md), external surface kept
internlm/Intern-Decision-4B at revision 0e5e6aa7d6d750e2b1504ba11a8136cb58aeb3cd, BF16 (Apache-2.0)
that snapshot's own inference.py (DecisionEngine.predict), sha256 c904e2c67ca0775621a22375ee373d2ba30b52117cda870c6c9ef74143b29863
torch 2.10.0+cu128, transformers 5.17.0, flash-linear-attention 0.5.2, causal-conv1d 1.7.0 (the 2026-09-30 bench stack)

intern-decision-serve: Intern-Decision-4B behind semif-serve's HTTP surface

Purpose

Prime, 2026-09-30: "replace semif with intern-decision now". The bench (docs/pfi/jev-candidates-bench-2026-09-30.md) picked Intern-Decision-4B on its own runtime. This service loads the model once and scores every request with the checkpoint's own inference.py (DecisionEngine.predict). It re-implements neither the prompt nor the readout. It keeps semif-serve's external surface, so a caller written for semif-serve works unchanged. It only maps semif-shaped requests onto the model's Jev request schema and maps the answers back. Every deliberate difference is listed under Deltas from semif-serve.

Endpoints (same as semif-serve)

Every POST takes and returns JSON and needs Authorization: Bearer <token>. GET /health is open.

method path body success
GET /health none 200 {status: "ok", model, vram_cap_gib, max_tokens, max_decisions, max_questions_per_call: 16, chunking, workloads: []}
POST /decide {id, state, question, options[2..16], orderings?, workload?} 200 one decision result
POST /decide/shared {state, decisions: [{id, question, options, orderings?}], workload?} 200 {results: [...], timing: {...}}, results in request order

options items are {id, description}. Validation is SemIf's, re-stated here because SemIf is gone. id and question are nonempty strings. state is a nonempty string, object or array, and must be finite JSON. There are 2..16 options, each with a string id and a string description, and option ids are unique within a decision. Decision ids are unique within a request. A violation is a 422.

Mapping onto the model (the seam)

  • One call is one predict() request: {"state": state, "questions": {<field>: {"type": "choice", "instructions": question, "criteria": {option.id: option.description, ...}}}}. The criteria keep the caller's option order. This is the format the bench measured.
  • Field names are positional: q when a call carries one question, and q1..qN in request order when it carries several. Decision ids never reach the prompt. Option ids do: the model prints A = <id>: <description>.
  • /decide is one call with one question.
  • /decide/shared is packed into calls of at most 16 questions, the model's own limit (validate_request). The packing is greedy and in request order: questions 1–16, then 17–32, and so on. Every call of a request runs back to back under the inference lock. /health reports max_questions_per_call: 16 and the rule in chunking.
  • Orderings. A decision may set orderings: "none" | "rotations" | "all", with semif's meaning. rotations gives the n cyclic shifts, the caller's order first. all gives the n! permutations, the caller's order first, and is 422 above 4 options. Ordering k of every decision that has more than k orderings forms wave k. Each wave is packed as above. So wave 0 is the request exactly as written, and no prompt ever holds two orderings of the same decision. Every ordering counts toward MAX_DECISIONS.
  • The result for one ordering (and for every plain decision):
    {id, option_ids, probabilities, top, confidence, calibration, native,
     input_tokens, prompt_sha256, prompt_version, model, readout, probability_status,
     call: {index, field, questions}}
    
    • native is the answer object predict() returned for that field, unchanged.
    • probabilities is native.probabilities listed in option_ids order.
    • top is native.choice and confidence is native.confidence.
    • calibration is the calibration object from predict().
    • input_tokens and prompt_sha256 describe the whole call.
    • /decide results (plain) also carry total_seconds and forward_seconds. forward_seconds is predict()'s timing.inference_ms / 1000.
  • Averaged result: semif's shape, {id, option_ids, combined: {method, orderings, probabilities, top, agreement, spread}, orderings: [...]}.
    • The per-ordering ids are <id>#o<k>.
    • combined.probabilities renormalises the per-option mean of log p. A probability of exactly 0 is floored at 1e-300 before the log.
    • combined.top is the first maximum in the caller's order.
    • agreement is the share of orderings whose top equals combined.top.
    • spread holds each option's min and max probabilities across orderings.
    • Temperature scaling is one monotone transform per ordering, so combined.top and agreement are what combining T=1 scores would give.
  • /decide/shared timing: {total_seconds, batch_size (orderings scored), calls, questions_per_call: [...], input_tokens: [...], inference_seconds}.

Invariants

  • INV-1 pass-through. Every number in native, probabilities, top, confidence, calibration and input_tokens is what predict() returned. The wrapper only re-keys it. If an answer lacks one of the decision's option ids, that is a 500 scoring_failed, never a guess.

  • INV-2 one model, one inference at a time. The model loads at startup, and a process-wide lock serialises every request's calls (all of a request's chunks run inside one hold). The app runs one worker. Calls run off the event loop, so /health answers during one.

  • INV-3 fail-closed startup. Before the service serves, all of these must hold:

    • inference.py in the checkpoint hashes to the pinned sha256 (it is executed code, loaded from a data mount);
    • the checkpoint path ends in snapshots/<pinned revision>;
    • the model sits on CUDA, and torch's arch list has the card's sm_XY;
    • one warm-up decision scores.

    device=cpu is allowed only when set explicitly.

  • INV-4 VRAM cap. VRAM_CAP_GIB, when set, becomes torch.cuda.set_per_process_memory_fraction before the weights load.

    • An OOM in any call makes the whole request a 503 out_of_memory. So does a RuntimeError whose first line says "out of memory". The engine then frees the failed call's frames, runs gc.collect() and empty_cache(), and raises unchained. The process stays up.
    • After every call, if reserved memory exceeds the post-warm-up baseline by more than RELEASE_SLACK_MIB (default 512), the engine runs empty_cache().
    • Any other failure except ValueError is logged with its traceback and released the same way, then raised unchained as ScoringFailed.
  • INV-5 no network. The entry point sets HF_HUB_OFFLINE=1 and TRANSFORMERS_OFFLINE=1 before torch or transformers load. The weights are read from the mounted, read-only HF cache.

  • INV-6 constant-time auth. The token is compared with hmac.compare_digest. It must be at least 32 visible ASCII characters (33–126), or startup refuses it.

  • INV-7 the text-only model is the same model.

    • The service takes no images, and neither did semif. So after the first warm-up the engine replaces the vision tower (model.model.visual, 0.62 GiB) with a stub that raises if it is ever called.
    • It then scores the warm-up again. It refuses to start unless the answer is bit-identical to the first one.
    • KEEP_VISION=1 keeps the tower.
  • INV-8 honest prompt hash. prompt_sha256 is the sha256 of the chat-template text rendered from inference.compile_row(...) with the same arguments HFBackend.encode uses. At startup, that text must tokenise to exactly the input_tokens predict() reports for the warm-up. Otherwise the service refuses to start rather than hash a prompt the model never saw.

Limits and errors

  • MAX_TOKENS (default 8192, the model's own default) is DecisionEngine(max_length=...), and applies to a whole call: the state plus all its questions. A longer call is a 422. It is never truncated; the model raises.
  • MAX_DECISIONS (default 64) caps the orderings scored per request, counted after expansion. A request must score 1..max of them, else 422.
  • The body may be at most MAX_BODY_BYTES (default 1 MiB), else 413. This is checked before each chunk is kept. A declared Content-Length is trusted only as ASCII digits.
  • Admission: at most MAX_QUEUE (default 32) POSTs may be queued or scoring at once. The next one gets 429 busy before its body is read.
  • workload: there is no per-workload table, and the model's own calibration always applies. So any non-null workload is a 422, exactly as semif-serve behaved with its deployed empty table.
status code when
401 unauthorized missing or wrong bearer
413 request_too_large body over the limit
422 invalid_request bad JSON or shape; a SemIf-rule violation; a model ValueError (token limit, reserved <decision> marker in the input, ...); a workload; all over 4 options; a row count outside 1..MAX_DECISIONS
429 busy MAX_QUEUE requests in progress
503 out_of_memory CUDA OOM in any call of the request
500 scoring_failed any other model failure, including one while building the response

The error body is {error: {code, message}}.

Configuration (env, prefix INTERN_DECISION_)

  • API_TOKEN is required.
  • CHECKPOINT defaults to /hf/hub/models--internlm--Intern-Decision-4B/snapshots/<revision>.
  • DEVICE defaults to cuda.
  • VRAM_CAP_GIB has no default: unset means uncapped, and when set it must be finite and > 0.
  • MAX_TOKENS, MAX_DECISIONS, MAX_BODY_BYTES, MAX_QUEUE and RELEASE_SLACK_MIB are integers; each must be ≥ 1, except the slack, which must be ≥ 0.
  • KEEP_VISION is 0 or 1.

A bad value is refused at startup with a ValueError naming the variable. In the stack's .env on the host, the cap is the single knob VRAM_CAP_GIB.

Deltas from semif-serve (deliberate)

  1. Prompt and model. The prompt is Intern-Decision's own Jev prompt.
    • Option ids are shown to the model (A = <id>: <description>); SemIf showed only the descriptions. An option id is therefore part of the question, so give options meaningful or neutral ids.
    • In the bench negative control, the ids-in-prompt cue made the top stay on 10/144 rows after the descriptions moved, against SemIf's 14/144.
  2. Shared requests are one prompt, not independent rows. The questions in a call are asked together.
    • A decision's answer can depend on the other questions in its call and on their order. In the bench, Wyrd scored 79/84 asked one decision at a time and 77/84 with a turn's 4 decisions in one prompt.
    • SemIf only shared a KV prefix, so each of its rows was independent.
    • Calls hold at most 16 questions, and the chunk boundaries follow request order.
  3. probabilities are temperature-scaled by the checkpoint's shipped calibration (T = 1.99241824).
    • probability_status says so, and calibration carries the method and T.
    • SemIf's were raw softmax, labelled uncalibrated. The argmax is the same either way.
    • This is the vendor's calibration on the vendor's data, not ours.
  4. option_logits does not exist. predict() does not expose logits. T=1 scores could only be derived up to a constant, which would be a derived number, not logits.
  5. New fields: top, confidence, calibration, native, call.
  6. prompt_sha256 and input_tokens describe the whole call, shared by every decision in it. SemIf's were per row.
  7. prompt_version names the pinned inference.py. readout and model describe Intern-Decision.
  8. workload is always a 422, and /health.workloads is []. There is no per-workload temperature table. This matches the deployed semif, whose table was empty.
  9. /health: semif_commit is gone; the model's pins live in model. It adds max_questions_per_call and chunking.
  10. /decide/shared timing: SemIf's prefix-cache fields cannot exist, because there is no prefix cache: prefix_tokens, prefill_seconds, replicate_seconds, suffix_forward_seconds, true_suffix_tokens, padded_suffix_tokens and encode_seconds. The timing adds calls, questions_per_call, input_tokens and inference_seconds. total_seconds and batch_size keep their meaning.
  11. MAX_TOKENS is per call (state plus up to 16 questions) and defaults to 8192. SemIf's was per row and defaulted to 4096.
  12. Orderings run in waves, one forward pass per wave per chunk. SemIf batched every ordering in one shared forward.
  13. Ties. Per-ordering top uses the model's argmax, whose exact tie goes to the smaller option id string. SemIf used the first maximum in the ordering. combined.top keeps semif's rule.
  14. Images. The vision tower is dropped (INV-7). The surface never took images.

Tests (TDD, fake engine: no torch, no model)

  • Auth. A POST without the right bearer is 401 and never reaches the engine. /health needs no auth. A short or non-visible-ASCII token is refused at startup.
  • Mapping.
    • /decide sends one call with field q and the criteria in the caller's order.
    • /decide/shared sends q1..qN over the shared state.
    • 17 questions become 2 calls (16 + 1), and 40 become 3 (16 + 16 + 8). Results come back in request order, and call.index and call.field are right.
    • Decision ids never appear in a call.
  • Result. probabilities follows option_ids. top, confidence, calibration, native and input_tokens are passed through. An answer missing an option id is a 500.
  • Orderings.
    • rotations puts each option in each position once, and the waves never repeat a decision within a call.
    • all is n!, and 422 above 4 options.
    • A position bias cancels exactly.
    • agreement and spread are computed from the orderings.
    • A mixed request keeps plain results unchanged, and wave 0 equals the plain request.
    • Orderings count toward the cap.
  • Validation (422). Fewer than 2 or more than 16 options, duplicate option ids, duplicate decision ids, an empty id or question, an empty or non-finite state, a workload, a row count outside 1..max, malformed JSON, and a model ValueError.
  • Limits. A body over the limit is 413. A queue past MAX_QUEUE is 429 before the body is read.
  • Engine failures. An engine OutOfMemory is 503, and any other failure is 500.
  • Concurrency. Requests are serialised: two never overlap inside the engine, and the chunks of one request are not interleaved with another's. /health answers while a call is blocked.
  • Engine against a fake torch.
    • An OOM is re-raised unchained, and empty_cache runs only after the failed call's tensors are freed.
    • A RuntimeError saying "out of memory" becomes OutOfMemory.
    • Another failure becomes ScoringFailed, unchained.
    • ValueError passes through.
    • A burst over baseline plus slack is released, and one at or under it is left alone.
  • Engine load (fake).
    • A wrong inference.py hash, or a checkpoint path that is not the pinned snapshot, refuses to start.
    • The cap is applied before the engine is constructed.
    • The vision swap refuses to start when the warm-up changes.
    • The prompt-hash check refuses to start on a token-count mismatch.
  • Config. Every value is validated.

Acceptance (fv-ml1, real model; not unit tests)

  1. Positive control: through the service, the bench's native numbers on the pooled 259 rows and on Wyrd, single ordering (bench: 240/259, Wyrd 79/84; floor: the pooled set resolves ±4 pts, and 0 labels moved across 4 restarts). Also a row-by-row comparison against the bench's own rows.
  2. Negative control: descriptions rotated one place. The top follows the moved description (bench: 122/144 follow, 10/144 same top).
  3. A-vs-A repeat stability, within the process and across restarts.
  4. Latency at our shape: 21 binary criteria, and 16 criteria over the ~3,900-token state, both server-side and from nh3-dev.
  5. Resident and peak VRAM (nvidia-smi and torch), and the cap chosen from them.
  6. An over-cap request is a 503, memory returns to baseline, and the service keeps answering.
  7. A /decide/shared with more than 16 questions is chunked, and its answers equal the same questions asked chunk by chunk.
  8. 401 without the token, and 429 past the queue.