# intern-decision > **LIVE since 2026-09-30 0941 PT** on fv-ml1 GPU 1, with cap 9.0 GiB and `MAX_TOKENS` 7,168. > Acceptance on the live service is below. `semif` is stopped and kept as the rollback. **Intern-Decision-4B typed decisions** on **fv-ml1 GPU 1**, the utility card beside `vllm-coder`, the erp/meromero seats and scriberr. It **replaced `semif`** on Prime's ruling of 2026-09-30, ~0510 PT: "replace semif with intern-decision now". The pick came from the previous night's bench, `docs/pfi/jev-candidates-bench-2026-09-30.md`: the only candidate at least as accurate as SemIf on our sets, faster at our shape, and inside the memory budget. [Intern-Decision-4B](https://huggingface.co/internlm/Intern-Decision-4B) (Shanghai AI Lab, Apache-2.0) is a Qwen3.5-4B finetune. It takes a state plus up to 16 named questions and scores every question's options in **one forward pass with no decoding**. It reads each answer from the logits just before a `` marker. `services/intern-decision-serve/` loads it once and scores through the checkpoint's **own** `inference.py` (`DecisionEngine.predict`). It keeps **semif-serve's HTTP surface**, so a semif caller changes only the URL and the token. The contract is `services/intern-decision-serve/intern-decision-serve.contract.md`. | | | |---|---| | **URL** | `http://10.251.50.54:8033` = `http://intern-decision.fv.internal:8033`. `/health` is open; POSTs need `Authorization: Bearer $(secret get intern-decision/api-token)` | | **Model** | `internlm/Intern-Decision-4B` @ `0e5e6aa7d6d750e2b1504ba11a8136cb58aeb3cd`, BF16, from `/tank/aimodels/huggingface`, mounted read-only and offline | | **Scoring code** | the snapshot's `inference.py`, sha256 `c904e2c6…43b29863`, checked before it is imported. Startup refuses any other file | | **Stack** | torch `2.10.0+cu128`, transformers `5.17.0`, fla `0.5.2`, causal-conv1d `1.7.0`: the bench's stack | | **Image** | `intern-decision-serve:`, built on fv-ml1 from `services/intern-decision-serve/` | | **State** | None. The service never downloads (`HF_HUB_OFFLINE=1`, read-only mount). If the weights are ever lost, re-pull the pinned revision by hand. | | **Rollback** | `stacks/semif` is stopped, not removed. See **Rollback** below. | ## API (semif-serve's) ```bash T=$(secret get intern-decision/api-token) curl -s -H "Authorization: Bearer $T" http://intern-decision.fv.internal:8033/decide -d '{ "id": "q1", "state": "Health checks passed in all three zones.", "question": "Did the deployment succeed?", "options": [{"id":"yes","description":"It succeeded."},{"id":"no","description":"It failed."}]}' ``` - `POST /decide` takes one decision, `{id, state, question, options[2..16]}`. - `POST /decide/shared` takes `{state, decisions: [{id, question, options}]}`. - `GET /health` reports the pins, the limits and the chunking rule. - **Order averaging:** `"orderings": "rotations"`, or `"all"` for ≤ 4 options, is still accepted, with semif's `combined` block. **It is barely needed:** the bench found Intern-Decision reorders 8 of 144 labels when the options are reversed, against SemIf's 30, and at one ordering it matched SemIf-with-rotations. - Each result carries semif's fields where they mean something: `option_ids`, `probabilities`, `input_tokens`, `prompt_sha256`, `prompt_version`, `model`, `readout`, `probability_status`, and on `/decide` also `total_seconds` and `forward_seconds`. New fields: - `top`; - `confidence`, the model's own; - `calibration`, the model's own temperature (T = 1.99241824); - `native`, the model's answer object, unchanged; - `call` (`index`, `field`, `questions`): which prompt the decision was asked in. - Past `MAX_QUEUE` (32) requests in progress, new POSTs get `429 busy` before their body is read. ### What a semif caller must know (every deliberate difference is in the contract, "Deltas") 1. **Option ids are shown to the model.** The prompt prints `A = : `, so an id is part of the question. Use meaningful or neutral ids, not misleading ones. 2. **Decisions in one `/decide/shared` call are asked together, in one prompt.** An answer can depend on the other questions in its call. - The bench measured this on Wyrd: 79/84 asked one decision at a time, 77/84 with a turn's 4 decisions in one prompt. - A call holds **at most 16 questions**. More are split greedily in request order: 1–16, 17–32, … . `/health` says so. 3. **`probabilities` are temperature-scaled** by the model's shipped calibration. SemIf's were raw. The argmax is the same either way. This calibration is the vendor's, fitted on its own data, not ours. 4. **Gone:** - `option_logits` (the model's runtime does not expose logits); - the prefix-cache timing fields; - `/health.semif_commit`; - per-workload calibration. Any `workload` is a 422, exactly as the deployed semif behaved with its empty table. 5. **`MAX_TOKENS` is 7,168 per call:** the state plus all its questions. It is lower than the model's own 8,192 so that every call the API accepts fits the VRAM cap. - The count is taken **before** the forward pass, so a longer call is a clear `422 invalid_request` ("Example has N tokens, above 7168; truncation is forbidden"). It is never truncated. - A request with more decisions is split into more calls, and each call must fit. ## VRAM: fits beside scriberr's peak, whatever the request **Budget (infra-ops, 2026-09-30):** GPU 1 needs nvidia-smi `Free` ≥ **15,400 MiB** before this service starts. That is our card peak of 9,876 MiB plus scriberr's peak of 5,496 MiB (with its 120 s slices), rounded up. So both can peak at the same moment. Read nvidia-smi's own `Free`, not total − used: the driver reserves 640 MiB on every card. The deploy at 0940 read 15,442. `torch.cuda.set_per_process_memory_fraction` caps only torch's allocator. The two knobs in `.env` are therefore chosen together and change together: ``` card footprint <= VRAM_CAP_GIB 9.0 (9,216 MiB) + 660 MiB outside the allocator (measured) = 9,876 MiB MAX_TOKENS 7168 = the largest call measured to fit under that cap, so no accepted call reaches the 503 ``` | on the live service, GPU 1 (per-process nvidia-smi, 0.1 s sampling, 2,939 samples over 316 s of acceptance load) | MiB | |---|---| | **at rest** after startup | **8,812** | | at rest after the acceptance (512 MiB release slack keeps a little cache) | 8,856 | | **peak**: largest requests, 64 decisions × 16 options, 4 calls × 7,168 tokens, N = 3 | **9,866**; torch `max_reserved` 8.996 of 9.0 GiB | | GPU 1 `Free` before the deploy / at rest after / lowest during the acceptance | 15,442 / 6,581 / 5,569, which stays above scriberr's 5,496 | - ⚠ **At 7,168 tokens the allocator reaches the cap with about 4 MiB to spare.** It fits because the allocator frees its cache and retries before it fails. - Measured: every call up to the limit answered 200. That is 145 shared calls on the live service, including the size search near the boundary, 3 × 4 calls at 7,168 tokens with 16 questions, and 3 at 7,168 tokens with one question. No 503 was seen. - The allocator's cache state depends on request history. If a max-size call ever does return 503, lower `MAX_TOKENS`. Do not raise the cap. - Raising either knob needs a fresh measurement (`acceptance/checks.py --checks fit,maxreq`) and the budget re-checked against scriberr. - **Over the cap the answer is 503 `out_of_memory`.** Memory returns to the resting baseline and the service keeps answering. This was proven on GPU 3 with a deliberately tight cap. Measured on GPU 3 before the deploy (alone on the card), which is where the knobs came from: | measured on fv-ml1 GPU 3 (alone on the card), image 0.1.0 | MiB | |---|---| | **at rest** (nvidia-smi, whole card less 2 MiB idle) | **8,820** (torch reserved 8,160 + outside 660) | | outside the allocator (CUDA context + loaded kernels), at rest / after 2×30 concurrent requests / after the latency shapes / after 4 × 8,191-token calls | 660 / 660 / 662 / 662 | | card peak at cap 9.0 GiB (fit search and a 64 × 8,191-token request, which got 503) | **9,876** | | peak with 4 × 8,191-token calls at cap 9.25 GiB (answered 200) / uncapped | 10,134 / 10,422 (allocator 9,760) | | largest 16-question call that answers 200: at cap 9.0 / at 9.25 GiB | 7,168 / 8,191 tokens | | startup peak (weights + vision tower until the swap + warm-up) | allocator 8.86 GiB. **A cap below ~8.9 GiB cannot start** (it fails closed; 8.5 was refused) | - **One inference thread is load-bearing** (contract INV-2). torch keeps CUDA state per host thread (cuBLAS handles and workspaces). The first build scored on anyio's 40-thread pool: a concurrent burst added 252 MiB outside the cap and 326 MiB inside it, and the peak reached 10,392 MiB. Load, warm-up and every call now run on one thread. - **Released after bursts:** when a call leaves reserved memory more than 512 MiB over the resting baseline, the engine calls `empty_cache()`. Measured: back to 7.97 GiB after every large request. - The rest footprint is 0.9 GB below the bench's 9,736 MiB. The service takes no images, so the vision tower (0.62 GiB) is swapped for a stub after the first warm-up, and the warm-up cache is released. Startup proves the swap changes nothing: the warm-up answer must be bit-identical (INV-7). ## Latency **Live, from nh3-dev** (0942 PT). The URL is `intern-decision.fv.internal:8033`, and the round trip is 22.9 ms on average (19.5–27.7 ms over 10 pings). GPU 1 is shared with the vLLM seats and scriberr. The harness is `bench_shape.py` from the Jev bench, 3 runs × 20 requests; each cell is the median with the run-median range. "server" is the service's own time. | shape | end to end from nh3-dev | server | GPU 3 loopback before deploy | SemIf 0.1.4 (README, from nh3-dev) | |---|---|---|---|---| | 1 decision, short (~180 tok) | 53.9 ms (53.3–54.2) | 35.2 | 36.4 | 69 | | **21 binary criteria**, one `/decide/shared` (2 calls: 16 + 5) | **114.3 ms** (114.1–114.3) | 80.3 | 82.3 | 159 | | 1 decision over the ~3,900-token state | 211.2 ms (210.8–211.6) | 186.8 | 191.4 | | | **16 criteria over the ~3,900-token state** (1 call, 4,579 tok) | **238.3 ms** (237.8–238.7) | 205.3 | 216.8 | (bench loopback: 505) | | largest request (4 calls × 7,168 tok), N = 3 | 1,519–1,543 ms | | | | Server-side time on GPU 1 matches GPU 3 within a few ms, so the vLLM neighbours were not measurably slowing it at 0942. The rest of the end-to-end time is the network (~23 ms round trip plus the HTTP exchange). ## Acceptance (2026-09-30) Raw results: `services/intern-decision-serve/acceptance/gpu3-2026-09-30/` (`compare.json`, one directory per process lifetime). **The harness is the bench's own.** `bench_sets.py --backend semif` speaks semif-serve's API, so it drove this service unchanged. `compare.py` scores it as the bench's `analyze.py` did. ### On GPU 3 before the deploy (the harness and the floor) | check (GPU 3, cap 9.25 GiB, N = 3 fresh processes unless noted) | result | bench / floor | |---|---|---| | **positive control**, pooled 259, single ordering | **240, 240, 240** | bench 240 (4 repeats); the pooled set resolves ±4 pts | | positive control, Wyrd /84, single | **79, 79, 79** | bench 79 | | row by row against the bench's native rows (560 rows, single / rotations / negative) | **0 top changes, max Δp 0.000** in every repeat: bit-identical | bench floor: 0 labels moved across 4 restarts | | rotations, pooled / Wyrd | 236 / 77 (×3) | bench 236 / 77 | | **negative control** (descriptions rotated, 144): same top / follows the description / right vs the original gold | **10 / 122 / 14** (×3) | bench 10 / 122 / 14 | | A-vs-A in the process (authored144 twice) | 0/144 flips, Δp 0.0 (×3) | | | A-vs-A across restarts (r1~r2, r1~r3, r2~r3; 560 rows, single and rotations) | 0 flips, Δp 0.0 | | | Wyrd, one `/decide/shared` per turn (4 decisions in one prompt) | 78/84 (×3); 1/84 row differs from the bench's 77 | expected: field names are positional here, decision names in the bench (contract) | | more than 16 questions: 20 decisions in one request, against the same decisions asked as 16 + 4 | 2 calls `[16, 4]`, all 20 tops equal, Δp 0, same prompt hashes | | | 401: no token / wrong token, both POSTs; `/health` open | 401 / 401; 200 | | | 429: 48 concurrent requests, `MAX_QUEUE` 32 | 32 answered, 16 × 429 `busy` | | | largest request at `MAX_TOKENS` 8,192 (64 × 16 options, 4 × 8,191 tok), cap 9.25 GiB | **200** ×3; card peak 10,134 MiB | superseded by the 9.0 / 7,168 knobs (VRAM) | | **over the cap** (the same request at a deliberately tight cap of 8.9 GiB) | **503 `out_of_memory`**; reserved 7.969 GiB before, 7.969 after, next request 200; then all 560 single rows answered, bit-identical to the bench | | | fit under the cap: the largest 16-question call that answers 200 | 8,191 tok (the whole API) at 9.25 GiB; 7,168 tok at 9.0 GiB | | | cap below startup | 8.5 GiB: startup refused (OOM loading), fails closed | | ### On the live service (GPU 1, 0942–0946 PT, through `http://intern-decision.fv.internal:8033` from nh3-dev) Raw results: `services/intern-decision-serve/acceptance/gpu1-live/`. | check | result | reference | |---|---|---| | **positive control**, pooled 259, single ordering | **240** | bench 240; GPU 3 240 ×3 | | positive control, Wyrd /84 | **79** | bench 79 | | row by row against the bench's native rows (560 single + 144 negative) | **0 top changes, max Δp 0.000** | | | negative control: same top / follows the description / right vs gold | **10 / 122 / 14** | bench 10 / 122 / 14 | | 401 without or with a wrong token (both POSTs); `/health` open | 401 / 401; 200 | | | largest accepted request: 64 decisions × 16 options, 4 calls × 7,168 tok | **200** ×3 | | | one question at 7,168 tok | 200 ×3; one token more (7,171) → **422** "above 7168" | | | any 503 across the whole live acceptance | **none** (server log: 128 + 145 × 200, 18 × 422, 4 × 401) | | The live pass is N = 1 for the positive control. It is read against the GPU 3 floor: the same image moved 0 of 560 rows across 3 restarts, and 0 against the bench's 4 restarts. ## Building and deploying ```bash # from nh3-dev. /opt/docker/src is root-owned, so create the version dir with sudo first. ssh infra-ops@10.251.50.54 'sudo -n install -d -o infra-ops -g infra-ops /opt/docker/src/intern-decision-serve-X.Y.Z | cat' tar -C services/intern-decision-serve -cf - --exclude=.venv --exclude=.pytest_cache --exclude=__pycache__ \ --exclude=acceptance --exclude='*.egg-info' . \ | ssh infra-ops@10.251.50.54 'tar -x -C /opt/docker/src/intern-decision-serve-X.Y.Z' # on fv-ml1 cd /opt/docker/src/intern-decision-serve-X.Y.Z && docker build -t intern-decision-serve:X.Y.Z . # from nh3-dev: compose.yaml, .env.example and this README (never .env) scripts/deploy-stack.sh fv-ml1 intern-decision # on fv-ml1: .env is infra-ops's, mode 0600. Rewrite it IN PLACE to keep the owner and mode. cd /opt/docker/compose/intern-decision && new=$(sed 's/^IMAGE=.*/IMAGE=intern-decision-serve:X.Y.Z/' .env) \ && printf '%s\n' "$new" > .env && docker compose config -q && docker compose up -d ``` **Before any deploy onto GPU 1**, check two things. If either fails, stop; do not squeeze scriberr. - nvidia-smi's own `Free` on GPU 1 must be at least **15,400 MiB**: `nvidia-smi -i 1 --query-gpu=memory.free --format=csv`. That is our card peak of 9,876 MiB plus scriberr's 5,496, rounded up. Do not use total − used, which misses the driver's 640 MiB reserve. - Scriberr must not be running a job. This command must print 0: `docker logs --since 2m scriberr | grep -c "Processing single-track job"`. Startup fails closed. A container that never reaches healthy did not pass its own checks: the `inference.py` hash, the pinned snapshot, the warm-up, the text-only swap and the prompt hash. Read `docker logs intern-decision`. Record the deploy with `scripts/ops-log`. **To move the model forward**, change `REVISION` and `INFERENCE_PY_SHA256` in `config.py`. Re-run the bench's sets through the service (`services/intern-decision-serve/acceptance/`), and re-measure the VRAM before changing the cap. ## Rollback `semif` is **stopped, not removed**, and is kept as the rollback. Only on Prime's word: ```bash cd /opt/docker/compose/intern-decision && docker compose stop # first: the two do not both fit GPU 1 cd /opt/docker/compose/semif && docker compose start ``` Then check GPU 1's free memory against scriberr's peak. Callers go back to `:8032` and `secret get semif/api-token`.