# intern-decision > ⚠ **STATUS 2026-09-30 0937 PT: built and accepted on GPU 3; the GPU 1 deploy is HELD.** GPU 1's > usable free memory is 15,442 MiB (nvidia-smi `Free`: 97,887 total − 640 driver-reserved − > 81,806 used), below the 15,800 MiB pre-deploy floor. The 10,300 MiB budget was derived from > total − used (16,081). The decision is with infra-ops/Prime; see "VRAM". Nothing is running > on port 8033 yet. **Intern-Decision-4B typed decisions** on **fv-ml1 GPU 1**, the utility card beside `vllm-coder`, the erp/meromero seats and scriberr. It **replaced `semif`** on Prime's ruling of 2026-09-30, ~0510 PT: "replace semif with intern-decision now". The pick came from the previous night's bench, `docs/pfi/jev-candidates-bench-2026-09-30.md`: the only candidate at least as accurate as SemIf on our sets, faster at our shape, and inside the memory budget. [Intern-Decision-4B](https://huggingface.co/internlm/Intern-Decision-4B) (Shanghai AI Lab, Apache-2.0) is a Qwen3.5-4B finetune. It takes a state plus up to 16 named questions and scores every question's options in **one forward pass with no decoding**. It reads each answer from the logits just before a `` marker. `services/intern-decision-serve/` loads it once and scores through the checkpoint's **own** `inference.py` (`DecisionEngine.predict`). It keeps **semif-serve's HTTP surface**, so a semif caller changes only the URL and the token. The contract is `services/intern-decision-serve/intern-decision-serve.contract.md`. | | | |---|---| | **URL** | `http://10.251.50.54:8033` = `http://intern-decision.fv.internal:8033`. `/health` is open; POSTs need `Authorization: Bearer $(secret get intern-decision/api-token)` | | **Model** | `internlm/Intern-Decision-4B` @ `0e5e6aa7d6d750e2b1504ba11a8136cb58aeb3cd`, BF16, from `/tank/aimodels/huggingface`, mounted read-only and offline | | **Scoring code** | the snapshot's `inference.py`, sha256 `c904e2c6…43b29863`, checked before it is imported. Startup refuses any other file | | **Stack** | torch `2.10.0+cu128`, transformers `5.17.0`, fla `0.5.2`, causal-conv1d `1.7.0`: the bench's stack | | **Image** | `intern-decision-serve:`, built on fv-ml1 from `services/intern-decision-serve/` | | **State** | None. The service never downloads (`HF_HUB_OFFLINE=1`, read-only mount). If the weights are ever lost, re-pull the pinned revision by hand. | | **Rollback** | `stacks/semif` is stopped, not removed. See **Rollback** below. | ## API (semif-serve's) ```bash T=$(secret get intern-decision/api-token) curl -s -H "Authorization: Bearer $T" http://intern-decision.fv.internal:8033/decide -d '{ "id": "q1", "state": "Health checks passed in all three zones.", "question": "Did the deployment succeed?", "options": [{"id":"yes","description":"It succeeded."},{"id":"no","description":"It failed."}]}' ``` - `POST /decide` takes one decision, `{id, state, question, options[2..16]}`. - `POST /decide/shared` takes `{state, decisions: [{id, question, options}]}`. - `GET /health` reports the pins, the limits and the chunking rule. - **Order averaging:** `"orderings": "rotations"`, or `"all"` for ≤ 4 options, is still accepted, with semif's `combined` block. **It is barely needed:** the bench found Intern-Decision reorders 8 of 144 labels when the options are reversed, against SemIf's 30, and at one ordering it matched SemIf-with-rotations. - Each result carries semif's fields where they mean something: `option_ids`, `probabilities`, `input_tokens`, `prompt_sha256`, `prompt_version`, `model`, `readout`, `probability_status`, and on `/decide` also `total_seconds` and `forward_seconds`. New fields: - `top`; - `confidence`, the model's own; - `calibration`, the model's own temperature (T = 1.99241824); - `native`, the model's answer object, unchanged; - `call` (`index`, `field`, `questions`): which prompt the decision was asked in. - Past `MAX_QUEUE` (32) requests in progress, new POSTs get `429 busy` before their body is read. ### What a semif caller must know (every deliberate difference is in the contract, "Deltas") 1. **Option ids are shown to the model.** The prompt prints `A = : `, so an id is part of the question. Use meaningful or neutral ids, not misleading ones. 2. **Decisions in one `/decide/shared` call are asked together, in one prompt.** An answer can depend on the other questions in its call. - The bench measured this on Wyrd: 79/84 asked one decision at a time, 77/84 with a turn's 4 decisions in one prompt. - A call holds **at most 16 questions**. More are split greedily in request order: 1–16, 17–32, … . `/health` says so. 3. **`probabilities` are temperature-scaled** by the model's shipped calibration. SemIf's were raw. The argmax is the same either way. This calibration is the vendor's, fitted on its own data, not ours. 4. **Gone:** - `option_logits` (the model's runtime does not expose logits); - the prefix-cache timing fields; - `/health.semif_commit`; - per-workload calibration. Any `workload` is a 422, exactly as the deployed semif behaved with its empty table. 5. **`MAX_TOKENS` (8192) is per call:** the state plus all its questions. A longer call is a 422 and is never truncated. Every call up to that limit fits the VRAM cap (measured). See VRAM below. ## VRAM: the whole container ≤ 10,300 MiB **Budget (infra-ops, 2026-09-30):** the container's **whole** nvidia-smi footprint, CUDA context included, must never exceed **10,300 MiB**, whatever the request. That leaves scriberr (peak 5,496 MiB with its 120 s slices) 295 MiB of spare even when both hit their peaks at once. `torch.cuda.set_per_process_memory_fraction` caps only torch's allocator, so the cap is the budget minus the measured non-allocator overhead: ``` footprint <= VRAM_CAP_GIB + overhead = 9,472 MiB (9.25 GiB) + 662 MiB = 10,134 MiB (166 MiB under budget) ``` `VRAM_CAP_GIB` in `.env` is **the single knob**. It is applied before the weights load. ⚠ **The budget assumed 16,081 MiB free on GPU 1. That is total − used.** nvidia-smi's own `Free` is **15,442 MiB**, because the driver reserves 640 MiB on every card (GPU 3 shows the same 640). Against the true free memory, a footprint that never collides with scriberr's 5,496 MiB peak must stay at or below **9,946 MiB**: | cap | footprint peak (measured, GPU 3) | with scriberr at its peak, against 15,442 free | calls that fit (16 questions × 16 options) | |---|---|---|---| | **9.25 GiB** (`.env.example` now) | 10,134 MiB | **188 MiB over** | every call the API accepts (8,191 tok) | | **9.0 GiB** | 9,876 MiB | 70 MiB spare | up to 7,168 tok; longer calls get 503. Every bench shape fits, at the same latency | | measured on fv-ml1 GPU 3 (alone on the card), image 0.1.0 | MiB | |---|---| | **at rest** (nvidia-smi, whole card less 2 MiB idle) | **8,820** (torch reserved 8,160 + outside 660) | | outside the allocator (CUDA context + loaded kernels), at rest / after 2×30 concurrent requests / after the latency shapes / after 4 × 8,191-token calls | 660 / 660 / 662 / 662 | | **peak, largest request the API accepts, capped at 9.25 GiB** (card, 0.1 s sampling) | **10,134**; torch `max_reserved` = 9.25 GiB exactly, so the cap is what held it | | the same request **uncapped**, for comparison | 10,422 (allocator peak 9,760): **over budget, so the cap is required** | | startup peak (weights + vision tower until the swap + warm-up) | allocator 8.86 GiB. **A cap below ~8.9 GiB cannot start** (it fails closed; 8.5 was refused) | - **The largest request the API accepts** is 64 decisions × 16 options, with each of its 4 calls at 8,191 tokens. Under the cap it answers **200** (1.67 s), not 503. Near the limit the allocator frees cached blocks before it fails, so the uncapped 9,760 MiB peak was partly cache. Every call up to `MAX_TOKENS` fits the cap. - **Over the cap the answer is 503 `out_of_memory`.** Memory goes back to the resting baseline and the service keeps answering. This was proven with a deliberately tight cap (acceptance). - **One inference thread is load-bearing** (contract INV-2). torch keeps CUDA state per host thread (cuBLAS handles and workspaces). The first build scored on anyio's 40-thread pool: a concurrent burst added 252 MiB outside the cap and 326 MiB inside it, and the peak reached 10,392 MiB. Load, warm-up and every call now run on one thread. - **Released after bursts:** when a call leaves reserved memory more than 512 MiB over the resting baseline, the engine calls `empty_cache()`. Measured: back to 7.97 GiB after every large request. - The rest footprint is 0.9 GB below the bench's 9,736 MiB. The service takes no images, so the vision tower (0.62 GiB) is swapped for a stub after the first warm-up, and the warm-up cache is released. Startup proves the swap changes nothing: the warm-up answer must be bit-identical (INV-7). ## Latency Loopback on fv-ml1 (GPU 3 alone on the card, image 0.1.0, cap 9.25 GiB). `bench_shape.py` from the Jev bench ran 3 runs × 20 requests, and the medians are shown with the run-median range. The bench's native Intern-Decision numbers are the reference. | shape | end to end | server | bench (native, loopback) | SemIf 0.1.4 (bench) | |---|---|---|---|---| | 1 decision, short (~180 tok) | 36.4 ms (36.4–36.5) | 34.7 | 39 | 37 | | **21 binary criteria**, one `/decide/shared` (2 calls: 16 + 5) | **82.3 ms** (82.3–82.4) | 78.9 | 88 | 131 | | 1 decision over the ~3,900-token state | 191.4 ms (191.4–191.5) | 187.9 | 190 | 186 | | **16 criteria over the ~3,900-token state** (1 call, 4,579 tok) | **216.8 ms** (216.5–217.0) | 211.1 | 215 | 505 | | largest accepted request (4 calls × 8,191 tok) | 1,672 ms (1,672–1,675, N=3) | | | | *From nh3-dev and on GPU 1, next to the vLLM seats: pending the GPU 1 deploy (held, see the status banner).* ## Acceptance (2026-09-30) Raw results: `services/intern-decision-serve/acceptance/gpu3-2026-09-30/` (`compare.json`, one directory per process lifetime). **The harness is the bench's own.** `bench_sets.py --backend semif` speaks semif-serve's API, so it drove this service unchanged. `compare.py` scores it as the bench's `analyze.py` did. | check (GPU 3, cap 9.25 GiB, N = 3 fresh processes unless noted) | result | bench / floor | |---|---|---| | **positive control**, pooled 259, single ordering | **240, 240, 240** | bench 240 (4 repeats); the pooled set resolves ±4 pts | | positive control, Wyrd /84, single | **79, 79, 79** | bench 79 | | row by row against the bench's native rows (560 rows, single / rotations / negative) | **0 top changes, max Δp 0.000** in every repeat: bit-identical | bench floor: 0 labels moved across 4 restarts | | rotations, pooled / Wyrd | 236 / 77 (×3) | bench 236 / 77 | | **negative control** (descriptions rotated, 144): same top / follows the description / right vs the original gold | **10 / 122 / 14** (×3) | bench 10 / 122 / 14 | | A-vs-A in the process (authored144 twice) | 0/144 flips, Δp 0.0 (×3) | | | A-vs-A across restarts (r1~r2, r1~r3, r2~r3; 560 rows, single and rotations) | 0 flips, Δp 0.0 | | | Wyrd, one `/decide/shared` per turn (4 decisions in one prompt) | 78/84 (×3); 1/84 row differs from the bench's 77 | expected: field names are positional here, decision names in the bench (contract) | | more than 16 questions: 20 decisions in one request, against the same decisions asked as 16 + 4 | 2 calls `[16, 4]`, all 20 tops equal, Δp 0, same prompt hashes | | | 401: no token / wrong token, both POSTs; `/health` open | 401 / 401; 200 | | | 429: 48 concurrent requests, `MAX_QUEUE` 32 | 32 answered, 16 × 429 `busy` | | | largest request the API accepts (64 × 16 options, 4 × 8,191 tok) at cap 9.25 GiB | **200** ×3; card peak 10,134 MiB | budget 10,300 | | **over the cap** (the same request at a deliberately tight cap of 8.9 GiB) | **503 `out_of_memory`**; reserved 7.969 GiB before, 7.969 after, next request 200; then all 560 single rows answered, bit-identical to the bench | | | fit under the cap: the largest 16-question call that answers 200 | 8,191 tok (the whole API) at 9.25 GiB; 7,168 tok at 9.0 GiB | | | cap below startup | 8.5 GiB: startup refused (OOM loading), fails closed | | *On the deployed service (GPU 1): pending the deploy (held, see the status banner).* ## Building and deploying ```bash # from nh3-dev. /opt/docker/src is root-owned, so create the version dir with sudo first. ssh infra-ops@10.251.50.54 'sudo -n install -d -o infra-ops -g infra-ops /opt/docker/src/intern-decision-serve-X.Y.Z | cat' tar -C services/intern-decision-serve -cf - --exclude=.venv --exclude=.pytest_cache --exclude=__pycache__ \ --exclude=acceptance --exclude='*.egg-info' . \ | ssh infra-ops@10.251.50.54 'tar -x -C /opt/docker/src/intern-decision-serve-X.Y.Z' # on fv-ml1 cd /opt/docker/src/intern-decision-serve-X.Y.Z && docker build -t intern-decision-serve:X.Y.Z . # from nh3-dev: compose.yaml, .env.example and this README (never .env) scripts/deploy-stack.sh fv-ml1 intern-decision # on fv-ml1: .env is infra-ops's, mode 0600. Rewrite it IN PLACE to keep the owner and mode. cd /opt/docker/compose/intern-decision && new=$(sed 's/^IMAGE=.*/IMAGE=intern-decision-serve:X.Y.Z/' .env) \ && printf '%s\n' "$new" > .env && docker compose config -q && docker compose up -d ``` **Before any deploy onto GPU 1**, check two things. If either fails, stop; do not squeeze scriberr. - nvidia-smi's own `Free` on GPU 1 must be at least **15,400 MiB**: `nvidia-smi -i 1 --query-gpu=memory.free --format=csv`. That is our card peak of 9,876 MiB plus scriberr's 5,496, rounded up. Do not use total − used, which misses the driver's 640 MiB reserve. - Scriberr must not be running a job. This command must print 0: `docker logs --since 2m scriberr | grep -c "Processing single-track job"`. Startup fails closed. A container that never reaches healthy did not pass its own checks: the `inference.py` hash, the pinned snapshot, the warm-up, the text-only swap and the prompt hash. Read `docker logs intern-decision`. Record the deploy with `scripts/ops-log`. **To move the model forward**, change `REVISION` and `INFERENCE_PY_SHA256` in `config.py`. Re-run the bench's sets through the service (`services/intern-decision-serve/acceptance/`), and re-measure the VRAM before changing the cap. ## Rollback `semif` is **stopped, not removed**, and is kept as the rollback. Only on Prime's word: ```bash cd /opt/docker/compose/intern-decision && docker compose stop # first: the two do not both fit GPU 1 cd /opt/docker/compose/semif && docker compose start ``` Then check GPU 1's free memory against scriberr's peak. Callers go back to `:8032` and `secret get semif/api-token`.