stacks/intern-decision: compose (GPU 1, :8033, hard VRAM cap as the single .env knob, healthcheck, Homepage group 'AI - Eval & Retrieval'), .env.example and README. dns: intern-decision.fv.internal -> fv-ml1 (synced to ana/esh/nh3). acceptance on fv-ml1 GPU 3, 3 fresh processes: bit-identical to the Jev bench's native rows (pooled 240/259, Wyrd 79/84, 0/560 flips, Δp 0), negative control 10/122/14, 0 flips across restarts; largest accepted request 200 at a 10,134 MiB card peak under a 9.25 GiB cap; 503 and recovery proven at a tight cap. GPU 1 deploy held: nvidia-smi Free on GPU 1 is 15,442 MiB.
14 KiB
intern-decision
⚠ STATUS 2026-09-30 0937 PT: built and accepted on GPU 3; the GPU 1 deploy is HELD. GPU 1's usable free memory is 15,442 MiB (nvidia-smi
Free: 97,887 total − 640 driver-reserved − 81,806 used), below the 15,800 MiB pre-deploy floor. The 10,300 MiB budget was derived from total − used (16,081). The decision is with infra-ops/Prime; see "VRAM". Nothing is running on port 8033 yet.
Intern-Decision-4B typed decisions on fv-ml1 GPU 1, the utility card beside vllm-coder,
the erp/meromero seats and scriberr. It replaced semif on Prime's ruling of 2026-09-30,
~0510 PT: "replace semif with intern-decision now". The pick came from the previous night's bench,
docs/pfi/jev-candidates-bench-2026-09-30.md: the only candidate at least as accurate as SemIf on
our sets, faster at our shape, and inside the memory budget.
Intern-Decision-4B (Shanghai AI Lab,
Apache-2.0) is a Qwen3.5-4B finetune. It takes a state plus up to 16 named questions and scores
every question's options in one forward pass with no decoding. It reads each answer from the
logits just before a <decision> marker. services/intern-decision-serve/ loads it once and
scores through the checkpoint's own inference.py (DecisionEngine.predict). It keeps
semif-serve's HTTP surface, so a semif caller changes only the URL and the token. The contract
is services/intern-decision-serve/intern-decision-serve.contract.md.
| URL | http://10.251.50.54:8033 = http://intern-decision.fv.internal:8033. /health is open; POSTs need Authorization: Bearer $(secret get intern-decision/api-token) |
| Model | internlm/Intern-Decision-4B @ 0e5e6aa7d6d750e2b1504ba11a8136cb58aeb3cd, BF16, from /tank/aimodels/huggingface, mounted read-only and offline |
| Scoring code | the snapshot's inference.py, sha256 c904e2c6…43b29863, checked before it is imported. Startup refuses any other file |
| Stack | torch 2.10.0+cu128, transformers 5.17.0, fla 0.5.2, causal-conv1d 1.7.0: the bench's stack |
| Image | intern-decision-serve:<version>, built on fv-ml1 from services/intern-decision-serve/ |
| State | None. The service never downloads (HF_HUB_OFFLINE=1, read-only mount). If the weights are ever lost, re-pull the pinned revision by hand. |
| Rollback | stacks/semif is stopped, not removed. See Rollback below. |
API (semif-serve's)
T=$(secret get intern-decision/api-token)
curl -s -H "Authorization: Bearer $T" http://intern-decision.fv.internal:8033/decide -d '{
"id": "q1", "state": "Health checks passed in all three zones.",
"question": "Did the deployment succeed?",
"options": [{"id":"yes","description":"It succeeded."},{"id":"no","description":"It failed."}]}'
POST /decidetakes one decision,{id, state, question, options[2..16]}.POST /decide/sharedtakes{state, decisions: [{id, question, options}]}.GET /healthreports the pins, the limits and the chunking rule.- Order averaging:
"orderings": "rotations", or"all"for ≤ 4 options, is still accepted, with semif'scombinedblock. It is barely needed: the bench found Intern-Decision reorders 8 of 144 labels when the options are reversed, against SemIf's 30, and at one ordering it matched SemIf-with-rotations. - Each result carries semif's fields where they mean something:
option_ids,probabilities,input_tokens,prompt_sha256,prompt_version,model,readout,probability_status, and on/decidealsototal_secondsandforward_seconds. New fields:top;confidence, the model's own;calibration, the model's own temperature (T = 1.99241824);native, the model's answer object, unchanged;call(index,field,questions): which prompt the decision was asked in.
- Past
MAX_QUEUE(32) requests in progress, new POSTs get429 busybefore their body is read.
What a semif caller must know (every deliberate difference is in the contract, "Deltas")
- Option ids are shown to the model. The prompt prints
A = <id>: <description>, so an id is part of the question. Use meaningful or neutral ids, not misleading ones. - Decisions in one
/decide/sharedcall are asked together, in one prompt. An answer can depend on the other questions in its call.- The bench measured this on Wyrd: 79/84 asked one decision at a time, 77/84 with a turn's 4 decisions in one prompt.
- A call holds at most 16 questions. More are split greedily in request order: 1–16,
17–32, … .
/healthsays so.
probabilitiesare temperature-scaled by the model's shipped calibration. SemIf's were raw. The argmax is the same either way. This calibration is the vendor's, fitted on its own data, not ours.- Gone:
option_logits(the model's runtime does not expose logits);- the prefix-cache timing fields;
/health.semif_commit;- per-workload calibration. Any
workloadis a 422, exactly as the deployed semif behaved with its empty table.
MAX_TOKENS(8192) is per call: the state plus all its questions. A longer call is a 422 and is never truncated. Every call up to that limit fits the VRAM cap (measured). See VRAM below.
VRAM: the whole container ≤ 10,300 MiB
Budget (infra-ops, 2026-09-30): the container's whole nvidia-smi footprint, CUDA context
included, must never exceed 10,300 MiB, whatever the request. That leaves scriberr (peak
5,496 MiB with its 120 s slices) 295 MiB of spare even when both hit their peaks at once.
torch.cuda.set_per_process_memory_fraction caps only torch's allocator, so the cap is the budget
minus the measured non-allocator overhead:
footprint <= VRAM_CAP_GIB + overhead = 9,472 MiB (9.25 GiB) + 662 MiB = 10,134 MiB (166 MiB under budget)
VRAM_CAP_GIB in .env is the single knob. It is applied before the weights load.
⚠ The budget assumed 16,081 MiB free on GPU 1. That is total − used. nvidia-smi's own Free is
15,442 MiB, because the driver reserves 640 MiB on every card (GPU 3 shows the same 640).
Against the true free memory, a footprint that never collides with scriberr's 5,496 MiB peak must
stay at or below 9,946 MiB:
| cap | footprint peak (measured, GPU 3) | with scriberr at its peak, against 15,442 free | calls that fit (16 questions × 16 options) |
|---|---|---|---|
9.25 GiB (.env.example now) |
10,134 MiB | 188 MiB over | every call the API accepts (8,191 tok) |
| 9.0 GiB | 9,876 MiB | 70 MiB spare | up to 7,168 tok; longer calls get 503. Every bench shape fits, at the same latency |
| measured on fv-ml1 GPU 3 (alone on the card), image 0.1.0 | MiB |
|---|---|
| at rest (nvidia-smi, whole card less 2 MiB idle) | 8,820 (torch reserved 8,160 + outside 660) |
| outside the allocator (CUDA context + loaded kernels), at rest / after 2×30 concurrent requests / after the latency shapes / after 4 × 8,191-token calls | 660 / 660 / 662 / 662 |
| peak, largest request the API accepts, capped at 9.25 GiB (card, 0.1 s sampling) | 10,134; torch max_reserved = 9.25 GiB exactly, so the cap is what held it |
| the same request uncapped, for comparison | 10,422 (allocator peak 9,760): over budget, so the cap is required |
| startup peak (weights + vision tower until the swap + warm-up) | allocator 8.86 GiB. A cap below ~8.9 GiB cannot start (it fails closed; 8.5 was refused) |
- The largest request the API accepts is 64 decisions × 16 options, with each of its 4 calls
at 8,191 tokens. Under the cap it answers 200 (1.67 s), not 503. Near the limit the
allocator frees cached blocks before it fails, so the uncapped 9,760 MiB peak was partly cache.
Every call up to
MAX_TOKENSfits the cap. - Over the cap the answer is 503
out_of_memory. Memory goes back to the resting baseline and the service keeps answering. This was proven with a deliberately tight cap (acceptance). - One inference thread is load-bearing (contract INV-2). torch keeps CUDA state per host thread (cuBLAS handles and workspaces). The first build scored on anyio's 40-thread pool: a concurrent burst added 252 MiB outside the cap and 326 MiB inside it, and the peak reached 10,392 MiB. Load, warm-up and every call now run on one thread.
- Released after bursts: when a call leaves reserved memory more than 512 MiB over the resting
baseline, the engine calls
empty_cache(). Measured: back to 7.97 GiB after every large request. - The rest footprint is 0.9 GB below the bench's 9,736 MiB. The service takes no images, so the vision tower (0.62 GiB) is swapped for a stub after the first warm-up, and the warm-up cache is released. Startup proves the swap changes nothing: the warm-up answer must be bit-identical (INV-7).
Latency
Loopback on fv-ml1 (GPU 3 alone on the card, image 0.1.0, cap 9.25 GiB). bench_shape.py
from the Jev bench ran 3 runs × 20 requests, and the medians are shown with the run-median
range. The bench's native Intern-Decision numbers are the reference.
| shape | end to end | server | bench (native, loopback) | SemIf 0.1.4 (bench) |
|---|---|---|---|---|
| 1 decision, short (~180 tok) | 36.4 ms (36.4–36.5) | 34.7 | 39 | 37 |
21 binary criteria, one /decide/shared (2 calls: 16 + 5) |
82.3 ms (82.3–82.4) | 78.9 | 88 | 131 |
| 1 decision over the ~3,900-token state | 191.4 ms (191.4–191.5) | 187.9 | 190 | 186 |
| 16 criteria over the ~3,900-token state (1 call, 4,579 tok) | 216.8 ms (216.5–217.0) | 211.1 | 215 | 505 |
| largest accepted request (4 calls × 8,191 tok) | 1,672 ms (1,672–1,675, N=3) |
From nh3-dev and on GPU 1, next to the vLLM seats: pending the GPU 1 deploy (held, see the status banner).
Acceptance (2026-09-30)
Raw results: services/intern-decision-serve/acceptance/gpu3-2026-09-30/ (compare.json, one
directory per process lifetime). The harness is the bench's own. bench_sets.py --backend semif speaks semif-serve's API, so it drove this service unchanged. compare.py
scores it as the bench's analyze.py did.
| check (GPU 3, cap 9.25 GiB, N = 3 fresh processes unless noted) | result | bench / floor |
|---|---|---|
| positive control, pooled 259, single ordering | 240, 240, 240 | bench 240 (4 repeats); the pooled set resolves ±4 pts |
| positive control, Wyrd /84, single | 79, 79, 79 | bench 79 |
| row by row against the bench's native rows (560 rows, single / rotations / negative) | 0 top changes, max Δp 0.000 in every repeat: bit-identical | bench floor: 0 labels moved across 4 restarts |
| rotations, pooled / Wyrd | 236 / 77 (×3) | bench 236 / 77 |
| negative control (descriptions rotated, 144): same top / follows the description / right vs the original gold | 10 / 122 / 14 (×3) | bench 10 / 122 / 14 |
| A-vs-A in the process (authored144 twice) | 0/144 flips, Δp 0.0 (×3) | |
| A-vs-A across restarts (r1 |
0 flips, Δp 0.0 | |
Wyrd, one /decide/shared per turn (4 decisions in one prompt) |
78/84 (×3); 1/84 row differs from the bench's 77 | expected: field names are positional here, decision names in the bench (contract) |
| more than 16 questions: 20 decisions in one request, against the same decisions asked as 16 + 4 | 2 calls [16, 4], all 20 tops equal, Δp 0, same prompt hashes |
|
401: no token / wrong token, both POSTs; /health open |
401 / 401; 200 | |
429: 48 concurrent requests, MAX_QUEUE 32 |
32 answered, 16 × 429 busy |
|
| largest request the API accepts (64 × 16 options, 4 × 8,191 tok) at cap 9.25 GiB | 200 ×3; card peak 10,134 MiB | budget 10,300 |
| over the cap (the same request at a deliberately tight cap of 8.9 GiB) | 503 out_of_memory; reserved 7.969 GiB before, 7.969 after, next request 200; then all 560 single rows answered, bit-identical to the bench |
|
| fit under the cap: the largest 16-question call that answers 200 | 8,191 tok (the whole API) at 9.25 GiB; 7,168 tok at 9.0 GiB | |
| cap below startup | 8.5 GiB: startup refused (OOM loading), fails closed |
On the deployed service (GPU 1): pending the deploy (held, see the status banner).
Building and deploying
# from nh3-dev. /opt/docker/src is root-owned, so create the version dir with sudo first.
ssh infra-ops@10.251.50.54 'sudo -n install -d -o infra-ops -g infra-ops /opt/docker/src/intern-decision-serve-X.Y.Z | cat'
tar -C services/intern-decision-serve -cf - --exclude=.venv --exclude=.pytest_cache --exclude=__pycache__ \
--exclude=acceptance --exclude='*.egg-info' . \
| ssh infra-ops@10.251.50.54 'tar -x -C /opt/docker/src/intern-decision-serve-X.Y.Z'
# on fv-ml1
cd /opt/docker/src/intern-decision-serve-X.Y.Z && docker build -t intern-decision-serve:X.Y.Z .
# from nh3-dev: compose.yaml, .env.example and this README (never .env)
scripts/deploy-stack.sh fv-ml1 intern-decision
# on fv-ml1: .env is infra-ops's, mode 0600. Rewrite it IN PLACE to keep the owner and mode.
cd /opt/docker/compose/intern-decision && new=$(sed 's/^IMAGE=.*/IMAGE=intern-decision-serve:X.Y.Z/' .env) \
&& printf '%s\n' "$new" > .env && docker compose config -q && docker compose up -d
Before any deploy onto GPU 1: GPU 1 must have at least 15,800 MiB free
(nvidia-smi -i 1 --query-gpu=memory.free --format=csv). If it has less, stop. Do not squeeze
scriberr.
Startup fails closed. A container that never reaches healthy did not pass its own checks: the
inference.py hash, the pinned snapshot, the warm-up, the text-only swap and the prompt hash.
Read docker logs intern-decision. Record the deploy with scripts/ops-log.
To move the model forward, change REVISION and INFERENCE_PY_SHA256 in config.py.
Re-run the bench's sets through the service (services/intern-decision-serve/acceptance/), and
re-measure the VRAM before changing the cap.
Rollback
semif is stopped, not removed, and is kept as the rollback. Only on Prime's word:
cd /opt/docker/compose/intern-decision && docker compose stop # first: the two do not both fit GPU 1
cd /opt/docker/compose/semif && docker compose start
Then check GPU 1's free memory against scriberr's peak. Callers go back to :8032 and
secret get semif/api-token.