Files
esh-pfi-infrastructure/stacks/intern-decision
vh 750675e391 feat(intern-decision): cap 9.0 GiB with MAX_TOKENS 7168, the largest call measured to fit
Both are required in compose because they are coupled: MAX_TOKENS is checked before the forward
pass, so an oversized call is a clear 422 instead of reaching the cap as a 503. Pre-deploy floor
is nvidia-smi Free >= 15,400 MiB on GPU 1 (card peak 9,876 + scriberr 5,496).
2026-09-30 09:40:08 -07:00
..

intern-decision

⚠ STATUS 2026-09-30 0937 PT: built and accepted on GPU 3; the GPU 1 deploy is HELD. GPU 1's usable free memory is 15,442 MiB (nvidia-smi Free: 97,887 total − 640 driver-reserved − 81,806 used), below the 15,800 MiB pre-deploy floor. The 10,300 MiB budget was derived from total − used (16,081). The decision is with infra-ops/Prime; see "VRAM". Nothing is running on port 8033 yet.

Intern-Decision-4B typed decisions on fv-ml1 GPU 1, the utility card beside vllm-coder, the erp/meromero seats and scriberr. It replaced semif on Prime's ruling of 2026-09-30, ~0510 PT: "replace semif with intern-decision now". The pick came from the previous night's bench, docs/pfi/jev-candidates-bench-2026-09-30.md: the only candidate at least as accurate as SemIf on our sets, faster at our shape, and inside the memory budget.

Intern-Decision-4B (Shanghai AI Lab, Apache-2.0) is a Qwen3.5-4B finetune. It takes a state plus up to 16 named questions and scores every question's options in one forward pass with no decoding. It reads each answer from the logits just before a <decision> marker. services/intern-decision-serve/ loads it once and scores through the checkpoint's own inference.py (DecisionEngine.predict). It keeps semif-serve's HTTP surface, so a semif caller changes only the URL and the token. The contract is services/intern-decision-serve/intern-decision-serve.contract.md.

URL http://10.251.50.54:8033 = http://intern-decision.fv.internal:8033. /health is open; POSTs need Authorization: Bearer $(secret get intern-decision/api-token)
Model internlm/Intern-Decision-4B @ 0e5e6aa7d6d750e2b1504ba11a8136cb58aeb3cd, BF16, from /tank/aimodels/huggingface, mounted read-only and offline
Scoring code the snapshot's inference.py, sha256 c904e2c6…43b29863, checked before it is imported. Startup refuses any other file
Stack torch 2.10.0+cu128, transformers 5.17.0, fla 0.5.2, causal-conv1d 1.7.0: the bench's stack
Image intern-decision-serve:<version>, built on fv-ml1 from services/intern-decision-serve/
State None. The service never downloads (HF_HUB_OFFLINE=1, read-only mount). If the weights are ever lost, re-pull the pinned revision by hand.
Rollback stacks/semif is stopped, not removed. See Rollback below.

API (semif-serve's)

T=$(secret get intern-decision/api-token)
curl -s -H "Authorization: Bearer $T" http://intern-decision.fv.internal:8033/decide -d '{
  "id": "q1", "state": "Health checks passed in all three zones.",
  "question": "Did the deployment succeed?",
  "options": [{"id":"yes","description":"It succeeded."},{"id":"no","description":"It failed."}]}'
  • POST /decide takes one decision, {id, state, question, options[2..16]}.
  • POST /decide/shared takes {state, decisions: [{id, question, options}]}.
  • GET /health reports the pins, the limits and the chunking rule.
  • Order averaging: "orderings": "rotations", or "all" for ≤ 4 options, is still accepted, with semif's combined block. It is barely needed: the bench found Intern-Decision reorders 8 of 144 labels when the options are reversed, against SemIf's 30, and at one ordering it matched SemIf-with-rotations.
  • Each result carries semif's fields where they mean something: option_ids, probabilities, input_tokens, prompt_sha256, prompt_version, model, readout, probability_status, and on /decide also total_seconds and forward_seconds. New fields:
    • top;
    • confidence, the model's own;
    • calibration, the model's own temperature (T = 1.99241824);
    • native, the model's answer object, unchanged;
    • call (index, field, questions): which prompt the decision was asked in.
  • Past MAX_QUEUE (32) requests in progress, new POSTs get 429 busy before their body is read.

What a semif caller must know (every deliberate difference is in the contract, "Deltas")

  1. Option ids are shown to the model. The prompt prints A = <id>: <description>, so an id is part of the question. Use meaningful or neutral ids, not misleading ones.
  2. Decisions in one /decide/shared call are asked together, in one prompt. An answer can depend on the other questions in its call.
    • The bench measured this on Wyrd: 79/84 asked one decision at a time, 77/84 with a turn's 4 decisions in one prompt.
    • A call holds at most 16 questions. More are split greedily in request order: 1–16, 17–32, … . /health says so.
  3. probabilities are temperature-scaled by the model's shipped calibration. SemIf's were raw. The argmax is the same either way. This calibration is the vendor's, fitted on its own data, not ours.
  4. Gone:
    • option_logits (the model's runtime does not expose logits);
    • the prefix-cache timing fields;
    • /health.semif_commit;
    • per-workload calibration. Any workload is a 422, exactly as the deployed semif behaved with its empty table.
  5. MAX_TOKENS (8192) is per call: the state plus all its questions. A longer call is a 422 and is never truncated. Every call up to that limit fits the VRAM cap (measured). See VRAM below.

VRAM: the whole container ≤ 10,300 MiB

Budget (infra-ops, 2026-09-30): the container's whole nvidia-smi footprint, CUDA context included, must never exceed 10,300 MiB, whatever the request. That leaves scriberr (peak 5,496 MiB with its 120 s slices) 295 MiB of spare even when both hit their peaks at once. torch.cuda.set_per_process_memory_fraction caps only torch's allocator, so the cap is the budget minus the measured non-allocator overhead:

footprint  <=  VRAM_CAP_GIB  +  overhead      = 9,472 MiB (9.25 GiB)  +  662 MiB  =  10,134 MiB  (166 MiB under budget)

VRAM_CAP_GIB in .env is the single knob. It is applied before the weights load.

⚠ The budget assumed 16,081 MiB free on GPU 1. That is total − used. nvidia-smi's own Free is 15,442 MiB, because the driver reserves 640 MiB on every card (GPU 3 shows the same 640). Against the true free memory, a footprint that never collides with scriberr's 5,496 MiB peak must stay at or below 9,946 MiB:

cap footprint peak (measured, GPU 3) with scriberr at its peak, against 15,442 free calls that fit (16 questions × 16 options)
9.25 GiB (.env.example now) 10,134 MiB 188 MiB over every call the API accepts (8,191 tok)
9.0 GiB 9,876 MiB 70 MiB spare up to 7,168 tok; longer calls get 503. Every bench shape fits, at the same latency
measured on fv-ml1 GPU 3 (alone on the card), image 0.1.0 MiB
at rest (nvidia-smi, whole card less 2 MiB idle) 8,820 (torch reserved 8,160 + outside 660)
outside the allocator (CUDA context + loaded kernels), at rest / after 2×30 concurrent requests / after the latency shapes / after 4 × 8,191-token calls 660 / 660 / 662 / 662
peak, largest request the API accepts, capped at 9.25 GiB (card, 0.1 s sampling) 10,134; torch max_reserved = 9.25 GiB exactly, so the cap is what held it
the same request uncapped, for comparison 10,422 (allocator peak 9,760): over budget, so the cap is required
startup peak (weights + vision tower until the swap + warm-up) allocator 8.86 GiB. A cap below ~8.9 GiB cannot start (it fails closed; 8.5 was refused)
  • The largest request the API accepts is 64 decisions × 16 options, with each of its 4 calls at 8,191 tokens. Under the cap it answers 200 (1.67 s), not 503. Near the limit the allocator frees cached blocks before it fails, so the uncapped 9,760 MiB peak was partly cache. Every call up to MAX_TOKENS fits the cap.
  • Over the cap the answer is 503 out_of_memory. Memory goes back to the resting baseline and the service keeps answering. This was proven with a deliberately tight cap (acceptance).
  • One inference thread is load-bearing (contract INV-2). torch keeps CUDA state per host thread (cuBLAS handles and workspaces). The first build scored on anyio's 40-thread pool: a concurrent burst added 252 MiB outside the cap and 326 MiB inside it, and the peak reached 10,392 MiB. Load, warm-up and every call now run on one thread.
  • Released after bursts: when a call leaves reserved memory more than 512 MiB over the resting baseline, the engine calls empty_cache(). Measured: back to 7.97 GiB after every large request.
  • The rest footprint is 0.9 GB below the bench's 9,736 MiB. The service takes no images, so the vision tower (0.62 GiB) is swapped for a stub after the first warm-up, and the warm-up cache is released. Startup proves the swap changes nothing: the warm-up answer must be bit-identical (INV-7).

Latency

Loopback on fv-ml1 (GPU 3 alone on the card, image 0.1.0, cap 9.25 GiB). bench_shape.py from the Jev bench ran 3 runs × 20 requests, and the medians are shown with the run-median range. The bench's native Intern-Decision numbers are the reference.

shape end to end server bench (native, loopback) SemIf 0.1.4 (bench)
1 decision, short (~180 tok) 36.4 ms (36.4–36.5) 34.7 39 37
21 binary criteria, one /decide/shared (2 calls: 16 + 5) 82.3 ms (82.3–82.4) 78.9 88 131
1 decision over the ~3,900-token state 191.4 ms (191.4–191.5) 187.9 190 186
16 criteria over the ~3,900-token state (1 call, 4,579 tok) 216.8 ms (216.5–217.0) 211.1 215 505
largest accepted request (4 calls × 8,191 tok) 1,672 ms (1,672–1,675, N=3)

From nh3-dev and on GPU 1, next to the vLLM seats: pending the GPU 1 deploy (held, see the status banner).

Acceptance (2026-09-30)

Raw results: services/intern-decision-serve/acceptance/gpu3-2026-09-30/ (compare.json, one directory per process lifetime). The harness is the bench's own. bench_sets.py --backend semif speaks semif-serve's API, so it drove this service unchanged. compare.py scores it as the bench's analyze.py did.

check (GPU 3, cap 9.25 GiB, N = 3 fresh processes unless noted) result bench / floor
positive control, pooled 259, single ordering 240, 240, 240 bench 240 (4 repeats); the pooled set resolves ±4 pts
positive control, Wyrd /84, single 79, 79, 79 bench 79
row by row against the bench's native rows (560 rows, single / rotations / negative) 0 top changes, max Δp 0.000 in every repeat: bit-identical bench floor: 0 labels moved across 4 restarts
rotations, pooled / Wyrd 236 / 77 (×3) bench 236 / 77
negative control (descriptions rotated, 144): same top / follows the description / right vs the original gold 10 / 122 / 14 (×3) bench 10 / 122 / 14
A-vs-A in the process (authored144 twice) 0/144 flips, Δp 0.0 (×3)
A-vs-A across restarts (r1r2, r1r3, r2~r3; 560 rows, single and rotations) 0 flips, Δp 0.0
Wyrd, one /decide/shared per turn (4 decisions in one prompt) 78/84 (×3); 1/84 row differs from the bench's 77 expected: field names are positional here, decision names in the bench (contract)
more than 16 questions: 20 decisions in one request, against the same decisions asked as 16 + 4 2 calls [16, 4], all 20 tops equal, Δp 0, same prompt hashes
401: no token / wrong token, both POSTs; /health open 401 / 401; 200
429: 48 concurrent requests, MAX_QUEUE 32 32 answered, 16 × 429 busy
largest request the API accepts (64 × 16 options, 4 × 8,191 tok) at cap 9.25 GiB 200 ×3; card peak 10,134 MiB budget 10,300
over the cap (the same request at a deliberately tight cap of 8.9 GiB) 503 out_of_memory; reserved 7.969 GiB before, 7.969 after, next request 200; then all 560 single rows answered, bit-identical to the bench
fit under the cap: the largest 16-question call that answers 200 8,191 tok (the whole API) at 9.25 GiB; 7,168 tok at 9.0 GiB
cap below startup 8.5 GiB: startup refused (OOM loading), fails closed

On the deployed service (GPU 1): pending the deploy (held, see the status banner).

Building and deploying

# from nh3-dev. /opt/docker/src is root-owned, so create the version dir with sudo first.
ssh infra-ops@10.251.50.54 'sudo -n install -d -o infra-ops -g infra-ops /opt/docker/src/intern-decision-serve-X.Y.Z | cat'
tar -C services/intern-decision-serve -cf - --exclude=.venv --exclude=.pytest_cache --exclude=__pycache__ \
    --exclude=acceptance --exclude='*.egg-info' . \
  | ssh infra-ops@10.251.50.54 'tar -x -C /opt/docker/src/intern-decision-serve-X.Y.Z'
# on fv-ml1
cd /opt/docker/src/intern-decision-serve-X.Y.Z && docker build -t intern-decision-serve:X.Y.Z .
# from nh3-dev: compose.yaml, .env.example and this README (never .env)
scripts/deploy-stack.sh fv-ml1 intern-decision
# on fv-ml1: .env is infra-ops's, mode 0600. Rewrite it IN PLACE to keep the owner and mode.
cd /opt/docker/compose/intern-decision && new=$(sed 's/^IMAGE=.*/IMAGE=intern-decision-serve:X.Y.Z/' .env) \
  && printf '%s\n' "$new" > .env && docker compose config -q && docker compose up -d

Before any deploy onto GPU 1, check two things. If either fails, stop; do not squeeze scriberr.

  • nvidia-smi's own Free on GPU 1 must be at least 15,400 MiB: nvidia-smi -i 1 --query-gpu=memory.free --format=csv. That is our card peak of 9,876 MiB plus scriberr's 5,496, rounded up. Do not use total − used, which misses the driver's 640 MiB reserve.
  • Scriberr must not be running a job. This command must print 0: docker logs --since 2m scriberr | grep -c "Processing single-track job".

Startup fails closed. A container that never reaches healthy did not pass its own checks: the inference.py hash, the pinned snapshot, the warm-up, the text-only swap and the prompt hash. Read docker logs intern-decision. Record the deploy with scripts/ops-log.

To move the model forward, change REVISION and INFERENCE_PY_SHA256 in config.py. Re-run the bench's sets through the service (services/intern-decision-serve/acceptance/), and re-measure the VRAM before changing the cap.

Rollback

semif is stopped, not removed, and is kept as the rollback. Only on Prime's word:

cd /opt/docker/compose/intern-decision && docker compose stop     # first: the two do not both fit GPU 1
cd /opt/docker/compose/semif && docker compose start

Then check GPU 1's free memory against scriberr's peak. Callers go back to :8032 and secret get semif/api-token.