Files
esh-pfi-infrastructure/stacks/intern-decision
vh 6b201e1d4a scriberr to fv-ml1 GPU 3 (on-demand, steps aside to irv-ml1 A6000); intern-decision 32k-token calls (cap 14.4 GiB)
Prime 2026-09-30: move scriberr to GPU 3 and extend the Jev endpoint to 32k tokens.
Scriberr holds 0 VRAM idle; verified a 20-min job on GPU 3 at 5,496 MiB. With GPU 1
freed, intern-decision's measured card peak at MAX_TOKENS=32768 is 15,220 MiB against
a 15,437 MiB budget (n=3, 1 and 16 questions); 32,769 tokens is refused 422 up front.
JevBench v1.2.16 via /v1/systemone unchanged: 202/231, 0 diffs vs the bench.
2026-09-30 13:35:32 -07:00
..

intern-decision

LIVE since 2026-09-30 0941 PT on fv-ml1 GPU 1, with cap 9.0 GiB and MAX_TOKENS 7,168. Acceptance on the live service is below. semif is stopped and kept as the rollback.

Intern-Decision-4B typed decisions on fv-ml1 GPU 1, the utility card beside vllm-coder, the erp/meromero seats and scriberr. It replaced semif on Prime's ruling of 2026-09-30, ~0510 PT: "replace semif with intern-decision now". The pick came from the previous night's bench, docs/pfi/jev-candidates-bench-2026-09-30.md: the only candidate at least as accurate as SemIf on our sets, faster at our shape, and inside the memory budget.

Intern-Decision-4B (Shanghai AI Lab, Apache-2.0) is a Qwen3.5-4B finetune. It takes a state plus up to 16 named questions and scores every question's options in one forward pass with no decoding. It reads each answer from the logits just before a <decision> marker. services/intern-decision-serve/ loads it once and scores through the checkpoint's own inference.py (DecisionEngine.predict). It keeps semif-serve's HTTP surface, so a semif caller changes only the URL and the token. The contract is services/intern-decision-serve/intern-decision-serve.contract.md.

URL http://10.251.50.54:8033 = http://intern-decision.fv.internal:8033. /health is open; POSTs need Authorization: Bearer $(secret get intern-decision/api-token)
Model internlm/Intern-Decision-4B @ 0e5e6aa7d6d750e2b1504ba11a8136cb58aeb3cd, BF16, from /tank/aimodels/huggingface, mounted read-only and offline
Scoring code the snapshot's inference.py, sha256 c904e2c6…43b29863, checked before it is imported. Startup refuses any other file
Stack torch 2.10.0+cu128, transformers 5.17.0, fla 0.5.2, causal-conv1d 1.7.0: the bench's stack
Image intern-decision-serve:<version>, built on fv-ml1 from services/intern-decision-serve/
State None. The service never downloads (HF_HUB_OFFLINE=1, read-only mount). If the weights are ever lost, re-pull the pinned revision by hand.
Rollback stacks/semif is stopped, not removed. See Rollback below.

API (semif-serve's)

T=$(secret get intern-decision/api-token)
curl -s -H "Authorization: Bearer $T" http://intern-decision.fv.internal:8033/decide -d '{
  "id": "q1", "state": "Health checks passed in all three zones.",
  "question": "Did the deployment succeed?",
  "options": [{"id":"yes","description":"It succeeded."},{"id":"no","description":"It failed."}]}'
  • POST /decide takes one decision, {id, state, question, options[2..16]}.
  • POST /decide/shared takes {state, decisions: [{id, question, options}]}.
  • GET /health reports the pins, the limits and the chunking rule.
  • Order averaging: "orderings": "rotations", or "all" for ≤ 4 options, is still accepted, with semif's combined block. It is barely needed: the bench found Intern-Decision reorders 8 of 144 labels when the options are reversed, against SemIf's 30, and at one ordering it matched SemIf-with-rotations.
  • Each result carries semif's fields where they mean something: option_ids, probabilities, input_tokens, prompt_sha256, prompt_version, model, readout, probability_status, and on /decide also total_seconds and forward_seconds. New fields:
    • top;
    • confidence, the model's own;
    • calibration, the model's own temperature (T = 1.99241824);
    • native, the model's answer object, unchanged;
    • call (index, field, questions): which prompt the decision was asked in.
  • Past MAX_QUEUE (32) requests in progress, new POSTs get 429 busy before their body is read.

What a semif caller must know (every deliberate difference is in the contract, "Deltas")

  1. Option ids are shown to the model. The prompt prints A = <id>: <description>, so an id is part of the question. Use meaningful or neutral ids, not misleading ones.
  2. Decisions in one /decide/shared call are asked together, in one prompt. An answer can depend on the other questions in its call.
    • The bench measured this on Wyrd: 79/84 asked one decision at a time, 77/84 with a turn's 4 decisions in one prompt.
    • A call holds at most 16 questions. More are split greedily in request order: 1–16, 17–32, … . /health says so.
  3. probabilities are temperature-scaled by the model's shipped calibration. SemIf's were raw. The argmax is the same either way. This calibration is the vendor's, fitted on its own data, not ours.
  4. Gone:
    • option_logits (the model's runtime does not expose logits);
    • the prefix-cache timing fields;
    • /health.semif_commit;
    • per-workload calibration. Any workload is a 422, exactly as the deployed semif behaved with its empty table.
  5. MAX_TOKENS is 7,168 per call: the state plus all its questions. It is lower than the model's own 8,192 so that every call the API accepts fits the VRAM cap.
    • The count is taken before the forward pass, so a longer call is a clear 422 invalid_request ("Example has N tokens, above 7168; truncation is forbidden"). It is never truncated.
    • A request with more decisions is split into more calls, and each call must fit.

VRAM: 32k-token calls since 2026-09-30 1330 (scriberr moved to GPU 3)

Current setting: VRAM_CAP_GIB=14.4, MAX_TOKENS=32768 (Prime: move scriberr to GPU 3 and "extend the jev endpoint to hit 32k tokens if possible"). With scriberr gone, GPU 1 holds only the static vLLM seats and this service. This container may therefore use its rest (8,812 MiB) plus GPU 1's nvidia-smi Free (6,625 MiB), which is 15,437 MiB. The 14.4 GiB cap plus the ~660 MiB outside the allocator puts the card ceiling at ~15,408 MiB.

Measured on the live service (per-process nvidia-smi every 0.1 s; single noul question unless noted; n=3, deterministic):

call tokens card peak MiB wall (warm)
3,187 9,306 0.15 s
12,187 11,206 0.68 s
24,187 13,552 1.49 s
29,987 14,692 1.92 s
32,768 (limit) 15,220 2.12 s
32,765, 16 questions 15,220 2.15 s
32,769 refused 422 before the forward 0.30 s

Spare at the limit: 217 MiB. JevBench v1.2.16 through /v1/systemone after the change: 202/231, with 0 answer and 0 probability diffs against the bench's r1..r4. /decide is unchanged.

⚠ Cold-shape latency: the first call in a new length bucket after a (re)start costs ~6.5 s extra (3,001 and 4,000 words were slow; 3,002–3,500 and 4,097–5,000 were not). The fast kernels autotune per shape bucket, and the result is cached in-process. Warm calls are as tabled. A startup warm-up across the buckets would remove it; that is not done yet.

The table below is the ORIGINAL budget (9.0 GiB / 7,168 tokens, beside scriberr) and is kept for history.

VRAM (history): fits beside scriberr's peak, whatever the request

Budget (infra-ops, 2026-09-30): GPU 1 needs nvidia-smi Free ≥ 15,400 MiB before this service starts. That is our card peak of 9,876 MiB plus scriberr's peak of 5,496 MiB (with its 120 s slices), rounded up. So both can peak at the same moment. Read nvidia-smi's own Free, not total − used: the driver reserves 640 MiB on every card. The deploy at 0940 read 15,442.

torch.cuda.set_per_process_memory_fraction caps only torch's allocator. The two knobs in .env are therefore chosen together and change together:

card footprint  <=  VRAM_CAP_GIB 9.0 (9,216 MiB)  +  660 MiB outside the allocator (measured)  =  9,876 MiB
MAX_TOKENS 7168  =  the largest call measured to fit under that cap, so no accepted call reaches the 503
on the live service, GPU 1 (per-process nvidia-smi, 0.1 s sampling, 2,939 samples over 316 s of acceptance load) MiB
at rest after startup 8,812
at rest after the acceptance (512 MiB release slack keeps a little cache) 8,856
peak: largest requests, 64 decisions × 16 options, 4 calls × 7,168 tokens, N = 3 9,866; torch max_reserved 8.996 of 9.0 GiB
GPU 1 Free before the deploy / at rest after / lowest during the acceptance 15,442 / 6,581 / 5,569, which stays above scriberr's 5,496
  • ⚠ At 7,168 tokens the allocator reaches the cap with about 4 MiB to spare. It fits because the allocator frees its cache and retries before it fails.
    • Measured: every call up to the limit answered 200. That is 145 shared calls on the live service, including the size search near the boundary, 3 × 4 calls at 7,168 tokens with 16 questions, and 3 at 7,168 tokens with one question. No 503 was seen.
    • The allocator's cache state depends on request history. If a max-size call ever does return 503, lower MAX_TOKENS. Do not raise the cap.
  • Raising either knob needs a fresh measurement (acceptance/checks.py --checks fit,maxreq) and the budget re-checked against scriberr.
  • Over the cap the answer is 503 out_of_memory. Memory returns to the resting baseline and the service keeps answering. This was proven on GPU 3 with a deliberately tight cap.

Measured on GPU 3 before the deploy (alone on the card), which is where the knobs came from:

measured on fv-ml1 GPU 3 (alone on the card), image 0.1.0 MiB
at rest (nvidia-smi, whole card less 2 MiB idle) 8,820 (torch reserved 8,160 + outside 660)
outside the allocator (CUDA context + loaded kernels), at rest / after 2×30 concurrent requests / after the latency shapes / after 4 × 8,191-token calls 660 / 660 / 662 / 662
card peak at cap 9.0 GiB (fit search and a 64 × 8,191-token request, which got 503) 9,876
peak with 4 × 8,191-token calls at cap 9.25 GiB (answered 200) / uncapped 10,134 / 10,422 (allocator 9,760)
largest 16-question call that answers 200: at cap 9.0 / at 9.25 GiB 7,168 / 8,191 tokens
startup peak (weights + vision tower until the swap + warm-up) allocator 8.86 GiB. A cap below ~8.9 GiB cannot start (it fails closed; 8.5 was refused)
  • One inference thread is load-bearing (contract INV-2). torch keeps CUDA state per host thread (cuBLAS handles and workspaces). The first build scored on anyio's 40-thread pool: a concurrent burst added 252 MiB outside the cap and 326 MiB inside it, and the peak reached 10,392 MiB. Load, warm-up and every call now run on one thread.
  • Released after bursts: when a call leaves reserved memory more than 512 MiB over the resting baseline, the engine calls empty_cache(). Measured: back to 7.97 GiB after every large request.
  • The rest footprint is 0.9 GB below the bench's 9,736 MiB. The service takes no images, so the vision tower (0.62 GiB) is swapped for a stub after the first warm-up, and the warm-up cache is released. Startup proves the swap changes nothing: the warm-up answer must be bit-identical (INV-7).

Latency

Live, from nh3-dev (0942 PT). The URL is intern-decision.fv.internal:8033, and the round trip is 22.9 ms on average (19.5–27.7 ms over 10 pings). GPU 1 is shared with the vLLM seats and scriberr. The harness is bench_shape.py from the Jev bench, 3 runs × 20 requests; each cell is the median with the run-median range. "server" is the service's own time.

shape end to end from nh3-dev server GPU 3 loopback before deploy SemIf 0.1.4 (README, from nh3-dev)
1 decision, short (~180 tok) 53.9 ms (53.3–54.2) 35.2 36.4 69
21 binary criteria, one /decide/shared (2 calls: 16 + 5) 114.3 ms (114.1–114.3) 80.3 82.3 159
1 decision over the ~3,900-token state 211.2 ms (210.8–211.6) 186.8 191.4
16 criteria over the ~3,900-token state (1 call, 4,579 tok) 238.3 ms (237.8–238.7) 205.3 216.8 (bench loopback: 505)
largest request (4 calls × 7,168 tok), N = 3 1,519–1,543 ms

Server-side time on GPU 1 matches GPU 3 within a few ms, so the vLLM neighbours were not measurably slowing it at 0942. The rest of the end-to-end time is the network (~23 ms round trip plus the HTTP exchange).

Acceptance (2026-09-30)

Raw results: services/intern-decision-serve/acceptance/gpu3-2026-09-30/ (compare.json, one directory per process lifetime). The harness is the bench's own. bench_sets.py --backend semif speaks semif-serve's API, so it drove this service unchanged. compare.py scores it as the bench's analyze.py did.

On GPU 3 before the deploy (the harness and the floor)

check (GPU 3, cap 9.25 GiB, N = 3 fresh processes unless noted) result bench / floor
positive control, pooled 259, single ordering 240, 240, 240 bench 240 (4 repeats); the pooled set resolves ±4 pts
positive control, Wyrd /84, single 79, 79, 79 bench 79
row by row against the bench's native rows (560 rows, single / rotations / negative) 0 top changes, max Δp 0.000 in every repeat: bit-identical bench floor: 0 labels moved across 4 restarts
rotations, pooled / Wyrd 236 / 77 (×3) bench 236 / 77
negative control (descriptions rotated, 144): same top / follows the description / right vs the original gold 10 / 122 / 14 (×3) bench 10 / 122 / 14
A-vs-A in the process (authored144 twice) 0/144 flips, Δp 0.0 (×3)
A-vs-A across restarts (r1r2, r1r3, r2~r3; 560 rows, single and rotations) 0 flips, Δp 0.0
Wyrd, one /decide/shared per turn (4 decisions in one prompt) 78/84 (×3); 1/84 row differs from the bench's 77 expected: field names are positional here, decision names in the bench (contract)
more than 16 questions: 20 decisions in one request, against the same decisions asked as 16 + 4 2 calls [16, 4], all 20 tops equal, Δp 0, same prompt hashes
401: no token / wrong token, both POSTs; /health open 401 / 401; 200
429: 48 concurrent requests, MAX_QUEUE 32 32 answered, 16 × 429 busy
largest request at MAX_TOKENS 8,192 (64 × 16 options, 4 × 8,191 tok), cap 9.25 GiB 200 ×3; card peak 10,134 MiB superseded by the 9.0 / 7,168 knobs (VRAM)
over the cap (the same request at a deliberately tight cap of 8.9 GiB) 503 out_of_memory; reserved 7.969 GiB before, 7.969 after, next request 200; then all 560 single rows answered, bit-identical to the bench
fit under the cap: the largest 16-question call that answers 200 8,191 tok (the whole API) at 9.25 GiB; 7,168 tok at 9.0 GiB
cap below startup 8.5 GiB: startup refused (OOM loading), fails closed

On the live service (GPU 1, 0942–0946 PT, through http://intern-decision.fv.internal:8033 from nh3-dev)

Raw results: services/intern-decision-serve/acceptance/gpu1-live/.

check result reference
positive control, pooled 259, single ordering 240 bench 240; GPU 3 240 ×3
positive control, Wyrd /84 79 bench 79
row by row against the bench's native rows (560 single + 144 negative) 0 top changes, max Δp 0.000
negative control: same top / follows the description / right vs gold 10 / 122 / 14 bench 10 / 122 / 14
401 without or with a wrong token (both POSTs); /health open 401 / 401; 200
largest accepted request: 64 decisions × 16 options, 4 calls × 7,168 tok 200 ×3
one question at 7,168 tok 200 ×3; one token more (7,171) → 422 "above 7168"
any 503 across the whole live acceptance none (server log: 128 + 145 × 200, 18 × 422, 4 × 401)

The live pass is N = 1 for the positive control. It is read against the GPU 3 floor: the same image moved 0 of 560 rows across 3 restarts, and 0 against the bench's 4 restarts.

Building and deploying

# from nh3-dev. /opt/docker/src is root-owned, so create the version dir with sudo first.
ssh infra-ops@10.251.50.54 'sudo -n install -d -o infra-ops -g infra-ops /opt/docker/src/intern-decision-serve-X.Y.Z | cat'
tar -C services/intern-decision-serve -cf - --exclude=.venv --exclude=.pytest_cache --exclude=__pycache__ \
    --exclude=acceptance --exclude='*.egg-info' . \
  | ssh infra-ops@10.251.50.54 'tar -x -C /opt/docker/src/intern-decision-serve-X.Y.Z'
# on fv-ml1
cd /opt/docker/src/intern-decision-serve-X.Y.Z && docker build -t intern-decision-serve:X.Y.Z .
# from nh3-dev: compose.yaml, .env.example and this README (never .env)
scripts/deploy-stack.sh fv-ml1 intern-decision
# on fv-ml1: .env is infra-ops's, mode 0600. Rewrite it IN PLACE to keep the owner and mode.
cd /opt/docker/compose/intern-decision && new=$(sed 's/^IMAGE=.*/IMAGE=intern-decision-serve:X.Y.Z/' .env) \
  && printf '%s\n' "$new" > .env && docker compose config -q && docker compose up -d

Before any deploy onto GPU 1, check two things. If either fails, stop; do not squeeze scriberr.

  • nvidia-smi's own Free on GPU 1 must be at least 15,400 MiB: nvidia-smi -i 1 --query-gpu=memory.free --format=csv. That is our card peak of 9,876 MiB plus scriberr's 5,496, rounded up. Do not use total − used, which misses the driver's 640 MiB reserve.
  • Scriberr must not be running a job. This command must print 0: docker logs --since 2m scriberr | grep -c "Processing single-track job".

Startup fails closed. A container that never reaches healthy did not pass its own checks: the inference.py hash, the pinned snapshot, the warm-up, the text-only swap and the prompt hash. Read docker logs intern-decision. Record the deploy with scripts/ops-log.

To move the model forward, change REVISION and INFERENCE_PY_SHA256 in config.py. Re-run the bench's sets through the service (services/intern-decision-serve/acceptance/), and re-measure the VRAM before changing the cap.

Rollback

semif is stopped, not removed, and is kept as the rollback. Only on Prime's word:

cd /opt/docker/compose/intern-decision && docker compose stop     # first: the two do not both fit GPU 1
cd /opt/docker/compose/semif && docker compose start

Then check GPU 1's free memory against scriberr's peak. Callers go back to :8032 and secret get semif/api-token.