Prime 2026-09-30: move scriberr to GPU 3 and extend the Jev endpoint to 32k tokens. Scriberr holds 0 VRAM idle; verified a 20-min job on GPU 3 at 5,496 MiB. With GPU 1 freed, intern-decision's measured card peak at MAX_TOKENS=32768 is 15,220 MiB against a 15,437 MiB budget (n=3, 1 and 16 questions); 32,769 tokens is refused 422 up front. JevBench v1.2.16 via /v1/systemone unchanged: 202/231, 0 diffs vs the bench.
intern-decision
LIVE since 2026-09-30 0941 PT on fv-ml1 GPU 1, with cap 9.0 GiB and
MAX_TOKENS7,168. Acceptance on the live service is below.semifis stopped and kept as the rollback.
Intern-Decision-4B typed decisions on fv-ml1 GPU 1, the utility card beside vllm-coder,
the erp/meromero seats and scriberr. It replaced semif on Prime's ruling of 2026-09-30,
~0510 PT: "replace semif with intern-decision now". The pick came from the previous night's bench,
docs/pfi/jev-candidates-bench-2026-09-30.md: the only candidate at least as accurate as SemIf on
our sets, faster at our shape, and inside the memory budget.
Intern-Decision-4B (Shanghai AI Lab,
Apache-2.0) is a Qwen3.5-4B finetune. It takes a state plus up to 16 named questions and scores
every question's options in one forward pass with no decoding. It reads each answer from the
logits just before a <decision> marker. services/intern-decision-serve/ loads it once and
scores through the checkpoint's own inference.py (DecisionEngine.predict). It keeps
semif-serve's HTTP surface, so a semif caller changes only the URL and the token. The contract
is services/intern-decision-serve/intern-decision-serve.contract.md.
| URL | http://10.251.50.54:8033 = http://intern-decision.fv.internal:8033. /health is open; POSTs need Authorization: Bearer $(secret get intern-decision/api-token) |
| Model | internlm/Intern-Decision-4B @ 0e5e6aa7d6d750e2b1504ba11a8136cb58aeb3cd, BF16, from /tank/aimodels/huggingface, mounted read-only and offline |
| Scoring code | the snapshot's inference.py, sha256 c904e2c6…43b29863, checked before it is imported. Startup refuses any other file |
| Stack | torch 2.10.0+cu128, transformers 5.17.0, fla 0.5.2, causal-conv1d 1.7.0: the bench's stack |
| Image | intern-decision-serve:<version>, built on fv-ml1 from services/intern-decision-serve/ |
| State | None. The service never downloads (HF_HUB_OFFLINE=1, read-only mount). If the weights are ever lost, re-pull the pinned revision by hand. |
| Rollback | stacks/semif is stopped, not removed. See Rollback below. |
API (semif-serve's)
T=$(secret get intern-decision/api-token)
curl -s -H "Authorization: Bearer $T" http://intern-decision.fv.internal:8033/decide -d '{
"id": "q1", "state": "Health checks passed in all three zones.",
"question": "Did the deployment succeed?",
"options": [{"id":"yes","description":"It succeeded."},{"id":"no","description":"It failed."}]}'
POST /decidetakes one decision,{id, state, question, options[2..16]}.POST /decide/sharedtakes{state, decisions: [{id, question, options}]}.GET /healthreports the pins, the limits and the chunking rule.- Order averaging:
"orderings": "rotations", or"all"for ≤ 4 options, is still accepted, with semif'scombinedblock. It is barely needed: the bench found Intern-Decision reorders 8 of 144 labels when the options are reversed, against SemIf's 30, and at one ordering it matched SemIf-with-rotations. - Each result carries semif's fields where they mean something:
option_ids,probabilities,input_tokens,prompt_sha256,prompt_version,model,readout,probability_status, and on/decidealsototal_secondsandforward_seconds. New fields:top;confidence, the model's own;calibration, the model's own temperature (T = 1.99241824);native, the model's answer object, unchanged;call(index,field,questions): which prompt the decision was asked in.
- Past
MAX_QUEUE(32) requests in progress, new POSTs get429 busybefore their body is read.
What a semif caller must know (every deliberate difference is in the contract, "Deltas")
- Option ids are shown to the model. The prompt prints
A = <id>: <description>, so an id is part of the question. Use meaningful or neutral ids, not misleading ones. - Decisions in one
/decide/sharedcall are asked together, in one prompt. An answer can depend on the other questions in its call.- The bench measured this on Wyrd: 79/84 asked one decision at a time, 77/84 with a turn's 4 decisions in one prompt.
- A call holds at most 16 questions. More are split greedily in request order: 1–16,
17–32, … .
/healthsays so.
probabilitiesare temperature-scaled by the model's shipped calibration. SemIf's were raw. The argmax is the same either way. This calibration is the vendor's, fitted on its own data, not ours.- Gone:
option_logits(the model's runtime does not expose logits);- the prefix-cache timing fields;
/health.semif_commit;- per-workload calibration. Any
workloadis a 422, exactly as the deployed semif behaved with its empty table.
MAX_TOKENSis 7,168 per call: the state plus all its questions. It is lower than the model's own 8,192 so that every call the API accepts fits the VRAM cap.- The count is taken before the forward pass, so a longer call is a clear
422 invalid_request("Example has N tokens, above 7168; truncation is forbidden"). It is never truncated. - A request with more decisions is split into more calls, and each call must fit.
- The count is taken before the forward pass, so a longer call is a clear
VRAM: 32k-token calls since 2026-09-30 1330 (scriberr moved to GPU 3)
Current setting: VRAM_CAP_GIB=14.4, MAX_TOKENS=32768 (Prime: move scriberr to GPU 3 and "extend the jev
endpoint to hit 32k tokens if possible"). With scriberr gone, GPU 1 holds only the static vLLM seats and this
service. This container may therefore use its rest (8,812 MiB) plus GPU 1's nvidia-smi Free (6,625 MiB), which is
15,437 MiB. The 14.4 GiB cap plus the ~660 MiB outside the allocator puts the card ceiling at ~15,408 MiB.
Measured on the live service (per-process nvidia-smi every 0.1 s; single noul question unless noted; n=3, deterministic):
| call tokens | card peak MiB | wall (warm) |
|---|---|---|
| 3,187 | 9,306 | 0.15 s |
| 12,187 | 11,206 | 0.68 s |
| 24,187 | 13,552 | 1.49 s |
| 29,987 | 14,692 | 1.92 s |
| 32,768 (limit) | 15,220 | 2.12 s |
| 32,765, 16 questions | 15,220 | 2.15 s |
| 32,769 | refused 422 before the forward | 0.30 s |
Spare at the limit: 217 MiB. JevBench v1.2.16 through /v1/systemone after the change: 202/231, with 0 answer and 0
probability diffs against the bench's r1..r4. /decide is unchanged.
⚠ Cold-shape latency: the first call in a new length bucket after a (re)start costs ~6.5 s extra (3,001 and 4,000 words were slow; 3,002–3,500 and 4,097–5,000 were not). The fast kernels autotune per shape bucket, and the result is cached in-process. Warm calls are as tabled. A startup warm-up across the buckets would remove it; that is not done yet.
The table below is the ORIGINAL budget (9.0 GiB / 7,168 tokens, beside scriberr) and is kept for history.
VRAM (history): fits beside scriberr's peak, whatever the request
Budget (infra-ops, 2026-09-30): GPU 1 needs nvidia-smi Free ≥ 15,400 MiB before this
service starts. That is our card peak of 9,876 MiB plus scriberr's peak of 5,496 MiB (with its
120 s slices), rounded up. So both can peak at the same moment. Read nvidia-smi's own Free, not
total − used: the driver reserves 640 MiB on every card. The deploy at 0940 read 15,442.
torch.cuda.set_per_process_memory_fraction caps only torch's allocator. The two knobs in .env
are therefore chosen together and change together:
card footprint <= VRAM_CAP_GIB 9.0 (9,216 MiB) + 660 MiB outside the allocator (measured) = 9,876 MiB
MAX_TOKENS 7168 = the largest call measured to fit under that cap, so no accepted call reaches the 503
| on the live service, GPU 1 (per-process nvidia-smi, 0.1 s sampling, 2,939 samples over 316 s of acceptance load) | MiB |
|---|---|
| at rest after startup | 8,812 |
| at rest after the acceptance (512 MiB release slack keeps a little cache) | 8,856 |
| peak: largest requests, 64 decisions × 16 options, 4 calls × 7,168 tokens, N = 3 | 9,866; torch max_reserved 8.996 of 9.0 GiB |
GPU 1 Free before the deploy / at rest after / lowest during the acceptance |
15,442 / 6,581 / 5,569, which stays above scriberr's 5,496 |
- ⚠ At 7,168 tokens the allocator reaches the cap with about 4 MiB to spare. It fits because
the allocator frees its cache and retries before it fails.
- Measured: every call up to the limit answered 200. That is 145 shared calls on the live service, including the size search near the boundary, 3 × 4 calls at 7,168 tokens with 16 questions, and 3 at 7,168 tokens with one question. No 503 was seen.
- The allocator's cache state depends on request history. If a max-size call ever does return
503, lower
MAX_TOKENS. Do not raise the cap.
- Raising either knob needs a fresh measurement (
acceptance/checks.py --checks fit,maxreq) and the budget re-checked against scriberr. - Over the cap the answer is 503
out_of_memory. Memory returns to the resting baseline and the service keeps answering. This was proven on GPU 3 with a deliberately tight cap.
Measured on GPU 3 before the deploy (alone on the card), which is where the knobs came from:
| measured on fv-ml1 GPU 3 (alone on the card), image 0.1.0 | MiB |
|---|---|
| at rest (nvidia-smi, whole card less 2 MiB idle) | 8,820 (torch reserved 8,160 + outside 660) |
| outside the allocator (CUDA context + loaded kernels), at rest / after 2×30 concurrent requests / after the latency shapes / after 4 × 8,191-token calls | 660 / 660 / 662 / 662 |
| card peak at cap 9.0 GiB (fit search and a 64 × 8,191-token request, which got 503) | 9,876 |
| peak with 4 × 8,191-token calls at cap 9.25 GiB (answered 200) / uncapped | 10,134 / 10,422 (allocator 9,760) |
| largest 16-question call that answers 200: at cap 9.0 / at 9.25 GiB | 7,168 / 8,191 tokens |
| startup peak (weights + vision tower until the swap + warm-up) | allocator 8.86 GiB. A cap below ~8.9 GiB cannot start (it fails closed; 8.5 was refused) |
- One inference thread is load-bearing (contract INV-2). torch keeps CUDA state per host thread (cuBLAS handles and workspaces). The first build scored on anyio's 40-thread pool: a concurrent burst added 252 MiB outside the cap and 326 MiB inside it, and the peak reached 10,392 MiB. Load, warm-up and every call now run on one thread.
- Released after bursts: when a call leaves reserved memory more than 512 MiB over the resting
baseline, the engine calls
empty_cache(). Measured: back to 7.97 GiB after every large request. - The rest footprint is 0.9 GB below the bench's 9,736 MiB. The service takes no images, so the vision tower (0.62 GiB) is swapped for a stub after the first warm-up, and the warm-up cache is released. Startup proves the swap changes nothing: the warm-up answer must be bit-identical (INV-7).
Latency
Live, from nh3-dev (0942 PT). The URL is intern-decision.fv.internal:8033, and the round trip
is 22.9 ms on average (19.5–27.7 ms over 10 pings). GPU 1 is shared with the vLLM seats and
scriberr. The harness is bench_shape.py from the Jev bench, 3 runs × 20 requests; each cell is
the median with the run-median range. "server" is the service's own time.
| shape | end to end from nh3-dev | server | GPU 3 loopback before deploy | SemIf 0.1.4 (README, from nh3-dev) |
|---|---|---|---|---|
| 1 decision, short (~180 tok) | 53.9 ms (53.3–54.2) | 35.2 | 36.4 | 69 |
21 binary criteria, one /decide/shared (2 calls: 16 + 5) |
114.3 ms (114.1–114.3) | 80.3 | 82.3 | 159 |
| 1 decision over the ~3,900-token state | 211.2 ms (210.8–211.6) | 186.8 | 191.4 | |
| 16 criteria over the ~3,900-token state (1 call, 4,579 tok) | 238.3 ms (237.8–238.7) | 205.3 | 216.8 | (bench loopback: 505) |
| largest request (4 calls × 7,168 tok), N = 3 | 1,519–1,543 ms |
Server-side time on GPU 1 matches GPU 3 within a few ms, so the vLLM neighbours were not measurably slowing it at 0942. The rest of the end-to-end time is the network (~23 ms round trip plus the HTTP exchange).
Acceptance (2026-09-30)
Raw results: services/intern-decision-serve/acceptance/gpu3-2026-09-30/ (compare.json, one
directory per process lifetime). The harness is the bench's own. bench_sets.py --backend semif speaks semif-serve's API, so it drove this service unchanged. compare.py
scores it as the bench's analyze.py did.
On GPU 3 before the deploy (the harness and the floor)
| check (GPU 3, cap 9.25 GiB, N = 3 fresh processes unless noted) | result | bench / floor |
|---|---|---|
| positive control, pooled 259, single ordering | 240, 240, 240 | bench 240 (4 repeats); the pooled set resolves ±4 pts |
| positive control, Wyrd /84, single | 79, 79, 79 | bench 79 |
| row by row against the bench's native rows (560 rows, single / rotations / negative) | 0 top changes, max Δp 0.000 in every repeat: bit-identical | bench floor: 0 labels moved across 4 restarts |
| rotations, pooled / Wyrd | 236 / 77 (×3) | bench 236 / 77 |
| negative control (descriptions rotated, 144): same top / follows the description / right vs the original gold | 10 / 122 / 14 (×3) | bench 10 / 122 / 14 |
| A-vs-A in the process (authored144 twice) | 0/144 flips, Δp 0.0 (×3) | |
| A-vs-A across restarts (r1 |
0 flips, Δp 0.0 | |
Wyrd, one /decide/shared per turn (4 decisions in one prompt) |
78/84 (×3); 1/84 row differs from the bench's 77 | expected: field names are positional here, decision names in the bench (contract) |
| more than 16 questions: 20 decisions in one request, against the same decisions asked as 16 + 4 | 2 calls [16, 4], all 20 tops equal, Δp 0, same prompt hashes |
|
401: no token / wrong token, both POSTs; /health open |
401 / 401; 200 | |
429: 48 concurrent requests, MAX_QUEUE 32 |
32 answered, 16 × 429 busy |
|
largest request at MAX_TOKENS 8,192 (64 × 16 options, 4 × 8,191 tok), cap 9.25 GiB |
200 ×3; card peak 10,134 MiB | superseded by the 9.0 / 7,168 knobs (VRAM) |
| over the cap (the same request at a deliberately tight cap of 8.9 GiB) | 503 out_of_memory; reserved 7.969 GiB before, 7.969 after, next request 200; then all 560 single rows answered, bit-identical to the bench |
|
| fit under the cap: the largest 16-question call that answers 200 | 8,191 tok (the whole API) at 9.25 GiB; 7,168 tok at 9.0 GiB | |
| cap below startup | 8.5 GiB: startup refused (OOM loading), fails closed |
On the live service (GPU 1, 0942–0946 PT, through http://intern-decision.fv.internal:8033 from nh3-dev)
Raw results: services/intern-decision-serve/acceptance/gpu1-live/.
| check | result | reference |
|---|---|---|
| positive control, pooled 259, single ordering | 240 | bench 240; GPU 3 240 ×3 |
| positive control, Wyrd /84 | 79 | bench 79 |
| row by row against the bench's native rows (560 single + 144 negative) | 0 top changes, max Δp 0.000 | |
| negative control: same top / follows the description / right vs gold | 10 / 122 / 14 | bench 10 / 122 / 14 |
401 without or with a wrong token (both POSTs); /health open |
401 / 401; 200 | |
| largest accepted request: 64 decisions × 16 options, 4 calls × 7,168 tok | 200 ×3 | |
| one question at 7,168 tok | 200 ×3; one token more (7,171) → 422 "above 7168" | |
| any 503 across the whole live acceptance | none (server log: 128 + 145 × 200, 18 × 422, 4 × 401) |
The live pass is N = 1 for the positive control. It is read against the GPU 3 floor: the same image moved 0 of 560 rows across 3 restarts, and 0 against the bench's 4 restarts.
Building and deploying
# from nh3-dev. /opt/docker/src is root-owned, so create the version dir with sudo first.
ssh infra-ops@10.251.50.54 'sudo -n install -d -o infra-ops -g infra-ops /opt/docker/src/intern-decision-serve-X.Y.Z | cat'
tar -C services/intern-decision-serve -cf - --exclude=.venv --exclude=.pytest_cache --exclude=__pycache__ \
--exclude=acceptance --exclude='*.egg-info' . \
| ssh infra-ops@10.251.50.54 'tar -x -C /opt/docker/src/intern-decision-serve-X.Y.Z'
# on fv-ml1
cd /opt/docker/src/intern-decision-serve-X.Y.Z && docker build -t intern-decision-serve:X.Y.Z .
# from nh3-dev: compose.yaml, .env.example and this README (never .env)
scripts/deploy-stack.sh fv-ml1 intern-decision
# on fv-ml1: .env is infra-ops's, mode 0600. Rewrite it IN PLACE to keep the owner and mode.
cd /opt/docker/compose/intern-decision && new=$(sed 's/^IMAGE=.*/IMAGE=intern-decision-serve:X.Y.Z/' .env) \
&& printf '%s\n' "$new" > .env && docker compose config -q && docker compose up -d
Before any deploy onto GPU 1, check two things. If either fails, stop; do not squeeze scriberr.
- nvidia-smi's own
Freeon GPU 1 must be at least 15,400 MiB:nvidia-smi -i 1 --query-gpu=memory.free --format=csv. That is our card peak of 9,876 MiB plus scriberr's 5,496, rounded up. Do not use total − used, which misses the driver's 640 MiB reserve. - Scriberr must not be running a job. This command must print 0:
docker logs --since 2m scriberr | grep -c "Processing single-track job".
Startup fails closed. A container that never reaches healthy did not pass its own checks: the
inference.py hash, the pinned snapshot, the warm-up, the text-only swap and the prompt hash.
Read docker logs intern-decision. Record the deploy with scripts/ops-log.
To move the model forward, change REVISION and INFERENCE_PY_SHA256 in config.py.
Re-run the bench's sets through the service (services/intern-decision-serve/acceptance/), and
re-measure the VRAM before changing the cap.
Rollback
semif is stopped, not removed, and is kept as the rollback. Only on Prime's word:
cd /opt/docker/compose/intern-decision && docker compose stop # first: the two do not both fit GPU 1
cd /opt/docker/compose/semif && docker compose start
Then check GPU 1's free memory against scriberr's peak. Callers go back to :8032 and
secret get semif/api-token.