Both are required in compose because they are coupled: MAX_TOKENS is checked before the forward pass, so an oversized call is a clear 422 instead of reaching the cap as a 503. Pre-deploy floor is nvidia-smi Free >= 15,400 MiB on GPU 1 (card peak 9,876 + scriberr 5,496).
221 lines
14 KiB
Markdown
221 lines
14 KiB
Markdown
# intern-decision
|
||
|
||
> ⚠ **STATUS 2026-09-30 0937 PT: built and accepted on GPU 3; the GPU 1 deploy is HELD.** GPU 1's
|
||
> usable free memory is 15,442 MiB (nvidia-smi `Free`: 97,887 total − 640 driver-reserved −
|
||
> 81,806 used), below the 15,800 MiB pre-deploy floor. The 10,300 MiB budget was derived from
|
||
> total − used (16,081). The decision is with infra-ops/Prime; see "VRAM". Nothing is running
|
||
> on port 8033 yet.
|
||
|
||
**Intern-Decision-4B typed decisions** on **fv-ml1 GPU 1**, the utility card beside `vllm-coder`,
|
||
the erp/meromero seats and scriberr. It **replaced `semif`** on Prime's ruling of 2026-09-30,
|
||
~0510 PT: "replace semif with intern-decision now". The pick came from the previous night's bench,
|
||
`docs/pfi/jev-candidates-bench-2026-09-30.md`: the only candidate at least as accurate as SemIf on
|
||
our sets, faster at our shape, and inside the memory budget.
|
||
|
||
[Intern-Decision-4B](https://huggingface.co/internlm/Intern-Decision-4B) (Shanghai AI Lab,
|
||
Apache-2.0) is a Qwen3.5-4B finetune. It takes a state plus up to 16 named questions and scores
|
||
every question's options in **one forward pass with no decoding**. It reads each answer from the
|
||
logits just before a `<decision>` marker. `services/intern-decision-serve/` loads it once and
|
||
scores through the checkpoint's **own** `inference.py` (`DecisionEngine.predict`). It keeps
|
||
**semif-serve's HTTP surface**, so a semif caller changes only the URL and the token. The contract
|
||
is `services/intern-decision-serve/intern-decision-serve.contract.md`.
|
||
|
||
| | |
|
||
|---|---|
|
||
| **URL** | `http://10.251.50.54:8033` = `http://intern-decision.fv.internal:8033`. `/health` is open; POSTs need `Authorization: Bearer $(secret get intern-decision/api-token)` |
|
||
| **Model** | `internlm/Intern-Decision-4B` @ `0e5e6aa7d6d750e2b1504ba11a8136cb58aeb3cd`, BF16, from `/tank/aimodels/huggingface`, mounted read-only and offline |
|
||
| **Scoring code** | the snapshot's `inference.py`, sha256 `c904e2c6…43b29863`, checked before it is imported. Startup refuses any other file |
|
||
| **Stack** | torch `2.10.0+cu128`, transformers `5.17.0`, fla `0.5.2`, causal-conv1d `1.7.0`: the bench's stack |
|
||
| **Image** | `intern-decision-serve:<version>`, built on fv-ml1 from `services/intern-decision-serve/` |
|
||
| **State** | None. The service never downloads (`HF_HUB_OFFLINE=1`, read-only mount). If the weights are ever lost, re-pull the pinned revision by hand. |
|
||
| **Rollback** | `stacks/semif` is stopped, not removed. See **Rollback** below. |
|
||
|
||
## API (semif-serve's)
|
||
|
||
```bash
|
||
T=$(secret get intern-decision/api-token)
|
||
curl -s -H "Authorization: Bearer $T" http://intern-decision.fv.internal:8033/decide -d '{
|
||
"id": "q1", "state": "Health checks passed in all three zones.",
|
||
"question": "Did the deployment succeed?",
|
||
"options": [{"id":"yes","description":"It succeeded."},{"id":"no","description":"It failed."}]}'
|
||
```
|
||
|
||
- `POST /decide` takes one decision, `{id, state, question, options[2..16]}`.
|
||
- `POST /decide/shared` takes `{state, decisions: [{id, question, options}]}`.
|
||
- `GET /health` reports the pins, the limits and the chunking rule.
|
||
- **Order averaging:** `"orderings": "rotations"`, or `"all"` for ≤ 4 options, is still accepted,
|
||
with semif's `combined` block. **It is barely needed:** the bench found Intern-Decision reorders
|
||
8 of 144 labels when the options are reversed, against SemIf's 30, and at one ordering it matched
|
||
SemIf-with-rotations.
|
||
- Each result carries semif's fields where they mean something: `option_ids`, `probabilities`,
|
||
`input_tokens`, `prompt_sha256`, `prompt_version`, `model`, `readout`, `probability_status`, and
|
||
on `/decide` also `total_seconds` and `forward_seconds`. New fields:
|
||
- `top`;
|
||
- `confidence`, the model's own;
|
||
- `calibration`, the model's own temperature (T = 1.99241824);
|
||
- `native`, the model's answer object, unchanged;
|
||
- `call` (`index`, `field`, `questions`): which prompt the decision was asked in.
|
||
- Past `MAX_QUEUE` (32) requests in progress, new POSTs get `429 busy` before their body is read.
|
||
|
||
### What a semif caller must know (every deliberate difference is in the contract, "Deltas")
|
||
|
||
1. **Option ids are shown to the model.** The prompt prints `A = <id>: <description>`, so an
|
||
id is part of the question. Use meaningful or neutral ids, not misleading ones.
|
||
2. **Decisions in one `/decide/shared` call are asked together, in one prompt.** An answer can
|
||
depend on the other questions in its call.
|
||
- The bench measured this on Wyrd: 79/84 asked one decision at a time, 77/84 with a turn's
|
||
4 decisions in one prompt.
|
||
- A call holds **at most 16 questions**. More are split greedily in request order: 1–16,
|
||
17–32, … . `/health` says so.
|
||
3. **`probabilities` are temperature-scaled** by the model's shipped calibration. SemIf's were
|
||
raw. The argmax is the same either way. This calibration is the vendor's, fitted on its own
|
||
data, not ours.
|
||
4. **Gone:**
|
||
- `option_logits` (the model's runtime does not expose logits);
|
||
- the prefix-cache timing fields;
|
||
- `/health.semif_commit`;
|
||
- per-workload calibration. Any `workload` is a 422, exactly as the deployed semif behaved
|
||
with its empty table.
|
||
5. **`MAX_TOKENS` (8192) is per call:** the state plus all its questions. A longer call is a 422
|
||
and is never truncated. Every call up to that limit fits the VRAM cap (measured). See VRAM
|
||
below.
|
||
|
||
## VRAM: the whole container ≤ 10,300 MiB
|
||
|
||
**Budget (infra-ops, 2026-09-30):** the container's **whole** nvidia-smi footprint, CUDA context
|
||
included, must never exceed **10,300 MiB**, whatever the request. That leaves scriberr (peak
|
||
5,496 MiB with its 120 s slices) 295 MiB of spare even when both hit their peaks at once.
|
||
`torch.cuda.set_per_process_memory_fraction` caps only torch's allocator, so the cap is the budget
|
||
minus the measured non-allocator overhead:
|
||
|
||
```
|
||
footprint <= VRAM_CAP_GIB + overhead = 9,472 MiB (9.25 GiB) + 662 MiB = 10,134 MiB (166 MiB under budget)
|
||
```
|
||
|
||
`VRAM_CAP_GIB` in `.env` is **the single knob**. It is applied before the weights load.
|
||
|
||
⚠ **The budget assumed 16,081 MiB free on GPU 1. That is total − used.** nvidia-smi's own `Free` is
|
||
**15,442 MiB**, because the driver reserves 640 MiB on every card (GPU 3 shows the same 640).
|
||
Against the true free memory, a footprint that never collides with scriberr's 5,496 MiB peak must
|
||
stay at or below **9,946 MiB**:
|
||
|
||
| cap | footprint peak (measured, GPU 3) | with scriberr at its peak, against 15,442 free | calls that fit (16 questions × 16 options) |
|
||
|---|---|---|---|
|
||
| **9.25 GiB** (`.env.example` now) | 10,134 MiB | **188 MiB over** | every call the API accepts (8,191 tok) |
|
||
| **9.0 GiB** | 9,876 MiB | 70 MiB spare | up to 7,168 tok; longer calls get 503. Every bench shape fits, at the same latency |
|
||
|
||
| measured on fv-ml1 GPU 3 (alone on the card), image 0.1.0 | MiB |
|
||
|---|---|
|
||
| **at rest** (nvidia-smi, whole card less 2 MiB idle) | **8,820** (torch reserved 8,160 + outside 660) |
|
||
| outside the allocator (CUDA context + loaded kernels), at rest / after 2×30 concurrent requests / after the latency shapes / after 4 × 8,191-token calls | 660 / 660 / 662 / 662 |
|
||
| **peak, largest request the API accepts, capped at 9.25 GiB** (card, 0.1 s sampling) | **10,134**; torch `max_reserved` = 9.25 GiB exactly, so the cap is what held it |
|
||
| the same request **uncapped**, for comparison | 10,422 (allocator peak 9,760): **over budget, so the cap is required** |
|
||
| startup peak (weights + vision tower until the swap + warm-up) | allocator 8.86 GiB. **A cap below ~8.9 GiB cannot start** (it fails closed; 8.5 was refused) |
|
||
|
||
- **The largest request the API accepts** is 64 decisions × 16 options, with each of its 4 calls
|
||
at 8,191 tokens. Under the cap it answers **200** (1.67 s), not 503. Near the limit the
|
||
allocator frees cached blocks before it fails, so the uncapped 9,760 MiB peak was partly cache.
|
||
Every call up to `MAX_TOKENS` fits the cap.
|
||
- **Over the cap the answer is 503 `out_of_memory`.** Memory goes back to the resting baseline
|
||
and the service keeps answering. This was proven with a deliberately tight cap (acceptance).
|
||
- **One inference thread is load-bearing** (contract INV-2). torch keeps CUDA state per host
|
||
thread (cuBLAS handles and workspaces). The first build scored on anyio's 40-thread pool: a
|
||
concurrent burst added 252 MiB outside the cap and 326 MiB inside it, and the peak reached
|
||
10,392 MiB. Load, warm-up and every call now run on one thread.
|
||
- **Released after bursts:** when a call leaves reserved memory more than 512 MiB over the resting
|
||
baseline, the engine calls `empty_cache()`. Measured: back to 7.97 GiB after every large
|
||
request.
|
||
- The rest footprint is 0.9 GB below the bench's 9,736 MiB. The service takes no images, so the
|
||
vision tower (0.62 GiB) is swapped for a stub after the first warm-up, and the warm-up cache is
|
||
released. Startup proves the swap changes nothing: the warm-up answer must be bit-identical
|
||
(INV-7).
|
||
|
||
## Latency
|
||
|
||
Loopback on fv-ml1 (GPU 3 alone on the card, image 0.1.0, cap 9.25 GiB). `bench_shape.py`
|
||
from the Jev bench ran 3 runs × 20 requests, and the medians are shown with the run-median
|
||
range. The bench's native Intern-Decision numbers are the reference.
|
||
|
||
| shape | end to end | server | bench (native, loopback) | SemIf 0.1.4 (bench) |
|
||
|---|---|---|---|---|
|
||
| 1 decision, short (~180 tok) | 36.4 ms (36.4–36.5) | 34.7 | 39 | 37 |
|
||
| **21 binary criteria**, one `/decide/shared` (2 calls: 16 + 5) | **82.3 ms** (82.3–82.4) | 78.9 | 88 | 131 |
|
||
| 1 decision over the ~3,900-token state | 191.4 ms (191.4–191.5) | 187.9 | 190 | 186 |
|
||
| **16 criteria over the ~3,900-token state** (1 call, 4,579 tok) | **216.8 ms** (216.5–217.0) | 211.1 | 215 | 505 |
|
||
| largest accepted request (4 calls × 8,191 tok) | 1,672 ms (1,672–1,675, N=3) | | | |
|
||
|
||
*From nh3-dev and on GPU 1, next to the vLLM seats: pending the GPU 1 deploy (held, see the status banner).*
|
||
|
||
## Acceptance (2026-09-30)
|
||
|
||
Raw results: `services/intern-decision-serve/acceptance/gpu3-2026-09-30/` (`compare.json`, one
|
||
directory per process lifetime). **The harness is the bench's own.** `bench_sets.py
|
||
--backend semif` speaks semif-serve's API, so it drove this service unchanged. `compare.py`
|
||
scores it as the bench's `analyze.py` did.
|
||
|
||
| check (GPU 3, cap 9.25 GiB, N = 3 fresh processes unless noted) | result | bench / floor |
|
||
|---|---|---|
|
||
| **positive control**, pooled 259, single ordering | **240, 240, 240** | bench 240 (4 repeats); the pooled set resolves ±4 pts |
|
||
| positive control, Wyrd /84, single | **79, 79, 79** | bench 79 |
|
||
| row by row against the bench's native rows (560 rows, single / rotations / negative) | **0 top changes, max Δp 0.000** in every repeat: bit-identical | bench floor: 0 labels moved across 4 restarts |
|
||
| rotations, pooled / Wyrd | 236 / 77 (×3) | bench 236 / 77 |
|
||
| **negative control** (descriptions rotated, 144): same top / follows the description / right vs the original gold | **10 / 122 / 14** (×3) | bench 10 / 122 / 14 |
|
||
| A-vs-A in the process (authored144 twice) | 0/144 flips, Δp 0.0 (×3) | |
|
||
| A-vs-A across restarts (r1~r2, r1~r3, r2~r3; 560 rows, single and rotations) | 0 flips, Δp 0.0 | |
|
||
| Wyrd, one `/decide/shared` per turn (4 decisions in one prompt) | 78/84 (×3); 1/84 row differs from the bench's 77 | expected: field names are positional here, decision names in the bench (contract) |
|
||
| more than 16 questions: 20 decisions in one request, against the same decisions asked as 16 + 4 | 2 calls `[16, 4]`, all 20 tops equal, Δp 0, same prompt hashes | |
|
||
| 401: no token / wrong token, both POSTs; `/health` open | 401 / 401; 200 | |
|
||
| 429: 48 concurrent requests, `MAX_QUEUE` 32 | 32 answered, 16 × 429 `busy` | |
|
||
| largest request the API accepts (64 × 16 options, 4 × 8,191 tok) at cap 9.25 GiB | **200** ×3; card peak 10,134 MiB | budget 10,300 |
|
||
| **over the cap** (the same request at a deliberately tight cap of 8.9 GiB) | **503 `out_of_memory`**; reserved 7.969 GiB before, 7.969 after, next request 200; then all 560 single rows answered, bit-identical to the bench | |
|
||
| fit under the cap: the largest 16-question call that answers 200 | 8,191 tok (the whole API) at 9.25 GiB; 7,168 tok at 9.0 GiB | |
|
||
| cap below startup | 8.5 GiB: startup refused (OOM loading), fails closed | |
|
||
|
||
*On the deployed service (GPU 1): pending the deploy (held, see the status banner).*
|
||
|
||
## Building and deploying
|
||
|
||
```bash
|
||
# from nh3-dev. /opt/docker/src is root-owned, so create the version dir with sudo first.
|
||
ssh infra-ops@10.251.50.54 'sudo -n install -d -o infra-ops -g infra-ops /opt/docker/src/intern-decision-serve-X.Y.Z | cat'
|
||
tar -C services/intern-decision-serve -cf - --exclude=.venv --exclude=.pytest_cache --exclude=__pycache__ \
|
||
--exclude=acceptance --exclude='*.egg-info' . \
|
||
| ssh infra-ops@10.251.50.54 'tar -x -C /opt/docker/src/intern-decision-serve-X.Y.Z'
|
||
# on fv-ml1
|
||
cd /opt/docker/src/intern-decision-serve-X.Y.Z && docker build -t intern-decision-serve:X.Y.Z .
|
||
# from nh3-dev: compose.yaml, .env.example and this README (never .env)
|
||
scripts/deploy-stack.sh fv-ml1 intern-decision
|
||
# on fv-ml1: .env is infra-ops's, mode 0600. Rewrite it IN PLACE to keep the owner and mode.
|
||
cd /opt/docker/compose/intern-decision && new=$(sed 's/^IMAGE=.*/IMAGE=intern-decision-serve:X.Y.Z/' .env) \
|
||
&& printf '%s\n' "$new" > .env && docker compose config -q && docker compose up -d
|
||
```
|
||
|
||
**Before any deploy onto GPU 1**, check two things. If either fails, stop; do not squeeze
|
||
scriberr.
|
||
- nvidia-smi's own `Free` on GPU 1 must be at least **15,400 MiB**:
|
||
`nvidia-smi -i 1 --query-gpu=memory.free --format=csv`. That is our card peak of 9,876 MiB plus
|
||
scriberr's 5,496, rounded up. Do not use total − used, which misses the driver's 640 MiB
|
||
reserve.
|
||
- Scriberr must not be running a job. This command must print 0:
|
||
`docker logs --since 2m scriberr | grep -c "Processing single-track job"`.
|
||
|
||
Startup fails closed. A container that never reaches healthy did not pass its own checks: the
|
||
`inference.py` hash, the pinned snapshot, the warm-up, the text-only swap and the prompt hash.
|
||
Read `docker logs intern-decision`. Record the deploy with `scripts/ops-log`.
|
||
|
||
**To move the model forward**, change `REVISION` and `INFERENCE_PY_SHA256` in `config.py`.
|
||
Re-run the bench's sets through the service (`services/intern-decision-serve/acceptance/`), and
|
||
re-measure the VRAM before changing the cap.
|
||
|
||
## Rollback
|
||
|
||
`semif` is **stopped, not removed**, and is kept as the rollback. Only on Prime's word:
|
||
|
||
```bash
|
||
cd /opt/docker/compose/intern-decision && docker compose stop # first: the two do not both fit GPU 1
|
||
cd /opt/docker/compose/semif && docker compose start
|
||
```
|
||
|
||
Then check GPU 1's free memory against scriberr's peak. Callers go back to `:8032` and
|
||
`secret get semif/api-token`.
|