Files
esh-pfi-infrastructure/stacks/intern-decision/README.md
T
vh 1cf763a7b1 docs(intern-decision): live on fv-ml1 GPU 1; semif marked REPLACED
intern-decision deployed 0941 PT (cap 9.0 GiB, MAX_TOKENS 7168). Live acceptance: positive
control 240/259 and Wyrd 79/84, bit-identical to the bench (0/560 rows, Δp 0); negative control
10/122/14; largest accepted requests 200 with no 503; per-process 8,812 MiB at rest and 9,866 peak;
GPU 1 Free 15,442 before and 6,581 after (lowest 5,569 under load). Latency from nh3-dev:
21 criteria 114 ms, 16 over ~3,900 tokens 238 ms. semif README banner now REPLACED with the
rollback; fv-ml1 GPU 1 note updated.
2026-09-30 09:48:30 -07:00

249 lines
16 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# intern-decision
> **LIVE since 2026-09-30 0941 PT** on fv-ml1 GPU 1, with cap 9.0 GiB and `MAX_TOKENS` 7,168.
> Acceptance on the live service is below. `semif` is stopped and kept as the rollback.
**Intern-Decision-4B typed decisions** on **fv-ml1 GPU 1**, the utility card beside `vllm-coder`,
the erp/meromero seats and scriberr. It **replaced `semif`** on Prime's ruling of 2026-09-30,
~0510 PT: "replace semif with intern-decision now". The pick came from the previous night's bench,
`docs/pfi/jev-candidates-bench-2026-09-30.md`: the only candidate at least as accurate as SemIf on
our sets, faster at our shape, and inside the memory budget.
[Intern-Decision-4B](https://huggingface.co/internlm/Intern-Decision-4B) (Shanghai AI Lab,
Apache-2.0) is a Qwen3.5-4B finetune. It takes a state plus up to 16 named questions and scores
every question's options in **one forward pass with no decoding**. It reads each answer from the
logits just before a `<decision>` marker. `services/intern-decision-serve/` loads it once and
scores through the checkpoint's **own** `inference.py` (`DecisionEngine.predict`). It keeps
**semif-serve's HTTP surface**, so a semif caller changes only the URL and the token. The contract
is `services/intern-decision-serve/intern-decision-serve.contract.md`.
| | |
|---|---|
| **URL** | `http://10.251.50.54:8033` = `http://intern-decision.fv.internal:8033`. `/health` is open; POSTs need `Authorization: Bearer $(secret get intern-decision/api-token)` |
| **Model** | `internlm/Intern-Decision-4B` @ `0e5e6aa7d6d750e2b1504ba11a8136cb58aeb3cd`, BF16, from `/tank/aimodels/huggingface`, mounted read-only and offline |
| **Scoring code** | the snapshot's `inference.py`, sha256 `c904e2c6…43b29863`, checked before it is imported. Startup refuses any other file |
| **Stack** | torch `2.10.0+cu128`, transformers `5.17.0`, fla `0.5.2`, causal-conv1d `1.7.0`: the bench's stack |
| **Image** | `intern-decision-serve:<version>`, built on fv-ml1 from `services/intern-decision-serve/` |
| **State** | None. The service never downloads (`HF_HUB_OFFLINE=1`, read-only mount). If the weights are ever lost, re-pull the pinned revision by hand. |
| **Rollback** | `stacks/semif` is stopped, not removed. See **Rollback** below. |
## API (semif-serve's)
```bash
T=$(secret get intern-decision/api-token)
curl -s -H "Authorization: Bearer $T" http://intern-decision.fv.internal:8033/decide -d '{
"id": "q1", "state": "Health checks passed in all three zones.",
"question": "Did the deployment succeed?",
"options": [{"id":"yes","description":"It succeeded."},{"id":"no","description":"It failed."}]}'
```
- `POST /decide` takes one decision, `{id, state, question, options[2..16]}`.
- `POST /decide/shared` takes `{state, decisions: [{id, question, options}]}`.
- `GET /health` reports the pins, the limits and the chunking rule.
- **Order averaging:** `"orderings": "rotations"`, or `"all"` for ≤ 4 options, is still accepted,
with semif's `combined` block. **It is barely needed:** the bench found Intern-Decision reorders
8 of 144 labels when the options are reversed, against SemIf's 30, and at one ordering it matched
SemIf-with-rotations.
- Each result carries semif's fields where they mean something: `option_ids`, `probabilities`,
`input_tokens`, `prompt_sha256`, `prompt_version`, `model`, `readout`, `probability_status`, and
on `/decide` also `total_seconds` and `forward_seconds`. New fields:
- `top`;
- `confidence`, the model's own;
- `calibration`, the model's own temperature (T = 1.99241824);
- `native`, the model's answer object, unchanged;
- `call` (`index`, `field`, `questions`): which prompt the decision was asked in.
- Past `MAX_QUEUE` (32) requests in progress, new POSTs get `429 busy` before their body is read.
### What a semif caller must know (every deliberate difference is in the contract, "Deltas")
1. **Option ids are shown to the model.** The prompt prints `A = <id>: <description>`, so an
id is part of the question. Use meaningful or neutral ids, not misleading ones.
2. **Decisions in one `/decide/shared` call are asked together, in one prompt.** An answer can
depend on the other questions in its call.
- The bench measured this on Wyrd: 79/84 asked one decision at a time, 77/84 with a turn's
4 decisions in one prompt.
- A call holds **at most 16 questions**. More are split greedily in request order: 1–16,
17–32, … . `/health` says so.
3. **`probabilities` are temperature-scaled** by the model's shipped calibration. SemIf's were
raw. The argmax is the same either way. This calibration is the vendor's, fitted on its own
data, not ours.
4. **Gone:**
- `option_logits` (the model's runtime does not expose logits);
- the prefix-cache timing fields;
- `/health.semif_commit`;
- per-workload calibration. Any `workload` is a 422, exactly as the deployed semif behaved
with its empty table.
5. **`MAX_TOKENS` is 7,168 per call:** the state plus all its questions. It is lower than the
model's own 8,192 so that every call the API accepts fits the VRAM cap.
- The count is taken **before** the forward pass, so a longer call is a clear
`422 invalid_request` ("Example has N tokens, above 7168; truncation is forbidden"). It is
never truncated.
- A request with more decisions is split into more calls, and each call must fit.
## VRAM: fits beside scriberr's peak, whatever the request
**Budget (infra-ops, 2026-09-30):** GPU 1 needs nvidia-smi `Free` ≥ **15,400 MiB** before this
service starts. That is our card peak of 9,876 MiB plus scriberr's peak of 5,496 MiB (with its
120 s slices), rounded up. So both can peak at the same moment. Read nvidia-smi's own `Free`, not
total − used: the driver reserves 640 MiB on every card. The deploy at 0940 read 15,442.
`torch.cuda.set_per_process_memory_fraction` caps only torch's allocator. The two knobs in `.env`
are therefore chosen together and change together:
```
card footprint <= VRAM_CAP_GIB 9.0 (9,216 MiB) + 660 MiB outside the allocator (measured) = 9,876 MiB
MAX_TOKENS 7168 = the largest call measured to fit under that cap, so no accepted call reaches the 503
```
| on the live service, GPU 1 (per-process nvidia-smi, 0.1 s sampling, 2,939 samples over 316 s of acceptance load) | MiB |
|---|---|
| **at rest** after startup | **8,812** |
| at rest after the acceptance (512 MiB release slack keeps a little cache) | 8,856 |
| **peak**: largest requests, 64 decisions × 16 options, 4 calls × 7,168 tokens, N = 3 | **9,866**; torch `max_reserved` 8.996 of 9.0 GiB |
| GPU 1 `Free` before the deploy / at rest after / lowest during the acceptance | 15,442 / 6,581 / 5,569, which stays above scriberr's 5,496 |
- ⚠ **At 7,168 tokens the allocator reaches the cap with about 4 MiB to spare.** It fits because
the allocator frees its cache and retries before it fails.
- Measured: every call up to the limit answered 200. That is 145 shared calls on the live
service, including the size search near the boundary, 3 × 4 calls at 7,168 tokens with
16 questions, and 3 at 7,168 tokens with one question. No 503 was seen.
- The allocator's cache state depends on request history. If a max-size call ever does return
503, lower `MAX_TOKENS`. Do not raise the cap.
- Raising either knob needs a fresh measurement (`acceptance/checks.py --checks fit,maxreq`) and
the budget re-checked against scriberr.
- **Over the cap the answer is 503 `out_of_memory`.** Memory returns to the resting baseline and
the service keeps answering. This was proven on GPU 3 with a deliberately tight cap.
Measured on GPU 3 before the deploy (alone on the card), which is where the knobs came from:
| measured on fv-ml1 GPU 3 (alone on the card), image 0.1.0 | MiB |
|---|---|
| **at rest** (nvidia-smi, whole card less 2 MiB idle) | **8,820** (torch reserved 8,160 + outside 660) |
| outside the allocator (CUDA context + loaded kernels), at rest / after 2×30 concurrent requests / after the latency shapes / after 4 × 8,191-token calls | 660 / 660 / 662 / 662 |
| card peak at cap 9.0 GiB (fit search and a 64 × 8,191-token request, which got 503) | **9,876** |
| peak with 4 × 8,191-token calls at cap 9.25 GiB (answered 200) / uncapped | 10,134 / 10,422 (allocator 9,760) |
| largest 16-question call that answers 200: at cap 9.0 / at 9.25 GiB | 7,168 / 8,191 tokens |
| startup peak (weights + vision tower until the swap + warm-up) | allocator 8.86 GiB. **A cap below ~8.9 GiB cannot start** (it fails closed; 8.5 was refused) |
- **One inference thread is load-bearing** (contract INV-2). torch keeps CUDA state per host
thread (cuBLAS handles and workspaces). The first build scored on anyio's 40-thread pool: a
concurrent burst added 252 MiB outside the cap and 326 MiB inside it, and the peak reached
10,392 MiB. Load, warm-up and every call now run on one thread.
- **Released after bursts:** when a call leaves reserved memory more than 512 MiB over the resting
baseline, the engine calls `empty_cache()`. Measured: back to 7.97 GiB after every large
request.
- The rest footprint is 0.9 GB below the bench's 9,736 MiB. The service takes no images, so the
vision tower (0.62 GiB) is swapped for a stub after the first warm-up, and the warm-up cache is
released. Startup proves the swap changes nothing: the warm-up answer must be bit-identical
(INV-7).
## Latency
**Live, from nh3-dev** (0942 PT). The URL is `intern-decision.fv.internal:8033`, and the round trip
is 22.9 ms on average (19.5–27.7 ms over 10 pings). GPU 1 is shared with the vLLM seats and
scriberr. The harness is `bench_shape.py` from the Jev bench, 3 runs × 20 requests; each cell is
the median with the run-median range. "server" is the service's own time.
| shape | end to end from nh3-dev | server | GPU 3 loopback before deploy | SemIf 0.1.4 (README, from nh3-dev) |
|---|---|---|---|---|
| 1 decision, short (~180 tok) | 53.9 ms (53.3–54.2) | 35.2 | 36.4 | 69 |
| **21 binary criteria**, one `/decide/shared` (2 calls: 16 + 5) | **114.3 ms** (114.1–114.3) | 80.3 | 82.3 | 159 |
| 1 decision over the ~3,900-token state | 211.2 ms (210.8–211.6) | 186.8 | 191.4 | |
| **16 criteria over the ~3,900-token state** (1 call, 4,579 tok) | **238.3 ms** (237.8–238.7) | 205.3 | 216.8 | (bench loopback: 505) |
| largest request (4 calls × 7,168 tok), N = 3 | 1,519–1,543 ms | | | |
Server-side time on GPU 1 matches GPU 3 within a few ms, so the vLLM neighbours were not
measurably slowing it at 0942. The rest of the end-to-end time is the network (~23 ms round trip
plus the HTTP exchange).
## Acceptance (2026-09-30)
Raw results: `services/intern-decision-serve/acceptance/gpu3-2026-09-30/` (`compare.json`, one
directory per process lifetime). **The harness is the bench's own.** `bench_sets.py
--backend semif` speaks semif-serve's API, so it drove this service unchanged. `compare.py`
scores it as the bench's `analyze.py` did.
### On GPU 3 before the deploy (the harness and the floor)
| check (GPU 3, cap 9.25 GiB, N = 3 fresh processes unless noted) | result | bench / floor |
|---|---|---|
| **positive control**, pooled 259, single ordering | **240, 240, 240** | bench 240 (4 repeats); the pooled set resolves ±4 pts |
| positive control, Wyrd /84, single | **79, 79, 79** | bench 79 |
| row by row against the bench's native rows (560 rows, single / rotations / negative) | **0 top changes, max Δp 0.000** in every repeat: bit-identical | bench floor: 0 labels moved across 4 restarts |
| rotations, pooled / Wyrd | 236 / 77 (×3) | bench 236 / 77 |
| **negative control** (descriptions rotated, 144): same top / follows the description / right vs the original gold | **10 / 122 / 14** (×3) | bench 10 / 122 / 14 |
| A-vs-A in the process (authored144 twice) | 0/144 flips, Δp 0.0 (×3) | |
| A-vs-A across restarts (r1~r2, r1~r3, r2~r3; 560 rows, single and rotations) | 0 flips, Δp 0.0 | |
| Wyrd, one `/decide/shared` per turn (4 decisions in one prompt) | 78/84 (×3); 1/84 row differs from the bench's 77 | expected: field names are positional here, decision names in the bench (contract) |
| more than 16 questions: 20 decisions in one request, against the same decisions asked as 16 + 4 | 2 calls `[16, 4]`, all 20 tops equal, Δp 0, same prompt hashes | |
| 401: no token / wrong token, both POSTs; `/health` open | 401 / 401; 200 | |
| 429: 48 concurrent requests, `MAX_QUEUE` 32 | 32 answered, 16 × 429 `busy` | |
| largest request at `MAX_TOKENS` 8,192 (64 × 16 options, 4 × 8,191 tok), cap 9.25 GiB | **200** ×3; card peak 10,134 MiB | superseded by the 9.0 / 7,168 knobs (VRAM) |
| **over the cap** (the same request at a deliberately tight cap of 8.9 GiB) | **503 `out_of_memory`**; reserved 7.969 GiB before, 7.969 after, next request 200; then all 560 single rows answered, bit-identical to the bench | |
| fit under the cap: the largest 16-question call that answers 200 | 8,191 tok (the whole API) at 9.25 GiB; 7,168 tok at 9.0 GiB | |
| cap below startup | 8.5 GiB: startup refused (OOM loading), fails closed | |
### On the live service (GPU 1, 0942–0946 PT, through `http://intern-decision.fv.internal:8033` from nh3-dev)
Raw results: `services/intern-decision-serve/acceptance/gpu1-live/`.
| check | result | reference |
|---|---|---|
| **positive control**, pooled 259, single ordering | **240** | bench 240; GPU 3 240 ×3 |
| positive control, Wyrd /84 | **79** | bench 79 |
| row by row against the bench's native rows (560 single + 144 negative) | **0 top changes, max Δp 0.000** | |
| negative control: same top / follows the description / right vs gold | **10 / 122 / 14** | bench 10 / 122 / 14 |
| 401 without or with a wrong token (both POSTs); `/health` open | 401 / 401; 200 | |
| largest accepted request: 64 decisions × 16 options, 4 calls × 7,168 tok | **200** ×3 | |
| one question at 7,168 tok | 200 ×3; one token more (7,171) → **422** "above 7168" | |
| any 503 across the whole live acceptance | **none** (server log: 128 + 145 × 200, 18 × 422, 4 × 401) | |
The live pass is N = 1 for the positive control. It is read against the GPU 3 floor: the same
image moved 0 of 560 rows across 3 restarts, and 0 against the bench's 4 restarts.
## Building and deploying
```bash
# from nh3-dev. /opt/docker/src is root-owned, so create the version dir with sudo first.
ssh infra-ops@10.251.50.54 'sudo -n install -d -o infra-ops -g infra-ops /opt/docker/src/intern-decision-serve-X.Y.Z | cat'
tar -C services/intern-decision-serve -cf - --exclude=.venv --exclude=.pytest_cache --exclude=__pycache__ \
--exclude=acceptance --exclude='*.egg-info' . \
| ssh infra-ops@10.251.50.54 'tar -x -C /opt/docker/src/intern-decision-serve-X.Y.Z'
# on fv-ml1
cd /opt/docker/src/intern-decision-serve-X.Y.Z && docker build -t intern-decision-serve:X.Y.Z .
# from nh3-dev: compose.yaml, .env.example and this README (never .env)
scripts/deploy-stack.sh fv-ml1 intern-decision
# on fv-ml1: .env is infra-ops's, mode 0600. Rewrite it IN PLACE to keep the owner and mode.
cd /opt/docker/compose/intern-decision && new=$(sed 's/^IMAGE=.*/IMAGE=intern-decision-serve:X.Y.Z/' .env) \
&& printf '%s\n' "$new" > .env && docker compose config -q && docker compose up -d
```
**Before any deploy onto GPU 1**, check two things. If either fails, stop; do not squeeze
scriberr.
- nvidia-smi's own `Free` on GPU 1 must be at least **15,400 MiB**:
`nvidia-smi -i 1 --query-gpu=memory.free --format=csv`. That is our card peak of 9,876 MiB plus
scriberr's 5,496, rounded up. Do not use total − used, which misses the driver's 640 MiB
reserve.
- Scriberr must not be running a job. This command must print 0:
`docker logs --since 2m scriberr | grep -c "Processing single-track job"`.
Startup fails closed. A container that never reaches healthy did not pass its own checks: the
`inference.py` hash, the pinned snapshot, the warm-up, the text-only swap and the prompt hash.
Read `docker logs intern-decision`. Record the deploy with `scripts/ops-log`.
**To move the model forward**, change `REVISION` and `INFERENCE_PY_SHA256` in `config.py`.
Re-run the bench's sets through the service (`services/intern-decision-serve/acceptance/`), and
re-measure the VRAM before changing the cap.
## Rollback
`semif` is **stopped, not removed**, and is kept as the rollback. Only on Prime's word:
```bash
cd /opt/docker/compose/intern-decision && docker compose stop # first: the two do not both fit GPU 1
cd /opt/docker/compose/semif && docker compose start
```
Then check GPU 1's free memory against scriberr's peak. Callers go back to `:8032` and
`secret get semif/api-token`.