Files
esh-pfi-infrastructure/stacks/intern-decision/README.md
T
vh 750675e391 feat(intern-decision): cap 9.0 GiB with MAX_TOKENS 7168, the largest call measured to fit
Both are required in compose because they are coupled: MAX_TOKENS is checked before the forward
pass, so an oversized call is a clear 422 instead of reaching the cap as a 503. Pre-deploy floor
is nvidia-smi Free >= 15,400 MiB on GPU 1 (card peak 9,876 + scriberr 5,496).
2026-09-30 09:40:08 -07:00

221 lines
14 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# intern-decision
> ⚠ **STATUS 2026-09-30 0937 PT: built and accepted on GPU 3; the GPU 1 deploy is HELD.** GPU 1's
> usable free memory is 15,442 MiB (nvidia-smi `Free`: 97,887 total − 640 driver-reserved −
> 81,806 used), below the 15,800 MiB pre-deploy floor. The 10,300 MiB budget was derived from
> total − used (16,081). The decision is with infra-ops/Prime; see "VRAM". Nothing is running
> on port 8033 yet.
**Intern-Decision-4B typed decisions** on **fv-ml1 GPU 1**, the utility card beside `vllm-coder`,
the erp/meromero seats and scriberr. It **replaced `semif`** on Prime's ruling of 2026-09-30,
~0510 PT: "replace semif with intern-decision now". The pick came from the previous night's bench,
`docs/pfi/jev-candidates-bench-2026-09-30.md`: the only candidate at least as accurate as SemIf on
our sets, faster at our shape, and inside the memory budget.
[Intern-Decision-4B](https://huggingface.co/internlm/Intern-Decision-4B) (Shanghai AI Lab,
Apache-2.0) is a Qwen3.5-4B finetune. It takes a state plus up to 16 named questions and scores
every question's options in **one forward pass with no decoding**. It reads each answer from the
logits just before a `<decision>` marker. `services/intern-decision-serve/` loads it once and
scores through the checkpoint's **own** `inference.py` (`DecisionEngine.predict`). It keeps
**semif-serve's HTTP surface**, so a semif caller changes only the URL and the token. The contract
is `services/intern-decision-serve/intern-decision-serve.contract.md`.
| | |
|---|---|
| **URL** | `http://10.251.50.54:8033` = `http://intern-decision.fv.internal:8033`. `/health` is open; POSTs need `Authorization: Bearer $(secret get intern-decision/api-token)` |
| **Model** | `internlm/Intern-Decision-4B` @ `0e5e6aa7d6d750e2b1504ba11a8136cb58aeb3cd`, BF16, from `/tank/aimodels/huggingface`, mounted read-only and offline |
| **Scoring code** | the snapshot's `inference.py`, sha256 `c904e2c6…43b29863`, checked before it is imported. Startup refuses any other file |
| **Stack** | torch `2.10.0+cu128`, transformers `5.17.0`, fla `0.5.2`, causal-conv1d `1.7.0`: the bench's stack |
| **Image** | `intern-decision-serve:<version>`, built on fv-ml1 from `services/intern-decision-serve/` |
| **State** | None. The service never downloads (`HF_HUB_OFFLINE=1`, read-only mount). If the weights are ever lost, re-pull the pinned revision by hand. |
| **Rollback** | `stacks/semif` is stopped, not removed. See **Rollback** below. |
## API (semif-serve's)
```bash
T=$(secret get intern-decision/api-token)
curl -s -H "Authorization: Bearer $T" http://intern-decision.fv.internal:8033/decide -d '{
"id": "q1", "state": "Health checks passed in all three zones.",
"question": "Did the deployment succeed?",
"options": [{"id":"yes","description":"It succeeded."},{"id":"no","description":"It failed."}]}'
```
- `POST /decide` takes one decision, `{id, state, question, options[2..16]}`.
- `POST /decide/shared` takes `{state, decisions: [{id, question, options}]}`.
- `GET /health` reports the pins, the limits and the chunking rule.
- **Order averaging:** `"orderings": "rotations"`, or `"all"` for ≤ 4 options, is still accepted,
with semif's `combined` block. **It is barely needed:** the bench found Intern-Decision reorders
8 of 144 labels when the options are reversed, against SemIf's 30, and at one ordering it matched
SemIf-with-rotations.
- Each result carries semif's fields where they mean something: `option_ids`, `probabilities`,
`input_tokens`, `prompt_sha256`, `prompt_version`, `model`, `readout`, `probability_status`, and
on `/decide` also `total_seconds` and `forward_seconds`. New fields:
- `top`;
- `confidence`, the model's own;
- `calibration`, the model's own temperature (T = 1.99241824);
- `native`, the model's answer object, unchanged;
- `call` (`index`, `field`, `questions`): which prompt the decision was asked in.
- Past `MAX_QUEUE` (32) requests in progress, new POSTs get `429 busy` before their body is read.
### What a semif caller must know (every deliberate difference is in the contract, "Deltas")
1. **Option ids are shown to the model.** The prompt prints `A = <id>: <description>`, so an
id is part of the question. Use meaningful or neutral ids, not misleading ones.
2. **Decisions in one `/decide/shared` call are asked together, in one prompt.** An answer can
depend on the other questions in its call.
- The bench measured this on Wyrd: 79/84 asked one decision at a time, 77/84 with a turn's
4 decisions in one prompt.
- A call holds **at most 16 questions**. More are split greedily in request order: 1–16,
17–32, … . `/health` says so.
3. **`probabilities` are temperature-scaled** by the model's shipped calibration. SemIf's were
raw. The argmax is the same either way. This calibration is the vendor's, fitted on its own
data, not ours.
4. **Gone:**
- `option_logits` (the model's runtime does not expose logits);
- the prefix-cache timing fields;
- `/health.semif_commit`;
- per-workload calibration. Any `workload` is a 422, exactly as the deployed semif behaved
with its empty table.
5. **`MAX_TOKENS` (8192) is per call:** the state plus all its questions. A longer call is a 422
and is never truncated. Every call up to that limit fits the VRAM cap (measured). See VRAM
below.
## VRAM: the whole container ≤ 10,300 MiB
**Budget (infra-ops, 2026-09-30):** the container's **whole** nvidia-smi footprint, CUDA context
included, must never exceed **10,300 MiB**, whatever the request. That leaves scriberr (peak
5,496 MiB with its 120 s slices) 295 MiB of spare even when both hit their peaks at once.
`torch.cuda.set_per_process_memory_fraction` caps only torch's allocator, so the cap is the budget
minus the measured non-allocator overhead:
```
footprint <= VRAM_CAP_GIB + overhead = 9,472 MiB (9.25 GiB) + 662 MiB = 10,134 MiB (166 MiB under budget)
```
`VRAM_CAP_GIB` in `.env` is **the single knob**. It is applied before the weights load.
⚠ **The budget assumed 16,081 MiB free on GPU 1. That is total − used.** nvidia-smi's own `Free` is
**15,442 MiB**, because the driver reserves 640 MiB on every card (GPU 3 shows the same 640).
Against the true free memory, a footprint that never collides with scriberr's 5,496 MiB peak must
stay at or below **9,946 MiB**:
| cap | footprint peak (measured, GPU 3) | with scriberr at its peak, against 15,442 free | calls that fit (16 questions × 16 options) |
|---|---|---|---|
| **9.25 GiB** (`.env.example` now) | 10,134 MiB | **188 MiB over** | every call the API accepts (8,191 tok) |
| **9.0 GiB** | 9,876 MiB | 70 MiB spare | up to 7,168 tok; longer calls get 503. Every bench shape fits, at the same latency |
| measured on fv-ml1 GPU 3 (alone on the card), image 0.1.0 | MiB |
|---|---|
| **at rest** (nvidia-smi, whole card less 2 MiB idle) | **8,820** (torch reserved 8,160 + outside 660) |
| outside the allocator (CUDA context + loaded kernels), at rest / after 2×30 concurrent requests / after the latency shapes / after 4 × 8,191-token calls | 660 / 660 / 662 / 662 |
| **peak, largest request the API accepts, capped at 9.25 GiB** (card, 0.1 s sampling) | **10,134**; torch `max_reserved` = 9.25 GiB exactly, so the cap is what held it |
| the same request **uncapped**, for comparison | 10,422 (allocator peak 9,760): **over budget, so the cap is required** |
| startup peak (weights + vision tower until the swap + warm-up) | allocator 8.86 GiB. **A cap below ~8.9 GiB cannot start** (it fails closed; 8.5 was refused) |
- **The largest request the API accepts** is 64 decisions × 16 options, with each of its 4 calls
at 8,191 tokens. Under the cap it answers **200** (1.67 s), not 503. Near the limit the
allocator frees cached blocks before it fails, so the uncapped 9,760 MiB peak was partly cache.
Every call up to `MAX_TOKENS` fits the cap.
- **Over the cap the answer is 503 `out_of_memory`.** Memory goes back to the resting baseline
and the service keeps answering. This was proven with a deliberately tight cap (acceptance).
- **One inference thread is load-bearing** (contract INV-2). torch keeps CUDA state per host
thread (cuBLAS handles and workspaces). The first build scored on anyio's 40-thread pool: a
concurrent burst added 252 MiB outside the cap and 326 MiB inside it, and the peak reached
10,392 MiB. Load, warm-up and every call now run on one thread.
- **Released after bursts:** when a call leaves reserved memory more than 512 MiB over the resting
baseline, the engine calls `empty_cache()`. Measured: back to 7.97 GiB after every large
request.
- The rest footprint is 0.9 GB below the bench's 9,736 MiB. The service takes no images, so the
vision tower (0.62 GiB) is swapped for a stub after the first warm-up, and the warm-up cache is
released. Startup proves the swap changes nothing: the warm-up answer must be bit-identical
(INV-7).
## Latency
Loopback on fv-ml1 (GPU 3 alone on the card, image 0.1.0, cap 9.25 GiB). `bench_shape.py`
from the Jev bench ran 3 runs × 20 requests, and the medians are shown with the run-median
range. The bench's native Intern-Decision numbers are the reference.
| shape | end to end | server | bench (native, loopback) | SemIf 0.1.4 (bench) |
|---|---|---|---|---|
| 1 decision, short (~180 tok) | 36.4 ms (36.4–36.5) | 34.7 | 39 | 37 |
| **21 binary criteria**, one `/decide/shared` (2 calls: 16 + 5) | **82.3 ms** (82.3–82.4) | 78.9 | 88 | 131 |
| 1 decision over the ~3,900-token state | 191.4 ms (191.4–191.5) | 187.9 | 190 | 186 |
| **16 criteria over the ~3,900-token state** (1 call, 4,579 tok) | **216.8 ms** (216.5–217.0) | 211.1 | 215 | 505 |
| largest accepted request (4 calls × 8,191 tok) | 1,672 ms (1,672–1,675, N=3) | | | |
*From nh3-dev and on GPU 1, next to the vLLM seats: pending the GPU 1 deploy (held, see the status banner).*
## Acceptance (2026-09-30)
Raw results: `services/intern-decision-serve/acceptance/gpu3-2026-09-30/` (`compare.json`, one
directory per process lifetime). **The harness is the bench's own.** `bench_sets.py
--backend semif` speaks semif-serve's API, so it drove this service unchanged. `compare.py`
scores it as the bench's `analyze.py` did.
| check (GPU 3, cap 9.25 GiB, N = 3 fresh processes unless noted) | result | bench / floor |
|---|---|---|
| **positive control**, pooled 259, single ordering | **240, 240, 240** | bench 240 (4 repeats); the pooled set resolves ±4 pts |
| positive control, Wyrd /84, single | **79, 79, 79** | bench 79 |
| row by row against the bench's native rows (560 rows, single / rotations / negative) | **0 top changes, max Δp 0.000** in every repeat: bit-identical | bench floor: 0 labels moved across 4 restarts |
| rotations, pooled / Wyrd | 236 / 77 (×3) | bench 236 / 77 |
| **negative control** (descriptions rotated, 144): same top / follows the description / right vs the original gold | **10 / 122 / 14** (×3) | bench 10 / 122 / 14 |
| A-vs-A in the process (authored144 twice) | 0/144 flips, Δp 0.0 (×3) | |
| A-vs-A across restarts (r1~r2, r1~r3, r2~r3; 560 rows, single and rotations) | 0 flips, Δp 0.0 | |
| Wyrd, one `/decide/shared` per turn (4 decisions in one prompt) | 78/84 (×3); 1/84 row differs from the bench's 77 | expected: field names are positional here, decision names in the bench (contract) |
| more than 16 questions: 20 decisions in one request, against the same decisions asked as 16 + 4 | 2 calls `[16, 4]`, all 20 tops equal, Δp 0, same prompt hashes | |
| 401: no token / wrong token, both POSTs; `/health` open | 401 / 401; 200 | |
| 429: 48 concurrent requests, `MAX_QUEUE` 32 | 32 answered, 16 × 429 `busy` | |
| largest request the API accepts (64 × 16 options, 4 × 8,191 tok) at cap 9.25 GiB | **200** ×3; card peak 10,134 MiB | budget 10,300 |
| **over the cap** (the same request at a deliberately tight cap of 8.9 GiB) | **503 `out_of_memory`**; reserved 7.969 GiB before, 7.969 after, next request 200; then all 560 single rows answered, bit-identical to the bench | |
| fit under the cap: the largest 16-question call that answers 200 | 8,191 tok (the whole API) at 9.25 GiB; 7,168 tok at 9.0 GiB | |
| cap below startup | 8.5 GiB: startup refused (OOM loading), fails closed | |
*On the deployed service (GPU 1): pending the deploy (held, see the status banner).*
## Building and deploying
```bash
# from nh3-dev. /opt/docker/src is root-owned, so create the version dir with sudo first.
ssh infra-ops@10.251.50.54 'sudo -n install -d -o infra-ops -g infra-ops /opt/docker/src/intern-decision-serve-X.Y.Z | cat'
tar -C services/intern-decision-serve -cf - --exclude=.venv --exclude=.pytest_cache --exclude=__pycache__ \
--exclude=acceptance --exclude='*.egg-info' . \
| ssh infra-ops@10.251.50.54 'tar -x -C /opt/docker/src/intern-decision-serve-X.Y.Z'
# on fv-ml1
cd /opt/docker/src/intern-decision-serve-X.Y.Z && docker build -t intern-decision-serve:X.Y.Z .
# from nh3-dev: compose.yaml, .env.example and this README (never .env)
scripts/deploy-stack.sh fv-ml1 intern-decision
# on fv-ml1: .env is infra-ops's, mode 0600. Rewrite it IN PLACE to keep the owner and mode.
cd /opt/docker/compose/intern-decision && new=$(sed 's/^IMAGE=.*/IMAGE=intern-decision-serve:X.Y.Z/' .env) \
&& printf '%s\n' "$new" > .env && docker compose config -q && docker compose up -d
```
**Before any deploy onto GPU 1**, check two things. If either fails, stop; do not squeeze
scriberr.
- nvidia-smi's own `Free` on GPU 1 must be at least **15,400 MiB**:
`nvidia-smi -i 1 --query-gpu=memory.free --format=csv`. That is our card peak of 9,876 MiB plus
scriberr's 5,496, rounded up. Do not use total − used, which misses the driver's 640 MiB
reserve.
- Scriberr must not be running a job. This command must print 0:
`docker logs --since 2m scriberr | grep -c "Processing single-track job"`.
Startup fails closed. A container that never reaches healthy did not pass its own checks: the
`inference.py` hash, the pinned snapshot, the warm-up, the text-only swap and the prompt hash.
Read `docker logs intern-decision`. Record the deploy with `scripts/ops-log`.
**To move the model forward**, change `REVISION` and `INFERENCE_PY_SHA256` in `config.py`.
Re-run the bench's sets through the service (`services/intern-decision-serve/acceptance/`), and
re-measure the VRAM before changing the cap.
## Rollback
`semif` is **stopped, not removed**, and is kept as the rollback. Only on Prime's word:
```bash
cd /opt/docker/compose/intern-decision && docker compose stop # first: the two do not both fit GPU 1
cd /opt/docker/compose/semif && docker compose start
```
Then check GPU 1's free memory against scriberr's peak. Callers go back to `:8032` and
`secret get semif/api-token`.