docs(intern-decision): live on fv-ml1 GPU 1; semif marked REPLACED

intern-decision deployed 0941 PT (cap 9.0 GiB, MAX_TOKENS 7168). Live acceptance: positive
control 240/259 and Wyrd 79/84, bit-identical to the bench (0/560 rows, Δp 0); negative control
10/122/14; largest accepted requests 200 with no 503; per-process 8,812 MiB at rest and 9,866 peak;
GPU 1 Free 15,442 before and 6,581 after (lowest 5,569 under load). Latency from nh3-dev:
21 criteria 114 ms, 16 over ~3,900 tokens 238 ms. semif README banner now REPLACED with the
rollback; fv-ml1 GPU 1 note updated.
This commit is contained in:
vh
2026-09-30 09:48:30 -07:00
parent 750675e391
commit 1cf763a7b1
10 changed files with 3957 additions and 58 deletions
+72 -44
View File
@@ -1,10 +1,7 @@
# intern-decision
> ⚠ **STATUS 2026-09-30 0937 PT: built and accepted on GPU 3; the GPU 1 deploy is HELD.** GPU 1's
> usable free memory is 15,442 MiB (nvidia-smi `Free`: 97,887 total − 640 driver-reserved −
> 81,806 used), below the 15,800 MiB pre-deploy floor. The 10,300 MiB budget was derived from
> total − used (16,081). The decision is with infra-ops/Prime; see "VRAM". Nothing is running
> on port 8033 yet.
> **LIVE since 2026-09-30 0941 PT** on fv-ml1 GPU 1, with cap 9.0 GiB and `MAX_TOKENS` 7,168.
> Acceptance on the live service is below. `semif` is stopped and kept as the rollback.
**Intern-Decision-4B typed decisions** on **fv-ml1 GPU 1**, the utility card beside `vllm-coder`,
the erp/meromero seats and scriberr. It **replaced `semif`** on Prime's ruling of 2026-09-30,
@@ -76,48 +73,58 @@ curl -s -H "Authorization: Bearer $T" http://intern-decision.fv.internal:8033/de
- `/health.semif_commit`;
- per-workload calibration. Any `workload` is a 422, exactly as the deployed semif behaved
with its empty table.
5. **`MAX_TOKENS` (8192) is per call:** the state plus all its questions. A longer call is a 422
and is never truncated. Every call up to that limit fits the VRAM cap (measured). See VRAM
below.
5. **`MAX_TOKENS` is 7,168 per call:** the state plus all its questions. It is lower than the
model's own 8,192 so that every call the API accepts fits the VRAM cap.
- The count is taken **before** the forward pass, so a longer call is a clear
`422 invalid_request` ("Example has N tokens, above 7168; truncation is forbidden"). It is
never truncated.
- A request with more decisions is split into more calls, and each call must fit.
## VRAM: the whole container ≤ 10,300 MiB
## VRAM: fits beside scriberr's peak, whatever the request
**Budget (infra-ops, 2026-09-30):** the container's **whole** nvidia-smi footprint, CUDA context
included, must never exceed **10,300 MiB**, whatever the request. That leaves scriberr (peak
5,496 MiB with its 120 s slices) 295 MiB of spare even when both hit their peaks at once.
`torch.cuda.set_per_process_memory_fraction` caps only torch's allocator, so the cap is the budget
minus the measured non-allocator overhead:
**Budget (infra-ops, 2026-09-30):** GPU 1 needs nvidia-smi `Free` ≥ **15,400 MiB** before this
service starts. That is our card peak of 9,876 MiB plus scriberr's peak of 5,496 MiB (with its
120 s slices), rounded up. So both can peak at the same moment. Read nvidia-smi's own `Free`, not
total − used: the driver reserves 640 MiB on every card. The deploy at 0940 read 15,442.
`torch.cuda.set_per_process_memory_fraction` caps only torch's allocator. The two knobs in `.env`
are therefore chosen together and change together:
```
footprint <= VRAM_CAP_GIB + overhead = 9,472 MiB (9.25 GiB) + 662 MiB = 10,134 MiB (166 MiB under budget)
card footprint <= VRAM_CAP_GIB 9.0 (9,216 MiB) + 660 MiB outside the allocator (measured) = 9,876 MiB
MAX_TOKENS 7168 = the largest call measured to fit under that cap, so no accepted call reaches the 503
```
`VRAM_CAP_GIB` in `.env` is **the single knob**. It is applied before the weights load.
| on the live service, GPU 1 (per-process nvidia-smi, 0.1 s sampling, 2,939 samples over 316 s of acceptance load) | MiB |
|---|---|
| **at rest** after startup | **8,812** |
| at rest after the acceptance (512 MiB release slack keeps a little cache) | 8,856 |
| **peak**: largest requests, 64 decisions × 16 options, 4 calls × 7,168 tokens, N = 3 | **9,866**; torch `max_reserved` 8.996 of 9.0 GiB |
| GPU 1 `Free` before the deploy / at rest after / lowest during the acceptance | 15,442 / 6,581 / 5,569, which stays above scriberr's 5,496 |
⚠ **The budget assumed 16,081 MiB free on GPU 1. That is total − used.** nvidia-smi's own `Free` is
**15,442 MiB**, because the driver reserves 640 MiB on every card (GPU 3 shows the same 640).
Against the true free memory, a footprint that never collides with scriberr's 5,496 MiB peak must
stay at or below **9,946 MiB**:
- ⚠ **At 7,168 tokens the allocator reaches the cap with about 4 MiB to spare.** It fits because
the allocator frees its cache and retries before it fails.
- Measured: every call up to the limit answered 200. That is 145 shared calls on the live
service, including the size search near the boundary, 3 × 4 calls at 7,168 tokens with
16 questions, and 3 at 7,168 tokens with one question. No 503 was seen.
- The allocator's cache state depends on request history. If a max-size call ever does return
503, lower `MAX_TOKENS`. Do not raise the cap.
- Raising either knob needs a fresh measurement (`acceptance/checks.py --checks fit,maxreq`) and
the budget re-checked against scriberr.
- **Over the cap the answer is 503 `out_of_memory`.** Memory returns to the resting baseline and
the service keeps answering. This was proven on GPU 3 with a deliberately tight cap.
| cap | footprint peak (measured, GPU 3) | with scriberr at its peak, against 15,442 free | calls that fit (16 questions × 16 options) |
|---|---|---|---|
| **9.25 GiB** (`.env.example` now) | 10,134 MiB | **188 MiB over** | every call the API accepts (8,191 tok) |
| **9.0 GiB** | 9,876 MiB | 70 MiB spare | up to 7,168 tok; longer calls get 503. Every bench shape fits, at the same latency |
Measured on GPU 3 before the deploy (alone on the card), which is where the knobs came from:
| measured on fv-ml1 GPU 3 (alone on the card), image 0.1.0 | MiB |
|---|---|
| **at rest** (nvidia-smi, whole card less 2 MiB idle) | **8,820** (torch reserved 8,160 + outside 660) |
| outside the allocator (CUDA context + loaded kernels), at rest / after 2×30 concurrent requests / after the latency shapes / after 4 × 8,191-token calls | 660 / 660 / 662 / 662 |
| **peak, largest request the API accepts, capped at 9.25 GiB** (card, 0.1 s sampling) | **10,134**; torch `max_reserved` = 9.25 GiB exactly, so the cap is what held it |
| the same request **uncapped**, for comparison | 10,422 (allocator peak 9,760): **over budget, so the cap is required** |
| card peak at cap 9.0 GiB (fit search and a 64 × 8,191-token request, which got 503) | **9,876** |
| peak with 4 × 8,191-token calls at cap 9.25 GiB (answered 200) / uncapped | 10,134 / 10,422 (allocator 9,760) |
| largest 16-question call that answers 200: at cap 9.0 / at 9.25 GiB | 7,168 / 8,191 tokens |
| startup peak (weights + vision tower until the swap + warm-up) | allocator 8.86 GiB. **A cap below ~8.9 GiB cannot start** (it fails closed; 8.5 was refused) |
- **The largest request the API accepts** is 64 decisions × 16 options, with each of its 4 calls
at 8,191 tokens. Under the cap it answers **200** (1.67 s), not 503. Near the limit the
allocator frees cached blocks before it fails, so the uncapped 9,760 MiB peak was partly cache.
Every call up to `MAX_TOKENS` fits the cap.
- **Over the cap the answer is 503 `out_of_memory`.** Memory goes back to the resting baseline
and the service keeps answering. This was proven with a deliberately tight cap (acceptance).
- **One inference thread is load-bearing** (contract INV-2). torch keeps CUDA state per host
thread (cuBLAS handles and workspaces). The first build scored on anyio's 40-thread pool: a
concurrent burst added 252 MiB outside the cap and 326 MiB inside it, and the peak reached
@@ -132,19 +139,22 @@ stay at or below **9,946 MiB**:
## Latency
Loopback on fv-ml1 (GPU 3 alone on the card, image 0.1.0, cap 9.25 GiB). `bench_shape.py`
from the Jev bench ran 3 runs × 20 requests, and the medians are shown with the run-median
range. The bench's native Intern-Decision numbers are the reference.
**Live, from nh3-dev** (0942 PT). The URL is `intern-decision.fv.internal:8033`, and the round trip
is 22.9 ms on average (19.5–27.7 ms over 10 pings). GPU 1 is shared with the vLLM seats and
scriberr. The harness is `bench_shape.py` from the Jev bench, 3 runs × 20 requests; each cell is
the median with the run-median range. "server" is the service's own time.
| shape | end to end | server | bench (native, loopback) | SemIf 0.1.4 (bench) |
| shape | end to end from nh3-dev | server | GPU 3 loopback before deploy | SemIf 0.1.4 (README, from nh3-dev) |
|---|---|---|---|---|
| 1 decision, short (~180 tok) | 36.4 ms (36.4–36.5) | 34.7 | 39 | 37 |
| **21 binary criteria**, one `/decide/shared` (2 calls: 16 + 5) | **82.3 ms** (82.3–82.4) | 78.9 | 88 | 131 |
| 1 decision over the ~3,900-token state | 191.4 ms (191.4–191.5) | 187.9 | 190 | 186 |
| **16 criteria over the ~3,900-token state** (1 call, 4,579 tok) | **216.8 ms** (216.5–217.0) | 211.1 | 215 | 505 |
| largest accepted request (4 calls × 8,191 tok) | 1,672 ms (1,672–1,675, N=3) | | | |
| 1 decision, short (~180 tok) | 53.9 ms (53.3–54.2) | 35.2 | 36.4 | 69 |
| **21 binary criteria**, one `/decide/shared` (2 calls: 16 + 5) | **114.3 ms** (114.1–114.3) | 80.3 | 82.3 | 159 |
| 1 decision over the ~3,900-token state | 211.2 ms (210.8–211.6) | 186.8 | 191.4 | |
| **16 criteria over the ~3,900-token state** (1 call, 4,579 tok) | **238.3 ms** (237.8–238.7) | 205.3 | 216.8 | (bench loopback: 505) |
| largest request (4 calls × 7,168 tok), N = 3 | 1,519–1,543 ms | | | |
*From nh3-dev and on GPU 1, next to the vLLM seats: pending the GPU 1 deploy (held, see the status banner).*
Server-side time on GPU 1 matches GPU 3 within a few ms, so the vLLM neighbours were not
measurably slowing it at 0942. The rest of the end-to-end time is the network (~23 ms round trip
plus the HTTP exchange).
## Acceptance (2026-09-30)
@@ -153,6 +163,8 @@ directory per process lifetime). **The harness is the bench's own.** `bench_sets
--backend semif` speaks semif-serve's API, so it drove this service unchanged. `compare.py`
scores it as the bench's `analyze.py` did.
### On GPU 3 before the deploy (the harness and the floor)
| check (GPU 3, cap 9.25 GiB, N = 3 fresh processes unless noted) | result | bench / floor |
|---|---|---|
| **positive control**, pooled 259, single ordering | **240, 240, 240** | bench 240 (4 repeats); the pooled set resolves ±4 pts |
@@ -166,12 +178,28 @@ scores it as the bench's `analyze.py` did.
| more than 16 questions: 20 decisions in one request, against the same decisions asked as 16 + 4 | 2 calls `[16, 4]`, all 20 tops equal, Δp 0, same prompt hashes | |
| 401: no token / wrong token, both POSTs; `/health` open | 401 / 401; 200 | |
| 429: 48 concurrent requests, `MAX_QUEUE` 32 | 32 answered, 16 × 429 `busy` | |
| largest request the API accepts (64 × 16 options, 4 × 8,191 tok) at cap 9.25 GiB | **200** ×3; card peak 10,134 MiB | budget 10,300 |
| largest request at `MAX_TOKENS` 8,192 (64 × 16 options, 4 × 8,191 tok), cap 9.25 GiB | **200** ×3; card peak 10,134 MiB | superseded by the 9.0 / 7,168 knobs (VRAM) |
| **over the cap** (the same request at a deliberately tight cap of 8.9 GiB) | **503 `out_of_memory`**; reserved 7.969 GiB before, 7.969 after, next request 200; then all 560 single rows answered, bit-identical to the bench | |
| fit under the cap: the largest 16-question call that answers 200 | 8,191 tok (the whole API) at 9.25 GiB; 7,168 tok at 9.0 GiB | |
| cap below startup | 8.5 GiB: startup refused (OOM loading), fails closed | |
*On the deployed service (GPU 1): pending the deploy (held, see the status banner).*
### On the live service (GPU 1, 0942–0946 PT, through `http://intern-decision.fv.internal:8033` from nh3-dev)
Raw results: `services/intern-decision-serve/acceptance/gpu1-live/`.
| check | result | reference |
|---|---|---|
| **positive control**, pooled 259, single ordering | **240** | bench 240; GPU 3 240 ×3 |
| positive control, Wyrd /84 | **79** | bench 79 |
| row by row against the bench's native rows (560 single + 144 negative) | **0 top changes, max Δp 0.000** | |
| negative control: same top / follows the description / right vs gold | **10 / 122 / 14** | bench 10 / 122 / 14 |
| 401 without or with a wrong token (both POSTs); `/health` open | 401 / 401; 200 | |
| largest accepted request: 64 decisions × 16 options, 4 calls × 7,168 tok | **200** ×3 | |
| one question at 7,168 tok | 200 ×3; one token more (7,171) → **422** "above 7168" | |
| any 503 across the whole live acceptance | **none** (server log: 128 + 145 × 200, 18 × 422, 4 × 401) | |
The live pass is N = 1 for the positive control. It is read against the GPU 3 floor: the same
image moved 0 of 560 rows across 3 restarts, and 0 against the bench's 4 restarts.
## Building and deploying
+20 -9
View File
@@ -1,14 +1,25 @@
# semif
> ⚠ **OFFLINE since 2026-09-30 0135 PT (Prime: "take semif offline for now").** It was
> stopped (`docker compose stop`, NOT removed) to give scriberr its GPU 1 headroom back:
> scriberr's Parakeet path cuts audio into 5-minute slices and needs over 6 GB, and with
> SemIf resident it had ~6.7 GB and hit CUDA OOM on a 35-minute file. Stopping it moved
> GPU 1 from 91,052 to 81,806 MiB used. The image, weights, config and token are all kept;
> `unless-stopped` keeps it down across a reboot. **Do not restart it without Prime's word.**
> To bring it back: `cd /opt/docker/compose/semif && docker compose start`, but first
> make sure scriberr has its room (the durable fix, shortening scriberr's slice length, is
> deferred by Prime to later). Everything below describes the service as it was deployed.
> ⚠ **REPLACED by intern-decision (Prime, 2026-09-30).** Prime ruled at ~0510 PT: "replace semif
> with intern-decision now". The replacement is `stacks/intern-decision` (Intern-Decision-4B, live since 0941 PT,
> `http://10.251.50.54:8033`, `intern-decision.fv.internal`). It keeps this service's HTTP surface
> (`/decide`, `/decide/shared`, `/health`), so callers only change the URL and the token
> (`secret get intern-decision/api-token`). Its README lists every deliberate difference.
>
> The `semif` container has been **stopped, not removed**, since 0135 PT that day, when it was
> taken offline to give scriberr its GPU 1 headroom back. Its image (`semif-serve:0.1.4`),
> weights, config and token are all kept, and `unless-stopped` keeps it down across a reboot.
> Port 8032 and `semif.fv.internal` stay reserved for it.
>
> **Rollback, only on Prime's word.**
> 1. Stop intern-decision first. The two services do not fit GPU 1 together next to scriberr.
> semif held 9.2 GB at rest and peaked at 12.9 GB.
> `cd /opt/docker/compose/intern-decision && docker compose stop`.
> 2. Start semif: `cd /opt/docker/compose/semif && docker compose start`.
> 3. Check nvidia-smi `Free` on GPU 1 against semif's 12.9 GB peak plus scriberr's 5.5 GB (its
> 120 s slices) before calling it done. If it does not cover both, semif collides with scriberr.
>
> Everything below describes the service as it was deployed.
**SemIf option-logit decisions** on **fv-ml1 GPU 1**, the utility card beside
`vllm-coder`, the erp/meromero seats and scriberr. Deployed 2026-09-27 at Prime's