docs(intern-decision): live on fv-ml1 GPU 1; semif marked REPLACED

intern-decision deployed 0941 PT (cap 9.0 GiB, MAX_TOKENS 7168). Live acceptance: positive
control 240/259 and Wyrd 79/84, bit-identical to the bench (0/560 rows, Δp 0); negative control
10/122/14; largest accepted requests 200 with no 503; per-process 8,812 MiB at rest and 9,866 peak;
GPU 1 Free 15,442 before and 6,581 after (lowest 5,569 under load). Latency from nh3-dev:
21 criteria 114 ms, 16 over ~3,900 tokens 238 ms. semif README banner now REPLACED with the
rollback; fv-ml1 GPU 1 note updated.
This commit is contained in:
vh
2026-09-30 09:48:30 -07:00
parent 750675e391
commit 1cf763a7b1
10 changed files with 3957 additions and 58 deletions
+2 -2
View File
@@ -127,5 +127,5 @@ aliases:
- {name: wherethef, site: nh3, target: nh3-dev, note: WhereTF :8093}
- {name: homepage, site: esh, target: esh-docker-vm, note: fleet dashboard :5100}
- {name: scriberr, site: fv, target: fv-ml1, note: transcription + diarization :8080 (GPU1)}
- {name: semif, site: fv, target: fv-ml1, note: SemIf option-logit decisions :8032 (GPU1) — stopped since 2026-09-30, being replaced by intern-decision (Prime); kept as rollback}
- {name: intern-decision, site: fv, target: fv-ml1, note: Intern-Decision-4B typed decisions :8033 (GPU1), semif-serve API; replaces semif (Prime 2026-09-30)}
- {name: semif, site: fv, target: fv-ml1, note: SemIf option-logit decisions :8032 (GPU1) — REPLACED by intern-decision 2026-09-30 (Prime), stopped, kept as rollback}
- {name: intern-decision, site: fv, target: fv-ml1, note: Intern-Decision-4B typed decisions :8033 (GPU1), semif-serve API; replaced semif 2026-09-30 (Prime)}
+9 -3
View File
@@ -181,9 +181,15 @@ embed/rerank/reward trio. GPUs are pinned per container via
**GPU 1 — light / eval / retrieval + char-RP GGUF (~91/98 GB, on-demand):**
> ⚠ **This table is stale (checked 2026-09-27).** Live GPU 1 residents were `scriberr`,
> `vllm-coder`, `vllm-erp-seat` and `vllm-meromero-rp`, plus **`semif`** (**OFFLINE since 2026-09-30, Prime; see `stacks/semif`**; :8032, ~8.7 GB resting when up,
> hard-capped at 12 GiB; `stacks/semif`, since 2026-09-27). Read the host
> ⚠ **This table is stale (checked 2026-09-27).** Live GPU 1 residents are `scriberr`,
> `vllm-coder`, `vllm-erp-seat` and `vllm-meromero-rp`.
> **`intern-decision`** joined them on 2026-09-30, 0941 PT: :8033, 8,812 MiB at rest, 9,866 MiB
> peak, hard-capped at 9.0 GiB with `MAX_TOKENS` 7,168; see `stacks/intern-decision`. It replaced
> **`semif`** (:8032), which is stopped and kept as the rollback, per Prime's ruling of 2026-09-30.
> **GPU 1 budget:** nvidia-smi `Free` read 15,442 MiB before intern-decision and 6,581 MiB after,
> at rest. That covers scriberr's 5,496 MiB peak even while intern-decision is at its own peak.
> Measure nvidia-smi `Free` (the driver reserves 640 MiB per card) before adding anything to
> this card. Read the host
> (`docker inspect … DeviceRequests`), not this table. (A fixtures-only `augaman` instance ran
> here for about an hour on 2026-09-27 for a speed bench, and was then removed on Prime's call.)
@@ -0,0 +1,139 @@
{
"url": "http://intern-decision.fv.internal:8033",
"started_utc": "2026-09-30T16:43:39Z",
"health": {
"status": "ok",
"model": {
"name": "Intern-Decision-4B",
"source": "internlm/Intern-Decision-4B",
"revision": "0e5e6aa7d6d750e2b1504ba11a8136cb58aeb3cd",
"checkpoint": "/hf/hub/models--internlm--Intern-Decision-4B/snapshots/0e5e6aa7d6d750e2b1504ba11a8136cb58aeb3cd",
"inference_py_sha256": "c904e2c67ca0775621a22375ee373d2ba30b52117cda870c6c9ef74143b29863",
"temperature": 1.99241824,
"dtype": "bfloat16",
"attn_implementation": "sdpa",
"device": "cuda",
"max_length": 7168,
"torch_version": "2.10.0+cu128",
"transformers_version": "5.17.0",
"vision_tower": "removed",
"device_name": "NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition",
"allocated_gib": 7.937,
"reserved_gib": 7.969,
"max_reserved_gib": 8.861
},
"vram_cap_gib": 9.0,
"max_tokens": 7168,
"max_decisions": 64,
"max_questions_per_call": 16,
"chunking": "/decide/shared questions are packed greedily, in request order, into calls of at most 16 (1-16, 17-32, ...); each call is one prompt, so the questions in a call are asked together. With orderings, ordering k of every decision forms wave k, packed the same way.",
"workloads": []
},
"auth": {
"/decide": {
"no_token": 401,
"wrong_token": 401,
"right_token": 200
},
"/decide/shared": {
"no_token": 401,
"wrong_token": 401,
"right_token": 200
},
"health_no_token": 200,
"pass": true,
"t_start": 1790786619.1466322,
"t_end": 1790786619.3572643
},
"maxreq": {
"state_chars": 4793,
"tokens_per_call": 7168,
"decisions": 64,
"options_each": 16,
"body_bytes": 96250,
"runs": [
{
"status": 200,
"e2e_ms": 1542.5,
"code": null,
"calls": 4,
"input_tokens": [
7168,
7168,
7168,
7168
],
"reserved_gib_after": 7.969,
"max_reserved_gib": 8.996,
"t_end": 1790786706.9185867
},
{
"status": 200,
"e2e_ms": 1518.7,
"code": null,
"calls": 4,
"input_tokens": [
7168,
7168,
7168,
7168
],
"reserved_gib_after": 7.969,
"max_reserved_gib": 8.996,
"t_end": 1790786708.4557755
},
{
"status": 200,
"e2e_ms": 1539.6,
"code": null,
"calls": 4,
"input_tokens": [
7168,
7168,
7168,
7168
],
"reserved_gib_after": 7.969,
"max_reserved_gib": 8.996,
"t_end": 1790786710.014339
}
],
"t_start": 1790786619.357301,
"t_end": 1790786710.014397
},
"maxone": {
"state_chars": 28727,
"tokens": 7168,
"one_more_is": [
422,
"Example has 7171 tokens, above 7168; truncation is forbidden"
],
"runs": [
{
"status": 200,
"e2e_ms": 394.2,
"input_tokens": 7168,
"code": null,
"t_end": 1790786715.324555
},
{
"status": 200,
"e2e_ms": 375.9,
"input_tokens": 7168,
"code": null,
"t_end": 1790786715.7005756
},
{
"status": 200,
"e2e_ms": 372.7,
"input_tokens": 7168,
"code": null,
"t_end": 1790786716.0734687
}
],
"pass": true,
"t_start": 1790786710.0144272,
"t_end": 1790786716.0735004
},
"finished_utc": "2026-09-30T16:45:16Z"
}
@@ -0,0 +1,671 @@
{
"ours": {
"live": {
"scores": {
"single/pooled": [
240,
259
],
"single/authored144": [
132,
144
],
"single/perturbations108": [
100,
108
],
"single/cicada-w1": [
29,
31
],
"single/cicada-w2": [
22,
31
],
"single/wyrd": [
79,
84
],
"single/wyrd:place": [
19,
21
],
"single/wyrd:place2": [
21,
21
],
"single/wyrd:exit": [
19,
21
],
"single/wyrd:exit2": [
20,
21
],
"single/failures": 0
},
"negative": {
"n": 144,
"same_top_as_unrotated": 10,
"follows_description": 122,
"vs_original_gold": 14
},
"a_vs_a_in_process_authored144": {
"rows": 0,
"top_differs": 0,
"flipped_ids": [],
"max_dp": null
}
}
},
"bench": {
"r1": {
"scores": {
"single/pooled": [
240,
259
],
"single/authored144": [
132,
144
],
"single/perturbations108": [
100,
108
],
"single/cicada-w1": [
29,
31
],
"single/cicada-w2": [
22,
31
],
"single/wyrd": [
79,
84
],
"single/wyrd:place": [
19,
21
],
"single/wyrd:place2": [
21,
21
],
"single/wyrd:exit": [
19,
21
],
"single/wyrd:exit2": [
20,
21
],
"single/failures": 0,
"rotations/pooled": [
236,
259
],
"rotations/authored144": [
130,
144
],
"rotations/perturbations108": [
99,
108
],
"rotations/cicada-w1": [
29,
31
],
"rotations/cicada-w2": [
22,
31
],
"rotations/wyrd": [
77,
84
],
"rotations/wyrd:place": [
18,
21
],
"rotations/wyrd:place2": [
21,
21
],
"rotations/wyrd:exit": [
19,
21
],
"rotations/wyrd:exit2": [
19,
21
],
"rotations/failures": 0
},
"negative": {
"n": 144,
"same_top_as_unrotated": 10,
"follows_description": 122,
"vs_original_gold": 14
}
},
"r2": {
"scores": {
"single/pooled": [
240,
259
],
"single/authored144": [
132,
144
],
"single/perturbations108": [
100,
108
],
"single/cicada-w1": [
29,
31
],
"single/cicada-w2": [
22,
31
],
"single/wyrd": [
79,
84
],
"single/wyrd:place": [
19,
21
],
"single/wyrd:place2": [
21,
21
],
"single/wyrd:exit": [
19,
21
],
"single/wyrd:exit2": [
20,
21
],
"single/failures": 0,
"rotations/pooled": [
236,
259
],
"rotations/authored144": [
130,
144
],
"rotations/perturbations108": [
99,
108
],
"rotations/cicada-w1": [
29,
31
],
"rotations/cicada-w2": [
22,
31
],
"rotations/wyrd": [
77,
84
],
"rotations/wyrd:place": [
18,
21
],
"rotations/wyrd:place2": [
21,
21
],
"rotations/wyrd:exit": [
19,
21
],
"rotations/wyrd:exit2": [
19,
21
],
"rotations/failures": 0,
"multifield/pooled": [
77,
84
],
"multifield/wyrd": [
77,
84
],
"multifield/wyrd:place": [
16,
21
],
"multifield/wyrd:place2": [
21,
21
],
"multifield/wyrd:exit": [
20,
21
],
"multifield/wyrd:exit2": [
20,
21
],
"multifield/failures": 0
},
"negative": {
"n": 144,
"same_top_as_unrotated": 10,
"follows_description": 122,
"vs_original_gold": 14
}
},
"r3": {
"scores": {
"single/pooled": [
240,
259
],
"single/authored144": [
132,
144
],
"single/perturbations108": [
100,
108
],
"single/cicada-w1": [
29,
31
],
"single/cicada-w2": [
22,
31
],
"single/wyrd": [
79,
84
],
"single/wyrd:place": [
19,
21
],
"single/wyrd:place2": [
21,
21
],
"single/wyrd:exit": [
19,
21
],
"single/wyrd:exit2": [
20,
21
],
"single/failures": 0,
"rotations/pooled": [
236,
259
],
"rotations/authored144": [
130,
144
],
"rotations/perturbations108": [
99,
108
],
"rotations/cicada-w1": [
29,
31
],
"rotations/cicada-w2": [
22,
31
],
"rotations/wyrd": [
77,
84
],
"rotations/wyrd:place": [
18,
21
],
"rotations/wyrd:place2": [
21,
21
],
"rotations/wyrd:exit": [
19,
21
],
"rotations/wyrd:exit2": [
19,
21
],
"rotations/failures": 0,
"multifield/pooled": [
77,
84
],
"multifield/wyrd": [
77,
84
],
"multifield/wyrd:place": [
16,
21
],
"multifield/wyrd:place2": [
21,
21
],
"multifield/wyrd:exit": [
20,
21
],
"multifield/wyrd:exit2": [
20,
21
],
"multifield/failures": 0
},
"negative": {
"n": 144,
"same_top_as_unrotated": 10,
"follows_description": 122,
"vs_original_gold": 14
}
},
"r4": {
"scores": {
"single/pooled": [
240,
259
],
"single/authored144": [
132,
144
],
"single/perturbations108": [
100,
108
],
"single/cicada-w1": [
29,
31
],
"single/cicada-w2": [
22,
31
],
"single/wyrd": [
79,
84
],
"single/wyrd:place": [
19,
21
],
"single/wyrd:place2": [
21,
21
],
"single/wyrd:exit": [
19,
21
],
"single/wyrd:exit2": [
20,
21
],
"single/failures": 0,
"rotations/pooled": [
236,
259
],
"rotations/authored144": [
130,
144
],
"rotations/perturbations108": [
99,
108
],
"rotations/cicada-w1": [
29,
31
],
"rotations/cicada-w2": [
22,
31
],
"rotations/wyrd": [
77,
84
],
"rotations/wyrd:place": [
18,
21
],
"rotations/wyrd:place2": [
21,
21
],
"rotations/wyrd:exit": [
19,
21
],
"rotations/wyrd:exit2": [
19,
21
],
"rotations/failures": 0,
"multifield/pooled": [
77,
84
],
"multifield/wyrd": [
77,
84
],
"multifield/wyrd:place": [
16,
21
],
"multifield/wyrd:place2": [
21,
21
],
"multifield/wyrd:exit": [
20,
21
],
"multifield/wyrd:exit2": [
20,
21
],
"multifield/failures": 0
},
"negative": {
"n": 144,
"same_top_as_unrotated": 10,
"follows_description": 122,
"vs_original_gold": 14
}
}
},
"vs_bench_row_by_row": {
"live": {
"single": {
"bench_repeat": "r1",
"rows": 560,
"top_differs": 0,
"flipped_ids": [],
"max_dp": 0.0
},
"rotations": {
"bench_repeat": "r1",
"rows": 0,
"top_differs": 0,
"flipped_ids": [],
"max_dp": null
},
"negative": {
"bench_repeat": "r1",
"rows": 144,
"top_differs": 0,
"flipped_ids": [],
"max_dp": 0.0
},
"multifield": {
"bench_repeat": "r2",
"rows": 0,
"top_differs": 0,
"flipped_ids": [],
"max_dp": null
}
}
},
"across_restarts": {},
"summary": {
"single/authored144": {
"ours_median": 132,
"ours_all": [
132
],
"bench_all": [
132,
132,
132,
132
],
"n": 144
},
"single/cicada-w1": {
"ours_median": 29,
"ours_all": [
29
],
"bench_all": [
29,
29,
29,
29
],
"n": 31
},
"single/cicada-w2": {
"ours_median": 22,
"ours_all": [
22
],
"bench_all": [
22,
22,
22,
22
],
"n": 31
},
"single/perturbations108": {
"ours_median": 100,
"ours_all": [
100
],
"bench_all": [
100,
100,
100,
100
],
"n": 108
},
"single/pooled": {
"ours_median": 240,
"ours_all": [
240
],
"bench_all": [
240,
240,
240,
240
],
"n": 259
},
"single/wyrd": {
"ours_median": 79,
"ours_all": [
79
],
"bench_all": [
79,
79,
79,
79
],
"n": 84
},
"single/wyrd:exit": {
"ours_median": 19,
"ours_all": [
19
],
"bench_all": [
19,
19,
19,
19
],
"n": 21
},
"single/wyrd:exit2": {
"ours_median": 20,
"ours_all": [
20
],
"bench_all": [
20,
20,
20,
20
],
"n": 21
},
"single/wyrd:place": {
"ours_median": 19,
"ours_all": [
19
],
"bench_all": [
19,
19,
19,
19
],
"n": 21
},
"single/wyrd:place2": {
"ours_median": 21,
"ours_all": [
21
],
"bench_all": [
21,
21,
21,
21
],
"n": 21
}
}
}
@@ -0,0 +1,6 @@
1790786561.567928123 shape:start
1790786613.903648987 shape:done
1790786619.093843341 checks:start
1790786716.079809795 checks:done
1790786739.656804839 sets:start
1790786809.130379221 sets:done
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because it is too large Load Diff
+72 -44
View File
@@ -1,10 +1,7 @@
# intern-decision
> ⚠ **STATUS 2026-09-30 0937 PT: built and accepted on GPU 3; the GPU 1 deploy is HELD.** GPU 1's
> usable free memory is 15,442 MiB (nvidia-smi `Free`: 97,887 total − 640 driver-reserved −
> 81,806 used), below the 15,800 MiB pre-deploy floor. The 10,300 MiB budget was derived from
> total − used (16,081). The decision is with infra-ops/Prime; see "VRAM". Nothing is running
> on port 8033 yet.
> **LIVE since 2026-09-30 0941 PT** on fv-ml1 GPU 1, with cap 9.0 GiB and `MAX_TOKENS` 7,168.
> Acceptance on the live service is below. `semif` is stopped and kept as the rollback.
**Intern-Decision-4B typed decisions** on **fv-ml1 GPU 1**, the utility card beside `vllm-coder`,
the erp/meromero seats and scriberr. It **replaced `semif`** on Prime's ruling of 2026-09-30,
@@ -76,48 +73,58 @@ curl -s -H "Authorization: Bearer $T" http://intern-decision.fv.internal:8033/de
- `/health.semif_commit`;
- per-workload calibration. Any `workload` is a 422, exactly as the deployed semif behaved
with its empty table.
5. **`MAX_TOKENS` (8192) is per call:** the state plus all its questions. A longer call is a 422
and is never truncated. Every call up to that limit fits the VRAM cap (measured). See VRAM
below.
5. **`MAX_TOKENS` is 7,168 per call:** the state plus all its questions. It is lower than the
model's own 8,192 so that every call the API accepts fits the VRAM cap.
- The count is taken **before** the forward pass, so a longer call is a clear
`422 invalid_request` ("Example has N tokens, above 7168; truncation is forbidden"). It is
never truncated.
- A request with more decisions is split into more calls, and each call must fit.
## VRAM: the whole container ≤ 10,300 MiB
## VRAM: fits beside scriberr's peak, whatever the request
**Budget (infra-ops, 2026-09-30):** the container's **whole** nvidia-smi footprint, CUDA context
included, must never exceed **10,300 MiB**, whatever the request. That leaves scriberr (peak
5,496 MiB with its 120 s slices) 295 MiB of spare even when both hit their peaks at once.
`torch.cuda.set_per_process_memory_fraction` caps only torch's allocator, so the cap is the budget
minus the measured non-allocator overhead:
**Budget (infra-ops, 2026-09-30):** GPU 1 needs nvidia-smi `Free` ≥ **15,400 MiB** before this
service starts. That is our card peak of 9,876 MiB plus scriberr's peak of 5,496 MiB (with its
120 s slices), rounded up. So both can peak at the same moment. Read nvidia-smi's own `Free`, not
total − used: the driver reserves 640 MiB on every card. The deploy at 0940 read 15,442.
`torch.cuda.set_per_process_memory_fraction` caps only torch's allocator. The two knobs in `.env`
are therefore chosen together and change together:
```
footprint <= VRAM_CAP_GIB + overhead = 9,472 MiB (9.25 GiB) + 662 MiB = 10,134 MiB (166 MiB under budget)
card footprint <= VRAM_CAP_GIB 9.0 (9,216 MiB) + 660 MiB outside the allocator (measured) = 9,876 MiB
MAX_TOKENS 7168 = the largest call measured to fit under that cap, so no accepted call reaches the 503
```
`VRAM_CAP_GIB` in `.env` is **the single knob**. It is applied before the weights load.
| on the live service, GPU 1 (per-process nvidia-smi, 0.1 s sampling, 2,939 samples over 316 s of acceptance load) | MiB |
|---|---|
| **at rest** after startup | **8,812** |
| at rest after the acceptance (512 MiB release slack keeps a little cache) | 8,856 |
| **peak**: largest requests, 64 decisions × 16 options, 4 calls × 7,168 tokens, N = 3 | **9,866**; torch `max_reserved` 8.996 of 9.0 GiB |
| GPU 1 `Free` before the deploy / at rest after / lowest during the acceptance | 15,442 / 6,581 / 5,569, which stays above scriberr's 5,496 |
⚠ **The budget assumed 16,081 MiB free on GPU 1. That is total − used.** nvidia-smi's own `Free` is
**15,442 MiB**, because the driver reserves 640 MiB on every card (GPU 3 shows the same 640).
Against the true free memory, a footprint that never collides with scriberr's 5,496 MiB peak must
stay at or below **9,946 MiB**:
- ⚠ **At 7,168 tokens the allocator reaches the cap with about 4 MiB to spare.** It fits because
the allocator frees its cache and retries before it fails.
- Measured: every call up to the limit answered 200. That is 145 shared calls on the live
service, including the size search near the boundary, 3 × 4 calls at 7,168 tokens with
16 questions, and 3 at 7,168 tokens with one question. No 503 was seen.
- The allocator's cache state depends on request history. If a max-size call ever does return
503, lower `MAX_TOKENS`. Do not raise the cap.
- Raising either knob needs a fresh measurement (`acceptance/checks.py --checks fit,maxreq`) and
the budget re-checked against scriberr.
- **Over the cap the answer is 503 `out_of_memory`.** Memory returns to the resting baseline and
the service keeps answering. This was proven on GPU 3 with a deliberately tight cap.
| cap | footprint peak (measured, GPU 3) | with scriberr at its peak, against 15,442 free | calls that fit (16 questions × 16 options) |
|---|---|---|---|
| **9.25 GiB** (`.env.example` now) | 10,134 MiB | **188 MiB over** | every call the API accepts (8,191 tok) |
| **9.0 GiB** | 9,876 MiB | 70 MiB spare | up to 7,168 tok; longer calls get 503. Every bench shape fits, at the same latency |
Measured on GPU 3 before the deploy (alone on the card), which is where the knobs came from:
| measured on fv-ml1 GPU 3 (alone on the card), image 0.1.0 | MiB |
|---|---|
| **at rest** (nvidia-smi, whole card less 2 MiB idle) | **8,820** (torch reserved 8,160 + outside 660) |
| outside the allocator (CUDA context + loaded kernels), at rest / after 2×30 concurrent requests / after the latency shapes / after 4 × 8,191-token calls | 660 / 660 / 662 / 662 |
| **peak, largest request the API accepts, capped at 9.25 GiB** (card, 0.1 s sampling) | **10,134**; torch `max_reserved` = 9.25 GiB exactly, so the cap is what held it |
| the same request **uncapped**, for comparison | 10,422 (allocator peak 9,760): **over budget, so the cap is required** |
| card peak at cap 9.0 GiB (fit search and a 64 × 8,191-token request, which got 503) | **9,876** |
| peak with 4 × 8,191-token calls at cap 9.25 GiB (answered 200) / uncapped | 10,134 / 10,422 (allocator 9,760) |
| largest 16-question call that answers 200: at cap 9.0 / at 9.25 GiB | 7,168 / 8,191 tokens |
| startup peak (weights + vision tower until the swap + warm-up) | allocator 8.86 GiB. **A cap below ~8.9 GiB cannot start** (it fails closed; 8.5 was refused) |
- **The largest request the API accepts** is 64 decisions × 16 options, with each of its 4 calls
at 8,191 tokens. Under the cap it answers **200** (1.67 s), not 503. Near the limit the
allocator frees cached blocks before it fails, so the uncapped 9,760 MiB peak was partly cache.
Every call up to `MAX_TOKENS` fits the cap.
- **Over the cap the answer is 503 `out_of_memory`.** Memory goes back to the resting baseline
and the service keeps answering. This was proven with a deliberately tight cap (acceptance).
- **One inference thread is load-bearing** (contract INV-2). torch keeps CUDA state per host
thread (cuBLAS handles and workspaces). The first build scored on anyio's 40-thread pool: a
concurrent burst added 252 MiB outside the cap and 326 MiB inside it, and the peak reached
@@ -132,19 +139,22 @@ stay at or below **9,946 MiB**:
## Latency
Loopback on fv-ml1 (GPU 3 alone on the card, image 0.1.0, cap 9.25 GiB). `bench_shape.py`
from the Jev bench ran 3 runs × 20 requests, and the medians are shown with the run-median
range. The bench's native Intern-Decision numbers are the reference.
**Live, from nh3-dev** (0942 PT). The URL is `intern-decision.fv.internal:8033`, and the round trip
is 22.9 ms on average (19.5–27.7 ms over 10 pings). GPU 1 is shared with the vLLM seats and
scriberr. The harness is `bench_shape.py` from the Jev bench, 3 runs × 20 requests; each cell is
the median with the run-median range. "server" is the service's own time.
| shape | end to end | server | bench (native, loopback) | SemIf 0.1.4 (bench) |
| shape | end to end from nh3-dev | server | GPU 3 loopback before deploy | SemIf 0.1.4 (README, from nh3-dev) |
|---|---|---|---|---|
| 1 decision, short (~180 tok) | 36.4 ms (36.4–36.5) | 34.7 | 39 | 37 |
| **21 binary criteria**, one `/decide/shared` (2 calls: 16 + 5) | **82.3 ms** (82.3–82.4) | 78.9 | 88 | 131 |
| 1 decision over the ~3,900-token state | 191.4 ms (191.4–191.5) | 187.9 | 190 | 186 |
| **16 criteria over the ~3,900-token state** (1 call, 4,579 tok) | **216.8 ms** (216.5–217.0) | 211.1 | 215 | 505 |
| largest accepted request (4 calls × 8,191 tok) | 1,672 ms (1,672–1,675, N=3) | | | |
| 1 decision, short (~180 tok) | 53.9 ms (53.3–54.2) | 35.2 | 36.4 | 69 |
| **21 binary criteria**, one `/decide/shared` (2 calls: 16 + 5) | **114.3 ms** (114.1–114.3) | 80.3 | 82.3 | 159 |
| 1 decision over the ~3,900-token state | 211.2 ms (210.8–211.6) | 186.8 | 191.4 | |
| **16 criteria over the ~3,900-token state** (1 call, 4,579 tok) | **238.3 ms** (237.8–238.7) | 205.3 | 216.8 | (bench loopback: 505) |
| largest request (4 calls × 7,168 tok), N = 3 | 1,519–1,543 ms | | | |
*From nh3-dev and on GPU 1, next to the vLLM seats: pending the GPU 1 deploy (held, see the status banner).*
Server-side time on GPU 1 matches GPU 3 within a few ms, so the vLLM neighbours were not
measurably slowing it at 0942. The rest of the end-to-end time is the network (~23 ms round trip
plus the HTTP exchange).
## Acceptance (2026-09-30)
@@ -153,6 +163,8 @@ directory per process lifetime). **The harness is the bench's own.** `bench_sets
--backend semif` speaks semif-serve's API, so it drove this service unchanged. `compare.py`
scores it as the bench's `analyze.py` did.
### On GPU 3 before the deploy (the harness and the floor)
| check (GPU 3, cap 9.25 GiB, N = 3 fresh processes unless noted) | result | bench / floor |
|---|---|---|
| **positive control**, pooled 259, single ordering | **240, 240, 240** | bench 240 (4 repeats); the pooled set resolves ±4 pts |
@@ -166,12 +178,28 @@ scores it as the bench's `analyze.py` did.
| more than 16 questions: 20 decisions in one request, against the same decisions asked as 16 + 4 | 2 calls `[16, 4]`, all 20 tops equal, Δp 0, same prompt hashes | |
| 401: no token / wrong token, both POSTs; `/health` open | 401 / 401; 200 | |
| 429: 48 concurrent requests, `MAX_QUEUE` 32 | 32 answered, 16 × 429 `busy` | |
| largest request the API accepts (64 × 16 options, 4 × 8,191 tok) at cap 9.25 GiB | **200** ×3; card peak 10,134 MiB | budget 10,300 |
| largest request at `MAX_TOKENS` 8,192 (64 × 16 options, 4 × 8,191 tok), cap 9.25 GiB | **200** ×3; card peak 10,134 MiB | superseded by the 9.0 / 7,168 knobs (VRAM) |
| **over the cap** (the same request at a deliberately tight cap of 8.9 GiB) | **503 `out_of_memory`**; reserved 7.969 GiB before, 7.969 after, next request 200; then all 560 single rows answered, bit-identical to the bench | |
| fit under the cap: the largest 16-question call that answers 200 | 8,191 tok (the whole API) at 9.25 GiB; 7,168 tok at 9.0 GiB | |
| cap below startup | 8.5 GiB: startup refused (OOM loading), fails closed | |
*On the deployed service (GPU 1): pending the deploy (held, see the status banner).*
### On the live service (GPU 1, 0942–0946 PT, through `http://intern-decision.fv.internal:8033` from nh3-dev)
Raw results: `services/intern-decision-serve/acceptance/gpu1-live/`.
| check | result | reference |
|---|---|---|
| **positive control**, pooled 259, single ordering | **240** | bench 240; GPU 3 240 ×3 |
| positive control, Wyrd /84 | **79** | bench 79 |
| row by row against the bench's native rows (560 single + 144 negative) | **0 top changes, max Δp 0.000** | |
| negative control: same top / follows the description / right vs gold | **10 / 122 / 14** | bench 10 / 122 / 14 |
| 401 without or with a wrong token (both POSTs); `/health` open | 401 / 401; 200 | |
| largest accepted request: 64 decisions × 16 options, 4 calls × 7,168 tok | **200** ×3 | |
| one question at 7,168 tok | 200 ×3; one token more (7,171) → **422** "above 7168" | |
| any 503 across the whole live acceptance | **none** (server log: 128 + 145 × 200, 18 × 422, 4 × 401) | |
The live pass is N = 1 for the positive control. It is read against the GPU 3 floor: the same
image moved 0 of 560 rows across 3 restarts, and 0 against the bench's 4 restarts.
## Building and deploying
+20 -9
View File
@@ -1,14 +1,25 @@
# semif
> ⚠ **OFFLINE since 2026-09-30 0135 PT (Prime: "take semif offline for now").** It was
> stopped (`docker compose stop`, NOT removed) to give scriberr its GPU 1 headroom back:
> scriberr's Parakeet path cuts audio into 5-minute slices and needs over 6 GB, and with
> SemIf resident it had ~6.7 GB and hit CUDA OOM on a 35-minute file. Stopping it moved
> GPU 1 from 91,052 to 81,806 MiB used. The image, weights, config and token are all kept;
> `unless-stopped` keeps it down across a reboot. **Do not restart it without Prime's word.**
> To bring it back: `cd /opt/docker/compose/semif && docker compose start`, but first
> make sure scriberr has its room (the durable fix, shortening scriberr's slice length, is
> deferred by Prime to later). Everything below describes the service as it was deployed.
> ⚠ **REPLACED by intern-decision (Prime, 2026-09-30).** Prime ruled at ~0510 PT: "replace semif
> with intern-decision now". The replacement is `stacks/intern-decision` (Intern-Decision-4B, live since 0941 PT,
> `http://10.251.50.54:8033`, `intern-decision.fv.internal`). It keeps this service's HTTP surface
> (`/decide`, `/decide/shared`, `/health`), so callers only change the URL and the token
> (`secret get intern-decision/api-token`). Its README lists every deliberate difference.
>
> The `semif` container has been **stopped, not removed**, since 0135 PT that day, when it was
> taken offline to give scriberr its GPU 1 headroom back. Its image (`semif-serve:0.1.4`),
> weights, config and token are all kept, and `unless-stopped` keeps it down across a reboot.
> Port 8032 and `semif.fv.internal` stay reserved for it.
>
> **Rollback, only on Prime's word.**
> 1. Stop intern-decision first. The two services do not fit GPU 1 together next to scriberr.
> semif held 9.2 GB at rest and peaked at 12.9 GB.
> `cd /opt/docker/compose/intern-decision && docker compose stop`.
> 2. Start semif: `cd /opt/docker/compose/semif && docker compose start`.
> 3. Check nvidia-smi `Free` on GPU 1 against semif's 12.9 GB peak plus scriberr's 5.5 GB (its
> 120 s slices) before calling it done. If it does not cover both, semif collides with scriberr.
>
> Everything below describes the service as it was deployed.
**SemIf option-logit decisions** on **fv-ml1 GPU 1**, the utility card beside
`vllm-coder`, the erp/meromero seats and scriberr. Deployed 2026-09-27 at Prime's