docs(intern-decision): live on fv-ml1 GPU 1; semif marked REPLACED
intern-decision deployed 0941 PT (cap 9.0 GiB, MAX_TOKENS 7168). Live acceptance: positive control 240/259 and Wyrd 79/84, bit-identical to the bench (0/560 rows, Δp 0); negative control 10/122/14; largest accepted requests 200 with no 503; per-process 8,812 MiB at rest and 9,866 peak; GPU 1 Free 15,442 before and 6,581 after (lowest 5,569 under load). Latency from nh3-dev: 21 criteria 114 ms, 16 over ~3,900 tokens 238 ms. semif README banner now REPLACED with the rollback; fv-ml1 GPU 1 note updated.
This commit is contained in:
+2
-2
@@ -127,5 +127,5 @@ aliases:
|
|||||||
- {name: wherethef, site: nh3, target: nh3-dev, note: WhereTF :8093}
|
- {name: wherethef, site: nh3, target: nh3-dev, note: WhereTF :8093}
|
||||||
- {name: homepage, site: esh, target: esh-docker-vm, note: fleet dashboard :5100}
|
- {name: homepage, site: esh, target: esh-docker-vm, note: fleet dashboard :5100}
|
||||||
- {name: scriberr, site: fv, target: fv-ml1, note: transcription + diarization :8080 (GPU1)}
|
- {name: scriberr, site: fv, target: fv-ml1, note: transcription + diarization :8080 (GPU1)}
|
||||||
- {name: semif, site: fv, target: fv-ml1, note: SemIf option-logit decisions :8032 (GPU1) — stopped since 2026-09-30, being replaced by intern-decision (Prime); kept as rollback}
|
- {name: semif, site: fv, target: fv-ml1, note: SemIf option-logit decisions :8032 (GPU1) — REPLACED by intern-decision 2026-09-30 (Prime), stopped, kept as rollback}
|
||||||
- {name: intern-decision, site: fv, target: fv-ml1, note: Intern-Decision-4B typed decisions :8033 (GPU1), semif-serve API; replaces semif (Prime 2026-09-30)}
|
- {name: intern-decision, site: fv, target: fv-ml1, note: Intern-Decision-4B typed decisions :8033 (GPU1), semif-serve API; replaced semif 2026-09-30 (Prime)}
|
||||||
|
|||||||
@@ -181,9 +181,15 @@ embed/rerank/reward trio. GPUs are pinned per container via
|
|||||||
|
|
||||||
**GPU 1 — light / eval / retrieval + char-RP GGUF (~91/98 GB, on-demand):**
|
**GPU 1 — light / eval / retrieval + char-RP GGUF (~91/98 GB, on-demand):**
|
||||||
|
|
||||||
> ⚠ **This table is stale (checked 2026-09-27).** Live GPU 1 residents were `scriberr`,
|
> ⚠ **This table is stale (checked 2026-09-27).** Live GPU 1 residents are `scriberr`,
|
||||||
> `vllm-coder`, `vllm-erp-seat` and `vllm-meromero-rp`, plus **`semif`** (**OFFLINE since 2026-09-30, Prime; see `stacks/semif`**; :8032, ~8.7 GB resting when up,
|
> `vllm-coder`, `vllm-erp-seat` and `vllm-meromero-rp`.
|
||||||
> hard-capped at 12 GiB; `stacks/semif`, since 2026-09-27). Read the host
|
> **`intern-decision`** joined them on 2026-09-30, 0941 PT: :8033, 8,812 MiB at rest, 9,866 MiB
|
||||||
|
> peak, hard-capped at 9.0 GiB with `MAX_TOKENS` 7,168; see `stacks/intern-decision`. It replaced
|
||||||
|
> **`semif`** (:8032), which is stopped and kept as the rollback, per Prime's ruling of 2026-09-30.
|
||||||
|
> **GPU 1 budget:** nvidia-smi `Free` read 15,442 MiB before intern-decision and 6,581 MiB after,
|
||||||
|
> at rest. That covers scriberr's 5,496 MiB peak even while intern-decision is at its own peak.
|
||||||
|
> Measure nvidia-smi `Free` (the driver reserves 640 MiB per card) before adding anything to
|
||||||
|
> this card. Read the host
|
||||||
> (`docker inspect … DeviceRequests`), not this table. (A fixtures-only `augaman` instance ran
|
> (`docker inspect … DeviceRequests`), not this table. (A fixtures-only `augaman` instance ran
|
||||||
> here for about an hour on 2026-09-27 for a speed bench, and was then removed on Prime's call.)
|
> here for about an hour on 2026-09-27 for a speed bench, and was then removed on Prime's call.)
|
||||||
|
|
||||||
|
|||||||
@@ -0,0 +1,139 @@
|
|||||||
|
{
|
||||||
|
"url": "http://intern-decision.fv.internal:8033",
|
||||||
|
"started_utc": "2026-09-30T16:43:39Z",
|
||||||
|
"health": {
|
||||||
|
"status": "ok",
|
||||||
|
"model": {
|
||||||
|
"name": "Intern-Decision-4B",
|
||||||
|
"source": "internlm/Intern-Decision-4B",
|
||||||
|
"revision": "0e5e6aa7d6d750e2b1504ba11a8136cb58aeb3cd",
|
||||||
|
"checkpoint": "/hf/hub/models--internlm--Intern-Decision-4B/snapshots/0e5e6aa7d6d750e2b1504ba11a8136cb58aeb3cd",
|
||||||
|
"inference_py_sha256": "c904e2c67ca0775621a22375ee373d2ba30b52117cda870c6c9ef74143b29863",
|
||||||
|
"temperature": 1.99241824,
|
||||||
|
"dtype": "bfloat16",
|
||||||
|
"attn_implementation": "sdpa",
|
||||||
|
"device": "cuda",
|
||||||
|
"max_length": 7168,
|
||||||
|
"torch_version": "2.10.0+cu128",
|
||||||
|
"transformers_version": "5.17.0",
|
||||||
|
"vision_tower": "removed",
|
||||||
|
"device_name": "NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition",
|
||||||
|
"allocated_gib": 7.937,
|
||||||
|
"reserved_gib": 7.969,
|
||||||
|
"max_reserved_gib": 8.861
|
||||||
|
},
|
||||||
|
"vram_cap_gib": 9.0,
|
||||||
|
"max_tokens": 7168,
|
||||||
|
"max_decisions": 64,
|
||||||
|
"max_questions_per_call": 16,
|
||||||
|
"chunking": "/decide/shared questions are packed greedily, in request order, into calls of at most 16 (1-16, 17-32, ...); each call is one prompt, so the questions in a call are asked together. With orderings, ordering k of every decision forms wave k, packed the same way.",
|
||||||
|
"workloads": []
|
||||||
|
},
|
||||||
|
"auth": {
|
||||||
|
"/decide": {
|
||||||
|
"no_token": 401,
|
||||||
|
"wrong_token": 401,
|
||||||
|
"right_token": 200
|
||||||
|
},
|
||||||
|
"/decide/shared": {
|
||||||
|
"no_token": 401,
|
||||||
|
"wrong_token": 401,
|
||||||
|
"right_token": 200
|
||||||
|
},
|
||||||
|
"health_no_token": 200,
|
||||||
|
"pass": true,
|
||||||
|
"t_start": 1790786619.1466322,
|
||||||
|
"t_end": 1790786619.3572643
|
||||||
|
},
|
||||||
|
"maxreq": {
|
||||||
|
"state_chars": 4793,
|
||||||
|
"tokens_per_call": 7168,
|
||||||
|
"decisions": 64,
|
||||||
|
"options_each": 16,
|
||||||
|
"body_bytes": 96250,
|
||||||
|
"runs": [
|
||||||
|
{
|
||||||
|
"status": 200,
|
||||||
|
"e2e_ms": 1542.5,
|
||||||
|
"code": null,
|
||||||
|
"calls": 4,
|
||||||
|
"input_tokens": [
|
||||||
|
7168,
|
||||||
|
7168,
|
||||||
|
7168,
|
||||||
|
7168
|
||||||
|
],
|
||||||
|
"reserved_gib_after": 7.969,
|
||||||
|
"max_reserved_gib": 8.996,
|
||||||
|
"t_end": 1790786706.9185867
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"status": 200,
|
||||||
|
"e2e_ms": 1518.7,
|
||||||
|
"code": null,
|
||||||
|
"calls": 4,
|
||||||
|
"input_tokens": [
|
||||||
|
7168,
|
||||||
|
7168,
|
||||||
|
7168,
|
||||||
|
7168
|
||||||
|
],
|
||||||
|
"reserved_gib_after": 7.969,
|
||||||
|
"max_reserved_gib": 8.996,
|
||||||
|
"t_end": 1790786708.4557755
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"status": 200,
|
||||||
|
"e2e_ms": 1539.6,
|
||||||
|
"code": null,
|
||||||
|
"calls": 4,
|
||||||
|
"input_tokens": [
|
||||||
|
7168,
|
||||||
|
7168,
|
||||||
|
7168,
|
||||||
|
7168
|
||||||
|
],
|
||||||
|
"reserved_gib_after": 7.969,
|
||||||
|
"max_reserved_gib": 8.996,
|
||||||
|
"t_end": 1790786710.014339
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"t_start": 1790786619.357301,
|
||||||
|
"t_end": 1790786710.014397
|
||||||
|
},
|
||||||
|
"maxone": {
|
||||||
|
"state_chars": 28727,
|
||||||
|
"tokens": 7168,
|
||||||
|
"one_more_is": [
|
||||||
|
422,
|
||||||
|
"Example has 7171 tokens, above 7168; truncation is forbidden"
|
||||||
|
],
|
||||||
|
"runs": [
|
||||||
|
{
|
||||||
|
"status": 200,
|
||||||
|
"e2e_ms": 394.2,
|
||||||
|
"input_tokens": 7168,
|
||||||
|
"code": null,
|
||||||
|
"t_end": 1790786715.324555
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"status": 200,
|
||||||
|
"e2e_ms": 375.9,
|
||||||
|
"input_tokens": 7168,
|
||||||
|
"code": null,
|
||||||
|
"t_end": 1790786715.7005756
|
||||||
|
},
|
||||||
|
{
|
||||||
|
"status": 200,
|
||||||
|
"e2e_ms": 372.7,
|
||||||
|
"input_tokens": 7168,
|
||||||
|
"code": null,
|
||||||
|
"t_end": 1790786716.0734687
|
||||||
|
}
|
||||||
|
],
|
||||||
|
"pass": true,
|
||||||
|
"t_start": 1790786710.0144272,
|
||||||
|
"t_end": 1790786716.0735004
|
||||||
|
},
|
||||||
|
"finished_utc": "2026-09-30T16:45:16Z"
|
||||||
|
}
|
||||||
@@ -0,0 +1,671 @@
|
|||||||
|
{
|
||||||
|
"ours": {
|
||||||
|
"live": {
|
||||||
|
"scores": {
|
||||||
|
"single/pooled": [
|
||||||
|
240,
|
||||||
|
259
|
||||||
|
],
|
||||||
|
"single/authored144": [
|
||||||
|
132,
|
||||||
|
144
|
||||||
|
],
|
||||||
|
"single/perturbations108": [
|
||||||
|
100,
|
||||||
|
108
|
||||||
|
],
|
||||||
|
"single/cicada-w1": [
|
||||||
|
29,
|
||||||
|
31
|
||||||
|
],
|
||||||
|
"single/cicada-w2": [
|
||||||
|
22,
|
||||||
|
31
|
||||||
|
],
|
||||||
|
"single/wyrd": [
|
||||||
|
79,
|
||||||
|
84
|
||||||
|
],
|
||||||
|
"single/wyrd:place": [
|
||||||
|
19,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"single/wyrd:place2": [
|
||||||
|
21,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"single/wyrd:exit": [
|
||||||
|
19,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"single/wyrd:exit2": [
|
||||||
|
20,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"single/failures": 0
|
||||||
|
},
|
||||||
|
"negative": {
|
||||||
|
"n": 144,
|
||||||
|
"same_top_as_unrotated": 10,
|
||||||
|
"follows_description": 122,
|
||||||
|
"vs_original_gold": 14
|
||||||
|
},
|
||||||
|
"a_vs_a_in_process_authored144": {
|
||||||
|
"rows": 0,
|
||||||
|
"top_differs": 0,
|
||||||
|
"flipped_ids": [],
|
||||||
|
"max_dp": null
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"bench": {
|
||||||
|
"r1": {
|
||||||
|
"scores": {
|
||||||
|
"single/pooled": [
|
||||||
|
240,
|
||||||
|
259
|
||||||
|
],
|
||||||
|
"single/authored144": [
|
||||||
|
132,
|
||||||
|
144
|
||||||
|
],
|
||||||
|
"single/perturbations108": [
|
||||||
|
100,
|
||||||
|
108
|
||||||
|
],
|
||||||
|
"single/cicada-w1": [
|
||||||
|
29,
|
||||||
|
31
|
||||||
|
],
|
||||||
|
"single/cicada-w2": [
|
||||||
|
22,
|
||||||
|
31
|
||||||
|
],
|
||||||
|
"single/wyrd": [
|
||||||
|
79,
|
||||||
|
84
|
||||||
|
],
|
||||||
|
"single/wyrd:place": [
|
||||||
|
19,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"single/wyrd:place2": [
|
||||||
|
21,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"single/wyrd:exit": [
|
||||||
|
19,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"single/wyrd:exit2": [
|
||||||
|
20,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"single/failures": 0,
|
||||||
|
"rotations/pooled": [
|
||||||
|
236,
|
||||||
|
259
|
||||||
|
],
|
||||||
|
"rotations/authored144": [
|
||||||
|
130,
|
||||||
|
144
|
||||||
|
],
|
||||||
|
"rotations/perturbations108": [
|
||||||
|
99,
|
||||||
|
108
|
||||||
|
],
|
||||||
|
"rotations/cicada-w1": [
|
||||||
|
29,
|
||||||
|
31
|
||||||
|
],
|
||||||
|
"rotations/cicada-w2": [
|
||||||
|
22,
|
||||||
|
31
|
||||||
|
],
|
||||||
|
"rotations/wyrd": [
|
||||||
|
77,
|
||||||
|
84
|
||||||
|
],
|
||||||
|
"rotations/wyrd:place": [
|
||||||
|
18,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"rotations/wyrd:place2": [
|
||||||
|
21,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"rotations/wyrd:exit": [
|
||||||
|
19,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"rotations/wyrd:exit2": [
|
||||||
|
19,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"rotations/failures": 0
|
||||||
|
},
|
||||||
|
"negative": {
|
||||||
|
"n": 144,
|
||||||
|
"same_top_as_unrotated": 10,
|
||||||
|
"follows_description": 122,
|
||||||
|
"vs_original_gold": 14
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"r2": {
|
||||||
|
"scores": {
|
||||||
|
"single/pooled": [
|
||||||
|
240,
|
||||||
|
259
|
||||||
|
],
|
||||||
|
"single/authored144": [
|
||||||
|
132,
|
||||||
|
144
|
||||||
|
],
|
||||||
|
"single/perturbations108": [
|
||||||
|
100,
|
||||||
|
108
|
||||||
|
],
|
||||||
|
"single/cicada-w1": [
|
||||||
|
29,
|
||||||
|
31
|
||||||
|
],
|
||||||
|
"single/cicada-w2": [
|
||||||
|
22,
|
||||||
|
31
|
||||||
|
],
|
||||||
|
"single/wyrd": [
|
||||||
|
79,
|
||||||
|
84
|
||||||
|
],
|
||||||
|
"single/wyrd:place": [
|
||||||
|
19,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"single/wyrd:place2": [
|
||||||
|
21,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"single/wyrd:exit": [
|
||||||
|
19,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"single/wyrd:exit2": [
|
||||||
|
20,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"single/failures": 0,
|
||||||
|
"rotations/pooled": [
|
||||||
|
236,
|
||||||
|
259
|
||||||
|
],
|
||||||
|
"rotations/authored144": [
|
||||||
|
130,
|
||||||
|
144
|
||||||
|
],
|
||||||
|
"rotations/perturbations108": [
|
||||||
|
99,
|
||||||
|
108
|
||||||
|
],
|
||||||
|
"rotations/cicada-w1": [
|
||||||
|
29,
|
||||||
|
31
|
||||||
|
],
|
||||||
|
"rotations/cicada-w2": [
|
||||||
|
22,
|
||||||
|
31
|
||||||
|
],
|
||||||
|
"rotations/wyrd": [
|
||||||
|
77,
|
||||||
|
84
|
||||||
|
],
|
||||||
|
"rotations/wyrd:place": [
|
||||||
|
18,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"rotations/wyrd:place2": [
|
||||||
|
21,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"rotations/wyrd:exit": [
|
||||||
|
19,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"rotations/wyrd:exit2": [
|
||||||
|
19,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"rotations/failures": 0,
|
||||||
|
"multifield/pooled": [
|
||||||
|
77,
|
||||||
|
84
|
||||||
|
],
|
||||||
|
"multifield/wyrd": [
|
||||||
|
77,
|
||||||
|
84
|
||||||
|
],
|
||||||
|
"multifield/wyrd:place": [
|
||||||
|
16,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"multifield/wyrd:place2": [
|
||||||
|
21,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"multifield/wyrd:exit": [
|
||||||
|
20,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"multifield/wyrd:exit2": [
|
||||||
|
20,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"multifield/failures": 0
|
||||||
|
},
|
||||||
|
"negative": {
|
||||||
|
"n": 144,
|
||||||
|
"same_top_as_unrotated": 10,
|
||||||
|
"follows_description": 122,
|
||||||
|
"vs_original_gold": 14
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"r3": {
|
||||||
|
"scores": {
|
||||||
|
"single/pooled": [
|
||||||
|
240,
|
||||||
|
259
|
||||||
|
],
|
||||||
|
"single/authored144": [
|
||||||
|
132,
|
||||||
|
144
|
||||||
|
],
|
||||||
|
"single/perturbations108": [
|
||||||
|
100,
|
||||||
|
108
|
||||||
|
],
|
||||||
|
"single/cicada-w1": [
|
||||||
|
29,
|
||||||
|
31
|
||||||
|
],
|
||||||
|
"single/cicada-w2": [
|
||||||
|
22,
|
||||||
|
31
|
||||||
|
],
|
||||||
|
"single/wyrd": [
|
||||||
|
79,
|
||||||
|
84
|
||||||
|
],
|
||||||
|
"single/wyrd:place": [
|
||||||
|
19,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"single/wyrd:place2": [
|
||||||
|
21,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"single/wyrd:exit": [
|
||||||
|
19,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"single/wyrd:exit2": [
|
||||||
|
20,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"single/failures": 0,
|
||||||
|
"rotations/pooled": [
|
||||||
|
236,
|
||||||
|
259
|
||||||
|
],
|
||||||
|
"rotations/authored144": [
|
||||||
|
130,
|
||||||
|
144
|
||||||
|
],
|
||||||
|
"rotations/perturbations108": [
|
||||||
|
99,
|
||||||
|
108
|
||||||
|
],
|
||||||
|
"rotations/cicada-w1": [
|
||||||
|
29,
|
||||||
|
31
|
||||||
|
],
|
||||||
|
"rotations/cicada-w2": [
|
||||||
|
22,
|
||||||
|
31
|
||||||
|
],
|
||||||
|
"rotations/wyrd": [
|
||||||
|
77,
|
||||||
|
84
|
||||||
|
],
|
||||||
|
"rotations/wyrd:place": [
|
||||||
|
18,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"rotations/wyrd:place2": [
|
||||||
|
21,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"rotations/wyrd:exit": [
|
||||||
|
19,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"rotations/wyrd:exit2": [
|
||||||
|
19,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"rotations/failures": 0,
|
||||||
|
"multifield/pooled": [
|
||||||
|
77,
|
||||||
|
84
|
||||||
|
],
|
||||||
|
"multifield/wyrd": [
|
||||||
|
77,
|
||||||
|
84
|
||||||
|
],
|
||||||
|
"multifield/wyrd:place": [
|
||||||
|
16,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"multifield/wyrd:place2": [
|
||||||
|
21,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"multifield/wyrd:exit": [
|
||||||
|
20,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"multifield/wyrd:exit2": [
|
||||||
|
20,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"multifield/failures": 0
|
||||||
|
},
|
||||||
|
"negative": {
|
||||||
|
"n": 144,
|
||||||
|
"same_top_as_unrotated": 10,
|
||||||
|
"follows_description": 122,
|
||||||
|
"vs_original_gold": 14
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"r4": {
|
||||||
|
"scores": {
|
||||||
|
"single/pooled": [
|
||||||
|
240,
|
||||||
|
259
|
||||||
|
],
|
||||||
|
"single/authored144": [
|
||||||
|
132,
|
||||||
|
144
|
||||||
|
],
|
||||||
|
"single/perturbations108": [
|
||||||
|
100,
|
||||||
|
108
|
||||||
|
],
|
||||||
|
"single/cicada-w1": [
|
||||||
|
29,
|
||||||
|
31
|
||||||
|
],
|
||||||
|
"single/cicada-w2": [
|
||||||
|
22,
|
||||||
|
31
|
||||||
|
],
|
||||||
|
"single/wyrd": [
|
||||||
|
79,
|
||||||
|
84
|
||||||
|
],
|
||||||
|
"single/wyrd:place": [
|
||||||
|
19,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"single/wyrd:place2": [
|
||||||
|
21,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"single/wyrd:exit": [
|
||||||
|
19,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"single/wyrd:exit2": [
|
||||||
|
20,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"single/failures": 0,
|
||||||
|
"rotations/pooled": [
|
||||||
|
236,
|
||||||
|
259
|
||||||
|
],
|
||||||
|
"rotations/authored144": [
|
||||||
|
130,
|
||||||
|
144
|
||||||
|
],
|
||||||
|
"rotations/perturbations108": [
|
||||||
|
99,
|
||||||
|
108
|
||||||
|
],
|
||||||
|
"rotations/cicada-w1": [
|
||||||
|
29,
|
||||||
|
31
|
||||||
|
],
|
||||||
|
"rotations/cicada-w2": [
|
||||||
|
22,
|
||||||
|
31
|
||||||
|
],
|
||||||
|
"rotations/wyrd": [
|
||||||
|
77,
|
||||||
|
84
|
||||||
|
],
|
||||||
|
"rotations/wyrd:place": [
|
||||||
|
18,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"rotations/wyrd:place2": [
|
||||||
|
21,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"rotations/wyrd:exit": [
|
||||||
|
19,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"rotations/wyrd:exit2": [
|
||||||
|
19,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"rotations/failures": 0,
|
||||||
|
"multifield/pooled": [
|
||||||
|
77,
|
||||||
|
84
|
||||||
|
],
|
||||||
|
"multifield/wyrd": [
|
||||||
|
77,
|
||||||
|
84
|
||||||
|
],
|
||||||
|
"multifield/wyrd:place": [
|
||||||
|
16,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"multifield/wyrd:place2": [
|
||||||
|
21,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"multifield/wyrd:exit": [
|
||||||
|
20,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"multifield/wyrd:exit2": [
|
||||||
|
20,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"multifield/failures": 0
|
||||||
|
},
|
||||||
|
"negative": {
|
||||||
|
"n": 144,
|
||||||
|
"same_top_as_unrotated": 10,
|
||||||
|
"follows_description": 122,
|
||||||
|
"vs_original_gold": 14
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"vs_bench_row_by_row": {
|
||||||
|
"live": {
|
||||||
|
"single": {
|
||||||
|
"bench_repeat": "r1",
|
||||||
|
"rows": 560,
|
||||||
|
"top_differs": 0,
|
||||||
|
"flipped_ids": [],
|
||||||
|
"max_dp": 0.0
|
||||||
|
},
|
||||||
|
"rotations": {
|
||||||
|
"bench_repeat": "r1",
|
||||||
|
"rows": 0,
|
||||||
|
"top_differs": 0,
|
||||||
|
"flipped_ids": [],
|
||||||
|
"max_dp": null
|
||||||
|
},
|
||||||
|
"negative": {
|
||||||
|
"bench_repeat": "r1",
|
||||||
|
"rows": 144,
|
||||||
|
"top_differs": 0,
|
||||||
|
"flipped_ids": [],
|
||||||
|
"max_dp": 0.0
|
||||||
|
},
|
||||||
|
"multifield": {
|
||||||
|
"bench_repeat": "r2",
|
||||||
|
"rows": 0,
|
||||||
|
"top_differs": 0,
|
||||||
|
"flipped_ids": [],
|
||||||
|
"max_dp": null
|
||||||
|
}
|
||||||
|
}
|
||||||
|
},
|
||||||
|
"across_restarts": {},
|
||||||
|
"summary": {
|
||||||
|
"single/authored144": {
|
||||||
|
"ours_median": 132,
|
||||||
|
"ours_all": [
|
||||||
|
132
|
||||||
|
],
|
||||||
|
"bench_all": [
|
||||||
|
132,
|
||||||
|
132,
|
||||||
|
132,
|
||||||
|
132
|
||||||
|
],
|
||||||
|
"n": 144
|
||||||
|
},
|
||||||
|
"single/cicada-w1": {
|
||||||
|
"ours_median": 29,
|
||||||
|
"ours_all": [
|
||||||
|
29
|
||||||
|
],
|
||||||
|
"bench_all": [
|
||||||
|
29,
|
||||||
|
29,
|
||||||
|
29,
|
||||||
|
29
|
||||||
|
],
|
||||||
|
"n": 31
|
||||||
|
},
|
||||||
|
"single/cicada-w2": {
|
||||||
|
"ours_median": 22,
|
||||||
|
"ours_all": [
|
||||||
|
22
|
||||||
|
],
|
||||||
|
"bench_all": [
|
||||||
|
22,
|
||||||
|
22,
|
||||||
|
22,
|
||||||
|
22
|
||||||
|
],
|
||||||
|
"n": 31
|
||||||
|
},
|
||||||
|
"single/perturbations108": {
|
||||||
|
"ours_median": 100,
|
||||||
|
"ours_all": [
|
||||||
|
100
|
||||||
|
],
|
||||||
|
"bench_all": [
|
||||||
|
100,
|
||||||
|
100,
|
||||||
|
100,
|
||||||
|
100
|
||||||
|
],
|
||||||
|
"n": 108
|
||||||
|
},
|
||||||
|
"single/pooled": {
|
||||||
|
"ours_median": 240,
|
||||||
|
"ours_all": [
|
||||||
|
240
|
||||||
|
],
|
||||||
|
"bench_all": [
|
||||||
|
240,
|
||||||
|
240,
|
||||||
|
240,
|
||||||
|
240
|
||||||
|
],
|
||||||
|
"n": 259
|
||||||
|
},
|
||||||
|
"single/wyrd": {
|
||||||
|
"ours_median": 79,
|
||||||
|
"ours_all": [
|
||||||
|
79
|
||||||
|
],
|
||||||
|
"bench_all": [
|
||||||
|
79,
|
||||||
|
79,
|
||||||
|
79,
|
||||||
|
79
|
||||||
|
],
|
||||||
|
"n": 84
|
||||||
|
},
|
||||||
|
"single/wyrd:exit": {
|
||||||
|
"ours_median": 19,
|
||||||
|
"ours_all": [
|
||||||
|
19
|
||||||
|
],
|
||||||
|
"bench_all": [
|
||||||
|
19,
|
||||||
|
19,
|
||||||
|
19,
|
||||||
|
19
|
||||||
|
],
|
||||||
|
"n": 21
|
||||||
|
},
|
||||||
|
"single/wyrd:exit2": {
|
||||||
|
"ours_median": 20,
|
||||||
|
"ours_all": [
|
||||||
|
20
|
||||||
|
],
|
||||||
|
"bench_all": [
|
||||||
|
20,
|
||||||
|
20,
|
||||||
|
20,
|
||||||
|
20
|
||||||
|
],
|
||||||
|
"n": 21
|
||||||
|
},
|
||||||
|
"single/wyrd:place": {
|
||||||
|
"ours_median": 19,
|
||||||
|
"ours_all": [
|
||||||
|
19
|
||||||
|
],
|
||||||
|
"bench_all": [
|
||||||
|
19,
|
||||||
|
19,
|
||||||
|
19,
|
||||||
|
19
|
||||||
|
],
|
||||||
|
"n": 21
|
||||||
|
},
|
||||||
|
"single/wyrd:place2": {
|
||||||
|
"ours_median": 21,
|
||||||
|
"ours_all": [
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"bench_all": [
|
||||||
|
21,
|
||||||
|
21,
|
||||||
|
21,
|
||||||
|
21
|
||||||
|
],
|
||||||
|
"n": 21
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
@@ -0,0 +1,6 @@
|
|||||||
|
1790786561.567928123 shape:start
|
||||||
|
1790786613.903648987 shape:done
|
||||||
|
1790786619.093843341 checks:start
|
||||||
|
1790786716.079809795 checks:done
|
||||||
|
1790786739.656804839 sets:start
|
||||||
|
1790786809.130379221 sets:done
|
||||||
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because it is too large
Load Diff
@@ -1,10 +1,7 @@
|
|||||||
# intern-decision
|
# intern-decision
|
||||||
|
|
||||||
> ⚠ **STATUS 2026-09-30 0937 PT: built and accepted on GPU 3; the GPU 1 deploy is HELD.** GPU 1's
|
> **LIVE since 2026-09-30 0941 PT** on fv-ml1 GPU 1, with cap 9.0 GiB and `MAX_TOKENS` 7,168.
|
||||||
> usable free memory is 15,442 MiB (nvidia-smi `Free`: 97,887 total − 640 driver-reserved −
|
> Acceptance on the live service is below. `semif` is stopped and kept as the rollback.
|
||||||
> 81,806 used), below the 15,800 MiB pre-deploy floor. The 10,300 MiB budget was derived from
|
|
||||||
> total − used (16,081). The decision is with infra-ops/Prime; see "VRAM". Nothing is running
|
|
||||||
> on port 8033 yet.
|
|
||||||
|
|
||||||
**Intern-Decision-4B typed decisions** on **fv-ml1 GPU 1**, the utility card beside `vllm-coder`,
|
**Intern-Decision-4B typed decisions** on **fv-ml1 GPU 1**, the utility card beside `vllm-coder`,
|
||||||
the erp/meromero seats and scriberr. It **replaced `semif`** on Prime's ruling of 2026-09-30,
|
the erp/meromero seats and scriberr. It **replaced `semif`** on Prime's ruling of 2026-09-30,
|
||||||
@@ -76,48 +73,58 @@ curl -s -H "Authorization: Bearer $T" http://intern-decision.fv.internal:8033/de
|
|||||||
- `/health.semif_commit`;
|
- `/health.semif_commit`;
|
||||||
- per-workload calibration. Any `workload` is a 422, exactly as the deployed semif behaved
|
- per-workload calibration. Any `workload` is a 422, exactly as the deployed semif behaved
|
||||||
with its empty table.
|
with its empty table.
|
||||||
5. **`MAX_TOKENS` (8192) is per call:** the state plus all its questions. A longer call is a 422
|
5. **`MAX_TOKENS` is 7,168 per call:** the state plus all its questions. It is lower than the
|
||||||
and is never truncated. Every call up to that limit fits the VRAM cap (measured). See VRAM
|
model's own 8,192 so that every call the API accepts fits the VRAM cap.
|
||||||
below.
|
- The count is taken **before** the forward pass, so a longer call is a clear
|
||||||
|
`422 invalid_request` ("Example has N tokens, above 7168; truncation is forbidden"). It is
|
||||||
|
never truncated.
|
||||||
|
- A request with more decisions is split into more calls, and each call must fit.
|
||||||
|
|
||||||
## VRAM: the whole container ≤ 10,300 MiB
|
## VRAM: fits beside scriberr's peak, whatever the request
|
||||||
|
|
||||||
**Budget (infra-ops, 2026-09-30):** the container's **whole** nvidia-smi footprint, CUDA context
|
**Budget (infra-ops, 2026-09-30):** GPU 1 needs nvidia-smi `Free` ≥ **15,400 MiB** before this
|
||||||
included, must never exceed **10,300 MiB**, whatever the request. That leaves scriberr (peak
|
service starts. That is our card peak of 9,876 MiB plus scriberr's peak of 5,496 MiB (with its
|
||||||
5,496 MiB with its 120 s slices) 295 MiB of spare even when both hit their peaks at once.
|
120 s slices), rounded up. So both can peak at the same moment. Read nvidia-smi's own `Free`, not
|
||||||
`torch.cuda.set_per_process_memory_fraction` caps only torch's allocator, so the cap is the budget
|
total − used: the driver reserves 640 MiB on every card. The deploy at 0940 read 15,442.
|
||||||
minus the measured non-allocator overhead:
|
|
||||||
|
`torch.cuda.set_per_process_memory_fraction` caps only torch's allocator. The two knobs in `.env`
|
||||||
|
are therefore chosen together and change together:
|
||||||
|
|
||||||
```
|
```
|
||||||
footprint <= VRAM_CAP_GIB + overhead = 9,472 MiB (9.25 GiB) + 662 MiB = 10,134 MiB (166 MiB under budget)
|
card footprint <= VRAM_CAP_GIB 9.0 (9,216 MiB) + 660 MiB outside the allocator (measured) = 9,876 MiB
|
||||||
|
MAX_TOKENS 7168 = the largest call measured to fit under that cap, so no accepted call reaches the 503
|
||||||
```
|
```
|
||||||
|
|
||||||
`VRAM_CAP_GIB` in `.env` is **the single knob**. It is applied before the weights load.
|
| on the live service, GPU 1 (per-process nvidia-smi, 0.1 s sampling, 2,939 samples over 316 s of acceptance load) | MiB |
|
||||||
|
|---|---|
|
||||||
|
| **at rest** after startup | **8,812** |
|
||||||
|
| at rest after the acceptance (512 MiB release slack keeps a little cache) | 8,856 |
|
||||||
|
| **peak**: largest requests, 64 decisions × 16 options, 4 calls × 7,168 tokens, N = 3 | **9,866**; torch `max_reserved` 8.996 of 9.0 GiB |
|
||||||
|
| GPU 1 `Free` before the deploy / at rest after / lowest during the acceptance | 15,442 / 6,581 / 5,569, which stays above scriberr's 5,496 |
|
||||||
|
|
||||||
⚠ **The budget assumed 16,081 MiB free on GPU 1. That is total − used.** nvidia-smi's own `Free` is
|
- ⚠ **At 7,168 tokens the allocator reaches the cap with about 4 MiB to spare.** It fits because
|
||||||
**15,442 MiB**, because the driver reserves 640 MiB on every card (GPU 3 shows the same 640).
|
the allocator frees its cache and retries before it fails.
|
||||||
Against the true free memory, a footprint that never collides with scriberr's 5,496 MiB peak must
|
- Measured: every call up to the limit answered 200. That is 145 shared calls on the live
|
||||||
stay at or below **9,946 MiB**:
|
service, including the size search near the boundary, 3 × 4 calls at 7,168 tokens with
|
||||||
|
16 questions, and 3 at 7,168 tokens with one question. No 503 was seen.
|
||||||
|
- The allocator's cache state depends on request history. If a max-size call ever does return
|
||||||
|
503, lower `MAX_TOKENS`. Do not raise the cap.
|
||||||
|
- Raising either knob needs a fresh measurement (`acceptance/checks.py --checks fit,maxreq`) and
|
||||||
|
the budget re-checked against scriberr.
|
||||||
|
- **Over the cap the answer is 503 `out_of_memory`.** Memory returns to the resting baseline and
|
||||||
|
the service keeps answering. This was proven on GPU 3 with a deliberately tight cap.
|
||||||
|
|
||||||
| cap | footprint peak (measured, GPU 3) | with scriberr at its peak, against 15,442 free | calls that fit (16 questions × 16 options) |
|
Measured on GPU 3 before the deploy (alone on the card), which is where the knobs came from:
|
||||||
|---|---|---|---|
|
|
||||||
| **9.25 GiB** (`.env.example` now) | 10,134 MiB | **188 MiB over** | every call the API accepts (8,191 tok) |
|
|
||||||
| **9.0 GiB** | 9,876 MiB | 70 MiB spare | up to 7,168 tok; longer calls get 503. Every bench shape fits, at the same latency |
|
|
||||||
|
|
||||||
| measured on fv-ml1 GPU 3 (alone on the card), image 0.1.0 | MiB |
|
| measured on fv-ml1 GPU 3 (alone on the card), image 0.1.0 | MiB |
|
||||||
|---|---|
|
|---|---|
|
||||||
| **at rest** (nvidia-smi, whole card less 2 MiB idle) | **8,820** (torch reserved 8,160 + outside 660) |
|
| **at rest** (nvidia-smi, whole card less 2 MiB idle) | **8,820** (torch reserved 8,160 + outside 660) |
|
||||||
| outside the allocator (CUDA context + loaded kernels), at rest / after 2×30 concurrent requests / after the latency shapes / after 4 × 8,191-token calls | 660 / 660 / 662 / 662 |
|
| outside the allocator (CUDA context + loaded kernels), at rest / after 2×30 concurrent requests / after the latency shapes / after 4 × 8,191-token calls | 660 / 660 / 662 / 662 |
|
||||||
| **peak, largest request the API accepts, capped at 9.25 GiB** (card, 0.1 s sampling) | **10,134**; torch `max_reserved` = 9.25 GiB exactly, so the cap is what held it |
|
| card peak at cap 9.0 GiB (fit search and a 64 × 8,191-token request, which got 503) | **9,876** |
|
||||||
| the same request **uncapped**, for comparison | 10,422 (allocator peak 9,760): **over budget, so the cap is required** |
|
| peak with 4 × 8,191-token calls at cap 9.25 GiB (answered 200) / uncapped | 10,134 / 10,422 (allocator 9,760) |
|
||||||
|
| largest 16-question call that answers 200: at cap 9.0 / at 9.25 GiB | 7,168 / 8,191 tokens |
|
||||||
| startup peak (weights + vision tower until the swap + warm-up) | allocator 8.86 GiB. **A cap below ~8.9 GiB cannot start** (it fails closed; 8.5 was refused) |
|
| startup peak (weights + vision tower until the swap + warm-up) | allocator 8.86 GiB. **A cap below ~8.9 GiB cannot start** (it fails closed; 8.5 was refused) |
|
||||||
|
|
||||||
- **The largest request the API accepts** is 64 decisions × 16 options, with each of its 4 calls
|
|
||||||
at 8,191 tokens. Under the cap it answers **200** (1.67 s), not 503. Near the limit the
|
|
||||||
allocator frees cached blocks before it fails, so the uncapped 9,760 MiB peak was partly cache.
|
|
||||||
Every call up to `MAX_TOKENS` fits the cap.
|
|
||||||
- **Over the cap the answer is 503 `out_of_memory`.** Memory goes back to the resting baseline
|
|
||||||
and the service keeps answering. This was proven with a deliberately tight cap (acceptance).
|
|
||||||
- **One inference thread is load-bearing** (contract INV-2). torch keeps CUDA state per host
|
- **One inference thread is load-bearing** (contract INV-2). torch keeps CUDA state per host
|
||||||
thread (cuBLAS handles and workspaces). The first build scored on anyio's 40-thread pool: a
|
thread (cuBLAS handles and workspaces). The first build scored on anyio's 40-thread pool: a
|
||||||
concurrent burst added 252 MiB outside the cap and 326 MiB inside it, and the peak reached
|
concurrent burst added 252 MiB outside the cap and 326 MiB inside it, and the peak reached
|
||||||
@@ -132,19 +139,22 @@ stay at or below **9,946 MiB**:
|
|||||||
|
|
||||||
## Latency
|
## Latency
|
||||||
|
|
||||||
Loopback on fv-ml1 (GPU 3 alone on the card, image 0.1.0, cap 9.25 GiB). `bench_shape.py`
|
**Live, from nh3-dev** (0942 PT). The URL is `intern-decision.fv.internal:8033`, and the round trip
|
||||||
from the Jev bench ran 3 runs × 20 requests, and the medians are shown with the run-median
|
is 22.9 ms on average (19.5–27.7 ms over 10 pings). GPU 1 is shared with the vLLM seats and
|
||||||
range. The bench's native Intern-Decision numbers are the reference.
|
scriberr. The harness is `bench_shape.py` from the Jev bench, 3 runs × 20 requests; each cell is
|
||||||
|
the median with the run-median range. "server" is the service's own time.
|
||||||
|
|
||||||
| shape | end to end | server | bench (native, loopback) | SemIf 0.1.4 (bench) |
|
| shape | end to end from nh3-dev | server | GPU 3 loopback before deploy | SemIf 0.1.4 (README, from nh3-dev) |
|
||||||
|---|---|---|---|---|
|
|---|---|---|---|---|
|
||||||
| 1 decision, short (~180 tok) | 36.4 ms (36.4–36.5) | 34.7 | 39 | 37 |
|
| 1 decision, short (~180 tok) | 53.9 ms (53.3–54.2) | 35.2 | 36.4 | 69 |
|
||||||
| **21 binary criteria**, one `/decide/shared` (2 calls: 16 + 5) | **82.3 ms** (82.3–82.4) | 78.9 | 88 | 131 |
|
| **21 binary criteria**, one `/decide/shared` (2 calls: 16 + 5) | **114.3 ms** (114.1–114.3) | 80.3 | 82.3 | 159 |
|
||||||
| 1 decision over the ~3,900-token state | 191.4 ms (191.4–191.5) | 187.9 | 190 | 186 |
|
| 1 decision over the ~3,900-token state | 211.2 ms (210.8–211.6) | 186.8 | 191.4 | |
|
||||||
| **16 criteria over the ~3,900-token state** (1 call, 4,579 tok) | **216.8 ms** (216.5–217.0) | 211.1 | 215 | 505 |
|
| **16 criteria over the ~3,900-token state** (1 call, 4,579 tok) | **238.3 ms** (237.8–238.7) | 205.3 | 216.8 | (bench loopback: 505) |
|
||||||
| largest accepted request (4 calls × 8,191 tok) | 1,672 ms (1,672–1,675, N=3) | | | |
|
| largest request (4 calls × 7,168 tok), N = 3 | 1,519–1,543 ms | | | |
|
||||||
|
|
||||||
*From nh3-dev and on GPU 1, next to the vLLM seats: pending the GPU 1 deploy (held, see the status banner).*
|
Server-side time on GPU 1 matches GPU 3 within a few ms, so the vLLM neighbours were not
|
||||||
|
measurably slowing it at 0942. The rest of the end-to-end time is the network (~23 ms round trip
|
||||||
|
plus the HTTP exchange).
|
||||||
|
|
||||||
## Acceptance (2026-09-30)
|
## Acceptance (2026-09-30)
|
||||||
|
|
||||||
@@ -153,6 +163,8 @@ directory per process lifetime). **The harness is the bench's own.** `bench_sets
|
|||||||
--backend semif` speaks semif-serve's API, so it drove this service unchanged. `compare.py`
|
--backend semif` speaks semif-serve's API, so it drove this service unchanged. `compare.py`
|
||||||
scores it as the bench's `analyze.py` did.
|
scores it as the bench's `analyze.py` did.
|
||||||
|
|
||||||
|
### On GPU 3 before the deploy (the harness and the floor)
|
||||||
|
|
||||||
| check (GPU 3, cap 9.25 GiB, N = 3 fresh processes unless noted) | result | bench / floor |
|
| check (GPU 3, cap 9.25 GiB, N = 3 fresh processes unless noted) | result | bench / floor |
|
||||||
|---|---|---|
|
|---|---|---|
|
||||||
| **positive control**, pooled 259, single ordering | **240, 240, 240** | bench 240 (4 repeats); the pooled set resolves ±4 pts |
|
| **positive control**, pooled 259, single ordering | **240, 240, 240** | bench 240 (4 repeats); the pooled set resolves ±4 pts |
|
||||||
@@ -166,12 +178,28 @@ scores it as the bench's `analyze.py` did.
|
|||||||
| more than 16 questions: 20 decisions in one request, against the same decisions asked as 16 + 4 | 2 calls `[16, 4]`, all 20 tops equal, Δp 0, same prompt hashes | |
|
| more than 16 questions: 20 decisions in one request, against the same decisions asked as 16 + 4 | 2 calls `[16, 4]`, all 20 tops equal, Δp 0, same prompt hashes | |
|
||||||
| 401: no token / wrong token, both POSTs; `/health` open | 401 / 401; 200 | |
|
| 401: no token / wrong token, both POSTs; `/health` open | 401 / 401; 200 | |
|
||||||
| 429: 48 concurrent requests, `MAX_QUEUE` 32 | 32 answered, 16 × 429 `busy` | |
|
| 429: 48 concurrent requests, `MAX_QUEUE` 32 | 32 answered, 16 × 429 `busy` | |
|
||||||
| largest request the API accepts (64 × 16 options, 4 × 8,191 tok) at cap 9.25 GiB | **200** ×3; card peak 10,134 MiB | budget 10,300 |
|
| largest request at `MAX_TOKENS` 8,192 (64 × 16 options, 4 × 8,191 tok), cap 9.25 GiB | **200** ×3; card peak 10,134 MiB | superseded by the 9.0 / 7,168 knobs (VRAM) |
|
||||||
| **over the cap** (the same request at a deliberately tight cap of 8.9 GiB) | **503 `out_of_memory`**; reserved 7.969 GiB before, 7.969 after, next request 200; then all 560 single rows answered, bit-identical to the bench | |
|
| **over the cap** (the same request at a deliberately tight cap of 8.9 GiB) | **503 `out_of_memory`**; reserved 7.969 GiB before, 7.969 after, next request 200; then all 560 single rows answered, bit-identical to the bench | |
|
||||||
| fit under the cap: the largest 16-question call that answers 200 | 8,191 tok (the whole API) at 9.25 GiB; 7,168 tok at 9.0 GiB | |
|
| fit under the cap: the largest 16-question call that answers 200 | 8,191 tok (the whole API) at 9.25 GiB; 7,168 tok at 9.0 GiB | |
|
||||||
| cap below startup | 8.5 GiB: startup refused (OOM loading), fails closed | |
|
| cap below startup | 8.5 GiB: startup refused (OOM loading), fails closed | |
|
||||||
|
|
||||||
*On the deployed service (GPU 1): pending the deploy (held, see the status banner).*
|
### On the live service (GPU 1, 0942–0946 PT, through `http://intern-decision.fv.internal:8033` from nh3-dev)
|
||||||
|
|
||||||
|
Raw results: `services/intern-decision-serve/acceptance/gpu1-live/`.
|
||||||
|
|
||||||
|
| check | result | reference |
|
||||||
|
|---|---|---|
|
||||||
|
| **positive control**, pooled 259, single ordering | **240** | bench 240; GPU 3 240 ×3 |
|
||||||
|
| positive control, Wyrd /84 | **79** | bench 79 |
|
||||||
|
| row by row against the bench's native rows (560 single + 144 negative) | **0 top changes, max Δp 0.000** | |
|
||||||
|
| negative control: same top / follows the description / right vs gold | **10 / 122 / 14** | bench 10 / 122 / 14 |
|
||||||
|
| 401 without or with a wrong token (both POSTs); `/health` open | 401 / 401; 200 | |
|
||||||
|
| largest accepted request: 64 decisions × 16 options, 4 calls × 7,168 tok | **200** ×3 | |
|
||||||
|
| one question at 7,168 tok | 200 ×3; one token more (7,171) → **422** "above 7168" | |
|
||||||
|
| any 503 across the whole live acceptance | **none** (server log: 128 + 145 × 200, 18 × 422, 4 × 401) | |
|
||||||
|
|
||||||
|
The live pass is N = 1 for the positive control. It is read against the GPU 3 floor: the same
|
||||||
|
image moved 0 of 560 rows across 3 restarts, and 0 against the bench's 4 restarts.
|
||||||
|
|
||||||
## Building and deploying
|
## Building and deploying
|
||||||
|
|
||||||
|
|||||||
+20
-9
@@ -1,14 +1,25 @@
|
|||||||
# semif
|
# semif
|
||||||
|
|
||||||
> ⚠ **OFFLINE since 2026-09-30 0135 PT (Prime: "take semif offline for now").** It was
|
> ⚠ **REPLACED by intern-decision (Prime, 2026-09-30).** Prime ruled at ~0510 PT: "replace semif
|
||||||
> stopped (`docker compose stop`, NOT removed) to give scriberr its GPU 1 headroom back:
|
> with intern-decision now". The replacement is `stacks/intern-decision` (Intern-Decision-4B, live since 0941 PT,
|
||||||
> scriberr's Parakeet path cuts audio into 5-minute slices and needs over 6 GB, and with
|
> `http://10.251.50.54:8033`, `intern-decision.fv.internal`). It keeps this service's HTTP surface
|
||||||
> SemIf resident it had ~6.7 GB and hit CUDA OOM on a 35-minute file. Stopping it moved
|
> (`/decide`, `/decide/shared`, `/health`), so callers only change the URL and the token
|
||||||
> GPU 1 from 91,052 to 81,806 MiB used. The image, weights, config and token are all kept;
|
> (`secret get intern-decision/api-token`). Its README lists every deliberate difference.
|
||||||
> `unless-stopped` keeps it down across a reboot. **Do not restart it without Prime's word.**
|
>
|
||||||
> To bring it back: `cd /opt/docker/compose/semif && docker compose start`, but first
|
> The `semif` container has been **stopped, not removed**, since 0135 PT that day, when it was
|
||||||
> make sure scriberr has its room (the durable fix, shortening scriberr's slice length, is
|
> taken offline to give scriberr its GPU 1 headroom back. Its image (`semif-serve:0.1.4`),
|
||||||
> deferred by Prime to later). Everything below describes the service as it was deployed.
|
> weights, config and token are all kept, and `unless-stopped` keeps it down across a reboot.
|
||||||
|
> Port 8032 and `semif.fv.internal` stay reserved for it.
|
||||||
|
>
|
||||||
|
> **Rollback, only on Prime's word.**
|
||||||
|
> 1. Stop intern-decision first. The two services do not fit GPU 1 together next to scriberr.
|
||||||
|
> semif held 9.2 GB at rest and peaked at 12.9 GB.
|
||||||
|
> `cd /opt/docker/compose/intern-decision && docker compose stop`.
|
||||||
|
> 2. Start semif: `cd /opt/docker/compose/semif && docker compose start`.
|
||||||
|
> 3. Check nvidia-smi `Free` on GPU 1 against semif's 12.9 GB peak plus scriberr's 5.5 GB (its
|
||||||
|
> 120 s slices) before calling it done. If it does not cover both, semif collides with scriberr.
|
||||||
|
>
|
||||||
|
> Everything below describes the service as it was deployed.
|
||||||
|
|
||||||
**SemIf option-logit decisions** on **fv-ml1 GPU 1**, the utility card beside
|
**SemIf option-logit decisions** on **fv-ml1 GPU 1**, the utility card beside
|
||||||
`vllm-coder`, the erp/meromero seats and scriberr. Deployed 2026-09-27 at Prime's
|
`vllm-coder`, the erp/meromero seats and scriberr. Deployed 2026-09-27 at Prime's
|
||||||
|
|||||||
Reference in New Issue
Block a user