docs(intern-decision): live on fv-ml1 GPU 1; semif marked REPLACED
intern-decision deployed 0941 PT (cap 9.0 GiB, MAX_TOKENS 7168). Live acceptance: positive control 240/259 and Wyrd 79/84, bit-identical to the bench (0/560 rows, Δp 0); negative control 10/122/14; largest accepted requests 200 with no 503; per-process 8,812 MiB at rest and 9,866 peak; GPU 1 Free 15,442 before and 6,581 after (lowest 5,569 under load). Latency from nh3-dev: 21 criteria 114 ms, 16 over ~3,900 tokens 238 ms. semif README banner now REPLACED with the rollback; fv-ml1 GPU 1 note updated.
This commit is contained in:
+2
-2
@@ -127,5 +127,5 @@ aliases:
|
||||
- {name: wherethef, site: nh3, target: nh3-dev, note: WhereTF :8093}
|
||||
- {name: homepage, site: esh, target: esh-docker-vm, note: fleet dashboard :5100}
|
||||
- {name: scriberr, site: fv, target: fv-ml1, note: transcription + diarization :8080 (GPU1)}
|
||||
- {name: semif, site: fv, target: fv-ml1, note: SemIf option-logit decisions :8032 (GPU1) — stopped since 2026-09-30, being replaced by intern-decision (Prime); kept as rollback}
|
||||
- {name: intern-decision, site: fv, target: fv-ml1, note: Intern-Decision-4B typed decisions :8033 (GPU1), semif-serve API; replaces semif (Prime 2026-09-30)}
|
||||
- {name: semif, site: fv, target: fv-ml1, note: SemIf option-logit decisions :8032 (GPU1) — REPLACED by intern-decision 2026-09-30 (Prime), stopped, kept as rollback}
|
||||
- {name: intern-decision, site: fv, target: fv-ml1, note: Intern-Decision-4B typed decisions :8033 (GPU1), semif-serve API; replaced semif 2026-09-30 (Prime)}
|
||||
|
||||
@@ -181,9 +181,15 @@ embed/rerank/reward trio. GPUs are pinned per container via
|
||||
|
||||
**GPU 1 — light / eval / retrieval + char-RP GGUF (~91/98 GB, on-demand):**
|
||||
|
||||
> ⚠ **This table is stale (checked 2026-09-27).** Live GPU 1 residents were `scriberr`,
|
||||
> `vllm-coder`, `vllm-erp-seat` and `vllm-meromero-rp`, plus **`semif`** (**OFFLINE since 2026-09-30, Prime; see `stacks/semif`**; :8032, ~8.7 GB resting when up,
|
||||
> hard-capped at 12 GiB; `stacks/semif`, since 2026-09-27). Read the host
|
||||
> ⚠ **This table is stale (checked 2026-09-27).** Live GPU 1 residents are `scriberr`,
|
||||
> `vllm-coder`, `vllm-erp-seat` and `vllm-meromero-rp`.
|
||||
> **`intern-decision`** joined them on 2026-09-30, 0941 PT: :8033, 8,812 MiB at rest, 9,866 MiB
|
||||
> peak, hard-capped at 9.0 GiB with `MAX_TOKENS` 7,168; see `stacks/intern-decision`. It replaced
|
||||
> **`semif`** (:8032), which is stopped and kept as the rollback, per Prime's ruling of 2026-09-30.
|
||||
> **GPU 1 budget:** nvidia-smi `Free` read 15,442 MiB before intern-decision and 6,581 MiB after,
|
||||
> at rest. That covers scriberr's 5,496 MiB peak even while intern-decision is at its own peak.
|
||||
> Measure nvidia-smi `Free` (the driver reserves 640 MiB per card) before adding anything to
|
||||
> this card. Read the host
|
||||
> (`docker inspect … DeviceRequests`), not this table. (A fixtures-only `augaman` instance ran
|
||||
> here for about an hour on 2026-09-27 for a speed bench, and was then removed on Prime's call.)
|
||||
|
||||
|
||||
@@ -0,0 +1,139 @@
|
||||
{
|
||||
"url": "http://intern-decision.fv.internal:8033",
|
||||
"started_utc": "2026-09-30T16:43:39Z",
|
||||
"health": {
|
||||
"status": "ok",
|
||||
"model": {
|
||||
"name": "Intern-Decision-4B",
|
||||
"source": "internlm/Intern-Decision-4B",
|
||||
"revision": "0e5e6aa7d6d750e2b1504ba11a8136cb58aeb3cd",
|
||||
"checkpoint": "/hf/hub/models--internlm--Intern-Decision-4B/snapshots/0e5e6aa7d6d750e2b1504ba11a8136cb58aeb3cd",
|
||||
"inference_py_sha256": "c904e2c67ca0775621a22375ee373d2ba30b52117cda870c6c9ef74143b29863",
|
||||
"temperature": 1.99241824,
|
||||
"dtype": "bfloat16",
|
||||
"attn_implementation": "sdpa",
|
||||
"device": "cuda",
|
||||
"max_length": 7168,
|
||||
"torch_version": "2.10.0+cu128",
|
||||
"transformers_version": "5.17.0",
|
||||
"vision_tower": "removed",
|
||||
"device_name": "NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition",
|
||||
"allocated_gib": 7.937,
|
||||
"reserved_gib": 7.969,
|
||||
"max_reserved_gib": 8.861
|
||||
},
|
||||
"vram_cap_gib": 9.0,
|
||||
"max_tokens": 7168,
|
||||
"max_decisions": 64,
|
||||
"max_questions_per_call": 16,
|
||||
"chunking": "/decide/shared questions are packed greedily, in request order, into calls of at most 16 (1-16, 17-32, ...); each call is one prompt, so the questions in a call are asked together. With orderings, ordering k of every decision forms wave k, packed the same way.",
|
||||
"workloads": []
|
||||
},
|
||||
"auth": {
|
||||
"/decide": {
|
||||
"no_token": 401,
|
||||
"wrong_token": 401,
|
||||
"right_token": 200
|
||||
},
|
||||
"/decide/shared": {
|
||||
"no_token": 401,
|
||||
"wrong_token": 401,
|
||||
"right_token": 200
|
||||
},
|
||||
"health_no_token": 200,
|
||||
"pass": true,
|
||||
"t_start": 1790786619.1466322,
|
||||
"t_end": 1790786619.3572643
|
||||
},
|
||||
"maxreq": {
|
||||
"state_chars": 4793,
|
||||
"tokens_per_call": 7168,
|
||||
"decisions": 64,
|
||||
"options_each": 16,
|
||||
"body_bytes": 96250,
|
||||
"runs": [
|
||||
{
|
||||
"status": 200,
|
||||
"e2e_ms": 1542.5,
|
||||
"code": null,
|
||||
"calls": 4,
|
||||
"input_tokens": [
|
||||
7168,
|
||||
7168,
|
||||
7168,
|
||||
7168
|
||||
],
|
||||
"reserved_gib_after": 7.969,
|
||||
"max_reserved_gib": 8.996,
|
||||
"t_end": 1790786706.9185867
|
||||
},
|
||||
{
|
||||
"status": 200,
|
||||
"e2e_ms": 1518.7,
|
||||
"code": null,
|
||||
"calls": 4,
|
||||
"input_tokens": [
|
||||
7168,
|
||||
7168,
|
||||
7168,
|
||||
7168
|
||||
],
|
||||
"reserved_gib_after": 7.969,
|
||||
"max_reserved_gib": 8.996,
|
||||
"t_end": 1790786708.4557755
|
||||
},
|
||||
{
|
||||
"status": 200,
|
||||
"e2e_ms": 1539.6,
|
||||
"code": null,
|
||||
"calls": 4,
|
||||
"input_tokens": [
|
||||
7168,
|
||||
7168,
|
||||
7168,
|
||||
7168
|
||||
],
|
||||
"reserved_gib_after": 7.969,
|
||||
"max_reserved_gib": 8.996,
|
||||
"t_end": 1790786710.014339
|
||||
}
|
||||
],
|
||||
"t_start": 1790786619.357301,
|
||||
"t_end": 1790786710.014397
|
||||
},
|
||||
"maxone": {
|
||||
"state_chars": 28727,
|
||||
"tokens": 7168,
|
||||
"one_more_is": [
|
||||
422,
|
||||
"Example has 7171 tokens, above 7168; truncation is forbidden"
|
||||
],
|
||||
"runs": [
|
||||
{
|
||||
"status": 200,
|
||||
"e2e_ms": 394.2,
|
||||
"input_tokens": 7168,
|
||||
"code": null,
|
||||
"t_end": 1790786715.324555
|
||||
},
|
||||
{
|
||||
"status": 200,
|
||||
"e2e_ms": 375.9,
|
||||
"input_tokens": 7168,
|
||||
"code": null,
|
||||
"t_end": 1790786715.7005756
|
||||
},
|
||||
{
|
||||
"status": 200,
|
||||
"e2e_ms": 372.7,
|
||||
"input_tokens": 7168,
|
||||
"code": null,
|
||||
"t_end": 1790786716.0734687
|
||||
}
|
||||
],
|
||||
"pass": true,
|
||||
"t_start": 1790786710.0144272,
|
||||
"t_end": 1790786716.0735004
|
||||
},
|
||||
"finished_utc": "2026-09-30T16:45:16Z"
|
||||
}
|
||||
@@ -0,0 +1,671 @@
|
||||
{
|
||||
"ours": {
|
||||
"live": {
|
||||
"scores": {
|
||||
"single/pooled": [
|
||||
240,
|
||||
259
|
||||
],
|
||||
"single/authored144": [
|
||||
132,
|
||||
144
|
||||
],
|
||||
"single/perturbations108": [
|
||||
100,
|
||||
108
|
||||
],
|
||||
"single/cicada-w1": [
|
||||
29,
|
||||
31
|
||||
],
|
||||
"single/cicada-w2": [
|
||||
22,
|
||||
31
|
||||
],
|
||||
"single/wyrd": [
|
||||
79,
|
||||
84
|
||||
],
|
||||
"single/wyrd:place": [
|
||||
19,
|
||||
21
|
||||
],
|
||||
"single/wyrd:place2": [
|
||||
21,
|
||||
21
|
||||
],
|
||||
"single/wyrd:exit": [
|
||||
19,
|
||||
21
|
||||
],
|
||||
"single/wyrd:exit2": [
|
||||
20,
|
||||
21
|
||||
],
|
||||
"single/failures": 0
|
||||
},
|
||||
"negative": {
|
||||
"n": 144,
|
||||
"same_top_as_unrotated": 10,
|
||||
"follows_description": 122,
|
||||
"vs_original_gold": 14
|
||||
},
|
||||
"a_vs_a_in_process_authored144": {
|
||||
"rows": 0,
|
||||
"top_differs": 0,
|
||||
"flipped_ids": [],
|
||||
"max_dp": null
|
||||
}
|
||||
}
|
||||
},
|
||||
"bench": {
|
||||
"r1": {
|
||||
"scores": {
|
||||
"single/pooled": [
|
||||
240,
|
||||
259
|
||||
],
|
||||
"single/authored144": [
|
||||
132,
|
||||
144
|
||||
],
|
||||
"single/perturbations108": [
|
||||
100,
|
||||
108
|
||||
],
|
||||
"single/cicada-w1": [
|
||||
29,
|
||||
31
|
||||
],
|
||||
"single/cicada-w2": [
|
||||
22,
|
||||
31
|
||||
],
|
||||
"single/wyrd": [
|
||||
79,
|
||||
84
|
||||
],
|
||||
"single/wyrd:place": [
|
||||
19,
|
||||
21
|
||||
],
|
||||
"single/wyrd:place2": [
|
||||
21,
|
||||
21
|
||||
],
|
||||
"single/wyrd:exit": [
|
||||
19,
|
||||
21
|
||||
],
|
||||
"single/wyrd:exit2": [
|
||||
20,
|
||||
21
|
||||
],
|
||||
"single/failures": 0,
|
||||
"rotations/pooled": [
|
||||
236,
|
||||
259
|
||||
],
|
||||
"rotations/authored144": [
|
||||
130,
|
||||
144
|
||||
],
|
||||
"rotations/perturbations108": [
|
||||
99,
|
||||
108
|
||||
],
|
||||
"rotations/cicada-w1": [
|
||||
29,
|
||||
31
|
||||
],
|
||||
"rotations/cicada-w2": [
|
||||
22,
|
||||
31
|
||||
],
|
||||
"rotations/wyrd": [
|
||||
77,
|
||||
84
|
||||
],
|
||||
"rotations/wyrd:place": [
|
||||
18,
|
||||
21
|
||||
],
|
||||
"rotations/wyrd:place2": [
|
||||
21,
|
||||
21
|
||||
],
|
||||
"rotations/wyrd:exit": [
|
||||
19,
|
||||
21
|
||||
],
|
||||
"rotations/wyrd:exit2": [
|
||||
19,
|
||||
21
|
||||
],
|
||||
"rotations/failures": 0
|
||||
},
|
||||
"negative": {
|
||||
"n": 144,
|
||||
"same_top_as_unrotated": 10,
|
||||
"follows_description": 122,
|
||||
"vs_original_gold": 14
|
||||
}
|
||||
},
|
||||
"r2": {
|
||||
"scores": {
|
||||
"single/pooled": [
|
||||
240,
|
||||
259
|
||||
],
|
||||
"single/authored144": [
|
||||
132,
|
||||
144
|
||||
],
|
||||
"single/perturbations108": [
|
||||
100,
|
||||
108
|
||||
],
|
||||
"single/cicada-w1": [
|
||||
29,
|
||||
31
|
||||
],
|
||||
"single/cicada-w2": [
|
||||
22,
|
||||
31
|
||||
],
|
||||
"single/wyrd": [
|
||||
79,
|
||||
84
|
||||
],
|
||||
"single/wyrd:place": [
|
||||
19,
|
||||
21
|
||||
],
|
||||
"single/wyrd:place2": [
|
||||
21,
|
||||
21
|
||||
],
|
||||
"single/wyrd:exit": [
|
||||
19,
|
||||
21
|
||||
],
|
||||
"single/wyrd:exit2": [
|
||||
20,
|
||||
21
|
||||
],
|
||||
"single/failures": 0,
|
||||
"rotations/pooled": [
|
||||
236,
|
||||
259
|
||||
],
|
||||
"rotations/authored144": [
|
||||
130,
|
||||
144
|
||||
],
|
||||
"rotations/perturbations108": [
|
||||
99,
|
||||
108
|
||||
],
|
||||
"rotations/cicada-w1": [
|
||||
29,
|
||||
31
|
||||
],
|
||||
"rotations/cicada-w2": [
|
||||
22,
|
||||
31
|
||||
],
|
||||
"rotations/wyrd": [
|
||||
77,
|
||||
84
|
||||
],
|
||||
"rotations/wyrd:place": [
|
||||
18,
|
||||
21
|
||||
],
|
||||
"rotations/wyrd:place2": [
|
||||
21,
|
||||
21
|
||||
],
|
||||
"rotations/wyrd:exit": [
|
||||
19,
|
||||
21
|
||||
],
|
||||
"rotations/wyrd:exit2": [
|
||||
19,
|
||||
21
|
||||
],
|
||||
"rotations/failures": 0,
|
||||
"multifield/pooled": [
|
||||
77,
|
||||
84
|
||||
],
|
||||
"multifield/wyrd": [
|
||||
77,
|
||||
84
|
||||
],
|
||||
"multifield/wyrd:place": [
|
||||
16,
|
||||
21
|
||||
],
|
||||
"multifield/wyrd:place2": [
|
||||
21,
|
||||
21
|
||||
],
|
||||
"multifield/wyrd:exit": [
|
||||
20,
|
||||
21
|
||||
],
|
||||
"multifield/wyrd:exit2": [
|
||||
20,
|
||||
21
|
||||
],
|
||||
"multifield/failures": 0
|
||||
},
|
||||
"negative": {
|
||||
"n": 144,
|
||||
"same_top_as_unrotated": 10,
|
||||
"follows_description": 122,
|
||||
"vs_original_gold": 14
|
||||
}
|
||||
},
|
||||
"r3": {
|
||||
"scores": {
|
||||
"single/pooled": [
|
||||
240,
|
||||
259
|
||||
],
|
||||
"single/authored144": [
|
||||
132,
|
||||
144
|
||||
],
|
||||
"single/perturbations108": [
|
||||
100,
|
||||
108
|
||||
],
|
||||
"single/cicada-w1": [
|
||||
29,
|
||||
31
|
||||
],
|
||||
"single/cicada-w2": [
|
||||
22,
|
||||
31
|
||||
],
|
||||
"single/wyrd": [
|
||||
79,
|
||||
84
|
||||
],
|
||||
"single/wyrd:place": [
|
||||
19,
|
||||
21
|
||||
],
|
||||
"single/wyrd:place2": [
|
||||
21,
|
||||
21
|
||||
],
|
||||
"single/wyrd:exit": [
|
||||
19,
|
||||
21
|
||||
],
|
||||
"single/wyrd:exit2": [
|
||||
20,
|
||||
21
|
||||
],
|
||||
"single/failures": 0,
|
||||
"rotations/pooled": [
|
||||
236,
|
||||
259
|
||||
],
|
||||
"rotations/authored144": [
|
||||
130,
|
||||
144
|
||||
],
|
||||
"rotations/perturbations108": [
|
||||
99,
|
||||
108
|
||||
],
|
||||
"rotations/cicada-w1": [
|
||||
29,
|
||||
31
|
||||
],
|
||||
"rotations/cicada-w2": [
|
||||
22,
|
||||
31
|
||||
],
|
||||
"rotations/wyrd": [
|
||||
77,
|
||||
84
|
||||
],
|
||||
"rotations/wyrd:place": [
|
||||
18,
|
||||
21
|
||||
],
|
||||
"rotations/wyrd:place2": [
|
||||
21,
|
||||
21
|
||||
],
|
||||
"rotations/wyrd:exit": [
|
||||
19,
|
||||
21
|
||||
],
|
||||
"rotations/wyrd:exit2": [
|
||||
19,
|
||||
21
|
||||
],
|
||||
"rotations/failures": 0,
|
||||
"multifield/pooled": [
|
||||
77,
|
||||
84
|
||||
],
|
||||
"multifield/wyrd": [
|
||||
77,
|
||||
84
|
||||
],
|
||||
"multifield/wyrd:place": [
|
||||
16,
|
||||
21
|
||||
],
|
||||
"multifield/wyrd:place2": [
|
||||
21,
|
||||
21
|
||||
],
|
||||
"multifield/wyrd:exit": [
|
||||
20,
|
||||
21
|
||||
],
|
||||
"multifield/wyrd:exit2": [
|
||||
20,
|
||||
21
|
||||
],
|
||||
"multifield/failures": 0
|
||||
},
|
||||
"negative": {
|
||||
"n": 144,
|
||||
"same_top_as_unrotated": 10,
|
||||
"follows_description": 122,
|
||||
"vs_original_gold": 14
|
||||
}
|
||||
},
|
||||
"r4": {
|
||||
"scores": {
|
||||
"single/pooled": [
|
||||
240,
|
||||
259
|
||||
],
|
||||
"single/authored144": [
|
||||
132,
|
||||
144
|
||||
],
|
||||
"single/perturbations108": [
|
||||
100,
|
||||
108
|
||||
],
|
||||
"single/cicada-w1": [
|
||||
29,
|
||||
31
|
||||
],
|
||||
"single/cicada-w2": [
|
||||
22,
|
||||
31
|
||||
],
|
||||
"single/wyrd": [
|
||||
79,
|
||||
84
|
||||
],
|
||||
"single/wyrd:place": [
|
||||
19,
|
||||
21
|
||||
],
|
||||
"single/wyrd:place2": [
|
||||
21,
|
||||
21
|
||||
],
|
||||
"single/wyrd:exit": [
|
||||
19,
|
||||
21
|
||||
],
|
||||
"single/wyrd:exit2": [
|
||||
20,
|
||||
21
|
||||
],
|
||||
"single/failures": 0,
|
||||
"rotations/pooled": [
|
||||
236,
|
||||
259
|
||||
],
|
||||
"rotations/authored144": [
|
||||
130,
|
||||
144
|
||||
],
|
||||
"rotations/perturbations108": [
|
||||
99,
|
||||
108
|
||||
],
|
||||
"rotations/cicada-w1": [
|
||||
29,
|
||||
31
|
||||
],
|
||||
"rotations/cicada-w2": [
|
||||
22,
|
||||
31
|
||||
],
|
||||
"rotations/wyrd": [
|
||||
77,
|
||||
84
|
||||
],
|
||||
"rotations/wyrd:place": [
|
||||
18,
|
||||
21
|
||||
],
|
||||
"rotations/wyrd:place2": [
|
||||
21,
|
||||
21
|
||||
],
|
||||
"rotations/wyrd:exit": [
|
||||
19,
|
||||
21
|
||||
],
|
||||
"rotations/wyrd:exit2": [
|
||||
19,
|
||||
21
|
||||
],
|
||||
"rotations/failures": 0,
|
||||
"multifield/pooled": [
|
||||
77,
|
||||
84
|
||||
],
|
||||
"multifield/wyrd": [
|
||||
77,
|
||||
84
|
||||
],
|
||||
"multifield/wyrd:place": [
|
||||
16,
|
||||
21
|
||||
],
|
||||
"multifield/wyrd:place2": [
|
||||
21,
|
||||
21
|
||||
],
|
||||
"multifield/wyrd:exit": [
|
||||
20,
|
||||
21
|
||||
],
|
||||
"multifield/wyrd:exit2": [
|
||||
20,
|
||||
21
|
||||
],
|
||||
"multifield/failures": 0
|
||||
},
|
||||
"negative": {
|
||||
"n": 144,
|
||||
"same_top_as_unrotated": 10,
|
||||
"follows_description": 122,
|
||||
"vs_original_gold": 14
|
||||
}
|
||||
}
|
||||
},
|
||||
"vs_bench_row_by_row": {
|
||||
"live": {
|
||||
"single": {
|
||||
"bench_repeat": "r1",
|
||||
"rows": 560,
|
||||
"top_differs": 0,
|
||||
"flipped_ids": [],
|
||||
"max_dp": 0.0
|
||||
},
|
||||
"rotations": {
|
||||
"bench_repeat": "r1",
|
||||
"rows": 0,
|
||||
"top_differs": 0,
|
||||
"flipped_ids": [],
|
||||
"max_dp": null
|
||||
},
|
||||
"negative": {
|
||||
"bench_repeat": "r1",
|
||||
"rows": 144,
|
||||
"top_differs": 0,
|
||||
"flipped_ids": [],
|
||||
"max_dp": 0.0
|
||||
},
|
||||
"multifield": {
|
||||
"bench_repeat": "r2",
|
||||
"rows": 0,
|
||||
"top_differs": 0,
|
||||
"flipped_ids": [],
|
||||
"max_dp": null
|
||||
}
|
||||
}
|
||||
},
|
||||
"across_restarts": {},
|
||||
"summary": {
|
||||
"single/authored144": {
|
||||
"ours_median": 132,
|
||||
"ours_all": [
|
||||
132
|
||||
],
|
||||
"bench_all": [
|
||||
132,
|
||||
132,
|
||||
132,
|
||||
132
|
||||
],
|
||||
"n": 144
|
||||
},
|
||||
"single/cicada-w1": {
|
||||
"ours_median": 29,
|
||||
"ours_all": [
|
||||
29
|
||||
],
|
||||
"bench_all": [
|
||||
29,
|
||||
29,
|
||||
29,
|
||||
29
|
||||
],
|
||||
"n": 31
|
||||
},
|
||||
"single/cicada-w2": {
|
||||
"ours_median": 22,
|
||||
"ours_all": [
|
||||
22
|
||||
],
|
||||
"bench_all": [
|
||||
22,
|
||||
22,
|
||||
22,
|
||||
22
|
||||
],
|
||||
"n": 31
|
||||
},
|
||||
"single/perturbations108": {
|
||||
"ours_median": 100,
|
||||
"ours_all": [
|
||||
100
|
||||
],
|
||||
"bench_all": [
|
||||
100,
|
||||
100,
|
||||
100,
|
||||
100
|
||||
],
|
||||
"n": 108
|
||||
},
|
||||
"single/pooled": {
|
||||
"ours_median": 240,
|
||||
"ours_all": [
|
||||
240
|
||||
],
|
||||
"bench_all": [
|
||||
240,
|
||||
240,
|
||||
240,
|
||||
240
|
||||
],
|
||||
"n": 259
|
||||
},
|
||||
"single/wyrd": {
|
||||
"ours_median": 79,
|
||||
"ours_all": [
|
||||
79
|
||||
],
|
||||
"bench_all": [
|
||||
79,
|
||||
79,
|
||||
79,
|
||||
79
|
||||
],
|
||||
"n": 84
|
||||
},
|
||||
"single/wyrd:exit": {
|
||||
"ours_median": 19,
|
||||
"ours_all": [
|
||||
19
|
||||
],
|
||||
"bench_all": [
|
||||
19,
|
||||
19,
|
||||
19,
|
||||
19
|
||||
],
|
||||
"n": 21
|
||||
},
|
||||
"single/wyrd:exit2": {
|
||||
"ours_median": 20,
|
||||
"ours_all": [
|
||||
20
|
||||
],
|
||||
"bench_all": [
|
||||
20,
|
||||
20,
|
||||
20,
|
||||
20
|
||||
],
|
||||
"n": 21
|
||||
},
|
||||
"single/wyrd:place": {
|
||||
"ours_median": 19,
|
||||
"ours_all": [
|
||||
19
|
||||
],
|
||||
"bench_all": [
|
||||
19,
|
||||
19,
|
||||
19,
|
||||
19
|
||||
],
|
||||
"n": 21
|
||||
},
|
||||
"single/wyrd:place2": {
|
||||
"ours_median": 21,
|
||||
"ours_all": [
|
||||
21
|
||||
],
|
||||
"bench_all": [
|
||||
21,
|
||||
21,
|
||||
21,
|
||||
21
|
||||
],
|
||||
"n": 21
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,6 @@
|
||||
1790786561.567928123 shape:start
|
||||
1790786613.903648987 shape:done
|
||||
1790786619.093843341 checks:start
|
||||
1790786716.079809795 checks:done
|
||||
1790786739.656804839 sets:start
|
||||
1790786809.130379221 sets:done
|
||||
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because it is too large
Load Diff
@@ -1,10 +1,7 @@
|
||||
# intern-decision
|
||||
|
||||
> ⚠ **STATUS 2026-09-30 0937 PT: built and accepted on GPU 3; the GPU 1 deploy is HELD.** GPU 1's
|
||||
> usable free memory is 15,442 MiB (nvidia-smi `Free`: 97,887 total − 640 driver-reserved −
|
||||
> 81,806 used), below the 15,800 MiB pre-deploy floor. The 10,300 MiB budget was derived from
|
||||
> total − used (16,081). The decision is with infra-ops/Prime; see "VRAM". Nothing is running
|
||||
> on port 8033 yet.
|
||||
> **LIVE since 2026-09-30 0941 PT** on fv-ml1 GPU 1, with cap 9.0 GiB and `MAX_TOKENS` 7,168.
|
||||
> Acceptance on the live service is below. `semif` is stopped and kept as the rollback.
|
||||
|
||||
**Intern-Decision-4B typed decisions** on **fv-ml1 GPU 1**, the utility card beside `vllm-coder`,
|
||||
the erp/meromero seats and scriberr. It **replaced `semif`** on Prime's ruling of 2026-09-30,
|
||||
@@ -76,48 +73,58 @@ curl -s -H "Authorization: Bearer $T" http://intern-decision.fv.internal:8033/de
|
||||
- `/health.semif_commit`;
|
||||
- per-workload calibration. Any `workload` is a 422, exactly as the deployed semif behaved
|
||||
with its empty table.
|
||||
5. **`MAX_TOKENS` (8192) is per call:** the state plus all its questions. A longer call is a 422
|
||||
and is never truncated. Every call up to that limit fits the VRAM cap (measured). See VRAM
|
||||
below.
|
||||
5. **`MAX_TOKENS` is 7,168 per call:** the state plus all its questions. It is lower than the
|
||||
model's own 8,192 so that every call the API accepts fits the VRAM cap.
|
||||
- The count is taken **before** the forward pass, so a longer call is a clear
|
||||
`422 invalid_request` ("Example has N tokens, above 7168; truncation is forbidden"). It is
|
||||
never truncated.
|
||||
- A request with more decisions is split into more calls, and each call must fit.
|
||||
|
||||
## VRAM: the whole container ≤ 10,300 MiB
|
||||
## VRAM: fits beside scriberr's peak, whatever the request
|
||||
|
||||
**Budget (infra-ops, 2026-09-30):** the container's **whole** nvidia-smi footprint, CUDA context
|
||||
included, must never exceed **10,300 MiB**, whatever the request. That leaves scriberr (peak
|
||||
5,496 MiB with its 120 s slices) 295 MiB of spare even when both hit their peaks at once.
|
||||
`torch.cuda.set_per_process_memory_fraction` caps only torch's allocator, so the cap is the budget
|
||||
minus the measured non-allocator overhead:
|
||||
**Budget (infra-ops, 2026-09-30):** GPU 1 needs nvidia-smi `Free` ≥ **15,400 MiB** before this
|
||||
service starts. That is our card peak of 9,876 MiB plus scriberr's peak of 5,496 MiB (with its
|
||||
120 s slices), rounded up. So both can peak at the same moment. Read nvidia-smi's own `Free`, not
|
||||
total − used: the driver reserves 640 MiB on every card. The deploy at 0940 read 15,442.
|
||||
|
||||
`torch.cuda.set_per_process_memory_fraction` caps only torch's allocator. The two knobs in `.env`
|
||||
are therefore chosen together and change together:
|
||||
|
||||
```
|
||||
footprint <= VRAM_CAP_GIB + overhead = 9,472 MiB (9.25 GiB) + 662 MiB = 10,134 MiB (166 MiB under budget)
|
||||
card footprint <= VRAM_CAP_GIB 9.0 (9,216 MiB) + 660 MiB outside the allocator (measured) = 9,876 MiB
|
||||
MAX_TOKENS 7168 = the largest call measured to fit under that cap, so no accepted call reaches the 503
|
||||
```
|
||||
|
||||
`VRAM_CAP_GIB` in `.env` is **the single knob**. It is applied before the weights load.
|
||||
| on the live service, GPU 1 (per-process nvidia-smi, 0.1 s sampling, 2,939 samples over 316 s of acceptance load) | MiB |
|
||||
|---|---|
|
||||
| **at rest** after startup | **8,812** |
|
||||
| at rest after the acceptance (512 MiB release slack keeps a little cache) | 8,856 |
|
||||
| **peak**: largest requests, 64 decisions × 16 options, 4 calls × 7,168 tokens, N = 3 | **9,866**; torch `max_reserved` 8.996 of 9.0 GiB |
|
||||
| GPU 1 `Free` before the deploy / at rest after / lowest during the acceptance | 15,442 / 6,581 / 5,569, which stays above scriberr's 5,496 |
|
||||
|
||||
⚠ **The budget assumed 16,081 MiB free on GPU 1. That is total − used.** nvidia-smi's own `Free` is
|
||||
**15,442 MiB**, because the driver reserves 640 MiB on every card (GPU 3 shows the same 640).
|
||||
Against the true free memory, a footprint that never collides with scriberr's 5,496 MiB peak must
|
||||
stay at or below **9,946 MiB**:
|
||||
- ⚠ **At 7,168 tokens the allocator reaches the cap with about 4 MiB to spare.** It fits because
|
||||
the allocator frees its cache and retries before it fails.
|
||||
- Measured: every call up to the limit answered 200. That is 145 shared calls on the live
|
||||
service, including the size search near the boundary, 3 × 4 calls at 7,168 tokens with
|
||||
16 questions, and 3 at 7,168 tokens with one question. No 503 was seen.
|
||||
- The allocator's cache state depends on request history. If a max-size call ever does return
|
||||
503, lower `MAX_TOKENS`. Do not raise the cap.
|
||||
- Raising either knob needs a fresh measurement (`acceptance/checks.py --checks fit,maxreq`) and
|
||||
the budget re-checked against scriberr.
|
||||
- **Over the cap the answer is 503 `out_of_memory`.** Memory returns to the resting baseline and
|
||||
the service keeps answering. This was proven on GPU 3 with a deliberately tight cap.
|
||||
|
||||
| cap | footprint peak (measured, GPU 3) | with scriberr at its peak, against 15,442 free | calls that fit (16 questions × 16 options) |
|
||||
|---|---|---|---|
|
||||
| **9.25 GiB** (`.env.example` now) | 10,134 MiB | **188 MiB over** | every call the API accepts (8,191 tok) |
|
||||
| **9.0 GiB** | 9,876 MiB | 70 MiB spare | up to 7,168 tok; longer calls get 503. Every bench shape fits, at the same latency |
|
||||
Measured on GPU 3 before the deploy (alone on the card), which is where the knobs came from:
|
||||
|
||||
| measured on fv-ml1 GPU 3 (alone on the card), image 0.1.0 | MiB |
|
||||
|---|---|
|
||||
| **at rest** (nvidia-smi, whole card less 2 MiB idle) | **8,820** (torch reserved 8,160 + outside 660) |
|
||||
| outside the allocator (CUDA context + loaded kernels), at rest / after 2×30 concurrent requests / after the latency shapes / after 4 × 8,191-token calls | 660 / 660 / 662 / 662 |
|
||||
| **peak, largest request the API accepts, capped at 9.25 GiB** (card, 0.1 s sampling) | **10,134**; torch `max_reserved` = 9.25 GiB exactly, so the cap is what held it |
|
||||
| the same request **uncapped**, for comparison | 10,422 (allocator peak 9,760): **over budget, so the cap is required** |
|
||||
| card peak at cap 9.0 GiB (fit search and a 64 × 8,191-token request, which got 503) | **9,876** |
|
||||
| peak with 4 × 8,191-token calls at cap 9.25 GiB (answered 200) / uncapped | 10,134 / 10,422 (allocator 9,760) |
|
||||
| largest 16-question call that answers 200: at cap 9.0 / at 9.25 GiB | 7,168 / 8,191 tokens |
|
||||
| startup peak (weights + vision tower until the swap + warm-up) | allocator 8.86 GiB. **A cap below ~8.9 GiB cannot start** (it fails closed; 8.5 was refused) |
|
||||
|
||||
- **The largest request the API accepts** is 64 decisions × 16 options, with each of its 4 calls
|
||||
at 8,191 tokens. Under the cap it answers **200** (1.67 s), not 503. Near the limit the
|
||||
allocator frees cached blocks before it fails, so the uncapped 9,760 MiB peak was partly cache.
|
||||
Every call up to `MAX_TOKENS` fits the cap.
|
||||
- **Over the cap the answer is 503 `out_of_memory`.** Memory goes back to the resting baseline
|
||||
and the service keeps answering. This was proven with a deliberately tight cap (acceptance).
|
||||
- **One inference thread is load-bearing** (contract INV-2). torch keeps CUDA state per host
|
||||
thread (cuBLAS handles and workspaces). The first build scored on anyio's 40-thread pool: a
|
||||
concurrent burst added 252 MiB outside the cap and 326 MiB inside it, and the peak reached
|
||||
@@ -132,19 +139,22 @@ stay at or below **9,946 MiB**:
|
||||
|
||||
## Latency
|
||||
|
||||
Loopback on fv-ml1 (GPU 3 alone on the card, image 0.1.0, cap 9.25 GiB). `bench_shape.py`
|
||||
from the Jev bench ran 3 runs × 20 requests, and the medians are shown with the run-median
|
||||
range. The bench's native Intern-Decision numbers are the reference.
|
||||
**Live, from nh3-dev** (0942 PT). The URL is `intern-decision.fv.internal:8033`, and the round trip
|
||||
is 22.9 ms on average (19.5–27.7 ms over 10 pings). GPU 1 is shared with the vLLM seats and
|
||||
scriberr. The harness is `bench_shape.py` from the Jev bench, 3 runs × 20 requests; each cell is
|
||||
the median with the run-median range. "server" is the service's own time.
|
||||
|
||||
| shape | end to end | server | bench (native, loopback) | SemIf 0.1.4 (bench) |
|
||||
| shape | end to end from nh3-dev | server | GPU 3 loopback before deploy | SemIf 0.1.4 (README, from nh3-dev) |
|
||||
|---|---|---|---|---|
|
||||
| 1 decision, short (~180 tok) | 36.4 ms (36.4–36.5) | 34.7 | 39 | 37 |
|
||||
| **21 binary criteria**, one `/decide/shared` (2 calls: 16 + 5) | **82.3 ms** (82.3–82.4) | 78.9 | 88 | 131 |
|
||||
| 1 decision over the ~3,900-token state | 191.4 ms (191.4–191.5) | 187.9 | 190 | 186 |
|
||||
| **16 criteria over the ~3,900-token state** (1 call, 4,579 tok) | **216.8 ms** (216.5–217.0) | 211.1 | 215 | 505 |
|
||||
| largest accepted request (4 calls × 8,191 tok) | 1,672 ms (1,672–1,675, N=3) | | | |
|
||||
| 1 decision, short (~180 tok) | 53.9 ms (53.3–54.2) | 35.2 | 36.4 | 69 |
|
||||
| **21 binary criteria**, one `/decide/shared` (2 calls: 16 + 5) | **114.3 ms** (114.1–114.3) | 80.3 | 82.3 | 159 |
|
||||
| 1 decision over the ~3,900-token state | 211.2 ms (210.8–211.6) | 186.8 | 191.4 | |
|
||||
| **16 criteria over the ~3,900-token state** (1 call, 4,579 tok) | **238.3 ms** (237.8–238.7) | 205.3 | 216.8 | (bench loopback: 505) |
|
||||
| largest request (4 calls × 7,168 tok), N = 3 | 1,519–1,543 ms | | | |
|
||||
|
||||
*From nh3-dev and on GPU 1, next to the vLLM seats: pending the GPU 1 deploy (held, see the status banner).*
|
||||
Server-side time on GPU 1 matches GPU 3 within a few ms, so the vLLM neighbours were not
|
||||
measurably slowing it at 0942. The rest of the end-to-end time is the network (~23 ms round trip
|
||||
plus the HTTP exchange).
|
||||
|
||||
## Acceptance (2026-09-30)
|
||||
|
||||
@@ -153,6 +163,8 @@ directory per process lifetime). **The harness is the bench's own.** `bench_sets
|
||||
--backend semif` speaks semif-serve's API, so it drove this service unchanged. `compare.py`
|
||||
scores it as the bench's `analyze.py` did.
|
||||
|
||||
### On GPU 3 before the deploy (the harness and the floor)
|
||||
|
||||
| check (GPU 3, cap 9.25 GiB, N = 3 fresh processes unless noted) | result | bench / floor |
|
||||
|---|---|---|
|
||||
| **positive control**, pooled 259, single ordering | **240, 240, 240** | bench 240 (4 repeats); the pooled set resolves ±4 pts |
|
||||
@@ -166,12 +178,28 @@ scores it as the bench's `analyze.py` did.
|
||||
| more than 16 questions: 20 decisions in one request, against the same decisions asked as 16 + 4 | 2 calls `[16, 4]`, all 20 tops equal, Δp 0, same prompt hashes | |
|
||||
| 401: no token / wrong token, both POSTs; `/health` open | 401 / 401; 200 | |
|
||||
| 429: 48 concurrent requests, `MAX_QUEUE` 32 | 32 answered, 16 × 429 `busy` | |
|
||||
| largest request the API accepts (64 × 16 options, 4 × 8,191 tok) at cap 9.25 GiB | **200** ×3; card peak 10,134 MiB | budget 10,300 |
|
||||
| largest request at `MAX_TOKENS` 8,192 (64 × 16 options, 4 × 8,191 tok), cap 9.25 GiB | **200** ×3; card peak 10,134 MiB | superseded by the 9.0 / 7,168 knobs (VRAM) |
|
||||
| **over the cap** (the same request at a deliberately tight cap of 8.9 GiB) | **503 `out_of_memory`**; reserved 7.969 GiB before, 7.969 after, next request 200; then all 560 single rows answered, bit-identical to the bench | |
|
||||
| fit under the cap: the largest 16-question call that answers 200 | 8,191 tok (the whole API) at 9.25 GiB; 7,168 tok at 9.0 GiB | |
|
||||
| cap below startup | 8.5 GiB: startup refused (OOM loading), fails closed | |
|
||||
|
||||
*On the deployed service (GPU 1): pending the deploy (held, see the status banner).*
|
||||
### On the live service (GPU 1, 0942–0946 PT, through `http://intern-decision.fv.internal:8033` from nh3-dev)
|
||||
|
||||
Raw results: `services/intern-decision-serve/acceptance/gpu1-live/`.
|
||||
|
||||
| check | result | reference |
|
||||
|---|---|---|
|
||||
| **positive control**, pooled 259, single ordering | **240** | bench 240; GPU 3 240 ×3 |
|
||||
| positive control, Wyrd /84 | **79** | bench 79 |
|
||||
| row by row against the bench's native rows (560 single + 144 negative) | **0 top changes, max Δp 0.000** | |
|
||||
| negative control: same top / follows the description / right vs gold | **10 / 122 / 14** | bench 10 / 122 / 14 |
|
||||
| 401 without or with a wrong token (both POSTs); `/health` open | 401 / 401; 200 | |
|
||||
| largest accepted request: 64 decisions × 16 options, 4 calls × 7,168 tok | **200** ×3 | |
|
||||
| one question at 7,168 tok | 200 ×3; one token more (7,171) → **422** "above 7168" | |
|
||||
| any 503 across the whole live acceptance | **none** (server log: 128 + 145 × 200, 18 × 422, 4 × 401) | |
|
||||
|
||||
The live pass is N = 1 for the positive control. It is read against the GPU 3 floor: the same
|
||||
image moved 0 of 560 rows across 3 restarts, and 0 against the bench's 4 restarts.
|
||||
|
||||
## Building and deploying
|
||||
|
||||
|
||||
+20
-9
@@ -1,14 +1,25 @@
|
||||
# semif
|
||||
|
||||
> ⚠ **OFFLINE since 2026-09-30 0135 PT (Prime: "take semif offline for now").** It was
|
||||
> stopped (`docker compose stop`, NOT removed) to give scriberr its GPU 1 headroom back:
|
||||
> scriberr's Parakeet path cuts audio into 5-minute slices and needs over 6 GB, and with
|
||||
> SemIf resident it had ~6.7 GB and hit CUDA OOM on a 35-minute file. Stopping it moved
|
||||
> GPU 1 from 91,052 to 81,806 MiB used. The image, weights, config and token are all kept;
|
||||
> `unless-stopped` keeps it down across a reboot. **Do not restart it without Prime's word.**
|
||||
> To bring it back: `cd /opt/docker/compose/semif && docker compose start`, but first
|
||||
> make sure scriberr has its room (the durable fix, shortening scriberr's slice length, is
|
||||
> deferred by Prime to later). Everything below describes the service as it was deployed.
|
||||
> ⚠ **REPLACED by intern-decision (Prime, 2026-09-30).** Prime ruled at ~0510 PT: "replace semif
|
||||
> with intern-decision now". The replacement is `stacks/intern-decision` (Intern-Decision-4B, live since 0941 PT,
|
||||
> `http://10.251.50.54:8033`, `intern-decision.fv.internal`). It keeps this service's HTTP surface
|
||||
> (`/decide`, `/decide/shared`, `/health`), so callers only change the URL and the token
|
||||
> (`secret get intern-decision/api-token`). Its README lists every deliberate difference.
|
||||
>
|
||||
> The `semif` container has been **stopped, not removed**, since 0135 PT that day, when it was
|
||||
> taken offline to give scriberr its GPU 1 headroom back. Its image (`semif-serve:0.1.4`),
|
||||
> weights, config and token are all kept, and `unless-stopped` keeps it down across a reboot.
|
||||
> Port 8032 and `semif.fv.internal` stay reserved for it.
|
||||
>
|
||||
> **Rollback, only on Prime's word.**
|
||||
> 1. Stop intern-decision first. The two services do not fit GPU 1 together next to scriberr.
|
||||
> semif held 9.2 GB at rest and peaked at 12.9 GB.
|
||||
> `cd /opt/docker/compose/intern-decision && docker compose stop`.
|
||||
> 2. Start semif: `cd /opt/docker/compose/semif && docker compose start`.
|
||||
> 3. Check nvidia-smi `Free` on GPU 1 against semif's 12.9 GB peak plus scriberr's 5.5 GB (its
|
||||
> 120 s slices) before calling it done. If it does not cover both, semif collides with scriberr.
|
||||
>
|
||||
> Everything below describes the service as it was deployed.
|
||||
|
||||
**SemIf option-logit decisions** on **fv-ml1 GPU 1**, the utility card beside
|
||||
`vllm-coder`, the erp/meromero seats and scriberr. Deployed 2026-09-27 at Prime's
|
||||
|
||||
Reference in New Issue
Block a user