Files
esh-pfi-infrastructure/services/intern-decision-serve/acceptance/systemone-2026-09-30/health-end.json
T
vh ff552abf1f feat(intern-decision-serve): 0.1.1 adds POST /v1/systemone (Jev wire shape)
Straight passthrough to the checkpoint's own DecisionEngine.predict — never the
semif mapping, whose different prompt would change the answers. Reuses the one
inference thread, bearer auth, MAX_QUEUE, VRAM cap and error envelope; no new
concurrency. 1..16 questions in ONE call (never chunked: Jev questions share a
prompt); images 422; over MAX_TOKENS 422 before the forward. /health advertises
the surface. Response 'model' is a string name@revision (JevBench's runner
hashes it; a dict broke its manifest step).

Acceptance on the live service (see acceptance/systemone-2026-09-30/): JevBench
v1.2.16 typesafe adapter over the 231 public items scores all 202/231, hard
83/111, with 0 changed answers across all 924 rows of the bench's own r1..r4;
controls 401/422x3 (token boundary proven at 7168 pass / 7169 refuse); GPU 1
per-process peak 9,866 MiB under the largest accepted request (budget 9,876);
/decide/shared positive control unchanged. 120 tests green.
2026-09-30 12:58:19 -07:00

1 line
1.2 KiB
JSON

{"status":"ok","model":{"name":"Intern-Decision-4B","source":"internlm/Intern-Decision-4B","revision":"0e5e6aa7d6d750e2b1504ba11a8136cb58aeb3cd","checkpoint":"/hf/hub/models--internlm--Intern-Decision-4B/snapshots/0e5e6aa7d6d750e2b1504ba11a8136cb58aeb3cd","inference_py_sha256":"c904e2c67ca0775621a22375ee373d2ba30b52117cda870c6c9ef74143b29863","temperature":1.99241824,"dtype":"bfloat16","attn_implementation":"sdpa","device":"cuda","max_length":7168,"torch_version":"2.10.0+cu128","transformers_version":"5.17.0","vision_tower":"removed","device_name":"NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition","allocated_gib":7.937,"reserved_gib":7.969,"max_reserved_gib":8.861},"vram_cap_gib":9.0,"max_tokens":7168,"max_decisions":64,"max_questions_per_call":16,"chunking":"/decide/shared questions are packed greedily, in request order, into calls of at most 16 (1-16, 17-32, ...); each call is one prompt, so the questions in a call are asked together. With orderings, ordering k of every decision forms wave k, packed the same way.","workloads":[],"endpoints":["/decide","/decide/shared","/v1/systemone"],"systemone":{"max_questions":16,"chunking":"none","images":"not supported","max_tokens":7168}}