Files
esh-pfi-infrastructure/services/intern-decision-serve/acceptance/gpu3-2026-09-30/probe/health-end.json
T
vh a262477a61 feat(intern-decision): stack, DNS and GPU 3 acceptance for the SemIf replacement
stacks/intern-decision: compose (GPU 1, :8033, hard VRAM cap as the single .env knob,
healthcheck, Homepage group 'AI - Eval & Retrieval'), .env.example and README.
dns: intern-decision.fv.internal -> fv-ml1 (synced to ana/esh/nh3).
acceptance on fv-ml1 GPU 3, 3 fresh processes: bit-identical to the Jev bench's native rows
(pooled 240/259, Wyrd 79/84, 0/560 flips, Δp 0), negative control 10/122/14, 0 flips across
restarts; largest accepted request 200 at a 10,134 MiB card peak under a 9.25 GiB cap; 503 and
recovery proven at a tight cap. GPU 1 deploy held: nvidia-smi Free on GPU 1 is 15,442 MiB.
2026-09-30 09:38:00 -07:00

1 line
1.0 KiB
JSON

{"status":"ok","model":{"name":"Intern-Decision-4B","source":"internlm/Intern-Decision-4B","revision":"0e5e6aa7d6d750e2b1504ba11a8136cb58aeb3cd","checkpoint":"/hf/hub/models--internlm--Intern-Decision-4B/snapshots/0e5e6aa7d6d750e2b1504ba11a8136cb58aeb3cd","inference_py_sha256":"c904e2c67ca0775621a22375ee373d2ba30b52117cda870c6c9ef74143b29863","temperature":1.99241824,"dtype":"bfloat16","attn_implementation":"sdpa","device":"cuda","max_length":8192,"torch_version":"2.10.0+cu128","transformers_version":"5.17.0","vision_tower":"removed","device_name":"NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition","allocated_gib":7.937,"reserved_gib":7.969,"max_reserved_gib":8.861},"vram_cap_gib":9.25,"max_tokens":8192,"max_decisions":64,"max_questions_per_call":16,"chunking":"/decide/shared questions are packed greedily, in request order, into calls of at most 16 (1-16, 17-32, ...); each call is one prompt, so the questions in a call are asked together. With orderings, ordering k of every decision forms wave k, packed the same way.","workloads":[]}