docs: Jev candidate bench vs SemIf (fv-ml1 GPU 3) — Intern-Decision-4B is the replacement if SemIf is displaced

Operator ask relayed by brokkr-smithy-dev. Positive control (SemIf 187/231,
hard 0.613) reproduced exactly; negative control and a 4-restart noise floor
(0 flips) measured. On our replaced-baseline sets no candidate beats SemIf-with-
rotations beyond the ~4-pt floor; Intern-Decision-4B native matches it at one
ordering, is better on Wyrd, fits 9.7/10.3 GB and is 1.5-2.3x faster. JevBench
rank does not transfer. Raw per-item data kept out of git.
This commit is contained in:
vh
2026-09-30 05:00:45 -07:00
parent 9a6ac59da7
commit 475d6d6bcb
27 changed files with 19068 additions and 0 deletions
@@ -0,0 +1,22 @@
"""Run a Python module or script under a hard per-process VRAM cap, the way semif-serve applies
SEMIF_VRAM_CAP_GIB (torch.cuda.set_per_process_memory_fraction BEFORE any weights load).
BENCH_VRAM_CAP_GIB=12 python capped.py <module-or-script.py> [args...]
Unset or 0 = no cap. The cap covers torch's allocator only, as in semif-serve; the CUDA context
(~0.5-0.9 GiB) sits outside it, so nvidia-smi reads higher than the cap would suggest."""
import os
import runpy
import sys
import torch
cap = float(os.environ.get("BENCH_VRAM_CAP_GIB") or 0)
if cap:
total = torch.cuda.get_device_properties(0).total_memory
torch.cuda.set_per_process_memory_fraction(cap * 2**30 / total, 0)
print(f"[capped] per-process VRAM cap {cap} GiB of {total / 2**30:.1f} GiB", flush=True)
target, sys.argv = sys.argv[1], sys.argv[1:]
if target.endswith(".py"):
sys.path.insert(0, os.path.dirname(os.path.abspath(target)))
runpy.run_path(target, run_name="__main__")
else:
runpy.run_module(target, run_name="__main__", alter_sys=True)