feat(intern-decision): stack, DNS and GPU 3 acceptance for the SemIf replacement

stacks/intern-decision: compose (GPU 1, :8033, hard VRAM cap as the single .env knob,
healthcheck, Homepage group 'AI - Eval & Retrieval'), .env.example and README.
dns: intern-decision.fv.internal -> fv-ml1 (synced to ana/esh/nh3).
acceptance on fv-ml1 GPU 3, 3 fresh processes: bit-identical to the Jev bench's native rows
(pooled 240/259, Wyrd 79/84, 0/560 flips, Δp 0), negative control 10/122/14, 0 flips across
restarts; largest accepted request 200 at a 10,134 MiB card peak under a 9.25 GiB cap; 503 and
recovery proven at a tight cap. GPU 1 deploy held: nvidia-smi Free on GPU 1 is 15,442 MiB.
This commit is contained in:
vh
2026-09-30 09:38:00 -07:00
parent 21d16d7ad8
commit a262477a61
59 changed files with 11969 additions and 1 deletions
@@ -0,0 +1,22 @@
"""Does GPU memory outside torch's allocator grow with the number of distinct threads that run a call?"""
import json, subprocess, sys, threading, time, urllib.request
URL = "http://127.0.0.1:18033"
TOKEN = open("token").read().strip()
BODY = json.dumps({"id": "x", "state": "The deploy passed.", "question": "Did it pass?",
"options": [{"id": "yes", "description": "Yes"}, {"id": "no", "description": "No"}]}).encode()
def card():
return int(subprocess.check_output(["nvidia-smi", "-i", "3", "--query-gpu=memory.used", "--format=csv,noheader,nounits"]).strip())
def reserved():
return json.load(urllib.request.urlopen(URL + "/health"))["model"]["reserved_gib"] * 1024
def post():
r = urllib.request.Request(URL + "/decide", data=BODY, headers={"Authorization": "Bearer " + TOKEN, "Content-Type": "application/json"})
urllib.request.urlopen(r).read()
def snap(tag):
time.sleep(1.5); c, r = card(), reserved(); print(f"{tag}: card {c} MiB, reserved {r:.0f} MiB, outside allocator {c - 2 - r:.0f} MiB", flush=True)
snap("start")
for i in range(20): post()
snap("after 20 sequential requests")
for burst in range(3):
ts = [threading.Thread(target=post) for _ in range(40)]
[t.start() for t in ts]; [t.join() for t in ts]
snap(f"after concurrent burst {burst + 1} (40 requests)")