Latency, measured from nh3-dev (3 runs x 20 per condition; network floor 31 ms): - /decide short: 71 ms end to end, 38 ms server-side; - /decide with a ~2,000-token state: 210 / 169 ms; - shared, 3 rotations: 113 / 79 ms; - shared, 6 orderings: 137 / 99 ms. Qwen3.5's fast kernels (causal_conv1d, flash-linear-attention) are not installed, so transformers falls back to its reference PyTorch paths. That is a speed lever, and using it needs a parity re-check. Averaging over option orderings, on SemIf authored144 + perturbations108 (252 rows, 72 groups): - a single ordering scores 78.6%; - log-mean over the 3 rotations scores 87.7% (+9.1 pts, group-bootstrap 95% CI +4.7 to +13.8); - all 6 permutations score 88.1%. Rotations capture nearly all of the gain. Rows where the rotations agree unanimously (161) are 94.4% accurate; split rows (91) are 75.8%.
59 lines
3.0 KiB
Python
59 lines
3.0 KiB
Python
"""semif-serve latency, 2026-09-27. Sequential client on nh3-dev, fresh connection per request,
|
|
3 runs x 20 timed after 3 warm-ups per condition, conditions interleaved per run.
|
|
e2e = wall time around the request; srv = what SemIf reports (direct: total_seconds; shared:
|
|
timing.total_seconds), i.e. no network or HTTP.
|
|
SEMIF_URL=... SEMIF_TOKEN=... uv run --with httpx python latency.py
|
|
"""
|
|
import os, statistics as st, time, httpx, itertools
|
|
|
|
U, H = os.environ["SEMIF_URL"], {"Authorization": f"Bearer {os.environ['SEMIF_TOKEN']}"}
|
|
OPTS = [{"id": "casual_outing", "description": "A casual outing"},
|
|
{"id": "romantic_date", "description": "A romantic date"},
|
|
{"id": "booty_call", "description": "A booty call"}]
|
|
CALL = "She calls up and says, hey, what're you doing right now? It's 2AM and I'm bored."
|
|
LONG = "The service logged a routine heartbeat from node alpha at the scheduled interval without incident. " * 120
|
|
Q = "What kind of invitation is this?"
|
|
BIN = [{"id": "yes", "description": "Yes"}, {"id": "no", "description": "No"}]
|
|
|
|
|
|
def rot(opts):
|
|
return [opts[i:] + opts[:i] for i in range(len(opts))]
|
|
|
|
|
|
CONDS = {
|
|
"health (floor)": ("GET", "/health", None),
|
|
"decide, short (~130 tok)": ("POST", "/decide", {"id": "a", "state": CALL, "question": Q, "options": OPTS}),
|
|
"decide, long (~2,000 tok)": ("POST", "/decide", {"id": "b", "state": LONG, "question": "Was there an incident?", "options": BIN}),
|
|
"shared, 3 rotations (short)": ("POST", "/decide/shared", {"state": CALL, "decisions": [
|
|
{"id": f"r{i}", "question": Q, "options": o} for i, o in enumerate(rot(OPTS))]}),
|
|
"shared, 6 orderings (short)": ("POST", "/decide/shared", {"state": CALL, "decisions": [
|
|
{"id": f"p{i}", "question": Q, "options": list(o)} for i, o in enumerate(itertools.permutations(OPTS))]}),
|
|
"shared, 3 rotations (long)": ("POST", "/decide/shared", {"state": LONG, "decisions": [
|
|
{"id": f"r{i}", "question": "Was there an incident?", "options": o} for i, o in enumerate(rot(BIN + [{"id": "unsure", "description": "Cannot tell"}]))]}),
|
|
}
|
|
|
|
|
|
def once(method, path, body):
|
|
t = time.perf_counter()
|
|
r = httpx.request(method, U + path, json=body, headers=H, timeout=120)
|
|
r.raise_for_status()
|
|
e2e = (time.perf_counter() - t) * 1000
|
|
j = r.json()
|
|
srv = (j["timing"]["total_seconds"] if "timing" in j else j.get("total_seconds")) if path != "/health" else None
|
|
return e2e, (srv * 1000 if srv is not None else None)
|
|
|
|
|
|
res = {k: {"e2e": [], "srv": []} for k in CONDS}
|
|
for run in range(3):
|
|
for name, (m, p, b) in CONDS.items():
|
|
for _ in range(3):
|
|
once(m, p, b)
|
|
e, s = zip(*(once(m, p, b) for _ in range(20)))
|
|
res[name]["e2e"].append(st.median(e))
|
|
if s[0] is not None:
|
|
res[name]["srv"].append(st.median(s))
|
|
print(f"{'condition':<30} {'e2e p50 (runs)':<28} {'server p50 (runs)'}")
|
|
for name, r in res.items():
|
|
fmt = lambda xs: f"{st.median(xs):6.1f} [{min(xs):.1f}-{max(xs):.1f}]" if xs else "-"
|
|
print(f"{name:<30} {fmt(r['e2e']):<28} {fmt(r['srv'])}")
|