feat(semif): 0.1.3 — order averaging, fast kernels, bug-hunt hardening (Prime)

Order averaging (Prime, after the 739aa03 spike):
- A decision may set orderings: rotations|all (all only for <= 4 options). Every
  ordering goes to the engine in one shared batch.
- The reply keeps each native result and adds combined {probabilities (log-mean),
  top, agreement, spread}.
- Through the service on SemIf's labelled sets (252 rows): 78.6% -> 88.1%
  (group-bootstrap 95% CI +5.1..+14.3). Unanimous agreement is 94.5% accurate.

Fast kernels: flash-linear-attention 0.5.2 and causal-conv1d 1.7.0 are now the
default build. A/B on the empty GPU 3:
- parity with upstream went from 142/144 to 144/144;
- a ~2k-token /decide went from 169 to 92 ms server-side;
- short 3-rotation batches cost ~3-6 ms more.
triton builds a C shim at runtime, so the image carries gcc. Without it the
warm-up failed and startup failed closed.

Heid bug-hunt panel (4/4 arms, thread 01M3H3F4RR7XBP90KQ3A39H4SX), folded:
- Startup validation: VRAM cap 0 no longer means uncapped (C1); limits must be
  >= 1 (S1); the token must be visible ASCII (S2); the calibration file must
  exist and parse, with T in [0.05, 20] (S8, and C3's NaN leg).
- The body limit is checked before a chunk is kept, and a Unicode-digit
  Content-Length no longer crashes (C2, S3).
- Failures while building the response now get the 500 envelope (C3).
- 429 busy past SEMIF_MAX_QUEUE requests in progress (C6).
- The engine releases memory on every non-validation failure, unchained after
  gc; an empty OOM message is handled; 'out of memory' RuntimeErrors map to 503
  (C4, C5, S9).
- The entry point forces HF_HUB_OFFLINE (S10). README wording fixed (S5, S6).
- New guard tests close the gaps the arms' mutation grids exposed: early stop of
  the body read, a shared-route lock, calibration pass-through, the gc cycle,
  the exact caps, TorchEngine.load's arch and device checks, and the offline
  entry point.
86 tests.

Deployed on fv-ml1 GPU 1: parity 144/144, OOM and burst release verified, shared
capacity 63/51/26/16 rows at ~140/520/1960/3900 prefix tokens.
This commit is contained in:
vh
2026-09-27 03:27:15 -07:00
parent d7ad235365
commit 77b8cb449c
22 changed files with 1018 additions and 92 deletions
+12 -3
View File
@@ -22,18 +22,26 @@ l.write_text(re.sub(r'(name = "semif-serve"\nversion = )"[^"]+"', r'\1"0.0.0"',
EOF
FROM python:3.12-slim-bookworm
# EXTRAS picks the optional dependency sets. Default (adopted 2026-09-27): the model plus
# Qwen3.5's fast kernels (fla + causal-conv1d). "--extra model" alone gives the reference
# PyTorch paths.
ARG EXTRAS="--extra model --extra fast"
COPY --from=ghcr.io/astral-sh/uv:0.6.9 /uv /bin/uv
ENV UV_COMPILE_BYTECODE=1 UV_LINK_MODE=copy UV_PYTHON_DOWNLOADS=never
# git: semif-phase1 installs from a pinned GitHub commit.
# git: semif-phase1 installs from a pinned GitHub commit. The fast extra also needs a C
# compiler AT RUNTIME: triton builds its CUDA driver shim on first use, and without gcc the
# warm-up dies with "Failed to find C compiler" (seen 2026-09-27; startup failed closed).
RUN apt-get update && apt-get install -y --no-install-recommends git ca-certificates \
&& if echo "$EXTRAS" | grep -q -- '--extra fast'; then \
apt-get install -y --no-install-recommends gcc libc6-dev; fi \
&& rm -rf /var/lib/apt/lists/*
WORKDIR /app
COPY --from=deps /deps/pyproject.toml /deps/uv.lock ./
RUN --mount=type=cache,target=/root/.cache/uv \
uv sync --frozen --no-dev --extra model --no-install-project
uv sync --frozen --no-dev $EXTRAS --no-install-project
COPY pyproject.toml uv.lock ./
COPY src ./src
RUN uv sync --frozen --no-dev --extra model --no-editable --no-cache
RUN uv sync --frozen --no-dev $EXTRAS --no-editable --no-cache
RUN groupadd --system --gid 10001 semif \
&& useradd --system --uid 10001 --gid 10001 --no-create-home --shell /usr/sbin/nologin semif
USER semif
@@ -41,6 +49,7 @@ ENV PATH=/app/.venv/bin:$PATH \
HF_HOME=/hf \
HF_HUB_OFFLINE=1 \
HF_HUB_DISABLE_TELEMETRY=1 \
TRITON_CACHE_DIR=/tmp/triton-cache \
NVIDIA_DRIVER_CAPABILITIES=compute,utility
EXPOSE 8000
# One worker (INV-2): the model and the inference lock live in this one process.
@@ -0,0 +1,18 @@
{
"rows": 252,
"groups": 72,
"accuracy_plain": 0.7857,
"accuracy_rotations": 0.881,
"delta_95ci": [
0.0512,
0.1429
],
"unanimous": {
"rows": 163,
"accuracy": 0.9448
},
"split": {
"rows": 89,
"accuracy": 0.764
}
}
@@ -0,0 +1,42 @@
"""Order-averaging acceptance through the SERVICE (0.1.3): the same 252 labelled rows as the spike
(SemIf authored144 + perturbations108), each scored twice: plain /decide (the caller's order), and
/decide with orderings=rotations (combined.top). Paired group bootstrap for the accuracy delta.
SEMIF_DIR=... SEMIF_URL=... SEMIF_TOKEN=... uv run --with httpx python averaging.py out.json
"""
import json, os, random, statistics as st, sys
from pathlib import Path
import httpx
S, U = Path(os.environ["SEMIF_DIR"]), os.environ["SEMIF_URL"]
H = {"Authorization": f"Bearer {os.environ['SEMIF_TOKEN']}"}
rows = []
for name in ("authored144", "perturbations108"):
rows += [dict(json.loads(l), set=name) for l in (S / f"benchmarks/data/{name}.jsonl").read_text().splitlines() if l.strip()]
out = []
with httpx.Client(timeout=120) as c:
for r in rows:
base = {k: r[k] for k in ("id", "state", "question", "options")}
plain = c.post(f"{U}/decide", headers=H, json=base).json()
avg = c.post(f"{U}/decide", headers=H, json={**base, "orderings": "rotations"}).json()
gold = r["options"][r["label"]]["id"]
top_plain = plain["option_ids"][plain["probabilities"].index(max(plain["probabilities"]))]
out.append({"group": r["group_id"], "plain": top_plain == gold, "rotations": avg["combined"]["top"] == gold,
"agreement": avg["combined"]["agreement"]})
groups = {}
for o in out:
groups.setdefault(o["group"], []).append(o)
rng, keys, deltas = random.Random(7), list(groups), []
for _ in range(10000):
sample = [o for g in (rng.choice(keys) for _ in keys) for o in groups[g]]
deltas.append(st.fmean(o["rotations"] for o in sample) - st.fmean(o["plain"] for o in sample))
deltas.sort()
unan = [o for o in out if o["agreement"] == 1.0]
split = [o for o in out if o["agreement"] < 1.0]
report = {"rows": len(out), "groups": len(groups),
"accuracy_plain": round(st.fmean(o["plain"] for o in out), 4),
"accuracy_rotations": round(st.fmean(o["rotations"] for o in out), 4),
"delta_95ci": [round(deltas[250], 4), round(deltas[9750], 4)],
"unanimous": {"rows": len(unan), "accuracy": round(st.fmean(o["rotations"] for o in unan), 4)},
"split": {"rows": len(split), "accuracy": round(st.fmean(o["rotations"] for o in split), 4) if split else None}}
json.dump(report, open(sys.argv[1], "w"), indent=1)
print(json.dumps(report))
@@ -0,0 +1,63 @@
{
"url": "http://10.251.50.54:8032",
"health": {
"status": "ok",
"semif_commit": "23cf1f39fc9534fe81437200959b6dfc7106e45a",
"model": {
"source": "Qwen/Qwen3.5-4B",
"revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a",
"dtype": "bfloat16",
"device": "cuda:0",
"torch_version": "2.10.0+cu128",
"transformers_version": "5.17.0",
"device_name": "NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition",
"allocated_gib": 7.84,
"reserved_gib": 8.12
},
"vram_cap_gib": 12.0,
"max_tokens": 4096,
"max_decisions": 64,
"workloads": []
},
"1_parity_vs_upstream": {
"rows": 144,
"top_choice_agree": 144,
"max_abs_prob_gap": 0.049096847865749194
},
"1_prompt_sha256_equal": 144,
"2_noise_floor_a_vs_b": {
"rows": 144,
"top_choice_agree": 144,
"max_abs_prob_gap": 0.0
},
"3_negative_rotated_options": {
"rows": 144,
"top_choice_agree": 14,
"max_abs_prob_gap": 0.9985836128543327
},
"4_shared_vs_direct": {
"groups": 36,
"rows": 72,
"top_choice_agree": 72,
"max_abs_prob_gap": 0.046346781311727814
},
"5_speed_21_binary": {
"prefix_tokens": 62,
"shared_s": {
"runs": [
0.139055563005968,
0.1375489159981953,
0.13897942200128455
],
"median": 0.13897942200128455
},
"sequential_decide_s": {
"runs": [
0.9357447380025405,
0.9384197259932989,
0.939685705001466
],
"median": 0.9384197259932989
}
}
}
+8 -1
View File
@@ -1,6 +1,6 @@
[project]
name = "semif-serve"
version = "0.1.2"
version = "0.1.3"
description = "HTTP wrapper around SemIf's direct and shared option-logit scorers"
requires-python = ">=3.12"
dependencies = [
@@ -17,6 +17,13 @@ model = [
"torch==2.10.0",
]
# Qwen3.5's fast kernels. Without them transformers runs its reference PyTorch paths
# ("correct but much slower"). Trialled 2026-09-27; adopted only if parity with upstream holds.
fast = [
"flash-linear-attention==0.5.2",
"causal-conv1d @ https://github.com/Dao-AILab/causal-conv1d/releases/download/v1.7.0/causal_conv1d-1.7.0+cu12torch2.10cxx11abiTRUE-cp312-cp312-linux_x86_64.whl ; sys_platform == 'linux' and platform_machine == 'x86_64'",
]
[dependency-groups]
dev = ["pytest==8.4.2", "httpx==0.28.1"]
+66 -6
View File
@@ -45,6 +45,37 @@ probabilities}` with `softmax(option_logits / T)`. The argmax never changes. The
native `probabilities` and `probability_status` stay untouched. An unknown
`workload` is a 422. With no `workload`, no `calibrated` key appears.
**Order averaging (0.1.3, Prime 2026-09-27).** A decision (the `/decide` body, or an
entry in `decisions`) may set `orderings`, whose default is `"none"`:
- `"rotations"`: the n cyclic shifts of the caller's option list, starting with
the caller's order, so every option sits in every position exactly once.
- `"all"`: every permutation (n!), the caller's order first. It is allowed only
when n ≤ 4, else 422.
Every ordering of every decision in the request becomes its own SemIf row
(`id` = `"<id>#o<k>"`, same state and question, reordered options). **All rows go
to the engine in ONE `shared` call**, which includes `/decide`. The result for an
averaged decision is:
```
{id, option_ids (caller's order),
combined: {method, orderings: n, probabilities, top, agreement, spread: {option_id: [min, max]}},
orderings: [the SemIf result for each ordering, unchanged]}
```
- `probabilities`: per ordering, log-softmax of `option_logits`; averaged per
option id; renormalised; reported in the caller's order.
- `agreement`: the fraction of orderings whose top option equals `combined.top`.
- `spread`: each option's min and max native probability across orderings.
Every expanded row counts toward `SEMIF_MAX_DECISIONS`. `workload` together with
`orderings` is a 422, because a temperature is fitted per method and none is
fitted on combined scores yet. Decisions without `orderings` keep the exact
pre-0.1.3 result shape. Averaging cancels any additive position bias exactly.
Measured by the 2026-09-27 spike: 3 rotations take SemIf's labelled sets from
78.6% to 87.7% accuracy, and unanimous agreement is 94.4% accurate.
## Invariants
- **INV-1 pass-through.** `option_ids`, `probabilities`, `option_logits`,
@@ -64,10 +95,18 @@ native `probabilities` and `probability_status` stay untouched. An unknown
baseline by more than 512 MiB, the engine calls `empty_cache()`. A burst must not
keep the card's shared headroom: on 2026-09-27 a 64-decision request left the
process holding 12.6 GB, leaving scriberr 3.5 GB.
**Any other scorer failure except `ValueError`** also releases before it is
reported: it is logged with its traceback, then re-raised unchained as
`ScoringFailed` after `gc.collect()` + `empty_cache()` (bug hunt C5). A
`RuntimeError` whose message says "out of memory" (cuBLAS/cuDNN allocation
failures) counts as an OOM → 503 (S9). An OOM with an empty message is reported
as "CUDA out of memory" rather than crashing the handler (C4).
- **INV-5 no network at runtime.** Weights come from the mounted HF cache at the
pinned revision (`HF_HUB_OFFLINE=1`).
pinned revision. The entry point sets `HF_HUB_OFFLINE=1` itself before torch or
transformers load, so this holds outside the image too (S10).
- **INV-6 constant-time auth.** Token comparison uses `hmac.compare_digest`. The
token is ≥ 32 characters, and startup refuses a shorter one.
token is ≥ 32 characters of visible ASCII (33–126). Startup refuses anything
else, because a CR, LF or NUL in the token can never arrive in a header (S2).
## Limits and errors
@@ -75,15 +114,21 @@ native `probabilities` and `probability_status` stay untouched. An unknown
422, never truncated (SemIf raises).
- `SEMIF_MAX_DECISIONS` (default 64) caps `decisions` per shared request. It must
hold 1..max entries, else 422.
- Request body ≤ `SEMIF_MAX_BODY_BYTES` (default 1 MiB), else 413.
- Request body ≤ `SEMIF_MAX_BODY_BYTES` (default 1 MiB), else 413. The limit is
checked **before** each chunk is kept, so no more than the limit is ever held
(C2). A declared `Content-Length` is trusted only if it is ASCII digits (S3).
- **Admission:** at most `SEMIF_MAX_QUEUE` (default 32) POSTs may be in progress,
counting both queued and scoring. The next one is refused with 429 `busy`
before its body is read (C6).
| status | code | when |
|---|---|---|
| 401 | `unauthorized` | missing or wrong bearer |
| 413 | `request_too_large` | body over the limit |
| 422 | `invalid_request` | bad JSON shape, a SemIf `ValueError` (validation, token limit, tokenisation), unknown workload, too many decisions |
| 429 | `busy` | `SEMIF_MAX_QUEUE` requests already in progress |
| 503 | `out_of_memory` | CUDA OOM during scoring |
| 500 | `scoring_failed` | any other scorer exception |
| 500 | `scoring_failed` | any other scorer exception, **or a failure while building the response** from a scorer result (calibration, averaging): always the envelope, never a bare 500 (C3) |
The error body is `{error: {code, message}}`.
@@ -92,7 +137,16 @@ The error body is `{error: {code, message}}`.
`SEMIF_API_TOKEN` (required), `SEMIF_MODEL` (default `Qwen/Qwen3.5-4B`),
`SEMIF_REVISION` (default the pinned SHA), `SEMIF_DEVICE` (default `cuda`),
`SEMIF_VRAM_CAP_GIB`, `SEMIF_MAX_TOKENS`, `SEMIF_MAX_DECISIONS`,
`SEMIF_MAX_BODY_BYTES`, `SEMIF_CALIBRATION` (path to a JSON `{workload: T}`; T > 0).
`SEMIF_MAX_BODY_BYTES`, `SEMIF_MAX_QUEUE`, `SEMIF_CALIBRATION` (path to a JSON
`{workload: T}`).
Startup validates every value and refuses a bad one with a `ValueError` naming
the variable (C1, S1, S8):
- the VRAM cap, when set, is finite and > 0 (`0` used to mean uncapped);
- every limit is an integer ≥ 1;
- each T is a finite number, not a bool, in [0.05, 20] (a tiny T overflowed to
NaN, and the response then failed to render);
- the calibration file must exist and parse.
## Tests (TDD, fake scorer: no torch, no model)
@@ -109,7 +163,13 @@ alone; any
other exception → 500; malformed
JSON or a wrong body shape → 422; too many decisions → 422; an oversized body → 413; requests
are serialised (two concurrent calls never overlap inside the scorer); `/health`
answers while a scorer call is blocked.
answers while a scorer call is blocked. **Averaging:** `rotations` sends n rows
in one shared call, each option once per position, with ids `<id>#o<k>`, and
cancels a position bias exactly; `all` sends n! rows and is 422 above 4 options;
agreement and spread are computed from the orderings; a mixed shared request
(averaged + plain) is one engine call, with results in request order and plain
results unchanged; expanded rows count toward the cap; `workload` + `orderings`
→ 422.
## Acceptance (on fv-ml1, real model; not unit tests)
@@ -0,0 +1,7 @@
condition e2e p50 (runs) server p50 (runs)
health (floor) 33.5 [31.7-34.7] -
decide, short (~130 tok) 69.3 [68.8-70.1] 35.3 [35.3-35.4]
decide, long (~2,000 tok) 130.5 [129.5-130.5] 92.1 [91.7-92.4]
shared, 3 rotations (short) 115.0 [112.4-115.7] 77.9 [77.1-79.3]
shared, 6 orderings (short) 118.3 [117.5-118.7] 81.2 [80.2-81.3]
shared, 3 rotations (long) 199.5 [196.6-201.0] 158.2 [157.7-158.3]
+144 -25
View File
@@ -2,9 +2,10 @@
from __future__ import annotations
import hmac
import itertools
import math
import threading
from typing import Any
from typing import Any, Literal
from fastapi import FastAPI, Request
from fastapi.concurrency import run_in_threadpool
@@ -29,6 +30,7 @@ class Decision(BaseModel):
id: str
question: str
options: list[Option]
orderings: Literal["none", "rotations", "all"] = "none"
def row(self, state: State) -> dict:
"""The SemIf row shape: exactly id, state, question, options."""
@@ -64,6 +66,75 @@ def _first_error(exc: ValidationError) -> str:
return f"{where}: {first.get('msg', 'invalid')}"
MAX_OPTIONS_FOR_ALL = 4
def ordering_perms(decision: Decision) -> list[tuple[int, ...]]:
"""Index permutations of the caller's options, the caller's own order first."""
n = len(decision.options)
if decision.orderings == "rotations":
return [tuple((start + k) % n for k in range(n)) for start in range(n)]
if n > MAX_OPTIONS_FOR_ALL:
raise ApiError(422, "invalid_request",
f"orderings 'all' allows at most {MAX_OPTIONS_FOR_ALL} options ({n} given); use 'rotations'")
return list(itertools.permutations(range(n)))
def expanded_rows(decision: Decision, state: State, perms: list[tuple[int, ...]]) -> list[dict]:
base = decision.row(state)
return [{**base, "id": f"{decision.id}#o{k}", "options": [base["options"][i] for i in perm]}
for k, perm in enumerate(perms)]
def combine(decision: Decision, results: list[dict]) -> dict:
"""Average per-ordering log-softmax by option id; the native results ride along unchanged."""
option_ids = [o.id for o in decision.options]
logp: dict[str, list[float]] = {i: [] for i in option_ids}
probs: dict[str, list[float]] = {i: [] for i in option_ids}
tops = []
for result in results:
logits = result["option_logits"]
top = max(logits)
lse = top + math.log(sum(math.exp(x - top) for x in logits))
for oid, x, p in zip(result["option_ids"], logits, result["probabilities"]):
logp[oid].append(x - lse)
probs[oid].append(p)
tops.append(result["option_ids"][logits.index(top)])
means = [sum(logp[i]) / len(logp[i]) for i in option_ids]
peak = max(means)
weights = [math.exp(m - peak) for m in means]
combined_p = [w / sum(weights) for w in weights]
winner = option_ids[combined_p.index(max(combined_p))]
return {
"id": decision.id,
"option_ids": option_ids,
"combined": {
"method": decision.orderings,
"orderings": len(results),
"probabilities": combined_p,
"top": winner,
"agreement": tops.count(winner) / len(tops),
"spread": {i: [min(probs[i]), max(probs[i])] for i in option_ids},
},
"orderings": results,
}
async def read_limited(stream, declared: str | None, limit: int) -> bytes:
"""Read a request body, refusing it once it would exceed `limit` bytes. The check runs BEFORE a
chunk is kept, and nothing after the crossing chunk is read (bug hunt C2). A declared length is
trusted only as ASCII digits: `"²".isdigit()` is True but `int("²")` raises (S3)."""
too_large = ApiError(413, "request_too_large", f"request body exceeds {limit} bytes")
if declared is not None and declared.isascii() and declared.isdigit() and int(declared) > limit:
raise too_large
body = bytearray()
async for chunk in stream:
if len(body) + len(chunk) > limit:
raise too_large
body.extend(chunk)
return bytes(body)
def calibrated_view(result: dict, workload: str, temperature: float) -> dict:
"""softmax(option_logits / T): the native fields are left exactly as SemIf returned them (INV-1)."""
scaled = [x / temperature for x in result["option_logits"]]
@@ -77,6 +148,27 @@ def create_app(settings: Settings, engine: Any) -> FastAPI:
app = FastAPI(title="semif-serve")
expected = f"Bearer {settings.api_token}".encode()
inference = threading.Lock() # INV-2: one scorer call at a time, off the event loop
in_progress = 0 # POSTs admitted and not yet answered (bug hunt C6)
def admit():
nonlocal in_progress
if in_progress >= settings.max_queue:
raise ApiError(429, "busy", f"{in_progress} requests already in progress (limit {settings.max_queue})")
in_progress += 1
def leave():
nonlocal in_progress
in_progress -= 1
def build(fn, *args):
"""Response construction from a scorer result (calibration, averaging) maps its failures to
the 500 envelope too, instead of escaping as a bare 500 (bug hunt C3)."""
try:
return fn(*args)
except ApiError:
raise
except Exception as exc: # noqa: BLE001
raise ApiError(500, "scoring_failed", f"building the response failed: {type(exc).__name__}: {exc}") from exc
def locked(fn, *args):
with inference:
@@ -94,22 +186,10 @@ def create_app(settings: Settings, engine: Any) -> FastAPI:
async def api_error(_request: Request, exc: ApiError):
return error(exc.status, exc.code, exc.message)
async def read_limited(request: Request) -> bytes:
limit = settings.max_body_bytes
too_large = ApiError(413, "request_too_large", f"request body exceeds {limit} bytes")
declared = request.headers.get("content-length")
if declared is not None and declared.isdigit() and int(declared) > limit:
raise too_large
body = bytearray()
async for chunk in request.stream(): # also caps bodies that declare no length
body.extend(chunk)
if len(body) > limit:
raise too_large
return bytes(body)
async def parse(request: Request, model: type[BaseModel]):
try:
return model.model_validate_json(await read_limited(request))
return model.model_validate_json(
await read_limited(request.stream(), request.headers.get("content-length"), settings.max_body_bytes))
except ValidationError as exc:
raise ApiError(422, "invalid_request", _first_error(exc)) from exc
@@ -142,20 +222,59 @@ def create_app(settings: Settings, engine: Any) -> FastAPI:
"vram_cap_gib": settings.vram_cap_gib, "max_tokens": settings.max_tokens,
"max_decisions": settings.max_decisions, "workloads": sorted(settings.calibration)}
async def score_batch(decisions: list[Decision], state: State, workload: str | None) -> tuple[list[dict], dict]:
"""One engine.shared call for every row of every decision; results in request order."""
plan = [] # (decision, perms or None, row count)
rows: list[dict] = []
for d in decisions:
if d.orderings == "none":
plan.append((d, None, 1))
rows.append(d.row(state))
else:
if workload is not None:
raise ApiError(422, "invalid_request",
"workload calibration is not available together with orderings")
perms = ordering_perms(d)
plan.append((d, perms, len(perms)))
rows.extend(expanded_rows(d, state, perms))
if not 1 <= len(rows) <= settings.max_decisions:
raise ApiError(422, "invalid_request",
f"this request expands to {len(rows)} scored rows; the limit is 1..{settings.max_decisions}")
temperature = temperature_for(workload)
results, timing = await score(engine.shared, rows)
out, cursor = [], 0
for d, perms, count in plan:
chunk = results[cursor:cursor + count]
cursor += count
out.append(build(with_calibration, chunk[0], workload, temperature) if perms is None
else build(combine, d, chunk))
return out, timing
@app.post("/decide")
async def decide(request: Request):
body = await parse(request, DecideBody)
temperature = temperature_for(body.workload)
return with_calibration(await score(engine.direct, body.row(body.state)), body.workload, temperature)
admit()
try:
body = await parse(request, DecideBody)
if body.orderings != "none":
results, _timing = await score_batch([body], body.state, body.workload)
return results[0]
temperature = temperature_for(body.workload)
result = await score(engine.direct, body.row(body.state))
return build(with_calibration, result, body.workload, temperature)
finally:
leave()
@app.post("/decide/shared")
async def decide_shared(request: Request):
body = await parse(request, SharedBody)
if not 1 <= len(body.decisions) <= settings.max_decisions:
raise ApiError(422, "invalid_request",
f"decisions must hold 1..{settings.max_decisions} entries, got {len(body.decisions)}")
temperature = temperature_for(body.workload)
results, timing = await score(engine.shared, [d.row(body.state) for d in body.decisions])
return {"results": [with_calibration(r, body.workload, temperature) for r in results], "timing": timing}
admit()
try:
body = await parse(request, SharedBody)
if not 1 <= len(body.decisions) <= settings.max_decisions:
raise ApiError(422, "invalid_request",
f"decisions must hold 1..{settings.max_decisions} entries, got {len(body.decisions)}")
results, timing = await score_batch(body.decisions, body.state, body.workload)
return {"results": results, "timing": timing}
finally:
leave()
return app
+54 -12
View File
@@ -1,4 +1,9 @@
"""Settings for semif-serve. Contract: semif-serve.contract.md § Configuration."""
"""Settings for semif-serve. Contract: semif-serve.contract.md § Configuration.
Every value is validated at startup and a bad one is refused with a ValueError naming the
variable: a service that starts and then rejects every request (or runs uncapped) is worse
than one that does not start (bug hunt 2026-09-27: C1, S1, S2, S8).
"""
from __future__ import annotations
import json
@@ -8,6 +13,7 @@ from dataclasses import dataclass, field
from pathlib import Path
MIN_TOKEN_CHARS = 32
MIN_TEMPERATURE, MAX_TEMPERATURE = 0.05, 20.0
SEMIF_COMMIT = "23cf1f39fc9534fe81437200959b6dfc7106e45a"
DEFAULT_MODEL = "Qwen/Qwen3.5-4B"
DEFAULT_REVISION = "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a"
@@ -23,35 +29,71 @@ class Settings:
max_tokens: int = 4096
max_decisions: int = 64
max_body_bytes: int = 1024 * 1024
max_queue: int = 32
calibration: dict[str, float] = field(default_factory=dict)
@classmethod
def from_env(cls, env: Mapping[str, str]) -> "Settings":
token = env.get("SEMIF_API_TOKEN", "")
if len(token) < MIN_TOKEN_CHARS: # INV-6
raise ValueError(f"SEMIF_API_TOKEN must be at least {MIN_TOKEN_CHARS} characters")
cap = env.get("SEMIF_VRAM_CAP_GIB")
# INV-6: visible ASCII only. A CR, LF or NUL can never arrive in a header, so a token
# carrying one would lock every caller out while /health still said ok.
if len(token) < MIN_TOKEN_CHARS or not all(33 <= ord(c) <= 126 for c in token):
raise ValueError(f"SEMIF_API_TOKEN must be at least {MIN_TOKEN_CHARS} visible ASCII characters")
return cls(
api_token=token,
model=env.get("SEMIF_MODEL", DEFAULT_MODEL),
revision=env.get("SEMIF_REVISION", DEFAULT_REVISION),
device=env.get("SEMIF_DEVICE", "cuda"),
vram_cap_gib=float(cap) if cap else None,
max_tokens=int(env.get("SEMIF_MAX_TOKENS", 4096)),
max_decisions=int(env.get("SEMIF_MAX_DECISIONS", 64)),
max_body_bytes=int(env.get("SEMIF_MAX_BODY_BYTES", 1024 * 1024)),
vram_cap_gib=_positive_float(env, "SEMIF_VRAM_CAP_GIB"),
max_tokens=_positive_int(env, "SEMIF_MAX_TOKENS", 4096),
max_decisions=_positive_int(env, "SEMIF_MAX_DECISIONS", 64),
max_body_bytes=_positive_int(env, "SEMIF_MAX_BODY_BYTES", 1024 * 1024),
max_queue=_positive_int(env, "SEMIF_MAX_QUEUE", 32),
calibration=_load_calibration(env.get("SEMIF_CALIBRATION")),
)
def _positive_int(env: Mapping[str, str], name: str, default: int) -> int:
raw = env.get(name)
if raw is None:
return default
try:
value = int(raw)
except ValueError:
raise ValueError(f"{name} must be an integer, got {raw!r}") from None
if value < 1:
raise ValueError(f"{name} must be >= 1, got {value}")
return value
def _positive_float(env: Mapping[str, str], name: str) -> float | None:
"""Unset means no cap. When set it must be finite and > 0: `0` used to slip through as 'no cap'."""
raw = env.get(name)
if raw is None or raw == "":
return None
try:
value = float(raw)
except ValueError:
raise ValueError(f"{name} must be a number, got {raw!r}") from None
if not math.isfinite(value) or value <= 0:
raise ValueError(f"{name} must be a finite number > 0, got {raw!r}")
return value
def _load_calibration(path: str | None) -> dict[str, float]:
"""{workload: T}, every T a finite number > 0 (T scales option logits before softmax)."""
"""{workload: T}; T scales option logits before softmax, so it is kept in a sane range
(a tiny T overflows to NaN and the response then fails to render)."""
if not path:
return {}
table = json.loads(Path(path).read_text())
try:
table = json.loads(Path(path).read_text())
except (OSError, ValueError) as exc:
raise ValueError(f"SEMIF_CALIBRATION {path!r} could not be read as JSON: {exc}") from None
if not isinstance(table, dict) or not all(
isinstance(t, (int, float)) and not isinstance(t, bool) and math.isfinite(t) and t > 0
isinstance(t, (int, float)) and not isinstance(t, bool) and math.isfinite(t)
and MIN_TEMPERATURE <= t <= MAX_TEMPERATURE
for t in table.values()
):
raise ValueError("SEMIF_CALIBRATION must be a JSON object of workload -> finite temperature > 0")
raise ValueError(f"SEMIF_CALIBRATION must be a JSON object of workload -> temperature in "
f"[{MIN_TEMPERATURE}, {MAX_TEMPERATURE}]")
return {str(k): float(v) for k, v in table.items()}
+27 -7
View File
@@ -7,10 +7,14 @@ unit-tested against a fake torch (tests/test_engine.py).
from __future__ import annotations
import gc
import logging
import traceback
from typing import Any, Callable
from .config import Settings
from .errors import OutOfMemory
from .errors import OutOfMemory, ScoringFailed
log = logging.getLogger("semif_serve.engine")
RELEASE_SLACK_BYTES = 512 * 2**20
WARMUP_ROW = {
@@ -24,6 +28,11 @@ WARMUP_ROW = {
}
def _first_line(exc: BaseException) -> str:
lines = str(exc).splitlines()
return lines[0] if lines else ""
class TorchEngine:
def __init__(self, torch: Any, model: Any, tokenizer: Any, metadata: dict, settings: Settings,
direct_fn: Callable, shared_fn: Callable, release_above_bytes: int | None = None):
@@ -46,7 +55,7 @@ class TorchEngine:
arch = f"sm_{major}{minor}"
if arch not in torch.cuda.get_arch_list(): # INV-3: no silent PTX/CPU fallback
raise RuntimeError(f"torch {torch.__version__} has no kernels for {arch}: {torch.cuda.get_arch_list()}")
if settings.vram_cap_gib: # INV-4: cap BEFORE the weights land
if settings.vram_cap_gib is not None: # INV-4: cap BEFORE the weights land
total = torch.cuda.get_device_properties(0).total_memory
fraction = settings.vram_cap_gib * 2**30 / total
if not 0 < fraction <= 1:
@@ -82,17 +91,28 @@ class TorchEngine:
def _guard(self, fn, *args):
try:
result = fn(*args)
except ValueError:
raise # validation: SemIf raises it before any GPU work
except self._torch.cuda.OutOfMemoryError as exc:
message = str(exc).splitlines()[0]
failure, message = OutOfMemory, _first_line(exc) or "CUDA out of memory"
except Exception as exc: # noqa: BLE001 — every other failure is released and reported below
message = _first_line(exc)
if "out of memory" in message.lower(): # cuBLAS/cuDNN allocation failures
failure = OutOfMemory
else:
failure, message = ScoringFailed, f"{type(exc).__name__}: {message}"
# Formatted text, not exc_info: a log record that keeps the traceback object alive
# (pytest's capture handler does; so would any buffering handler) pins the tensors.
log.error("scorer failed:\n%s", traceback.format_exc())
else:
self._release_burst()
return result
# INV-4, outside the except block on purpose: the torch exception's traceback holds the
# failed scorer's frames, and with them its tensors (the replicated prefix cache). Raising
# inside the block, or `from exc`, would chain to it and keep GiBs allocated after the 503.
# INV-4, outside the except block on purpose: the exception's traceback holds the failed
# scorer's frames, and with them its tensors (the replicated prefix cache). Raising inside
# the block, or `from exc`, would chain to it and keep GiBs allocated after the response.
gc.collect()
self._torch.cuda.empty_cache()
raise OutOfMemory(message)
raise failure(message)
def direct(self, row: dict) -> dict:
return self._guard(self._direct, self._model, self._tokenizer, row, self._metadata, self._settings.max_tokens)
@@ -3,3 +3,8 @@
class OutOfMemory(RuntimeError):
"""The engine ran out of GPU memory during a request and has already released its cache (INV-4)."""
class ScoringFailed(RuntimeError):
"""A scorer call failed for a reason other than validation or OOM. Raised unchained, after the
failed call's memory has been released; the original traceback is logged, not carried (INV-4)."""
@@ -11,6 +11,10 @@ from .config import Settings
def app_from_env() -> FastAPI:
settings = Settings.from_env(os.environ)
# INV-5: never download at runtime, inside the image or out of it (bug hunt S10). Set before
# torch / transformers / huggingface_hub are imported, since they read it at import time.
os.environ["HF_HUB_OFFLINE"] = "1"
os.environ["TRANSFORMERS_OFFLINE"] = "1"
from .engine import TorchEngine # torch loads only here, never in the unit tests
return create_app(settings, TorchEngine.load(settings))
+103
View File
@@ -1,5 +1,7 @@
"""semif-serve HTTP behaviour against a fake engine (no torch, no model).
Contract: services/semif-serve/semif-serve.contract.md"""
import json
import pytest
from fastapi.testclient import TestClient
@@ -231,3 +233,104 @@ def test_health_reports_the_pins_limits_and_workloads():
body = client.get("/health").json()
assert body == {"status": "ok", "semif_commit": SEMIF_COMMIT, "model": FakeEngine().health(),
"vram_cap_gib": 12.0, "max_tokens": 4096, "max_decisions": 8, "workloads": ["alerts", "triage"]}
def test_a_body_of_exactly_the_limit_is_accepted():
body = json.dumps(ROW).encode()
client = make_client(max_body_bytes=len(body))
assert client.post("/decide", content=body, headers={**AUTH, "content-type": "application/json"}).status_code == 200
def test_a_non_ascii_digit_content_length_is_ignored_not_a_crash():
"""HTTP clients cannot send one (httpx refuses; h11 rejects it), so check the reader directly."""
import asyncio
from semif_serve.app import read_limited
async def body():
yield b"{}"
assert asyncio.run(read_limited(body(), "²", 100)) == b"{}" # int("²") would raise
def test_exactly_max_decisions_is_accepted():
decisions = [{"id": str(i), "question": "Q?", "options": OPTIONS} for i in range(3)]
client = make_client(max_decisions=3)
assert client.post("/decide/shared", json={"state": "s", "decisions": decisions}, headers=AUTH).status_code == 200
def test_calibration_leaves_every_native_field_alone():
engine = FakeEngine(logits=(3.0, 1.0))
body = make_client(engine, calibration={"triage": 2.0}).post(
"/decide", json={**ROW, "workload": "triage"}, headers=AUTH).json()
body.pop("calibrated")
assert body == engine.direct(ROW)
class MalformedEngine(FakeEngine):
def direct(self, row):
return {"id": row["id"], "option_ids": ["yes", "no"], "probabilities": [0.5, 0.5]} # no option_logits
def shared(self, rows):
return [self.direct(r) for r in rows], {}
@pytest.mark.parametrize("path, body", [
("/decide", {**ROW, "workload": "triage"}),
("/decide", {**ROW, "orderings": "rotations"}),
])
def test_a_malformed_scorer_result_is_an_envelope_500_not_a_bare_one(path, body):
client = make_client(MalformedEngine(), calibration={"triage": 2.0})
response = client.post(path, json=body, headers=AUTH)
assert response.status_code == 500
assert response.json()["error"]["code"] == "scoring_failed"
class SharedSlowEngine(SlowEngine):
def shared(self, rows):
return [self.direct(r) for r in rows], {}
def test_shared_requests_are_serialised_too():
from concurrent.futures import ThreadPoolExecutor
engine = SharedSlowEngine()
body = {"state": "s", "decisions": [{"id": "a", "question": "Q?", "options": OPTIONS}]}
with make_client(engine) as client, ThreadPoolExecutor(3) as pool:
futures = [pool.submit(client.post, "/decide/shared", json=body, headers=AUTH) for _ in range(3)]
assert engine.entered.wait(5)
engine.release.set()
assert [f.result().status_code for f in futures] == [200] * 3
assert engine.peak == 1
def test_more_than_max_queue_requests_in_progress_get_429_busy():
from concurrent.futures import ThreadPoolExecutor
engine = SlowEngine()
with make_client(engine, max_queue=2) as client, ThreadPoolExecutor(3) as pool:
held = [pool.submit(client.post, "/decide", json={**ROW, "id": f"r{i}"}, headers=AUTH) for i in range(2)]
assert engine.entered.wait(5)
import time
deadline = time.monotonic() + 5
while engine.inside + 0 < 1 and time.monotonic() < deadline:
time.sleep(0.01)
time.sleep(0.2) # let the second request reach the queue
extra = client.post("/decide", json={**ROW, "id": "extra"}, headers=AUTH)
assert extra.status_code == 429 and extra.json()["error"]["code"] == "busy"
engine.release.set()
assert [f.result().status_code for f in held] == [200, 200]
assert make_client(FakeEngine(), max_queue=2).post("/decide", json=ROW, headers=AUTH).status_code == 200
def test_read_limited_stops_reading_at_the_crossing_chunk():
import asyncio
from semif_serve.app import ApiError, read_limited
consumed = []
async def chunks():
for i in range(10):
consumed.append(i)
yield b"x" * 100
with pytest.raises(ApiError) as info:
asyncio.run(read_limited(chunks(), None, 250))
assert info.value.status == 413
assert consumed == [0, 1, 2] # the third chunk crosses 250 and is never kept; nothing after is read
+42
View File
@@ -34,3 +34,45 @@ def test_a_calibration_table_needs_positive_finite_numbers(tmp_path, table):
cal.write_text(json.dumps(table))
with pytest.raises(ValueError, match="SEMIF_CALIBRATION"):
Settings.from_env({"SEMIF_API_TOKEN": TOKEN, "SEMIF_CALIBRATION": str(cal)})
@pytest.mark.parametrize("var, value", [
("SEMIF_VRAM_CAP_GIB", "0"), ("SEMIF_VRAM_CAP_GIB", "-4"), ("SEMIF_VRAM_CAP_GIB", "nan"), ("SEMIF_VRAM_CAP_GIB", "inf"),
("SEMIF_MAX_TOKENS", "0"), ("SEMIF_MAX_DECISIONS", "-1"), ("SEMIF_MAX_BODY_BYTES", "0"), ("SEMIF_MAX_QUEUE", "0"),
("SEMIF_MAX_TOKENS", "lots"), ("SEMIF_VRAM_CAP_GIB", "twelve"),
])
def test_out_of_range_or_unparseable_values_are_refused_naming_the_variable(var, value):
with pytest.raises(ValueError, match=var):
Settings.from_env({"SEMIF_API_TOKEN": TOKEN, var: value})
@pytest.mark.parametrize("token", ["x" * 31 + "\n", "x" * 30 + "\r\n", "x" * 32 + "\x00", "x" * 16 + " " + "x" * 16,
"é" * 32])
def test_a_token_with_non_visible_ascii_is_refused(token):
with pytest.raises(ValueError, match="SEMIF_API_TOKEN"):
Settings.from_env({"SEMIF_API_TOKEN": token})
@pytest.mark.parametrize("table", [{"w": True}, {"w": 1e-300}, {"w": 0.04}, {"w": 21}])
def test_a_temperature_must_be_a_real_number_in_range(tmp_path, table):
cal = tmp_path / "cal.json"
cal.write_text(json.dumps(table))
with pytest.raises(ValueError, match="SEMIF_CALIBRATION"):
Settings.from_env({"SEMIF_API_TOKEN": TOKEN, "SEMIF_CALIBRATION": str(cal)})
@pytest.mark.parametrize("content", [None, "{not json"])
def test_a_missing_or_malformed_calibration_file_is_a_named_startup_error(tmp_path, content):
cal = tmp_path / "cal.json"
if content is not None:
cal.write_text(content)
with pytest.raises(ValueError, match="SEMIF_CALIBRATION"):
Settings.from_env({"SEMIF_API_TOKEN": TOKEN, "SEMIF_CALIBRATION": str(cal)})
def test_in_range_edges_are_accepted(tmp_path):
cal = tmp_path / "cal.json"
cal.write_text(json.dumps({"lo": 0.05, "hi": 20}))
s = Settings.from_env({"SEMIF_API_TOKEN": "!" + "~" * 31, "SEMIF_VRAM_CAP_GIB": "0.5", "SEMIF_MAX_QUEUE": "1",
"SEMIF_CALIBRATION": str(cal)})
assert (s.vram_cap_gib, s.max_queue, s.calibration) == (0.5, 1, {"lo": 0.05, "hi": 20.0})
+31 -1
View File
@@ -7,7 +7,7 @@ import pytest
from semif_serve.config import Settings
from semif_serve.engine import TorchEngine
from semif_serve.errors import OutOfMemory
from semif_serve.errors import OutOfMemory, ScoringFailed
class FakeTorch:
@@ -81,3 +81,33 @@ def test_a_burst_is_returned_to_the_driver_after_the_call(reserved_after, releas
direct_fn=scorer, shared_fn=scorer, release_above_bytes=8 * 2**30 + 512 * 2**20)
assert engine.direct({}) == {"ok": True}
assert ReservingTorch.cuda.emptied == released
def cyclic_tensor():
"""A tensor held in a reference cycle, as real frames and tensors often are: only gc frees it."""
t = Tensor()
t.self_ref = t
FakeTorch.cuda.watched.append(weakref.ref(t))
return t
@pytest.mark.parametrize("raised, expected_type, expected_message", [
(lambda: FakeTorch.cuda.OutOfMemoryError(""), OutOfMemory, "CUDA out of memory"),
(lambda: RuntimeError("CUBLAS_STATUS_ALLOC_FAILED: CUDA error: out of memory"), OutOfMemory,
"CUBLAS_STATUS_ALLOC_FAILED: CUDA error: out of memory"),
(lambda: RuntimeError("Invalid native prefix cache"), ScoringFailed, "RuntimeError: Invalid native prefix cache"),
(lambda: KeyError("option_logits"), ScoringFailed, "KeyError: 'option_logits'"),
])
def test_every_non_validation_failure_is_released_unchained_after_gc(raised, expected_type, expected_message):
FakeTorch.cuda.empties.clear(), FakeTorch.cuda.watched.clear()
def scorer(*_args):
kv_cache = cyclic_tensor() # noqa: F841 — alive in this frame when it raises
raise raised()
engine = TorchEngine(FakeTorch, None, None, {}, Settings(api_token="t" * 40), direct_fn=scorer, shared_fn=scorer)
with pytest.raises(expected_type) as info:
engine.shared([])
assert str(info.value) == expected_message
assert info.value.__cause__ is None and info.value.__context__ is None
assert FakeTorch.cuda.empties == [True] # gc freed the cycle BEFORE the cache was emptied
@@ -0,0 +1,110 @@
"""TorchEngine.load and the entry point, against fake torch/semif modules (INV-3, INV-4, INV-5).
The bug hunt found load() had no test at all: removing the arch check survived every test."""
import os
import sys
import types
import pytest
from semif_serve.config import Settings
TOKEN = "t" * 40
class Param:
def __init__(self, device):
self.device = types.SimpleNamespace(type=device)
class Model:
def __init__(self, device):
self._device = device
def parameters(self):
yield Param(self._device)
@pytest.fixture
def fakes(monkeypatch):
calls = []
cuda = types.SimpleNamespace(
is_available=lambda: True,
get_device_capability=lambda _i=0: (12, 0),
get_arch_list=lambda: ["sm_90", "sm_120"],
get_device_properties=lambda _i=0: types.SimpleNamespace(total_memory=96 * 2**30),
set_per_process_memory_fraction=lambda f, _i=0: calls.append(("cap", round(f, 4))),
memory_reserved=lambda _i=0: 8 * 2**30,
OutOfMemoryError=type("OutOfMemoryError", (RuntimeError,), {}),
empty_cache=lambda: calls.append(("empty",)),
)
torch = types.SimpleNamespace(cuda=cuda, __version__="2.10.0+cu128")
state = {"device": "cuda", "warmup_raises": None}
def load_causal_model(model, revision, device, dtype):
calls.append(("load", model, revision, device, dtype))
return Model(state["device"]), object(), {"source": model}
def score(model, tok, row, meta, max_tokens):
calls.append(("score", row["id"]))
if state["warmup_raises"]:
raise state["warmup_raises"]
return {"id": row["id"]}
core = types.ModuleType("semif_phase1.core")
core.load_causal_model = load_causal_model
direct = types.ModuleType("semif_phase1.direct")
direct.score = score
shared = types.ModuleType("semif_phase1.shared")
shared.score_shared = lambda *a: ([], {})
pkg = types.ModuleType("semif_phase1")
for name, mod in {"torch": torch, "semif_phase1": pkg, "semif_phase1.core": core,
"semif_phase1.direct": direct, "semif_phase1.shared": shared}.items():
monkeypatch.setitem(sys.modules, name, mod)
return torch, calls, state
def test_load_caps_before_the_weights_land_then_warms_up(fakes):
from semif_serve.engine import TorchEngine
_torch, calls, _ = fakes
TorchEngine.load(Settings(api_token=TOKEN, vram_cap_gib=12.0))
assert [c[0] for c in calls] == ["cap", "load", "score"]
assert calls[0] == ("cap", round(12 / 96, 4))
assert calls[2] == ("score", "semif-serve-warmup")
def test_load_refuses_a_card_torch_has_no_kernels_for(fakes):
from semif_serve.engine import TorchEngine
torch, calls, _ = fakes
torch.cuda.get_arch_list = lambda: ["sm_80", "sm_90"]
with pytest.raises(RuntimeError, match="sm_120"):
TorchEngine.load(Settings(api_token=TOKEN))
assert not any(c[0] == "load" for c in calls)
def test_load_refuses_a_model_that_landed_on_the_wrong_device(fakes):
from semif_serve.engine import TorchEngine
_torch, _calls, state = fakes
state["device"] = "cpu"
with pytest.raises(RuntimeError, match="landed on cpu"):
TorchEngine.load(Settings(api_token=TOKEN))
def test_load_fails_closed_when_the_warmup_decision_fails(fakes):
from semif_serve.engine import TorchEngine
from semif_serve.errors import ScoringFailed
_torch, _calls, state = fakes
state["warmup_raises"] = RuntimeError("Failed to find C compiler")
with pytest.raises(ScoringFailed, match="C compiler"):
TorchEngine.load(Settings(api_token=TOKEN))
def test_the_entry_point_forces_offline_mode_before_the_engine_loads(fakes, monkeypatch):
import semif_serve.engine as engine_mod
from semif_serve import main
seen = {}
monkeypatch.delenv("HF_HUB_OFFLINE", raising=False)
monkeypatch.setenv("SEMIF_API_TOKEN", TOKEN)
monkeypatch.setattr(engine_mod.TorchEngine, "load",
classmethod(lambda cls, s: seen.update(offline=os.environ.get("HF_HUB_OFFLINE")) or object()))
main.app_from_env()
assert seen["offline"] == "1"
@@ -0,0 +1,132 @@
"""Order averaging (0.1.3). Contract: semif-serve.contract.md § Order averaging."""
import math
import pytest
from fastapi.testclient import TestClient
from semif_serve.app import create_app
from semif_serve.config import Settings
TOKEN = "t" * 40
AUTH = {"Authorization": f"Bearer {TOKEN}"}
OPTS = [{"id": "casual", "description": "A casual outing"},
{"id": "date", "description": "A romantic date"},
{"id": "booty", "description": "A booty call"}]
ROW = {"id": "q", "state": "It's 2AM and I'm bored.", "question": "What is this?", "options": OPTS}
BASE = {"casual": 1.0, "date": -1.0, "booty": 0.5} # what the model "really" thinks
BIAS = 2.5 # added to whichever option is listed first
def softmax(xs):
top = max(xs)
w = [math.exp(x - top) for x in xs]
return [v / sum(w) for v in w]
class BiasedEngine:
"""Scores each option as BASE[id], plus BIAS for the first-listed option, the way a small model leans."""
def __init__(self):
self.shared_calls, self.direct_calls = [], []
def health(self):
return {}
def _score(self, row):
ids = [o["id"] for o in row["options"]]
logits = [BASE[i] + (BIAS if k == 0 else 0.0) for k, i in enumerate(ids)]
return {"id": row["id"], "option_ids": ids, "probabilities": softmax(logits), "option_logits": logits,
"prompt_sha256": row["id"], "probability_status": "uncalibrated"}
def direct(self, row):
self.direct_calls.append(row)
return self._score(row)
def shared(self, rows):
self.shared_calls.append(rows)
return [self._score(r) for r in rows], {"batch_size": len(rows)}
def client(engine, **kw):
return TestClient(create_app(Settings(api_token=TOKEN, **kw), engine))
def test_rotations_send_one_shared_call_with_every_option_once_per_position_and_cancel_the_bias():
engine = BiasedEngine()
body = client(engine).post("/decide", json={**ROW, "orderings": "rotations"}, headers=AUTH).json()
assert engine.direct_calls == [] and len(engine.shared_calls) == 1
rows = engine.shared_calls[0]
assert [r["id"] for r in rows] == ["q#o0", "q#o1", "q#o2"]
assert [[o["id"] for o in r["options"]] for r in rows] == [
["casual", "date", "booty"], ["date", "booty", "casual"], ["booty", "casual", "date"]]
assert all(r["state"] == ROW["state"] and r["question"] == ROW["question"] for r in rows)
assert body["id"] == "q" and body["option_ids"] == ["casual", "date", "booty"]
combined = body["combined"]
assert combined["method"] == "rotations" and combined["orderings"] == 3
# the first-position bias is additive, and every option sat first exactly once, so it cancels exactly
assert combined["probabilities"] == pytest.approx(softmax([BASE["casual"], BASE["date"], BASE["booty"]]))
assert combined["top"] == "casual"
# the bias is strong enough that every ordering's first option wins: casual, date, booty
tops = [max(zip(r["option_logits"], r["option_ids"]))[1] for r in body["orderings"]]
assert tops == ["casual", "date", "booty"]
assert combined["agreement"] == pytest.approx(1 / 3)
assert body["orderings"] == [engine._score(r) for r in rows] # native results, unchanged
for oid in ("casual", "date", "booty"):
ps = [dict(zip(r["option_ids"], r["probabilities"]))[oid] for r in body["orderings"]]
assert combined["spread"][oid] == pytest.approx([min(ps), max(ps)])
def test_all_sends_every_permutation_with_the_callers_order_first():
engine = BiasedEngine()
body = client(engine).post("/decide", json={**ROW, "orderings": "all"}, headers=AUTH).json()
rows = engine.shared_calls[0]
orders = [tuple(o["id"] for o in r["options"]) for r in rows]
assert len(orders) == 6 and len(set(orders)) == 6 and orders[0] == ("casual", "date", "booty")
assert body["combined"]["orderings"] == 6 and body["combined"]["method"] == "all"
def test_all_above_four_options_is_422_before_the_engine_runs():
engine = BiasedEngine()
five = [{"id": f"o{i}", "description": f"Option {i}"} for i in range(5)]
r = client(engine).post("/decide", json={**ROW, "options": five, "orderings": "all"}, headers=AUTH)
assert r.status_code == 422 and "rotations" in r.json()["error"]["message"]
assert engine.shared_calls == []
def test_a_mixed_shared_request_is_one_engine_call_with_results_in_request_order():
engine = BiasedEngine()
body = {"state": ROW["state"], "decisions": [
{"id": "plain", "question": "Q1?", "options": OPTS},
{"id": "avg", "question": "Q2?", "options": OPTS, "orderings": "rotations"},
{"id": "plain2", "question": "Q3?", "options": OPTS[:2]}]}
out = client(engine).post("/decide/shared", json=body, headers=AUTH).json()
assert len(engine.shared_calls) == 1
assert [r["id"] for r in engine.shared_calls[0]] == ["plain", "avg#o0", "avg#o1", "avg#o2", "plain2"]
ids = [r["id"] for r in out["results"]]
assert ids == ["plain", "avg", "plain2"]
assert out["results"][0] == engine._score(engine.shared_calls[0][0]) # plain results keep the old shape
assert "combined" in out["results"][1] and "combined" not in out["results"][2]
assert out["timing"] == {"batch_size": 5}
def test_expanded_rows_count_toward_the_cap():
engine = BiasedEngine()
r = client(engine, max_decisions=5).post("/decide/shared", json={"state": "s", "decisions": [
{"id": "a", "question": "Q?", "options": OPTS, "orderings": "rotations"},
{"id": "b", "question": "Q?", "options": OPTS, "orderings": "rotations"}]}, headers=AUTH)
assert r.status_code == 422 and "6 scored rows" in r.json()["error"]["message"]
assert engine.shared_calls == []
def test_workload_with_orderings_is_422_and_workload_alone_still_calibrates():
engine = BiasedEngine()
c = client(engine, calibration={"triage": 2.0})
r = c.post("/decide", json={**ROW, "orderings": "rotations", "workload": "triage"}, headers=AUTH)
assert r.status_code == 422 and engine.shared_calls == []
assert "calibrated" in c.post("/decide", json={**ROW, "workload": "triage"}, headers=AUTH).json()
def test_an_unknown_orderings_value_is_422():
r = client(BiasedEngine()).post("/decide", json={**ROW, "orderings": "shuffle"}, headers=AUTH)
assert r.status_code == 422
+80 -2
View File
@@ -51,6 +51,26 @@ wheels = [
{ url = "https://files.pythonhosted.org/packages/12/b8/4bd346e22b28902df4d651910f5242c28d84e4a5c2435ca5c3f797ed7e2e/anyio-4.15.1-py3-none-any.whl", hash = "sha256:6152fdbbf9a77fdec97731721bebf7c4c44f7c29b424b0065826173efc7ed101", size = 132079 },
]
[[package]]
name = "causal-conv1d"
version = "1.7.0"
source = { url = "https://github.com/Dao-AILab/causal-conv1d/releases/download/v1.7.0/causal_conv1d-1.7.0+cu12torch2.10cxx11abiTRUE-cp312-cp312-linux_x86_64.whl" }
dependencies = [
{ name = "ninja" },
{ name = "packaging" },
{ name = "torch" },
]
wheels = [
{ url = "https://github.com/Dao-AILab/causal-conv1d/releases/download/v1.7.0/causal_conv1d-1.7.0+cu12torch2.10cxx11abiTRUE-cp312-cp312-linux_x86_64.whl", hash = "sha256:8e81f8c76435ad31aa41edc6c0c9d26de971a293b190183d2816da71547614d7" },
]
[package.metadata]
requires-dist = [
{ name = "ninja" },
{ name = "packaging" },
{ name = "torch" },
]
[[package]]
name = "certifi"
version = "2026.7.22"
@@ -106,6 +126,15 @@ wheels = [
{ url = "https://files.pythonhosted.org/packages/98/59/239c7259e669c46ddbcac0aa60e3a0ef00bfeaaa687f905b24dd6a7a10fe/cuda_pathfinder-1.8.2-py3-none-any.whl", hash = "sha256:4e65059febdb4d19d5cbc4798677e19db2b582f2f702f457b609e571690d357e", size = 62551 },
]
[[package]]
name = "einops"
version = "0.8.2"
source = { registry = "https://pypi.org/simple" }
sdist = { url = "https://files.pythonhosted.org/packages/2c/77/850bef8d72ffb9219f0b1aac23fbc1bf7d038ee6ea666f331fa273031aa2/einops-0.8.2.tar.gz", hash = "sha256:609da665570e5e265e27283aab09e7f279ade90c4f01bcfca111f3d3e13f2827", size = 56261 }
wheels = [
{ url = "https://files.pythonhosted.org/packages/2a/09/f8d8f8f31e4483c10a906437b4ce31bdf3d6d417b73fe33f1a8b59e34228/einops-0.8.2-py3-none-any.whl", hash = "sha256:54058201ac7087911181bfec4af6091bb59380360f069276601256a76af08193", size = 65638 },
]
[[package]]
name = "fastapi"
version = "0.118.0"
@@ -129,6 +158,31 @@ wheels = [
{ url = "https://files.pythonhosted.org/packages/bc/ac/8b9c6dc2aa7e9cf5582e20a337006c3428c3d752c2255dbaca2c78a8be45/filelock-4.0.4-py3-none-any.whl", hash = "sha256:0df72be195ca7892216d16f2edce8d9b93a571f02402972020a8cff84c594c7b", size = 108629 },
]
[[package]]
name = "fla-core"
version = "0.5.2"
source = { registry = "https://pypi.org/simple" }
dependencies = [
{ name = "einops" },
]
sdist = { url = "https://files.pythonhosted.org/packages/94/85/20bdc0fbbbaeec27b5b198e0dbbbc51161a267066a4aebd9e4b413637c6c/fla_core-0.5.2.tar.gz", hash = "sha256:9360fc412f784c1c8f05c320ec2902d5d968764178c9b8a92efc919e17a39680", size = 587835 }
wheels = [
{ url = "https://files.pythonhosted.org/packages/2d/ed/dfe19c4da779957eb6a42a26812f9b4e2280bf757a17a71933ff59ffcb98/fla_core-0.5.2-py3-none-any.whl", hash = "sha256:5e830c85bad3d0d34677f98ac7074d08687a3756f0f0499d95ceb96eb6920761", size = 819225 },
]
[[package]]
name = "flash-linear-attention"
version = "0.5.2"
source = { registry = "https://pypi.org/simple" }
dependencies = [
{ name = "fla-core" },
{ name = "transformers" },
]
sdist = { url = "https://files.pythonhosted.org/packages/d2/76/c180949eae5161b9fcf4c928cab52f0c257ce84c1a59b4462b1d2b6e5883/flash_linear_attention-0.5.2.tar.gz", hash = "sha256:c053d3a75c8f5b725063f719ae6bce0f4d9dac6da32735be8ae7d485a6c9820f", size = 208110 }
wheels = [
{ url = "https://files.pythonhosted.org/packages/90/d2/2070e3cf2148c5cce99ca4876633c4b7b89ec085323600bd3d47aeacd306/flash_linear_attention-0.5.2-py3-none-any.whl", hash = "sha256:dcf405d81f5426393b59037097aa700d0f4a841465d5028d5aa543f4502f2400", size = 399590 },
]
[[package]]
name = "fsspec"
version = "2026.9.0"
@@ -351,6 +405,24 @@ wheels = [
{ url = "https://files.pythonhosted.org/packages/7e/cd/fe58041e9011f307c490e3e17dd48cc516448f7c698a3f2d9d9d65d7e6a8/networkx-3.7-py3-none-any.whl", hash = "sha256:e3fd2c13a7814cee3746340d8d7f8598a67f16a58bf47fb7f8793fab6efca1b0", size = 2142205 },
]
[[package]]
name = "ninja"
version = "1.13.2"
source = { registry = "https://pypi.org/simple" }
sdist = { url = "https://files.pythonhosted.org/packages/ac/92/410b7917d16ab54c04b05cc32b9284803671d91cf79d33be6009c28d4ea8/ninja-1.13.2.tar.gz", hash = "sha256:525bfa3fc88aa30a4467df270fd5be6f9fcae8061d54d4df74ea1dc5abd5a975", size = 243739 }
wheels = [
{ url = "https://files.pythonhosted.org/packages/48/23/fcbe234a66966e35928c47b86336f92a7612db4781665f4e5f5fddef9630/ninja-1.13.2-py3-none-manylinux2014_ppc64le.manylinux_2_17_ppc64le.whl", hash = "sha256:227cbc3ae3e5e429692388103cae8c09451df086cd2d342dae0795af0d162547", size = 197676 },
{ url = "https://files.pythonhosted.org/packages/24/eb/a6ca97ef0ff7bb8bdcb395ec65a716e65d7c1f40896c3afe0090bb3e1535/ninja-1.13.2-py3-none-manylinux2014_s390x.manylinux_2_17_s390x.whl", hash = "sha256:1684c60d031c54c1d049541b64243c0c567dca5463dbd77682a8901780af293d", size = 187980 },
{ url = "https://files.pythonhosted.org/packages/6e/53/ebfed7b689c338dd8ebeec9c0730c8d56821292f14e2536e5f3ef1a05744/ninja-1.13.2-py3-none-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:65a24341b5ac09fcadcc37082660be40a94174e51a937fabf6e2cae26225fa2c", size = 183365 },
{ url = "https://files.pythonhosted.org/packages/c7/d6/dcf06d7ab44ade992ae5aa1228feff317684b463a1bd47e8642b30ac922e/ninja-1.13.2-py3-none-manylinux_2_28_armv7l.manylinux_2_31_armv7l.whl", hash = "sha256:aa3d2ae5706a2c4d1e93edc951d1c6cbb45107413c404f8fde1741239efbc9a0", size = 155089 },
{ url = "https://files.pythonhosted.org/packages/e1/6b/6513c09c33382b17c05b4349b8e81437b18680d0d7ec6fb8f7edc29adda1/ninja-1.13.2-py3-none-manylinux_2_31_riscv64.whl", hash = "sha256:919572cbc3f233261ecd41fe1f3efc9d44aa02464a4588867e06a8b4f6f416ea", size = 152149 },
{ url = "https://files.pythonhosted.org/packages/37/04/c8c2dc5b2f5fee79a1691d490256b178b7e1af97d56769117422ae8a23cc/ninja-1.13.2-py3-none-musllinux_1_2_armv7l.whl", hash = "sha256:59d71c3e15b6b6f3d903eb0c27285544e0747ca59925ada7037bb1af781ad4b3", size = 466268 },
{ url = "https://files.pythonhosted.org/packages/10/a2/d8eedd25d0ae80b9e874aea362013e67416877c8f732eb4d5e7c971fdb9c/ninja-1.13.2-py3-none-musllinux_1_2_ppc64le.whl", hash = "sha256:b2f687437fac460b27b7eadc99039b1163016fb4ba7276e2782a192d9f24ee0e", size = 610806 },
{ url = "https://files.pythonhosted.org/packages/14/0f/696d96821fad1b5767fd311c1569dde8881a57412369bfe7b11bcbfde036/ninja-1.13.2-py3-none-musllinux_1_2_riscv64.whl", hash = "sha256:09de9ab04f7352f51570c73fd4913acb1e6c24be0a72cd8b80243d4d3ed04925", size = 533978 },
{ url = "https://files.pythonhosted.org/packages/5d/69/28844ca579156776a202217a7cd66f60d06a0710a935e879bb89ce396ecc/ninja-1.13.2-py3-none-musllinux_1_2_s390x.whl", hash = "sha256:6a87bf42b123abe2f37737300185f0a303a891899da85d73a3613ee80547e578", size = 653822 },
{ url = "https://files.pythonhosted.org/packages/f5/5f/c511f2952f94ab2966d60edd9c34e744ea32f2724b1184b62270bde55b3a/ninja-1.13.2-py3-none-musllinux_1_2_x86_64.whl", hash = "sha256:915bd482c4be41c75120fd67a22e0bb3f0fbb3bbc5f95b89787deadd59e27ef2", size = 544460 },
]
[[package]]
name = "numpy"
version = "2.2.6"
@@ -919,7 +991,7 @@ dependencies = [
[[package]]
name = "semif-serve"
version = "0.1.2"
version = "0.1.3"
source = { editable = "." }
dependencies = [
{ name = "fastapi" },
@@ -927,6 +999,10 @@ dependencies = [
]
[package.optional-dependencies]
fast = [
{ name = "causal-conv1d", marker = "platform_machine == 'x86_64' and sys_platform == 'linux'" },
{ name = "flash-linear-attention" },
]
model = [
{ name = "semif-phase1" },
{ name = "torch" },
@@ -940,12 +1016,14 @@ dev = [
[package.metadata]
requires-dist = [
{ name = "causal-conv1d", marker = "platform_machine == 'x86_64' and sys_platform == 'linux' and extra == 'fast'", url = "https://github.com/Dao-AILab/causal-conv1d/releases/download/v1.7.0/causal_conv1d-1.7.0+cu12torch2.10cxx11abiTRUE-cp312-cp312-linux_x86_64.whl" },
{ name = "fastapi", specifier = "==0.118.0" },
{ name = "flash-linear-attention", marker = "extra == 'fast'", specifier = "==0.5.2" },
{ name = "semif-phase1", marker = "extra == 'model'", git = "https://github.com/TheoLeeCJ/SemIf-OpenJev?rev=23cf1f39fc9534fe81437200959b6dfc7106e45a" },
{ name = "torch", marker = "extra == 'model'", specifier = "==2.10.0", index = "https://download.pytorch.org/whl/cu128" },
{ name = "uvicorn", specifier = "==0.37.0" },
]
provides-extras = ["model"]
provides-extras = ["model", "fast"]
[package.metadata.requires-dev]
dev = [