feat(semif): 0.1.3 — order averaging, fast kernels, bug-hunt hardening (Prime)

Order averaging (Prime, after the 739aa03 spike):
- A decision may set orderings: rotations|all (all only for <= 4 options). Every
  ordering goes to the engine in one shared batch.
- The reply keeps each native result and adds combined {probabilities (log-mean),
  top, agreement, spread}.
- Through the service on SemIf's labelled sets (252 rows): 78.6% -> 88.1%
  (group-bootstrap 95% CI +5.1..+14.3). Unanimous agreement is 94.5% accurate.

Fast kernels: flash-linear-attention 0.5.2 and causal-conv1d 1.7.0 are now the
default build. A/B on the empty GPU 3:
- parity with upstream went from 142/144 to 144/144;
- a ~2k-token /decide went from 169 to 92 ms server-side;
- short 3-rotation batches cost ~3-6 ms more.
triton builds a C shim at runtime, so the image carries gcc. Without it the
warm-up failed and startup failed closed.

Heid bug-hunt panel (4/4 arms, thread 01M3H3F4RR7XBP90KQ3A39H4SX), folded:
- Startup validation: VRAM cap 0 no longer means uncapped (C1); limits must be
  >= 1 (S1); the token must be visible ASCII (S2); the calibration file must
  exist and parse, with T in [0.05, 20] (S8, and C3's NaN leg).
- The body limit is checked before a chunk is kept, and a Unicode-digit
  Content-Length no longer crashes (C2, S3).
- Failures while building the response now get the 500 envelope (C3).
- 429 busy past SEMIF_MAX_QUEUE requests in progress (C6).
- The engine releases memory on every non-validation failure, unchained after
  gc; an empty OOM message is handled; 'out of memory' RuntimeErrors map to 503
  (C4, C5, S9).
- The entry point forces HF_HUB_OFFLINE (S10). README wording fixed (S5, S6).
- New guard tests close the gaps the arms' mutation grids exposed: early stop of
  the body read, a shared-route lock, calibration pass-through, the gc cycle,
  the exact caps, TorchEngine.load's arch and device checks, and the offline
  entry point.
86 tests.

Deployed on fv-ml1 GPU 1: parity 144/144, OOM and burst release verified, shared
capacity 63/51/26/16 rows at ~140/520/1960/3900 prefix tokens.
This commit is contained in:
vh
2026-09-27 03:27:15 -07:00
parent d7ad235365
commit 77b8cb449c
22 changed files with 1018 additions and 92 deletions
+11 -27
View File
@@ -144,33 +144,17 @@ _As of 2026-09-27 ~0255 PT._
### SemIf on fv-ml1 GPU 1 (2026-09-27, Prime) ### SemIf on fv-ml1 GPU 1 (2026-09-27, Prime)
- **LIVE: `semif-serve` 0.1.2** at `http://10.251.50.54:8032` (`semif.fv.internal`). It wraps the SemIf - **LIVE: `semif-serve` 0.1.3** at `http://10.251.50.54:8032` (`semif.fv.internal`), with order averaging
option-logit scorers (pinned `23cf1f39`, Qwen3.5-4B `851bf6e8`, BF16); token `semif/api-token`. and the fast kernels. SemIf `23cf1f39`, Qwen3.5-4B `851bf6e8`, BF16; token `semif/api-token`. Code +
Code + contract: `services/semif-serve/`; stack `stacks/semif`. **No consumer yet.** contract: `services/semif-serve/`; stack `stacks/semif`. **No consumer yet.**
- Acceptance: 142/144 top choice vs upstream (both misses are exact bf16 ties), 144/144 prompt hashes, - 0.1.3 acceptance: 144/144 parity with upstream; averaging through the service 78.6% → 88.1%
deterministic, negative control 14/144. It holds 12 GiB hard with release after a burst (~8.7 GB resting). (95% CI +5.1..+14.3); unanimous rotations 94.5% accurate. Latency, envelope and the fast-kernel A/B
- Latency from nh3-dev: `/decide` short 71 ms end to end (38 ms server); 3 rotations 113 ms; ~2k-token are in `stacks/semif/README.md`.
state 210 ms (`services/semif-serve/spike/latency-2026-09-27.txt`). - The heid bug-hunt panel (thread `01M3H3F4RR7XBP90KQ3A39H4SX`) is triaged and folded into 0.1.3:
- **NEXT, in this order (Prime 0250):** C1–C6, S1–S3, S5, S6, S8–S10 fixed with tests; S4 settled; S7 (publish on all interfaces) accepted
1. **Build order-averaging into semif-serve as 0.1.3.** Spike (`739aa03`): 3 rotations, log-mean, as LAN + auth.
take accuracy 78.6% → 87.7% (group-bootstrap 95% CI +4.7..+13.8 pts); unanimous rows are 94.4% - NVFP4 is not worth it (see the SemIf detail files). Prime's probes (2AM/2PM, dragon/lottery) are
accurate, split rows 75.8%. Design: opt-in `orderings: rotations|all` (all only ≤ 4 options), recorded in [[2026-09-27-semif-order-averaging]].
every ordering in one shared batch, per-ordering SemIf results returned unchanged + a
`combined` block (log-mean probabilities, top, agreement, per-option spread). Orderings count
toward `max_decisions`. Contract amendment → TDD → deploy → acceptance (authored144 accuracy
with/without; 2AM/2PM agreement as controls).
2. **Trial the fast kernels** (`causal_conv1d`, `flash-linear-attention`). Today transformers runs
Qwen3.5's reference PyTorch paths and logs that they are "much slower". Adopt only if the
authored144 parity re-check holds.
3. **heid bug-hunt panel** on 0.1.2 is pending (thread `01M3H3F4RR7XBP90KQ3A39H4SX`). Triage and fold
its findings into 0.1.3.
- NVFP4 is not worth it: the SemIf scorer runs on transformers, not vLLM; quant noise lands on the scores
and would need its own calibration; BF16 already fits. SemIf's own MLX 4-bit run moved authored144
0.813 → 0.789 (one run each, so indicative only).
- Prime's probes (single scenarios, not benchmarks): at 2AM vs 2PM, booty-call probability averaged over
all orderings goes 0.42 → 0.05; a baby dragon reads life-changing 0.65 / world-changing 0.25 with only
2/3 rotations agreeing (one ordering said 0.925), the lottery ticket reads life-changing unanimously,
and the paperclip control reads trivial 0.999.
### restic: credential leak fixed (2026-09-27, Prime) ### restic: credential leak fixed (2026-09-27, Prime)
+12 -3
View File
@@ -22,18 +22,26 @@ l.write_text(re.sub(r'(name = "semif-serve"\nversion = )"[^"]+"', r'\1"0.0.0"',
EOF EOF
FROM python:3.12-slim-bookworm FROM python:3.12-slim-bookworm
# EXTRAS picks the optional dependency sets. Default (adopted 2026-09-27): the model plus
# Qwen3.5's fast kernels (fla + causal-conv1d). "--extra model" alone gives the reference
# PyTorch paths.
ARG EXTRAS="--extra model --extra fast"
COPY --from=ghcr.io/astral-sh/uv:0.6.9 /uv /bin/uv COPY --from=ghcr.io/astral-sh/uv:0.6.9 /uv /bin/uv
ENV UV_COMPILE_BYTECODE=1 UV_LINK_MODE=copy UV_PYTHON_DOWNLOADS=never ENV UV_COMPILE_BYTECODE=1 UV_LINK_MODE=copy UV_PYTHON_DOWNLOADS=never
# git: semif-phase1 installs from a pinned GitHub commit. # git: semif-phase1 installs from a pinned GitHub commit. The fast extra also needs a C
# compiler AT RUNTIME: triton builds its CUDA driver shim on first use, and without gcc the
# warm-up dies with "Failed to find C compiler" (seen 2026-09-27; startup failed closed).
RUN apt-get update && apt-get install -y --no-install-recommends git ca-certificates \ RUN apt-get update && apt-get install -y --no-install-recommends git ca-certificates \
&& if echo "$EXTRAS" | grep -q -- '--extra fast'; then \
apt-get install -y --no-install-recommends gcc libc6-dev; fi \
&& rm -rf /var/lib/apt/lists/* && rm -rf /var/lib/apt/lists/*
WORKDIR /app WORKDIR /app
COPY --from=deps /deps/pyproject.toml /deps/uv.lock ./ COPY --from=deps /deps/pyproject.toml /deps/uv.lock ./
RUN --mount=type=cache,target=/root/.cache/uv \ RUN --mount=type=cache,target=/root/.cache/uv \
uv sync --frozen --no-dev --extra model --no-install-project uv sync --frozen --no-dev $EXTRAS --no-install-project
COPY pyproject.toml uv.lock ./ COPY pyproject.toml uv.lock ./
COPY src ./src COPY src ./src
RUN uv sync --frozen --no-dev --extra model --no-editable --no-cache RUN uv sync --frozen --no-dev $EXTRAS --no-editable --no-cache
RUN groupadd --system --gid 10001 semif \ RUN groupadd --system --gid 10001 semif \
&& useradd --system --uid 10001 --gid 10001 --no-create-home --shell /usr/sbin/nologin semif && useradd --system --uid 10001 --gid 10001 --no-create-home --shell /usr/sbin/nologin semif
USER semif USER semif
@@ -41,6 +49,7 @@ ENV PATH=/app/.venv/bin:$PATH \
HF_HOME=/hf \ HF_HOME=/hf \
HF_HUB_OFFLINE=1 \ HF_HUB_OFFLINE=1 \
HF_HUB_DISABLE_TELEMETRY=1 \ HF_HUB_DISABLE_TELEMETRY=1 \
TRITON_CACHE_DIR=/tmp/triton-cache \
NVIDIA_DRIVER_CAPABILITIES=compute,utility NVIDIA_DRIVER_CAPABILITIES=compute,utility
EXPOSE 8000 EXPOSE 8000
# One worker (INV-2): the model and the inference lock live in this one process. # One worker (INV-2): the model and the inference lock live in this one process.
@@ -0,0 +1,18 @@
{
"rows": 252,
"groups": 72,
"accuracy_plain": 0.7857,
"accuracy_rotations": 0.881,
"delta_95ci": [
0.0512,
0.1429
],
"unanimous": {
"rows": 163,
"accuracy": 0.9448
},
"split": {
"rows": 89,
"accuracy": 0.764
}
}
@@ -0,0 +1,42 @@
"""Order-averaging acceptance through the SERVICE (0.1.3): the same 252 labelled rows as the spike
(SemIf authored144 + perturbations108), each scored twice: plain /decide (the caller's order), and
/decide with orderings=rotations (combined.top). Paired group bootstrap for the accuracy delta.
SEMIF_DIR=... SEMIF_URL=... SEMIF_TOKEN=... uv run --with httpx python averaging.py out.json
"""
import json, os, random, statistics as st, sys
from pathlib import Path
import httpx
S, U = Path(os.environ["SEMIF_DIR"]), os.environ["SEMIF_URL"]
H = {"Authorization": f"Bearer {os.environ['SEMIF_TOKEN']}"}
rows = []
for name in ("authored144", "perturbations108"):
rows += [dict(json.loads(l), set=name) for l in (S / f"benchmarks/data/{name}.jsonl").read_text().splitlines() if l.strip()]
out = []
with httpx.Client(timeout=120) as c:
for r in rows:
base = {k: r[k] for k in ("id", "state", "question", "options")}
plain = c.post(f"{U}/decide", headers=H, json=base).json()
avg = c.post(f"{U}/decide", headers=H, json={**base, "orderings": "rotations"}).json()
gold = r["options"][r["label"]]["id"]
top_plain = plain["option_ids"][plain["probabilities"].index(max(plain["probabilities"]))]
out.append({"group": r["group_id"], "plain": top_plain == gold, "rotations": avg["combined"]["top"] == gold,
"agreement": avg["combined"]["agreement"]})
groups = {}
for o in out:
groups.setdefault(o["group"], []).append(o)
rng, keys, deltas = random.Random(7), list(groups), []
for _ in range(10000):
sample = [o for g in (rng.choice(keys) for _ in keys) for o in groups[g]]
deltas.append(st.fmean(o["rotations"] for o in sample) - st.fmean(o["plain"] for o in sample))
deltas.sort()
unan = [o for o in out if o["agreement"] == 1.0]
split = [o for o in out if o["agreement"] < 1.0]
report = {"rows": len(out), "groups": len(groups),
"accuracy_plain": round(st.fmean(o["plain"] for o in out), 4),
"accuracy_rotations": round(st.fmean(o["rotations"] for o in out), 4),
"delta_95ci": [round(deltas[250], 4), round(deltas[9750], 4)],
"unanimous": {"rows": len(unan), "accuracy": round(st.fmean(o["rotations"] for o in unan), 4)},
"split": {"rows": len(split), "accuracy": round(st.fmean(o["rotations"] for o in split), 4) if split else None}}
json.dump(report, open(sys.argv[1], "w"), indent=1)
print(json.dumps(report))
@@ -0,0 +1,63 @@
{
"url": "http://10.251.50.54:8032",
"health": {
"status": "ok",
"semif_commit": "23cf1f39fc9534fe81437200959b6dfc7106e45a",
"model": {
"source": "Qwen/Qwen3.5-4B",
"revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a",
"dtype": "bfloat16",
"device": "cuda:0",
"torch_version": "2.10.0+cu128",
"transformers_version": "5.17.0",
"device_name": "NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition",
"allocated_gib": 7.84,
"reserved_gib": 8.12
},
"vram_cap_gib": 12.0,
"max_tokens": 4096,
"max_decisions": 64,
"workloads": []
},
"1_parity_vs_upstream": {
"rows": 144,
"top_choice_agree": 144,
"max_abs_prob_gap": 0.049096847865749194
},
"1_prompt_sha256_equal": 144,
"2_noise_floor_a_vs_b": {
"rows": 144,
"top_choice_agree": 144,
"max_abs_prob_gap": 0.0
},
"3_negative_rotated_options": {
"rows": 144,
"top_choice_agree": 14,
"max_abs_prob_gap": 0.9985836128543327
},
"4_shared_vs_direct": {
"groups": 36,
"rows": 72,
"top_choice_agree": 72,
"max_abs_prob_gap": 0.046346781311727814
},
"5_speed_21_binary": {
"prefix_tokens": 62,
"shared_s": {
"runs": [
0.139055563005968,
0.1375489159981953,
0.13897942200128455
],
"median": 0.13897942200128455
},
"sequential_decide_s": {
"runs": [
0.9357447380025405,
0.9384197259932989,
0.939685705001466
],
"median": 0.9384197259932989
}
}
}
+8 -1
View File
@@ -1,6 +1,6 @@
[project] [project]
name = "semif-serve" name = "semif-serve"
version = "0.1.2" version = "0.1.3"
description = "HTTP wrapper around SemIf's direct and shared option-logit scorers" description = "HTTP wrapper around SemIf's direct and shared option-logit scorers"
requires-python = ">=3.12" requires-python = ">=3.12"
dependencies = [ dependencies = [
@@ -17,6 +17,13 @@ model = [
"torch==2.10.0", "torch==2.10.0",
] ]
# Qwen3.5's fast kernels. Without them transformers runs its reference PyTorch paths
# ("correct but much slower"). Trialled 2026-09-27; adopted only if parity with upstream holds.
fast = [
"flash-linear-attention==0.5.2",
"causal-conv1d @ https://github.com/Dao-AILab/causal-conv1d/releases/download/v1.7.0/causal_conv1d-1.7.0+cu12torch2.10cxx11abiTRUE-cp312-cp312-linux_x86_64.whl ; sys_platform == 'linux' and platform_machine == 'x86_64'",
]
[dependency-groups] [dependency-groups]
dev = ["pytest==8.4.2", "httpx==0.28.1"] dev = ["pytest==8.4.2", "httpx==0.28.1"]
+66 -6
View File
@@ -45,6 +45,37 @@ probabilities}` with `softmax(option_logits / T)`. The argmax never changes. The
native `probabilities` and `probability_status` stay untouched. An unknown native `probabilities` and `probability_status` stay untouched. An unknown
`workload` is a 422. With no `workload`, no `calibrated` key appears. `workload` is a 422. With no `workload`, no `calibrated` key appears.
**Order averaging (0.1.3, Prime 2026-09-27).** A decision (the `/decide` body, or an
entry in `decisions`) may set `orderings`, whose default is `"none"`:
- `"rotations"`: the n cyclic shifts of the caller's option list, starting with
the caller's order, so every option sits in every position exactly once.
- `"all"`: every permutation (n!), the caller's order first. It is allowed only
when n ≤ 4, else 422.
Every ordering of every decision in the request becomes its own SemIf row
(`id` = `"<id>#o<k>"`, same state and question, reordered options). **All rows go
to the engine in ONE `shared` call**, which includes `/decide`. The result for an
averaged decision is:
```
{id, option_ids (caller's order),
combined: {method, orderings: n, probabilities, top, agreement, spread: {option_id: [min, max]}},
orderings: [the SemIf result for each ordering, unchanged]}
```
- `probabilities`: per ordering, log-softmax of `option_logits`; averaged per
option id; renormalised; reported in the caller's order.
- `agreement`: the fraction of orderings whose top option equals `combined.top`.
- `spread`: each option's min and max native probability across orderings.
Every expanded row counts toward `SEMIF_MAX_DECISIONS`. `workload` together with
`orderings` is a 422, because a temperature is fitted per method and none is
fitted on combined scores yet. Decisions without `orderings` keep the exact
pre-0.1.3 result shape. Averaging cancels any additive position bias exactly.
Measured by the 2026-09-27 spike: 3 rotations take SemIf's labelled sets from
78.6% to 87.7% accuracy, and unanimous agreement is 94.4% accurate.
## Invariants ## Invariants
- **INV-1 pass-through.** `option_ids`, `probabilities`, `option_logits`, - **INV-1 pass-through.** `option_ids`, `probabilities`, `option_logits`,
@@ -64,10 +95,18 @@ native `probabilities` and `probability_status` stay untouched. An unknown
baseline by more than 512 MiB, the engine calls `empty_cache()`. A burst must not baseline by more than 512 MiB, the engine calls `empty_cache()`. A burst must not
keep the card's shared headroom: on 2026-09-27 a 64-decision request left the keep the card's shared headroom: on 2026-09-27 a 64-decision request left the
process holding 12.6 GB, leaving scriberr 3.5 GB. process holding 12.6 GB, leaving scriberr 3.5 GB.
**Any other scorer failure except `ValueError`** also releases before it is
reported: it is logged with its traceback, then re-raised unchained as
`ScoringFailed` after `gc.collect()` + `empty_cache()` (bug hunt C5). A
`RuntimeError` whose message says "out of memory" (cuBLAS/cuDNN allocation
failures) counts as an OOM → 503 (S9). An OOM with an empty message is reported
as "CUDA out of memory" rather than crashing the handler (C4).
- **INV-5 no network at runtime.** Weights come from the mounted HF cache at the - **INV-5 no network at runtime.** Weights come from the mounted HF cache at the
pinned revision (`HF_HUB_OFFLINE=1`). pinned revision. The entry point sets `HF_HUB_OFFLINE=1` itself before torch or
transformers load, so this holds outside the image too (S10).
- **INV-6 constant-time auth.** Token comparison uses `hmac.compare_digest`. The - **INV-6 constant-time auth.** Token comparison uses `hmac.compare_digest`. The
token is ≥ 32 characters, and startup refuses a shorter one. token is ≥ 32 characters of visible ASCII (33–126). Startup refuses anything
else, because a CR, LF or NUL in the token can never arrive in a header (S2).
## Limits and errors ## Limits and errors
@@ -75,15 +114,21 @@ native `probabilities` and `probability_status` stay untouched. An unknown
422, never truncated (SemIf raises). 422, never truncated (SemIf raises).
- `SEMIF_MAX_DECISIONS` (default 64) caps `decisions` per shared request. It must - `SEMIF_MAX_DECISIONS` (default 64) caps `decisions` per shared request. It must
hold 1..max entries, else 422. hold 1..max entries, else 422.
- Request body ≤ `SEMIF_MAX_BODY_BYTES` (default 1 MiB), else 413. - Request body ≤ `SEMIF_MAX_BODY_BYTES` (default 1 MiB), else 413. The limit is
checked **before** each chunk is kept, so no more than the limit is ever held
(C2). A declared `Content-Length` is trusted only if it is ASCII digits (S3).
- **Admission:** at most `SEMIF_MAX_QUEUE` (default 32) POSTs may be in progress,
counting both queued and scoring. The next one is refused with 429 `busy`
before its body is read (C6).
| status | code | when | | status | code | when |
|---|---|---| |---|---|---|
| 401 | `unauthorized` | missing or wrong bearer | | 401 | `unauthorized` | missing or wrong bearer |
| 413 | `request_too_large` | body over the limit | | 413 | `request_too_large` | body over the limit |
| 422 | `invalid_request` | bad JSON shape, a SemIf `ValueError` (validation, token limit, tokenisation), unknown workload, too many decisions | | 422 | `invalid_request` | bad JSON shape, a SemIf `ValueError` (validation, token limit, tokenisation), unknown workload, too many decisions |
| 429 | `busy` | `SEMIF_MAX_QUEUE` requests already in progress |
| 503 | `out_of_memory` | CUDA OOM during scoring | | 503 | `out_of_memory` | CUDA OOM during scoring |
| 500 | `scoring_failed` | any other scorer exception | | 500 | `scoring_failed` | any other scorer exception, **or a failure while building the response** from a scorer result (calibration, averaging): always the envelope, never a bare 500 (C3) |
The error body is `{error: {code, message}}`. The error body is `{error: {code, message}}`.
@@ -92,7 +137,16 @@ The error body is `{error: {code, message}}`.
`SEMIF_API_TOKEN` (required), `SEMIF_MODEL` (default `Qwen/Qwen3.5-4B`), `SEMIF_API_TOKEN` (required), `SEMIF_MODEL` (default `Qwen/Qwen3.5-4B`),
`SEMIF_REVISION` (default the pinned SHA), `SEMIF_DEVICE` (default `cuda`), `SEMIF_REVISION` (default the pinned SHA), `SEMIF_DEVICE` (default `cuda`),
`SEMIF_VRAM_CAP_GIB`, `SEMIF_MAX_TOKENS`, `SEMIF_MAX_DECISIONS`, `SEMIF_VRAM_CAP_GIB`, `SEMIF_MAX_TOKENS`, `SEMIF_MAX_DECISIONS`,
`SEMIF_MAX_BODY_BYTES`, `SEMIF_CALIBRATION` (path to a JSON `{workload: T}`; T > 0). `SEMIF_MAX_BODY_BYTES`, `SEMIF_MAX_QUEUE`, `SEMIF_CALIBRATION` (path to a JSON
`{workload: T}`).
Startup validates every value and refuses a bad one with a `ValueError` naming
the variable (C1, S1, S8):
- the VRAM cap, when set, is finite and > 0 (`0` used to mean uncapped);
- every limit is an integer ≥ 1;
- each T is a finite number, not a bool, in [0.05, 20] (a tiny T overflowed to
NaN, and the response then failed to render);
- the calibration file must exist and parse.
## Tests (TDD, fake scorer: no torch, no model) ## Tests (TDD, fake scorer: no torch, no model)
@@ -109,7 +163,13 @@ alone; any
other exception → 500; malformed other exception → 500; malformed
JSON or a wrong body shape → 422; too many decisions → 422; an oversized body → 413; requests JSON or a wrong body shape → 422; too many decisions → 422; an oversized body → 413; requests
are serialised (two concurrent calls never overlap inside the scorer); `/health` are serialised (two concurrent calls never overlap inside the scorer); `/health`
answers while a scorer call is blocked. answers while a scorer call is blocked. **Averaging:** `rotations` sends n rows
in one shared call, each option once per position, with ids `<id>#o<k>`, and
cancels a position bias exactly; `all` sends n! rows and is 422 above 4 options;
agreement and spread are computed from the orderings; a mixed shared request
(averaged + plain) is one engine call, with results in request order and plain
results unchanged; expanded rows count toward the cap; `workload` + `orderings`
→ 422.
## Acceptance (on fv-ml1, real model; not unit tests) ## Acceptance (on fv-ml1, real model; not unit tests)
@@ -0,0 +1,7 @@
condition e2e p50 (runs) server p50 (runs)
health (floor) 33.5 [31.7-34.7] -
decide, short (~130 tok) 69.3 [68.8-70.1] 35.3 [35.3-35.4]
decide, long (~2,000 tok) 130.5 [129.5-130.5] 92.1 [91.7-92.4]
shared, 3 rotations (short) 115.0 [112.4-115.7] 77.9 [77.1-79.3]
shared, 6 orderings (short) 118.3 [117.5-118.7] 81.2 [80.2-81.3]
shared, 3 rotations (long) 199.5 [196.6-201.0] 158.2 [157.7-158.3]
+144 -25
View File
@@ -2,9 +2,10 @@
from __future__ import annotations from __future__ import annotations
import hmac import hmac
import itertools
import math import math
import threading import threading
from typing import Any from typing import Any, Literal
from fastapi import FastAPI, Request from fastapi import FastAPI, Request
from fastapi.concurrency import run_in_threadpool from fastapi.concurrency import run_in_threadpool
@@ -29,6 +30,7 @@ class Decision(BaseModel):
id: str id: str
question: str question: str
options: list[Option] options: list[Option]
orderings: Literal["none", "rotations", "all"] = "none"
def row(self, state: State) -> dict: def row(self, state: State) -> dict:
"""The SemIf row shape: exactly id, state, question, options.""" """The SemIf row shape: exactly id, state, question, options."""
@@ -64,6 +66,75 @@ def _first_error(exc: ValidationError) -> str:
return f"{where}: {first.get('msg', 'invalid')}" return f"{where}: {first.get('msg', 'invalid')}"
MAX_OPTIONS_FOR_ALL = 4
def ordering_perms(decision: Decision) -> list[tuple[int, ...]]:
"""Index permutations of the caller's options, the caller's own order first."""
n = len(decision.options)
if decision.orderings == "rotations":
return [tuple((start + k) % n for k in range(n)) for start in range(n)]
if n > MAX_OPTIONS_FOR_ALL:
raise ApiError(422, "invalid_request",
f"orderings 'all' allows at most {MAX_OPTIONS_FOR_ALL} options ({n} given); use 'rotations'")
return list(itertools.permutations(range(n)))
def expanded_rows(decision: Decision, state: State, perms: list[tuple[int, ...]]) -> list[dict]:
base = decision.row(state)
return [{**base, "id": f"{decision.id}#o{k}", "options": [base["options"][i] for i in perm]}
for k, perm in enumerate(perms)]
def combine(decision: Decision, results: list[dict]) -> dict:
"""Average per-ordering log-softmax by option id; the native results ride along unchanged."""
option_ids = [o.id for o in decision.options]
logp: dict[str, list[float]] = {i: [] for i in option_ids}
probs: dict[str, list[float]] = {i: [] for i in option_ids}
tops = []
for result in results:
logits = result["option_logits"]
top = max(logits)
lse = top + math.log(sum(math.exp(x - top) for x in logits))
for oid, x, p in zip(result["option_ids"], logits, result["probabilities"]):
logp[oid].append(x - lse)
probs[oid].append(p)
tops.append(result["option_ids"][logits.index(top)])
means = [sum(logp[i]) / len(logp[i]) for i in option_ids]
peak = max(means)
weights = [math.exp(m - peak) for m in means]
combined_p = [w / sum(weights) for w in weights]
winner = option_ids[combined_p.index(max(combined_p))]
return {
"id": decision.id,
"option_ids": option_ids,
"combined": {
"method": decision.orderings,
"orderings": len(results),
"probabilities": combined_p,
"top": winner,
"agreement": tops.count(winner) / len(tops),
"spread": {i: [min(probs[i]), max(probs[i])] for i in option_ids},
},
"orderings": results,
}
async def read_limited(stream, declared: str | None, limit: int) -> bytes:
"""Read a request body, refusing it once it would exceed `limit` bytes. The check runs BEFORE a
chunk is kept, and nothing after the crossing chunk is read (bug hunt C2). A declared length is
trusted only as ASCII digits: `"²".isdigit()` is True but `int("²")` raises (S3)."""
too_large = ApiError(413, "request_too_large", f"request body exceeds {limit} bytes")
if declared is not None and declared.isascii() and declared.isdigit() and int(declared) > limit:
raise too_large
body = bytearray()
async for chunk in stream:
if len(body) + len(chunk) > limit:
raise too_large
body.extend(chunk)
return bytes(body)
def calibrated_view(result: dict, workload: str, temperature: float) -> dict: def calibrated_view(result: dict, workload: str, temperature: float) -> dict:
"""softmax(option_logits / T): the native fields are left exactly as SemIf returned them (INV-1).""" """softmax(option_logits / T): the native fields are left exactly as SemIf returned them (INV-1)."""
scaled = [x / temperature for x in result["option_logits"]] scaled = [x / temperature for x in result["option_logits"]]
@@ -77,6 +148,27 @@ def create_app(settings: Settings, engine: Any) -> FastAPI:
app = FastAPI(title="semif-serve") app = FastAPI(title="semif-serve")
expected = f"Bearer {settings.api_token}".encode() expected = f"Bearer {settings.api_token}".encode()
inference = threading.Lock() # INV-2: one scorer call at a time, off the event loop inference = threading.Lock() # INV-2: one scorer call at a time, off the event loop
in_progress = 0 # POSTs admitted and not yet answered (bug hunt C6)
def admit():
nonlocal in_progress
if in_progress >= settings.max_queue:
raise ApiError(429, "busy", f"{in_progress} requests already in progress (limit {settings.max_queue})")
in_progress += 1
def leave():
nonlocal in_progress
in_progress -= 1
def build(fn, *args):
"""Response construction from a scorer result (calibration, averaging) maps its failures to
the 500 envelope too, instead of escaping as a bare 500 (bug hunt C3)."""
try:
return fn(*args)
except ApiError:
raise
except Exception as exc: # noqa: BLE001
raise ApiError(500, "scoring_failed", f"building the response failed: {type(exc).__name__}: {exc}") from exc
def locked(fn, *args): def locked(fn, *args):
with inference: with inference:
@@ -94,22 +186,10 @@ def create_app(settings: Settings, engine: Any) -> FastAPI:
async def api_error(_request: Request, exc: ApiError): async def api_error(_request: Request, exc: ApiError):
return error(exc.status, exc.code, exc.message) return error(exc.status, exc.code, exc.message)
async def read_limited(request: Request) -> bytes:
limit = settings.max_body_bytes
too_large = ApiError(413, "request_too_large", f"request body exceeds {limit} bytes")
declared = request.headers.get("content-length")
if declared is not None and declared.isdigit() and int(declared) > limit:
raise too_large
body = bytearray()
async for chunk in request.stream(): # also caps bodies that declare no length
body.extend(chunk)
if len(body) > limit:
raise too_large
return bytes(body)
async def parse(request: Request, model: type[BaseModel]): async def parse(request: Request, model: type[BaseModel]):
try: try:
return model.model_validate_json(await read_limited(request)) return model.model_validate_json(
await read_limited(request.stream(), request.headers.get("content-length"), settings.max_body_bytes))
except ValidationError as exc: except ValidationError as exc:
raise ApiError(422, "invalid_request", _first_error(exc)) from exc raise ApiError(422, "invalid_request", _first_error(exc)) from exc
@@ -142,20 +222,59 @@ def create_app(settings: Settings, engine: Any) -> FastAPI:
"vram_cap_gib": settings.vram_cap_gib, "max_tokens": settings.max_tokens, "vram_cap_gib": settings.vram_cap_gib, "max_tokens": settings.max_tokens,
"max_decisions": settings.max_decisions, "workloads": sorted(settings.calibration)} "max_decisions": settings.max_decisions, "workloads": sorted(settings.calibration)}
async def score_batch(decisions: list[Decision], state: State, workload: str | None) -> tuple[list[dict], dict]:
"""One engine.shared call for every row of every decision; results in request order."""
plan = [] # (decision, perms or None, row count)
rows: list[dict] = []
for d in decisions:
if d.orderings == "none":
plan.append((d, None, 1))
rows.append(d.row(state))
else:
if workload is not None:
raise ApiError(422, "invalid_request",
"workload calibration is not available together with orderings")
perms = ordering_perms(d)
plan.append((d, perms, len(perms)))
rows.extend(expanded_rows(d, state, perms))
if not 1 <= len(rows) <= settings.max_decisions:
raise ApiError(422, "invalid_request",
f"this request expands to {len(rows)} scored rows; the limit is 1..{settings.max_decisions}")
temperature = temperature_for(workload)
results, timing = await score(engine.shared, rows)
out, cursor = [], 0
for d, perms, count in plan:
chunk = results[cursor:cursor + count]
cursor += count
out.append(build(with_calibration, chunk[0], workload, temperature) if perms is None
else build(combine, d, chunk))
return out, timing
@app.post("/decide") @app.post("/decide")
async def decide(request: Request): async def decide(request: Request):
body = await parse(request, DecideBody) admit()
temperature = temperature_for(body.workload) try:
return with_calibration(await score(engine.direct, body.row(body.state)), body.workload, temperature) body = await parse(request, DecideBody)
if body.orderings != "none":
results, _timing = await score_batch([body], body.state, body.workload)
return results[0]
temperature = temperature_for(body.workload)
result = await score(engine.direct, body.row(body.state))
return build(with_calibration, result, body.workload, temperature)
finally:
leave()
@app.post("/decide/shared") @app.post("/decide/shared")
async def decide_shared(request: Request): async def decide_shared(request: Request):
body = await parse(request, SharedBody) admit()
if not 1 <= len(body.decisions) <= settings.max_decisions: try:
raise ApiError(422, "invalid_request", body = await parse(request, SharedBody)
f"decisions must hold 1..{settings.max_decisions} entries, got {len(body.decisions)}") if not 1 <= len(body.decisions) <= settings.max_decisions:
temperature = temperature_for(body.workload) raise ApiError(422, "invalid_request",
results, timing = await score(engine.shared, [d.row(body.state) for d in body.decisions]) f"decisions must hold 1..{settings.max_decisions} entries, got {len(body.decisions)}")
return {"results": [with_calibration(r, body.workload, temperature) for r in results], "timing": timing} results, timing = await score_batch(body.decisions, body.state, body.workload)
return {"results": results, "timing": timing}
finally:
leave()
return app return app
+54 -12
View File
@@ -1,4 +1,9 @@
"""Settings for semif-serve. Contract: semif-serve.contract.md § Configuration.""" """Settings for semif-serve. Contract: semif-serve.contract.md § Configuration.
Every value is validated at startup and a bad one is refused with a ValueError naming the
variable: a service that starts and then rejects every request (or runs uncapped) is worse
than one that does not start (bug hunt 2026-09-27: C1, S1, S2, S8).
"""
from __future__ import annotations from __future__ import annotations
import json import json
@@ -8,6 +13,7 @@ from dataclasses import dataclass, field
from pathlib import Path from pathlib import Path
MIN_TOKEN_CHARS = 32 MIN_TOKEN_CHARS = 32
MIN_TEMPERATURE, MAX_TEMPERATURE = 0.05, 20.0
SEMIF_COMMIT = "23cf1f39fc9534fe81437200959b6dfc7106e45a" SEMIF_COMMIT = "23cf1f39fc9534fe81437200959b6dfc7106e45a"
DEFAULT_MODEL = "Qwen/Qwen3.5-4B" DEFAULT_MODEL = "Qwen/Qwen3.5-4B"
DEFAULT_REVISION = "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a" DEFAULT_REVISION = "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a"
@@ -23,35 +29,71 @@ class Settings:
max_tokens: int = 4096 max_tokens: int = 4096
max_decisions: int = 64 max_decisions: int = 64
max_body_bytes: int = 1024 * 1024 max_body_bytes: int = 1024 * 1024
max_queue: int = 32
calibration: dict[str, float] = field(default_factory=dict) calibration: dict[str, float] = field(default_factory=dict)
@classmethod @classmethod
def from_env(cls, env: Mapping[str, str]) -> "Settings": def from_env(cls, env: Mapping[str, str]) -> "Settings":
token = env.get("SEMIF_API_TOKEN", "") token = env.get("SEMIF_API_TOKEN", "")
if len(token) < MIN_TOKEN_CHARS: # INV-6 # INV-6: visible ASCII only. A CR, LF or NUL can never arrive in a header, so a token
raise ValueError(f"SEMIF_API_TOKEN must be at least {MIN_TOKEN_CHARS} characters") # carrying one would lock every caller out while /health still said ok.
cap = env.get("SEMIF_VRAM_CAP_GIB") if len(token) < MIN_TOKEN_CHARS or not all(33 <= ord(c) <= 126 for c in token):
raise ValueError(f"SEMIF_API_TOKEN must be at least {MIN_TOKEN_CHARS} visible ASCII characters")
return cls( return cls(
api_token=token, api_token=token,
model=env.get("SEMIF_MODEL", DEFAULT_MODEL), model=env.get("SEMIF_MODEL", DEFAULT_MODEL),
revision=env.get("SEMIF_REVISION", DEFAULT_REVISION), revision=env.get("SEMIF_REVISION", DEFAULT_REVISION),
device=env.get("SEMIF_DEVICE", "cuda"), device=env.get("SEMIF_DEVICE", "cuda"),
vram_cap_gib=float(cap) if cap else None, vram_cap_gib=_positive_float(env, "SEMIF_VRAM_CAP_GIB"),
max_tokens=int(env.get("SEMIF_MAX_TOKENS", 4096)), max_tokens=_positive_int(env, "SEMIF_MAX_TOKENS", 4096),
max_decisions=int(env.get("SEMIF_MAX_DECISIONS", 64)), max_decisions=_positive_int(env, "SEMIF_MAX_DECISIONS", 64),
max_body_bytes=int(env.get("SEMIF_MAX_BODY_BYTES", 1024 * 1024)), max_body_bytes=_positive_int(env, "SEMIF_MAX_BODY_BYTES", 1024 * 1024),
max_queue=_positive_int(env, "SEMIF_MAX_QUEUE", 32),
calibration=_load_calibration(env.get("SEMIF_CALIBRATION")), calibration=_load_calibration(env.get("SEMIF_CALIBRATION")),
) )
def _positive_int(env: Mapping[str, str], name: str, default: int) -> int:
raw = env.get(name)
if raw is None:
return default
try:
value = int(raw)
except ValueError:
raise ValueError(f"{name} must be an integer, got {raw!r}") from None
if value < 1:
raise ValueError(f"{name} must be >= 1, got {value}")
return value
def _positive_float(env: Mapping[str, str], name: str) -> float | None:
"""Unset means no cap. When set it must be finite and > 0: `0` used to slip through as 'no cap'."""
raw = env.get(name)
if raw is None or raw == "":
return None
try:
value = float(raw)
except ValueError:
raise ValueError(f"{name} must be a number, got {raw!r}") from None
if not math.isfinite(value) or value <= 0:
raise ValueError(f"{name} must be a finite number > 0, got {raw!r}")
return value
def _load_calibration(path: str | None) -> dict[str, float]: def _load_calibration(path: str | None) -> dict[str, float]:
"""{workload: T}, every T a finite number > 0 (T scales option logits before softmax).""" """{workload: T}; T scales option logits before softmax, so it is kept in a sane range
(a tiny T overflows to NaN and the response then fails to render)."""
if not path: if not path:
return {} return {}
table = json.loads(Path(path).read_text()) try:
table = json.loads(Path(path).read_text())
except (OSError, ValueError) as exc:
raise ValueError(f"SEMIF_CALIBRATION {path!r} could not be read as JSON: {exc}") from None
if not isinstance(table, dict) or not all( if not isinstance(table, dict) or not all(
isinstance(t, (int, float)) and not isinstance(t, bool) and math.isfinite(t) and t > 0 isinstance(t, (int, float)) and not isinstance(t, bool) and math.isfinite(t)
and MIN_TEMPERATURE <= t <= MAX_TEMPERATURE
for t in table.values() for t in table.values()
): ):
raise ValueError("SEMIF_CALIBRATION must be a JSON object of workload -> finite temperature > 0") raise ValueError(f"SEMIF_CALIBRATION must be a JSON object of workload -> temperature in "
f"[{MIN_TEMPERATURE}, {MAX_TEMPERATURE}]")
return {str(k): float(v) for k, v in table.items()} return {str(k): float(v) for k, v in table.items()}
+27 -7
View File
@@ -7,10 +7,14 @@ unit-tested against a fake torch (tests/test_engine.py).
from __future__ import annotations from __future__ import annotations
import gc import gc
import logging
import traceback
from typing import Any, Callable from typing import Any, Callable
from .config import Settings from .config import Settings
from .errors import OutOfMemory from .errors import OutOfMemory, ScoringFailed
log = logging.getLogger("semif_serve.engine")
RELEASE_SLACK_BYTES = 512 * 2**20 RELEASE_SLACK_BYTES = 512 * 2**20
WARMUP_ROW = { WARMUP_ROW = {
@@ -24,6 +28,11 @@ WARMUP_ROW = {
} }
def _first_line(exc: BaseException) -> str:
lines = str(exc).splitlines()
return lines[0] if lines else ""
class TorchEngine: class TorchEngine:
def __init__(self, torch: Any, model: Any, tokenizer: Any, metadata: dict, settings: Settings, def __init__(self, torch: Any, model: Any, tokenizer: Any, metadata: dict, settings: Settings,
direct_fn: Callable, shared_fn: Callable, release_above_bytes: int | None = None): direct_fn: Callable, shared_fn: Callable, release_above_bytes: int | None = None):
@@ -46,7 +55,7 @@ class TorchEngine:
arch = f"sm_{major}{minor}" arch = f"sm_{major}{minor}"
if arch not in torch.cuda.get_arch_list(): # INV-3: no silent PTX/CPU fallback if arch not in torch.cuda.get_arch_list(): # INV-3: no silent PTX/CPU fallback
raise RuntimeError(f"torch {torch.__version__} has no kernels for {arch}: {torch.cuda.get_arch_list()}") raise RuntimeError(f"torch {torch.__version__} has no kernels for {arch}: {torch.cuda.get_arch_list()}")
if settings.vram_cap_gib: # INV-4: cap BEFORE the weights land if settings.vram_cap_gib is not None: # INV-4: cap BEFORE the weights land
total = torch.cuda.get_device_properties(0).total_memory total = torch.cuda.get_device_properties(0).total_memory
fraction = settings.vram_cap_gib * 2**30 / total fraction = settings.vram_cap_gib * 2**30 / total
if not 0 < fraction <= 1: if not 0 < fraction <= 1:
@@ -82,17 +91,28 @@ class TorchEngine:
def _guard(self, fn, *args): def _guard(self, fn, *args):
try: try:
result = fn(*args) result = fn(*args)
except ValueError:
raise # validation: SemIf raises it before any GPU work
except self._torch.cuda.OutOfMemoryError as exc: except self._torch.cuda.OutOfMemoryError as exc:
message = str(exc).splitlines()[0] failure, message = OutOfMemory, _first_line(exc) or "CUDA out of memory"
except Exception as exc: # noqa: BLE001 — every other failure is released and reported below
message = _first_line(exc)
if "out of memory" in message.lower(): # cuBLAS/cuDNN allocation failures
failure = OutOfMemory
else:
failure, message = ScoringFailed, f"{type(exc).__name__}: {message}"
# Formatted text, not exc_info: a log record that keeps the traceback object alive
# (pytest's capture handler does; so would any buffering handler) pins the tensors.
log.error("scorer failed:\n%s", traceback.format_exc())
else: else:
self._release_burst() self._release_burst()
return result return result
# INV-4, outside the except block on purpose: the torch exception's traceback holds the # INV-4, outside the except block on purpose: the exception's traceback holds the failed
# failed scorer's frames, and with them its tensors (the replicated prefix cache). Raising # scorer's frames, and with them its tensors (the replicated prefix cache). Raising inside
# inside the block, or `from exc`, would chain to it and keep GiBs allocated after the 503. # the block, or `from exc`, would chain to it and keep GiBs allocated after the response.
gc.collect() gc.collect()
self._torch.cuda.empty_cache() self._torch.cuda.empty_cache()
raise OutOfMemory(message) raise failure(message)
def direct(self, row: dict) -> dict: def direct(self, row: dict) -> dict:
return self._guard(self._direct, self._model, self._tokenizer, row, self._metadata, self._settings.max_tokens) return self._guard(self._direct, self._model, self._tokenizer, row, self._metadata, self._settings.max_tokens)
@@ -3,3 +3,8 @@
class OutOfMemory(RuntimeError): class OutOfMemory(RuntimeError):
"""The engine ran out of GPU memory during a request and has already released its cache (INV-4).""" """The engine ran out of GPU memory during a request and has already released its cache (INV-4)."""
class ScoringFailed(RuntimeError):
"""A scorer call failed for a reason other than validation or OOM. Raised unchained, after the
failed call's memory has been released; the original traceback is logged, not carried (INV-4)."""
@@ -11,6 +11,10 @@ from .config import Settings
def app_from_env() -> FastAPI: def app_from_env() -> FastAPI:
settings = Settings.from_env(os.environ) settings = Settings.from_env(os.environ)
# INV-5: never download at runtime, inside the image or out of it (bug hunt S10). Set before
# torch / transformers / huggingface_hub are imported, since they read it at import time.
os.environ["HF_HUB_OFFLINE"] = "1"
os.environ["TRANSFORMERS_OFFLINE"] = "1"
from .engine import TorchEngine # torch loads only here, never in the unit tests from .engine import TorchEngine # torch loads only here, never in the unit tests
return create_app(settings, TorchEngine.load(settings)) return create_app(settings, TorchEngine.load(settings))
+103
View File
@@ -1,5 +1,7 @@
"""semif-serve HTTP behaviour against a fake engine (no torch, no model). """semif-serve HTTP behaviour against a fake engine (no torch, no model).
Contract: services/semif-serve/semif-serve.contract.md""" Contract: services/semif-serve/semif-serve.contract.md"""
import json
import pytest import pytest
from fastapi.testclient import TestClient from fastapi.testclient import TestClient
@@ -231,3 +233,104 @@ def test_health_reports_the_pins_limits_and_workloads():
body = client.get("/health").json() body = client.get("/health").json()
assert body == {"status": "ok", "semif_commit": SEMIF_COMMIT, "model": FakeEngine().health(), assert body == {"status": "ok", "semif_commit": SEMIF_COMMIT, "model": FakeEngine().health(),
"vram_cap_gib": 12.0, "max_tokens": 4096, "max_decisions": 8, "workloads": ["alerts", "triage"]} "vram_cap_gib": 12.0, "max_tokens": 4096, "max_decisions": 8, "workloads": ["alerts", "triage"]}
def test_a_body_of_exactly_the_limit_is_accepted():
body = json.dumps(ROW).encode()
client = make_client(max_body_bytes=len(body))
assert client.post("/decide", content=body, headers={**AUTH, "content-type": "application/json"}).status_code == 200
def test_a_non_ascii_digit_content_length_is_ignored_not_a_crash():
"""HTTP clients cannot send one (httpx refuses; h11 rejects it), so check the reader directly."""
import asyncio
from semif_serve.app import read_limited
async def body():
yield b"{}"
assert asyncio.run(read_limited(body(), "²", 100)) == b"{}" # int("²") would raise
def test_exactly_max_decisions_is_accepted():
decisions = [{"id": str(i), "question": "Q?", "options": OPTIONS} for i in range(3)]
client = make_client(max_decisions=3)
assert client.post("/decide/shared", json={"state": "s", "decisions": decisions}, headers=AUTH).status_code == 200
def test_calibration_leaves_every_native_field_alone():
engine = FakeEngine(logits=(3.0, 1.0))
body = make_client(engine, calibration={"triage": 2.0}).post(
"/decide", json={**ROW, "workload": "triage"}, headers=AUTH).json()
body.pop("calibrated")
assert body == engine.direct(ROW)
class MalformedEngine(FakeEngine):
def direct(self, row):
return {"id": row["id"], "option_ids": ["yes", "no"], "probabilities": [0.5, 0.5]} # no option_logits
def shared(self, rows):
return [self.direct(r) for r in rows], {}
@pytest.mark.parametrize("path, body", [
("/decide", {**ROW, "workload": "triage"}),
("/decide", {**ROW, "orderings": "rotations"}),
])
def test_a_malformed_scorer_result_is_an_envelope_500_not_a_bare_one(path, body):
client = make_client(MalformedEngine(), calibration={"triage": 2.0})
response = client.post(path, json=body, headers=AUTH)
assert response.status_code == 500
assert response.json()["error"]["code"] == "scoring_failed"
class SharedSlowEngine(SlowEngine):
def shared(self, rows):
return [self.direct(r) for r in rows], {}
def test_shared_requests_are_serialised_too():
from concurrent.futures import ThreadPoolExecutor
engine = SharedSlowEngine()
body = {"state": "s", "decisions": [{"id": "a", "question": "Q?", "options": OPTIONS}]}
with make_client(engine) as client, ThreadPoolExecutor(3) as pool:
futures = [pool.submit(client.post, "/decide/shared", json=body, headers=AUTH) for _ in range(3)]
assert engine.entered.wait(5)
engine.release.set()
assert [f.result().status_code for f in futures] == [200] * 3
assert engine.peak == 1
def test_more_than_max_queue_requests_in_progress_get_429_busy():
from concurrent.futures import ThreadPoolExecutor
engine = SlowEngine()
with make_client(engine, max_queue=2) as client, ThreadPoolExecutor(3) as pool:
held = [pool.submit(client.post, "/decide", json={**ROW, "id": f"r{i}"}, headers=AUTH) for i in range(2)]
assert engine.entered.wait(5)
import time
deadline = time.monotonic() + 5
while engine.inside + 0 < 1 and time.monotonic() < deadline:
time.sleep(0.01)
time.sleep(0.2) # let the second request reach the queue
extra = client.post("/decide", json={**ROW, "id": "extra"}, headers=AUTH)
assert extra.status_code == 429 and extra.json()["error"]["code"] == "busy"
engine.release.set()
assert [f.result().status_code for f in held] == [200, 200]
assert make_client(FakeEngine(), max_queue=2).post("/decide", json=ROW, headers=AUTH).status_code == 200
def test_read_limited_stops_reading_at_the_crossing_chunk():
import asyncio
from semif_serve.app import ApiError, read_limited
consumed = []
async def chunks():
for i in range(10):
consumed.append(i)
yield b"x" * 100
with pytest.raises(ApiError) as info:
asyncio.run(read_limited(chunks(), None, 250))
assert info.value.status == 413
assert consumed == [0, 1, 2] # the third chunk crosses 250 and is never kept; nothing after is read
+42
View File
@@ -34,3 +34,45 @@ def test_a_calibration_table_needs_positive_finite_numbers(tmp_path, table):
cal.write_text(json.dumps(table)) cal.write_text(json.dumps(table))
with pytest.raises(ValueError, match="SEMIF_CALIBRATION"): with pytest.raises(ValueError, match="SEMIF_CALIBRATION"):
Settings.from_env({"SEMIF_API_TOKEN": TOKEN, "SEMIF_CALIBRATION": str(cal)}) Settings.from_env({"SEMIF_API_TOKEN": TOKEN, "SEMIF_CALIBRATION": str(cal)})
@pytest.mark.parametrize("var, value", [
("SEMIF_VRAM_CAP_GIB", "0"), ("SEMIF_VRAM_CAP_GIB", "-4"), ("SEMIF_VRAM_CAP_GIB", "nan"), ("SEMIF_VRAM_CAP_GIB", "inf"),
("SEMIF_MAX_TOKENS", "0"), ("SEMIF_MAX_DECISIONS", "-1"), ("SEMIF_MAX_BODY_BYTES", "0"), ("SEMIF_MAX_QUEUE", "0"),
("SEMIF_MAX_TOKENS", "lots"), ("SEMIF_VRAM_CAP_GIB", "twelve"),
])
def test_out_of_range_or_unparseable_values_are_refused_naming_the_variable(var, value):
with pytest.raises(ValueError, match=var):
Settings.from_env({"SEMIF_API_TOKEN": TOKEN, var: value})
@pytest.mark.parametrize("token", ["x" * 31 + "\n", "x" * 30 + "\r\n", "x" * 32 + "\x00", "x" * 16 + " " + "x" * 16,
"é" * 32])
def test_a_token_with_non_visible_ascii_is_refused(token):
with pytest.raises(ValueError, match="SEMIF_API_TOKEN"):
Settings.from_env({"SEMIF_API_TOKEN": token})
@pytest.mark.parametrize("table", [{"w": True}, {"w": 1e-300}, {"w": 0.04}, {"w": 21}])
def test_a_temperature_must_be_a_real_number_in_range(tmp_path, table):
cal = tmp_path / "cal.json"
cal.write_text(json.dumps(table))
with pytest.raises(ValueError, match="SEMIF_CALIBRATION"):
Settings.from_env({"SEMIF_API_TOKEN": TOKEN, "SEMIF_CALIBRATION": str(cal)})
@pytest.mark.parametrize("content", [None, "{not json"])
def test_a_missing_or_malformed_calibration_file_is_a_named_startup_error(tmp_path, content):
cal = tmp_path / "cal.json"
if content is not None:
cal.write_text(content)
with pytest.raises(ValueError, match="SEMIF_CALIBRATION"):
Settings.from_env({"SEMIF_API_TOKEN": TOKEN, "SEMIF_CALIBRATION": str(cal)})
def test_in_range_edges_are_accepted(tmp_path):
cal = tmp_path / "cal.json"
cal.write_text(json.dumps({"lo": 0.05, "hi": 20}))
s = Settings.from_env({"SEMIF_API_TOKEN": "!" + "~" * 31, "SEMIF_VRAM_CAP_GIB": "0.5", "SEMIF_MAX_QUEUE": "1",
"SEMIF_CALIBRATION": str(cal)})
assert (s.vram_cap_gib, s.max_queue, s.calibration) == (0.5, 1, {"lo": 0.05, "hi": 20.0})
+31 -1
View File
@@ -7,7 +7,7 @@ import pytest
from semif_serve.config import Settings from semif_serve.config import Settings
from semif_serve.engine import TorchEngine from semif_serve.engine import TorchEngine
from semif_serve.errors import OutOfMemory from semif_serve.errors import OutOfMemory, ScoringFailed
class FakeTorch: class FakeTorch:
@@ -81,3 +81,33 @@ def test_a_burst_is_returned_to_the_driver_after_the_call(reserved_after, releas
direct_fn=scorer, shared_fn=scorer, release_above_bytes=8 * 2**30 + 512 * 2**20) direct_fn=scorer, shared_fn=scorer, release_above_bytes=8 * 2**30 + 512 * 2**20)
assert engine.direct({}) == {"ok": True} assert engine.direct({}) == {"ok": True}
assert ReservingTorch.cuda.emptied == released assert ReservingTorch.cuda.emptied == released
def cyclic_tensor():
"""A tensor held in a reference cycle, as real frames and tensors often are: only gc frees it."""
t = Tensor()
t.self_ref = t
FakeTorch.cuda.watched.append(weakref.ref(t))
return t
@pytest.mark.parametrize("raised, expected_type, expected_message", [
(lambda: FakeTorch.cuda.OutOfMemoryError(""), OutOfMemory, "CUDA out of memory"),
(lambda: RuntimeError("CUBLAS_STATUS_ALLOC_FAILED: CUDA error: out of memory"), OutOfMemory,
"CUBLAS_STATUS_ALLOC_FAILED: CUDA error: out of memory"),
(lambda: RuntimeError("Invalid native prefix cache"), ScoringFailed, "RuntimeError: Invalid native prefix cache"),
(lambda: KeyError("option_logits"), ScoringFailed, "KeyError: 'option_logits'"),
])
def test_every_non_validation_failure_is_released_unchained_after_gc(raised, expected_type, expected_message):
FakeTorch.cuda.empties.clear(), FakeTorch.cuda.watched.clear()
def scorer(*_args):
kv_cache = cyclic_tensor() # noqa: F841 — alive in this frame when it raises
raise raised()
engine = TorchEngine(FakeTorch, None, None, {}, Settings(api_token="t" * 40), direct_fn=scorer, shared_fn=scorer)
with pytest.raises(expected_type) as info:
engine.shared([])
assert str(info.value) == expected_message
assert info.value.__cause__ is None and info.value.__context__ is None
assert FakeTorch.cuda.empties == [True] # gc freed the cycle BEFORE the cache was emptied
@@ -0,0 +1,110 @@
"""TorchEngine.load and the entry point, against fake torch/semif modules (INV-3, INV-4, INV-5).
The bug hunt found load() had no test at all: removing the arch check survived every test."""
import os
import sys
import types
import pytest
from semif_serve.config import Settings
TOKEN = "t" * 40
class Param:
def __init__(self, device):
self.device = types.SimpleNamespace(type=device)
class Model:
def __init__(self, device):
self._device = device
def parameters(self):
yield Param(self._device)
@pytest.fixture
def fakes(monkeypatch):
calls = []
cuda = types.SimpleNamespace(
is_available=lambda: True,
get_device_capability=lambda _i=0: (12, 0),
get_arch_list=lambda: ["sm_90", "sm_120"],
get_device_properties=lambda _i=0: types.SimpleNamespace(total_memory=96 * 2**30),
set_per_process_memory_fraction=lambda f, _i=0: calls.append(("cap", round(f, 4))),
memory_reserved=lambda _i=0: 8 * 2**30,
OutOfMemoryError=type("OutOfMemoryError", (RuntimeError,), {}),
empty_cache=lambda: calls.append(("empty",)),
)
torch = types.SimpleNamespace(cuda=cuda, __version__="2.10.0+cu128")
state = {"device": "cuda", "warmup_raises": None}
def load_causal_model(model, revision, device, dtype):
calls.append(("load", model, revision, device, dtype))
return Model(state["device"]), object(), {"source": model}
def score(model, tok, row, meta, max_tokens):
calls.append(("score", row["id"]))
if state["warmup_raises"]:
raise state["warmup_raises"]
return {"id": row["id"]}
core = types.ModuleType("semif_phase1.core")
core.load_causal_model = load_causal_model
direct = types.ModuleType("semif_phase1.direct")
direct.score = score
shared = types.ModuleType("semif_phase1.shared")
shared.score_shared = lambda *a: ([], {})
pkg = types.ModuleType("semif_phase1")
for name, mod in {"torch": torch, "semif_phase1": pkg, "semif_phase1.core": core,
"semif_phase1.direct": direct, "semif_phase1.shared": shared}.items():
monkeypatch.setitem(sys.modules, name, mod)
return torch, calls, state
def test_load_caps_before_the_weights_land_then_warms_up(fakes):
from semif_serve.engine import TorchEngine
_torch, calls, _ = fakes
TorchEngine.load(Settings(api_token=TOKEN, vram_cap_gib=12.0))
assert [c[0] for c in calls] == ["cap", "load", "score"]
assert calls[0] == ("cap", round(12 / 96, 4))
assert calls[2] == ("score", "semif-serve-warmup")
def test_load_refuses_a_card_torch_has_no_kernels_for(fakes):
from semif_serve.engine import TorchEngine
torch, calls, _ = fakes
torch.cuda.get_arch_list = lambda: ["sm_80", "sm_90"]
with pytest.raises(RuntimeError, match="sm_120"):
TorchEngine.load(Settings(api_token=TOKEN))
assert not any(c[0] == "load" for c in calls)
def test_load_refuses_a_model_that_landed_on_the_wrong_device(fakes):
from semif_serve.engine import TorchEngine
_torch, _calls, state = fakes
state["device"] = "cpu"
with pytest.raises(RuntimeError, match="landed on cpu"):
TorchEngine.load(Settings(api_token=TOKEN))
def test_load_fails_closed_when_the_warmup_decision_fails(fakes):
from semif_serve.engine import TorchEngine
from semif_serve.errors import ScoringFailed
_torch, _calls, state = fakes
state["warmup_raises"] = RuntimeError("Failed to find C compiler")
with pytest.raises(ScoringFailed, match="C compiler"):
TorchEngine.load(Settings(api_token=TOKEN))
def test_the_entry_point_forces_offline_mode_before_the_engine_loads(fakes, monkeypatch):
import semif_serve.engine as engine_mod
from semif_serve import main
seen = {}
monkeypatch.delenv("HF_HUB_OFFLINE", raising=False)
monkeypatch.setenv("SEMIF_API_TOKEN", TOKEN)
monkeypatch.setattr(engine_mod.TorchEngine, "load",
classmethod(lambda cls, s: seen.update(offline=os.environ.get("HF_HUB_OFFLINE")) or object()))
main.app_from_env()
assert seen["offline"] == "1"
@@ -0,0 +1,132 @@
"""Order averaging (0.1.3). Contract: semif-serve.contract.md § Order averaging."""
import math
import pytest
from fastapi.testclient import TestClient
from semif_serve.app import create_app
from semif_serve.config import Settings
TOKEN = "t" * 40
AUTH = {"Authorization": f"Bearer {TOKEN}"}
OPTS = [{"id": "casual", "description": "A casual outing"},
{"id": "date", "description": "A romantic date"},
{"id": "booty", "description": "A booty call"}]
ROW = {"id": "q", "state": "It's 2AM and I'm bored.", "question": "What is this?", "options": OPTS}
BASE = {"casual": 1.0, "date": -1.0, "booty": 0.5} # what the model "really" thinks
BIAS = 2.5 # added to whichever option is listed first
def softmax(xs):
top = max(xs)
w = [math.exp(x - top) for x in xs]
return [v / sum(w) for v in w]
class BiasedEngine:
"""Scores each option as BASE[id], plus BIAS for the first-listed option, the way a small model leans."""
def __init__(self):
self.shared_calls, self.direct_calls = [], []
def health(self):
return {}
def _score(self, row):
ids = [o["id"] for o in row["options"]]
logits = [BASE[i] + (BIAS if k == 0 else 0.0) for k, i in enumerate(ids)]
return {"id": row["id"], "option_ids": ids, "probabilities": softmax(logits), "option_logits": logits,
"prompt_sha256": row["id"], "probability_status": "uncalibrated"}
def direct(self, row):
self.direct_calls.append(row)
return self._score(row)
def shared(self, rows):
self.shared_calls.append(rows)
return [self._score(r) for r in rows], {"batch_size": len(rows)}
def client(engine, **kw):
return TestClient(create_app(Settings(api_token=TOKEN, **kw), engine))
def test_rotations_send_one_shared_call_with_every_option_once_per_position_and_cancel_the_bias():
engine = BiasedEngine()
body = client(engine).post("/decide", json={**ROW, "orderings": "rotations"}, headers=AUTH).json()
assert engine.direct_calls == [] and len(engine.shared_calls) == 1
rows = engine.shared_calls[0]
assert [r["id"] for r in rows] == ["q#o0", "q#o1", "q#o2"]
assert [[o["id"] for o in r["options"]] for r in rows] == [
["casual", "date", "booty"], ["date", "booty", "casual"], ["booty", "casual", "date"]]
assert all(r["state"] == ROW["state"] and r["question"] == ROW["question"] for r in rows)
assert body["id"] == "q" and body["option_ids"] == ["casual", "date", "booty"]
combined = body["combined"]
assert combined["method"] == "rotations" and combined["orderings"] == 3
# the first-position bias is additive, and every option sat first exactly once, so it cancels exactly
assert combined["probabilities"] == pytest.approx(softmax([BASE["casual"], BASE["date"], BASE["booty"]]))
assert combined["top"] == "casual"
# the bias is strong enough that every ordering's first option wins: casual, date, booty
tops = [max(zip(r["option_logits"], r["option_ids"]))[1] for r in body["orderings"]]
assert tops == ["casual", "date", "booty"]
assert combined["agreement"] == pytest.approx(1 / 3)
assert body["orderings"] == [engine._score(r) for r in rows] # native results, unchanged
for oid in ("casual", "date", "booty"):
ps = [dict(zip(r["option_ids"], r["probabilities"]))[oid] for r in body["orderings"]]
assert combined["spread"][oid] == pytest.approx([min(ps), max(ps)])
def test_all_sends_every_permutation_with_the_callers_order_first():
engine = BiasedEngine()
body = client(engine).post("/decide", json={**ROW, "orderings": "all"}, headers=AUTH).json()
rows = engine.shared_calls[0]
orders = [tuple(o["id"] for o in r["options"]) for r in rows]
assert len(orders) == 6 and len(set(orders)) == 6 and orders[0] == ("casual", "date", "booty")
assert body["combined"]["orderings"] == 6 and body["combined"]["method"] == "all"
def test_all_above_four_options_is_422_before_the_engine_runs():
engine = BiasedEngine()
five = [{"id": f"o{i}", "description": f"Option {i}"} for i in range(5)]
r = client(engine).post("/decide", json={**ROW, "options": five, "orderings": "all"}, headers=AUTH)
assert r.status_code == 422 and "rotations" in r.json()["error"]["message"]
assert engine.shared_calls == []
def test_a_mixed_shared_request_is_one_engine_call_with_results_in_request_order():
engine = BiasedEngine()
body = {"state": ROW["state"], "decisions": [
{"id": "plain", "question": "Q1?", "options": OPTS},
{"id": "avg", "question": "Q2?", "options": OPTS, "orderings": "rotations"},
{"id": "plain2", "question": "Q3?", "options": OPTS[:2]}]}
out = client(engine).post("/decide/shared", json=body, headers=AUTH).json()
assert len(engine.shared_calls) == 1
assert [r["id"] for r in engine.shared_calls[0]] == ["plain", "avg#o0", "avg#o1", "avg#o2", "plain2"]
ids = [r["id"] for r in out["results"]]
assert ids == ["plain", "avg", "plain2"]
assert out["results"][0] == engine._score(engine.shared_calls[0][0]) # plain results keep the old shape
assert "combined" in out["results"][1] and "combined" not in out["results"][2]
assert out["timing"] == {"batch_size": 5}
def test_expanded_rows_count_toward_the_cap():
engine = BiasedEngine()
r = client(engine, max_decisions=5).post("/decide/shared", json={"state": "s", "decisions": [
{"id": "a", "question": "Q?", "options": OPTS, "orderings": "rotations"},
{"id": "b", "question": "Q?", "options": OPTS, "orderings": "rotations"}]}, headers=AUTH)
assert r.status_code == 422 and "6 scored rows" in r.json()["error"]["message"]
assert engine.shared_calls == []
def test_workload_with_orderings_is_422_and_workload_alone_still_calibrates():
engine = BiasedEngine()
c = client(engine, calibration={"triage": 2.0})
r = c.post("/decide", json={**ROW, "orderings": "rotations", "workload": "triage"}, headers=AUTH)
assert r.status_code == 422 and engine.shared_calls == []
assert "calibrated" in c.post("/decide", json={**ROW, "workload": "triage"}, headers=AUTH).json()
def test_an_unknown_orderings_value_is_422():
r = client(BiasedEngine()).post("/decide", json={**ROW, "orderings": "shuffle"}, headers=AUTH)
assert r.status_code == 422
+80 -2
View File
@@ -51,6 +51,26 @@ wheels = [
{ url = "https://files.pythonhosted.org/packages/12/b8/4bd346e22b28902df4d651910f5242c28d84e4a5c2435ca5c3f797ed7e2e/anyio-4.15.1-py3-none-any.whl", hash = "sha256:6152fdbbf9a77fdec97731721bebf7c4c44f7c29b424b0065826173efc7ed101", size = 132079 }, { url = "https://files.pythonhosted.org/packages/12/b8/4bd346e22b28902df4d651910f5242c28d84e4a5c2435ca5c3f797ed7e2e/anyio-4.15.1-py3-none-any.whl", hash = "sha256:6152fdbbf9a77fdec97731721bebf7c4c44f7c29b424b0065826173efc7ed101", size = 132079 },
] ]
[[package]]
name = "causal-conv1d"
version = "1.7.0"
source = { url = "https://github.com/Dao-AILab/causal-conv1d/releases/download/v1.7.0/causal_conv1d-1.7.0+cu12torch2.10cxx11abiTRUE-cp312-cp312-linux_x86_64.whl" }
dependencies = [
{ name = "ninja" },
{ name = "packaging" },
{ name = "torch" },
]
wheels = [
{ url = "https://github.com/Dao-AILab/causal-conv1d/releases/download/v1.7.0/causal_conv1d-1.7.0+cu12torch2.10cxx11abiTRUE-cp312-cp312-linux_x86_64.whl", hash = "sha256:8e81f8c76435ad31aa41edc6c0c9d26de971a293b190183d2816da71547614d7" },
]
[package.metadata]
requires-dist = [
{ name = "ninja" },
{ name = "packaging" },
{ name = "torch" },
]
[[package]] [[package]]
name = "certifi" name = "certifi"
version = "2026.7.22" version = "2026.7.22"
@@ -106,6 +126,15 @@ wheels = [
{ url = "https://files.pythonhosted.org/packages/98/59/239c7259e669c46ddbcac0aa60e3a0ef00bfeaaa687f905b24dd6a7a10fe/cuda_pathfinder-1.8.2-py3-none-any.whl", hash = "sha256:4e65059febdb4d19d5cbc4798677e19db2b582f2f702f457b609e571690d357e", size = 62551 }, { url = "https://files.pythonhosted.org/packages/98/59/239c7259e669c46ddbcac0aa60e3a0ef00bfeaaa687f905b24dd6a7a10fe/cuda_pathfinder-1.8.2-py3-none-any.whl", hash = "sha256:4e65059febdb4d19d5cbc4798677e19db2b582f2f702f457b609e571690d357e", size = 62551 },
] ]
[[package]]
name = "einops"
version = "0.8.2"
source = { registry = "https://pypi.org/simple" }
sdist = { url = "https://files.pythonhosted.org/packages/2c/77/850bef8d72ffb9219f0b1aac23fbc1bf7d038ee6ea666f331fa273031aa2/einops-0.8.2.tar.gz", hash = "sha256:609da665570e5e265e27283aab09e7f279ade90c4f01bcfca111f3d3e13f2827", size = 56261 }
wheels = [
{ url = "https://files.pythonhosted.org/packages/2a/09/f8d8f8f31e4483c10a906437b4ce31bdf3d6d417b73fe33f1a8b59e34228/einops-0.8.2-py3-none-any.whl", hash = "sha256:54058201ac7087911181bfec4af6091bb59380360f069276601256a76af08193", size = 65638 },
]
[[package]] [[package]]
name = "fastapi" name = "fastapi"
version = "0.118.0" version = "0.118.0"
@@ -129,6 +158,31 @@ wheels = [
{ url = "https://files.pythonhosted.org/packages/bc/ac/8b9c6dc2aa7e9cf5582e20a337006c3428c3d752c2255dbaca2c78a8be45/filelock-4.0.4-py3-none-any.whl", hash = "sha256:0df72be195ca7892216d16f2edce8d9b93a571f02402972020a8cff84c594c7b", size = 108629 }, { url = "https://files.pythonhosted.org/packages/bc/ac/8b9c6dc2aa7e9cf5582e20a337006c3428c3d752c2255dbaca2c78a8be45/filelock-4.0.4-py3-none-any.whl", hash = "sha256:0df72be195ca7892216d16f2edce8d9b93a571f02402972020a8cff84c594c7b", size = 108629 },
] ]
[[package]]
name = "fla-core"
version = "0.5.2"
source = { registry = "https://pypi.org/simple" }
dependencies = [
{ name = "einops" },
]
sdist = { url = "https://files.pythonhosted.org/packages/94/85/20bdc0fbbbaeec27b5b198e0dbbbc51161a267066a4aebd9e4b413637c6c/fla_core-0.5.2.tar.gz", hash = "sha256:9360fc412f784c1c8f05c320ec2902d5d968764178c9b8a92efc919e17a39680", size = 587835 }
wheels = [
{ url = "https://files.pythonhosted.org/packages/2d/ed/dfe19c4da779957eb6a42a26812f9b4e2280bf757a17a71933ff59ffcb98/fla_core-0.5.2-py3-none-any.whl", hash = "sha256:5e830c85bad3d0d34677f98ac7074d08687a3756f0f0499d95ceb96eb6920761", size = 819225 },
]
[[package]]
name = "flash-linear-attention"
version = "0.5.2"
source = { registry = "https://pypi.org/simple" }
dependencies = [
{ name = "fla-core" },
{ name = "transformers" },
]
sdist = { url = "https://files.pythonhosted.org/packages/d2/76/c180949eae5161b9fcf4c928cab52f0c257ce84c1a59b4462b1d2b6e5883/flash_linear_attention-0.5.2.tar.gz", hash = "sha256:c053d3a75c8f5b725063f719ae6bce0f4d9dac6da32735be8ae7d485a6c9820f", size = 208110 }
wheels = [
{ url = "https://files.pythonhosted.org/packages/90/d2/2070e3cf2148c5cce99ca4876633c4b7b89ec085323600bd3d47aeacd306/flash_linear_attention-0.5.2-py3-none-any.whl", hash = "sha256:dcf405d81f5426393b59037097aa700d0f4a841465d5028d5aa543f4502f2400", size = 399590 },
]
[[package]] [[package]]
name = "fsspec" name = "fsspec"
version = "2026.9.0" version = "2026.9.0"
@@ -351,6 +405,24 @@ wheels = [
{ url = "https://files.pythonhosted.org/packages/7e/cd/fe58041e9011f307c490e3e17dd48cc516448f7c698a3f2d9d9d65d7e6a8/networkx-3.7-py3-none-any.whl", hash = "sha256:e3fd2c13a7814cee3746340d8d7f8598a67f16a58bf47fb7f8793fab6efca1b0", size = 2142205 }, { url = "https://files.pythonhosted.org/packages/7e/cd/fe58041e9011f307c490e3e17dd48cc516448f7c698a3f2d9d9d65d7e6a8/networkx-3.7-py3-none-any.whl", hash = "sha256:e3fd2c13a7814cee3746340d8d7f8598a67f16a58bf47fb7f8793fab6efca1b0", size = 2142205 },
] ]
[[package]]
name = "ninja"
version = "1.13.2"
source = { registry = "https://pypi.org/simple" }
sdist = { url = "https://files.pythonhosted.org/packages/ac/92/410b7917d16ab54c04b05cc32b9284803671d91cf79d33be6009c28d4ea8/ninja-1.13.2.tar.gz", hash = "sha256:525bfa3fc88aa30a4467df270fd5be6f9fcae8061d54d4df74ea1dc5abd5a975", size = 243739 }
wheels = [
{ url = "https://files.pythonhosted.org/packages/48/23/fcbe234a66966e35928c47b86336f92a7612db4781665f4e5f5fddef9630/ninja-1.13.2-py3-none-manylinux2014_ppc64le.manylinux_2_17_ppc64le.whl", hash = "sha256:227cbc3ae3e5e429692388103cae8c09451df086cd2d342dae0795af0d162547", size = 197676 },
{ url = "https://files.pythonhosted.org/packages/24/eb/a6ca97ef0ff7bb8bdcb395ec65a716e65d7c1f40896c3afe0090bb3e1535/ninja-1.13.2-py3-none-manylinux2014_s390x.manylinux_2_17_s390x.whl", hash = "sha256:1684c60d031c54c1d049541b64243c0c567dca5463dbd77682a8901780af293d", size = 187980 },
{ url = "https://files.pythonhosted.org/packages/6e/53/ebfed7b689c338dd8ebeec9c0730c8d56821292f14e2536e5f3ef1a05744/ninja-1.13.2-py3-none-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:65a24341b5ac09fcadcc37082660be40a94174e51a937fabf6e2cae26225fa2c", size = 183365 },
{ url = "https://files.pythonhosted.org/packages/c7/d6/dcf06d7ab44ade992ae5aa1228feff317684b463a1bd47e8642b30ac922e/ninja-1.13.2-py3-none-manylinux_2_28_armv7l.manylinux_2_31_armv7l.whl", hash = "sha256:aa3d2ae5706a2c4d1e93edc951d1c6cbb45107413c404f8fde1741239efbc9a0", size = 155089 },
{ url = "https://files.pythonhosted.org/packages/e1/6b/6513c09c33382b17c05b4349b8e81437b18680d0d7ec6fb8f7edc29adda1/ninja-1.13.2-py3-none-manylinux_2_31_riscv64.whl", hash = "sha256:919572cbc3f233261ecd41fe1f3efc9d44aa02464a4588867e06a8b4f6f416ea", size = 152149 },
{ url = "https://files.pythonhosted.org/packages/37/04/c8c2dc5b2f5fee79a1691d490256b178b7e1af97d56769117422ae8a23cc/ninja-1.13.2-py3-none-musllinux_1_2_armv7l.whl", hash = "sha256:59d71c3e15b6b6f3d903eb0c27285544e0747ca59925ada7037bb1af781ad4b3", size = 466268 },
{ url = "https://files.pythonhosted.org/packages/10/a2/d8eedd25d0ae80b9e874aea362013e67416877c8f732eb4d5e7c971fdb9c/ninja-1.13.2-py3-none-musllinux_1_2_ppc64le.whl", hash = "sha256:b2f687437fac460b27b7eadc99039b1163016fb4ba7276e2782a192d9f24ee0e", size = 610806 },
{ url = "https://files.pythonhosted.org/packages/14/0f/696d96821fad1b5767fd311c1569dde8881a57412369bfe7b11bcbfde036/ninja-1.13.2-py3-none-musllinux_1_2_riscv64.whl", hash = "sha256:09de9ab04f7352f51570c73fd4913acb1e6c24be0a72cd8b80243d4d3ed04925", size = 533978 },
{ url = "https://files.pythonhosted.org/packages/5d/69/28844ca579156776a202217a7cd66f60d06a0710a935e879bb89ce396ecc/ninja-1.13.2-py3-none-musllinux_1_2_s390x.whl", hash = "sha256:6a87bf42b123abe2f37737300185f0a303a891899da85d73a3613ee80547e578", size = 653822 },
{ url = "https://files.pythonhosted.org/packages/f5/5f/c511f2952f94ab2966d60edd9c34e744ea32f2724b1184b62270bde55b3a/ninja-1.13.2-py3-none-musllinux_1_2_x86_64.whl", hash = "sha256:915bd482c4be41c75120fd67a22e0bb3f0fbb3bbc5f95b89787deadd59e27ef2", size = 544460 },
]
[[package]] [[package]]
name = "numpy" name = "numpy"
version = "2.2.6" version = "2.2.6"
@@ -919,7 +991,7 @@ dependencies = [
[[package]] [[package]]
name = "semif-serve" name = "semif-serve"
version = "0.1.2" version = "0.1.3"
source = { editable = "." } source = { editable = "." }
dependencies = [ dependencies = [
{ name = "fastapi" }, { name = "fastapi" },
@@ -927,6 +999,10 @@ dependencies = [
] ]
[package.optional-dependencies] [package.optional-dependencies]
fast = [
{ name = "causal-conv1d", marker = "platform_machine == 'x86_64' and sys_platform == 'linux'" },
{ name = "flash-linear-attention" },
]
model = [ model = [
{ name = "semif-phase1" }, { name = "semif-phase1" },
{ name = "torch" }, { name = "torch" },
@@ -940,12 +1016,14 @@ dev = [
[package.metadata] [package.metadata]
requires-dist = [ requires-dist = [
{ name = "causal-conv1d", marker = "platform_machine == 'x86_64' and sys_platform == 'linux' and extra == 'fast'", url = "https://github.com/Dao-AILab/causal-conv1d/releases/download/v1.7.0/causal_conv1d-1.7.0+cu12torch2.10cxx11abiTRUE-cp312-cp312-linux_x86_64.whl" },
{ name = "fastapi", specifier = "==0.118.0" }, { name = "fastapi", specifier = "==0.118.0" },
{ name = "flash-linear-attention", marker = "extra == 'fast'", specifier = "==0.5.2" },
{ name = "semif-phase1", marker = "extra == 'model'", git = "https://github.com/TheoLeeCJ/SemIf-OpenJev?rev=23cf1f39fc9534fe81437200959b6dfc7106e45a" }, { name = "semif-phase1", marker = "extra == 'model'", git = "https://github.com/TheoLeeCJ/SemIf-OpenJev?rev=23cf1f39fc9534fe81437200959b6dfc7106e45a" },
{ name = "torch", marker = "extra == 'model'", specifier = "==2.10.0", index = "https://download.pytorch.org/whl/cu128" }, { name = "torch", marker = "extra == 'model'", specifier = "==2.10.0", index = "https://download.pytorch.org/whl/cu128" },
{ name = "uvicorn", specifier = "==0.37.0" }, { name = "uvicorn", specifier = "==0.37.0" },
] ]
provides-extras = ["model"] provides-extras = ["model", "fast"]
[package.metadata.requires-dev] [package.metadata.requires-dev]
dev = [ dev = [
+1 -1
View File
@@ -1,6 +1,6 @@
# semif — copy to /opt/docker/compose/semif/.env on fv-ml1 (mode 0600). # semif — copy to /opt/docker/compose/semif/.env on fv-ml1 (mode 0600).
# Built on fv-ml1 from services/semif-serve (see README "Building"). # Built on fv-ml1 from services/semif-serve (see README "Building").
IMAGE=semif-serve:0.1.2 IMAGE=semif-serve:0.1.3
PORT=8032 PORT=8032
HOST_IP=10.251.50.54 HOST_IP=10.251.50.54
# fv-ml1 GPU 1 = the utility card (vllm-coder, erp, meromero, scriberr). # fv-ml1 GPU 1 = the utility card (vllm-coder, erp, meromero, scriberr).
+56 -7
View File
@@ -16,7 +16,7 @@ service. The contract is `services/semif-serve/semif-serve.contract.md`.
| **Model** | `Qwen/Qwen3.5-4B` @ `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a`, BF16, from `/tank/aimodels/huggingface` (read-only, offline) | | **Model** | `Qwen/Qwen3.5-4B` @ `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a`, BF16, from `/tank/aimodels/huggingface` (read-only, offline) |
| **SemIf** | commit `23cf1f39fc9534fe81437200959b6dfc7106e45a`; torch `2.10.0+cu128`, transformers `5.17.0`, the same stack SemIf's committed predictions were made on | | **SemIf** | commit `23cf1f39fc9534fe81437200959b6dfc7106e45a`; torch `2.10.0+cu128`, transformers `5.17.0`, the same stack SemIf's committed predictions were made on |
| **Image** | `semif-serve:<version>`, built on fv-ml1 from `services/semif-serve/` | | **Image** | `semif-serve:<version>`, built on fv-ml1 from `services/semif-serve/` |
| **State** | none. Weights re-download at the pinned revision. No backup needed beyond the host's `/opt/docker` restic. | | **State** | none. If the weights are ever lost, re-pull them by hand at the pinned revision (the service itself never downloads: `HF_HUB_OFFLINE=1`, read-only mount). No backup needed beyond the host's `/opt/docker` restic. |
## API ## API
@@ -34,6 +34,18 @@ curl -s -H "Authorization: Bearer $T" http://10.251.50.54:8032/decide -d '{
the state once and scores every criterion in one batch. Use it when many questions the state once and scores every criterion in one batch. Use it when many questions
share one long state. share one long state.
- `GET /health`: the pins, limits and calibrated workloads. - `GET /health`: the pins, limits and calibrated workloads.
- **Order averaging (0.1.3):** add `"orderings": "rotations"` to a decision, or
`"all"` for ≤ 4 options. The option list is asked in every rotation inside ONE
shared batch. The reply keeps each ordering's native result under `orderings` and adds
`combined: {method, orderings, probabilities, top, agreement, spread}`. **Use it for
anything real:** a small model leans toward the first-listed option on ambiguous
inputs, and averaging cancels that. Through the service on SemIf's labelled sets,
accuracy goes from 78.6% to **88.1%** (group-bootstrap 95% CI +5.1..+14.3 pts, 252
rows). **`agreement` is the cheap confidence signal**: unanimous rows are 94.5%
accurate, split rows 76.4%. The orderings count toward the decision cap. `workload`
calibration is not available together with `orderings` yet (422).
- Past `MAX_QUEUE` (32) requests in progress, new POSTs get `429 busy` before their
body is read.
## ⚠ Probabilities are uncalibrated ## ⚠ Probabilities are uncalibrated
@@ -57,18 +69,55 @@ not squeeze scriberr, which shares GPU 1. A request that would exceed the cap ge
behaviours were verified on the card (0.1.0 held 11.9 GiB after an OOM, and 12.6 GB behaviours were verified on the card (0.1.0 held 11.9 GiB after an OOM, and 12.6 GB
after a large request; 0.1.2 returns to 7.85 GiB in both cases). after a large request; 0.1.2 returns to 7.85 GiB in both cases).
**What fits under 12 GiB** (measured, `/decide/shared`, binary decisions): **What fits under 12 GiB** (measured on 0.1.3, `/decide/shared`, binary decisions;
0.1.2 figures in brackets, before the fast kernels):
| state size (prefix tokens) | max decisions in one request | | state size (prefix tokens) | max rows in one request |
|---|---| |---|---|
| ~140 | 52 | | ~140 | 63 (52) |
| ~520 | 43 | | ~520 | 51 (43) |
| ~1,960 | 19 | | ~1,960 | 26 (19) |
| ~3,900 | 13 | | ~3,900 | 16 (13) |
Rows = decisions × orderings, so `rotations` over 3 options uses 3 rows per decision.
`/decide` fits at the full 4,096-token limit. Past the table you get a 503, so split `/decide` fits at the full 4,096-token limit. Past the table you get a 503, so split
the decisions across requests. the decisions across requests.
## Fast kernels (0.1.3)
The image ships Qwen3.5's fast kernels, `flash-linear-attention` 0.5.2 and
`causal-conv1d` 1.7.0 (a prebuilt cu12/torch2.10 wheel). Without them transformers
logs that it falls back to "much slower" reference PyTorch paths. They are adopted
because an A/B on the empty GPU 3 showed:
- **parity improved**: 144/144 vs upstream (142/144 without), so upstream evidently
ran with them;
- **long inputs got much faster**: a ~2,000-token `/decide` went from 169 to 92 ms
server-side.
Short 3-rotation batches cost ~3–6 ms more; everything else is equal or faster.
⚠ triton builds a C shim at runtime, so the image carries `gcc`. Without it the
warm-up fails, and startup fails closed. Build without the kernels:
`--build-arg EXTRAS="--extra model"`.
## Latency (0.1.3, from nh3-dev, 3 runs × 20, network floor ~33 ms)
| request | end to end | server |
|---|---|---|
| `/decide`, short (~130 tok) | 69 ms | 35 ms |
| `/decide`, ~2,000-token state | 131 ms | 92 ms |
| 3 rotations, short | 115 ms | 78 ms |
| 6 orderings, short | 118 ms | 81 ms |
| 3 rotations, ~2,000-token state | 200 ms | 158 ms |
## Acceptance (2026-09-27, v0.1.3)
Raw: `services/semif-serve/acceptance/result-2026-09-27-v0.1.3.json`,
`averaging-2026-09-27-v0.1.3.json`. v0.1.3 matches upstream on **144/144** (identical
prompt hashes, max prob gap 0.049), is deterministic, fails the negative control as it
should (14/144), and shared matches direct on 72/72. The v0.1.2 table below is the
reference-kernel baseline.
## Acceptance (2026-09-27, v0.1.2) ## Acceptance (2026-09-27, v0.1.2)
Raw: `services/semif-serve/acceptance/result-2026-09-27-v0.1.2.json`. Raw: `services/semif-serve/acceptance/result-2026-09-27-v0.1.2.json`.
+2
View File
@@ -30,6 +30,8 @@ services:
SEMIF_VRAM_CAP_GIB: ${VRAM_CAP_GIB:?set VRAM_CAP_GIB} SEMIF_VRAM_CAP_GIB: ${VRAM_CAP_GIB:?set VRAM_CAP_GIB}
SEMIF_MAX_TOKENS: ${MAX_TOKENS:-4096} SEMIF_MAX_TOKENS: ${MAX_TOKENS:-4096}
SEMIF_MAX_DECISIONS: ${MAX_DECISIONS:-64} SEMIF_MAX_DECISIONS: ${MAX_DECISIONS:-64}
# POSTs in progress (queued + scoring) before new ones get 429 busy.
SEMIF_MAX_QUEUE: ${MAX_QUEUE:-32}
SEMIF_CALIBRATION: /conf/calibration.json SEMIF_CALIBRATION: /conf/calibration.json
volumes: volumes:
# Pinned weights, read offline (HF_HUB_OFFLINE=1 in the image). Never downloads. # Pinned weights, read offline (HF_HUB_OFFLINE=1 in the image). Never downloads.