feat(semif): 0.1.3 — order averaging, fast kernels, bug-hunt hardening (Prime)
Order averaging (Prime, after the 739aa03 spike):
- A decision may set orderings: rotations|all (all only for <= 4 options). Every
ordering goes to the engine in one shared batch.
- The reply keeps each native result and adds combined {probabilities (log-mean),
top, agreement, spread}.
- Through the service on SemIf's labelled sets (252 rows): 78.6% -> 88.1%
(group-bootstrap 95% CI +5.1..+14.3). Unanimous agreement is 94.5% accurate.
Fast kernels: flash-linear-attention 0.5.2 and causal-conv1d 1.7.0 are now the
default build. A/B on the empty GPU 3:
- parity with upstream went from 142/144 to 144/144;
- a ~2k-token /decide went from 169 to 92 ms server-side;
- short 3-rotation batches cost ~3-6 ms more.
triton builds a C shim at runtime, so the image carries gcc. Without it the
warm-up failed and startup failed closed.
Heid bug-hunt panel (4/4 arms, thread 01M3H3F4RR7XBP90KQ3A39H4SX), folded:
- Startup validation: VRAM cap 0 no longer means uncapped (C1); limits must be
>= 1 (S1); the token must be visible ASCII (S2); the calibration file must
exist and parse, with T in [0.05, 20] (S8, and C3's NaN leg).
- The body limit is checked before a chunk is kept, and a Unicode-digit
Content-Length no longer crashes (C2, S3).
- Failures while building the response now get the 500 envelope (C3).
- 429 busy past SEMIF_MAX_QUEUE requests in progress (C6).
- The engine releases memory on every non-validation failure, unchained after
gc; an empty OOM message is handled; 'out of memory' RuntimeErrors map to 503
(C4, C5, S9).
- The entry point forces HF_HUB_OFFLINE (S10). README wording fixed (S5, S6).
- New guard tests close the gaps the arms' mutation grids exposed: early stop of
the body read, a shared-route lock, calibration pass-through, the gc cycle,
the exact caps, TorchEngine.load's arch and device checks, and the offline
entry point.
86 tests.
Deployed on fv-ml1 GPU 1: parity 144/144, OOM and burst release verified, shared
capacity 63/51/26/16 rows at ~140/520/1960/3900 prefix tokens.
This commit is contained in:
+11
-27
@@ -144,33 +144,17 @@ _As of 2026-09-27 ~0255 PT._
|
||||
|
||||
### SemIf on fv-ml1 GPU 1 (2026-09-27, Prime)
|
||||
|
||||
- **LIVE: `semif-serve` 0.1.2** at `http://10.251.50.54:8032` (`semif.fv.internal`). It wraps the SemIf
|
||||
option-logit scorers (pinned `23cf1f39`, Qwen3.5-4B `851bf6e8`, BF16); token `semif/api-token`.
|
||||
Code + contract: `services/semif-serve/`; stack `stacks/semif`. **No consumer yet.**
|
||||
- Acceptance: 142/144 top choice vs upstream (both misses are exact bf16 ties), 144/144 prompt hashes,
|
||||
deterministic, negative control 14/144. It holds 12 GiB hard with release after a burst (~8.7 GB resting).
|
||||
- Latency from nh3-dev: `/decide` short 71 ms end to end (38 ms server); 3 rotations 113 ms; ~2k-token
|
||||
state 210 ms (`services/semif-serve/spike/latency-2026-09-27.txt`).
|
||||
- **NEXT, in this order (Prime 0250):**
|
||||
1. **Build order-averaging into semif-serve as 0.1.3.** Spike (`739aa03`): 3 rotations, log-mean,
|
||||
take accuracy 78.6% → 87.7% (group-bootstrap 95% CI +4.7..+13.8 pts); unanimous rows are 94.4%
|
||||
accurate, split rows 75.8%. Design: opt-in `orderings: rotations|all` (all only ≤ 4 options),
|
||||
every ordering in one shared batch, per-ordering SemIf results returned unchanged + a
|
||||
`combined` block (log-mean probabilities, top, agreement, per-option spread). Orderings count
|
||||
toward `max_decisions`. Contract amendment → TDD → deploy → acceptance (authored144 accuracy
|
||||
with/without; 2AM/2PM agreement as controls).
|
||||
2. **Trial the fast kernels** (`causal_conv1d`, `flash-linear-attention`). Today transformers runs
|
||||
Qwen3.5's reference PyTorch paths and logs that they are "much slower". Adopt only if the
|
||||
authored144 parity re-check holds.
|
||||
3. **heid bug-hunt panel** on 0.1.2 is pending (thread `01M3H3F4RR7XBP90KQ3A39H4SX`). Triage and fold
|
||||
its findings into 0.1.3.
|
||||
- NVFP4 is not worth it: the SemIf scorer runs on transformers, not vLLM; quant noise lands on the scores
|
||||
and would need its own calibration; BF16 already fits. SemIf's own MLX 4-bit run moved authored144
|
||||
0.813 → 0.789 (one run each, so indicative only).
|
||||
- Prime's probes (single scenarios, not benchmarks): at 2AM vs 2PM, booty-call probability averaged over
|
||||
all orderings goes 0.42 → 0.05; a baby dragon reads life-changing 0.65 / world-changing 0.25 with only
|
||||
2/3 rotations agreeing (one ordering said 0.925), the lottery ticket reads life-changing unanimously,
|
||||
and the paperclip control reads trivial 0.999.
|
||||
- **LIVE: `semif-serve` 0.1.3** at `http://10.251.50.54:8032` (`semif.fv.internal`), with order averaging
|
||||
and the fast kernels. SemIf `23cf1f39`, Qwen3.5-4B `851bf6e8`, BF16; token `semif/api-token`. Code +
|
||||
contract: `services/semif-serve/`; stack `stacks/semif`. **No consumer yet.**
|
||||
- 0.1.3 acceptance: 144/144 parity with upstream; averaging through the service 78.6% → 88.1%
|
||||
(95% CI +5.1..+14.3); unanimous rotations 94.5% accurate. Latency, envelope and the fast-kernel A/B
|
||||
are in `stacks/semif/README.md`.
|
||||
- The heid bug-hunt panel (thread `01M3H3F4RR7XBP90KQ3A39H4SX`) is triaged and folded into 0.1.3:
|
||||
C1–C6, S1–S3, S5, S6, S8–S10 fixed with tests; S4 settled; S7 (publish on all interfaces) accepted
|
||||
as LAN + auth.
|
||||
- NVFP4 is not worth it (see the SemIf detail files). Prime's probes (2AM/2PM, dragon/lottery) are
|
||||
recorded in [[2026-09-27-semif-order-averaging]].
|
||||
|
||||
### restic: credential leak fixed (2026-09-27, Prime)
|
||||
|
||||
|
||||
@@ -22,18 +22,26 @@ l.write_text(re.sub(r'(name = "semif-serve"\nversion = )"[^"]+"', r'\1"0.0.0"',
|
||||
EOF
|
||||
|
||||
FROM python:3.12-slim-bookworm
|
||||
# EXTRAS picks the optional dependency sets. Default (adopted 2026-09-27): the model plus
|
||||
# Qwen3.5's fast kernels (fla + causal-conv1d). "--extra model" alone gives the reference
|
||||
# PyTorch paths.
|
||||
ARG EXTRAS="--extra model --extra fast"
|
||||
COPY --from=ghcr.io/astral-sh/uv:0.6.9 /uv /bin/uv
|
||||
ENV UV_COMPILE_BYTECODE=1 UV_LINK_MODE=copy UV_PYTHON_DOWNLOADS=never
|
||||
# git: semif-phase1 installs from a pinned GitHub commit.
|
||||
# git: semif-phase1 installs from a pinned GitHub commit. The fast extra also needs a C
|
||||
# compiler AT RUNTIME: triton builds its CUDA driver shim on first use, and without gcc the
|
||||
# warm-up dies with "Failed to find C compiler" (seen 2026-09-27; startup failed closed).
|
||||
RUN apt-get update && apt-get install -y --no-install-recommends git ca-certificates \
|
||||
&& if echo "$EXTRAS" | grep -q -- '--extra fast'; then \
|
||||
apt-get install -y --no-install-recommends gcc libc6-dev; fi \
|
||||
&& rm -rf /var/lib/apt/lists/*
|
||||
WORKDIR /app
|
||||
COPY --from=deps /deps/pyproject.toml /deps/uv.lock ./
|
||||
RUN --mount=type=cache,target=/root/.cache/uv \
|
||||
uv sync --frozen --no-dev --extra model --no-install-project
|
||||
uv sync --frozen --no-dev $EXTRAS --no-install-project
|
||||
COPY pyproject.toml uv.lock ./
|
||||
COPY src ./src
|
||||
RUN uv sync --frozen --no-dev --extra model --no-editable --no-cache
|
||||
RUN uv sync --frozen --no-dev $EXTRAS --no-editable --no-cache
|
||||
RUN groupadd --system --gid 10001 semif \
|
||||
&& useradd --system --uid 10001 --gid 10001 --no-create-home --shell /usr/sbin/nologin semif
|
||||
USER semif
|
||||
@@ -41,6 +49,7 @@ ENV PATH=/app/.venv/bin:$PATH \
|
||||
HF_HOME=/hf \
|
||||
HF_HUB_OFFLINE=1 \
|
||||
HF_HUB_DISABLE_TELEMETRY=1 \
|
||||
TRITON_CACHE_DIR=/tmp/triton-cache \
|
||||
NVIDIA_DRIVER_CAPABILITIES=compute,utility
|
||||
EXPOSE 8000
|
||||
# One worker (INV-2): the model and the inference lock live in this one process.
|
||||
|
||||
@@ -0,0 +1,18 @@
|
||||
{
|
||||
"rows": 252,
|
||||
"groups": 72,
|
||||
"accuracy_plain": 0.7857,
|
||||
"accuracy_rotations": 0.881,
|
||||
"delta_95ci": [
|
||||
0.0512,
|
||||
0.1429
|
||||
],
|
||||
"unanimous": {
|
||||
"rows": 163,
|
||||
"accuracy": 0.9448
|
||||
},
|
||||
"split": {
|
||||
"rows": 89,
|
||||
"accuracy": 0.764
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,42 @@
|
||||
"""Order-averaging acceptance through the SERVICE (0.1.3): the same 252 labelled rows as the spike
|
||||
(SemIf authored144 + perturbations108), each scored twice: plain /decide (the caller's order), and
|
||||
/decide with orderings=rotations (combined.top). Paired group bootstrap for the accuracy delta.
|
||||
SEMIF_DIR=... SEMIF_URL=... SEMIF_TOKEN=... uv run --with httpx python averaging.py out.json
|
||||
"""
|
||||
import json, os, random, statistics as st, sys
|
||||
from pathlib import Path
|
||||
import httpx
|
||||
|
||||
S, U = Path(os.environ["SEMIF_DIR"]), os.environ["SEMIF_URL"]
|
||||
H = {"Authorization": f"Bearer {os.environ['SEMIF_TOKEN']}"}
|
||||
rows = []
|
||||
for name in ("authored144", "perturbations108"):
|
||||
rows += [dict(json.loads(l), set=name) for l in (S / f"benchmarks/data/{name}.jsonl").read_text().splitlines() if l.strip()]
|
||||
out = []
|
||||
with httpx.Client(timeout=120) as c:
|
||||
for r in rows:
|
||||
base = {k: r[k] for k in ("id", "state", "question", "options")}
|
||||
plain = c.post(f"{U}/decide", headers=H, json=base).json()
|
||||
avg = c.post(f"{U}/decide", headers=H, json={**base, "orderings": "rotations"}).json()
|
||||
gold = r["options"][r["label"]]["id"]
|
||||
top_plain = plain["option_ids"][plain["probabilities"].index(max(plain["probabilities"]))]
|
||||
out.append({"group": r["group_id"], "plain": top_plain == gold, "rotations": avg["combined"]["top"] == gold,
|
||||
"agreement": avg["combined"]["agreement"]})
|
||||
groups = {}
|
||||
for o in out:
|
||||
groups.setdefault(o["group"], []).append(o)
|
||||
rng, keys, deltas = random.Random(7), list(groups), []
|
||||
for _ in range(10000):
|
||||
sample = [o for g in (rng.choice(keys) for _ in keys) for o in groups[g]]
|
||||
deltas.append(st.fmean(o["rotations"] for o in sample) - st.fmean(o["plain"] for o in sample))
|
||||
deltas.sort()
|
||||
unan = [o for o in out if o["agreement"] == 1.0]
|
||||
split = [o for o in out if o["agreement"] < 1.0]
|
||||
report = {"rows": len(out), "groups": len(groups),
|
||||
"accuracy_plain": round(st.fmean(o["plain"] for o in out), 4),
|
||||
"accuracy_rotations": round(st.fmean(o["rotations"] for o in out), 4),
|
||||
"delta_95ci": [round(deltas[250], 4), round(deltas[9750], 4)],
|
||||
"unanimous": {"rows": len(unan), "accuracy": round(st.fmean(o["rotations"] for o in unan), 4)},
|
||||
"split": {"rows": len(split), "accuracy": round(st.fmean(o["rotations"] for o in split), 4) if split else None}}
|
||||
json.dump(report, open(sys.argv[1], "w"), indent=1)
|
||||
print(json.dumps(report))
|
||||
@@ -0,0 +1,63 @@
|
||||
{
|
||||
"url": "http://10.251.50.54:8032",
|
||||
"health": {
|
||||
"status": "ok",
|
||||
"semif_commit": "23cf1f39fc9534fe81437200959b6dfc7106e45a",
|
||||
"model": {
|
||||
"source": "Qwen/Qwen3.5-4B",
|
||||
"revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a",
|
||||
"dtype": "bfloat16",
|
||||
"device": "cuda:0",
|
||||
"torch_version": "2.10.0+cu128",
|
||||
"transformers_version": "5.17.0",
|
||||
"device_name": "NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition",
|
||||
"allocated_gib": 7.84,
|
||||
"reserved_gib": 8.12
|
||||
},
|
||||
"vram_cap_gib": 12.0,
|
||||
"max_tokens": 4096,
|
||||
"max_decisions": 64,
|
||||
"workloads": []
|
||||
},
|
||||
"1_parity_vs_upstream": {
|
||||
"rows": 144,
|
||||
"top_choice_agree": 144,
|
||||
"max_abs_prob_gap": 0.049096847865749194
|
||||
},
|
||||
"1_prompt_sha256_equal": 144,
|
||||
"2_noise_floor_a_vs_b": {
|
||||
"rows": 144,
|
||||
"top_choice_agree": 144,
|
||||
"max_abs_prob_gap": 0.0
|
||||
},
|
||||
"3_negative_rotated_options": {
|
||||
"rows": 144,
|
||||
"top_choice_agree": 14,
|
||||
"max_abs_prob_gap": 0.9985836128543327
|
||||
},
|
||||
"4_shared_vs_direct": {
|
||||
"groups": 36,
|
||||
"rows": 72,
|
||||
"top_choice_agree": 72,
|
||||
"max_abs_prob_gap": 0.046346781311727814
|
||||
},
|
||||
"5_speed_21_binary": {
|
||||
"prefix_tokens": 62,
|
||||
"shared_s": {
|
||||
"runs": [
|
||||
0.139055563005968,
|
||||
0.1375489159981953,
|
||||
0.13897942200128455
|
||||
],
|
||||
"median": 0.13897942200128455
|
||||
},
|
||||
"sequential_decide_s": {
|
||||
"runs": [
|
||||
0.9357447380025405,
|
||||
0.9384197259932989,
|
||||
0.939685705001466
|
||||
],
|
||||
"median": 0.9384197259932989
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -1,6 +1,6 @@
|
||||
[project]
|
||||
name = "semif-serve"
|
||||
version = "0.1.2"
|
||||
version = "0.1.3"
|
||||
description = "HTTP wrapper around SemIf's direct and shared option-logit scorers"
|
||||
requires-python = ">=3.12"
|
||||
dependencies = [
|
||||
@@ -17,6 +17,13 @@ model = [
|
||||
"torch==2.10.0",
|
||||
]
|
||||
|
||||
# Qwen3.5's fast kernels. Without them transformers runs its reference PyTorch paths
|
||||
# ("correct but much slower"). Trialled 2026-09-27; adopted only if parity with upstream holds.
|
||||
fast = [
|
||||
"flash-linear-attention==0.5.2",
|
||||
"causal-conv1d @ https://github.com/Dao-AILab/causal-conv1d/releases/download/v1.7.0/causal_conv1d-1.7.0+cu12torch2.10cxx11abiTRUE-cp312-cp312-linux_x86_64.whl ; sys_platform == 'linux' and platform_machine == 'x86_64'",
|
||||
]
|
||||
|
||||
[dependency-groups]
|
||||
dev = ["pytest==8.4.2", "httpx==0.28.1"]
|
||||
|
||||
|
||||
@@ -45,6 +45,37 @@ probabilities}` with `softmax(option_logits / T)`. The argmax never changes. The
|
||||
native `probabilities` and `probability_status` stay untouched. An unknown
|
||||
`workload` is a 422. With no `workload`, no `calibrated` key appears.
|
||||
|
||||
**Order averaging (0.1.3, Prime 2026-09-27).** A decision (the `/decide` body, or an
|
||||
entry in `decisions`) may set `orderings`, whose default is `"none"`:
|
||||
|
||||
- `"rotations"`: the n cyclic shifts of the caller's option list, starting with
|
||||
the caller's order, so every option sits in every position exactly once.
|
||||
- `"all"`: every permutation (n!), the caller's order first. It is allowed only
|
||||
when n ≤ 4, else 422.
|
||||
|
||||
Every ordering of every decision in the request becomes its own SemIf row
|
||||
(`id` = `"<id>#o<k>"`, same state and question, reordered options). **All rows go
|
||||
to the engine in ONE `shared` call**, which includes `/decide`. The result for an
|
||||
averaged decision is:
|
||||
|
||||
```
|
||||
{id, option_ids (caller's order),
|
||||
combined: {method, orderings: n, probabilities, top, agreement, spread: {option_id: [min, max]}},
|
||||
orderings: [the SemIf result for each ordering, unchanged]}
|
||||
```
|
||||
|
||||
- `probabilities`: per ordering, log-softmax of `option_logits`; averaged per
|
||||
option id; renormalised; reported in the caller's order.
|
||||
- `agreement`: the fraction of orderings whose top option equals `combined.top`.
|
||||
- `spread`: each option's min and max native probability across orderings.
|
||||
|
||||
Every expanded row counts toward `SEMIF_MAX_DECISIONS`. `workload` together with
|
||||
`orderings` is a 422, because a temperature is fitted per method and none is
|
||||
fitted on combined scores yet. Decisions without `orderings` keep the exact
|
||||
pre-0.1.3 result shape. Averaging cancels any additive position bias exactly.
|
||||
Measured by the 2026-09-27 spike: 3 rotations take SemIf's labelled sets from
|
||||
78.6% to 87.7% accuracy, and unanimous agreement is 94.4% accurate.
|
||||
|
||||
## Invariants
|
||||
|
||||
- **INV-1 pass-through.** `option_ids`, `probabilities`, `option_logits`,
|
||||
@@ -64,10 +95,18 @@ native `probabilities` and `probability_status` stay untouched. An unknown
|
||||
baseline by more than 512 MiB, the engine calls `empty_cache()`. A burst must not
|
||||
keep the card's shared headroom: on 2026-09-27 a 64-decision request left the
|
||||
process holding 12.6 GB, leaving scriberr 3.5 GB.
|
||||
**Any other scorer failure except `ValueError`** also releases before it is
|
||||
reported: it is logged with its traceback, then re-raised unchained as
|
||||
`ScoringFailed` after `gc.collect()` + `empty_cache()` (bug hunt C5). A
|
||||
`RuntimeError` whose message says "out of memory" (cuBLAS/cuDNN allocation
|
||||
failures) counts as an OOM → 503 (S9). An OOM with an empty message is reported
|
||||
as "CUDA out of memory" rather than crashing the handler (C4).
|
||||
- **INV-5 no network at runtime.** Weights come from the mounted HF cache at the
|
||||
pinned revision (`HF_HUB_OFFLINE=1`).
|
||||
pinned revision. The entry point sets `HF_HUB_OFFLINE=1` itself before torch or
|
||||
transformers load, so this holds outside the image too (S10).
|
||||
- **INV-6 constant-time auth.** Token comparison uses `hmac.compare_digest`. The
|
||||
token is ≥ 32 characters, and startup refuses a shorter one.
|
||||
token is ≥ 32 characters of visible ASCII (33–126). Startup refuses anything
|
||||
else, because a CR, LF or NUL in the token can never arrive in a header (S2).
|
||||
|
||||
## Limits and errors
|
||||
|
||||
@@ -75,15 +114,21 @@ native `probabilities` and `probability_status` stay untouched. An unknown
|
||||
422, never truncated (SemIf raises).
|
||||
- `SEMIF_MAX_DECISIONS` (default 64) caps `decisions` per shared request. It must
|
||||
hold 1..max entries, else 422.
|
||||
- Request body ≤ `SEMIF_MAX_BODY_BYTES` (default 1 MiB), else 413.
|
||||
- Request body ≤ `SEMIF_MAX_BODY_BYTES` (default 1 MiB), else 413. The limit is
|
||||
checked **before** each chunk is kept, so no more than the limit is ever held
|
||||
(C2). A declared `Content-Length` is trusted only if it is ASCII digits (S3).
|
||||
- **Admission:** at most `SEMIF_MAX_QUEUE` (default 32) POSTs may be in progress,
|
||||
counting both queued and scoring. The next one is refused with 429 `busy`
|
||||
before its body is read (C6).
|
||||
|
||||
| status | code | when |
|
||||
|---|---|---|
|
||||
| 401 | `unauthorized` | missing or wrong bearer |
|
||||
| 413 | `request_too_large` | body over the limit |
|
||||
| 422 | `invalid_request` | bad JSON shape, a SemIf `ValueError` (validation, token limit, tokenisation), unknown workload, too many decisions |
|
||||
| 429 | `busy` | `SEMIF_MAX_QUEUE` requests already in progress |
|
||||
| 503 | `out_of_memory` | CUDA OOM during scoring |
|
||||
| 500 | `scoring_failed` | any other scorer exception |
|
||||
| 500 | `scoring_failed` | any other scorer exception, **or a failure while building the response** from a scorer result (calibration, averaging): always the envelope, never a bare 500 (C3) |
|
||||
|
||||
The error body is `{error: {code, message}}`.
|
||||
|
||||
@@ -92,7 +137,16 @@ The error body is `{error: {code, message}}`.
|
||||
`SEMIF_API_TOKEN` (required), `SEMIF_MODEL` (default `Qwen/Qwen3.5-4B`),
|
||||
`SEMIF_REVISION` (default the pinned SHA), `SEMIF_DEVICE` (default `cuda`),
|
||||
`SEMIF_VRAM_CAP_GIB`, `SEMIF_MAX_TOKENS`, `SEMIF_MAX_DECISIONS`,
|
||||
`SEMIF_MAX_BODY_BYTES`, `SEMIF_CALIBRATION` (path to a JSON `{workload: T}`; T > 0).
|
||||
`SEMIF_MAX_BODY_BYTES`, `SEMIF_MAX_QUEUE`, `SEMIF_CALIBRATION` (path to a JSON
|
||||
`{workload: T}`).
|
||||
|
||||
Startup validates every value and refuses a bad one with a `ValueError` naming
|
||||
the variable (C1, S1, S8):
|
||||
- the VRAM cap, when set, is finite and > 0 (`0` used to mean uncapped);
|
||||
- every limit is an integer ≥ 1;
|
||||
- each T is a finite number, not a bool, in [0.05, 20] (a tiny T overflowed to
|
||||
NaN, and the response then failed to render);
|
||||
- the calibration file must exist and parse.
|
||||
|
||||
## Tests (TDD, fake scorer: no torch, no model)
|
||||
|
||||
@@ -109,7 +163,13 @@ alone; any
|
||||
other exception → 500; malformed
|
||||
JSON or a wrong body shape → 422; too many decisions → 422; an oversized body → 413; requests
|
||||
are serialised (two concurrent calls never overlap inside the scorer); `/health`
|
||||
answers while a scorer call is blocked.
|
||||
answers while a scorer call is blocked. **Averaging:** `rotations` sends n rows
|
||||
in one shared call, each option once per position, with ids `<id>#o<k>`, and
|
||||
cancels a position bias exactly; `all` sends n! rows and is 422 above 4 options;
|
||||
agreement and spread are computed from the orderings; a mixed shared request
|
||||
(averaged + plain) is one engine call, with results in request order and plain
|
||||
results unchanged; expanded rows count toward the cap; `workload` + `orderings`
|
||||
→ 422.
|
||||
|
||||
## Acceptance (on fv-ml1, real model; not unit tests)
|
||||
|
||||
|
||||
@@ -0,0 +1,7 @@
|
||||
condition e2e p50 (runs) server p50 (runs)
|
||||
health (floor) 33.5 [31.7-34.7] -
|
||||
decide, short (~130 tok) 69.3 [68.8-70.1] 35.3 [35.3-35.4]
|
||||
decide, long (~2,000 tok) 130.5 [129.5-130.5] 92.1 [91.7-92.4]
|
||||
shared, 3 rotations (short) 115.0 [112.4-115.7] 77.9 [77.1-79.3]
|
||||
shared, 6 orderings (short) 118.3 [117.5-118.7] 81.2 [80.2-81.3]
|
||||
shared, 3 rotations (long) 199.5 [196.6-201.0] 158.2 [157.7-158.3]
|
||||
@@ -2,9 +2,10 @@
|
||||
from __future__ import annotations
|
||||
|
||||
import hmac
|
||||
import itertools
|
||||
import math
|
||||
import threading
|
||||
from typing import Any
|
||||
from typing import Any, Literal
|
||||
|
||||
from fastapi import FastAPI, Request
|
||||
from fastapi.concurrency import run_in_threadpool
|
||||
@@ -29,6 +30,7 @@ class Decision(BaseModel):
|
||||
id: str
|
||||
question: str
|
||||
options: list[Option]
|
||||
orderings: Literal["none", "rotations", "all"] = "none"
|
||||
|
||||
def row(self, state: State) -> dict:
|
||||
"""The SemIf row shape: exactly id, state, question, options."""
|
||||
@@ -64,6 +66,75 @@ def _first_error(exc: ValidationError) -> str:
|
||||
return f"{where}: {first.get('msg', 'invalid')}"
|
||||
|
||||
|
||||
MAX_OPTIONS_FOR_ALL = 4
|
||||
|
||||
|
||||
def ordering_perms(decision: Decision) -> list[tuple[int, ...]]:
|
||||
"""Index permutations of the caller's options, the caller's own order first."""
|
||||
n = len(decision.options)
|
||||
if decision.orderings == "rotations":
|
||||
return [tuple((start + k) % n for k in range(n)) for start in range(n)]
|
||||
if n > MAX_OPTIONS_FOR_ALL:
|
||||
raise ApiError(422, "invalid_request",
|
||||
f"orderings 'all' allows at most {MAX_OPTIONS_FOR_ALL} options ({n} given); use 'rotations'")
|
||||
return list(itertools.permutations(range(n)))
|
||||
|
||||
|
||||
def expanded_rows(decision: Decision, state: State, perms: list[tuple[int, ...]]) -> list[dict]:
|
||||
base = decision.row(state)
|
||||
return [{**base, "id": f"{decision.id}#o{k}", "options": [base["options"][i] for i in perm]}
|
||||
for k, perm in enumerate(perms)]
|
||||
|
||||
|
||||
def combine(decision: Decision, results: list[dict]) -> dict:
|
||||
"""Average per-ordering log-softmax by option id; the native results ride along unchanged."""
|
||||
option_ids = [o.id for o in decision.options]
|
||||
logp: dict[str, list[float]] = {i: [] for i in option_ids}
|
||||
probs: dict[str, list[float]] = {i: [] for i in option_ids}
|
||||
tops = []
|
||||
for result in results:
|
||||
logits = result["option_logits"]
|
||||
top = max(logits)
|
||||
lse = top + math.log(sum(math.exp(x - top) for x in logits))
|
||||
for oid, x, p in zip(result["option_ids"], logits, result["probabilities"]):
|
||||
logp[oid].append(x - lse)
|
||||
probs[oid].append(p)
|
||||
tops.append(result["option_ids"][logits.index(top)])
|
||||
means = [sum(logp[i]) / len(logp[i]) for i in option_ids]
|
||||
peak = max(means)
|
||||
weights = [math.exp(m - peak) for m in means]
|
||||
combined_p = [w / sum(weights) for w in weights]
|
||||
winner = option_ids[combined_p.index(max(combined_p))]
|
||||
return {
|
||||
"id": decision.id,
|
||||
"option_ids": option_ids,
|
||||
"combined": {
|
||||
"method": decision.orderings,
|
||||
"orderings": len(results),
|
||||
"probabilities": combined_p,
|
||||
"top": winner,
|
||||
"agreement": tops.count(winner) / len(tops),
|
||||
"spread": {i: [min(probs[i]), max(probs[i])] for i in option_ids},
|
||||
},
|
||||
"orderings": results,
|
||||
}
|
||||
|
||||
|
||||
async def read_limited(stream, declared: str | None, limit: int) -> bytes:
|
||||
"""Read a request body, refusing it once it would exceed `limit` bytes. The check runs BEFORE a
|
||||
chunk is kept, and nothing after the crossing chunk is read (bug hunt C2). A declared length is
|
||||
trusted only as ASCII digits: `"²".isdigit()` is True but `int("²")` raises (S3)."""
|
||||
too_large = ApiError(413, "request_too_large", f"request body exceeds {limit} bytes")
|
||||
if declared is not None and declared.isascii() and declared.isdigit() and int(declared) > limit:
|
||||
raise too_large
|
||||
body = bytearray()
|
||||
async for chunk in stream:
|
||||
if len(body) + len(chunk) > limit:
|
||||
raise too_large
|
||||
body.extend(chunk)
|
||||
return bytes(body)
|
||||
|
||||
|
||||
def calibrated_view(result: dict, workload: str, temperature: float) -> dict:
|
||||
"""softmax(option_logits / T): the native fields are left exactly as SemIf returned them (INV-1)."""
|
||||
scaled = [x / temperature for x in result["option_logits"]]
|
||||
@@ -77,6 +148,27 @@ def create_app(settings: Settings, engine: Any) -> FastAPI:
|
||||
app = FastAPI(title="semif-serve")
|
||||
expected = f"Bearer {settings.api_token}".encode()
|
||||
inference = threading.Lock() # INV-2: one scorer call at a time, off the event loop
|
||||
in_progress = 0 # POSTs admitted and not yet answered (bug hunt C6)
|
||||
|
||||
def admit():
|
||||
nonlocal in_progress
|
||||
if in_progress >= settings.max_queue:
|
||||
raise ApiError(429, "busy", f"{in_progress} requests already in progress (limit {settings.max_queue})")
|
||||
in_progress += 1
|
||||
|
||||
def leave():
|
||||
nonlocal in_progress
|
||||
in_progress -= 1
|
||||
|
||||
def build(fn, *args):
|
||||
"""Response construction from a scorer result (calibration, averaging) maps its failures to
|
||||
the 500 envelope too, instead of escaping as a bare 500 (bug hunt C3)."""
|
||||
try:
|
||||
return fn(*args)
|
||||
except ApiError:
|
||||
raise
|
||||
except Exception as exc: # noqa: BLE001
|
||||
raise ApiError(500, "scoring_failed", f"building the response failed: {type(exc).__name__}: {exc}") from exc
|
||||
|
||||
def locked(fn, *args):
|
||||
with inference:
|
||||
@@ -94,22 +186,10 @@ def create_app(settings: Settings, engine: Any) -> FastAPI:
|
||||
async def api_error(_request: Request, exc: ApiError):
|
||||
return error(exc.status, exc.code, exc.message)
|
||||
|
||||
async def read_limited(request: Request) -> bytes:
|
||||
limit = settings.max_body_bytes
|
||||
too_large = ApiError(413, "request_too_large", f"request body exceeds {limit} bytes")
|
||||
declared = request.headers.get("content-length")
|
||||
if declared is not None and declared.isdigit() and int(declared) > limit:
|
||||
raise too_large
|
||||
body = bytearray()
|
||||
async for chunk in request.stream(): # also caps bodies that declare no length
|
||||
body.extend(chunk)
|
||||
if len(body) > limit:
|
||||
raise too_large
|
||||
return bytes(body)
|
||||
|
||||
async def parse(request: Request, model: type[BaseModel]):
|
||||
try:
|
||||
return model.model_validate_json(await read_limited(request))
|
||||
return model.model_validate_json(
|
||||
await read_limited(request.stream(), request.headers.get("content-length"), settings.max_body_bytes))
|
||||
except ValidationError as exc:
|
||||
raise ApiError(422, "invalid_request", _first_error(exc)) from exc
|
||||
|
||||
@@ -142,20 +222,59 @@ def create_app(settings: Settings, engine: Any) -> FastAPI:
|
||||
"vram_cap_gib": settings.vram_cap_gib, "max_tokens": settings.max_tokens,
|
||||
"max_decisions": settings.max_decisions, "workloads": sorted(settings.calibration)}
|
||||
|
||||
async def score_batch(decisions: list[Decision], state: State, workload: str | None) -> tuple[list[dict], dict]:
|
||||
"""One engine.shared call for every row of every decision; results in request order."""
|
||||
plan = [] # (decision, perms or None, row count)
|
||||
rows: list[dict] = []
|
||||
for d in decisions:
|
||||
if d.orderings == "none":
|
||||
plan.append((d, None, 1))
|
||||
rows.append(d.row(state))
|
||||
else:
|
||||
if workload is not None:
|
||||
raise ApiError(422, "invalid_request",
|
||||
"workload calibration is not available together with orderings")
|
||||
perms = ordering_perms(d)
|
||||
plan.append((d, perms, len(perms)))
|
||||
rows.extend(expanded_rows(d, state, perms))
|
||||
if not 1 <= len(rows) <= settings.max_decisions:
|
||||
raise ApiError(422, "invalid_request",
|
||||
f"this request expands to {len(rows)} scored rows; the limit is 1..{settings.max_decisions}")
|
||||
temperature = temperature_for(workload)
|
||||
results, timing = await score(engine.shared, rows)
|
||||
out, cursor = [], 0
|
||||
for d, perms, count in plan:
|
||||
chunk = results[cursor:cursor + count]
|
||||
cursor += count
|
||||
out.append(build(with_calibration, chunk[0], workload, temperature) if perms is None
|
||||
else build(combine, d, chunk))
|
||||
return out, timing
|
||||
|
||||
@app.post("/decide")
|
||||
async def decide(request: Request):
|
||||
body = await parse(request, DecideBody)
|
||||
temperature = temperature_for(body.workload)
|
||||
return with_calibration(await score(engine.direct, body.row(body.state)), body.workload, temperature)
|
||||
admit()
|
||||
try:
|
||||
body = await parse(request, DecideBody)
|
||||
if body.orderings != "none":
|
||||
results, _timing = await score_batch([body], body.state, body.workload)
|
||||
return results[0]
|
||||
temperature = temperature_for(body.workload)
|
||||
result = await score(engine.direct, body.row(body.state))
|
||||
return build(with_calibration, result, body.workload, temperature)
|
||||
finally:
|
||||
leave()
|
||||
|
||||
@app.post("/decide/shared")
|
||||
async def decide_shared(request: Request):
|
||||
body = await parse(request, SharedBody)
|
||||
if not 1 <= len(body.decisions) <= settings.max_decisions:
|
||||
raise ApiError(422, "invalid_request",
|
||||
f"decisions must hold 1..{settings.max_decisions} entries, got {len(body.decisions)}")
|
||||
temperature = temperature_for(body.workload)
|
||||
results, timing = await score(engine.shared, [d.row(body.state) for d in body.decisions])
|
||||
return {"results": [with_calibration(r, body.workload, temperature) for r in results], "timing": timing}
|
||||
admit()
|
||||
try:
|
||||
body = await parse(request, SharedBody)
|
||||
if not 1 <= len(body.decisions) <= settings.max_decisions:
|
||||
raise ApiError(422, "invalid_request",
|
||||
f"decisions must hold 1..{settings.max_decisions} entries, got {len(body.decisions)}")
|
||||
results, timing = await score_batch(body.decisions, body.state, body.workload)
|
||||
return {"results": results, "timing": timing}
|
||||
finally:
|
||||
leave()
|
||||
|
||||
return app
|
||||
|
||||
@@ -1,4 +1,9 @@
|
||||
"""Settings for semif-serve. Contract: semif-serve.contract.md § Configuration."""
|
||||
"""Settings for semif-serve. Contract: semif-serve.contract.md § Configuration.
|
||||
|
||||
Every value is validated at startup and a bad one is refused with a ValueError naming the
|
||||
variable: a service that starts and then rejects every request (or runs uncapped) is worse
|
||||
than one that does not start (bug hunt 2026-09-27: C1, S1, S2, S8).
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
@@ -8,6 +13,7 @@ from dataclasses import dataclass, field
|
||||
from pathlib import Path
|
||||
|
||||
MIN_TOKEN_CHARS = 32
|
||||
MIN_TEMPERATURE, MAX_TEMPERATURE = 0.05, 20.0
|
||||
SEMIF_COMMIT = "23cf1f39fc9534fe81437200959b6dfc7106e45a"
|
||||
DEFAULT_MODEL = "Qwen/Qwen3.5-4B"
|
||||
DEFAULT_REVISION = "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a"
|
||||
@@ -23,35 +29,71 @@ class Settings:
|
||||
max_tokens: int = 4096
|
||||
max_decisions: int = 64
|
||||
max_body_bytes: int = 1024 * 1024
|
||||
max_queue: int = 32
|
||||
calibration: dict[str, float] = field(default_factory=dict)
|
||||
|
||||
@classmethod
|
||||
def from_env(cls, env: Mapping[str, str]) -> "Settings":
|
||||
token = env.get("SEMIF_API_TOKEN", "")
|
||||
if len(token) < MIN_TOKEN_CHARS: # INV-6
|
||||
raise ValueError(f"SEMIF_API_TOKEN must be at least {MIN_TOKEN_CHARS} characters")
|
||||
cap = env.get("SEMIF_VRAM_CAP_GIB")
|
||||
# INV-6: visible ASCII only. A CR, LF or NUL can never arrive in a header, so a token
|
||||
# carrying one would lock every caller out while /health still said ok.
|
||||
if len(token) < MIN_TOKEN_CHARS or not all(33 <= ord(c) <= 126 for c in token):
|
||||
raise ValueError(f"SEMIF_API_TOKEN must be at least {MIN_TOKEN_CHARS} visible ASCII characters")
|
||||
return cls(
|
||||
api_token=token,
|
||||
model=env.get("SEMIF_MODEL", DEFAULT_MODEL),
|
||||
revision=env.get("SEMIF_REVISION", DEFAULT_REVISION),
|
||||
device=env.get("SEMIF_DEVICE", "cuda"),
|
||||
vram_cap_gib=float(cap) if cap else None,
|
||||
max_tokens=int(env.get("SEMIF_MAX_TOKENS", 4096)),
|
||||
max_decisions=int(env.get("SEMIF_MAX_DECISIONS", 64)),
|
||||
max_body_bytes=int(env.get("SEMIF_MAX_BODY_BYTES", 1024 * 1024)),
|
||||
vram_cap_gib=_positive_float(env, "SEMIF_VRAM_CAP_GIB"),
|
||||
max_tokens=_positive_int(env, "SEMIF_MAX_TOKENS", 4096),
|
||||
max_decisions=_positive_int(env, "SEMIF_MAX_DECISIONS", 64),
|
||||
max_body_bytes=_positive_int(env, "SEMIF_MAX_BODY_BYTES", 1024 * 1024),
|
||||
max_queue=_positive_int(env, "SEMIF_MAX_QUEUE", 32),
|
||||
calibration=_load_calibration(env.get("SEMIF_CALIBRATION")),
|
||||
)
|
||||
|
||||
|
||||
def _positive_int(env: Mapping[str, str], name: str, default: int) -> int:
|
||||
raw = env.get(name)
|
||||
if raw is None:
|
||||
return default
|
||||
try:
|
||||
value = int(raw)
|
||||
except ValueError:
|
||||
raise ValueError(f"{name} must be an integer, got {raw!r}") from None
|
||||
if value < 1:
|
||||
raise ValueError(f"{name} must be >= 1, got {value}")
|
||||
return value
|
||||
|
||||
|
||||
def _positive_float(env: Mapping[str, str], name: str) -> float | None:
|
||||
"""Unset means no cap. When set it must be finite and > 0: `0` used to slip through as 'no cap'."""
|
||||
raw = env.get(name)
|
||||
if raw is None or raw == "":
|
||||
return None
|
||||
try:
|
||||
value = float(raw)
|
||||
except ValueError:
|
||||
raise ValueError(f"{name} must be a number, got {raw!r}") from None
|
||||
if not math.isfinite(value) or value <= 0:
|
||||
raise ValueError(f"{name} must be a finite number > 0, got {raw!r}")
|
||||
return value
|
||||
|
||||
|
||||
def _load_calibration(path: str | None) -> dict[str, float]:
|
||||
"""{workload: T}, every T a finite number > 0 (T scales option logits before softmax)."""
|
||||
"""{workload: T}; T scales option logits before softmax, so it is kept in a sane range
|
||||
(a tiny T overflows to NaN and the response then fails to render)."""
|
||||
if not path:
|
||||
return {}
|
||||
table = json.loads(Path(path).read_text())
|
||||
try:
|
||||
table = json.loads(Path(path).read_text())
|
||||
except (OSError, ValueError) as exc:
|
||||
raise ValueError(f"SEMIF_CALIBRATION {path!r} could not be read as JSON: {exc}") from None
|
||||
if not isinstance(table, dict) or not all(
|
||||
isinstance(t, (int, float)) and not isinstance(t, bool) and math.isfinite(t) and t > 0
|
||||
isinstance(t, (int, float)) and not isinstance(t, bool) and math.isfinite(t)
|
||||
and MIN_TEMPERATURE <= t <= MAX_TEMPERATURE
|
||||
for t in table.values()
|
||||
):
|
||||
raise ValueError("SEMIF_CALIBRATION must be a JSON object of workload -> finite temperature > 0")
|
||||
raise ValueError(f"SEMIF_CALIBRATION must be a JSON object of workload -> temperature in "
|
||||
f"[{MIN_TEMPERATURE}, {MAX_TEMPERATURE}]")
|
||||
return {str(k): float(v) for k, v in table.items()}
|
||||
|
||||
@@ -7,10 +7,14 @@ unit-tested against a fake torch (tests/test_engine.py).
|
||||
from __future__ import annotations
|
||||
|
||||
import gc
|
||||
import logging
|
||||
import traceback
|
||||
from typing import Any, Callable
|
||||
|
||||
from .config import Settings
|
||||
from .errors import OutOfMemory
|
||||
from .errors import OutOfMemory, ScoringFailed
|
||||
|
||||
log = logging.getLogger("semif_serve.engine")
|
||||
|
||||
RELEASE_SLACK_BYTES = 512 * 2**20
|
||||
WARMUP_ROW = {
|
||||
@@ -24,6 +28,11 @@ WARMUP_ROW = {
|
||||
}
|
||||
|
||||
|
||||
def _first_line(exc: BaseException) -> str:
|
||||
lines = str(exc).splitlines()
|
||||
return lines[0] if lines else ""
|
||||
|
||||
|
||||
class TorchEngine:
|
||||
def __init__(self, torch: Any, model: Any, tokenizer: Any, metadata: dict, settings: Settings,
|
||||
direct_fn: Callable, shared_fn: Callable, release_above_bytes: int | None = None):
|
||||
@@ -46,7 +55,7 @@ class TorchEngine:
|
||||
arch = f"sm_{major}{minor}"
|
||||
if arch not in torch.cuda.get_arch_list(): # INV-3: no silent PTX/CPU fallback
|
||||
raise RuntimeError(f"torch {torch.__version__} has no kernels for {arch}: {torch.cuda.get_arch_list()}")
|
||||
if settings.vram_cap_gib: # INV-4: cap BEFORE the weights land
|
||||
if settings.vram_cap_gib is not None: # INV-4: cap BEFORE the weights land
|
||||
total = torch.cuda.get_device_properties(0).total_memory
|
||||
fraction = settings.vram_cap_gib * 2**30 / total
|
||||
if not 0 < fraction <= 1:
|
||||
@@ -82,17 +91,28 @@ class TorchEngine:
|
||||
def _guard(self, fn, *args):
|
||||
try:
|
||||
result = fn(*args)
|
||||
except ValueError:
|
||||
raise # validation: SemIf raises it before any GPU work
|
||||
except self._torch.cuda.OutOfMemoryError as exc:
|
||||
message = str(exc).splitlines()[0]
|
||||
failure, message = OutOfMemory, _first_line(exc) or "CUDA out of memory"
|
||||
except Exception as exc: # noqa: BLE001 — every other failure is released and reported below
|
||||
message = _first_line(exc)
|
||||
if "out of memory" in message.lower(): # cuBLAS/cuDNN allocation failures
|
||||
failure = OutOfMemory
|
||||
else:
|
||||
failure, message = ScoringFailed, f"{type(exc).__name__}: {message}"
|
||||
# Formatted text, not exc_info: a log record that keeps the traceback object alive
|
||||
# (pytest's capture handler does; so would any buffering handler) pins the tensors.
|
||||
log.error("scorer failed:\n%s", traceback.format_exc())
|
||||
else:
|
||||
self._release_burst()
|
||||
return result
|
||||
# INV-4, outside the except block on purpose: the torch exception's traceback holds the
|
||||
# failed scorer's frames, and with them its tensors (the replicated prefix cache). Raising
|
||||
# inside the block, or `from exc`, would chain to it and keep GiBs allocated after the 503.
|
||||
# INV-4, outside the except block on purpose: the exception's traceback holds the failed
|
||||
# scorer's frames, and with them its tensors (the replicated prefix cache). Raising inside
|
||||
# the block, or `from exc`, would chain to it and keep GiBs allocated after the response.
|
||||
gc.collect()
|
||||
self._torch.cuda.empty_cache()
|
||||
raise OutOfMemory(message)
|
||||
raise failure(message)
|
||||
|
||||
def direct(self, row: dict) -> dict:
|
||||
return self._guard(self._direct, self._model, self._tokenizer, row, self._metadata, self._settings.max_tokens)
|
||||
|
||||
@@ -3,3 +3,8 @@
|
||||
|
||||
class OutOfMemory(RuntimeError):
|
||||
"""The engine ran out of GPU memory during a request and has already released its cache (INV-4)."""
|
||||
|
||||
|
||||
class ScoringFailed(RuntimeError):
|
||||
"""A scorer call failed for a reason other than validation or OOM. Raised unchained, after the
|
||||
failed call's memory has been released; the original traceback is logged, not carried (INV-4)."""
|
||||
|
||||
@@ -11,6 +11,10 @@ from .config import Settings
|
||||
|
||||
def app_from_env() -> FastAPI:
|
||||
settings = Settings.from_env(os.environ)
|
||||
# INV-5: never download at runtime, inside the image or out of it (bug hunt S10). Set before
|
||||
# torch / transformers / huggingface_hub are imported, since they read it at import time.
|
||||
os.environ["HF_HUB_OFFLINE"] = "1"
|
||||
os.environ["TRANSFORMERS_OFFLINE"] = "1"
|
||||
from .engine import TorchEngine # torch loads only here, never in the unit tests
|
||||
|
||||
return create_app(settings, TorchEngine.load(settings))
|
||||
|
||||
@@ -1,5 +1,7 @@
|
||||
"""semif-serve HTTP behaviour against a fake engine (no torch, no model).
|
||||
Contract: services/semif-serve/semif-serve.contract.md"""
|
||||
import json
|
||||
|
||||
import pytest
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
@@ -231,3 +233,104 @@ def test_health_reports_the_pins_limits_and_workloads():
|
||||
body = client.get("/health").json()
|
||||
assert body == {"status": "ok", "semif_commit": SEMIF_COMMIT, "model": FakeEngine().health(),
|
||||
"vram_cap_gib": 12.0, "max_tokens": 4096, "max_decisions": 8, "workloads": ["alerts", "triage"]}
|
||||
|
||||
|
||||
def test_a_body_of_exactly_the_limit_is_accepted():
|
||||
body = json.dumps(ROW).encode()
|
||||
client = make_client(max_body_bytes=len(body))
|
||||
assert client.post("/decide", content=body, headers={**AUTH, "content-type": "application/json"}).status_code == 200
|
||||
|
||||
|
||||
def test_a_non_ascii_digit_content_length_is_ignored_not_a_crash():
|
||||
"""HTTP clients cannot send one (httpx refuses; h11 rejects it), so check the reader directly."""
|
||||
import asyncio
|
||||
from semif_serve.app import read_limited
|
||||
|
||||
async def body():
|
||||
yield b"{}"
|
||||
|
||||
assert asyncio.run(read_limited(body(), "²", 100)) == b"{}" # int("²") would raise
|
||||
|
||||
|
||||
def test_exactly_max_decisions_is_accepted():
|
||||
decisions = [{"id": str(i), "question": "Q?", "options": OPTIONS} for i in range(3)]
|
||||
client = make_client(max_decisions=3)
|
||||
assert client.post("/decide/shared", json={"state": "s", "decisions": decisions}, headers=AUTH).status_code == 200
|
||||
|
||||
|
||||
def test_calibration_leaves_every_native_field_alone():
|
||||
engine = FakeEngine(logits=(3.0, 1.0))
|
||||
body = make_client(engine, calibration={"triage": 2.0}).post(
|
||||
"/decide", json={**ROW, "workload": "triage"}, headers=AUTH).json()
|
||||
body.pop("calibrated")
|
||||
assert body == engine.direct(ROW)
|
||||
|
||||
|
||||
class MalformedEngine(FakeEngine):
|
||||
def direct(self, row):
|
||||
return {"id": row["id"], "option_ids": ["yes", "no"], "probabilities": [0.5, 0.5]} # no option_logits
|
||||
|
||||
def shared(self, rows):
|
||||
return [self.direct(r) for r in rows], {}
|
||||
|
||||
|
||||
@pytest.mark.parametrize("path, body", [
|
||||
("/decide", {**ROW, "workload": "triage"}),
|
||||
("/decide", {**ROW, "orderings": "rotations"}),
|
||||
])
|
||||
def test_a_malformed_scorer_result_is_an_envelope_500_not_a_bare_one(path, body):
|
||||
client = make_client(MalformedEngine(), calibration={"triage": 2.0})
|
||||
response = client.post(path, json=body, headers=AUTH)
|
||||
assert response.status_code == 500
|
||||
assert response.json()["error"]["code"] == "scoring_failed"
|
||||
|
||||
|
||||
class SharedSlowEngine(SlowEngine):
|
||||
def shared(self, rows):
|
||||
return [self.direct(r) for r in rows], {}
|
||||
|
||||
|
||||
def test_shared_requests_are_serialised_too():
|
||||
from concurrent.futures import ThreadPoolExecutor
|
||||
engine = SharedSlowEngine()
|
||||
body = {"state": "s", "decisions": [{"id": "a", "question": "Q?", "options": OPTIONS}]}
|
||||
with make_client(engine) as client, ThreadPoolExecutor(3) as pool:
|
||||
futures = [pool.submit(client.post, "/decide/shared", json=body, headers=AUTH) for _ in range(3)]
|
||||
assert engine.entered.wait(5)
|
||||
engine.release.set()
|
||||
assert [f.result().status_code for f in futures] == [200] * 3
|
||||
assert engine.peak == 1
|
||||
|
||||
|
||||
def test_more_than_max_queue_requests_in_progress_get_429_busy():
|
||||
from concurrent.futures import ThreadPoolExecutor
|
||||
engine = SlowEngine()
|
||||
with make_client(engine, max_queue=2) as client, ThreadPoolExecutor(3) as pool:
|
||||
held = [pool.submit(client.post, "/decide", json={**ROW, "id": f"r{i}"}, headers=AUTH) for i in range(2)]
|
||||
assert engine.entered.wait(5)
|
||||
import time
|
||||
deadline = time.monotonic() + 5
|
||||
while engine.inside + 0 < 1 and time.monotonic() < deadline:
|
||||
time.sleep(0.01)
|
||||
time.sleep(0.2) # let the second request reach the queue
|
||||
extra = client.post("/decide", json={**ROW, "id": "extra"}, headers=AUTH)
|
||||
assert extra.status_code == 429 and extra.json()["error"]["code"] == "busy"
|
||||
engine.release.set()
|
||||
assert [f.result().status_code for f in held] == [200, 200]
|
||||
assert make_client(FakeEngine(), max_queue=2).post("/decide", json=ROW, headers=AUTH).status_code == 200
|
||||
|
||||
|
||||
def test_read_limited_stops_reading_at_the_crossing_chunk():
|
||||
import asyncio
|
||||
from semif_serve.app import ApiError, read_limited
|
||||
consumed = []
|
||||
|
||||
async def chunks():
|
||||
for i in range(10):
|
||||
consumed.append(i)
|
||||
yield b"x" * 100
|
||||
|
||||
with pytest.raises(ApiError) as info:
|
||||
asyncio.run(read_limited(chunks(), None, 250))
|
||||
assert info.value.status == 413
|
||||
assert consumed == [0, 1, 2] # the third chunk crosses 250 and is never kept; nothing after is read
|
||||
|
||||
@@ -34,3 +34,45 @@ def test_a_calibration_table_needs_positive_finite_numbers(tmp_path, table):
|
||||
cal.write_text(json.dumps(table))
|
||||
with pytest.raises(ValueError, match="SEMIF_CALIBRATION"):
|
||||
Settings.from_env({"SEMIF_API_TOKEN": TOKEN, "SEMIF_CALIBRATION": str(cal)})
|
||||
|
||||
|
||||
@pytest.mark.parametrize("var, value", [
|
||||
("SEMIF_VRAM_CAP_GIB", "0"), ("SEMIF_VRAM_CAP_GIB", "-4"), ("SEMIF_VRAM_CAP_GIB", "nan"), ("SEMIF_VRAM_CAP_GIB", "inf"),
|
||||
("SEMIF_MAX_TOKENS", "0"), ("SEMIF_MAX_DECISIONS", "-1"), ("SEMIF_MAX_BODY_BYTES", "0"), ("SEMIF_MAX_QUEUE", "0"),
|
||||
("SEMIF_MAX_TOKENS", "lots"), ("SEMIF_VRAM_CAP_GIB", "twelve"),
|
||||
])
|
||||
def test_out_of_range_or_unparseable_values_are_refused_naming_the_variable(var, value):
|
||||
with pytest.raises(ValueError, match=var):
|
||||
Settings.from_env({"SEMIF_API_TOKEN": TOKEN, var: value})
|
||||
|
||||
|
||||
@pytest.mark.parametrize("token", ["x" * 31 + "\n", "x" * 30 + "\r\n", "x" * 32 + "\x00", "x" * 16 + " " + "x" * 16,
|
||||
"é" * 32])
|
||||
def test_a_token_with_non_visible_ascii_is_refused(token):
|
||||
with pytest.raises(ValueError, match="SEMIF_API_TOKEN"):
|
||||
Settings.from_env({"SEMIF_API_TOKEN": token})
|
||||
|
||||
|
||||
@pytest.mark.parametrize("table", [{"w": True}, {"w": 1e-300}, {"w": 0.04}, {"w": 21}])
|
||||
def test_a_temperature_must_be_a_real_number_in_range(tmp_path, table):
|
||||
cal = tmp_path / "cal.json"
|
||||
cal.write_text(json.dumps(table))
|
||||
with pytest.raises(ValueError, match="SEMIF_CALIBRATION"):
|
||||
Settings.from_env({"SEMIF_API_TOKEN": TOKEN, "SEMIF_CALIBRATION": str(cal)})
|
||||
|
||||
|
||||
@pytest.mark.parametrize("content", [None, "{not json"])
|
||||
def test_a_missing_or_malformed_calibration_file_is_a_named_startup_error(tmp_path, content):
|
||||
cal = tmp_path / "cal.json"
|
||||
if content is not None:
|
||||
cal.write_text(content)
|
||||
with pytest.raises(ValueError, match="SEMIF_CALIBRATION"):
|
||||
Settings.from_env({"SEMIF_API_TOKEN": TOKEN, "SEMIF_CALIBRATION": str(cal)})
|
||||
|
||||
|
||||
def test_in_range_edges_are_accepted(tmp_path):
|
||||
cal = tmp_path / "cal.json"
|
||||
cal.write_text(json.dumps({"lo": 0.05, "hi": 20}))
|
||||
s = Settings.from_env({"SEMIF_API_TOKEN": "!" + "~" * 31, "SEMIF_VRAM_CAP_GIB": "0.5", "SEMIF_MAX_QUEUE": "1",
|
||||
"SEMIF_CALIBRATION": str(cal)})
|
||||
assert (s.vram_cap_gib, s.max_queue, s.calibration) == (0.5, 1, {"lo": 0.05, "hi": 20.0})
|
||||
|
||||
@@ -7,7 +7,7 @@ import pytest
|
||||
|
||||
from semif_serve.config import Settings
|
||||
from semif_serve.engine import TorchEngine
|
||||
from semif_serve.errors import OutOfMemory
|
||||
from semif_serve.errors import OutOfMemory, ScoringFailed
|
||||
|
||||
|
||||
class FakeTorch:
|
||||
@@ -81,3 +81,33 @@ def test_a_burst_is_returned_to_the_driver_after_the_call(reserved_after, releas
|
||||
direct_fn=scorer, shared_fn=scorer, release_above_bytes=8 * 2**30 + 512 * 2**20)
|
||||
assert engine.direct({}) == {"ok": True}
|
||||
assert ReservingTorch.cuda.emptied == released
|
||||
|
||||
|
||||
def cyclic_tensor():
|
||||
"""A tensor held in a reference cycle, as real frames and tensors often are: only gc frees it."""
|
||||
t = Tensor()
|
||||
t.self_ref = t
|
||||
FakeTorch.cuda.watched.append(weakref.ref(t))
|
||||
return t
|
||||
|
||||
|
||||
@pytest.mark.parametrize("raised, expected_type, expected_message", [
|
||||
(lambda: FakeTorch.cuda.OutOfMemoryError(""), OutOfMemory, "CUDA out of memory"),
|
||||
(lambda: RuntimeError("CUBLAS_STATUS_ALLOC_FAILED: CUDA error: out of memory"), OutOfMemory,
|
||||
"CUBLAS_STATUS_ALLOC_FAILED: CUDA error: out of memory"),
|
||||
(lambda: RuntimeError("Invalid native prefix cache"), ScoringFailed, "RuntimeError: Invalid native prefix cache"),
|
||||
(lambda: KeyError("option_logits"), ScoringFailed, "KeyError: 'option_logits'"),
|
||||
])
|
||||
def test_every_non_validation_failure_is_released_unchained_after_gc(raised, expected_type, expected_message):
|
||||
FakeTorch.cuda.empties.clear(), FakeTorch.cuda.watched.clear()
|
||||
|
||||
def scorer(*_args):
|
||||
kv_cache = cyclic_tensor() # noqa: F841 — alive in this frame when it raises
|
||||
raise raised()
|
||||
|
||||
engine = TorchEngine(FakeTorch, None, None, {}, Settings(api_token="t" * 40), direct_fn=scorer, shared_fn=scorer)
|
||||
with pytest.raises(expected_type) as info:
|
||||
engine.shared([])
|
||||
assert str(info.value) == expected_message
|
||||
assert info.value.__cause__ is None and info.value.__context__ is None
|
||||
assert FakeTorch.cuda.empties == [True] # gc freed the cycle BEFORE the cache was emptied
|
||||
|
||||
@@ -0,0 +1,110 @@
|
||||
"""TorchEngine.load and the entry point, against fake torch/semif modules (INV-3, INV-4, INV-5).
|
||||
The bug hunt found load() had no test at all: removing the arch check survived every test."""
|
||||
import os
|
||||
import sys
|
||||
import types
|
||||
|
||||
import pytest
|
||||
|
||||
from semif_serve.config import Settings
|
||||
|
||||
TOKEN = "t" * 40
|
||||
|
||||
|
||||
class Param:
|
||||
def __init__(self, device):
|
||||
self.device = types.SimpleNamespace(type=device)
|
||||
|
||||
|
||||
class Model:
|
||||
def __init__(self, device):
|
||||
self._device = device
|
||||
|
||||
def parameters(self):
|
||||
yield Param(self._device)
|
||||
|
||||
|
||||
@pytest.fixture
|
||||
def fakes(monkeypatch):
|
||||
calls = []
|
||||
cuda = types.SimpleNamespace(
|
||||
is_available=lambda: True,
|
||||
get_device_capability=lambda _i=0: (12, 0),
|
||||
get_arch_list=lambda: ["sm_90", "sm_120"],
|
||||
get_device_properties=lambda _i=0: types.SimpleNamespace(total_memory=96 * 2**30),
|
||||
set_per_process_memory_fraction=lambda f, _i=0: calls.append(("cap", round(f, 4))),
|
||||
memory_reserved=lambda _i=0: 8 * 2**30,
|
||||
OutOfMemoryError=type("OutOfMemoryError", (RuntimeError,), {}),
|
||||
empty_cache=lambda: calls.append(("empty",)),
|
||||
)
|
||||
torch = types.SimpleNamespace(cuda=cuda, __version__="2.10.0+cu128")
|
||||
state = {"device": "cuda", "warmup_raises": None}
|
||||
|
||||
def load_causal_model(model, revision, device, dtype):
|
||||
calls.append(("load", model, revision, device, dtype))
|
||||
return Model(state["device"]), object(), {"source": model}
|
||||
|
||||
def score(model, tok, row, meta, max_tokens):
|
||||
calls.append(("score", row["id"]))
|
||||
if state["warmup_raises"]:
|
||||
raise state["warmup_raises"]
|
||||
return {"id": row["id"]}
|
||||
|
||||
core = types.ModuleType("semif_phase1.core")
|
||||
core.load_causal_model = load_causal_model
|
||||
direct = types.ModuleType("semif_phase1.direct")
|
||||
direct.score = score
|
||||
shared = types.ModuleType("semif_phase1.shared")
|
||||
shared.score_shared = lambda *a: ([], {})
|
||||
pkg = types.ModuleType("semif_phase1")
|
||||
for name, mod in {"torch": torch, "semif_phase1": pkg, "semif_phase1.core": core,
|
||||
"semif_phase1.direct": direct, "semif_phase1.shared": shared}.items():
|
||||
monkeypatch.setitem(sys.modules, name, mod)
|
||||
return torch, calls, state
|
||||
|
||||
|
||||
def test_load_caps_before_the_weights_land_then_warms_up(fakes):
|
||||
from semif_serve.engine import TorchEngine
|
||||
_torch, calls, _ = fakes
|
||||
TorchEngine.load(Settings(api_token=TOKEN, vram_cap_gib=12.0))
|
||||
assert [c[0] for c in calls] == ["cap", "load", "score"]
|
||||
assert calls[0] == ("cap", round(12 / 96, 4))
|
||||
assert calls[2] == ("score", "semif-serve-warmup")
|
||||
|
||||
|
||||
def test_load_refuses_a_card_torch_has_no_kernels_for(fakes):
|
||||
from semif_serve.engine import TorchEngine
|
||||
torch, calls, _ = fakes
|
||||
torch.cuda.get_arch_list = lambda: ["sm_80", "sm_90"]
|
||||
with pytest.raises(RuntimeError, match="sm_120"):
|
||||
TorchEngine.load(Settings(api_token=TOKEN))
|
||||
assert not any(c[0] == "load" for c in calls)
|
||||
|
||||
|
||||
def test_load_refuses_a_model_that_landed_on_the_wrong_device(fakes):
|
||||
from semif_serve.engine import TorchEngine
|
||||
_torch, _calls, state = fakes
|
||||
state["device"] = "cpu"
|
||||
with pytest.raises(RuntimeError, match="landed on cpu"):
|
||||
TorchEngine.load(Settings(api_token=TOKEN))
|
||||
|
||||
|
||||
def test_load_fails_closed_when_the_warmup_decision_fails(fakes):
|
||||
from semif_serve.engine import TorchEngine
|
||||
from semif_serve.errors import ScoringFailed
|
||||
_torch, _calls, state = fakes
|
||||
state["warmup_raises"] = RuntimeError("Failed to find C compiler")
|
||||
with pytest.raises(ScoringFailed, match="C compiler"):
|
||||
TorchEngine.load(Settings(api_token=TOKEN))
|
||||
|
||||
|
||||
def test_the_entry_point_forces_offline_mode_before_the_engine_loads(fakes, monkeypatch):
|
||||
import semif_serve.engine as engine_mod
|
||||
from semif_serve import main
|
||||
seen = {}
|
||||
monkeypatch.delenv("HF_HUB_OFFLINE", raising=False)
|
||||
monkeypatch.setenv("SEMIF_API_TOKEN", TOKEN)
|
||||
monkeypatch.setattr(engine_mod.TorchEngine, "load",
|
||||
classmethod(lambda cls, s: seen.update(offline=os.environ.get("HF_HUB_OFFLINE")) or object()))
|
||||
main.app_from_env()
|
||||
assert seen["offline"] == "1"
|
||||
@@ -0,0 +1,132 @@
|
||||
"""Order averaging (0.1.3). Contract: semif-serve.contract.md § Order averaging."""
|
||||
import math
|
||||
|
||||
import pytest
|
||||
from fastapi.testclient import TestClient
|
||||
|
||||
from semif_serve.app import create_app
|
||||
from semif_serve.config import Settings
|
||||
|
||||
TOKEN = "t" * 40
|
||||
AUTH = {"Authorization": f"Bearer {TOKEN}"}
|
||||
OPTS = [{"id": "casual", "description": "A casual outing"},
|
||||
{"id": "date", "description": "A romantic date"},
|
||||
{"id": "booty", "description": "A booty call"}]
|
||||
ROW = {"id": "q", "state": "It's 2AM and I'm bored.", "question": "What is this?", "options": OPTS}
|
||||
BASE = {"casual": 1.0, "date": -1.0, "booty": 0.5} # what the model "really" thinks
|
||||
BIAS = 2.5 # added to whichever option is listed first
|
||||
|
||||
|
||||
def softmax(xs):
|
||||
top = max(xs)
|
||||
w = [math.exp(x - top) for x in xs]
|
||||
return [v / sum(w) for v in w]
|
||||
|
||||
|
||||
class BiasedEngine:
|
||||
"""Scores each option as BASE[id], plus BIAS for the first-listed option, the way a small model leans."""
|
||||
|
||||
def __init__(self):
|
||||
self.shared_calls, self.direct_calls = [], []
|
||||
|
||||
def health(self):
|
||||
return {}
|
||||
|
||||
def _score(self, row):
|
||||
ids = [o["id"] for o in row["options"]]
|
||||
logits = [BASE[i] + (BIAS if k == 0 else 0.0) for k, i in enumerate(ids)]
|
||||
return {"id": row["id"], "option_ids": ids, "probabilities": softmax(logits), "option_logits": logits,
|
||||
"prompt_sha256": row["id"], "probability_status": "uncalibrated"}
|
||||
|
||||
def direct(self, row):
|
||||
self.direct_calls.append(row)
|
||||
return self._score(row)
|
||||
|
||||
def shared(self, rows):
|
||||
self.shared_calls.append(rows)
|
||||
return [self._score(r) for r in rows], {"batch_size": len(rows)}
|
||||
|
||||
|
||||
def client(engine, **kw):
|
||||
return TestClient(create_app(Settings(api_token=TOKEN, **kw), engine))
|
||||
|
||||
|
||||
def test_rotations_send_one_shared_call_with_every_option_once_per_position_and_cancel_the_bias():
|
||||
engine = BiasedEngine()
|
||||
body = client(engine).post("/decide", json={**ROW, "orderings": "rotations"}, headers=AUTH).json()
|
||||
assert engine.direct_calls == [] and len(engine.shared_calls) == 1
|
||||
rows = engine.shared_calls[0]
|
||||
assert [r["id"] for r in rows] == ["q#o0", "q#o1", "q#o2"]
|
||||
assert [[o["id"] for o in r["options"]] for r in rows] == [
|
||||
["casual", "date", "booty"], ["date", "booty", "casual"], ["booty", "casual", "date"]]
|
||||
assert all(r["state"] == ROW["state"] and r["question"] == ROW["question"] for r in rows)
|
||||
|
||||
assert body["id"] == "q" and body["option_ids"] == ["casual", "date", "booty"]
|
||||
combined = body["combined"]
|
||||
assert combined["method"] == "rotations" and combined["orderings"] == 3
|
||||
# the first-position bias is additive, and every option sat first exactly once, so it cancels exactly
|
||||
assert combined["probabilities"] == pytest.approx(softmax([BASE["casual"], BASE["date"], BASE["booty"]]))
|
||||
assert combined["top"] == "casual"
|
||||
# the bias is strong enough that every ordering's first option wins: casual, date, booty
|
||||
tops = [max(zip(r["option_logits"], r["option_ids"]))[1] for r in body["orderings"]]
|
||||
assert tops == ["casual", "date", "booty"]
|
||||
assert combined["agreement"] == pytest.approx(1 / 3)
|
||||
assert body["orderings"] == [engine._score(r) for r in rows] # native results, unchanged
|
||||
for oid in ("casual", "date", "booty"):
|
||||
ps = [dict(zip(r["option_ids"], r["probabilities"]))[oid] for r in body["orderings"]]
|
||||
assert combined["spread"][oid] == pytest.approx([min(ps), max(ps)])
|
||||
|
||||
|
||||
def test_all_sends_every_permutation_with_the_callers_order_first():
|
||||
engine = BiasedEngine()
|
||||
body = client(engine).post("/decide", json={**ROW, "orderings": "all"}, headers=AUTH).json()
|
||||
rows = engine.shared_calls[0]
|
||||
orders = [tuple(o["id"] for o in r["options"]) for r in rows]
|
||||
assert len(orders) == 6 and len(set(orders)) == 6 and orders[0] == ("casual", "date", "booty")
|
||||
assert body["combined"]["orderings"] == 6 and body["combined"]["method"] == "all"
|
||||
|
||||
|
||||
def test_all_above_four_options_is_422_before_the_engine_runs():
|
||||
engine = BiasedEngine()
|
||||
five = [{"id": f"o{i}", "description": f"Option {i}"} for i in range(5)]
|
||||
r = client(engine).post("/decide", json={**ROW, "options": five, "orderings": "all"}, headers=AUTH)
|
||||
assert r.status_code == 422 and "rotations" in r.json()["error"]["message"]
|
||||
assert engine.shared_calls == []
|
||||
|
||||
|
||||
def test_a_mixed_shared_request_is_one_engine_call_with_results_in_request_order():
|
||||
engine = BiasedEngine()
|
||||
body = {"state": ROW["state"], "decisions": [
|
||||
{"id": "plain", "question": "Q1?", "options": OPTS},
|
||||
{"id": "avg", "question": "Q2?", "options": OPTS, "orderings": "rotations"},
|
||||
{"id": "plain2", "question": "Q3?", "options": OPTS[:2]}]}
|
||||
out = client(engine).post("/decide/shared", json=body, headers=AUTH).json()
|
||||
assert len(engine.shared_calls) == 1
|
||||
assert [r["id"] for r in engine.shared_calls[0]] == ["plain", "avg#o0", "avg#o1", "avg#o2", "plain2"]
|
||||
ids = [r["id"] for r in out["results"]]
|
||||
assert ids == ["plain", "avg", "plain2"]
|
||||
assert out["results"][0] == engine._score(engine.shared_calls[0][0]) # plain results keep the old shape
|
||||
assert "combined" in out["results"][1] and "combined" not in out["results"][2]
|
||||
assert out["timing"] == {"batch_size": 5}
|
||||
|
||||
|
||||
def test_expanded_rows_count_toward_the_cap():
|
||||
engine = BiasedEngine()
|
||||
r = client(engine, max_decisions=5).post("/decide/shared", json={"state": "s", "decisions": [
|
||||
{"id": "a", "question": "Q?", "options": OPTS, "orderings": "rotations"},
|
||||
{"id": "b", "question": "Q?", "options": OPTS, "orderings": "rotations"}]}, headers=AUTH)
|
||||
assert r.status_code == 422 and "6 scored rows" in r.json()["error"]["message"]
|
||||
assert engine.shared_calls == []
|
||||
|
||||
|
||||
def test_workload_with_orderings_is_422_and_workload_alone_still_calibrates():
|
||||
engine = BiasedEngine()
|
||||
c = client(engine, calibration={"triage": 2.0})
|
||||
r = c.post("/decide", json={**ROW, "orderings": "rotations", "workload": "triage"}, headers=AUTH)
|
||||
assert r.status_code == 422 and engine.shared_calls == []
|
||||
assert "calibrated" in c.post("/decide", json={**ROW, "workload": "triage"}, headers=AUTH).json()
|
||||
|
||||
|
||||
def test_an_unknown_orderings_value_is_422():
|
||||
r = client(BiasedEngine()).post("/decide", json={**ROW, "orderings": "shuffle"}, headers=AUTH)
|
||||
assert r.status_code == 422
|
||||
Generated
+80
-2
@@ -51,6 +51,26 @@ wheels = [
|
||||
{ url = "https://files.pythonhosted.org/packages/12/b8/4bd346e22b28902df4d651910f5242c28d84e4a5c2435ca5c3f797ed7e2e/anyio-4.15.1-py3-none-any.whl", hash = "sha256:6152fdbbf9a77fdec97731721bebf7c4c44f7c29b424b0065826173efc7ed101", size = 132079 },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "causal-conv1d"
|
||||
version = "1.7.0"
|
||||
source = { url = "https://github.com/Dao-AILab/causal-conv1d/releases/download/v1.7.0/causal_conv1d-1.7.0+cu12torch2.10cxx11abiTRUE-cp312-cp312-linux_x86_64.whl" }
|
||||
dependencies = [
|
||||
{ name = "ninja" },
|
||||
{ name = "packaging" },
|
||||
{ name = "torch" },
|
||||
]
|
||||
wheels = [
|
||||
{ url = "https://github.com/Dao-AILab/causal-conv1d/releases/download/v1.7.0/causal_conv1d-1.7.0+cu12torch2.10cxx11abiTRUE-cp312-cp312-linux_x86_64.whl", hash = "sha256:8e81f8c76435ad31aa41edc6c0c9d26de971a293b190183d2816da71547614d7" },
|
||||
]
|
||||
|
||||
[package.metadata]
|
||||
requires-dist = [
|
||||
{ name = "ninja" },
|
||||
{ name = "packaging" },
|
||||
{ name = "torch" },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "certifi"
|
||||
version = "2026.7.22"
|
||||
@@ -106,6 +126,15 @@ wheels = [
|
||||
{ url = "https://files.pythonhosted.org/packages/98/59/239c7259e669c46ddbcac0aa60e3a0ef00bfeaaa687f905b24dd6a7a10fe/cuda_pathfinder-1.8.2-py3-none-any.whl", hash = "sha256:4e65059febdb4d19d5cbc4798677e19db2b582f2f702f457b609e571690d357e", size = 62551 },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "einops"
|
||||
version = "0.8.2"
|
||||
source = { registry = "https://pypi.org/simple" }
|
||||
sdist = { url = "https://files.pythonhosted.org/packages/2c/77/850bef8d72ffb9219f0b1aac23fbc1bf7d038ee6ea666f331fa273031aa2/einops-0.8.2.tar.gz", hash = "sha256:609da665570e5e265e27283aab09e7f279ade90c4f01bcfca111f3d3e13f2827", size = 56261 }
|
||||
wheels = [
|
||||
{ url = "https://files.pythonhosted.org/packages/2a/09/f8d8f8f31e4483c10a906437b4ce31bdf3d6d417b73fe33f1a8b59e34228/einops-0.8.2-py3-none-any.whl", hash = "sha256:54058201ac7087911181bfec4af6091bb59380360f069276601256a76af08193", size = 65638 },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "fastapi"
|
||||
version = "0.118.0"
|
||||
@@ -129,6 +158,31 @@ wheels = [
|
||||
{ url = "https://files.pythonhosted.org/packages/bc/ac/8b9c6dc2aa7e9cf5582e20a337006c3428c3d752c2255dbaca2c78a8be45/filelock-4.0.4-py3-none-any.whl", hash = "sha256:0df72be195ca7892216d16f2edce8d9b93a571f02402972020a8cff84c594c7b", size = 108629 },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "fla-core"
|
||||
version = "0.5.2"
|
||||
source = { registry = "https://pypi.org/simple" }
|
||||
dependencies = [
|
||||
{ name = "einops" },
|
||||
]
|
||||
sdist = { url = "https://files.pythonhosted.org/packages/94/85/20bdc0fbbbaeec27b5b198e0dbbbc51161a267066a4aebd9e4b413637c6c/fla_core-0.5.2.tar.gz", hash = "sha256:9360fc412f784c1c8f05c320ec2902d5d968764178c9b8a92efc919e17a39680", size = 587835 }
|
||||
wheels = [
|
||||
{ url = "https://files.pythonhosted.org/packages/2d/ed/dfe19c4da779957eb6a42a26812f9b4e2280bf757a17a71933ff59ffcb98/fla_core-0.5.2-py3-none-any.whl", hash = "sha256:5e830c85bad3d0d34677f98ac7074d08687a3756f0f0499d95ceb96eb6920761", size = 819225 },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "flash-linear-attention"
|
||||
version = "0.5.2"
|
||||
source = { registry = "https://pypi.org/simple" }
|
||||
dependencies = [
|
||||
{ name = "fla-core" },
|
||||
{ name = "transformers" },
|
||||
]
|
||||
sdist = { url = "https://files.pythonhosted.org/packages/d2/76/c180949eae5161b9fcf4c928cab52f0c257ce84c1a59b4462b1d2b6e5883/flash_linear_attention-0.5.2.tar.gz", hash = "sha256:c053d3a75c8f5b725063f719ae6bce0f4d9dac6da32735be8ae7d485a6c9820f", size = 208110 }
|
||||
wheels = [
|
||||
{ url = "https://files.pythonhosted.org/packages/90/d2/2070e3cf2148c5cce99ca4876633c4b7b89ec085323600bd3d47aeacd306/flash_linear_attention-0.5.2-py3-none-any.whl", hash = "sha256:dcf405d81f5426393b59037097aa700d0f4a841465d5028d5aa543f4502f2400", size = 399590 },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "fsspec"
|
||||
version = "2026.9.0"
|
||||
@@ -351,6 +405,24 @@ wheels = [
|
||||
{ url = "https://files.pythonhosted.org/packages/7e/cd/fe58041e9011f307c490e3e17dd48cc516448f7c698a3f2d9d9d65d7e6a8/networkx-3.7-py3-none-any.whl", hash = "sha256:e3fd2c13a7814cee3746340d8d7f8598a67f16a58bf47fb7f8793fab6efca1b0", size = 2142205 },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "ninja"
|
||||
version = "1.13.2"
|
||||
source = { registry = "https://pypi.org/simple" }
|
||||
sdist = { url = "https://files.pythonhosted.org/packages/ac/92/410b7917d16ab54c04b05cc32b9284803671d91cf79d33be6009c28d4ea8/ninja-1.13.2.tar.gz", hash = "sha256:525bfa3fc88aa30a4467df270fd5be6f9fcae8061d54d4df74ea1dc5abd5a975", size = 243739 }
|
||||
wheels = [
|
||||
{ url = "https://files.pythonhosted.org/packages/48/23/fcbe234a66966e35928c47b86336f92a7612db4781665f4e5f5fddef9630/ninja-1.13.2-py3-none-manylinux2014_ppc64le.manylinux_2_17_ppc64le.whl", hash = "sha256:227cbc3ae3e5e429692388103cae8c09451df086cd2d342dae0795af0d162547", size = 197676 },
|
||||
{ url = "https://files.pythonhosted.org/packages/24/eb/a6ca97ef0ff7bb8bdcb395ec65a716e65d7c1f40896c3afe0090bb3e1535/ninja-1.13.2-py3-none-manylinux2014_s390x.manylinux_2_17_s390x.whl", hash = "sha256:1684c60d031c54c1d049541b64243c0c567dca5463dbd77682a8901780af293d", size = 187980 },
|
||||
{ url = "https://files.pythonhosted.org/packages/6e/53/ebfed7b689c338dd8ebeec9c0730c8d56821292f14e2536e5f3ef1a05744/ninja-1.13.2-py3-none-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:65a24341b5ac09fcadcc37082660be40a94174e51a937fabf6e2cae26225fa2c", size = 183365 },
|
||||
{ url = "https://files.pythonhosted.org/packages/c7/d6/dcf06d7ab44ade992ae5aa1228feff317684b463a1bd47e8642b30ac922e/ninja-1.13.2-py3-none-manylinux_2_28_armv7l.manylinux_2_31_armv7l.whl", hash = "sha256:aa3d2ae5706a2c4d1e93edc951d1c6cbb45107413c404f8fde1741239efbc9a0", size = 155089 },
|
||||
{ url = "https://files.pythonhosted.org/packages/e1/6b/6513c09c33382b17c05b4349b8e81437b18680d0d7ec6fb8f7edc29adda1/ninja-1.13.2-py3-none-manylinux_2_31_riscv64.whl", hash = "sha256:919572cbc3f233261ecd41fe1f3efc9d44aa02464a4588867e06a8b4f6f416ea", size = 152149 },
|
||||
{ url = "https://files.pythonhosted.org/packages/37/04/c8c2dc5b2f5fee79a1691d490256b178b7e1af97d56769117422ae8a23cc/ninja-1.13.2-py3-none-musllinux_1_2_armv7l.whl", hash = "sha256:59d71c3e15b6b6f3d903eb0c27285544e0747ca59925ada7037bb1af781ad4b3", size = 466268 },
|
||||
{ url = "https://files.pythonhosted.org/packages/10/a2/d8eedd25d0ae80b9e874aea362013e67416877c8f732eb4d5e7c971fdb9c/ninja-1.13.2-py3-none-musllinux_1_2_ppc64le.whl", hash = "sha256:b2f687437fac460b27b7eadc99039b1163016fb4ba7276e2782a192d9f24ee0e", size = 610806 },
|
||||
{ url = "https://files.pythonhosted.org/packages/14/0f/696d96821fad1b5767fd311c1569dde8881a57412369bfe7b11bcbfde036/ninja-1.13.2-py3-none-musllinux_1_2_riscv64.whl", hash = "sha256:09de9ab04f7352f51570c73fd4913acb1e6c24be0a72cd8b80243d4d3ed04925", size = 533978 },
|
||||
{ url = "https://files.pythonhosted.org/packages/5d/69/28844ca579156776a202217a7cd66f60d06a0710a935e879bb89ce396ecc/ninja-1.13.2-py3-none-musllinux_1_2_s390x.whl", hash = "sha256:6a87bf42b123abe2f37737300185f0a303a891899da85d73a3613ee80547e578", size = 653822 },
|
||||
{ url = "https://files.pythonhosted.org/packages/f5/5f/c511f2952f94ab2966d60edd9c34e744ea32f2724b1184b62270bde55b3a/ninja-1.13.2-py3-none-musllinux_1_2_x86_64.whl", hash = "sha256:915bd482c4be41c75120fd67a22e0bb3f0fbb3bbc5f95b89787deadd59e27ef2", size = 544460 },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "numpy"
|
||||
version = "2.2.6"
|
||||
@@ -919,7 +991,7 @@ dependencies = [
|
||||
|
||||
[[package]]
|
||||
name = "semif-serve"
|
||||
version = "0.1.2"
|
||||
version = "0.1.3"
|
||||
source = { editable = "." }
|
||||
dependencies = [
|
||||
{ name = "fastapi" },
|
||||
@@ -927,6 +999,10 @@ dependencies = [
|
||||
]
|
||||
|
||||
[package.optional-dependencies]
|
||||
fast = [
|
||||
{ name = "causal-conv1d", marker = "platform_machine == 'x86_64' and sys_platform == 'linux'" },
|
||||
{ name = "flash-linear-attention" },
|
||||
]
|
||||
model = [
|
||||
{ name = "semif-phase1" },
|
||||
{ name = "torch" },
|
||||
@@ -940,12 +1016,14 @@ dev = [
|
||||
|
||||
[package.metadata]
|
||||
requires-dist = [
|
||||
{ name = "causal-conv1d", marker = "platform_machine == 'x86_64' and sys_platform == 'linux' and extra == 'fast'", url = "https://github.com/Dao-AILab/causal-conv1d/releases/download/v1.7.0/causal_conv1d-1.7.0+cu12torch2.10cxx11abiTRUE-cp312-cp312-linux_x86_64.whl" },
|
||||
{ name = "fastapi", specifier = "==0.118.0" },
|
||||
{ name = "flash-linear-attention", marker = "extra == 'fast'", specifier = "==0.5.2" },
|
||||
{ name = "semif-phase1", marker = "extra == 'model'", git = "https://github.com/TheoLeeCJ/SemIf-OpenJev?rev=23cf1f39fc9534fe81437200959b6dfc7106e45a" },
|
||||
{ name = "torch", marker = "extra == 'model'", specifier = "==2.10.0", index = "https://download.pytorch.org/whl/cu128" },
|
||||
{ name = "uvicorn", specifier = "==0.37.0" },
|
||||
]
|
||||
provides-extras = ["model"]
|
||||
provides-extras = ["model", "fast"]
|
||||
|
||||
[package.metadata.requires-dev]
|
||||
dev = [
|
||||
|
||||
@@ -1,6 +1,6 @@
|
||||
# semif — copy to /opt/docker/compose/semif/.env on fv-ml1 (mode 0600).
|
||||
# Built on fv-ml1 from services/semif-serve (see README "Building").
|
||||
IMAGE=semif-serve:0.1.2
|
||||
IMAGE=semif-serve:0.1.3
|
||||
PORT=8032
|
||||
HOST_IP=10.251.50.54
|
||||
# fv-ml1 GPU 1 = the utility card (vllm-coder, erp, meromero, scriberr).
|
||||
|
||||
+56
-7
@@ -16,7 +16,7 @@ service. The contract is `services/semif-serve/semif-serve.contract.md`.
|
||||
| **Model** | `Qwen/Qwen3.5-4B` @ `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a`, BF16, from `/tank/aimodels/huggingface` (read-only, offline) |
|
||||
| **SemIf** | commit `23cf1f39fc9534fe81437200959b6dfc7106e45a`; torch `2.10.0+cu128`, transformers `5.17.0`, the same stack SemIf's committed predictions were made on |
|
||||
| **Image** | `semif-serve:<version>`, built on fv-ml1 from `services/semif-serve/` |
|
||||
| **State** | none. Weights re-download at the pinned revision. No backup needed beyond the host's `/opt/docker` restic. |
|
||||
| **State** | none. If the weights are ever lost, re-pull them by hand at the pinned revision (the service itself never downloads: `HF_HUB_OFFLINE=1`, read-only mount). No backup needed beyond the host's `/opt/docker` restic. |
|
||||
|
||||
## API
|
||||
|
||||
@@ -34,6 +34,18 @@ curl -s -H "Authorization: Bearer $T" http://10.251.50.54:8032/decide -d '{
|
||||
the state once and scores every criterion in one batch. Use it when many questions
|
||||
share one long state.
|
||||
- `GET /health`: the pins, limits and calibrated workloads.
|
||||
- **Order averaging (0.1.3):** add `"orderings": "rotations"` to a decision, or
|
||||
`"all"` for ≤ 4 options. The option list is asked in every rotation inside ONE
|
||||
shared batch. The reply keeps each ordering's native result under `orderings` and adds
|
||||
`combined: {method, orderings, probabilities, top, agreement, spread}`. **Use it for
|
||||
anything real:** a small model leans toward the first-listed option on ambiguous
|
||||
inputs, and averaging cancels that. Through the service on SemIf's labelled sets,
|
||||
accuracy goes from 78.6% to **88.1%** (group-bootstrap 95% CI +5.1..+14.3 pts, 252
|
||||
rows). **`agreement` is the cheap confidence signal**: unanimous rows are 94.5%
|
||||
accurate, split rows 76.4%. The orderings count toward the decision cap. `workload`
|
||||
calibration is not available together with `orderings` yet (422).
|
||||
- Past `MAX_QUEUE` (32) requests in progress, new POSTs get `429 busy` before their
|
||||
body is read.
|
||||
|
||||
## ⚠ Probabilities are uncalibrated
|
||||
|
||||
@@ -57,18 +69,55 @@ not squeeze scriberr, which shares GPU 1. A request that would exceed the cap ge
|
||||
behaviours were verified on the card (0.1.0 held 11.9 GiB after an OOM, and 12.6 GB
|
||||
after a large request; 0.1.2 returns to 7.85 GiB in both cases).
|
||||
|
||||
**What fits under 12 GiB** (measured, `/decide/shared`, binary decisions):
|
||||
**What fits under 12 GiB** (measured on 0.1.3, `/decide/shared`, binary decisions;
|
||||
0.1.2 figures in brackets, before the fast kernels):
|
||||
|
||||
| state size (prefix tokens) | max decisions in one request |
|
||||
| state size (prefix tokens) | max rows in one request |
|
||||
|---|---|
|
||||
| ~140 | 52 |
|
||||
| ~520 | 43 |
|
||||
| ~1,960 | 19 |
|
||||
| ~3,900 | 13 |
|
||||
| ~140 | 63 (52) |
|
||||
| ~520 | 51 (43) |
|
||||
| ~1,960 | 26 (19) |
|
||||
| ~3,900 | 16 (13) |
|
||||
|
||||
Rows = decisions × orderings, so `rotations` over 3 options uses 3 rows per decision.
|
||||
|
||||
`/decide` fits at the full 4,096-token limit. Past the table you get a 503, so split
|
||||
the decisions across requests.
|
||||
|
||||
## Fast kernels (0.1.3)
|
||||
|
||||
The image ships Qwen3.5's fast kernels, `flash-linear-attention` 0.5.2 and
|
||||
`causal-conv1d` 1.7.0 (a prebuilt cu12/torch2.10 wheel). Without them transformers
|
||||
logs that it falls back to "much slower" reference PyTorch paths. They are adopted
|
||||
because an A/B on the empty GPU 3 showed:
|
||||
- **parity improved**: 144/144 vs upstream (142/144 without), so upstream evidently
|
||||
ran with them;
|
||||
- **long inputs got much faster**: a ~2,000-token `/decide` went from 169 to 92 ms
|
||||
server-side.
|
||||
|
||||
Short 3-rotation batches cost ~3–6 ms more; everything else is equal or faster.
|
||||
⚠ triton builds a C shim at runtime, so the image carries `gcc`. Without it the
|
||||
warm-up fails, and startup fails closed. Build without the kernels:
|
||||
`--build-arg EXTRAS="--extra model"`.
|
||||
|
||||
## Latency (0.1.3, from nh3-dev, 3 runs × 20, network floor ~33 ms)
|
||||
|
||||
| request | end to end | server |
|
||||
|---|---|---|
|
||||
| `/decide`, short (~130 tok) | 69 ms | 35 ms |
|
||||
| `/decide`, ~2,000-token state | 131 ms | 92 ms |
|
||||
| 3 rotations, short | 115 ms | 78 ms |
|
||||
| 6 orderings, short | 118 ms | 81 ms |
|
||||
| 3 rotations, ~2,000-token state | 200 ms | 158 ms |
|
||||
|
||||
## Acceptance (2026-09-27, v0.1.3)
|
||||
|
||||
Raw: `services/semif-serve/acceptance/result-2026-09-27-v0.1.3.json`,
|
||||
`averaging-2026-09-27-v0.1.3.json`. v0.1.3 matches upstream on **144/144** (identical
|
||||
prompt hashes, max prob gap 0.049), is deterministic, fails the negative control as it
|
||||
should (14/144), and shared matches direct on 72/72. The v0.1.2 table below is the
|
||||
reference-kernel baseline.
|
||||
|
||||
## Acceptance (2026-09-27, v0.1.2)
|
||||
|
||||
Raw: `services/semif-serve/acceptance/result-2026-09-27-v0.1.2.json`.
|
||||
|
||||
@@ -30,6 +30,8 @@ services:
|
||||
SEMIF_VRAM_CAP_GIB: ${VRAM_CAP_GIB:?set VRAM_CAP_GIB}
|
||||
SEMIF_MAX_TOKENS: ${MAX_TOKENS:-4096}
|
||||
SEMIF_MAX_DECISIONS: ${MAX_DECISIONS:-64}
|
||||
# POSTs in progress (queued + scoring) before new ones get 429 busy.
|
||||
SEMIF_MAX_QUEUE: ${MAX_QUEUE:-32}
|
||||
SEMIF_CALIBRATION: /conf/calibration.json
|
||||
volumes:
|
||||
# Pinned weights, read offline (HF_HUB_OFFLINE=1 in the image). Never downloads.
|
||||
|
||||
Reference in New Issue
Block a user