From 77b8cb449c5ae48d28a57846195b7f9cfa6f635c Mon Sep 17 00:00:00 2001 From: Vuong Hoang Date: Sun, 27 Sep 2026 03:27:15 -0700 Subject: [PATCH] =?UTF-8?q?feat(semif):=200.1.3=20=E2=80=94=20order=20aver?= =?UTF-8?q?aging,=20fast=20kernels,=20bug-hunt=20hardening=20(Prime)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Order averaging (Prime, after the 739aa03 spike): - A decision may set orderings: rotations|all (all only for <= 4 options). Every ordering goes to the engine in one shared batch. - The reply keeps each native result and adds combined {probabilities (log-mean), top, agreement, spread}. - Through the service on SemIf's labelled sets (252 rows): 78.6% -> 88.1% (group-bootstrap 95% CI +5.1..+14.3). Unanimous agreement is 94.5% accurate. Fast kernels: flash-linear-attention 0.5.2 and causal-conv1d 1.7.0 are now the default build. A/B on the empty GPU 3: - parity with upstream went from 142/144 to 144/144; - a ~2k-token /decide went from 169 to 92 ms server-side; - short 3-rotation batches cost ~3-6 ms more. triton builds a C shim at runtime, so the image carries gcc. Without it the warm-up failed and startup failed closed. Heid bug-hunt panel (4/4 arms, thread 01M3H3F4RR7XBP90KQ3A39H4SX), folded: - Startup validation: VRAM cap 0 no longer means uncapped (C1); limits must be >= 1 (S1); the token must be visible ASCII (S2); the calibration file must exist and parse, with T in [0.05, 20] (S8, and C3's NaN leg). - The body limit is checked before a chunk is kept, and a Unicode-digit Content-Length no longer crashes (C2, S3). - Failures while building the response now get the 500 envelope (C3). - 429 busy past SEMIF_MAX_QUEUE requests in progress (C6). - The engine releases memory on every non-validation failure, unchained after gc; an empty OOM message is handled; 'out of memory' RuntimeErrors map to 503 (C4, C5, S9). - The entry point forces HF_HUB_OFFLINE (S10). README wording fixed (S5, S6). - New guard tests close the gaps the arms' mutation grids exposed: early stop of the body read, a shared-route lock, calibration pass-through, the gc cycle, the exact caps, TorchEngine.load's arch and device checks, and the offline entry point. 86 tests. Deployed on fv-ml1 GPU 1: parity 144/144, OOM and burst release verified, shared capacity 63/51/26/16 rows at ~140/520/1960/3900 prefix tokens. --- persistent-memory.md | 38 ++-- services/semif-serve/Dockerfile | 15 +- .../averaging-2026-09-27-v0.1.3.json | 18 ++ services/semif-serve/acceptance/averaging.py | 42 +++++ .../acceptance/result-2026-09-27-v0.1.3.json | 63 +++++++ services/semif-serve/pyproject.toml | 9 +- services/semif-serve/semif-serve.contract.md | 72 +++++++- .../spike/latency-2026-09-27-v0.1.3.txt | 7 + services/semif-serve/src/semif_serve/app.py | 169 +++++++++++++++--- .../semif-serve/src/semif_serve/config.py | 66 +++++-- .../semif-serve/src/semif_serve/engine.py | 34 +++- .../semif-serve/src/semif_serve/errors.py | 5 + services/semif-serve/src/semif_serve/main.py | 4 + services/semif-serve/tests/test_app.py | 103 +++++++++++ services/semif-serve/tests/test_config.py | 42 +++++ services/semif-serve/tests/test_engine.py | 32 +++- .../semif-serve/tests/test_engine_load.py | 110 ++++++++++++ services/semif-serve/tests/test_orderings.py | 132 ++++++++++++++ services/semif-serve/uv.lock | 82 ++++++++- stacks/semif/.env.example | 2 +- stacks/semif/README.md | 63 ++++++- stacks/semif/compose.yaml | 2 + 22 files changed, 1018 insertions(+), 92 deletions(-) create mode 100644 services/semif-serve/acceptance/averaging-2026-09-27-v0.1.3.json create mode 100644 services/semif-serve/acceptance/averaging.py create mode 100644 services/semif-serve/acceptance/result-2026-09-27-v0.1.3.json create mode 100644 services/semif-serve/spike/latency-2026-09-27-v0.1.3.txt create mode 100644 services/semif-serve/tests/test_engine_load.py create mode 100644 services/semif-serve/tests/test_orderings.py diff --git a/persistent-memory.md b/persistent-memory.md index e720326..ed827ef 100644 --- a/persistent-memory.md +++ b/persistent-memory.md @@ -144,33 +144,17 @@ _As of 2026-09-27 ~0255 PT._ ### SemIf on fv-ml1 GPU 1 (2026-09-27, Prime) -- **LIVE: `semif-serve` 0.1.2** at `http://10.251.50.54:8032` (`semif.fv.internal`). It wraps the SemIf - option-logit scorers (pinned `23cf1f39`, Qwen3.5-4B `851bf6e8`, BF16); token `semif/api-token`. - Code + contract: `services/semif-serve/`; stack `stacks/semif`. **No consumer yet.** -- Acceptance: 142/144 top choice vs upstream (both misses are exact bf16 ties), 144/144 prompt hashes, - deterministic, negative control 14/144. It holds 12 GiB hard with release after a burst (~8.7 GB resting). -- Latency from nh3-dev: `/decide` short 71 ms end to end (38 ms server); 3 rotations 113 ms; ~2k-token - state 210 ms (`services/semif-serve/spike/latency-2026-09-27.txt`). -- **NEXT, in this order (Prime 0250):** - 1. **Build order-averaging into semif-serve as 0.1.3.** Spike (`739aa03`): 3 rotations, log-mean, - take accuracy 78.6% → 87.7% (group-bootstrap 95% CI +4.7..+13.8 pts); unanimous rows are 94.4% - accurate, split rows 75.8%. Design: opt-in `orderings: rotations|all` (all only ≤ 4 options), - every ordering in one shared batch, per-ordering SemIf results returned unchanged + a - `combined` block (log-mean probabilities, top, agreement, per-option spread). Orderings count - toward `max_decisions`. Contract amendment → TDD → deploy → acceptance (authored144 accuracy - with/without; 2AM/2PM agreement as controls). - 2. **Trial the fast kernels** (`causal_conv1d`, `flash-linear-attention`). Today transformers runs - Qwen3.5's reference PyTorch paths and logs that they are "much slower". Adopt only if the - authored144 parity re-check holds. - 3. **heid bug-hunt panel** on 0.1.2 is pending (thread `01M3H3F4RR7XBP90KQ3A39H4SX`). Triage and fold - its findings into 0.1.3. -- NVFP4 is not worth it: the SemIf scorer runs on transformers, not vLLM; quant noise lands on the scores - and would need its own calibration; BF16 already fits. SemIf's own MLX 4-bit run moved authored144 - 0.813 → 0.789 (one run each, so indicative only). -- Prime's probes (single scenarios, not benchmarks): at 2AM vs 2PM, booty-call probability averaged over - all orderings goes 0.42 → 0.05; a baby dragon reads life-changing 0.65 / world-changing 0.25 with only - 2/3 rotations agreeing (one ordering said 0.925), the lottery ticket reads life-changing unanimously, - and the paperclip control reads trivial 0.999. +- **LIVE: `semif-serve` 0.1.3** at `http://10.251.50.54:8032` (`semif.fv.internal`), with order averaging + and the fast kernels. SemIf `23cf1f39`, Qwen3.5-4B `851bf6e8`, BF16; token `semif/api-token`. Code + + contract: `services/semif-serve/`; stack `stacks/semif`. **No consumer yet.** +- 0.1.3 acceptance: 144/144 parity with upstream; averaging through the service 78.6% → 88.1% + (95% CI +5.1..+14.3); unanimous rotations 94.5% accurate. Latency, envelope and the fast-kernel A/B + are in `stacks/semif/README.md`. +- The heid bug-hunt panel (thread `01M3H3F4RR7XBP90KQ3A39H4SX`) is triaged and folded into 0.1.3: + C1–C6, S1–S3, S5, S6, S8–S10 fixed with tests; S4 settled; S7 (publish on all interfaces) accepted + as LAN + auth. +- NVFP4 is not worth it (see the SemIf detail files). Prime's probes (2AM/2PM, dragon/lottery) are + recorded in [[2026-09-27-semif-order-averaging]]. ### restic: credential leak fixed (2026-09-27, Prime) diff --git a/services/semif-serve/Dockerfile b/services/semif-serve/Dockerfile index 7c1f758..5ece93e 100644 --- a/services/semif-serve/Dockerfile +++ b/services/semif-serve/Dockerfile @@ -22,18 +22,26 @@ l.write_text(re.sub(r'(name = "semif-serve"\nversion = )"[^"]+"', r'\1"0.0.0"', EOF FROM python:3.12-slim-bookworm +# EXTRAS picks the optional dependency sets. Default (adopted 2026-09-27): the model plus +# Qwen3.5's fast kernels (fla + causal-conv1d). "--extra model" alone gives the reference +# PyTorch paths. +ARG EXTRAS="--extra model --extra fast" COPY --from=ghcr.io/astral-sh/uv:0.6.9 /uv /bin/uv ENV UV_COMPILE_BYTECODE=1 UV_LINK_MODE=copy UV_PYTHON_DOWNLOADS=never -# git: semif-phase1 installs from a pinned GitHub commit. +# git: semif-phase1 installs from a pinned GitHub commit. The fast extra also needs a C +# compiler AT RUNTIME: triton builds its CUDA driver shim on first use, and without gcc the +# warm-up dies with "Failed to find C compiler" (seen 2026-09-27; startup failed closed). RUN apt-get update && apt-get install -y --no-install-recommends git ca-certificates \ + && if echo "$EXTRAS" | grep -q -- '--extra fast'; then \ + apt-get install -y --no-install-recommends gcc libc6-dev; fi \ && rm -rf /var/lib/apt/lists/* WORKDIR /app COPY --from=deps /deps/pyproject.toml /deps/uv.lock ./ RUN --mount=type=cache,target=/root/.cache/uv \ - uv sync --frozen --no-dev --extra model --no-install-project + uv sync --frozen --no-dev $EXTRAS --no-install-project COPY pyproject.toml uv.lock ./ COPY src ./src -RUN uv sync --frozen --no-dev --extra model --no-editable --no-cache +RUN uv sync --frozen --no-dev $EXTRAS --no-editable --no-cache RUN groupadd --system --gid 10001 semif \ && useradd --system --uid 10001 --gid 10001 --no-create-home --shell /usr/sbin/nologin semif USER semif @@ -41,6 +49,7 @@ ENV PATH=/app/.venv/bin:$PATH \ HF_HOME=/hf \ HF_HUB_OFFLINE=1 \ HF_HUB_DISABLE_TELEMETRY=1 \ + TRITON_CACHE_DIR=/tmp/triton-cache \ NVIDIA_DRIVER_CAPABILITIES=compute,utility EXPOSE 8000 # One worker (INV-2): the model and the inference lock live in this one process. diff --git a/services/semif-serve/acceptance/averaging-2026-09-27-v0.1.3.json b/services/semif-serve/acceptance/averaging-2026-09-27-v0.1.3.json new file mode 100644 index 0000000..00a342b --- /dev/null +++ b/services/semif-serve/acceptance/averaging-2026-09-27-v0.1.3.json @@ -0,0 +1,18 @@ +{ + "rows": 252, + "groups": 72, + "accuracy_plain": 0.7857, + "accuracy_rotations": 0.881, + "delta_95ci": [ + 0.0512, + 0.1429 + ], + "unanimous": { + "rows": 163, + "accuracy": 0.9448 + }, + "split": { + "rows": 89, + "accuracy": 0.764 + } +} \ No newline at end of file diff --git a/services/semif-serve/acceptance/averaging.py b/services/semif-serve/acceptance/averaging.py new file mode 100644 index 0000000..32f34c2 --- /dev/null +++ b/services/semif-serve/acceptance/averaging.py @@ -0,0 +1,42 @@ +"""Order-averaging acceptance through the SERVICE (0.1.3): the same 252 labelled rows as the spike +(SemIf authored144 + perturbations108), each scored twice: plain /decide (the caller's order), and +/decide with orderings=rotations (combined.top). Paired group bootstrap for the accuracy delta. + SEMIF_DIR=... SEMIF_URL=... SEMIF_TOKEN=... uv run --with httpx python averaging.py out.json +""" +import json, os, random, statistics as st, sys +from pathlib import Path +import httpx + +S, U = Path(os.environ["SEMIF_DIR"]), os.environ["SEMIF_URL"] +H = {"Authorization": f"Bearer {os.environ['SEMIF_TOKEN']}"} +rows = [] +for name in ("authored144", "perturbations108"): + rows += [dict(json.loads(l), set=name) for l in (S / f"benchmarks/data/{name}.jsonl").read_text().splitlines() if l.strip()] +out = [] +with httpx.Client(timeout=120) as c: + for r in rows: + base = {k: r[k] for k in ("id", "state", "question", "options")} + plain = c.post(f"{U}/decide", headers=H, json=base).json() + avg = c.post(f"{U}/decide", headers=H, json={**base, "orderings": "rotations"}).json() + gold = r["options"][r["label"]]["id"] + top_plain = plain["option_ids"][plain["probabilities"].index(max(plain["probabilities"]))] + out.append({"group": r["group_id"], "plain": top_plain == gold, "rotations": avg["combined"]["top"] == gold, + "agreement": avg["combined"]["agreement"]}) +groups = {} +for o in out: + groups.setdefault(o["group"], []).append(o) +rng, keys, deltas = random.Random(7), list(groups), [] +for _ in range(10000): + sample = [o for g in (rng.choice(keys) for _ in keys) for o in groups[g]] + deltas.append(st.fmean(o["rotations"] for o in sample) - st.fmean(o["plain"] for o in sample)) +deltas.sort() +unan = [o for o in out if o["agreement"] == 1.0] +split = [o for o in out if o["agreement"] < 1.0] +report = {"rows": len(out), "groups": len(groups), + "accuracy_plain": round(st.fmean(o["plain"] for o in out), 4), + "accuracy_rotations": round(st.fmean(o["rotations"] for o in out), 4), + "delta_95ci": [round(deltas[250], 4), round(deltas[9750], 4)], + "unanimous": {"rows": len(unan), "accuracy": round(st.fmean(o["rotations"] for o in unan), 4)}, + "split": {"rows": len(split), "accuracy": round(st.fmean(o["rotations"] for o in split), 4) if split else None}} +json.dump(report, open(sys.argv[1], "w"), indent=1) +print(json.dumps(report)) diff --git a/services/semif-serve/acceptance/result-2026-09-27-v0.1.3.json b/services/semif-serve/acceptance/result-2026-09-27-v0.1.3.json new file mode 100644 index 0000000..d558874 --- /dev/null +++ b/services/semif-serve/acceptance/result-2026-09-27-v0.1.3.json @@ -0,0 +1,63 @@ +{ + "url": "http://10.251.50.54:8032", + "health": { + "status": "ok", + "semif_commit": "23cf1f39fc9534fe81437200959b6dfc7106e45a", + "model": { + "source": "Qwen/Qwen3.5-4B", + "revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a", + "dtype": "bfloat16", + "device": "cuda:0", + "torch_version": "2.10.0+cu128", + "transformers_version": "5.17.0", + "device_name": "NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition", + "allocated_gib": 7.84, + "reserved_gib": 8.12 + }, + "vram_cap_gib": 12.0, + "max_tokens": 4096, + "max_decisions": 64, + "workloads": [] + }, + "1_parity_vs_upstream": { + "rows": 144, + "top_choice_agree": 144, + "max_abs_prob_gap": 0.049096847865749194 + }, + "1_prompt_sha256_equal": 144, + "2_noise_floor_a_vs_b": { + "rows": 144, + "top_choice_agree": 144, + "max_abs_prob_gap": 0.0 + }, + "3_negative_rotated_options": { + "rows": 144, + "top_choice_agree": 14, + "max_abs_prob_gap": 0.9985836128543327 + }, + "4_shared_vs_direct": { + "groups": 36, + "rows": 72, + "top_choice_agree": 72, + "max_abs_prob_gap": 0.046346781311727814 + }, + "5_speed_21_binary": { + "prefix_tokens": 62, + "shared_s": { + "runs": [ + 0.139055563005968, + 0.1375489159981953, + 0.13897942200128455 + ], + "median": 0.13897942200128455 + }, + "sequential_decide_s": { + "runs": [ + 0.9357447380025405, + 0.9384197259932989, + 0.939685705001466 + ], + "median": 0.9384197259932989 + } + } +} \ No newline at end of file diff --git a/services/semif-serve/pyproject.toml b/services/semif-serve/pyproject.toml index ad245af..7bab3e5 100644 --- a/services/semif-serve/pyproject.toml +++ b/services/semif-serve/pyproject.toml @@ -1,6 +1,6 @@ [project] name = "semif-serve" -version = "0.1.2" +version = "0.1.3" description = "HTTP wrapper around SemIf's direct and shared option-logit scorers" requires-python = ">=3.12" dependencies = [ @@ -17,6 +17,13 @@ model = [ "torch==2.10.0", ] +# Qwen3.5's fast kernels. Without them transformers runs its reference PyTorch paths +# ("correct but much slower"). Trialled 2026-09-27; adopted only if parity with upstream holds. +fast = [ + "flash-linear-attention==0.5.2", + "causal-conv1d @ https://github.com/Dao-AILab/causal-conv1d/releases/download/v1.7.0/causal_conv1d-1.7.0+cu12torch2.10cxx11abiTRUE-cp312-cp312-linux_x86_64.whl ; sys_platform == 'linux' and platform_machine == 'x86_64'", +] + [dependency-groups] dev = ["pytest==8.4.2", "httpx==0.28.1"] diff --git a/services/semif-serve/semif-serve.contract.md b/services/semif-serve/semif-serve.contract.md index 7488d75..f4b365f 100644 --- a/services/semif-serve/semif-serve.contract.md +++ b/services/semif-serve/semif-serve.contract.md @@ -45,6 +45,37 @@ probabilities}` with `softmax(option_logits / T)`. The argmax never changes. The native `probabilities` and `probability_status` stay untouched. An unknown `workload` is a 422. With no `workload`, no `calibrated` key appears. +**Order averaging (0.1.3, Prime 2026-09-27).** A decision (the `/decide` body, or an +entry in `decisions`) may set `orderings`, whose default is `"none"`: + +- `"rotations"`: the n cyclic shifts of the caller's option list, starting with + the caller's order, so every option sits in every position exactly once. +- `"all"`: every permutation (n!), the caller's order first. It is allowed only + when n ≤ 4, else 422. + +Every ordering of every decision in the request becomes its own SemIf row +(`id` = `"#o"`, same state and question, reordered options). **All rows go +to the engine in ONE `shared` call**, which includes `/decide`. The result for an +averaged decision is: + +``` +{id, option_ids (caller's order), + combined: {method, orderings: n, probabilities, top, agreement, spread: {option_id: [min, max]}}, + orderings: [the SemIf result for each ordering, unchanged]} +``` + +- `probabilities`: per ordering, log-softmax of `option_logits`; averaged per + option id; renormalised; reported in the caller's order. +- `agreement`: the fraction of orderings whose top option equals `combined.top`. +- `spread`: each option's min and max native probability across orderings. + +Every expanded row counts toward `SEMIF_MAX_DECISIONS`. `workload` together with +`orderings` is a 422, because a temperature is fitted per method and none is +fitted on combined scores yet. Decisions without `orderings` keep the exact +pre-0.1.3 result shape. Averaging cancels any additive position bias exactly. +Measured by the 2026-09-27 spike: 3 rotations take SemIf's labelled sets from +78.6% to 87.7% accuracy, and unanimous agreement is 94.4% accurate. + ## Invariants - **INV-1 pass-through.** `option_ids`, `probabilities`, `option_logits`, @@ -64,10 +95,18 @@ native `probabilities` and `probability_status` stay untouched. An unknown baseline by more than 512 MiB, the engine calls `empty_cache()`. A burst must not keep the card's shared headroom: on 2026-09-27 a 64-decision request left the process holding 12.6 GB, leaving scriberr 3.5 GB. + **Any other scorer failure except `ValueError`** also releases before it is + reported: it is logged with its traceback, then re-raised unchained as + `ScoringFailed` after `gc.collect()` + `empty_cache()` (bug hunt C5). A + `RuntimeError` whose message says "out of memory" (cuBLAS/cuDNN allocation + failures) counts as an OOM → 503 (S9). An OOM with an empty message is reported + as "CUDA out of memory" rather than crashing the handler (C4). - **INV-5 no network at runtime.** Weights come from the mounted HF cache at the - pinned revision (`HF_HUB_OFFLINE=1`). + pinned revision. The entry point sets `HF_HUB_OFFLINE=1` itself before torch or + transformers load, so this holds outside the image too (S10). - **INV-6 constant-time auth.** Token comparison uses `hmac.compare_digest`. The - token is ≥ 32 characters, and startup refuses a shorter one. + token is ≥ 32 characters of visible ASCII (33–126). Startup refuses anything + else, because a CR, LF or NUL in the token can never arrive in a header (S2). ## Limits and errors @@ -75,15 +114,21 @@ native `probabilities` and `probability_status` stay untouched. An unknown 422, never truncated (SemIf raises). - `SEMIF_MAX_DECISIONS` (default 64) caps `decisions` per shared request. It must hold 1..max entries, else 422. -- Request body ≤ `SEMIF_MAX_BODY_BYTES` (default 1 MiB), else 413. +- Request body ≤ `SEMIF_MAX_BODY_BYTES` (default 1 MiB), else 413. The limit is + checked **before** each chunk is kept, so no more than the limit is ever held + (C2). A declared `Content-Length` is trusted only if it is ASCII digits (S3). +- **Admission:** at most `SEMIF_MAX_QUEUE` (default 32) POSTs may be in progress, + counting both queued and scoring. The next one is refused with 429 `busy` + before its body is read (C6). | status | code | when | |---|---|---| | 401 | `unauthorized` | missing or wrong bearer | | 413 | `request_too_large` | body over the limit | | 422 | `invalid_request` | bad JSON shape, a SemIf `ValueError` (validation, token limit, tokenisation), unknown workload, too many decisions | +| 429 | `busy` | `SEMIF_MAX_QUEUE` requests already in progress | | 503 | `out_of_memory` | CUDA OOM during scoring | -| 500 | `scoring_failed` | any other scorer exception | +| 500 | `scoring_failed` | any other scorer exception, **or a failure while building the response** from a scorer result (calibration, averaging): always the envelope, never a bare 500 (C3) | The error body is `{error: {code, message}}`. @@ -92,7 +137,16 @@ The error body is `{error: {code, message}}`. `SEMIF_API_TOKEN` (required), `SEMIF_MODEL` (default `Qwen/Qwen3.5-4B`), `SEMIF_REVISION` (default the pinned SHA), `SEMIF_DEVICE` (default `cuda`), `SEMIF_VRAM_CAP_GIB`, `SEMIF_MAX_TOKENS`, `SEMIF_MAX_DECISIONS`, -`SEMIF_MAX_BODY_BYTES`, `SEMIF_CALIBRATION` (path to a JSON `{workload: T}`; T > 0). +`SEMIF_MAX_BODY_BYTES`, `SEMIF_MAX_QUEUE`, `SEMIF_CALIBRATION` (path to a JSON +`{workload: T}`). + +Startup validates every value and refuses a bad one with a `ValueError` naming +the variable (C1, S1, S8): +- the VRAM cap, when set, is finite and > 0 (`0` used to mean uncapped); +- every limit is an integer ≥ 1; +- each T is a finite number, not a bool, in [0.05, 20] (a tiny T overflowed to + NaN, and the response then failed to render); +- the calibration file must exist and parse. ## Tests (TDD, fake scorer: no torch, no model) @@ -109,7 +163,13 @@ alone; any other exception → 500; malformed JSON or a wrong body shape → 422; too many decisions → 422; an oversized body → 413; requests are serialised (two concurrent calls never overlap inside the scorer); `/health` -answers while a scorer call is blocked. +answers while a scorer call is blocked. **Averaging:** `rotations` sends n rows +in one shared call, each option once per position, with ids `#o`, and +cancels a position bias exactly; `all` sends n! rows and is 422 above 4 options; +agreement and spread are computed from the orderings; a mixed shared request +(averaged + plain) is one engine call, with results in request order and plain +results unchanged; expanded rows count toward the cap; `workload` + `orderings` +→ 422. ## Acceptance (on fv-ml1, real model; not unit tests) diff --git a/services/semif-serve/spike/latency-2026-09-27-v0.1.3.txt b/services/semif-serve/spike/latency-2026-09-27-v0.1.3.txt new file mode 100644 index 0000000..597e5c3 --- /dev/null +++ b/services/semif-serve/spike/latency-2026-09-27-v0.1.3.txt @@ -0,0 +1,7 @@ +condition e2e p50 (runs) server p50 (runs) +health (floor) 33.5 [31.7-34.7] - +decide, short (~130 tok) 69.3 [68.8-70.1] 35.3 [35.3-35.4] +decide, long (~2,000 tok) 130.5 [129.5-130.5] 92.1 [91.7-92.4] +shared, 3 rotations (short) 115.0 [112.4-115.7] 77.9 [77.1-79.3] +shared, 6 orderings (short) 118.3 [117.5-118.7] 81.2 [80.2-81.3] +shared, 3 rotations (long) 199.5 [196.6-201.0] 158.2 [157.7-158.3] diff --git a/services/semif-serve/src/semif_serve/app.py b/services/semif-serve/src/semif_serve/app.py index 26e543b..d906a94 100644 --- a/services/semif-serve/src/semif_serve/app.py +++ b/services/semif-serve/src/semif_serve/app.py @@ -2,9 +2,10 @@ from __future__ import annotations import hmac +import itertools import math import threading -from typing import Any +from typing import Any, Literal from fastapi import FastAPI, Request from fastapi.concurrency import run_in_threadpool @@ -29,6 +30,7 @@ class Decision(BaseModel): id: str question: str options: list[Option] + orderings: Literal["none", "rotations", "all"] = "none" def row(self, state: State) -> dict: """The SemIf row shape: exactly id, state, question, options.""" @@ -64,6 +66,75 @@ def _first_error(exc: ValidationError) -> str: return f"{where}: {first.get('msg', 'invalid')}" +MAX_OPTIONS_FOR_ALL = 4 + + +def ordering_perms(decision: Decision) -> list[tuple[int, ...]]: + """Index permutations of the caller's options, the caller's own order first.""" + n = len(decision.options) + if decision.orderings == "rotations": + return [tuple((start + k) % n for k in range(n)) for start in range(n)] + if n > MAX_OPTIONS_FOR_ALL: + raise ApiError(422, "invalid_request", + f"orderings 'all' allows at most {MAX_OPTIONS_FOR_ALL} options ({n} given); use 'rotations'") + return list(itertools.permutations(range(n))) + + +def expanded_rows(decision: Decision, state: State, perms: list[tuple[int, ...]]) -> list[dict]: + base = decision.row(state) + return [{**base, "id": f"{decision.id}#o{k}", "options": [base["options"][i] for i in perm]} + for k, perm in enumerate(perms)] + + +def combine(decision: Decision, results: list[dict]) -> dict: + """Average per-ordering log-softmax by option id; the native results ride along unchanged.""" + option_ids = [o.id for o in decision.options] + logp: dict[str, list[float]] = {i: [] for i in option_ids} + probs: dict[str, list[float]] = {i: [] for i in option_ids} + tops = [] + for result in results: + logits = result["option_logits"] + top = max(logits) + lse = top + math.log(sum(math.exp(x - top) for x in logits)) + for oid, x, p in zip(result["option_ids"], logits, result["probabilities"]): + logp[oid].append(x - lse) + probs[oid].append(p) + tops.append(result["option_ids"][logits.index(top)]) + means = [sum(logp[i]) / len(logp[i]) for i in option_ids] + peak = max(means) + weights = [math.exp(m - peak) for m in means] + combined_p = [w / sum(weights) for w in weights] + winner = option_ids[combined_p.index(max(combined_p))] + return { + "id": decision.id, + "option_ids": option_ids, + "combined": { + "method": decision.orderings, + "orderings": len(results), + "probabilities": combined_p, + "top": winner, + "agreement": tops.count(winner) / len(tops), + "spread": {i: [min(probs[i]), max(probs[i])] for i in option_ids}, + }, + "orderings": results, + } + + +async def read_limited(stream, declared: str | None, limit: int) -> bytes: + """Read a request body, refusing it once it would exceed `limit` bytes. The check runs BEFORE a + chunk is kept, and nothing after the crossing chunk is read (bug hunt C2). A declared length is + trusted only as ASCII digits: `"²".isdigit()` is True but `int("²")` raises (S3).""" + too_large = ApiError(413, "request_too_large", f"request body exceeds {limit} bytes") + if declared is not None and declared.isascii() and declared.isdigit() and int(declared) > limit: + raise too_large + body = bytearray() + async for chunk in stream: + if len(body) + len(chunk) > limit: + raise too_large + body.extend(chunk) + return bytes(body) + + def calibrated_view(result: dict, workload: str, temperature: float) -> dict: """softmax(option_logits / T): the native fields are left exactly as SemIf returned them (INV-1).""" scaled = [x / temperature for x in result["option_logits"]] @@ -77,6 +148,27 @@ def create_app(settings: Settings, engine: Any) -> FastAPI: app = FastAPI(title="semif-serve") expected = f"Bearer {settings.api_token}".encode() inference = threading.Lock() # INV-2: one scorer call at a time, off the event loop + in_progress = 0 # POSTs admitted and not yet answered (bug hunt C6) + + def admit(): + nonlocal in_progress + if in_progress >= settings.max_queue: + raise ApiError(429, "busy", f"{in_progress} requests already in progress (limit {settings.max_queue})") + in_progress += 1 + + def leave(): + nonlocal in_progress + in_progress -= 1 + + def build(fn, *args): + """Response construction from a scorer result (calibration, averaging) maps its failures to + the 500 envelope too, instead of escaping as a bare 500 (bug hunt C3).""" + try: + return fn(*args) + except ApiError: + raise + except Exception as exc: # noqa: BLE001 + raise ApiError(500, "scoring_failed", f"building the response failed: {type(exc).__name__}: {exc}") from exc def locked(fn, *args): with inference: @@ -94,22 +186,10 @@ def create_app(settings: Settings, engine: Any) -> FastAPI: async def api_error(_request: Request, exc: ApiError): return error(exc.status, exc.code, exc.message) - async def read_limited(request: Request) -> bytes: - limit = settings.max_body_bytes - too_large = ApiError(413, "request_too_large", f"request body exceeds {limit} bytes") - declared = request.headers.get("content-length") - if declared is not None and declared.isdigit() and int(declared) > limit: - raise too_large - body = bytearray() - async for chunk in request.stream(): # also caps bodies that declare no length - body.extend(chunk) - if len(body) > limit: - raise too_large - return bytes(body) - async def parse(request: Request, model: type[BaseModel]): try: - return model.model_validate_json(await read_limited(request)) + return model.model_validate_json( + await read_limited(request.stream(), request.headers.get("content-length"), settings.max_body_bytes)) except ValidationError as exc: raise ApiError(422, "invalid_request", _first_error(exc)) from exc @@ -142,20 +222,59 @@ def create_app(settings: Settings, engine: Any) -> FastAPI: "vram_cap_gib": settings.vram_cap_gib, "max_tokens": settings.max_tokens, "max_decisions": settings.max_decisions, "workloads": sorted(settings.calibration)} + async def score_batch(decisions: list[Decision], state: State, workload: str | None) -> tuple[list[dict], dict]: + """One engine.shared call for every row of every decision; results in request order.""" + plan = [] # (decision, perms or None, row count) + rows: list[dict] = [] + for d in decisions: + if d.orderings == "none": + plan.append((d, None, 1)) + rows.append(d.row(state)) + else: + if workload is not None: + raise ApiError(422, "invalid_request", + "workload calibration is not available together with orderings") + perms = ordering_perms(d) + plan.append((d, perms, len(perms))) + rows.extend(expanded_rows(d, state, perms)) + if not 1 <= len(rows) <= settings.max_decisions: + raise ApiError(422, "invalid_request", + f"this request expands to {len(rows)} scored rows; the limit is 1..{settings.max_decisions}") + temperature = temperature_for(workload) + results, timing = await score(engine.shared, rows) + out, cursor = [], 0 + for d, perms, count in plan: + chunk = results[cursor:cursor + count] + cursor += count + out.append(build(with_calibration, chunk[0], workload, temperature) if perms is None + else build(combine, d, chunk)) + return out, timing + @app.post("/decide") async def decide(request: Request): - body = await parse(request, DecideBody) - temperature = temperature_for(body.workload) - return with_calibration(await score(engine.direct, body.row(body.state)), body.workload, temperature) + admit() + try: + body = await parse(request, DecideBody) + if body.orderings != "none": + results, _timing = await score_batch([body], body.state, body.workload) + return results[0] + temperature = temperature_for(body.workload) + result = await score(engine.direct, body.row(body.state)) + return build(with_calibration, result, body.workload, temperature) + finally: + leave() @app.post("/decide/shared") async def decide_shared(request: Request): - body = await parse(request, SharedBody) - if not 1 <= len(body.decisions) <= settings.max_decisions: - raise ApiError(422, "invalid_request", - f"decisions must hold 1..{settings.max_decisions} entries, got {len(body.decisions)}") - temperature = temperature_for(body.workload) - results, timing = await score(engine.shared, [d.row(body.state) for d in body.decisions]) - return {"results": [with_calibration(r, body.workload, temperature) for r in results], "timing": timing} + admit() + try: + body = await parse(request, SharedBody) + if not 1 <= len(body.decisions) <= settings.max_decisions: + raise ApiError(422, "invalid_request", + f"decisions must hold 1..{settings.max_decisions} entries, got {len(body.decisions)}") + results, timing = await score_batch(body.decisions, body.state, body.workload) + return {"results": results, "timing": timing} + finally: + leave() return app diff --git a/services/semif-serve/src/semif_serve/config.py b/services/semif-serve/src/semif_serve/config.py index 3cdb295..f0ac50e 100644 --- a/services/semif-serve/src/semif_serve/config.py +++ b/services/semif-serve/src/semif_serve/config.py @@ -1,4 +1,9 @@ -"""Settings for semif-serve. Contract: semif-serve.contract.md § Configuration.""" +"""Settings for semif-serve. Contract: semif-serve.contract.md § Configuration. + +Every value is validated at startup and a bad one is refused with a ValueError naming the +variable: a service that starts and then rejects every request (or runs uncapped) is worse +than one that does not start (bug hunt 2026-09-27: C1, S1, S2, S8). +""" from __future__ import annotations import json @@ -8,6 +13,7 @@ from dataclasses import dataclass, field from pathlib import Path MIN_TOKEN_CHARS = 32 +MIN_TEMPERATURE, MAX_TEMPERATURE = 0.05, 20.0 SEMIF_COMMIT = "23cf1f39fc9534fe81437200959b6dfc7106e45a" DEFAULT_MODEL = "Qwen/Qwen3.5-4B" DEFAULT_REVISION = "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a" @@ -23,35 +29,71 @@ class Settings: max_tokens: int = 4096 max_decisions: int = 64 max_body_bytes: int = 1024 * 1024 + max_queue: int = 32 calibration: dict[str, float] = field(default_factory=dict) @classmethod def from_env(cls, env: Mapping[str, str]) -> "Settings": token = env.get("SEMIF_API_TOKEN", "") - if len(token) < MIN_TOKEN_CHARS: # INV-6 - raise ValueError(f"SEMIF_API_TOKEN must be at least {MIN_TOKEN_CHARS} characters") - cap = env.get("SEMIF_VRAM_CAP_GIB") + # INV-6: visible ASCII only. A CR, LF or NUL can never arrive in a header, so a token + # carrying one would lock every caller out while /health still said ok. + if len(token) < MIN_TOKEN_CHARS or not all(33 <= ord(c) <= 126 for c in token): + raise ValueError(f"SEMIF_API_TOKEN must be at least {MIN_TOKEN_CHARS} visible ASCII characters") return cls( api_token=token, model=env.get("SEMIF_MODEL", DEFAULT_MODEL), revision=env.get("SEMIF_REVISION", DEFAULT_REVISION), device=env.get("SEMIF_DEVICE", "cuda"), - vram_cap_gib=float(cap) if cap else None, - max_tokens=int(env.get("SEMIF_MAX_TOKENS", 4096)), - max_decisions=int(env.get("SEMIF_MAX_DECISIONS", 64)), - max_body_bytes=int(env.get("SEMIF_MAX_BODY_BYTES", 1024 * 1024)), + vram_cap_gib=_positive_float(env, "SEMIF_VRAM_CAP_GIB"), + max_tokens=_positive_int(env, "SEMIF_MAX_TOKENS", 4096), + max_decisions=_positive_int(env, "SEMIF_MAX_DECISIONS", 64), + max_body_bytes=_positive_int(env, "SEMIF_MAX_BODY_BYTES", 1024 * 1024), + max_queue=_positive_int(env, "SEMIF_MAX_QUEUE", 32), calibration=_load_calibration(env.get("SEMIF_CALIBRATION")), ) +def _positive_int(env: Mapping[str, str], name: str, default: int) -> int: + raw = env.get(name) + if raw is None: + return default + try: + value = int(raw) + except ValueError: + raise ValueError(f"{name} must be an integer, got {raw!r}") from None + if value < 1: + raise ValueError(f"{name} must be >= 1, got {value}") + return value + + +def _positive_float(env: Mapping[str, str], name: str) -> float | None: + """Unset means no cap. When set it must be finite and > 0: `0` used to slip through as 'no cap'.""" + raw = env.get(name) + if raw is None or raw == "": + return None + try: + value = float(raw) + except ValueError: + raise ValueError(f"{name} must be a number, got {raw!r}") from None + if not math.isfinite(value) or value <= 0: + raise ValueError(f"{name} must be a finite number > 0, got {raw!r}") + return value + + def _load_calibration(path: str | None) -> dict[str, float]: - """{workload: T}, every T a finite number > 0 (T scales option logits before softmax).""" + """{workload: T}; T scales option logits before softmax, so it is kept in a sane range + (a tiny T overflows to NaN and the response then fails to render).""" if not path: return {} - table = json.loads(Path(path).read_text()) + try: + table = json.loads(Path(path).read_text()) + except (OSError, ValueError) as exc: + raise ValueError(f"SEMIF_CALIBRATION {path!r} could not be read as JSON: {exc}") from None if not isinstance(table, dict) or not all( - isinstance(t, (int, float)) and not isinstance(t, bool) and math.isfinite(t) and t > 0 + isinstance(t, (int, float)) and not isinstance(t, bool) and math.isfinite(t) + and MIN_TEMPERATURE <= t <= MAX_TEMPERATURE for t in table.values() ): - raise ValueError("SEMIF_CALIBRATION must be a JSON object of workload -> finite temperature > 0") + raise ValueError(f"SEMIF_CALIBRATION must be a JSON object of workload -> temperature in " + f"[{MIN_TEMPERATURE}, {MAX_TEMPERATURE}]") return {str(k): float(v) for k, v in table.items()} diff --git a/services/semif-serve/src/semif_serve/engine.py b/services/semif-serve/src/semif_serve/engine.py index 5894585..c75eda6 100644 --- a/services/semif-serve/src/semif_serve/engine.py +++ b/services/semif-serve/src/semif_serve/engine.py @@ -7,10 +7,14 @@ unit-tested against a fake torch (tests/test_engine.py). from __future__ import annotations import gc +import logging +import traceback from typing import Any, Callable from .config import Settings -from .errors import OutOfMemory +from .errors import OutOfMemory, ScoringFailed + +log = logging.getLogger("semif_serve.engine") RELEASE_SLACK_BYTES = 512 * 2**20 WARMUP_ROW = { @@ -24,6 +28,11 @@ WARMUP_ROW = { } +def _first_line(exc: BaseException) -> str: + lines = str(exc).splitlines() + return lines[0] if lines else "" + + class TorchEngine: def __init__(self, torch: Any, model: Any, tokenizer: Any, metadata: dict, settings: Settings, direct_fn: Callable, shared_fn: Callable, release_above_bytes: int | None = None): @@ -46,7 +55,7 @@ class TorchEngine: arch = f"sm_{major}{minor}" if arch not in torch.cuda.get_arch_list(): # INV-3: no silent PTX/CPU fallback raise RuntimeError(f"torch {torch.__version__} has no kernels for {arch}: {torch.cuda.get_arch_list()}") - if settings.vram_cap_gib: # INV-4: cap BEFORE the weights land + if settings.vram_cap_gib is not None: # INV-4: cap BEFORE the weights land total = torch.cuda.get_device_properties(0).total_memory fraction = settings.vram_cap_gib * 2**30 / total if not 0 < fraction <= 1: @@ -82,17 +91,28 @@ class TorchEngine: def _guard(self, fn, *args): try: result = fn(*args) + except ValueError: + raise # validation: SemIf raises it before any GPU work except self._torch.cuda.OutOfMemoryError as exc: - message = str(exc).splitlines()[0] + failure, message = OutOfMemory, _first_line(exc) or "CUDA out of memory" + except Exception as exc: # noqa: BLE001 — every other failure is released and reported below + message = _first_line(exc) + if "out of memory" in message.lower(): # cuBLAS/cuDNN allocation failures + failure = OutOfMemory + else: + failure, message = ScoringFailed, f"{type(exc).__name__}: {message}" + # Formatted text, not exc_info: a log record that keeps the traceback object alive + # (pytest's capture handler does; so would any buffering handler) pins the tensors. + log.error("scorer failed:\n%s", traceback.format_exc()) else: self._release_burst() return result - # INV-4, outside the except block on purpose: the torch exception's traceback holds the - # failed scorer's frames, and with them its tensors (the replicated prefix cache). Raising - # inside the block, or `from exc`, would chain to it and keep GiBs allocated after the 503. + # INV-4, outside the except block on purpose: the exception's traceback holds the failed + # scorer's frames, and with them its tensors (the replicated prefix cache). Raising inside + # the block, or `from exc`, would chain to it and keep GiBs allocated after the response. gc.collect() self._torch.cuda.empty_cache() - raise OutOfMemory(message) + raise failure(message) def direct(self, row: dict) -> dict: return self._guard(self._direct, self._model, self._tokenizer, row, self._metadata, self._settings.max_tokens) diff --git a/services/semif-serve/src/semif_serve/errors.py b/services/semif-serve/src/semif_serve/errors.py index 3a36479..6d2d853 100644 --- a/services/semif-serve/src/semif_serve/errors.py +++ b/services/semif-serve/src/semif_serve/errors.py @@ -3,3 +3,8 @@ class OutOfMemory(RuntimeError): """The engine ran out of GPU memory during a request and has already released its cache (INV-4).""" + + +class ScoringFailed(RuntimeError): + """A scorer call failed for a reason other than validation or OOM. Raised unchained, after the + failed call's memory has been released; the original traceback is logged, not carried (INV-4).""" diff --git a/services/semif-serve/src/semif_serve/main.py b/services/semif-serve/src/semif_serve/main.py index e5cf351..c3391bf 100644 --- a/services/semif-serve/src/semif_serve/main.py +++ b/services/semif-serve/src/semif_serve/main.py @@ -11,6 +11,10 @@ from .config import Settings def app_from_env() -> FastAPI: settings = Settings.from_env(os.environ) + # INV-5: never download at runtime, inside the image or out of it (bug hunt S10). Set before + # torch / transformers / huggingface_hub are imported, since they read it at import time. + os.environ["HF_HUB_OFFLINE"] = "1" + os.environ["TRANSFORMERS_OFFLINE"] = "1" from .engine import TorchEngine # torch loads only here, never in the unit tests return create_app(settings, TorchEngine.load(settings)) diff --git a/services/semif-serve/tests/test_app.py b/services/semif-serve/tests/test_app.py index c55b57c..e7a1515 100644 --- a/services/semif-serve/tests/test_app.py +++ b/services/semif-serve/tests/test_app.py @@ -1,5 +1,7 @@ """semif-serve HTTP behaviour against a fake engine (no torch, no model). Contract: services/semif-serve/semif-serve.contract.md""" +import json + import pytest from fastapi.testclient import TestClient @@ -231,3 +233,104 @@ def test_health_reports_the_pins_limits_and_workloads(): body = client.get("/health").json() assert body == {"status": "ok", "semif_commit": SEMIF_COMMIT, "model": FakeEngine().health(), "vram_cap_gib": 12.0, "max_tokens": 4096, "max_decisions": 8, "workloads": ["alerts", "triage"]} + + +def test_a_body_of_exactly_the_limit_is_accepted(): + body = json.dumps(ROW).encode() + client = make_client(max_body_bytes=len(body)) + assert client.post("/decide", content=body, headers={**AUTH, "content-type": "application/json"}).status_code == 200 + + +def test_a_non_ascii_digit_content_length_is_ignored_not_a_crash(): + """HTTP clients cannot send one (httpx refuses; h11 rejects it), so check the reader directly.""" + import asyncio + from semif_serve.app import read_limited + + async def body(): + yield b"{}" + + assert asyncio.run(read_limited(body(), "²", 100)) == b"{}" # int("²") would raise + + +def test_exactly_max_decisions_is_accepted(): + decisions = [{"id": str(i), "question": "Q?", "options": OPTIONS} for i in range(3)] + client = make_client(max_decisions=3) + assert client.post("/decide/shared", json={"state": "s", "decisions": decisions}, headers=AUTH).status_code == 200 + + +def test_calibration_leaves_every_native_field_alone(): + engine = FakeEngine(logits=(3.0, 1.0)) + body = make_client(engine, calibration={"triage": 2.0}).post( + "/decide", json={**ROW, "workload": "triage"}, headers=AUTH).json() + body.pop("calibrated") + assert body == engine.direct(ROW) + + +class MalformedEngine(FakeEngine): + def direct(self, row): + return {"id": row["id"], "option_ids": ["yes", "no"], "probabilities": [0.5, 0.5]} # no option_logits + + def shared(self, rows): + return [self.direct(r) for r in rows], {} + + +@pytest.mark.parametrize("path, body", [ + ("/decide", {**ROW, "workload": "triage"}), + ("/decide", {**ROW, "orderings": "rotations"}), +]) +def test_a_malformed_scorer_result_is_an_envelope_500_not_a_bare_one(path, body): + client = make_client(MalformedEngine(), calibration={"triage": 2.0}) + response = client.post(path, json=body, headers=AUTH) + assert response.status_code == 500 + assert response.json()["error"]["code"] == "scoring_failed" + + +class SharedSlowEngine(SlowEngine): + def shared(self, rows): + return [self.direct(r) for r in rows], {} + + +def test_shared_requests_are_serialised_too(): + from concurrent.futures import ThreadPoolExecutor + engine = SharedSlowEngine() + body = {"state": "s", "decisions": [{"id": "a", "question": "Q?", "options": OPTIONS}]} + with make_client(engine) as client, ThreadPoolExecutor(3) as pool: + futures = [pool.submit(client.post, "/decide/shared", json=body, headers=AUTH) for _ in range(3)] + assert engine.entered.wait(5) + engine.release.set() + assert [f.result().status_code for f in futures] == [200] * 3 + assert engine.peak == 1 + + +def test_more_than_max_queue_requests_in_progress_get_429_busy(): + from concurrent.futures import ThreadPoolExecutor + engine = SlowEngine() + with make_client(engine, max_queue=2) as client, ThreadPoolExecutor(3) as pool: + held = [pool.submit(client.post, "/decide", json={**ROW, "id": f"r{i}"}, headers=AUTH) for i in range(2)] + assert engine.entered.wait(5) + import time + deadline = time.monotonic() + 5 + while engine.inside + 0 < 1 and time.monotonic() < deadline: + time.sleep(0.01) + time.sleep(0.2) # let the second request reach the queue + extra = client.post("/decide", json={**ROW, "id": "extra"}, headers=AUTH) + assert extra.status_code == 429 and extra.json()["error"]["code"] == "busy" + engine.release.set() + assert [f.result().status_code for f in held] == [200, 200] + assert make_client(FakeEngine(), max_queue=2).post("/decide", json=ROW, headers=AUTH).status_code == 200 + + +def test_read_limited_stops_reading_at_the_crossing_chunk(): + import asyncio + from semif_serve.app import ApiError, read_limited + consumed = [] + + async def chunks(): + for i in range(10): + consumed.append(i) + yield b"x" * 100 + + with pytest.raises(ApiError) as info: + asyncio.run(read_limited(chunks(), None, 250)) + assert info.value.status == 413 + assert consumed == [0, 1, 2] # the third chunk crosses 250 and is never kept; nothing after is read diff --git a/services/semif-serve/tests/test_config.py b/services/semif-serve/tests/test_config.py index 00352fb..dd461a1 100644 --- a/services/semif-serve/tests/test_config.py +++ b/services/semif-serve/tests/test_config.py @@ -34,3 +34,45 @@ def test_a_calibration_table_needs_positive_finite_numbers(tmp_path, table): cal.write_text(json.dumps(table)) with pytest.raises(ValueError, match="SEMIF_CALIBRATION"): Settings.from_env({"SEMIF_API_TOKEN": TOKEN, "SEMIF_CALIBRATION": str(cal)}) + + +@pytest.mark.parametrize("var, value", [ + ("SEMIF_VRAM_CAP_GIB", "0"), ("SEMIF_VRAM_CAP_GIB", "-4"), ("SEMIF_VRAM_CAP_GIB", "nan"), ("SEMIF_VRAM_CAP_GIB", "inf"), + ("SEMIF_MAX_TOKENS", "0"), ("SEMIF_MAX_DECISIONS", "-1"), ("SEMIF_MAX_BODY_BYTES", "0"), ("SEMIF_MAX_QUEUE", "0"), + ("SEMIF_MAX_TOKENS", "lots"), ("SEMIF_VRAM_CAP_GIB", "twelve"), +]) +def test_out_of_range_or_unparseable_values_are_refused_naming_the_variable(var, value): + with pytest.raises(ValueError, match=var): + Settings.from_env({"SEMIF_API_TOKEN": TOKEN, var: value}) + + +@pytest.mark.parametrize("token", ["x" * 31 + "\n", "x" * 30 + "\r\n", "x" * 32 + "\x00", "x" * 16 + " " + "x" * 16, + "é" * 32]) +def test_a_token_with_non_visible_ascii_is_refused(token): + with pytest.raises(ValueError, match="SEMIF_API_TOKEN"): + Settings.from_env({"SEMIF_API_TOKEN": token}) + + +@pytest.mark.parametrize("table", [{"w": True}, {"w": 1e-300}, {"w": 0.04}, {"w": 21}]) +def test_a_temperature_must_be_a_real_number_in_range(tmp_path, table): + cal = tmp_path / "cal.json" + cal.write_text(json.dumps(table)) + with pytest.raises(ValueError, match="SEMIF_CALIBRATION"): + Settings.from_env({"SEMIF_API_TOKEN": TOKEN, "SEMIF_CALIBRATION": str(cal)}) + + +@pytest.mark.parametrize("content", [None, "{not json"]) +def test_a_missing_or_malformed_calibration_file_is_a_named_startup_error(tmp_path, content): + cal = tmp_path / "cal.json" + if content is not None: + cal.write_text(content) + with pytest.raises(ValueError, match="SEMIF_CALIBRATION"): + Settings.from_env({"SEMIF_API_TOKEN": TOKEN, "SEMIF_CALIBRATION": str(cal)}) + + +def test_in_range_edges_are_accepted(tmp_path): + cal = tmp_path / "cal.json" + cal.write_text(json.dumps({"lo": 0.05, "hi": 20})) + s = Settings.from_env({"SEMIF_API_TOKEN": "!" + "~" * 31, "SEMIF_VRAM_CAP_GIB": "0.5", "SEMIF_MAX_QUEUE": "1", + "SEMIF_CALIBRATION": str(cal)}) + assert (s.vram_cap_gib, s.max_queue, s.calibration) == (0.5, 1, {"lo": 0.05, "hi": 20.0}) diff --git a/services/semif-serve/tests/test_engine.py b/services/semif-serve/tests/test_engine.py index 8fabc8a..364b9fe 100644 --- a/services/semif-serve/tests/test_engine.py +++ b/services/semif-serve/tests/test_engine.py @@ -7,7 +7,7 @@ import pytest from semif_serve.config import Settings from semif_serve.engine import TorchEngine -from semif_serve.errors import OutOfMemory +from semif_serve.errors import OutOfMemory, ScoringFailed class FakeTorch: @@ -81,3 +81,33 @@ def test_a_burst_is_returned_to_the_driver_after_the_call(reserved_after, releas direct_fn=scorer, shared_fn=scorer, release_above_bytes=8 * 2**30 + 512 * 2**20) assert engine.direct({}) == {"ok": True} assert ReservingTorch.cuda.emptied == released + + +def cyclic_tensor(): + """A tensor held in a reference cycle, as real frames and tensors often are: only gc frees it.""" + t = Tensor() + t.self_ref = t + FakeTorch.cuda.watched.append(weakref.ref(t)) + return t + + +@pytest.mark.parametrize("raised, expected_type, expected_message", [ + (lambda: FakeTorch.cuda.OutOfMemoryError(""), OutOfMemory, "CUDA out of memory"), + (lambda: RuntimeError("CUBLAS_STATUS_ALLOC_FAILED: CUDA error: out of memory"), OutOfMemory, + "CUBLAS_STATUS_ALLOC_FAILED: CUDA error: out of memory"), + (lambda: RuntimeError("Invalid native prefix cache"), ScoringFailed, "RuntimeError: Invalid native prefix cache"), + (lambda: KeyError("option_logits"), ScoringFailed, "KeyError: 'option_logits'"), +]) +def test_every_non_validation_failure_is_released_unchained_after_gc(raised, expected_type, expected_message): + FakeTorch.cuda.empties.clear(), FakeTorch.cuda.watched.clear() + + def scorer(*_args): + kv_cache = cyclic_tensor() # noqa: F841 — alive in this frame when it raises + raise raised() + + engine = TorchEngine(FakeTorch, None, None, {}, Settings(api_token="t" * 40), direct_fn=scorer, shared_fn=scorer) + with pytest.raises(expected_type) as info: + engine.shared([]) + assert str(info.value) == expected_message + assert info.value.__cause__ is None and info.value.__context__ is None + assert FakeTorch.cuda.empties == [True] # gc freed the cycle BEFORE the cache was emptied diff --git a/services/semif-serve/tests/test_engine_load.py b/services/semif-serve/tests/test_engine_load.py new file mode 100644 index 0000000..9143270 --- /dev/null +++ b/services/semif-serve/tests/test_engine_load.py @@ -0,0 +1,110 @@ +"""TorchEngine.load and the entry point, against fake torch/semif modules (INV-3, INV-4, INV-5). +The bug hunt found load() had no test at all: removing the arch check survived every test.""" +import os +import sys +import types + +import pytest + +from semif_serve.config import Settings + +TOKEN = "t" * 40 + + +class Param: + def __init__(self, device): + self.device = types.SimpleNamespace(type=device) + + +class Model: + def __init__(self, device): + self._device = device + + def parameters(self): + yield Param(self._device) + + +@pytest.fixture +def fakes(monkeypatch): + calls = [] + cuda = types.SimpleNamespace( + is_available=lambda: True, + get_device_capability=lambda _i=0: (12, 0), + get_arch_list=lambda: ["sm_90", "sm_120"], + get_device_properties=lambda _i=0: types.SimpleNamespace(total_memory=96 * 2**30), + set_per_process_memory_fraction=lambda f, _i=0: calls.append(("cap", round(f, 4))), + memory_reserved=lambda _i=0: 8 * 2**30, + OutOfMemoryError=type("OutOfMemoryError", (RuntimeError,), {}), + empty_cache=lambda: calls.append(("empty",)), + ) + torch = types.SimpleNamespace(cuda=cuda, __version__="2.10.0+cu128") + state = {"device": "cuda", "warmup_raises": None} + + def load_causal_model(model, revision, device, dtype): + calls.append(("load", model, revision, device, dtype)) + return Model(state["device"]), object(), {"source": model} + + def score(model, tok, row, meta, max_tokens): + calls.append(("score", row["id"])) + if state["warmup_raises"]: + raise state["warmup_raises"] + return {"id": row["id"]} + + core = types.ModuleType("semif_phase1.core") + core.load_causal_model = load_causal_model + direct = types.ModuleType("semif_phase1.direct") + direct.score = score + shared = types.ModuleType("semif_phase1.shared") + shared.score_shared = lambda *a: ([], {}) + pkg = types.ModuleType("semif_phase1") + for name, mod in {"torch": torch, "semif_phase1": pkg, "semif_phase1.core": core, + "semif_phase1.direct": direct, "semif_phase1.shared": shared}.items(): + monkeypatch.setitem(sys.modules, name, mod) + return torch, calls, state + + +def test_load_caps_before_the_weights_land_then_warms_up(fakes): + from semif_serve.engine import TorchEngine + _torch, calls, _ = fakes + TorchEngine.load(Settings(api_token=TOKEN, vram_cap_gib=12.0)) + assert [c[0] for c in calls] == ["cap", "load", "score"] + assert calls[0] == ("cap", round(12 / 96, 4)) + assert calls[2] == ("score", "semif-serve-warmup") + + +def test_load_refuses_a_card_torch_has_no_kernels_for(fakes): + from semif_serve.engine import TorchEngine + torch, calls, _ = fakes + torch.cuda.get_arch_list = lambda: ["sm_80", "sm_90"] + with pytest.raises(RuntimeError, match="sm_120"): + TorchEngine.load(Settings(api_token=TOKEN)) + assert not any(c[0] == "load" for c in calls) + + +def test_load_refuses_a_model_that_landed_on_the_wrong_device(fakes): + from semif_serve.engine import TorchEngine + _torch, _calls, state = fakes + state["device"] = "cpu" + with pytest.raises(RuntimeError, match="landed on cpu"): + TorchEngine.load(Settings(api_token=TOKEN)) + + +def test_load_fails_closed_when_the_warmup_decision_fails(fakes): + from semif_serve.engine import TorchEngine + from semif_serve.errors import ScoringFailed + _torch, _calls, state = fakes + state["warmup_raises"] = RuntimeError("Failed to find C compiler") + with pytest.raises(ScoringFailed, match="C compiler"): + TorchEngine.load(Settings(api_token=TOKEN)) + + +def test_the_entry_point_forces_offline_mode_before_the_engine_loads(fakes, monkeypatch): + import semif_serve.engine as engine_mod + from semif_serve import main + seen = {} + monkeypatch.delenv("HF_HUB_OFFLINE", raising=False) + monkeypatch.setenv("SEMIF_API_TOKEN", TOKEN) + monkeypatch.setattr(engine_mod.TorchEngine, "load", + classmethod(lambda cls, s: seen.update(offline=os.environ.get("HF_HUB_OFFLINE")) or object())) + main.app_from_env() + assert seen["offline"] == "1" diff --git a/services/semif-serve/tests/test_orderings.py b/services/semif-serve/tests/test_orderings.py new file mode 100644 index 0000000..53168fe --- /dev/null +++ b/services/semif-serve/tests/test_orderings.py @@ -0,0 +1,132 @@ +"""Order averaging (0.1.3). Contract: semif-serve.contract.md § Order averaging.""" +import math + +import pytest +from fastapi.testclient import TestClient + +from semif_serve.app import create_app +from semif_serve.config import Settings + +TOKEN = "t" * 40 +AUTH = {"Authorization": f"Bearer {TOKEN}"} +OPTS = [{"id": "casual", "description": "A casual outing"}, + {"id": "date", "description": "A romantic date"}, + {"id": "booty", "description": "A booty call"}] +ROW = {"id": "q", "state": "It's 2AM and I'm bored.", "question": "What is this?", "options": OPTS} +BASE = {"casual": 1.0, "date": -1.0, "booty": 0.5} # what the model "really" thinks +BIAS = 2.5 # added to whichever option is listed first + + +def softmax(xs): + top = max(xs) + w = [math.exp(x - top) for x in xs] + return [v / sum(w) for v in w] + + +class BiasedEngine: + """Scores each option as BASE[id], plus BIAS for the first-listed option, the way a small model leans.""" + + def __init__(self): + self.shared_calls, self.direct_calls = [], [] + + def health(self): + return {} + + def _score(self, row): + ids = [o["id"] for o in row["options"]] + logits = [BASE[i] + (BIAS if k == 0 else 0.0) for k, i in enumerate(ids)] + return {"id": row["id"], "option_ids": ids, "probabilities": softmax(logits), "option_logits": logits, + "prompt_sha256": row["id"], "probability_status": "uncalibrated"} + + def direct(self, row): + self.direct_calls.append(row) + return self._score(row) + + def shared(self, rows): + self.shared_calls.append(rows) + return [self._score(r) for r in rows], {"batch_size": len(rows)} + + +def client(engine, **kw): + return TestClient(create_app(Settings(api_token=TOKEN, **kw), engine)) + + +def test_rotations_send_one_shared_call_with_every_option_once_per_position_and_cancel_the_bias(): + engine = BiasedEngine() + body = client(engine).post("/decide", json={**ROW, "orderings": "rotations"}, headers=AUTH).json() + assert engine.direct_calls == [] and len(engine.shared_calls) == 1 + rows = engine.shared_calls[0] + assert [r["id"] for r in rows] == ["q#o0", "q#o1", "q#o2"] + assert [[o["id"] for o in r["options"]] for r in rows] == [ + ["casual", "date", "booty"], ["date", "booty", "casual"], ["booty", "casual", "date"]] + assert all(r["state"] == ROW["state"] and r["question"] == ROW["question"] for r in rows) + + assert body["id"] == "q" and body["option_ids"] == ["casual", "date", "booty"] + combined = body["combined"] + assert combined["method"] == "rotations" and combined["orderings"] == 3 + # the first-position bias is additive, and every option sat first exactly once, so it cancels exactly + assert combined["probabilities"] == pytest.approx(softmax([BASE["casual"], BASE["date"], BASE["booty"]])) + assert combined["top"] == "casual" + # the bias is strong enough that every ordering's first option wins: casual, date, booty + tops = [max(zip(r["option_logits"], r["option_ids"]))[1] for r in body["orderings"]] + assert tops == ["casual", "date", "booty"] + assert combined["agreement"] == pytest.approx(1 / 3) + assert body["orderings"] == [engine._score(r) for r in rows] # native results, unchanged + for oid in ("casual", "date", "booty"): + ps = [dict(zip(r["option_ids"], r["probabilities"]))[oid] for r in body["orderings"]] + assert combined["spread"][oid] == pytest.approx([min(ps), max(ps)]) + + +def test_all_sends_every_permutation_with_the_callers_order_first(): + engine = BiasedEngine() + body = client(engine).post("/decide", json={**ROW, "orderings": "all"}, headers=AUTH).json() + rows = engine.shared_calls[0] + orders = [tuple(o["id"] for o in r["options"]) for r in rows] + assert len(orders) == 6 and len(set(orders)) == 6 and orders[0] == ("casual", "date", "booty") + assert body["combined"]["orderings"] == 6 and body["combined"]["method"] == "all" + + +def test_all_above_four_options_is_422_before_the_engine_runs(): + engine = BiasedEngine() + five = [{"id": f"o{i}", "description": f"Option {i}"} for i in range(5)] + r = client(engine).post("/decide", json={**ROW, "options": five, "orderings": "all"}, headers=AUTH) + assert r.status_code == 422 and "rotations" in r.json()["error"]["message"] + assert engine.shared_calls == [] + + +def test_a_mixed_shared_request_is_one_engine_call_with_results_in_request_order(): + engine = BiasedEngine() + body = {"state": ROW["state"], "decisions": [ + {"id": "plain", "question": "Q1?", "options": OPTS}, + {"id": "avg", "question": "Q2?", "options": OPTS, "orderings": "rotations"}, + {"id": "plain2", "question": "Q3?", "options": OPTS[:2]}]} + out = client(engine).post("/decide/shared", json=body, headers=AUTH).json() + assert len(engine.shared_calls) == 1 + assert [r["id"] for r in engine.shared_calls[0]] == ["plain", "avg#o0", "avg#o1", "avg#o2", "plain2"] + ids = [r["id"] for r in out["results"]] + assert ids == ["plain", "avg", "plain2"] + assert out["results"][0] == engine._score(engine.shared_calls[0][0]) # plain results keep the old shape + assert "combined" in out["results"][1] and "combined" not in out["results"][2] + assert out["timing"] == {"batch_size": 5} + + +def test_expanded_rows_count_toward_the_cap(): + engine = BiasedEngine() + r = client(engine, max_decisions=5).post("/decide/shared", json={"state": "s", "decisions": [ + {"id": "a", "question": "Q?", "options": OPTS, "orderings": "rotations"}, + {"id": "b", "question": "Q?", "options": OPTS, "orderings": "rotations"}]}, headers=AUTH) + assert r.status_code == 422 and "6 scored rows" in r.json()["error"]["message"] + assert engine.shared_calls == [] + + +def test_workload_with_orderings_is_422_and_workload_alone_still_calibrates(): + engine = BiasedEngine() + c = client(engine, calibration={"triage": 2.0}) + r = c.post("/decide", json={**ROW, "orderings": "rotations", "workload": "triage"}, headers=AUTH) + assert r.status_code == 422 and engine.shared_calls == [] + assert "calibrated" in c.post("/decide", json={**ROW, "workload": "triage"}, headers=AUTH).json() + + +def test_an_unknown_orderings_value_is_422(): + r = client(BiasedEngine()).post("/decide", json={**ROW, "orderings": "shuffle"}, headers=AUTH) + assert r.status_code == 422 diff --git a/services/semif-serve/uv.lock b/services/semif-serve/uv.lock index 9a367f0..c68f554 100644 --- a/services/semif-serve/uv.lock +++ b/services/semif-serve/uv.lock @@ -51,6 +51,26 @@ wheels = [ { url = "https://files.pythonhosted.org/packages/12/b8/4bd346e22b28902df4d651910f5242c28d84e4a5c2435ca5c3f797ed7e2e/anyio-4.15.1-py3-none-any.whl", hash = "sha256:6152fdbbf9a77fdec97731721bebf7c4c44f7c29b424b0065826173efc7ed101", size = 132079 }, ] +[[package]] +name = "causal-conv1d" +version = "1.7.0" +source = { url = "https://github.com/Dao-AILab/causal-conv1d/releases/download/v1.7.0/causal_conv1d-1.7.0+cu12torch2.10cxx11abiTRUE-cp312-cp312-linux_x86_64.whl" } +dependencies = [ + { name = "ninja" }, + { name = "packaging" }, + { name = "torch" }, +] +wheels = [ + { url = "https://github.com/Dao-AILab/causal-conv1d/releases/download/v1.7.0/causal_conv1d-1.7.0+cu12torch2.10cxx11abiTRUE-cp312-cp312-linux_x86_64.whl", hash = "sha256:8e81f8c76435ad31aa41edc6c0c9d26de971a293b190183d2816da71547614d7" }, +] + +[package.metadata] +requires-dist = [ + { name = "ninja" }, + { name = "packaging" }, + { name = "torch" }, +] + [[package]] name = "certifi" version = "2026.7.22" @@ -106,6 +126,15 @@ wheels = [ { url = "https://files.pythonhosted.org/packages/98/59/239c7259e669c46ddbcac0aa60e3a0ef00bfeaaa687f905b24dd6a7a10fe/cuda_pathfinder-1.8.2-py3-none-any.whl", hash = "sha256:4e65059febdb4d19d5cbc4798677e19db2b582f2f702f457b609e571690d357e", size = 62551 }, ] +[[package]] +name = "einops" +version = "0.8.2" +source = { registry = "https://pypi.org/simple" } +sdist = { url = "https://files.pythonhosted.org/packages/2c/77/850bef8d72ffb9219f0b1aac23fbc1bf7d038ee6ea666f331fa273031aa2/einops-0.8.2.tar.gz", hash = "sha256:609da665570e5e265e27283aab09e7f279ade90c4f01bcfca111f3d3e13f2827", size = 56261 } +wheels = [ + { url = "https://files.pythonhosted.org/packages/2a/09/f8d8f8f31e4483c10a906437b4ce31bdf3d6d417b73fe33f1a8b59e34228/einops-0.8.2-py3-none-any.whl", hash = "sha256:54058201ac7087911181bfec4af6091bb59380360f069276601256a76af08193", size = 65638 }, +] + [[package]] name = "fastapi" version = "0.118.0" @@ -129,6 +158,31 @@ wheels = [ { url = "https://files.pythonhosted.org/packages/bc/ac/8b9c6dc2aa7e9cf5582e20a337006c3428c3d752c2255dbaca2c78a8be45/filelock-4.0.4-py3-none-any.whl", hash = "sha256:0df72be195ca7892216d16f2edce8d9b93a571f02402972020a8cff84c594c7b", size = 108629 }, ] +[[package]] +name = "fla-core" +version = "0.5.2" +source = { registry = "https://pypi.org/simple" } +dependencies = [ + { name = "einops" }, +] +sdist = { url = "https://files.pythonhosted.org/packages/94/85/20bdc0fbbbaeec27b5b198e0dbbbc51161a267066a4aebd9e4b413637c6c/fla_core-0.5.2.tar.gz", hash = "sha256:9360fc412f784c1c8f05c320ec2902d5d968764178c9b8a92efc919e17a39680", size = 587835 } +wheels = [ + { url = "https://files.pythonhosted.org/packages/2d/ed/dfe19c4da779957eb6a42a26812f9b4e2280bf757a17a71933ff59ffcb98/fla_core-0.5.2-py3-none-any.whl", hash = "sha256:5e830c85bad3d0d34677f98ac7074d08687a3756f0f0499d95ceb96eb6920761", size = 819225 }, +] + +[[package]] +name = "flash-linear-attention" +version = "0.5.2" +source = { registry = "https://pypi.org/simple" } +dependencies = [ + { name = "fla-core" }, + { name = "transformers" }, +] +sdist = { url = "https://files.pythonhosted.org/packages/d2/76/c180949eae5161b9fcf4c928cab52f0c257ce84c1a59b4462b1d2b6e5883/flash_linear_attention-0.5.2.tar.gz", hash = "sha256:c053d3a75c8f5b725063f719ae6bce0f4d9dac6da32735be8ae7d485a6c9820f", size = 208110 } +wheels = [ + { url = "https://files.pythonhosted.org/packages/90/d2/2070e3cf2148c5cce99ca4876633c4b7b89ec085323600bd3d47aeacd306/flash_linear_attention-0.5.2-py3-none-any.whl", hash = "sha256:dcf405d81f5426393b59037097aa700d0f4a841465d5028d5aa543f4502f2400", size = 399590 }, +] + [[package]] name = "fsspec" version = "2026.9.0" @@ -351,6 +405,24 @@ wheels = [ { url = "https://files.pythonhosted.org/packages/7e/cd/fe58041e9011f307c490e3e17dd48cc516448f7c698a3f2d9d9d65d7e6a8/networkx-3.7-py3-none-any.whl", hash = "sha256:e3fd2c13a7814cee3746340d8d7f8598a67f16a58bf47fb7f8793fab6efca1b0", size = 2142205 }, ] +[[package]] +name = "ninja" +version = "1.13.2" +source = { registry = "https://pypi.org/simple" } +sdist = { url = "https://files.pythonhosted.org/packages/ac/92/410b7917d16ab54c04b05cc32b9284803671d91cf79d33be6009c28d4ea8/ninja-1.13.2.tar.gz", hash = "sha256:525bfa3fc88aa30a4467df270fd5be6f9fcae8061d54d4df74ea1dc5abd5a975", size = 243739 } +wheels = [ + { url = "https://files.pythonhosted.org/packages/48/23/fcbe234a66966e35928c47b86336f92a7612db4781665f4e5f5fddef9630/ninja-1.13.2-py3-none-manylinux2014_ppc64le.manylinux_2_17_ppc64le.whl", hash = "sha256:227cbc3ae3e5e429692388103cae8c09451df086cd2d342dae0795af0d162547", size = 197676 }, + { url = "https://files.pythonhosted.org/packages/24/eb/a6ca97ef0ff7bb8bdcb395ec65a716e65d7c1f40896c3afe0090bb3e1535/ninja-1.13.2-py3-none-manylinux2014_s390x.manylinux_2_17_s390x.whl", hash = "sha256:1684c60d031c54c1d049541b64243c0c567dca5463dbd77682a8901780af293d", size = 187980 }, + { url = "https://files.pythonhosted.org/packages/6e/53/ebfed7b689c338dd8ebeec9c0730c8d56821292f14e2536e5f3ef1a05744/ninja-1.13.2-py3-none-manylinux2014_x86_64.manylinux_2_17_x86_64.whl", hash = "sha256:65a24341b5ac09fcadcc37082660be40a94174e51a937fabf6e2cae26225fa2c", size = 183365 }, + { url = "https://files.pythonhosted.org/packages/c7/d6/dcf06d7ab44ade992ae5aa1228feff317684b463a1bd47e8642b30ac922e/ninja-1.13.2-py3-none-manylinux_2_28_armv7l.manylinux_2_31_armv7l.whl", hash = "sha256:aa3d2ae5706a2c4d1e93edc951d1c6cbb45107413c404f8fde1741239efbc9a0", size = 155089 }, + { url = "https://files.pythonhosted.org/packages/e1/6b/6513c09c33382b17c05b4349b8e81437b18680d0d7ec6fb8f7edc29adda1/ninja-1.13.2-py3-none-manylinux_2_31_riscv64.whl", hash = "sha256:919572cbc3f233261ecd41fe1f3efc9d44aa02464a4588867e06a8b4f6f416ea", size = 152149 }, + { url = "https://files.pythonhosted.org/packages/37/04/c8c2dc5b2f5fee79a1691d490256b178b7e1af97d56769117422ae8a23cc/ninja-1.13.2-py3-none-musllinux_1_2_armv7l.whl", hash = "sha256:59d71c3e15b6b6f3d903eb0c27285544e0747ca59925ada7037bb1af781ad4b3", size = 466268 }, + { url = "https://files.pythonhosted.org/packages/10/a2/d8eedd25d0ae80b9e874aea362013e67416877c8f732eb4d5e7c971fdb9c/ninja-1.13.2-py3-none-musllinux_1_2_ppc64le.whl", hash = "sha256:b2f687437fac460b27b7eadc99039b1163016fb4ba7276e2782a192d9f24ee0e", size = 610806 }, + { url = "https://files.pythonhosted.org/packages/14/0f/696d96821fad1b5767fd311c1569dde8881a57412369bfe7b11bcbfde036/ninja-1.13.2-py3-none-musllinux_1_2_riscv64.whl", hash = "sha256:09de9ab04f7352f51570c73fd4913acb1e6c24be0a72cd8b80243d4d3ed04925", size = 533978 }, + { url = "https://files.pythonhosted.org/packages/5d/69/28844ca579156776a202217a7cd66f60d06a0710a935e879bb89ce396ecc/ninja-1.13.2-py3-none-musllinux_1_2_s390x.whl", hash = "sha256:6a87bf42b123abe2f37737300185f0a303a891899da85d73a3613ee80547e578", size = 653822 }, + { url = "https://files.pythonhosted.org/packages/f5/5f/c511f2952f94ab2966d60edd9c34e744ea32f2724b1184b62270bde55b3a/ninja-1.13.2-py3-none-musllinux_1_2_x86_64.whl", hash = "sha256:915bd482c4be41c75120fd67a22e0bb3f0fbb3bbc5f95b89787deadd59e27ef2", size = 544460 }, +] + [[package]] name = "numpy" version = "2.2.6" @@ -919,7 +991,7 @@ dependencies = [ [[package]] name = "semif-serve" -version = "0.1.2" +version = "0.1.3" source = { editable = "." } dependencies = [ { name = "fastapi" }, @@ -927,6 +999,10 @@ dependencies = [ ] [package.optional-dependencies] +fast = [ + { name = "causal-conv1d", marker = "platform_machine == 'x86_64' and sys_platform == 'linux'" }, + { name = "flash-linear-attention" }, +] model = [ { name = "semif-phase1" }, { name = "torch" }, @@ -940,12 +1016,14 @@ dev = [ [package.metadata] requires-dist = [ + { name = "causal-conv1d", marker = "platform_machine == 'x86_64' and sys_platform == 'linux' and extra == 'fast'", url = "https://github.com/Dao-AILab/causal-conv1d/releases/download/v1.7.0/causal_conv1d-1.7.0+cu12torch2.10cxx11abiTRUE-cp312-cp312-linux_x86_64.whl" }, { name = "fastapi", specifier = "==0.118.0" }, + { name = "flash-linear-attention", marker = "extra == 'fast'", specifier = "==0.5.2" }, { name = "semif-phase1", marker = "extra == 'model'", git = "https://github.com/TheoLeeCJ/SemIf-OpenJev?rev=23cf1f39fc9534fe81437200959b6dfc7106e45a" }, { name = "torch", marker = "extra == 'model'", specifier = "==2.10.0", index = "https://download.pytorch.org/whl/cu128" }, { name = "uvicorn", specifier = "==0.37.0" }, ] -provides-extras = ["model"] +provides-extras = ["model", "fast"] [package.metadata.requires-dev] dev = [ diff --git a/stacks/semif/.env.example b/stacks/semif/.env.example index 67f84dd..08899fa 100644 --- a/stacks/semif/.env.example +++ b/stacks/semif/.env.example @@ -1,6 +1,6 @@ # semif — copy to /opt/docker/compose/semif/.env on fv-ml1 (mode 0600). # Built on fv-ml1 from services/semif-serve (see README "Building"). -IMAGE=semif-serve:0.1.2 +IMAGE=semif-serve:0.1.3 PORT=8032 HOST_IP=10.251.50.54 # fv-ml1 GPU 1 = the utility card (vllm-coder, erp, meromero, scriberr). diff --git a/stacks/semif/README.md b/stacks/semif/README.md index e70340a..0cc5b45 100644 --- a/stacks/semif/README.md +++ b/stacks/semif/README.md @@ -16,7 +16,7 @@ service. The contract is `services/semif-serve/semif-serve.contract.md`. | **Model** | `Qwen/Qwen3.5-4B` @ `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a`, BF16, from `/tank/aimodels/huggingface` (read-only, offline) | | **SemIf** | commit `23cf1f39fc9534fe81437200959b6dfc7106e45a`; torch `2.10.0+cu128`, transformers `5.17.0`, the same stack SemIf's committed predictions were made on | | **Image** | `semif-serve:`, built on fv-ml1 from `services/semif-serve/` | -| **State** | none. Weights re-download at the pinned revision. No backup needed beyond the host's `/opt/docker` restic. | +| **State** | none. If the weights are ever lost, re-pull them by hand at the pinned revision (the service itself never downloads: `HF_HUB_OFFLINE=1`, read-only mount). No backup needed beyond the host's `/opt/docker` restic. | ## API @@ -34,6 +34,18 @@ curl -s -H "Authorization: Bearer $T" http://10.251.50.54:8032/decide -d '{ the state once and scores every criterion in one batch. Use it when many questions share one long state. - `GET /health`: the pins, limits and calibrated workloads. +- **Order averaging (0.1.3):** add `"orderings": "rotations"` to a decision, or + `"all"` for ≤ 4 options. The option list is asked in every rotation inside ONE + shared batch. The reply keeps each ordering's native result under `orderings` and adds + `combined: {method, orderings, probabilities, top, agreement, spread}`. **Use it for + anything real:** a small model leans toward the first-listed option on ambiguous + inputs, and averaging cancels that. Through the service on SemIf's labelled sets, + accuracy goes from 78.6% to **88.1%** (group-bootstrap 95% CI +5.1..+14.3 pts, 252 + rows). **`agreement` is the cheap confidence signal**: unanimous rows are 94.5% + accurate, split rows 76.4%. The orderings count toward the decision cap. `workload` + calibration is not available together with `orderings` yet (422). +- Past `MAX_QUEUE` (32) requests in progress, new POSTs get `429 busy` before their + body is read. ## ⚠ Probabilities are uncalibrated @@ -57,18 +69,55 @@ not squeeze scriberr, which shares GPU 1. A request that would exceed the cap ge behaviours were verified on the card (0.1.0 held 11.9 GiB after an OOM, and 12.6 GB after a large request; 0.1.2 returns to 7.85 GiB in both cases). -**What fits under 12 GiB** (measured, `/decide/shared`, binary decisions): +**What fits under 12 GiB** (measured on 0.1.3, `/decide/shared`, binary decisions; +0.1.2 figures in brackets, before the fast kernels): -| state size (prefix tokens) | max decisions in one request | +| state size (prefix tokens) | max rows in one request | |---|---| -| ~140 | 52 | -| ~520 | 43 | -| ~1,960 | 19 | -| ~3,900 | 13 | +| ~140 | 63 (52) | +| ~520 | 51 (43) | +| ~1,960 | 26 (19) | +| ~3,900 | 16 (13) | + +Rows = decisions × orderings, so `rotations` over 3 options uses 3 rows per decision. `/decide` fits at the full 4,096-token limit. Past the table you get a 503, so split the decisions across requests. +## Fast kernels (0.1.3) + +The image ships Qwen3.5's fast kernels, `flash-linear-attention` 0.5.2 and +`causal-conv1d` 1.7.0 (a prebuilt cu12/torch2.10 wheel). Without them transformers +logs that it falls back to "much slower" reference PyTorch paths. They are adopted +because an A/B on the empty GPU 3 showed: +- **parity improved**: 144/144 vs upstream (142/144 without), so upstream evidently + ran with them; +- **long inputs got much faster**: a ~2,000-token `/decide` went from 169 to 92 ms + server-side. + +Short 3-rotation batches cost ~3–6 ms more; everything else is equal or faster. +⚠ triton builds a C shim at runtime, so the image carries `gcc`. Without it the +warm-up fails, and startup fails closed. Build without the kernels: +`--build-arg EXTRAS="--extra model"`. + +## Latency (0.1.3, from nh3-dev, 3 runs × 20, network floor ~33 ms) + +| request | end to end | server | +|---|---|---| +| `/decide`, short (~130 tok) | 69 ms | 35 ms | +| `/decide`, ~2,000-token state | 131 ms | 92 ms | +| 3 rotations, short | 115 ms | 78 ms | +| 6 orderings, short | 118 ms | 81 ms | +| 3 rotations, ~2,000-token state | 200 ms | 158 ms | + +## Acceptance (2026-09-27, v0.1.3) + +Raw: `services/semif-serve/acceptance/result-2026-09-27-v0.1.3.json`, +`averaging-2026-09-27-v0.1.3.json`. v0.1.3 matches upstream on **144/144** (identical +prompt hashes, max prob gap 0.049), is deterministic, fails the negative control as it +should (14/144), and shared matches direct on 72/72. The v0.1.2 table below is the +reference-kernel baseline. + ## Acceptance (2026-09-27, v0.1.2) Raw: `services/semif-serve/acceptance/result-2026-09-27-v0.1.2.json`. diff --git a/stacks/semif/compose.yaml b/stacks/semif/compose.yaml index cdfa326..73938ad 100644 --- a/stacks/semif/compose.yaml +++ b/stacks/semif/compose.yaml @@ -30,6 +30,8 @@ services: SEMIF_VRAM_CAP_GIB: ${VRAM_CAP_GIB:?set VRAM_CAP_GIB} SEMIF_MAX_TOKENS: ${MAX_TOKENS:-4096} SEMIF_MAX_DECISIONS: ${MAX_DECISIONS:-64} + # POSTs in progress (queued + scoring) before new ones get 429 busy. + SEMIF_MAX_QUEUE: ${MAX_QUEUE:-32} SEMIF_CALIBRATION: /conf/calibration.json volumes: # Pinned weights, read offline (HF_HUB_OFFLINE=1 in the image). Never downloads.