services/semif-serve is a FastAPI wrapper around SemIf's direct and shared torch scorers (SemIf-OpenJev @ 23cf1f39, MIT). Upstream ships only a batch CLI. The wrapper loads the pinned Qwen3.5-4B (851bf6e8, BF16) once from the offline HF cache and returns SemIf's result dicts unchanged, with an optional per-workload temperature-calibrated view. Contract: semif-serve.contract.md. Built with a short contract, TDD (39 tests, fake engine and fake torch, no GPU) and a heid bug-hunt panel (pending). On the card: - torch 2.10.0+cu128 with sm_120 kernels, which is SemIf's own stack; - a hard 12 GiB VRAM cap. Two defects surfaced only on the card, and each fix is covered by a test: - 0.1.1: an OOM raised as a chained exception kept the failed request's tensors alive (11.9 GiB after the 503). It is now raised unchained, after gc. - 0.1.2: a large request left 12.6 GB reserved on the shared card. After each call, reserved memory over the baseline + 512 MiB is now released. Acceptance against SemIf's committed torch predictions (authored144): - 142/144 same top choice; both misses are exact bf16 ties; - 144/144 identical prompt hashes; - deterministic A-vs-A; - negative control 14/144; - shared vs direct 72/72. 21 binary criteria over one state take 159 ms. The shared-mode capacity table under the cap is in stacks/semif/README.md. The Dockerfile installs dependencies from a manifest with the project version blanked, so a version bump reuses the ~4 GB torch layer. Verified: 41 s rebuild, dependency layer CACHED. DNS: semif.fv.internal. Token: vault semif/api-token.
6.3 KiB
title, kind, status, owner, created, depends_on
| title | kind | status | owner | created | depends_on | ||
|---|---|---|---|---|---|---|---|
| semif-serve | module-contract | draft | infra-ops | 2026-09-27 |
|
semif-serve: an HTTP wrapper around SemIf's direct and shared scorers
Purpose
SemIf decides by reading the logits of the option letters after one forward pass.
It ships as a batch CLI only. semif-serve loads the model once and exposes the
two torch scorers over HTTP, so fleet callers can ask typed questions without
decoding. It adds nothing to the scoring itself: every score it returns is what
semif_phase1.direct.score or semif_phase1.shared.score_shared returned,
unchanged, and it adds an optional calibrated view next to it.
Operator decisions (Prime, 2026-09-27): runs on fv-ml1 GPU 1 under a hard VRAM cap; built by infra-ops with a light process (this contract → TDD → heid bug hunt); no consumer is named yet.
Endpoints
Every POST takes and returns JSON. Every endpoint except GET /health requires
Authorization: Bearer <token>.
| method | path | body | success |
|---|---|---|---|
| GET | /health |
— | 200 {status: "ok", semif_commit, model, vram_cap_gib, max_tokens, max_decisions, workloads} |
| POST | /decide |
one SemIf row {id, state, question, options[2..16]} + optional workload |
200 the direct.score dict + optional calibrated |
| POST | /decide/shared |
{state, decisions: [{id, question, options}], workload?} |
200 {results: [...], timing: {...}} from score_shared |
options items are {id, description}, as in SemIf. state is a nonempty
string, object or array. For /decide/shared, each decision becomes a SemIf row
by adding the shared state.
Calibration. workload is optional. When it names an entry in the
calibration table, each result gains calibrated: {workload, temperature, probabilities} with softmax(option_logits / T). The argmax never changes. The
native probabilities and probability_status stay untouched. An unknown
workload is a 422. With no workload, no calibrated key appears.
Invariants
- INV-1 pass-through.
option_ids,probabilities,option_logits,prompt_sha256,prompt_version,modelandprobability_statusare exactly what SemIf returned. The wrapper never rewrites a score. - INV-2 one model, one inference at a time. The model is loaded at startup,
and a process-wide lock serialises every scorer call. The app runs as one worker.
Scorer calls run off the event loop, so
/healthanswers while one is in progress. - INV-3 fail-closed startup. Startup refuses to serve unless the model sits
on a CUDA device, torch's arch list includes the card's
sm_XY, and one warm-up decision scores.device=cpuis allowed only when set explicitly. - INV-4 VRAM cap. When
SEMIF_VRAM_CAP_GIBis set, the process is capped at that share of the card (torch.cuda.set_per_process_memory_fraction) before the model loads. An out-of-memory error during a request is a 503out_of_memory, followed bytorch.cuda.empty_cache(). The process stays up. After every scorer call, when the reserved memory exceeds the post-warm-up baseline by more than 512 MiB, the engine callsempty_cache(). A burst must not keep the card's shared headroom: on 2026-09-27 a 64-decision request left the process holding 12.6 GB, leaving scriberr 3.5 GB. - INV-5 no network at runtime. Weights come from the mounted HF cache at the
pinned revision (
HF_HUB_OFFLINE=1). - INV-6 constant-time auth. Token comparison uses
hmac.compare_digest. The token is ≥ 32 characters, and startup refuses a shorter one.
Limits and errors
SEMIF_MAX_TOKENS(default 4096) is passed to the scorers. A longer prompt is a 422, never truncated (SemIf raises).SEMIF_MAX_DECISIONS(default 64) capsdecisionsper shared request. It must hold 1..max entries, else 422.- Request body ≤
SEMIF_MAX_BODY_BYTES(default 1 MiB), else 413.
| status | code | when |
|---|---|---|
| 401 | unauthorized |
missing or wrong bearer |
| 413 | request_too_large |
body over the limit |
| 422 | invalid_request |
bad JSON shape, a SemIf ValueError (validation, token limit, tokenisation), unknown workload, too many decisions |
| 503 | out_of_memory |
CUDA OOM during scoring |
| 500 | scoring_failed |
any other scorer exception |
The error body is {error: {code, message}}.
Configuration (env)
SEMIF_API_TOKEN (required), SEMIF_MODEL (default Qwen/Qwen3.5-4B),
SEMIF_REVISION (default the pinned SHA), SEMIF_DEVICE (default cuda),
SEMIF_VRAM_CAP_GIB, SEMIF_MAX_TOKENS, SEMIF_MAX_DECISIONS,
SEMIF_MAX_BODY_BYTES, SEMIF_CALIBRATION (path to a JSON {workload: T}; T > 0).
Tests (TDD, fake scorer: no torch, no model)
auth required on POSTs and not on /health; a short token is refused at startup;
/decide passes the row through and returns the scorer dict unchanged;
/decide/shared builds rows with the shared state and returns results + timing;
calibration adds calibrated and keeps the argmax; an unknown workload → 422; a
scorer ValueError → 422; the engine's OutOfMemory → 503; the torch engine,
against a fake torch, turns torch.cuda.OutOfMemoryError into an unchained
OutOfMemory and calls empty_cache() only after the failed call's tensors are freed
(found on the card: a chained exception kept 11.9 GiB allocated after the 503); after
a call, reserved memory over baseline + 512 MiB is released and at or under it is left
alone; any
other exception → 500; malformed
JSON or a wrong body shape → 422; too many decisions → 422; an oversized body → 413; requests
are serialised (two concurrent calls never overlap inside the scorer); /health
answers while a scorer call is blocked.
Acceptance (on fv-ml1, real model; not unit tests)
- Parity: our
/decideover SemIf'sauthored144against their committed torch predictions (top choice and max probability gap). - Noise floor: the same run twice (A-vs-A).
- Negative control: shuffled option descriptions must break agreement.
- Shared vs direct: the same rows agree within the A-vs-A floor.
- Speed: 21 binary criteria over one state, N ≥ 3, p50 + spread.
- VRAM: the peak at a 4096-token input sets
SEMIF_VRAM_CAP_GIB.