Order averaging (Prime, after the 739aa03 spike):
- A decision may set orderings: rotations|all (all only for <= 4 options). Every
ordering goes to the engine in one shared batch.
- The reply keeps each native result and adds combined {probabilities (log-mean),
top, agreement, spread}.
- Through the service on SemIf's labelled sets (252 rows): 78.6% -> 88.1%
(group-bootstrap 95% CI +5.1..+14.3). Unanimous agreement is 94.5% accurate.
Fast kernels: flash-linear-attention 0.5.2 and causal-conv1d 1.7.0 are now the
default build. A/B on the empty GPU 3:
- parity with upstream went from 142/144 to 144/144;
- a ~2k-token /decide went from 169 to 92 ms server-side;
- short 3-rotation batches cost ~3-6 ms more.
triton builds a C shim at runtime, so the image carries gcc. Without it the
warm-up failed and startup failed closed.
Heid bug-hunt panel (4/4 arms, thread 01M3H3F4RR7XBP90KQ3A39H4SX), folded:
- Startup validation: VRAM cap 0 no longer means uncapped (C1); limits must be
>= 1 (S1); the token must be visible ASCII (S2); the calibration file must
exist and parse, with T in [0.05, 20] (S8, and C3's NaN leg).
- The body limit is checked before a chunk is kept, and a Unicode-digit
Content-Length no longer crashes (C2, S3).
- Failures while building the response now get the 500 envelope (C3).
- 429 busy past SEMIF_MAX_QUEUE requests in progress (C6).
- The engine releases memory on every non-validation failure, unchained after
gc; an empty OOM message is handled; 'out of memory' RuntimeErrors map to 503
(C4, C5, S9).
- The entry point forces HF_HUB_OFFLINE (S10). README wording fixed (S5, S6).
- New guard tests close the gaps the arms' mutation grids exposed: early stop of
the body read, a shared-route lock, calibration pass-through, the gc cycle,
the exact caps, TorchEngine.load's arch and device checks, and the offline
entry point.
86 tests.
Deployed on fv-ml1 GPU 1: parity 144/144, OOM and burst release verified, shared
capacity 63/51/26/16 rows at ~140/520/1960/3900 prefix tokens.
9.9 KiB
title, kind, status, owner, created, depends_on
| title | kind | status | owner | created | depends_on | ||
|---|---|---|---|---|---|---|---|
| semif-serve | module-contract | draft | infra-ops | 2026-09-27 |
|
semif-serve: an HTTP wrapper around SemIf's direct and shared scorers
Purpose
SemIf decides by reading the logits of the option letters after one forward pass.
It ships as a batch CLI only. semif-serve loads the model once and exposes the
two torch scorers over HTTP, so fleet callers can ask typed questions without
decoding. It adds nothing to the scoring itself: every score it returns is what
semif_phase1.direct.score or semif_phase1.shared.score_shared returned,
unchanged, and it adds an optional calibrated view next to it.
Operator decisions (Prime, 2026-09-27): runs on fv-ml1 GPU 1 under a hard VRAM cap; built by infra-ops with a light process (this contract → TDD → heid bug hunt); no consumer is named yet.
Endpoints
Every POST takes and returns JSON. Every endpoint except GET /health requires
Authorization: Bearer <token>.
| method | path | body | success |
|---|---|---|---|
| GET | /health |
— | 200 {status: "ok", semif_commit, model, vram_cap_gib, max_tokens, max_decisions, workloads} |
| POST | /decide |
one SemIf row {id, state, question, options[2..16]} + optional workload |
200 the direct.score dict + optional calibrated |
| POST | /decide/shared |
{state, decisions: [{id, question, options}], workload?} |
200 {results: [...], timing: {...}} from score_shared |
options items are {id, description}, as in SemIf. state is a nonempty
string, object or array. For /decide/shared, each decision becomes a SemIf row
by adding the shared state.
Calibration. workload is optional. When it names an entry in the
calibration table, each result gains calibrated: {workload, temperature, probabilities} with softmax(option_logits / T). The argmax never changes. The
native probabilities and probability_status stay untouched. An unknown
workload is a 422. With no workload, no calibrated key appears.
Order averaging (0.1.3, Prime 2026-09-27). A decision (the /decide body, or an
entry in decisions) may set orderings, whose default is "none":
"rotations": the n cyclic shifts of the caller's option list, starting with the caller's order, so every option sits in every position exactly once."all": every permutation (n!), the caller's order first. It is allowed only when n ≤ 4, else 422.
Every ordering of every decision in the request becomes its own SemIf row
(id = "<id>#o<k>", same state and question, reordered options). All rows go
to the engine in ONE shared call, which includes /decide. The result for an
averaged decision is:
{id, option_ids (caller's order),
combined: {method, orderings: n, probabilities, top, agreement, spread: {option_id: [min, max]}},
orderings: [the SemIf result for each ordering, unchanged]}
probabilities: per ordering, log-softmax ofoption_logits; averaged per option id; renormalised; reported in the caller's order.agreement: the fraction of orderings whose top option equalscombined.top.spread: each option's min and max native probability across orderings.
Every expanded row counts toward SEMIF_MAX_DECISIONS. workload together with
orderings is a 422, because a temperature is fitted per method and none is
fitted on combined scores yet. Decisions without orderings keep the exact
pre-0.1.3 result shape. Averaging cancels any additive position bias exactly.
Measured by the 2026-09-27 spike: 3 rotations take SemIf's labelled sets from
78.6% to 87.7% accuracy, and unanimous agreement is 94.4% accurate.
Invariants
- INV-1 pass-through.
option_ids,probabilities,option_logits,prompt_sha256,prompt_version,modelandprobability_statusare exactly what SemIf returned. The wrapper never rewrites a score. - INV-2 one model, one inference at a time. The model is loaded at startup,
and a process-wide lock serialises every scorer call. The app runs as one worker.
Scorer calls run off the event loop, so
/healthanswers while one is in progress. - INV-3 fail-closed startup. Startup refuses to serve unless the model sits
on a CUDA device, torch's arch list includes the card's
sm_XY, and one warm-up decision scores.device=cpuis allowed only when set explicitly. - INV-4 VRAM cap. When
SEMIF_VRAM_CAP_GIBis set, the process is capped at that share of the card (torch.cuda.set_per_process_memory_fraction) before the model loads. An out-of-memory error during a request is a 503out_of_memory, followed bytorch.cuda.empty_cache(). The process stays up. After every scorer call, when the reserved memory exceeds the post-warm-up baseline by more than 512 MiB, the engine callsempty_cache(). A burst must not keep the card's shared headroom: on 2026-09-27 a 64-decision request left the process holding 12.6 GB, leaving scriberr 3.5 GB. Any other scorer failure exceptValueErroralso releases before it is reported: it is logged with its traceback, then re-raised unchained asScoringFailedaftergc.collect()+empty_cache()(bug hunt C5). ARuntimeErrorwhose message says "out of memory" (cuBLAS/cuDNN allocation failures) counts as an OOM → 503 (S9). An OOM with an empty message is reported as "CUDA out of memory" rather than crashing the handler (C4). - INV-5 no network at runtime. Weights come from the mounted HF cache at the
pinned revision. The entry point sets
HF_HUB_OFFLINE=1itself before torch or transformers load, so this holds outside the image too (S10). - INV-6 constant-time auth. Token comparison uses
hmac.compare_digest. The token is ≥ 32 characters of visible ASCII (33–126). Startup refuses anything else, because a CR, LF or NUL in the token can never arrive in a header (S2).
Limits and errors
SEMIF_MAX_TOKENS(default 4096) is passed to the scorers. A longer prompt is a 422, never truncated (SemIf raises).SEMIF_MAX_DECISIONS(default 64) capsdecisionsper shared request. It must hold 1..max entries, else 422.- Request body ≤
SEMIF_MAX_BODY_BYTES(default 1 MiB), else 413. The limit is checked before each chunk is kept, so no more than the limit is ever held (C2). A declaredContent-Lengthis trusted only if it is ASCII digits (S3). - Admission: at most
SEMIF_MAX_QUEUE(default 32) POSTs may be in progress, counting both queued and scoring. The next one is refused with 429busybefore its body is read (C6).
| status | code | when |
|---|---|---|
| 401 | unauthorized |
missing or wrong bearer |
| 413 | request_too_large |
body over the limit |
| 422 | invalid_request |
bad JSON shape, a SemIf ValueError (validation, token limit, tokenisation), unknown workload, too many decisions |
| 429 | busy |
SEMIF_MAX_QUEUE requests already in progress |
| 503 | out_of_memory |
CUDA OOM during scoring |
| 500 | scoring_failed |
any other scorer exception, or a failure while building the response from a scorer result (calibration, averaging): always the envelope, never a bare 500 (C3) |
The error body is {error: {code, message}}.
Configuration (env)
SEMIF_API_TOKEN (required), SEMIF_MODEL (default Qwen/Qwen3.5-4B),
SEMIF_REVISION (default the pinned SHA), SEMIF_DEVICE (default cuda),
SEMIF_VRAM_CAP_GIB, SEMIF_MAX_TOKENS, SEMIF_MAX_DECISIONS,
SEMIF_MAX_BODY_BYTES, SEMIF_MAX_QUEUE, SEMIF_CALIBRATION (path to a JSON
{workload: T}).
Startup validates every value and refuses a bad one with a ValueError naming
the variable (C1, S1, S8):
- the VRAM cap, when set, is finite and > 0 (
0used to mean uncapped); - every limit is an integer ≥ 1;
- each T is a finite number, not a bool, in [0.05, 20] (a tiny T overflowed to NaN, and the response then failed to render);
- the calibration file must exist and parse.
Tests (TDD, fake scorer: no torch, no model)
auth required on POSTs and not on /health; a short token is refused at startup;
/decide passes the row through and returns the scorer dict unchanged;
/decide/shared builds rows with the shared state and returns results + timing;
calibration adds calibrated and keeps the argmax; an unknown workload → 422; a
scorer ValueError → 422; the engine's OutOfMemory → 503; the torch engine,
against a fake torch, turns torch.cuda.OutOfMemoryError into an unchained
OutOfMemory and calls empty_cache() only after the failed call's tensors are freed
(found on the card: a chained exception kept 11.9 GiB allocated after the 503); after
a call, reserved memory over baseline + 512 MiB is released and at or under it is left
alone; any
other exception → 500; malformed
JSON or a wrong body shape → 422; too many decisions → 422; an oversized body → 413; requests
are serialised (two concurrent calls never overlap inside the scorer); /health
answers while a scorer call is blocked. Averaging: rotations sends n rows
in one shared call, each option once per position, with ids <id>#o<k>, and
cancels a position bias exactly; all sends n! rows and is 422 above 4 options;
agreement and spread are computed from the orderings; a mixed shared request
(averaged + plain) is one engine call, with results in request order and plain
results unchanged; expanded rows count toward the cap; workload + orderings
→ 422.
Acceptance (on fv-ml1, real model; not unit tests)
- Parity: our
/decideover SemIf'sauthored144against their committed torch predictions (top choice and max probability gap). - Noise floor: the same run twice (A-vs-A).
- Negative control: shuffled option descriptions must break agreement.
- Shared vs direct: the same rows agree within the A-vs-A floor.
- Speed: 21 binary criteria over one state, N ≥ 3, p50 + spread.
- VRAM: the peak at a 4096-token input sets
SEMIF_VRAM_CAP_GIB.