SemIf's shared scorer trims one token at the state boundary. When an object state's last value ends in ')', ';' or '}', the JSON that follows re-merges two tokens back, so score_shared refused the request with 422. The engine now wraps semif_phase1.shared._state_prefix to keep only the tokens the full prompts share. Each row scores the same token sequence; only the prefill/suffix split moves. Startup proves the fix is in effect, not just installed (heid bug hunt SKAL, folded). It checks that the hook is callable and is what score_shared resolves, that an ordinary state keeps upstream's whole prefix, and that a merge-prone state scores through the shared path. Real tokenizer: 154 states, 23 refused before and 0 after, with no ordinary or authored144 prefix changed. Acceptance: 144/144 parity. Shared vs direct 71/72; the miss is a bf16 tie that flipped across a plain restart (see README).
12 KiB
title, kind, status, owner, created, depends_on
| title | kind | status | owner | created | depends_on | ||
|---|---|---|---|---|---|---|---|
| semif-serve | module-contract | draft | infra-ops | 2026-09-27 |
|
semif-serve: an HTTP wrapper around SemIf's direct and shared scorers
Purpose
SemIf decides by reading the logits of the option letters after one forward pass.
It ships as a batch CLI only. semif-serve loads the model once and exposes the
two torch scorers over HTTP, so fleet callers can ask typed questions without
decoding. It adds nothing to the scoring itself: every score it returns is what
semif_phase1.direct.score or semif_phase1.shared.score_shared returned,
unchanged, and it adds an optional calibrated view next to it.
Operator decisions (Prime, 2026-09-27): runs on fv-ml1 GPU 1 under a hard VRAM cap; built by infra-ops with a light process (this contract → TDD → heid bug hunt); no consumer is named yet.
Endpoints
Every POST takes and returns JSON. Every endpoint except GET /health requires
Authorization: Bearer <token>.
| method | path | body | success |
|---|---|---|---|
| GET | /health |
— | 200 {status: "ok", semif_commit, model, vram_cap_gib, max_tokens, max_decisions, workloads} |
| POST | /decide |
one SemIf row {id, state, question, options[2..16]} + optional workload |
200 the direct.score dict + optional calibrated |
| POST | /decide/shared |
{state, decisions: [{id, question, options}], workload?} |
200 {results: [...], timing: {...}} from score_shared |
options items are {id, description}, as in SemIf. state is a nonempty
string, object or array. For /decide/shared, each decision becomes a SemIf row
by adding the shared state.
Calibration. workload is optional. When it names an entry in the
calibration table, each result gains calibrated: {workload, temperature, probabilities} with softmax(option_logits / T). The argmax never changes. The
native probabilities and probability_status stay untouched. An unknown
workload is a 422. With no workload, no calibrated key appears.
Order averaging (0.1.3, Prime 2026-09-27). A decision (the /decide body, or an
entry in decisions) may set orderings, whose default is "none":
"rotations": the n cyclic shifts of the caller's option list, starting with the caller's order, so every option sits in every position exactly once."all": every permutation (n!), the caller's order first. It is allowed only when n ≤ 4, else 422.
Every ordering of every decision in the request becomes its own SemIf row
(id = "<id>#o<k>", same state and question, reordered options). All rows go
to the engine in ONE shared call, which includes /decide. The result for an
averaged decision is:
{id, option_ids (caller's order),
combined: {method, orderings: n, probabilities, top, agreement, spread: {option_id: [min, max]}},
orderings: [the SemIf result for each ordering, unchanged]}
probabilities: per ordering, log-softmax ofoption_logits; averaged per option id; renormalised; reported in the caller's order.agreement: the fraction of orderings whose top option equalscombined.top.spread: each option's min and max native probability across orderings.
Every expanded row counts toward SEMIF_MAX_DECISIONS. workload together with
orderings is a 422, because a temperature is fitted per method and none is
fitted on combined scores yet. Decisions without orderings keep the exact
pre-0.1.3 result shape. Averaging cancels any additive position bias exactly.
Measured by the 2026-09-27 spike: 3 rotations take SemIf's labelled sets from
78.6% to 87.7% accuracy, and unanimous agreement is 94.4% accurate.
Invariants
- INV-1 pass-through.
option_ids,probabilities,option_logits,prompt_sha256,prompt_version,modelandprobability_statusare exactly what SemIf returned. The wrapper never rewrites a score. - INV-2 one model, one inference at a time. The model is loaded at startup,
and a process-wide lock serialises every scorer call. The app runs as one worker.
Scorer calls run off the event loop, so
/healthanswers while one is in progress. - INV-3 fail-closed startup. Startup refuses to serve unless the model sits
on a CUDA device, torch's arch list includes the card's
sm_XY, and one warm-up decision scores.device=cpuis allowed only when set explicitly. - INV-4 VRAM cap. When
SEMIF_VRAM_CAP_GIBis set, the process is capped at that share of the card (torch.cuda.set_per_process_memory_fraction) before the model loads. An out-of-memory error during a request is a 503out_of_memory, followed bytorch.cuda.empty_cache(). The process stays up. After every scorer call, when the reserved memory exceeds the post-warm-up baseline by more than 512 MiB, the engine callsempty_cache(). A burst must not keep the card's shared headroom: on 2026-09-27 a 64-decision request left the process holding 12.6 GB, leaving scriberr 3.5 GB. Any other scorer failure exceptValueErroralso releases before it is reported: it is logged with its traceback, then re-raised unchained asScoringFailedaftergc.collect()+empty_cache()(bug hunt C5). ARuntimeErrorwhose message says "out of memory" (cuBLAS/cuDNN allocation failures) counts as an OOM → 503 (S9). An OOM with an empty message is reported as "CUDA out of memory" rather than crashing the handler (C4). - INV-5 no network at runtime. Weights come from the mounted HF cache at the
pinned revision. The entry point sets
HF_HUB_OFFLINE=1itself before torch or transformers load, so this holds outside the image too (S10). - INV-7 boundary-safe shared prefix (0.1.4). At load, the engine wraps SemIf's
module-global
semif_phase1.shared._state_prefix, whichscore_sharedcalls. The wrapper keeps only the leading tokens that upstream's prefix shares with a real full prompt for the same state (a probe row: evidence, then a placeholder criterion).- Why: upstream drops just one token at the state boundary. An object state whose
last value ends in
),;or}re-tokenises two tokens back once, "criterion"follows, so it was refused with 422 "The fixed state prefix does not match every full prompt" (found 2026-09-27). - Effect: each row still scores the same token sequence; only the prefill/suffix split
moves, costing a few tokens of sharing. An ordinary state keeps upstream's prefix
unchanged.
score_sharedstill checks every real row and fails closed. - Startup proves the fix is in effect, not just installed (heid bug hunt SKAL,
2026-09-27). Before the weights load, it refuses to start if
_state_prefixis missing or not callable, or ifscore_shareddoes not resolve it fromsemif_phase1.shared's globals (a SemIf bump that moves or re-exports it). After the warm-up, it checks two things. The wrapper must return upstream's exact prefix for the warm-up state, since a wrapper rendering the wrong prompt would silently drop all sharing. And a state upstream alone refuses,{"person_said": "ok :)"}, must score throughscore_shareditself; a hook bound before the patch would fail here. A reload wraps the original again rather than stacking wrappers.
- Why: upstream drops just one token at the state boundary. An object state whose
last value ends in
- INV-6 constant-time auth. Token comparison uses
hmac.compare_digest. The token is ≥ 32 characters of visible ASCII (33–126). Startup refuses anything else, because a CR, LF or NUL in the token can never arrive in a header (S2).
Limits and errors
SEMIF_MAX_TOKENS(default 4096) is passed to the scorers. A longer prompt is a 422, never truncated (SemIf raises).SEMIF_MAX_DECISIONS(default 64) capsdecisionsper shared request. It must hold 1..max entries, else 422.- Request body ≤
SEMIF_MAX_BODY_BYTES(default 1 MiB), else 413. The limit is checked before each chunk is kept, so no more than the limit is ever held (C2). A declaredContent-Lengthis trusted only if it is ASCII digits (S3). - Admission: at most
SEMIF_MAX_QUEUE(default 32) POSTs may be in progress, counting both queued and scoring. The next one is refused with 429busybefore its body is read (C6).
| status | code | when |
|---|---|---|
| 401 | unauthorized |
missing or wrong bearer |
| 413 | request_too_large |
body over the limit |
| 422 | invalid_request |
bad JSON shape, a SemIf ValueError (validation, token limit, tokenisation), unknown workload, too many decisions |
| 429 | busy |
SEMIF_MAX_QUEUE requests already in progress |
| 503 | out_of_memory |
CUDA OOM during scoring |
| 500 | scoring_failed |
any other scorer exception, or a failure while building the response from a scorer result (calibration, averaging): always the envelope, never a bare 500 (C3) |
The error body is {error: {code, message}}.
Configuration (env)
SEMIF_API_TOKEN (required), SEMIF_MODEL (default Qwen/Qwen3.5-4B),
SEMIF_REVISION (default the pinned SHA), SEMIF_DEVICE (default cuda),
SEMIF_VRAM_CAP_GIB, SEMIF_MAX_TOKENS, SEMIF_MAX_DECISIONS,
SEMIF_MAX_BODY_BYTES, SEMIF_MAX_QUEUE, SEMIF_CALIBRATION (path to a JSON
{workload: T}).
Startup validates every value and refuses a bad one with a ValueError naming
the variable (C1, S1, S8):
- the VRAM cap, when set, is finite and > 0 (
0used to mean uncapped); - every limit is an integer ≥ 1;
- each T is a finite number, not a bool, in [0.05, 20] (a tiny T overflowed to NaN, and the response then failed to render);
- the calibration file must exist and parse.
Tests (TDD, fake scorer: no torch, no model)
auth required on POSTs and not on /health; a short token is refused at startup;
/decide passes the row through and returns the scorer dict unchanged;
/decide/shared builds rows with the shared state and returns results + timing;
calibration adds calibrated and keeps the argmax; an unknown workload → 422; a
scorer ValueError → 422; the engine's OutOfMemory → 503; the torch engine,
against a fake torch, turns torch.cuda.OutOfMemoryError into an unchained
OutOfMemory and calls empty_cache() only after the failed call's tensors are freed
(found on the card: a chained exception kept 11.9 GiB allocated after the 503); after
a call, reserved memory over baseline + 512 MiB is released and at or under it is left
alone; any
other exception → 500; malformed
JSON or a wrong body shape → 422; too many decisions → 422; an oversized body → 413; requests
are serialised (two concurrent calls never overlap inside the scorer); /health
answers while a scorer call is blocked. Averaging: rotations sends n rows
in one shared call, each option once per position, with ids <id>#o<k>, and
cancels a position bias exactly; all sends n! rows and is 422 above 4 options;
agreement and spread are computed from the orderings; a mixed shared request
(averaged + plain) is one engine call, with results in request order and plain
results unchanged; expanded rows count toward the cap; workload + orderings
→ 422. Prefix (INV-7): against a tokenizer whose merge reaches two tokens back, the
wrapped prefix is a token-prefix of every row's full prompt and gives up exactly the
mismatched token; an ordinary state keeps the whole upstream prefix; load() installs
one wrapper however many times it runs, and refuses to start when _state_prefix is
gone or not callable, when score_shared binds it early or resolves its globals
elsewhere, or when the wrapper renders the wrong prompt. The fake score_shared is
compiled into the fake module, so that it resolves its globals as the real one does.
Acceptance (on fv-ml1, real model; not unit tests)
- Parity: our
/decideover SemIf'sauthored144against their committed torch predictions (top choice and max probability gap). - Noise floor: the same run twice (A-vs-A).
- Negative control: shuffled option descriptions must break agreement.
- Shared vs direct: the same rows agree within the A-vs-A floor.
- Speed: 21 binary criteria over one state, N ≥ 3, p50 + spread.
- VRAM: the peak at a 4096-token input sets
SEMIF_VRAM_CAP_GIB. - Boundary (0.1.4): states whose last value ends in
),;and}, as objects, are all answered by/decide/shared, with the same top choice as a string state holding the same text; parity (1) still holds.