SemIf's shared scorer trims one token at the state boundary. When an object state's last value ends in ')', ';' or '}', the JSON that follows re-merges two tokens back, so score_shared refused the request with 422. The engine now wraps semif_phase1.shared._state_prefix to keep only the tokens the full prompts share. Each row scores the same token sequence; only the prefill/suffix split moves. Startup proves the fix is in effect, not just installed (heid bug hunt SKAL, folded). It checks that the hook is callable and is what score_shared resolves, that an ordinary state keeps upstream's whole prefix, and that a merge-prone state scores through the shared path. Real tokenizer: 154 states, 23 refused before and 0 after, with no ordinary or authored144 prefix changed. Acceptance: 144/144 parity. Shared vs direct 71/72; the miss is a bf16 tie that flipped across a plain restart (see README).
212 lines
12 KiB
Markdown
212 lines
12 KiB
Markdown
---
|
||
title: semif-serve
|
||
kind: module-contract
|
||
status: draft
|
||
owner: infra-ops
|
||
created: 2026-09-27
|
||
depends_on:
|
||
- SemIf-OpenJev (MIT) at commit 23cf1f39fc9534fe81437200959b6dfc7106e45a, package semif_phase1
|
||
- Qwen/Qwen3.5-4B at revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a, BF16
|
||
---
|
||
|
||
# semif-serve: an HTTP wrapper around SemIf's direct and shared scorers
|
||
|
||
## Purpose
|
||
|
||
SemIf decides by reading the logits of the option letters after one forward pass.
|
||
It ships as a batch CLI only. semif-serve loads the model **once** and exposes the
|
||
two torch scorers over HTTP, so fleet callers can ask typed questions without
|
||
decoding. It adds nothing to the scoring itself: every score it returns is what
|
||
`semif_phase1.direct.score` or `semif_phase1.shared.score_shared` returned,
|
||
unchanged, and it adds an optional calibrated view next to it.
|
||
|
||
Operator decisions (Prime, 2026-09-27): runs on fv-ml1 GPU 1 under a hard VRAM
|
||
cap; built by infra-ops with a light process (this contract → TDD → heid bug
|
||
hunt); no consumer is named yet.
|
||
|
||
## Endpoints
|
||
|
||
Every POST takes and returns JSON. Every endpoint except `GET /health` requires
|
||
`Authorization: Bearer <token>`.
|
||
|
||
| method | path | body | success |
|
||
|---|---|---|---|
|
||
| GET | `/health` | — | 200 `{status: "ok", semif_commit, model, vram_cap_gib, max_tokens, max_decisions, workloads}` |
|
||
| POST | `/decide` | one SemIf row `{id, state, question, options[2..16]}` + optional `workload` | 200 the `direct.score` dict + optional `calibrated` |
|
||
| POST | `/decide/shared` | `{state, decisions: [{id, question, options}], workload?}` | 200 `{results: [...], timing: {...}}` from `score_shared` |
|
||
|
||
`options` items are `{id, description}`, as in SemIf. `state` is a nonempty
|
||
string, object or array. For `/decide/shared`, each decision becomes a SemIf row
|
||
by adding the shared `state`.
|
||
|
||
**Calibration.** `workload` is optional. When it names an entry in the
|
||
calibration table, each result gains `calibrated: {workload, temperature,
|
||
probabilities}` with `softmax(option_logits / T)`. The argmax never changes. The
|
||
native `probabilities` and `probability_status` stay untouched. An unknown
|
||
`workload` is a 422. With no `workload`, no `calibrated` key appears.
|
||
|
||
**Order averaging (0.1.3, Prime 2026-09-27).** A decision (the `/decide` body, or an
|
||
entry in `decisions`) may set `orderings`, whose default is `"none"`:
|
||
|
||
- `"rotations"`: the n cyclic shifts of the caller's option list, starting with
|
||
the caller's order, so every option sits in every position exactly once.
|
||
- `"all"`: every permutation (n!), the caller's order first. It is allowed only
|
||
when n ≤ 4, else 422.
|
||
|
||
Every ordering of every decision in the request becomes its own SemIf row
|
||
(`id` = `"<id>#o<k>"`, same state and question, reordered options). **All rows go
|
||
to the engine in ONE `shared` call**, which includes `/decide`. The result for an
|
||
averaged decision is:
|
||
|
||
```
|
||
{id, option_ids (caller's order),
|
||
combined: {method, orderings: n, probabilities, top, agreement, spread: {option_id: [min, max]}},
|
||
orderings: [the SemIf result for each ordering, unchanged]}
|
||
```
|
||
|
||
- `probabilities`: per ordering, log-softmax of `option_logits`; averaged per
|
||
option id; renormalised; reported in the caller's order.
|
||
- `agreement`: the fraction of orderings whose top option equals `combined.top`.
|
||
- `spread`: each option's min and max native probability across orderings.
|
||
|
||
Every expanded row counts toward `SEMIF_MAX_DECISIONS`. `workload` together with
|
||
`orderings` is a 422, because a temperature is fitted per method and none is
|
||
fitted on combined scores yet. Decisions without `orderings` keep the exact
|
||
pre-0.1.3 result shape. Averaging cancels any additive position bias exactly.
|
||
Measured by the 2026-09-27 spike: 3 rotations take SemIf's labelled sets from
|
||
78.6% to 87.7% accuracy, and unanimous agreement is 94.4% accurate.
|
||
|
||
## Invariants
|
||
|
||
- **INV-1 pass-through.** `option_ids`, `probabilities`, `option_logits`,
|
||
`prompt_sha256`, `prompt_version`, `model` and `probability_status` are exactly
|
||
what SemIf returned. The wrapper never rewrites a score.
|
||
- **INV-2 one model, one inference at a time.** The model is loaded at startup,
|
||
and a process-wide lock serialises every scorer call. The app runs as one worker.
|
||
Scorer calls run off the event loop, so `/health` answers while one is in progress.
|
||
- **INV-3 fail-closed startup.** Startup refuses to serve unless the model sits
|
||
on a CUDA device, torch's arch list includes the card's `sm_XY`, and one
|
||
warm-up decision scores. `device=cpu` is allowed only when set explicitly.
|
||
- **INV-4 VRAM cap.** When `SEMIF_VRAM_CAP_GIB` is set, the process is capped at
|
||
that share of the card (`torch.cuda.set_per_process_memory_fraction`) before
|
||
the model loads. An out-of-memory error during a request is a 503
|
||
`out_of_memory`, followed by `torch.cuda.empty_cache()`. The process stays up.
|
||
**After every scorer call**, when the reserved memory exceeds the post-warm-up
|
||
baseline by more than 512 MiB, the engine calls `empty_cache()`. A burst must not
|
||
keep the card's shared headroom: on 2026-09-27 a 64-decision request left the
|
||
process holding 12.6 GB, leaving scriberr 3.5 GB.
|
||
**Any other scorer failure except `ValueError`** also releases before it is
|
||
reported: it is logged with its traceback, then re-raised unchained as
|
||
`ScoringFailed` after `gc.collect()` + `empty_cache()` (bug hunt C5). A
|
||
`RuntimeError` whose message says "out of memory" (cuBLAS/cuDNN allocation
|
||
failures) counts as an OOM → 503 (S9). An OOM with an empty message is reported
|
||
as "CUDA out of memory" rather than crashing the handler (C4).
|
||
- **INV-5 no network at runtime.** Weights come from the mounted HF cache at the
|
||
pinned revision. The entry point sets `HF_HUB_OFFLINE=1` itself before torch or
|
||
transformers load, so this holds outside the image too (S10).
|
||
- **INV-7 boundary-safe shared prefix (0.1.4).** At load, the engine wraps SemIf's
|
||
module-global `semif_phase1.shared._state_prefix`, which `score_shared` calls. The wrapper
|
||
keeps only the leading tokens that upstream's prefix shares with a real full prompt for
|
||
the same state (a probe row: evidence, then a placeholder criterion).
|
||
- **Why:** upstream drops just one token at the state boundary. An object state whose
|
||
last value ends in `)`, `;` or `}` re-tokenises two tokens back once `, "criterion"`
|
||
follows, so it was refused with 422 "The fixed state prefix does not match every full
|
||
prompt" (found 2026-09-27).
|
||
- **Effect:** each row still scores the same token sequence; only the prefill/suffix split
|
||
moves, costing a few tokens of sharing. An ordinary state keeps upstream's prefix
|
||
unchanged. `score_shared` still checks every real row and fails closed.
|
||
- **Startup proves the fix is in effect, not just installed** (heid bug hunt SKAL,
|
||
2026-09-27). Before the weights load, it refuses to start if `_state_prefix` is
|
||
missing or not callable, or if `score_shared` does not resolve it from
|
||
`semif_phase1.shared`'s globals (a SemIf bump that moves or re-exports it). After the
|
||
warm-up, it checks two things. The wrapper must return upstream's exact prefix for
|
||
the warm-up state, since a wrapper rendering the wrong prompt would silently drop all
|
||
sharing. And a state upstream alone refuses, `{"person_said": "ok :)"}`, must score
|
||
through `score_shared` itself; a hook bound before the patch would fail here. A
|
||
reload wraps the original again rather than stacking wrappers.
|
||
- **INV-6 constant-time auth.** Token comparison uses `hmac.compare_digest`. The
|
||
token is ≥ 32 characters of visible ASCII (33–126). Startup refuses anything
|
||
else, because a CR, LF or NUL in the token can never arrive in a header (S2).
|
||
|
||
## Limits and errors
|
||
|
||
- `SEMIF_MAX_TOKENS` (default 4096) is passed to the scorers. A longer prompt is a
|
||
422, never truncated (SemIf raises).
|
||
- `SEMIF_MAX_DECISIONS` (default 64) caps `decisions` per shared request. It must
|
||
hold 1..max entries, else 422.
|
||
- Request body ≤ `SEMIF_MAX_BODY_BYTES` (default 1 MiB), else 413. The limit is
|
||
checked **before** each chunk is kept, so no more than the limit is ever held
|
||
(C2). A declared `Content-Length` is trusted only if it is ASCII digits (S3).
|
||
- **Admission:** at most `SEMIF_MAX_QUEUE` (default 32) POSTs may be in progress,
|
||
counting both queued and scoring. The next one is refused with 429 `busy`
|
||
before its body is read (C6).
|
||
|
||
| status | code | when |
|
||
|---|---|---|
|
||
| 401 | `unauthorized` | missing or wrong bearer |
|
||
| 413 | `request_too_large` | body over the limit |
|
||
| 422 | `invalid_request` | bad JSON shape, a SemIf `ValueError` (validation, token limit, tokenisation), unknown workload, too many decisions |
|
||
| 429 | `busy` | `SEMIF_MAX_QUEUE` requests already in progress |
|
||
| 503 | `out_of_memory` | CUDA OOM during scoring |
|
||
| 500 | `scoring_failed` | any other scorer exception, **or a failure while building the response** from a scorer result (calibration, averaging): always the envelope, never a bare 500 (C3) |
|
||
|
||
The error body is `{error: {code, message}}`.
|
||
|
||
## Configuration (env)
|
||
|
||
`SEMIF_API_TOKEN` (required), `SEMIF_MODEL` (default `Qwen/Qwen3.5-4B`),
|
||
`SEMIF_REVISION` (default the pinned SHA), `SEMIF_DEVICE` (default `cuda`),
|
||
`SEMIF_VRAM_CAP_GIB`, `SEMIF_MAX_TOKENS`, `SEMIF_MAX_DECISIONS`,
|
||
`SEMIF_MAX_BODY_BYTES`, `SEMIF_MAX_QUEUE`, `SEMIF_CALIBRATION` (path to a JSON
|
||
`{workload: T}`).
|
||
|
||
Startup validates every value and refuses a bad one with a `ValueError` naming
|
||
the variable (C1, S1, S8):
|
||
- the VRAM cap, when set, is finite and > 0 (`0` used to mean uncapped);
|
||
- every limit is an integer ≥ 1;
|
||
- each T is a finite number, not a bool, in [0.05, 20] (a tiny T overflowed to
|
||
NaN, and the response then failed to render);
|
||
- the calibration file must exist and parse.
|
||
|
||
## Tests (TDD, fake scorer: no torch, no model)
|
||
|
||
auth required on POSTs and not on /health; a short token is refused at startup;
|
||
`/decide` passes the row through and returns the scorer dict unchanged;
|
||
`/decide/shared` builds rows with the shared state and returns results + timing;
|
||
calibration adds `calibrated` and keeps the argmax; an unknown workload → 422; a
|
||
scorer `ValueError` → 422; the engine's `OutOfMemory` → 503; the torch engine,
|
||
against a fake torch, turns `torch.cuda.OutOfMemoryError` into an **unchained**
|
||
`OutOfMemory` and calls `empty_cache()` only after the failed call's tensors are freed
|
||
(found on the card: a chained exception kept 11.9 GiB allocated after the 503); after
|
||
a call, reserved memory over baseline + 512 MiB is released and at or under it is left
|
||
alone; any
|
||
other exception → 500; malformed
|
||
JSON or a wrong body shape → 422; too many decisions → 422; an oversized body → 413; requests
|
||
are serialised (two concurrent calls never overlap inside the scorer); `/health`
|
||
answers while a scorer call is blocked. **Averaging:** `rotations` sends n rows
|
||
in one shared call, each option once per position, with ids `<id>#o<k>`, and
|
||
cancels a position bias exactly; `all` sends n! rows and is 422 above 4 options;
|
||
agreement and spread are computed from the orderings; a mixed shared request
|
||
(averaged + plain) is one engine call, with results in request order and plain
|
||
results unchanged; expanded rows count toward the cap; `workload` + `orderings`
|
||
→ 422. **Prefix (INV-7):** against a tokenizer whose merge reaches two tokens back, the
|
||
wrapped prefix is a token-prefix of every row's full prompt and gives up exactly the
|
||
mismatched token; an ordinary state keeps the whole upstream prefix; load() installs
|
||
one wrapper however many times it runs, and refuses to start when `_state_prefix` is
|
||
gone or not callable, when `score_shared` binds it early or resolves its globals
|
||
elsewhere, or when the wrapper renders the wrong prompt. The fake `score_shared` is
|
||
compiled into the fake module, so that it resolves its globals as the real one does.
|
||
|
||
## Acceptance (on fv-ml1, real model; not unit tests)
|
||
|
||
1. **Parity:** our `/decide` over SemIf's `authored144` against their committed
|
||
torch predictions (top choice and max probability gap).
|
||
2. **Noise floor:** the same run twice (A-vs-A).
|
||
3. **Negative control:** shuffled option descriptions must break agreement.
|
||
4. **Shared vs direct:** the same rows agree within the A-vs-A floor.
|
||
5. **Speed:** 21 binary criteria over one state, N ≥ 3, p50 + spread.
|
||
6. **VRAM:** the peak at a 4096-token input sets `SEMIF_VRAM_CAP_GIB`.
|
||
7. **Boundary (0.1.4):** states whose last value ends in `)`, `;` and `}`, as objects,
|
||
are all answered by `/decide/shared`, with the same top choice as a string state
|
||
holding the same text; parity (1) still holds.
|