Files
esh-pfi-infrastructure/services/semif-serve/semif-serve.contract.md
T
vh 47cad33dd1 fix(semif): 0.1.4 — object states ending in ) ; } no longer 422 (INV-7)
SemIf's shared scorer trims one token at the state boundary. When an object
state's last value ends in ')', ';' or '}', the JSON that follows re-merges two
tokens back, so score_shared refused the request with 422. The engine now wraps
semif_phase1.shared._state_prefix to keep only the tokens the full prompts
share. Each row scores the same token sequence; only the prefill/suffix split
moves.

Startup proves the fix is in effect, not just installed (heid bug hunt SKAL,
folded). It checks that the hook is callable and is what score_shared resolves,
that an ordinary state keeps upstream's whole prefix, and that a merge-prone
state scores through the shared path.

Real tokenizer: 154 states, 23 refused before and 0 after, with no ordinary or
authored144 prefix changed. Acceptance: 144/144 parity. Shared vs direct 71/72;
the miss is a bf16 tie that flipped across a plain restart (see README).
2026-09-27 10:23:14 -07:00

212 lines
12 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
title: semif-serve
kind: module-contract
status: draft
owner: infra-ops
created: 2026-09-27
depends_on:
- SemIf-OpenJev (MIT) at commit 23cf1f39fc9534fe81437200959b6dfc7106e45a, package semif_phase1
- Qwen/Qwen3.5-4B at revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a, BF16
---
# semif-serve: an HTTP wrapper around SemIf's direct and shared scorers
## Purpose
SemIf decides by reading the logits of the option letters after one forward pass.
It ships as a batch CLI only. semif-serve loads the model **once** and exposes the
two torch scorers over HTTP, so fleet callers can ask typed questions without
decoding. It adds nothing to the scoring itself: every score it returns is what
`semif_phase1.direct.score` or `semif_phase1.shared.score_shared` returned,
unchanged, and it adds an optional calibrated view next to it.
Operator decisions (Prime, 2026-09-27): runs on fv-ml1 GPU 1 under a hard VRAM
cap; built by infra-ops with a light process (this contract → TDD → heid bug
hunt); no consumer is named yet.
## Endpoints
Every POST takes and returns JSON. Every endpoint except `GET /health` requires
`Authorization: Bearer <token>`.
| method | path | body | success |
|---|---|---|---|
| GET | `/health` | — | 200 `{status: "ok", semif_commit, model, vram_cap_gib, max_tokens, max_decisions, workloads}` |
| POST | `/decide` | one SemIf row `{id, state, question, options[2..16]}` + optional `workload` | 200 the `direct.score` dict + optional `calibrated` |
| POST | `/decide/shared` | `{state, decisions: [{id, question, options}], workload?}` | 200 `{results: [...], timing: {...}}` from `score_shared` |
`options` items are `{id, description}`, as in SemIf. `state` is a nonempty
string, object or array. For `/decide/shared`, each decision becomes a SemIf row
by adding the shared `state`.
**Calibration.** `workload` is optional. When it names an entry in the
calibration table, each result gains `calibrated: {workload, temperature,
probabilities}` with `softmax(option_logits / T)`. The argmax never changes. The
native `probabilities` and `probability_status` stay untouched. An unknown
`workload` is a 422. With no `workload`, no `calibrated` key appears.
**Order averaging (0.1.3, Prime 2026-09-27).** A decision (the `/decide` body, or an
entry in `decisions`) may set `orderings`, whose default is `"none"`:
- `"rotations"`: the n cyclic shifts of the caller's option list, starting with
the caller's order, so every option sits in every position exactly once.
- `"all"`: every permutation (n!), the caller's order first. It is allowed only
when n ≤ 4, else 422.
Every ordering of every decision in the request becomes its own SemIf row
(`id` = `"<id>#o<k>"`, same state and question, reordered options). **All rows go
to the engine in ONE `shared` call**, which includes `/decide`. The result for an
averaged decision is:
```
{id, option_ids (caller's order),
combined: {method, orderings: n, probabilities, top, agreement, spread: {option_id: [min, max]}},
orderings: [the SemIf result for each ordering, unchanged]}
```
- `probabilities`: per ordering, log-softmax of `option_logits`; averaged per
option id; renormalised; reported in the caller's order.
- `agreement`: the fraction of orderings whose top option equals `combined.top`.
- `spread`: each option's min and max native probability across orderings.
Every expanded row counts toward `SEMIF_MAX_DECISIONS`. `workload` together with
`orderings` is a 422, because a temperature is fitted per method and none is
fitted on combined scores yet. Decisions without `orderings` keep the exact
pre-0.1.3 result shape. Averaging cancels any additive position bias exactly.
Measured by the 2026-09-27 spike: 3 rotations take SemIf's labelled sets from
78.6% to 87.7% accuracy, and unanimous agreement is 94.4% accurate.
## Invariants
- **INV-1 pass-through.** `option_ids`, `probabilities`, `option_logits`,
`prompt_sha256`, `prompt_version`, `model` and `probability_status` are exactly
what SemIf returned. The wrapper never rewrites a score.
- **INV-2 one model, one inference at a time.** The model is loaded at startup,
and a process-wide lock serialises every scorer call. The app runs as one worker.
Scorer calls run off the event loop, so `/health` answers while one is in progress.
- **INV-3 fail-closed startup.** Startup refuses to serve unless the model sits
on a CUDA device, torch's arch list includes the card's `sm_XY`, and one
warm-up decision scores. `device=cpu` is allowed only when set explicitly.
- **INV-4 VRAM cap.** When `SEMIF_VRAM_CAP_GIB` is set, the process is capped at
that share of the card (`torch.cuda.set_per_process_memory_fraction`) before
the model loads. An out-of-memory error during a request is a 503
`out_of_memory`, followed by `torch.cuda.empty_cache()`. The process stays up.
**After every scorer call**, when the reserved memory exceeds the post-warm-up
baseline by more than 512 MiB, the engine calls `empty_cache()`. A burst must not
keep the card's shared headroom: on 2026-09-27 a 64-decision request left the
process holding 12.6 GB, leaving scriberr 3.5 GB.
**Any other scorer failure except `ValueError`** also releases before it is
reported: it is logged with its traceback, then re-raised unchained as
`ScoringFailed` after `gc.collect()` + `empty_cache()` (bug hunt C5). A
`RuntimeError` whose message says "out of memory" (cuBLAS/cuDNN allocation
failures) counts as an OOM → 503 (S9). An OOM with an empty message is reported
as "CUDA out of memory" rather than crashing the handler (C4).
- **INV-5 no network at runtime.** Weights come from the mounted HF cache at the
pinned revision. The entry point sets `HF_HUB_OFFLINE=1` itself before torch or
transformers load, so this holds outside the image too (S10).
- **INV-7 boundary-safe shared prefix (0.1.4).** At load, the engine wraps SemIf's
module-global `semif_phase1.shared._state_prefix`, which `score_shared` calls. The wrapper
keeps only the leading tokens that upstream's prefix shares with a real full prompt for
the same state (a probe row: evidence, then a placeholder criterion).
- **Why:** upstream drops just one token at the state boundary. An object state whose
last value ends in `)`, `;` or `}` re-tokenises two tokens back once `, "criterion"`
follows, so it was refused with 422 "The fixed state prefix does not match every full
prompt" (found 2026-09-27).
- **Effect:** each row still scores the same token sequence; only the prefill/suffix split
moves, costing a few tokens of sharing. An ordinary state keeps upstream's prefix
unchanged. `score_shared` still checks every real row and fails closed.
- **Startup proves the fix is in effect, not just installed** (heid bug hunt SKAL,
2026-09-27). Before the weights load, it refuses to start if `_state_prefix` is
missing or not callable, or if `score_shared` does not resolve it from
`semif_phase1.shared`'s globals (a SemIf bump that moves or re-exports it). After the
warm-up, it checks two things. The wrapper must return upstream's exact prefix for
the warm-up state, since a wrapper rendering the wrong prompt would silently drop all
sharing. And a state upstream alone refuses, `{"person_said": "ok :)"}`, must score
through `score_shared` itself; a hook bound before the patch would fail here. A
reload wraps the original again rather than stacking wrappers.
- **INV-6 constant-time auth.** Token comparison uses `hmac.compare_digest`. The
token is ≥ 32 characters of visible ASCII (33–126). Startup refuses anything
else, because a CR, LF or NUL in the token can never arrive in a header (S2).
## Limits and errors
- `SEMIF_MAX_TOKENS` (default 4096) is passed to the scorers. A longer prompt is a
422, never truncated (SemIf raises).
- `SEMIF_MAX_DECISIONS` (default 64) caps `decisions` per shared request. It must
hold 1..max entries, else 422.
- Request body ≤ `SEMIF_MAX_BODY_BYTES` (default 1 MiB), else 413. The limit is
checked **before** each chunk is kept, so no more than the limit is ever held
(C2). A declared `Content-Length` is trusted only if it is ASCII digits (S3).
- **Admission:** at most `SEMIF_MAX_QUEUE` (default 32) POSTs may be in progress,
counting both queued and scoring. The next one is refused with 429 `busy`
before its body is read (C6).
| status | code | when |
|---|---|---|
| 401 | `unauthorized` | missing or wrong bearer |
| 413 | `request_too_large` | body over the limit |
| 422 | `invalid_request` | bad JSON shape, a SemIf `ValueError` (validation, token limit, tokenisation), unknown workload, too many decisions |
| 429 | `busy` | `SEMIF_MAX_QUEUE` requests already in progress |
| 503 | `out_of_memory` | CUDA OOM during scoring |
| 500 | `scoring_failed` | any other scorer exception, **or a failure while building the response** from a scorer result (calibration, averaging): always the envelope, never a bare 500 (C3) |
The error body is `{error: {code, message}}`.
## Configuration (env)
`SEMIF_API_TOKEN` (required), `SEMIF_MODEL` (default `Qwen/Qwen3.5-4B`),
`SEMIF_REVISION` (default the pinned SHA), `SEMIF_DEVICE` (default `cuda`),
`SEMIF_VRAM_CAP_GIB`, `SEMIF_MAX_TOKENS`, `SEMIF_MAX_DECISIONS`,
`SEMIF_MAX_BODY_BYTES`, `SEMIF_MAX_QUEUE`, `SEMIF_CALIBRATION` (path to a JSON
`{workload: T}`).
Startup validates every value and refuses a bad one with a `ValueError` naming
the variable (C1, S1, S8):
- the VRAM cap, when set, is finite and > 0 (`0` used to mean uncapped);
- every limit is an integer ≥ 1;
- each T is a finite number, not a bool, in [0.05, 20] (a tiny T overflowed to
NaN, and the response then failed to render);
- the calibration file must exist and parse.
## Tests (TDD, fake scorer: no torch, no model)
auth required on POSTs and not on /health; a short token is refused at startup;
`/decide` passes the row through and returns the scorer dict unchanged;
`/decide/shared` builds rows with the shared state and returns results + timing;
calibration adds `calibrated` and keeps the argmax; an unknown workload → 422; a
scorer `ValueError` → 422; the engine's `OutOfMemory` → 503; the torch engine,
against a fake torch, turns `torch.cuda.OutOfMemoryError` into an **unchained**
`OutOfMemory` and calls `empty_cache()` only after the failed call's tensors are freed
(found on the card: a chained exception kept 11.9 GiB allocated after the 503); after
a call, reserved memory over baseline + 512 MiB is released and at or under it is left
alone; any
other exception → 500; malformed
JSON or a wrong body shape → 422; too many decisions → 422; an oversized body → 413; requests
are serialised (two concurrent calls never overlap inside the scorer); `/health`
answers while a scorer call is blocked. **Averaging:** `rotations` sends n rows
in one shared call, each option once per position, with ids `<id>#o<k>`, and
cancels a position bias exactly; `all` sends n! rows and is 422 above 4 options;
agreement and spread are computed from the orderings; a mixed shared request
(averaged + plain) is one engine call, with results in request order and plain
results unchanged; expanded rows count toward the cap; `workload` + `orderings`
→ 422. **Prefix (INV-7):** against a tokenizer whose merge reaches two tokens back, the
wrapped prefix is a token-prefix of every row's full prompt and gives up exactly the
mismatched token; an ordinary state keeps the whole upstream prefix; load() installs
one wrapper however many times it runs, and refuses to start when `_state_prefix` is
gone or not callable, when `score_shared` binds it early or resolves its globals
elsewhere, or when the wrapper renders the wrong prompt. The fake `score_shared` is
compiled into the fake module, so that it resolves its globals as the real one does.
## Acceptance (on fv-ml1, real model; not unit tests)
1. **Parity:** our `/decide` over SemIf's `authored144` against their committed
torch predictions (top choice and max probability gap).
2. **Noise floor:** the same run twice (A-vs-A).
3. **Negative control:** shuffled option descriptions must break agreement.
4. **Shared vs direct:** the same rows agree within the A-vs-A floor.
5. **Speed:** 21 binary criteria over one state, N ≥ 3, p50 + spread.
6. **VRAM:** the peak at a 4096-token input sets `SEMIF_VRAM_CAP_GIB`.
7. **Boundary (0.1.4):** states whose last value ends in `)`, `;` and `}`, as objects,
are all answered by `/decide/shared`, with the same top choice as a string state
holding the same text; parity (1) still holds.