Files
esh-pfi-infrastructure/services/semif-serve/semif-serve.contract.md
T
vh 77b8cb449c feat(semif): 0.1.3 — order averaging, fast kernels, bug-hunt hardening (Prime)
Order averaging (Prime, after the 739aa03 spike):
- A decision may set orderings: rotations|all (all only for <= 4 options). Every
  ordering goes to the engine in one shared batch.
- The reply keeps each native result and adds combined {probabilities (log-mean),
  top, agreement, spread}.
- Through the service on SemIf's labelled sets (252 rows): 78.6% -> 88.1%
  (group-bootstrap 95% CI +5.1..+14.3). Unanimous agreement is 94.5% accurate.

Fast kernels: flash-linear-attention 0.5.2 and causal-conv1d 1.7.0 are now the
default build. A/B on the empty GPU 3:
- parity with upstream went from 142/144 to 144/144;
- a ~2k-token /decide went from 169 to 92 ms server-side;
- short 3-rotation batches cost ~3-6 ms more.
triton builds a C shim at runtime, so the image carries gcc. Without it the
warm-up failed and startup failed closed.

Heid bug-hunt panel (4/4 arms, thread 01M3H3F4RR7XBP90KQ3A39H4SX), folded:
- Startup validation: VRAM cap 0 no longer means uncapped (C1); limits must be
  >= 1 (S1); the token must be visible ASCII (S2); the calibration file must
  exist and parse, with T in [0.05, 20] (S8, and C3's NaN leg).
- The body limit is checked before a chunk is kept, and a Unicode-digit
  Content-Length no longer crashes (C2, S3).
- Failures while building the response now get the 500 envelope (C3).
- 429 busy past SEMIF_MAX_QUEUE requests in progress (C6).
- The engine releases memory on every non-validation failure, unchained after
  gc; an empty OOM message is handled; 'out of memory' RuntimeErrors map to 503
  (C4, C5, S9).
- The entry point forces HF_HUB_OFFLINE (S10). README wording fixed (S5, S6).
- New guard tests close the gaps the arms' mutation grids exposed: early stop of
  the body read, a shared-route lock, calibration pass-through, the gc cycle,
  the exact caps, TorchEngine.load's arch and device checks, and the offline
  entry point.
86 tests.

Deployed on fv-ml1 GPU 1: parity 144/144, OOM and burst release verified, shared
capacity 63/51/26/16 rows at ~140/520/1960/3900 prefix tokens.
2026-09-27 03:27:15 -07:00

183 lines
9.9 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
title: semif-serve
kind: module-contract
status: draft
owner: infra-ops
created: 2026-09-27
depends_on:
- SemIf-OpenJev (MIT) at commit 23cf1f39fc9534fe81437200959b6dfc7106e45a, package semif_phase1
- Qwen/Qwen3.5-4B at revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a, BF16
---
# semif-serve: an HTTP wrapper around SemIf's direct and shared scorers
## Purpose
SemIf decides by reading the logits of the option letters after one forward pass.
It ships as a batch CLI only. semif-serve loads the model **once** and exposes the
two torch scorers over HTTP, so fleet callers can ask typed questions without
decoding. It adds nothing to the scoring itself: every score it returns is what
`semif_phase1.direct.score` or `semif_phase1.shared.score_shared` returned,
unchanged, and it adds an optional calibrated view next to it.
Operator decisions (Prime, 2026-09-27): runs on fv-ml1 GPU 1 under a hard VRAM
cap; built by infra-ops with a light process (this contract → TDD → heid bug
hunt); no consumer is named yet.
## Endpoints
Every POST takes and returns JSON. Every endpoint except `GET /health` requires
`Authorization: Bearer <token>`.
| method | path | body | success |
|---|---|---|---|
| GET | `/health` | — | 200 `{status: "ok", semif_commit, model, vram_cap_gib, max_tokens, max_decisions, workloads}` |
| POST | `/decide` | one SemIf row `{id, state, question, options[2..16]}` + optional `workload` | 200 the `direct.score` dict + optional `calibrated` |
| POST | `/decide/shared` | `{state, decisions: [{id, question, options}], workload?}` | 200 `{results: [...], timing: {...}}` from `score_shared` |
`options` items are `{id, description}`, as in SemIf. `state` is a nonempty
string, object or array. For `/decide/shared`, each decision becomes a SemIf row
by adding the shared `state`.
**Calibration.** `workload` is optional. When it names an entry in the
calibration table, each result gains `calibrated: {workload, temperature,
probabilities}` with `softmax(option_logits / T)`. The argmax never changes. The
native `probabilities` and `probability_status` stay untouched. An unknown
`workload` is a 422. With no `workload`, no `calibrated` key appears.
**Order averaging (0.1.3, Prime 2026-09-27).** A decision (the `/decide` body, or an
entry in `decisions`) may set `orderings`, whose default is `"none"`:
- `"rotations"`: the n cyclic shifts of the caller's option list, starting with
the caller's order, so every option sits in every position exactly once.
- `"all"`: every permutation (n!), the caller's order first. It is allowed only
when n ≤ 4, else 422.
Every ordering of every decision in the request becomes its own SemIf row
(`id` = `"<id>#o<k>"`, same state and question, reordered options). **All rows go
to the engine in ONE `shared` call**, which includes `/decide`. The result for an
averaged decision is:
```
{id, option_ids (caller's order),
combined: {method, orderings: n, probabilities, top, agreement, spread: {option_id: [min, max]}},
orderings: [the SemIf result for each ordering, unchanged]}
```
- `probabilities`: per ordering, log-softmax of `option_logits`; averaged per
option id; renormalised; reported in the caller's order.
- `agreement`: the fraction of orderings whose top option equals `combined.top`.
- `spread`: each option's min and max native probability across orderings.
Every expanded row counts toward `SEMIF_MAX_DECISIONS`. `workload` together with
`orderings` is a 422, because a temperature is fitted per method and none is
fitted on combined scores yet. Decisions without `orderings` keep the exact
pre-0.1.3 result shape. Averaging cancels any additive position bias exactly.
Measured by the 2026-09-27 spike: 3 rotations take SemIf's labelled sets from
78.6% to 87.7% accuracy, and unanimous agreement is 94.4% accurate.
## Invariants
- **INV-1 pass-through.** `option_ids`, `probabilities`, `option_logits`,
`prompt_sha256`, `prompt_version`, `model` and `probability_status` are exactly
what SemIf returned. The wrapper never rewrites a score.
- **INV-2 one model, one inference at a time.** The model is loaded at startup,
and a process-wide lock serialises every scorer call. The app runs as one worker.
Scorer calls run off the event loop, so `/health` answers while one is in progress.
- **INV-3 fail-closed startup.** Startup refuses to serve unless the model sits
on a CUDA device, torch's arch list includes the card's `sm_XY`, and one
warm-up decision scores. `device=cpu` is allowed only when set explicitly.
- **INV-4 VRAM cap.** When `SEMIF_VRAM_CAP_GIB` is set, the process is capped at
that share of the card (`torch.cuda.set_per_process_memory_fraction`) before
the model loads. An out-of-memory error during a request is a 503
`out_of_memory`, followed by `torch.cuda.empty_cache()`. The process stays up.
**After every scorer call**, when the reserved memory exceeds the post-warm-up
baseline by more than 512 MiB, the engine calls `empty_cache()`. A burst must not
keep the card's shared headroom: on 2026-09-27 a 64-decision request left the
process holding 12.6 GB, leaving scriberr 3.5 GB.
**Any other scorer failure except `ValueError`** also releases before it is
reported: it is logged with its traceback, then re-raised unchained as
`ScoringFailed` after `gc.collect()` + `empty_cache()` (bug hunt C5). A
`RuntimeError` whose message says "out of memory" (cuBLAS/cuDNN allocation
failures) counts as an OOM → 503 (S9). An OOM with an empty message is reported
as "CUDA out of memory" rather than crashing the handler (C4).
- **INV-5 no network at runtime.** Weights come from the mounted HF cache at the
pinned revision. The entry point sets `HF_HUB_OFFLINE=1` itself before torch or
transformers load, so this holds outside the image too (S10).
- **INV-6 constant-time auth.** Token comparison uses `hmac.compare_digest`. The
token is ≥ 32 characters of visible ASCII (33–126). Startup refuses anything
else, because a CR, LF or NUL in the token can never arrive in a header (S2).
## Limits and errors
- `SEMIF_MAX_TOKENS` (default 4096) is passed to the scorers. A longer prompt is a
422, never truncated (SemIf raises).
- `SEMIF_MAX_DECISIONS` (default 64) caps `decisions` per shared request. It must
hold 1..max entries, else 422.
- Request body ≤ `SEMIF_MAX_BODY_BYTES` (default 1 MiB), else 413. The limit is
checked **before** each chunk is kept, so no more than the limit is ever held
(C2). A declared `Content-Length` is trusted only if it is ASCII digits (S3).
- **Admission:** at most `SEMIF_MAX_QUEUE` (default 32) POSTs may be in progress,
counting both queued and scoring. The next one is refused with 429 `busy`
before its body is read (C6).
| status | code | when |
|---|---|---|
| 401 | `unauthorized` | missing or wrong bearer |
| 413 | `request_too_large` | body over the limit |
| 422 | `invalid_request` | bad JSON shape, a SemIf `ValueError` (validation, token limit, tokenisation), unknown workload, too many decisions |
| 429 | `busy` | `SEMIF_MAX_QUEUE` requests already in progress |
| 503 | `out_of_memory` | CUDA OOM during scoring |
| 500 | `scoring_failed` | any other scorer exception, **or a failure while building the response** from a scorer result (calibration, averaging): always the envelope, never a bare 500 (C3) |
The error body is `{error: {code, message}}`.
## Configuration (env)
`SEMIF_API_TOKEN` (required), `SEMIF_MODEL` (default `Qwen/Qwen3.5-4B`),
`SEMIF_REVISION` (default the pinned SHA), `SEMIF_DEVICE` (default `cuda`),
`SEMIF_VRAM_CAP_GIB`, `SEMIF_MAX_TOKENS`, `SEMIF_MAX_DECISIONS`,
`SEMIF_MAX_BODY_BYTES`, `SEMIF_MAX_QUEUE`, `SEMIF_CALIBRATION` (path to a JSON
`{workload: T}`).
Startup validates every value and refuses a bad one with a `ValueError` naming
the variable (C1, S1, S8):
- the VRAM cap, when set, is finite and > 0 (`0` used to mean uncapped);
- every limit is an integer ≥ 1;
- each T is a finite number, not a bool, in [0.05, 20] (a tiny T overflowed to
NaN, and the response then failed to render);
- the calibration file must exist and parse.
## Tests (TDD, fake scorer: no torch, no model)
auth required on POSTs and not on /health; a short token is refused at startup;
`/decide` passes the row through and returns the scorer dict unchanged;
`/decide/shared` builds rows with the shared state and returns results + timing;
calibration adds `calibrated` and keeps the argmax; an unknown workload → 422; a
scorer `ValueError` → 422; the engine's `OutOfMemory` → 503; the torch engine,
against a fake torch, turns `torch.cuda.OutOfMemoryError` into an **unchained**
`OutOfMemory` and calls `empty_cache()` only after the failed call's tensors are freed
(found on the card: a chained exception kept 11.9 GiB allocated after the 503); after
a call, reserved memory over baseline + 512 MiB is released and at or under it is left
alone; any
other exception → 500; malformed
JSON or a wrong body shape → 422; too many decisions → 422; an oversized body → 413; requests
are serialised (two concurrent calls never overlap inside the scorer); `/health`
answers while a scorer call is blocked. **Averaging:** `rotations` sends n rows
in one shared call, each option once per position, with ids `<id>#o<k>`, and
cancels a position bias exactly; `all` sends n! rows and is 422 above 4 options;
agreement and spread are computed from the orderings; a mixed shared request
(averaged + plain) is one engine call, with results in request order and plain
results unchanged; expanded rows count toward the cap; `workload` + `orderings`
→ 422.
## Acceptance (on fv-ml1, real model; not unit tests)
1. **Parity:** our `/decide` over SemIf's `authored144` against their committed
torch predictions (top choice and max probability gap).
2. **Noise floor:** the same run twice (A-vs-A).
3. **Negative control:** shuffled option descriptions must break agreement.
4. **Shared vs direct:** the same rows agree within the A-vs-A floor.
5. **Speed:** 21 binary criteria over one state, N ≥ 3, p50 + spread.
6. **VRAM:** the peak at a 4096-token input sets `SEMIF_VRAM_CAP_GIB`.