Files
esh-pfi-infrastructure/services/semif-serve/semif-serve.contract.md
T
vh 77b8cb449c feat(semif): 0.1.3 — order averaging, fast kernels, bug-hunt hardening (Prime)
Order averaging (Prime, after the 739aa03 spike):
- A decision may set orderings: rotations|all (all only for <= 4 options). Every
  ordering goes to the engine in one shared batch.
- The reply keeps each native result and adds combined {probabilities (log-mean),
  top, agreement, spread}.
- Through the service on SemIf's labelled sets (252 rows): 78.6% -> 88.1%
  (group-bootstrap 95% CI +5.1..+14.3). Unanimous agreement is 94.5% accurate.

Fast kernels: flash-linear-attention 0.5.2 and causal-conv1d 1.7.0 are now the
default build. A/B on the empty GPU 3:
- parity with upstream went from 142/144 to 144/144;
- a ~2k-token /decide went from 169 to 92 ms server-side;
- short 3-rotation batches cost ~3-6 ms more.
triton builds a C shim at runtime, so the image carries gcc. Without it the
warm-up failed and startup failed closed.

Heid bug-hunt panel (4/4 arms, thread 01M3H3F4RR7XBP90KQ3A39H4SX), folded:
- Startup validation: VRAM cap 0 no longer means uncapped (C1); limits must be
  >= 1 (S1); the token must be visible ASCII (S2); the calibration file must
  exist and parse, with T in [0.05, 20] (S8, and C3's NaN leg).
- The body limit is checked before a chunk is kept, and a Unicode-digit
  Content-Length no longer crashes (C2, S3).
- Failures while building the response now get the 500 envelope (C3).
- 429 busy past SEMIF_MAX_QUEUE requests in progress (C6).
- The engine releases memory on every non-validation failure, unchained after
  gc; an empty OOM message is handled; 'out of memory' RuntimeErrors map to 503
  (C4, C5, S9).
- The entry point forces HF_HUB_OFFLINE (S10). README wording fixed (S5, S6).
- New guard tests close the gaps the arms' mutation grids exposed: early stop of
  the body read, a shared-route lock, calibration pass-through, the gc cycle,
  the exact caps, TorchEngine.load's arch and device checks, and the offline
  entry point.
86 tests.

Deployed on fv-ml1 GPU 1: parity 144/144, OOM and burst release verified, shared
capacity 63/51/26/16 rows at ~140/520/1960/3900 prefix tokens.
2026-09-27 03:27:15 -07:00

9.9 KiB
Raw Blame History

title, kind, status, owner, created, depends_on
title kind status owner created depends_on
semif-serve module-contract draft infra-ops 2026-09-27
SemIf-OpenJev (MIT) at commit 23cf1f39fc9534fe81437200959b6dfc7106e45a, package semif_phase1
Qwen/Qwen3.5-4B at revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a, BF16

semif-serve: an HTTP wrapper around SemIf's direct and shared scorers

Purpose

SemIf decides by reading the logits of the option letters after one forward pass. It ships as a batch CLI only. semif-serve loads the model once and exposes the two torch scorers over HTTP, so fleet callers can ask typed questions without decoding. It adds nothing to the scoring itself: every score it returns is what semif_phase1.direct.score or semif_phase1.shared.score_shared returned, unchanged, and it adds an optional calibrated view next to it.

Operator decisions (Prime, 2026-09-27): runs on fv-ml1 GPU 1 under a hard VRAM cap; built by infra-ops with a light process (this contract → TDD → heid bug hunt); no consumer is named yet.

Endpoints

Every POST takes and returns JSON. Every endpoint except GET /health requires Authorization: Bearer <token>.

method path body success
GET /health — 200 {status: "ok", semif_commit, model, vram_cap_gib, max_tokens, max_decisions, workloads}
POST /decide one SemIf row {id, state, question, options[2..16]} + optional workload 200 the direct.score dict + optional calibrated
POST /decide/shared {state, decisions: [{id, question, options}], workload?} 200 {results: [...], timing: {...}} from score_shared

options items are {id, description}, as in SemIf. state is a nonempty string, object or array. For /decide/shared, each decision becomes a SemIf row by adding the shared state.

Calibration. workload is optional. When it names an entry in the calibration table, each result gains calibrated: {workload, temperature, probabilities} with softmax(option_logits / T). The argmax never changes. The native probabilities and probability_status stay untouched. An unknown workload is a 422. With no workload, no calibrated key appears.

Order averaging (0.1.3, Prime 2026-09-27). A decision (the /decide body, or an entry in decisions) may set orderings, whose default is "none":

  • "rotations": the n cyclic shifts of the caller's option list, starting with the caller's order, so every option sits in every position exactly once.
  • "all": every permutation (n!), the caller's order first. It is allowed only when n ≤ 4, else 422.

Every ordering of every decision in the request becomes its own SemIf row (id = "<id>#o<k>", same state and question, reordered options). All rows go to the engine in ONE shared call, which includes /decide. The result for an averaged decision is:

{id, option_ids (caller's order),
 combined: {method, orderings: n, probabilities, top, agreement, spread: {option_id: [min, max]}},
 orderings: [the SemIf result for each ordering, unchanged]}
  • probabilities: per ordering, log-softmax of option_logits; averaged per option id; renormalised; reported in the caller's order.
  • agreement: the fraction of orderings whose top option equals combined.top.
  • spread: each option's min and max native probability across orderings.

Every expanded row counts toward SEMIF_MAX_DECISIONS. workload together with orderings is a 422, because a temperature is fitted per method and none is fitted on combined scores yet. Decisions without orderings keep the exact pre-0.1.3 result shape. Averaging cancels any additive position bias exactly. Measured by the 2026-09-27 spike: 3 rotations take SemIf's labelled sets from 78.6% to 87.7% accuracy, and unanimous agreement is 94.4% accurate.

Invariants

  • INV-1 pass-through. option_ids, probabilities, option_logits, prompt_sha256, prompt_version, model and probability_status are exactly what SemIf returned. The wrapper never rewrites a score.
  • INV-2 one model, one inference at a time. The model is loaded at startup, and a process-wide lock serialises every scorer call. The app runs as one worker. Scorer calls run off the event loop, so /health answers while one is in progress.
  • INV-3 fail-closed startup. Startup refuses to serve unless the model sits on a CUDA device, torch's arch list includes the card's sm_XY, and one warm-up decision scores. device=cpu is allowed only when set explicitly.
  • INV-4 VRAM cap. When SEMIF_VRAM_CAP_GIB is set, the process is capped at that share of the card (torch.cuda.set_per_process_memory_fraction) before the model loads. An out-of-memory error during a request is a 503 out_of_memory, followed by torch.cuda.empty_cache(). The process stays up. After every scorer call, when the reserved memory exceeds the post-warm-up baseline by more than 512 MiB, the engine calls empty_cache(). A burst must not keep the card's shared headroom: on 2026-09-27 a 64-decision request left the process holding 12.6 GB, leaving scriberr 3.5 GB. Any other scorer failure except ValueError also releases before it is reported: it is logged with its traceback, then re-raised unchained as ScoringFailed after gc.collect() + empty_cache() (bug hunt C5). A RuntimeError whose message says "out of memory" (cuBLAS/cuDNN allocation failures) counts as an OOM → 503 (S9). An OOM with an empty message is reported as "CUDA out of memory" rather than crashing the handler (C4).
  • INV-5 no network at runtime. Weights come from the mounted HF cache at the pinned revision. The entry point sets HF_HUB_OFFLINE=1 itself before torch or transformers load, so this holds outside the image too (S10).
  • INV-6 constant-time auth. Token comparison uses hmac.compare_digest. The token is ≥ 32 characters of visible ASCII (33–126). Startup refuses anything else, because a CR, LF or NUL in the token can never arrive in a header (S2).

Limits and errors

  • SEMIF_MAX_TOKENS (default 4096) is passed to the scorers. A longer prompt is a 422, never truncated (SemIf raises).
  • SEMIF_MAX_DECISIONS (default 64) caps decisions per shared request. It must hold 1..max entries, else 422.
  • Request body ≤ SEMIF_MAX_BODY_BYTES (default 1 MiB), else 413. The limit is checked before each chunk is kept, so no more than the limit is ever held (C2). A declared Content-Length is trusted only if it is ASCII digits (S3).
  • Admission: at most SEMIF_MAX_QUEUE (default 32) POSTs may be in progress, counting both queued and scoring. The next one is refused with 429 busy before its body is read (C6).
status code when
401 unauthorized missing or wrong bearer
413 request_too_large body over the limit
422 invalid_request bad JSON shape, a SemIf ValueError (validation, token limit, tokenisation), unknown workload, too many decisions
429 busy SEMIF_MAX_QUEUE requests already in progress
503 out_of_memory CUDA OOM during scoring
500 scoring_failed any other scorer exception, or a failure while building the response from a scorer result (calibration, averaging): always the envelope, never a bare 500 (C3)

The error body is {error: {code, message}}.

Configuration (env)

SEMIF_API_TOKEN (required), SEMIF_MODEL (default Qwen/Qwen3.5-4B), SEMIF_REVISION (default the pinned SHA), SEMIF_DEVICE (default cuda), SEMIF_VRAM_CAP_GIB, SEMIF_MAX_TOKENS, SEMIF_MAX_DECISIONS, SEMIF_MAX_BODY_BYTES, SEMIF_MAX_QUEUE, SEMIF_CALIBRATION (path to a JSON {workload: T}).

Startup validates every value and refuses a bad one with a ValueError naming the variable (C1, S1, S8):

  • the VRAM cap, when set, is finite and > 0 (0 used to mean uncapped);
  • every limit is an integer ≥ 1;
  • each T is a finite number, not a bool, in [0.05, 20] (a tiny T overflowed to NaN, and the response then failed to render);
  • the calibration file must exist and parse.

Tests (TDD, fake scorer: no torch, no model)

auth required on POSTs and not on /health; a short token is refused at startup; /decide passes the row through and returns the scorer dict unchanged; /decide/shared builds rows with the shared state and returns results + timing; calibration adds calibrated and keeps the argmax; an unknown workload → 422; a scorer ValueError → 422; the engine's OutOfMemory → 503; the torch engine, against a fake torch, turns torch.cuda.OutOfMemoryError into an unchained OutOfMemory and calls empty_cache() only after the failed call's tensors are freed (found on the card: a chained exception kept 11.9 GiB allocated after the 503); after a call, reserved memory over baseline + 512 MiB is released and at or under it is left alone; any other exception → 500; malformed JSON or a wrong body shape → 422; too many decisions → 422; an oversized body → 413; requests are serialised (two concurrent calls never overlap inside the scorer); /health answers while a scorer call is blocked. Averaging: rotations sends n rows in one shared call, each option once per position, with ids <id>#o<k>, and cancels a position bias exactly; all sends n! rows and is 422 above 4 options; agreement and spread are computed from the orderings; a mixed shared request (averaged + plain) is one engine call, with results in request order and plain results unchanged; expanded rows count toward the cap; workload + orderings → 422.

Acceptance (on fv-ml1, real model; not unit tests)

  1. Parity: our /decide over SemIf's authored144 against their committed torch predictions (top choice and max probability gap).
  2. Noise floor: the same run twice (A-vs-A).
  3. Negative control: shuffled option descriptions must break agreement.
  4. Shared vs direct: the same rows agree within the A-vs-A floor.
  5. Speed: 21 binary criteria over one state, N ≥ 3, p50 + spread.
  6. VRAM: the peak at a 4096-token input sets SEMIF_VRAM_CAP_GIB.