Files
esh-pfi-infrastructure/services/semif-serve/semif-serve.contract.md
T
vh 069725c4b3 feat(semif): SemIf option-logit decisions on fv-ml1 GPU 1 (Prime)
services/semif-serve is a FastAPI wrapper around SemIf's direct and shared torch
scorers (SemIf-OpenJev @ 23cf1f39, MIT). Upstream ships only a batch CLI. The
wrapper loads the pinned Qwen3.5-4B (851bf6e8, BF16) once from the offline HF
cache and returns SemIf's result dicts unchanged, with an optional per-workload
temperature-calibrated view. Contract: semif-serve.contract.md. Built with a
short contract, TDD (39 tests, fake engine and fake torch, no GPU) and a heid
bug-hunt panel (pending).

On the card:
- torch 2.10.0+cu128 with sm_120 kernels, which is SemIf's own stack;
- a hard 12 GiB VRAM cap.
Two defects surfaced only on the card, and each fix is covered by a test:
- 0.1.1: an OOM raised as a chained exception kept the failed request's tensors
  alive (11.9 GiB after the 503). It is now raised unchained, after gc.
- 0.1.2: a large request left 12.6 GB reserved on the shared card. After each
  call, reserved memory over the baseline + 512 MiB is now released.

Acceptance against SemIf's committed torch predictions (authored144):
- 142/144 same top choice; both misses are exact bf16 ties;
- 144/144 identical prompt hashes;
- deterministic A-vs-A;
- negative control 14/144;
- shared vs direct 72/72.
21 binary criteria over one state take 159 ms. The shared-mode capacity table
under the cap is in stacks/semif/README.md.

The Dockerfile installs dependencies from a manifest with the project version
blanked, so a version bump reuses the ~4 GB torch layer. Verified: 41 s rebuild,
dependency layer CACHED.

DNS: semif.fv.internal. Token: vault semif/api-token.
2026-09-27 02:36:56 -07:00

6.3 KiB

title, kind, status, owner, created, depends_on
title kind status owner created depends_on
semif-serve module-contract draft infra-ops 2026-09-27
SemIf-OpenJev (MIT) at commit 23cf1f39fc9534fe81437200959b6dfc7106e45a, package semif_phase1
Qwen/Qwen3.5-4B at revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a, BF16

semif-serve: an HTTP wrapper around SemIf's direct and shared scorers

Purpose

SemIf decides by reading the logits of the option letters after one forward pass. It ships as a batch CLI only. semif-serve loads the model once and exposes the two torch scorers over HTTP, so fleet callers can ask typed questions without decoding. It adds nothing to the scoring itself: every score it returns is what semif_phase1.direct.score or semif_phase1.shared.score_shared returned, unchanged, and it adds an optional calibrated view next to it.

Operator decisions (Prime, 2026-09-27): runs on fv-ml1 GPU 1 under a hard VRAM cap; built by infra-ops with a light process (this contract → TDD → heid bug hunt); no consumer is named yet.

Endpoints

Every POST takes and returns JSON. Every endpoint except GET /health requires Authorization: Bearer <token>.

method path body success
GET /health — 200 {status: "ok", semif_commit, model, vram_cap_gib, max_tokens, max_decisions, workloads}
POST /decide one SemIf row {id, state, question, options[2..16]} + optional workload 200 the direct.score dict + optional calibrated
POST /decide/shared {state, decisions: [{id, question, options}], workload?} 200 {results: [...], timing: {...}} from score_shared

options items are {id, description}, as in SemIf. state is a nonempty string, object or array. For /decide/shared, each decision becomes a SemIf row by adding the shared state.

Calibration. workload is optional. When it names an entry in the calibration table, each result gains calibrated: {workload, temperature, probabilities} with softmax(option_logits / T). The argmax never changes. The native probabilities and probability_status stay untouched. An unknown workload is a 422. With no workload, no calibrated key appears.

Invariants

  • INV-1 pass-through. option_ids, probabilities, option_logits, prompt_sha256, prompt_version, model and probability_status are exactly what SemIf returned. The wrapper never rewrites a score.
  • INV-2 one model, one inference at a time. The model is loaded at startup, and a process-wide lock serialises every scorer call. The app runs as one worker. Scorer calls run off the event loop, so /health answers while one is in progress.
  • INV-3 fail-closed startup. Startup refuses to serve unless the model sits on a CUDA device, torch's arch list includes the card's sm_XY, and one warm-up decision scores. device=cpu is allowed only when set explicitly.
  • INV-4 VRAM cap. When SEMIF_VRAM_CAP_GIB is set, the process is capped at that share of the card (torch.cuda.set_per_process_memory_fraction) before the model loads. An out-of-memory error during a request is a 503 out_of_memory, followed by torch.cuda.empty_cache(). The process stays up. After every scorer call, when the reserved memory exceeds the post-warm-up baseline by more than 512 MiB, the engine calls empty_cache(). A burst must not keep the card's shared headroom: on 2026-09-27 a 64-decision request left the process holding 12.6 GB, leaving scriberr 3.5 GB.
  • INV-5 no network at runtime. Weights come from the mounted HF cache at the pinned revision (HF_HUB_OFFLINE=1).
  • INV-6 constant-time auth. Token comparison uses hmac.compare_digest. The token is ≥ 32 characters, and startup refuses a shorter one.

Limits and errors

  • SEMIF_MAX_TOKENS (default 4096) is passed to the scorers. A longer prompt is a 422, never truncated (SemIf raises).
  • SEMIF_MAX_DECISIONS (default 64) caps decisions per shared request. It must hold 1..max entries, else 422.
  • Request body ≤ SEMIF_MAX_BODY_BYTES (default 1 MiB), else 413.
status code when
401 unauthorized missing or wrong bearer
413 request_too_large body over the limit
422 invalid_request bad JSON shape, a SemIf ValueError (validation, token limit, tokenisation), unknown workload, too many decisions
503 out_of_memory CUDA OOM during scoring
500 scoring_failed any other scorer exception

The error body is {error: {code, message}}.

Configuration (env)

SEMIF_API_TOKEN (required), SEMIF_MODEL (default Qwen/Qwen3.5-4B), SEMIF_REVISION (default the pinned SHA), SEMIF_DEVICE (default cuda), SEMIF_VRAM_CAP_GIB, SEMIF_MAX_TOKENS, SEMIF_MAX_DECISIONS, SEMIF_MAX_BODY_BYTES, SEMIF_CALIBRATION (path to a JSON {workload: T}; T > 0).

Tests (TDD, fake scorer: no torch, no model)

auth required on POSTs and not on /health; a short token is refused at startup; /decide passes the row through and returns the scorer dict unchanged; /decide/shared builds rows with the shared state and returns results + timing; calibration adds calibrated and keeps the argmax; an unknown workload → 422; a scorer ValueError → 422; the engine's OutOfMemory → 503; the torch engine, against a fake torch, turns torch.cuda.OutOfMemoryError into an unchained OutOfMemory and calls empty_cache() only after the failed call's tensors are freed (found on the card: a chained exception kept 11.9 GiB allocated after the 503); after a call, reserved memory over baseline + 512 MiB is released and at or under it is left alone; any other exception → 500; malformed JSON or a wrong body shape → 422; too many decisions → 422; an oversized body → 413; requests are serialised (two concurrent calls never overlap inside the scorer); /health answers while a scorer call is blocked.

Acceptance (on fv-ml1, real model; not unit tests)

  1. Parity: our /decide over SemIf's authored144 against their committed torch predictions (top choice and max probability gap).
  2. Noise floor: the same run twice (A-vs-A).
  3. Negative control: shuffled option descriptions must break agreement.
  4. Shared vs direct: the same rows agree within the A-vs-A floor.
  5. Speed: 21 binary criteria over one state, N ≥ 3, p50 + spread.
  6. VRAM: the peak at a 4096-token input sets SEMIF_VRAM_CAP_GIB.