feat(semif): SemIf option-logit decisions on fv-ml1 GPU 1 (Prime)
services/semif-serve is a FastAPI wrapper around SemIf's direct and shared torch scorers (SemIf-OpenJev @ 23cf1f39, MIT). Upstream ships only a batch CLI. The wrapper loads the pinned Qwen3.5-4B (851bf6e8, BF16) once from the offline HF cache and returns SemIf's result dicts unchanged, with an optional per-workload temperature-calibrated view. Contract: semif-serve.contract.md. Built with a short contract, TDD (39 tests, fake engine and fake torch, no GPU) and a heid bug-hunt panel (pending). On the card: - torch 2.10.0+cu128 with sm_120 kernels, which is SemIf's own stack; - a hard 12 GiB VRAM cap. Two defects surfaced only on the card, and each fix is covered by a test: - 0.1.1: an OOM raised as a chained exception kept the failed request's tensors alive (11.9 GiB after the 503). It is now raised unchained, after gc. - 0.1.2: a large request left 12.6 GB reserved on the shared card. After each call, reserved memory over the baseline + 512 MiB is now released. Acceptance against SemIf's committed torch predictions (authored144): - 142/144 same top choice; both misses are exact bf16 ties; - 144/144 identical prompt hashes; - deterministic A-vs-A; - negative control 14/144; - shared vs direct 72/72. 21 binary criteria over one state take 159 ms. The shared-mode capacity table under the cap is in stacks/semif/README.md. The Dockerfile installs dependencies from a manifest with the project version blanked, so a version bump reuses the ~4 GB torch layer. Verified: 41 s rebuild, dependency layer CACHED. DNS: semif.fv.internal. Token: vault semif/api-token.
This commit is contained in:
@@ -0,0 +1,122 @@
|
||||
---
|
||||
title: semif-serve
|
||||
kind: module-contract
|
||||
status: draft
|
||||
owner: infra-ops
|
||||
created: 2026-09-27
|
||||
depends_on:
|
||||
- SemIf-OpenJev (MIT) at commit 23cf1f39fc9534fe81437200959b6dfc7106e45a, package semif_phase1
|
||||
- Qwen/Qwen3.5-4B at revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a, BF16
|
||||
---
|
||||
|
||||
# semif-serve: an HTTP wrapper around SemIf's direct and shared scorers
|
||||
|
||||
## Purpose
|
||||
|
||||
SemIf decides by reading the logits of the option letters after one forward pass.
|
||||
It ships as a batch CLI only. semif-serve loads the model **once** and exposes the
|
||||
two torch scorers over HTTP, so fleet callers can ask typed questions without
|
||||
decoding. It adds nothing to the scoring itself: every score it returns is what
|
||||
`semif_phase1.direct.score` or `semif_phase1.shared.score_shared` returned,
|
||||
unchanged, and it adds an optional calibrated view next to it.
|
||||
|
||||
Operator decisions (Prime, 2026-09-27): runs on fv-ml1 GPU 1 under a hard VRAM
|
||||
cap; built by infra-ops with a light process (this contract → TDD → heid bug
|
||||
hunt); no consumer is named yet.
|
||||
|
||||
## Endpoints
|
||||
|
||||
Every POST takes and returns JSON. Every endpoint except `GET /health` requires
|
||||
`Authorization: Bearer <token>`.
|
||||
|
||||
| method | path | body | success |
|
||||
|---|---|---|---|
|
||||
| GET | `/health` | — | 200 `{status: "ok", semif_commit, model, vram_cap_gib, max_tokens, max_decisions, workloads}` |
|
||||
| POST | `/decide` | one SemIf row `{id, state, question, options[2..16]}` + optional `workload` | 200 the `direct.score` dict + optional `calibrated` |
|
||||
| POST | `/decide/shared` | `{state, decisions: [{id, question, options}], workload?}` | 200 `{results: [...], timing: {...}}` from `score_shared` |
|
||||
|
||||
`options` items are `{id, description}`, as in SemIf. `state` is a nonempty
|
||||
string, object or array. For `/decide/shared`, each decision becomes a SemIf row
|
||||
by adding the shared `state`.
|
||||
|
||||
**Calibration.** `workload` is optional. When it names an entry in the
|
||||
calibration table, each result gains `calibrated: {workload, temperature,
|
||||
probabilities}` with `softmax(option_logits / T)`. The argmax never changes. The
|
||||
native `probabilities` and `probability_status` stay untouched. An unknown
|
||||
`workload` is a 422. With no `workload`, no `calibrated` key appears.
|
||||
|
||||
## Invariants
|
||||
|
||||
- **INV-1 pass-through.** `option_ids`, `probabilities`, `option_logits`,
|
||||
`prompt_sha256`, `prompt_version`, `model` and `probability_status` are exactly
|
||||
what SemIf returned. The wrapper never rewrites a score.
|
||||
- **INV-2 one model, one inference at a time.** The model is loaded at startup,
|
||||
and a process-wide lock serialises every scorer call. The app runs as one worker.
|
||||
Scorer calls run off the event loop, so `/health` answers while one is in progress.
|
||||
- **INV-3 fail-closed startup.** Startup refuses to serve unless the model sits
|
||||
on a CUDA device, torch's arch list includes the card's `sm_XY`, and one
|
||||
warm-up decision scores. `device=cpu` is allowed only when set explicitly.
|
||||
- **INV-4 VRAM cap.** When `SEMIF_VRAM_CAP_GIB` is set, the process is capped at
|
||||
that share of the card (`torch.cuda.set_per_process_memory_fraction`) before
|
||||
the model loads. An out-of-memory error during a request is a 503
|
||||
`out_of_memory`, followed by `torch.cuda.empty_cache()`. The process stays up.
|
||||
**After every scorer call**, when the reserved memory exceeds the post-warm-up
|
||||
baseline by more than 512 MiB, the engine calls `empty_cache()`. A burst must not
|
||||
keep the card's shared headroom: on 2026-09-27 a 64-decision request left the
|
||||
process holding 12.6 GB, leaving scriberr 3.5 GB.
|
||||
- **INV-5 no network at runtime.** Weights come from the mounted HF cache at the
|
||||
pinned revision (`HF_HUB_OFFLINE=1`).
|
||||
- **INV-6 constant-time auth.** Token comparison uses `hmac.compare_digest`. The
|
||||
token is ≥ 32 characters, and startup refuses a shorter one.
|
||||
|
||||
## Limits and errors
|
||||
|
||||
- `SEMIF_MAX_TOKENS` (default 4096) is passed to the scorers. A longer prompt is a
|
||||
422, never truncated (SemIf raises).
|
||||
- `SEMIF_MAX_DECISIONS` (default 64) caps `decisions` per shared request. It must
|
||||
hold 1..max entries, else 422.
|
||||
- Request body ≤ `SEMIF_MAX_BODY_BYTES` (default 1 MiB), else 413.
|
||||
|
||||
| status | code | when |
|
||||
|---|---|---|
|
||||
| 401 | `unauthorized` | missing or wrong bearer |
|
||||
| 413 | `request_too_large` | body over the limit |
|
||||
| 422 | `invalid_request` | bad JSON shape, a SemIf `ValueError` (validation, token limit, tokenisation), unknown workload, too many decisions |
|
||||
| 503 | `out_of_memory` | CUDA OOM during scoring |
|
||||
| 500 | `scoring_failed` | any other scorer exception |
|
||||
|
||||
The error body is `{error: {code, message}}`.
|
||||
|
||||
## Configuration (env)
|
||||
|
||||
`SEMIF_API_TOKEN` (required), `SEMIF_MODEL` (default `Qwen/Qwen3.5-4B`),
|
||||
`SEMIF_REVISION` (default the pinned SHA), `SEMIF_DEVICE` (default `cuda`),
|
||||
`SEMIF_VRAM_CAP_GIB`, `SEMIF_MAX_TOKENS`, `SEMIF_MAX_DECISIONS`,
|
||||
`SEMIF_MAX_BODY_BYTES`, `SEMIF_CALIBRATION` (path to a JSON `{workload: T}`; T > 0).
|
||||
|
||||
## Tests (TDD, fake scorer: no torch, no model)
|
||||
|
||||
auth required on POSTs and not on /health; a short token is refused at startup;
|
||||
`/decide` passes the row through and returns the scorer dict unchanged;
|
||||
`/decide/shared` builds rows with the shared state and returns results + timing;
|
||||
calibration adds `calibrated` and keeps the argmax; an unknown workload → 422; a
|
||||
scorer `ValueError` → 422; the engine's `OutOfMemory` → 503; the torch engine,
|
||||
against a fake torch, turns `torch.cuda.OutOfMemoryError` into an **unchained**
|
||||
`OutOfMemory` and calls `empty_cache()` only after the failed call's tensors are freed
|
||||
(found on the card: a chained exception kept 11.9 GiB allocated after the 503); after
|
||||
a call, reserved memory over baseline + 512 MiB is released and at or under it is left
|
||||
alone; any
|
||||
other exception → 500; malformed
|
||||
JSON or a wrong body shape → 422; too many decisions → 422; an oversized body → 413; requests
|
||||
are serialised (two concurrent calls never overlap inside the scorer); `/health`
|
||||
answers while a scorer call is blocked.
|
||||
|
||||
## Acceptance (on fv-ml1, real model; not unit tests)
|
||||
|
||||
1. **Parity:** our `/decide` over SemIf's `authored144` against their committed
|
||||
torch predictions (top choice and max probability gap).
|
||||
2. **Noise floor:** the same run twice (A-vs-A).
|
||||
3. **Negative control:** shuffled option descriptions must break agreement.
|
||||
4. **Shared vs direct:** the same rows agree within the A-vs-A floor.
|
||||
5. **Speed:** 21 binary criteria over one state, N ≥ 3, p50 + spread.
|
||||
6. **VRAM:** the peak at a 4096-token input sets `SEMIF_VRAM_CAP_GIB`.
|
||||
Reference in New Issue
Block a user