Files
esh-pfi-infrastructure/stacks/semif/README.md
T
vh 069725c4b3 feat(semif): SemIf option-logit decisions on fv-ml1 GPU 1 (Prime)
services/semif-serve is a FastAPI wrapper around SemIf's direct and shared torch
scorers (SemIf-OpenJev @ 23cf1f39, MIT). Upstream ships only a batch CLI. The
wrapper loads the pinned Qwen3.5-4B (851bf6e8, BF16) once from the offline HF
cache and returns SemIf's result dicts unchanged, with an optional per-workload
temperature-calibrated view. Contract: semif-serve.contract.md. Built with a
short contract, TDD (39 tests, fake engine and fake torch, no GPU) and a heid
bug-hunt panel (pending).

On the card:
- torch 2.10.0+cu128 with sm_120 kernels, which is SemIf's own stack;
- a hard 12 GiB VRAM cap.
Two defects surfaced only on the card, and each fix is covered by a test:
- 0.1.1: an OOM raised as a chained exception kept the failed request's tensors
  alive (11.9 GiB after the 503). It is now raised unchained, after gc.
- 0.1.2: a large request left 12.6 GB reserved on the shared card. After each
  call, reserved memory over the baseline + 512 MiB is now released.

Acceptance against SemIf's committed torch predictions (authored144):
- 142/144 same top choice; both misses are exact bf16 ties;
- 144/144 identical prompt hashes;
- deterministic A-vs-A;
- negative control 14/144;
- shared vs direct 72/72.
21 binary criteria over one state take 159 ms. The shared-mode capacity table
under the cap is in stacks/semif/README.md.

The Dockerfile installs dependencies from a manifest with the project version
blanked, so a version bump reuses the ~4 GB torch layer. Verified: 41 s rebuild,
dependency layer CACHED.

DNS: semif.fv.internal. Token: vault semif/api-token.
2026-09-27 02:36:56 -07:00

104 lines
5.2 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# semif
**SemIf option-logit decisions** on **fv-ml1 GPU 1**, the utility card beside
`vllm-coder`, the erp/meromero seats and scriberr. Deployed 2026-09-27 at Prime's
request. Hosting assessment: infra-hermes, 2026-09-25.
SemIf ([TheoLeeCJ/SemIf-OpenJev](https://github.com/TheoLeeCJ/SemIf-OpenJev), MIT)
asks a small model a typed question and reads the answer straight from the logits of
the option letters, after **one forward pass with no decoding**. Upstream ships only a
batch CLI, so `services/semif-serve/` wraps its two torch scorers in a small FastAPI
service. The contract is `services/semif-serve/semif-serve.contract.md`.
| | |
|---|---|
| **URL** | `http://10.251.50.54:8032` (`/health` is open; POSTs need `Authorization: Bearer $(secret get semif/api-token)`) |
| **Model** | `Qwen/Qwen3.5-4B` @ `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a`, BF16, from `/tank/aimodels/huggingface` (read-only, offline) |
| **SemIf** | commit `23cf1f39fc9534fe81437200959b6dfc7106e45a`; torch `2.10.0+cu128`, transformers `5.17.0`, the same stack SemIf's committed predictions were made on |
| **Image** | `semif-serve:<version>`, built on fv-ml1 from `services/semif-serve/` |
| **State** | none. Weights re-download at the pinned revision. No backup needed beyond the host's `/opt/docker` restic. |
## API
```bash
T=$(secret get semif/api-token)
curl -s -H "Authorization: Bearer $T" http://10.251.50.54:8032/decide -d '{
"id": "q1", "state": "Health checks passed in all three zones.",
"question": "Did the deployment succeed?",
"options": [{"id":"yes","description":"It succeeded."},{"id":"no","description":"It failed."}]}'
```
- `POST /decide`: one decision, returning SemIf's result dict unchanged (`option_ids`,
`probabilities`, `option_logits`, `prompt_sha256`, `model`, …).
- `POST /decide/shared`: `{state, decisions: [{id, question, options}]}`. It prefills
the state once and scores every criterion in one batch. Use it when many questions
share one long state.
- `GET /health`: the pins, limits and calibrated workloads.
## ⚠ Probabilities are uncalibrated
SemIf labels its output `"conditional option score; uncalibrated as decision
confidence"`, and means it: on WANLI the model is right ~64% of the time while
reporting far higher confidence. **Before a caller thresholds on `probabilities`, it
brings labelled rows (≥ ~150) for its workload.** We fit one temperature `T` with
SemIf's `benchmarks/calibrate.py`, add `{"<workload>": T}` to
`/opt/docker/conf/semif/calibration.json` (`stacks/semif/conf/`), and restart. The
caller then passes `"workload": "<name>"` and gets a `calibrated` block beside the
native scores. The argmax never changes.
## VRAM: a hard cap, released after every burst
`VRAM_CAP_GIB=12` becomes `torch.cuda.set_per_process_memory_fraction` before the
weights load. At rest the process holds **~8.7 GB** (nvidia-smi); the weights are 7.84
GiB. After any call that grows torch's reserved memory past the post-warm-up baseline
+ 512 MiB, the engine calls `empty_cache()`, so a burst returns to the card and does
not squeeze scriberr, which shares GPU 1. A request that would exceed the cap gets
`503 out_of_memory`, memory returns to baseline, and the service stays up. Both
behaviours were verified on the card (0.1.0 held 11.9 GiB after an OOM, and 12.6 GB
after a large request; 0.1.2 returns to 7.85 GiB in both cases).
**What fits under 12 GiB** (measured, `/decide/shared`, binary decisions):
| state size (prefix tokens) | max decisions in one request |
|---|---|
| ~140 | 52 |
| ~520 | 43 |
| ~1,960 | 19 |
| ~3,900 | 13 |
`/decide` fits at the full 4,096-token limit. Past the table you get a 503, so split
the decisions across requests.
## Acceptance (2026-09-27, v0.1.2)
Raw: `services/semif-serve/acceptance/result-2026-09-27-v0.1.2.json`.
| check | result |
|---|---|
| parity with SemIf's committed torch predictions (authored144) | **142/144** same top choice; **144/144** identical prompt SHA-256; max prob gap 0.093 |
| noise floor (same 144 twice) | 144/144, gap 0.0: deterministic |
| negative control (option descriptions rotated) | 14/144: the check catches a wrong answer |
| shared vs direct (36 shared states, 72 rows) | 72/72, max gap 0.045 |
| 21 binary criteria over one state, from nh3-dev | shared **159 ms** (3 runs, 159–160) vs 21 sequential calls 981 ms |
The two parity misses are **exact bf16 ties in our output** (top-2 margin 0.000),
where upstream's prefix-cache path gave margins of 0.054 and 0.185; one of the two
now matches the label. So the misses come from the numeric path, not the wrapper. The
speed figure includes one network round trip (~27 ms). The burst release costs ~16 ms
on it (0.1.1 measured 143 ms without the release).
## Building
```bash
# from nh3-dev
tar -C services/semif-serve -cf - --exclude=.venv --exclude=.pytest_cache --exclude=__pycache__ --exclude=acceptance . \
| ssh infra-ops@10.251.50.54 'mkdir -p /opt/docker/src/semif-serve-X.Y.Z && tar -x -C /opt/docker/src/semif-serve-X.Y.Z'
# on fv-ml1
cd /opt/docker/src/semif-serve-X.Y.Z && docker build -t semif-serve:X.Y.Z .
```
To move SemIf forward: bump the commit in `services/semif-serve/pyproject.toml`
(and `SEMIF_COMMIT` in `config.py`), run `uv lock`, rebuild, and **re-run the
acceptance**. Upstream is research code that changes weekly, which is why it is
pinned.