# semif **SemIf option-logit decisions** on **fv-ml1 GPU 1**, the utility card beside `vllm-coder`, the erp/meromero seats and scriberr. Deployed 2026-09-27 at Prime's request. Hosting assessment: infra-hermes, 2026-09-25. SemIf ([TheoLeeCJ/SemIf-OpenJev](https://github.com/TheoLeeCJ/SemIf-OpenJev), MIT) asks a small model a typed question and reads the answer straight from the logits of the option letters, after **one forward pass with no decoding**. Upstream ships only a batch CLI, so `services/semif-serve/` wraps its two torch scorers in a small FastAPI service. The contract is `services/semif-serve/semif-serve.contract.md`. | | | |---|---| | **URL** | `http://10.251.50.54:8032` (`/health` is open; POSTs need `Authorization: Bearer $(secret get semif/api-token)`) | | **Model** | `Qwen/Qwen3.5-4B` @ `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a`, BF16, from `/tank/aimodels/huggingface` (read-only, offline) | | **SemIf** | commit `23cf1f39fc9534fe81437200959b6dfc7106e45a`; torch `2.10.0+cu128`, transformers `5.17.0`, the same stack SemIf's committed predictions were made on | | **Image** | `semif-serve:`, built on fv-ml1 from `services/semif-serve/` | | **State** | none. If the weights are ever lost, re-pull them by hand at the pinned revision (the service itself never downloads: `HF_HUB_OFFLINE=1`, read-only mount). No backup needed beyond the host's `/opt/docker` restic. | ## API ```bash T=$(secret get semif/api-token) curl -s -H "Authorization: Bearer $T" http://10.251.50.54:8032/decide -d '{ "id": "q1", "state": "Health checks passed in all three zones.", "question": "Did the deployment succeed?", "options": [{"id":"yes","description":"It succeeded."},{"id":"no","description":"It failed."}]}' ``` - `POST /decide`: one decision, returning SemIf's result dict unchanged (`option_ids`, `probabilities`, `option_logits`, `prompt_sha256`, `model`, …). - `POST /decide/shared`: `{state, decisions: [{id, question, options}]}`. It prefills the state once and scores every criterion in one batch. Use it when many questions share one long state. - `GET /health`: the pins, limits and calibrated workloads. - **Order averaging (0.1.3):** add `"orderings": "rotations"` to a decision, or `"all"` for ≤ 4 options. The option list is asked in every rotation inside ONE shared batch. The reply keeps each ordering's native result under `orderings` and adds `combined: {method, orderings, probabilities, top, agreement, spread}`. **Use it for anything real:** a small model leans toward the first-listed option on ambiguous inputs, and averaging cancels that. Through the service on SemIf's labelled sets, accuracy goes from 78.6% to **88.1%** (group-bootstrap 95% CI +5.1..+14.3 pts, 252 rows). **`agreement` is the cheap confidence signal**: unanimous rows are 94.5% accurate, split rows 76.4%. The orderings count toward the decision cap. `workload` calibration is not available together with `orderings` yet (422). - Past `MAX_QUEUE` (32) requests in progress, new POSTs get `429 busy` before their body is read. - ⚠ **Rotations cost options², not options.** The options live in each row's suffix, and the shared prefix is only the state. So `rotations` over n options sends n rows each carrying all n options. Measured 2026-09-27 on a short state, over 16 options of about 40 tokens each: **850 ms with rotations against 109 ms for one ordering** (10,304 against 644 suffix tokens, 5 calls after warm-up). One cold call at that size returned 503. The VRAM table below covers binary decisions only. For many options, use one ordering or shorter option text. - ⚠ **`/decide/shared` refuses some object states.** If the state is an object whose LAST value ends in `)`, `;` or `}`, the service returns 422 "The fixed state prefix does not match every full prompt". The closing `"}` merges with that character into one token. The same text as a plain string state works, and `.`, `!`, `?`, `]`, `…` and `—` endings work. Not fixed yet. Callers that pass user text last should append a full stop or send a string state. ## ⚠ Probabilities are uncalibrated SemIf labels its output `"conditional option score; uncalibrated as decision confidence"`, and means it: on WANLI the model is right ~64% of the time while reporting far higher confidence. **Before a caller thresholds on `probabilities`, it brings labelled rows (≥ ~150) for its workload.** We fit one temperature `T` with SemIf's `benchmarks/calibrate.py`, add `{"": T}` to `/opt/docker/conf/semif/calibration.json` (`stacks/semif/conf/`), and restart. The caller then passes `"workload": ""` and gets a `calibrated` block beside the native scores. The argmax never changes. ## VRAM: a hard cap, released after every burst `VRAM_CAP_GIB=12` becomes `torch.cuda.set_per_process_memory_fraction` before the weights load. At rest the process holds **~8.7 GB** (nvidia-smi); the weights are 7.84 GiB. After any call that grows torch's reserved memory past the post-warm-up baseline + 512 MiB, the engine calls `empty_cache()`, so a burst returns to the card and does not squeeze scriberr, which shares GPU 1. A request that would exceed the cap gets `503 out_of_memory`, memory returns to baseline, and the service stays up. Both behaviours were verified on the card (0.1.0 held 11.9 GiB after an OOM, and 12.6 GB after a large request; 0.1.2 returns to 7.85 GiB in both cases). **What fits under 12 GiB** (measured on 0.1.3, `/decide/shared`, binary decisions; 0.1.2 figures in brackets, before the fast kernels): | state size (prefix tokens) | max rows in one request | |---|---| | ~140 | 63 (52) | | ~520 | 51 (43) | | ~1,960 | 26 (19) | | ~3,900 | 16 (13) | Rows = decisions × orderings, so `rotations` over 3 options uses 3 rows per decision. `/decide` fits at the full 4,096-token limit. Past the table you get a 503, so split the decisions across requests. ## Fast kernels (0.1.3) The image ships Qwen3.5's fast kernels, `flash-linear-attention` 0.5.2 and `causal-conv1d` 1.7.0 (a prebuilt cu12/torch2.10 wheel). Without them transformers logs that it falls back to "much slower" reference PyTorch paths. They are adopted because an A/B on the empty GPU 3 showed: - **parity improved**: 144/144 vs upstream (142/144 without), so upstream evidently ran with them; - **long inputs got much faster**: a ~2,000-token `/decide` went from 169 to 92 ms server-side. Short 3-rotation batches cost ~3–6 ms more; everything else is equal or faster. ⚠ triton builds a C shim at runtime, so the image carries `gcc`. Without it the warm-up fails, and startup fails closed. Build without the kernels: `--build-arg EXTRAS="--extra model"`. ## Latency (0.1.3, from nh3-dev, 3 runs × 20, network floor ~33 ms) | request | end to end | server | |---|---|---| | `/decide`, short (~130 tok) | 69 ms | 35 ms | | `/decide`, ~2,000-token state | 131 ms | 92 ms | | 3 rotations, short | 115 ms | 78 ms | | 6 orderings, short | 118 ms | 81 ms | | 3 rotations, ~2,000-token state | 200 ms | 158 ms | ## Acceptance (2026-09-27, v0.1.3) Raw: `services/semif-serve/acceptance/result-2026-09-27-v0.1.3.json`, `averaging-2026-09-27-v0.1.3.json`. v0.1.3 matches upstream on **144/144** (identical prompt hashes, max prob gap 0.049), is deterministic, fails the negative control as it should (14/144), and shared matches direct on 72/72. The v0.1.2 table below is the reference-kernel baseline. ## Acceptance (2026-09-27, v0.1.2) Raw: `services/semif-serve/acceptance/result-2026-09-27-v0.1.2.json`. | check | result | |---|---| | parity with SemIf's committed torch predictions (authored144) | **142/144** same top choice; **144/144** identical prompt SHA-256; max prob gap 0.093 | | noise floor (same 144 twice) | 144/144, gap 0.0: deterministic | | negative control (option descriptions rotated) | 14/144: the check catches a wrong answer | | shared vs direct (36 shared states, 72 rows) | 72/72, max gap 0.045 | | 21 binary criteria over one state, from nh3-dev | shared **159 ms** (3 runs, 159–160) vs 21 sequential calls 981 ms | The two parity misses are **exact bf16 ties in our output** (top-2 margin 0.000), where upstream's prefix-cache path gave margins of 0.054 and 0.185; one of the two now matches the label. So the misses come from the numeric path, not the wrapper. The speed figure includes one network round trip (~27 ms). The burst release costs ~16 ms on it (0.1.1 measured 143 ms without the release). ## Building ```bash # from nh3-dev tar -C services/semif-serve -cf - --exclude=.venv --exclude=.pytest_cache --exclude=__pycache__ --exclude=acceptance . \ | ssh infra-ops@10.251.50.54 'mkdir -p /opt/docker/src/semif-serve-X.Y.Z && tar -x -C /opt/docker/src/semif-serve-X.Y.Z' # on fv-ml1 cd /opt/docker/src/semif-serve-X.Y.Z && docker build -t semif-serve:X.Y.Z . ``` To move SemIf forward: bump the commit in `services/semif-serve/pyproject.toml` (and `SEMIF_COMMIT` in `config.py`), run `uv lock`, rebuild, and **re-run the acceptance**. Upstream is research code that changes weekly, which is why it is pinned.