Against talk /face's guided pose (first paragraph 246 ms median), SemIf in parallel adds 32 ms and SemIf-first adds 94 ms (n=72 each, noise floor 16.5 ms). Removing the pose header saves only ~31 ms, and SemIf shares GPU 1 with the LLM. Acceptable pose 67% vs 92% on clear-emotion lines, and the mood carried through mundane follow-ups 7/15 vs 14/15. SemIf gestures far less (13% vs 58%). README: rotations cost options^2 in suffix tokens, and /decide/shared returns 422 when an object state's last value ends in ) ; or }.
semif
SemIf option-logit decisions on fv-ml1 GPU 1, the utility card beside
vllm-coder, the erp/meromero seats and scriberr. Deployed 2026-09-27 at Prime's
request. Hosting assessment: infra-hermes, 2026-09-25.
SemIf (TheoLeeCJ/SemIf-OpenJev, MIT)
asks a small model a typed question and reads the answer straight from the logits of
the option letters, after one forward pass with no decoding. Upstream ships only a
batch CLI, so services/semif-serve/ wraps its two torch scorers in a small FastAPI
service. The contract is services/semif-serve/semif-serve.contract.md.
| URL | http://10.251.50.54:8032 (/health is open; POSTs need Authorization: Bearer $(secret get semif/api-token)) |
| Model | Qwen/Qwen3.5-4B @ 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a, BF16, from /tank/aimodels/huggingface (read-only, offline) |
| SemIf | commit 23cf1f39fc9534fe81437200959b6dfc7106e45a; torch 2.10.0+cu128, transformers 5.17.0, the same stack SemIf's committed predictions were made on |
| Image | semif-serve:<version>, built on fv-ml1 from services/semif-serve/ |
| State | none. If the weights are ever lost, re-pull them by hand at the pinned revision (the service itself never downloads: HF_HUB_OFFLINE=1, read-only mount). No backup needed beyond the host's /opt/docker restic. |
API
T=$(secret get semif/api-token)
curl -s -H "Authorization: Bearer $T" http://10.251.50.54:8032/decide -d '{
"id": "q1", "state": "Health checks passed in all three zones.",
"question": "Did the deployment succeed?",
"options": [{"id":"yes","description":"It succeeded."},{"id":"no","description":"It failed."}]}'
POST /decide: one decision, returning SemIf's result dict unchanged (option_ids,probabilities,option_logits,prompt_sha256,model, …).POST /decide/shared:{state, decisions: [{id, question, options}]}. It prefills the state once and scores every criterion in one batch. Use it when many questions share one long state.GET /health: the pins, limits and calibrated workloads.- Order averaging (0.1.3): add
"orderings": "rotations"to a decision, or"all"for ≤ 4 options. The option list is asked in every rotation inside ONE shared batch. The reply keeps each ordering's native result underorderingsand addscombined: {method, orderings, probabilities, top, agreement, spread}. Use it for anything real: a small model leans toward the first-listed option on ambiguous inputs, and averaging cancels that. Through the service on SemIf's labelled sets, accuracy goes from 78.6% to 88.1% (group-bootstrap 95% CI +5.1..+14.3 pts, 252 rows).agreementis the cheap confidence signal: unanimous rows are 94.5% accurate, split rows 76.4%. The orderings count toward the decision cap.workloadcalibration is not available together withorderingsyet (422). - Past
MAX_QUEUE(32) requests in progress, new POSTs get429 busybefore their body is read. - ⚠ Rotations cost options², not options. The options live in each row's suffix, and
the shared prefix is only the state. So
rotationsover n options sends n rows each carrying all n options. Measured 2026-09-27 on a short state, over 16 options of about 40 tokens each: 850 ms with rotations against 109 ms for one ordering (10,304 against 644 suffix tokens, 5 calls after warm-up). One cold call at that size returned 503. The VRAM table below covers binary decisions only. For many options, use one ordering or shorter option text. - ⚠
/decide/sharedrefuses some object states. If the state is an object whose LAST value ends in),;or}, the service returns 422 "The fixed state prefix does not match every full prompt". The closing"}merges with that character into one token. The same text as a plain string state works, and.,!,?,],…and—endings work. Not fixed yet. Callers that pass user text last should append a full stop or send a string state.
⚠ Probabilities are uncalibrated
SemIf labels its output "conditional option score; uncalibrated as decision confidence", and means it: on WANLI the model is right ~64% of the time while
reporting far higher confidence. Before a caller thresholds on probabilities, it
brings labelled rows (≥ ~150) for its workload. We fit one temperature T with
SemIf's benchmarks/calibrate.py, add {"<workload>": T} to
/opt/docker/conf/semif/calibration.json (stacks/semif/conf/), and restart. The
caller then passes "workload": "<name>" and gets a calibrated block beside the
native scores. The argmax never changes.
VRAM: a hard cap, released after every burst
VRAM_CAP_GIB=12 becomes torch.cuda.set_per_process_memory_fraction before the
weights load. At rest the process holds ~8.7 GB (nvidia-smi); the weights are 7.84
GiB. After any call that grows torch's reserved memory past the post-warm-up baseline
- 512 MiB, the engine calls
empty_cache(), so a burst returns to the card and does not squeeze scriberr, which shares GPU 1. A request that would exceed the cap gets503 out_of_memory, memory returns to baseline, and the service stays up. Both behaviours were verified on the card (0.1.0 held 11.9 GiB after an OOM, and 12.6 GB after a large request; 0.1.2 returns to 7.85 GiB in both cases).
What fits under 12 GiB (measured on 0.1.3, /decide/shared, binary decisions;
0.1.2 figures in brackets, before the fast kernels):
| state size (prefix tokens) | max rows in one request |
|---|---|
| ~140 | 63 (52) |
| ~520 | 51 (43) |
| ~1,960 | 26 (19) |
| ~3,900 | 16 (13) |
Rows = decisions × orderings, so rotations over 3 options uses 3 rows per decision.
/decide fits at the full 4,096-token limit. Past the table you get a 503, so split
the decisions across requests.
Fast kernels (0.1.3)
The image ships Qwen3.5's fast kernels, flash-linear-attention 0.5.2 and
causal-conv1d 1.7.0 (a prebuilt cu12/torch2.10 wheel). Without them transformers
logs that it falls back to "much slower" reference PyTorch paths. They are adopted
because an A/B on the empty GPU 3 showed:
- parity improved: 144/144 vs upstream (142/144 without), so upstream evidently ran with them;
- long inputs got much faster: a ~2,000-token
/decidewent from 169 to 92 ms server-side.
Short 3-rotation batches cost ~3–6 ms more; everything else is equal or faster.
⚠ triton builds a C shim at runtime, so the image carries gcc. Without it the
warm-up fails, and startup fails closed. Build without the kernels:
--build-arg EXTRAS="--extra model".
Latency (0.1.3, from nh3-dev, 3 runs × 20, network floor ~33 ms)
| request | end to end | server |
|---|---|---|
/decide, short (~130 tok) |
69 ms | 35 ms |
/decide, ~2,000-token state |
131 ms | 92 ms |
| 3 rotations, short | 115 ms | 78 ms |
| 6 orderings, short | 118 ms | 81 ms |
| 3 rotations, ~2,000-token state | 200 ms | 158 ms |
Acceptance (2026-09-27, v0.1.3)
Raw: services/semif-serve/acceptance/result-2026-09-27-v0.1.3.json,
averaging-2026-09-27-v0.1.3.json. v0.1.3 matches upstream on 144/144 (identical
prompt hashes, max prob gap 0.049), is deterministic, fails the negative control as it
should (14/144), and shared matches direct on 72/72. The v0.1.2 table below is the
reference-kernel baseline.
Acceptance (2026-09-27, v0.1.2)
Raw: services/semif-serve/acceptance/result-2026-09-27-v0.1.2.json.
| check | result |
|---|---|
| parity with SemIf's committed torch predictions (authored144) | 142/144 same top choice; 144/144 identical prompt SHA-256; max prob gap 0.093 |
| noise floor (same 144 twice) | 144/144, gap 0.0: deterministic |
| negative control (option descriptions rotated) | 14/144: the check catches a wrong answer |
| shared vs direct (36 shared states, 72 rows) | 72/72, max gap 0.045 |
| 21 binary criteria over one state, from nh3-dev | shared 159 ms (3 runs, 159–160) vs 21 sequential calls 981 ms |
The two parity misses are exact bf16 ties in our output (top-2 margin 0.000), where upstream's prefix-cache path gave margins of 0.054 and 0.185; one of the two now matches the label. So the misses come from the numeric path, not the wrapper. The speed figure includes one network round trip (~27 ms). The burst release costs ~16 ms on it (0.1.1 measured 143 ms without the release).
Building
# from nh3-dev
tar -C services/semif-serve -cf - --exclude=.venv --exclude=.pytest_cache --exclude=__pycache__ --exclude=acceptance . \
| ssh infra-ops@10.251.50.54 'mkdir -p /opt/docker/src/semif-serve-X.Y.Z && tar -x -C /opt/docker/src/semif-serve-X.Y.Z'
# on fv-ml1
cd /opt/docker/src/semif-serve-X.Y.Z && docker build -t semif-serve:X.Y.Z .
To move SemIf forward: bump the commit in services/semif-serve/pyproject.toml
(and SEMIF_COMMIT in config.py), run uv lock, rebuild, and re-run the
acceptance. Upstream is research code that changes weekly, which is why it is
pinned.