Against talk /face's guided pose (first paragraph 246 ms median), SemIf in parallel adds 32 ms and SemIf-first adds 94 ms (n=72 each, noise floor 16.5 ms). Removing the pose header saves only ~31 ms, and SemIf shares GPU 1 with the LLM. Acceptable pose 67% vs 92% on clear-emotion lines, and the mood carried through mundane follow-ups 7/15 vs 14/15. SemIf gestures far less (13% vs 58%). README: rotations cost options^2 in suffix tokens, and /decide/shared returns 422 when an object state's last value ends in ) ; or }.
166 lines
8.9 KiB
Markdown
166 lines
8.9 KiB
Markdown
# semif
|
||
|
||
**SemIf option-logit decisions** on **fv-ml1 GPU 1**, the utility card beside
|
||
`vllm-coder`, the erp/meromero seats and scriberr. Deployed 2026-09-27 at Prime's
|
||
request. Hosting assessment: infra-hermes, 2026-09-25.
|
||
|
||
SemIf ([TheoLeeCJ/SemIf-OpenJev](https://github.com/TheoLeeCJ/SemIf-OpenJev), MIT)
|
||
asks a small model a typed question and reads the answer straight from the logits of
|
||
the option letters, after **one forward pass with no decoding**. Upstream ships only a
|
||
batch CLI, so `services/semif-serve/` wraps its two torch scorers in a small FastAPI
|
||
service. The contract is `services/semif-serve/semif-serve.contract.md`.
|
||
|
||
| | |
|
||
|---|---|
|
||
| **URL** | `http://10.251.50.54:8032` (`/health` is open; POSTs need `Authorization: Bearer $(secret get semif/api-token)`) |
|
||
| **Model** | `Qwen/Qwen3.5-4B` @ `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a`, BF16, from `/tank/aimodels/huggingface` (read-only, offline) |
|
||
| **SemIf** | commit `23cf1f39fc9534fe81437200959b6dfc7106e45a`; torch `2.10.0+cu128`, transformers `5.17.0`, the same stack SemIf's committed predictions were made on |
|
||
| **Image** | `semif-serve:<version>`, built on fv-ml1 from `services/semif-serve/` |
|
||
| **State** | none. If the weights are ever lost, re-pull them by hand at the pinned revision (the service itself never downloads: `HF_HUB_OFFLINE=1`, read-only mount). No backup needed beyond the host's `/opt/docker` restic. |
|
||
|
||
## API
|
||
|
||
```bash
|
||
T=$(secret get semif/api-token)
|
||
curl -s -H "Authorization: Bearer $T" http://10.251.50.54:8032/decide -d '{
|
||
"id": "q1", "state": "Health checks passed in all three zones.",
|
||
"question": "Did the deployment succeed?",
|
||
"options": [{"id":"yes","description":"It succeeded."},{"id":"no","description":"It failed."}]}'
|
||
```
|
||
|
||
- `POST /decide`: one decision, returning SemIf's result dict unchanged (`option_ids`,
|
||
`probabilities`, `option_logits`, `prompt_sha256`, `model`, …).
|
||
- `POST /decide/shared`: `{state, decisions: [{id, question, options}]}`. It prefills
|
||
the state once and scores every criterion in one batch. Use it when many questions
|
||
share one long state.
|
||
- `GET /health`: the pins, limits and calibrated workloads.
|
||
- **Order averaging (0.1.3):** add `"orderings": "rotations"` to a decision, or
|
||
`"all"` for ≤ 4 options. The option list is asked in every rotation inside ONE
|
||
shared batch. The reply keeps each ordering's native result under `orderings` and adds
|
||
`combined: {method, orderings, probabilities, top, agreement, spread}`. **Use it for
|
||
anything real:** a small model leans toward the first-listed option on ambiguous
|
||
inputs, and averaging cancels that. Through the service on SemIf's labelled sets,
|
||
accuracy goes from 78.6% to **88.1%** (group-bootstrap 95% CI +5.1..+14.3 pts, 252
|
||
rows). **`agreement` is the cheap confidence signal**: unanimous rows are 94.5%
|
||
accurate, split rows 76.4%. The orderings count toward the decision cap. `workload`
|
||
calibration is not available together with `orderings` yet (422).
|
||
- Past `MAX_QUEUE` (32) requests in progress, new POSTs get `429 busy` before their
|
||
body is read.
|
||
- ⚠ **Rotations cost options², not options.** The options live in each row's suffix, and
|
||
the shared prefix is only the state. So `rotations` over n options sends n rows each
|
||
carrying all n options. Measured 2026-09-27 on a short state, over 16 options of about
|
||
40 tokens each: **850 ms with rotations against 109 ms for one ordering** (10,304 against
|
||
644 suffix tokens, 5 calls after warm-up). One cold call at that size returned 503. The
|
||
VRAM table below covers binary decisions only. For many options, use one ordering or
|
||
shorter option text.
|
||
- ⚠ **`/decide/shared` refuses some object states.** If the state is an object whose
|
||
LAST value ends in `)`, `;` or `}`, the service returns 422 "The fixed state prefix
|
||
does not match every full prompt". The closing `"}` merges with that character into one
|
||
token. The same text as a plain string state works, and `.`, `!`, `?`, `]`, `…` and
|
||
`—` endings work. Not fixed yet. Callers that pass user text last should
|
||
append a full stop or send a string state.
|
||
|
||
## ⚠ Probabilities are uncalibrated
|
||
|
||
SemIf labels its output `"conditional option score; uncalibrated as decision
|
||
confidence"`, and means it: on WANLI the model is right ~64% of the time while
|
||
reporting far higher confidence. **Before a caller thresholds on `probabilities`, it
|
||
brings labelled rows (≥ ~150) for its workload.** We fit one temperature `T` with
|
||
SemIf's `benchmarks/calibrate.py`, add `{"<workload>": T}` to
|
||
`/opt/docker/conf/semif/calibration.json` (`stacks/semif/conf/`), and restart. The
|
||
caller then passes `"workload": "<name>"` and gets a `calibrated` block beside the
|
||
native scores. The argmax never changes.
|
||
|
||
## VRAM: a hard cap, released after every burst
|
||
|
||
`VRAM_CAP_GIB=12` becomes `torch.cuda.set_per_process_memory_fraction` before the
|
||
weights load. At rest the process holds **~8.7 GB** (nvidia-smi); the weights are 7.84
|
||
GiB. After any call that grows torch's reserved memory past the post-warm-up baseline
|
||
+ 512 MiB, the engine calls `empty_cache()`, so a burst returns to the card and does
|
||
not squeeze scriberr, which shares GPU 1. A request that would exceed the cap gets
|
||
`503 out_of_memory`, memory returns to baseline, and the service stays up. Both
|
||
behaviours were verified on the card (0.1.0 held 11.9 GiB after an OOM, and 12.6 GB
|
||
after a large request; 0.1.2 returns to 7.85 GiB in both cases).
|
||
|
||
**What fits under 12 GiB** (measured on 0.1.3, `/decide/shared`, binary decisions;
|
||
0.1.2 figures in brackets, before the fast kernels):
|
||
|
||
| state size (prefix tokens) | max rows in one request |
|
||
|---|---|
|
||
| ~140 | 63 (52) |
|
||
| ~520 | 51 (43) |
|
||
| ~1,960 | 26 (19) |
|
||
| ~3,900 | 16 (13) |
|
||
|
||
Rows = decisions × orderings, so `rotations` over 3 options uses 3 rows per decision.
|
||
|
||
`/decide` fits at the full 4,096-token limit. Past the table you get a 503, so split
|
||
the decisions across requests.
|
||
|
||
## Fast kernels (0.1.3)
|
||
|
||
The image ships Qwen3.5's fast kernels, `flash-linear-attention` 0.5.2 and
|
||
`causal-conv1d` 1.7.0 (a prebuilt cu12/torch2.10 wheel). Without them transformers
|
||
logs that it falls back to "much slower" reference PyTorch paths. They are adopted
|
||
because an A/B on the empty GPU 3 showed:
|
||
- **parity improved**: 144/144 vs upstream (142/144 without), so upstream evidently
|
||
ran with them;
|
||
- **long inputs got much faster**: a ~2,000-token `/decide` went from 169 to 92 ms
|
||
server-side.
|
||
|
||
Short 3-rotation batches cost ~3–6 ms more; everything else is equal or faster.
|
||
⚠ triton builds a C shim at runtime, so the image carries `gcc`. Without it the
|
||
warm-up fails, and startup fails closed. Build without the kernels:
|
||
`--build-arg EXTRAS="--extra model"`.
|
||
|
||
## Latency (0.1.3, from nh3-dev, 3 runs × 20, network floor ~33 ms)
|
||
|
||
| request | end to end | server |
|
||
|---|---|---|
|
||
| `/decide`, short (~130 tok) | 69 ms | 35 ms |
|
||
| `/decide`, ~2,000-token state | 131 ms | 92 ms |
|
||
| 3 rotations, short | 115 ms | 78 ms |
|
||
| 6 orderings, short | 118 ms | 81 ms |
|
||
| 3 rotations, ~2,000-token state | 200 ms | 158 ms |
|
||
|
||
## Acceptance (2026-09-27, v0.1.3)
|
||
|
||
Raw: `services/semif-serve/acceptance/result-2026-09-27-v0.1.3.json`,
|
||
`averaging-2026-09-27-v0.1.3.json`. v0.1.3 matches upstream on **144/144** (identical
|
||
prompt hashes, max prob gap 0.049), is deterministic, fails the negative control as it
|
||
should (14/144), and shared matches direct on 72/72. The v0.1.2 table below is the
|
||
reference-kernel baseline.
|
||
|
||
## Acceptance (2026-09-27, v0.1.2)
|
||
|
||
Raw: `services/semif-serve/acceptance/result-2026-09-27-v0.1.2.json`.
|
||
|
||
| check | result |
|
||
|---|---|
|
||
| parity with SemIf's committed torch predictions (authored144) | **142/144** same top choice; **144/144** identical prompt SHA-256; max prob gap 0.093 |
|
||
| noise floor (same 144 twice) | 144/144, gap 0.0: deterministic |
|
||
| negative control (option descriptions rotated) | 14/144: the check catches a wrong answer |
|
||
| shared vs direct (36 shared states, 72 rows) | 72/72, max gap 0.045 |
|
||
| 21 binary criteria over one state, from nh3-dev | shared **159 ms** (3 runs, 159–160) vs 21 sequential calls 981 ms |
|
||
|
||
The two parity misses are **exact bf16 ties in our output** (top-2 margin 0.000),
|
||
where upstream's prefix-cache path gave margins of 0.054 and 0.185; one of the two
|
||
now matches the label. So the misses come from the numeric path, not the wrapper. The
|
||
speed figure includes one network round trip (~27 ms). The burst release costs ~16 ms
|
||
on it (0.1.1 measured 143 ms without the release).
|
||
|
||
## Building
|
||
|
||
```bash
|
||
# from nh3-dev
|
||
tar -C services/semif-serve -cf - --exclude=.venv --exclude=.pytest_cache --exclude=__pycache__ --exclude=acceptance . \
|
||
| ssh infra-ops@10.251.50.54 'mkdir -p /opt/docker/src/semif-serve-X.Y.Z && tar -x -C /opt/docker/src/semif-serve-X.Y.Z'
|
||
# on fv-ml1
|
||
cd /opt/docker/src/semif-serve-X.Y.Z && docker build -t semif-serve:X.Y.Z .
|
||
```
|
||
|
||
To move SemIf forward: bump the commit in `services/semif-serve/pyproject.toml`
|
||
(and `SEMIF_COMMIT` in `config.py`), run `uv lock`, rebuild, and **re-run the
|
||
acceptance**. Upstream is research code that changes weekly, which is why it is
|
||
pinned.
|