Files
esh-pfi-infrastructure/stacks/semif/README.md
T
vh 77b8cb449c feat(semif): 0.1.3 — order averaging, fast kernels, bug-hunt hardening (Prime)
Order averaging (Prime, after the 739aa03 spike):
- A decision may set orderings: rotations|all (all only for <= 4 options). Every
  ordering goes to the engine in one shared batch.
- The reply keeps each native result and adds combined {probabilities (log-mean),
  top, agreement, spread}.
- Through the service on SemIf's labelled sets (252 rows): 78.6% -> 88.1%
  (group-bootstrap 95% CI +5.1..+14.3). Unanimous agreement is 94.5% accurate.

Fast kernels: flash-linear-attention 0.5.2 and causal-conv1d 1.7.0 are now the
default build. A/B on the empty GPU 3:
- parity with upstream went from 142/144 to 144/144;
- a ~2k-token /decide went from 169 to 92 ms server-side;
- short 3-rotation batches cost ~3-6 ms more.
triton builds a C shim at runtime, so the image carries gcc. Without it the
warm-up failed and startup failed closed.

Heid bug-hunt panel (4/4 arms, thread 01M3H3F4RR7XBP90KQ3A39H4SX), folded:
- Startup validation: VRAM cap 0 no longer means uncapped (C1); limits must be
  >= 1 (S1); the token must be visible ASCII (S2); the calibration file must
  exist and parse, with T in [0.05, 20] (S8, and C3's NaN leg).
- The body limit is checked before a chunk is kept, and a Unicode-digit
  Content-Length no longer crashes (C2, S3).
- Failures while building the response now get the 500 envelope (C3).
- 429 busy past SEMIF_MAX_QUEUE requests in progress (C6).
- The engine releases memory on every non-validation failure, unchained after
  gc; an empty OOM message is handled; 'out of memory' RuntimeErrors map to 503
  (C4, C5, S9).
- The entry point forces HF_HUB_OFFLINE (S10). README wording fixed (S5, S6).
- New guard tests close the gaps the arms' mutation grids exposed: early stop of
  the body read, a shared-route lock, calibration pass-through, the gc cycle,
  the exact caps, TorchEngine.load's arch and device checks, and the offline
  entry point.
86 tests.

Deployed on fv-ml1 GPU 1: parity 144/144, OOM and burst release verified, shared
capacity 63/51/26/16 rows at ~140/520/1960/3900 prefix tokens.
2026-09-27 03:27:15 -07:00

153 lines
7.9 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# semif
**SemIf option-logit decisions** on **fv-ml1 GPU 1**, the utility card beside
`vllm-coder`, the erp/meromero seats and scriberr. Deployed 2026-09-27 at Prime's
request. Hosting assessment: infra-hermes, 2026-09-25.
SemIf ([TheoLeeCJ/SemIf-OpenJev](https://github.com/TheoLeeCJ/SemIf-OpenJev), MIT)
asks a small model a typed question and reads the answer straight from the logits of
the option letters, after **one forward pass with no decoding**. Upstream ships only a
batch CLI, so `services/semif-serve/` wraps its two torch scorers in a small FastAPI
service. The contract is `services/semif-serve/semif-serve.contract.md`.
| | |
|---|---|
| **URL** | `http://10.251.50.54:8032` (`/health` is open; POSTs need `Authorization: Bearer $(secret get semif/api-token)`) |
| **Model** | `Qwen/Qwen3.5-4B` @ `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a`, BF16, from `/tank/aimodels/huggingface` (read-only, offline) |
| **SemIf** | commit `23cf1f39fc9534fe81437200959b6dfc7106e45a`; torch `2.10.0+cu128`, transformers `5.17.0`, the same stack SemIf's committed predictions were made on |
| **Image** | `semif-serve:<version>`, built on fv-ml1 from `services/semif-serve/` |
| **State** | none. If the weights are ever lost, re-pull them by hand at the pinned revision (the service itself never downloads: `HF_HUB_OFFLINE=1`, read-only mount). No backup needed beyond the host's `/opt/docker` restic. |
## API
```bash
T=$(secret get semif/api-token)
curl -s -H "Authorization: Bearer $T" http://10.251.50.54:8032/decide -d '{
"id": "q1", "state": "Health checks passed in all three zones.",
"question": "Did the deployment succeed?",
"options": [{"id":"yes","description":"It succeeded."},{"id":"no","description":"It failed."}]}'
```
- `POST /decide`: one decision, returning SemIf's result dict unchanged (`option_ids`,
`probabilities`, `option_logits`, `prompt_sha256`, `model`, …).
- `POST /decide/shared`: `{state, decisions: [{id, question, options}]}`. It prefills
the state once and scores every criterion in one batch. Use it when many questions
share one long state.
- `GET /health`: the pins, limits and calibrated workloads.
- **Order averaging (0.1.3):** add `"orderings": "rotations"` to a decision, or
`"all"` for ≤ 4 options. The option list is asked in every rotation inside ONE
shared batch. The reply keeps each ordering's native result under `orderings` and adds
`combined: {method, orderings, probabilities, top, agreement, spread}`. **Use it for
anything real:** a small model leans toward the first-listed option on ambiguous
inputs, and averaging cancels that. Through the service on SemIf's labelled sets,
accuracy goes from 78.6% to **88.1%** (group-bootstrap 95% CI +5.1..+14.3 pts, 252
rows). **`agreement` is the cheap confidence signal**: unanimous rows are 94.5%
accurate, split rows 76.4%. The orderings count toward the decision cap. `workload`
calibration is not available together with `orderings` yet (422).
- Past `MAX_QUEUE` (32) requests in progress, new POSTs get `429 busy` before their
body is read.
## ⚠ Probabilities are uncalibrated
SemIf labels its output `"conditional option score; uncalibrated as decision
confidence"`, and means it: on WANLI the model is right ~64% of the time while
reporting far higher confidence. **Before a caller thresholds on `probabilities`, it
brings labelled rows (≥ ~150) for its workload.** We fit one temperature `T` with
SemIf's `benchmarks/calibrate.py`, add `{"<workload>": T}` to
`/opt/docker/conf/semif/calibration.json` (`stacks/semif/conf/`), and restart. The
caller then passes `"workload": "<name>"` and gets a `calibrated` block beside the
native scores. The argmax never changes.
## VRAM: a hard cap, released after every burst
`VRAM_CAP_GIB=12` becomes `torch.cuda.set_per_process_memory_fraction` before the
weights load. At rest the process holds **~8.7 GB** (nvidia-smi); the weights are 7.84
GiB. After any call that grows torch's reserved memory past the post-warm-up baseline
+ 512 MiB, the engine calls `empty_cache()`, so a burst returns to the card and does
not squeeze scriberr, which shares GPU 1. A request that would exceed the cap gets
`503 out_of_memory`, memory returns to baseline, and the service stays up. Both
behaviours were verified on the card (0.1.0 held 11.9 GiB after an OOM, and 12.6 GB
after a large request; 0.1.2 returns to 7.85 GiB in both cases).
**What fits under 12 GiB** (measured on 0.1.3, `/decide/shared`, binary decisions;
0.1.2 figures in brackets, before the fast kernels):
| state size (prefix tokens) | max rows in one request |
|---|---|
| ~140 | 63 (52) |
| ~520 | 51 (43) |
| ~1,960 | 26 (19) |
| ~3,900 | 16 (13) |
Rows = decisions × orderings, so `rotations` over 3 options uses 3 rows per decision.
`/decide` fits at the full 4,096-token limit. Past the table you get a 503, so split
the decisions across requests.
## Fast kernels (0.1.3)
The image ships Qwen3.5's fast kernels, `flash-linear-attention` 0.5.2 and
`causal-conv1d` 1.7.0 (a prebuilt cu12/torch2.10 wheel). Without them transformers
logs that it falls back to "much slower" reference PyTorch paths. They are adopted
because an A/B on the empty GPU 3 showed:
- **parity improved**: 144/144 vs upstream (142/144 without), so upstream evidently
ran with them;
- **long inputs got much faster**: a ~2,000-token `/decide` went from 169 to 92 ms
server-side.
Short 3-rotation batches cost ~3–6 ms more; everything else is equal or faster.
⚠ triton builds a C shim at runtime, so the image carries `gcc`. Without it the
warm-up fails, and startup fails closed. Build without the kernels:
`--build-arg EXTRAS="--extra model"`.
## Latency (0.1.3, from nh3-dev, 3 runs × 20, network floor ~33 ms)
| request | end to end | server |
|---|---|---|
| `/decide`, short (~130 tok) | 69 ms | 35 ms |
| `/decide`, ~2,000-token state | 131 ms | 92 ms |
| 3 rotations, short | 115 ms | 78 ms |
| 6 orderings, short | 118 ms | 81 ms |
| 3 rotations, ~2,000-token state | 200 ms | 158 ms |
## Acceptance (2026-09-27, v0.1.3)
Raw: `services/semif-serve/acceptance/result-2026-09-27-v0.1.3.json`,
`averaging-2026-09-27-v0.1.3.json`. v0.1.3 matches upstream on **144/144** (identical
prompt hashes, max prob gap 0.049), is deterministic, fails the negative control as it
should (14/144), and shared matches direct on 72/72. The v0.1.2 table below is the
reference-kernel baseline.
## Acceptance (2026-09-27, v0.1.2)
Raw: `services/semif-serve/acceptance/result-2026-09-27-v0.1.2.json`.
| check | result |
|---|---|
| parity with SemIf's committed torch predictions (authored144) | **142/144** same top choice; **144/144** identical prompt SHA-256; max prob gap 0.093 |
| noise floor (same 144 twice) | 144/144, gap 0.0: deterministic |
| negative control (option descriptions rotated) | 14/144: the check catches a wrong answer |
| shared vs direct (36 shared states, 72 rows) | 72/72, max gap 0.045 |
| 21 binary criteria over one state, from nh3-dev | shared **159 ms** (3 runs, 159–160) vs 21 sequential calls 981 ms |
The two parity misses are **exact bf16 ties in our output** (top-2 margin 0.000),
where upstream's prefix-cache path gave margins of 0.054 and 0.185; one of the two
now matches the label. So the misses come from the numeric path, not the wrapper. The
speed figure includes one network round trip (~27 ms). The burst release costs ~16 ms
on it (0.1.1 measured 143 ms without the release).
## Building
```bash
# from nh3-dev
tar -C services/semif-serve -cf - --exclude=.venv --exclude=.pytest_cache --exclude=__pycache__ --exclude=acceptance . \
| ssh infra-ops@10.251.50.54 'mkdir -p /opt/docker/src/semif-serve-X.Y.Z && tar -x -C /opt/docker/src/semif-serve-X.Y.Z'
# on fv-ml1
cd /opt/docker/src/semif-serve-X.Y.Z && docker build -t semif-serve:X.Y.Z .
```
To move SemIf forward: bump the commit in `services/semif-serve/pyproject.toml`
(and `SEMIF_COMMIT` in `config.py`), run `uv lock`, rebuild, and **re-run the
acceptance**. Upstream is research code that changes weekly, which is why it is
pinned.