Files
esh-pfi-infrastructure/stacks/semif
vh 77b8cb449c feat(semif): 0.1.3 — order averaging, fast kernels, bug-hunt hardening (Prime)
Order averaging (Prime, after the 739aa03 spike):
- A decision may set orderings: rotations|all (all only for <= 4 options). Every
  ordering goes to the engine in one shared batch.
- The reply keeps each native result and adds combined {probabilities (log-mean),
  top, agreement, spread}.
- Through the service on SemIf's labelled sets (252 rows): 78.6% -> 88.1%
  (group-bootstrap 95% CI +5.1..+14.3). Unanimous agreement is 94.5% accurate.

Fast kernels: flash-linear-attention 0.5.2 and causal-conv1d 1.7.0 are now the
default build. A/B on the empty GPU 3:
- parity with upstream went from 142/144 to 144/144;
- a ~2k-token /decide went from 169 to 92 ms server-side;
- short 3-rotation batches cost ~3-6 ms more.
triton builds a C shim at runtime, so the image carries gcc. Without it the
warm-up failed and startup failed closed.

Heid bug-hunt panel (4/4 arms, thread 01M3H3F4RR7XBP90KQ3A39H4SX), folded:
- Startup validation: VRAM cap 0 no longer means uncapped (C1); limits must be
  >= 1 (S1); the token must be visible ASCII (S2); the calibration file must
  exist and parse, with T in [0.05, 20] (S8, and C3's NaN leg).
- The body limit is checked before a chunk is kept, and a Unicode-digit
  Content-Length no longer crashes (C2, S3).
- Failures while building the response now get the 500 envelope (C3).
- 429 busy past SEMIF_MAX_QUEUE requests in progress (C6).
- The engine releases memory on every non-validation failure, unchained after
  gc; an empty OOM message is handled; 'out of memory' RuntimeErrors map to 503
  (C4, C5, S9).
- The entry point forces HF_HUB_OFFLINE (S10). README wording fixed (S5, S6).
- New guard tests close the gaps the arms' mutation grids exposed: early stop of
  the body read, a shared-route lock, calibration pass-through, the gc cycle,
  the exact caps, TorchEngine.load's arch and device checks, and the offline
  entry point.
86 tests.

Deployed on fv-ml1 GPU 1: parity 144/144, OOM and burst release verified, shared
capacity 63/51/26/16 rows at ~140/520/1960/3900 prefix tokens.
2026-09-27 03:27:15 -07:00
..

semif

SemIf option-logit decisions on fv-ml1 GPU 1, the utility card beside vllm-coder, the erp/meromero seats and scriberr. Deployed 2026-09-27 at Prime's request. Hosting assessment: infra-hermes, 2026-09-25.

SemIf (TheoLeeCJ/SemIf-OpenJev, MIT) asks a small model a typed question and reads the answer straight from the logits of the option letters, after one forward pass with no decoding. Upstream ships only a batch CLI, so services/semif-serve/ wraps its two torch scorers in a small FastAPI service. The contract is services/semif-serve/semif-serve.contract.md.

URL http://10.251.50.54:8032 (/health is open; POSTs need Authorization: Bearer $(secret get semif/api-token))
Model Qwen/Qwen3.5-4B @ 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a, BF16, from /tank/aimodels/huggingface (read-only, offline)
SemIf commit 23cf1f39fc9534fe81437200959b6dfc7106e45a; torch 2.10.0+cu128, transformers 5.17.0, the same stack SemIf's committed predictions were made on
Image semif-serve:<version>, built on fv-ml1 from services/semif-serve/
State none. If the weights are ever lost, re-pull them by hand at the pinned revision (the service itself never downloads: HF_HUB_OFFLINE=1, read-only mount). No backup needed beyond the host's /opt/docker restic.

API

T=$(secret get semif/api-token)
curl -s -H "Authorization: Bearer $T" http://10.251.50.54:8032/decide -d '{
  "id": "q1", "state": "Health checks passed in all three zones.",
  "question": "Did the deployment succeed?",
  "options": [{"id":"yes","description":"It succeeded."},{"id":"no","description":"It failed."}]}'
  • POST /decide: one decision, returning SemIf's result dict unchanged (option_ids, probabilities, option_logits, prompt_sha256, model, …).
  • POST /decide/shared: {state, decisions: [{id, question, options}]}. It prefills the state once and scores every criterion in one batch. Use it when many questions share one long state.
  • GET /health: the pins, limits and calibrated workloads.
  • Order averaging (0.1.3): add "orderings": "rotations" to a decision, or "all" for ≤ 4 options. The option list is asked in every rotation inside ONE shared batch. The reply keeps each ordering's native result under orderings and adds combined: {method, orderings, probabilities, top, agreement, spread}. Use it for anything real: a small model leans toward the first-listed option on ambiguous inputs, and averaging cancels that. Through the service on SemIf's labelled sets, accuracy goes from 78.6% to 88.1% (group-bootstrap 95% CI +5.1..+14.3 pts, 252 rows). agreement is the cheap confidence signal: unanimous rows are 94.5% accurate, split rows 76.4%. The orderings count toward the decision cap. workload calibration is not available together with orderings yet (422).
  • Past MAX_QUEUE (32) requests in progress, new POSTs get 429 busy before their body is read.

⚠ Probabilities are uncalibrated

SemIf labels its output "conditional option score; uncalibrated as decision confidence", and means it: on WANLI the model is right ~64% of the time while reporting far higher confidence. Before a caller thresholds on probabilities, it brings labelled rows (≥ ~150) for its workload. We fit one temperature T with SemIf's benchmarks/calibrate.py, add {"<workload>": T} to /opt/docker/conf/semif/calibration.json (stacks/semif/conf/), and restart. The caller then passes "workload": "<name>" and gets a calibrated block beside the native scores. The argmax never changes.

VRAM: a hard cap, released after every burst

VRAM_CAP_GIB=12 becomes torch.cuda.set_per_process_memory_fraction before the weights load. At rest the process holds ~8.7 GB (nvidia-smi); the weights are 7.84 GiB. After any call that grows torch's reserved memory past the post-warm-up baseline

  • 512 MiB, the engine calls empty_cache(), so a burst returns to the card and does not squeeze scriberr, which shares GPU 1. A request that would exceed the cap gets 503 out_of_memory, memory returns to baseline, and the service stays up. Both behaviours were verified on the card (0.1.0 held 11.9 GiB after an OOM, and 12.6 GB after a large request; 0.1.2 returns to 7.85 GiB in both cases).

What fits under 12 GiB (measured on 0.1.3, /decide/shared, binary decisions; 0.1.2 figures in brackets, before the fast kernels):

state size (prefix tokens) max rows in one request
~140 63 (52)
~520 51 (43)
~1,960 26 (19)
~3,900 16 (13)

Rows = decisions × orderings, so rotations over 3 options uses 3 rows per decision.

/decide fits at the full 4,096-token limit. Past the table you get a 503, so split the decisions across requests.

Fast kernels (0.1.3)

The image ships Qwen3.5's fast kernels, flash-linear-attention 0.5.2 and causal-conv1d 1.7.0 (a prebuilt cu12/torch2.10 wheel). Without them transformers logs that it falls back to "much slower" reference PyTorch paths. They are adopted because an A/B on the empty GPU 3 showed:

  • parity improved: 144/144 vs upstream (142/144 without), so upstream evidently ran with them;
  • long inputs got much faster: a ~2,000-token /decide went from 169 to 92 ms server-side.

Short 3-rotation batches cost ~3–6 ms more; everything else is equal or faster. ⚠ triton builds a C shim at runtime, so the image carries gcc. Without it the warm-up fails, and startup fails closed. Build without the kernels: --build-arg EXTRAS="--extra model".

Latency (0.1.3, from nh3-dev, 3 runs × 20, network floor ~33 ms)

request end to end server
/decide, short (~130 tok) 69 ms 35 ms
/decide, ~2,000-token state 131 ms 92 ms
3 rotations, short 115 ms 78 ms
6 orderings, short 118 ms 81 ms
3 rotations, ~2,000-token state 200 ms 158 ms

Acceptance (2026-09-27, v0.1.3)

Raw: services/semif-serve/acceptance/result-2026-09-27-v0.1.3.json, averaging-2026-09-27-v0.1.3.json. v0.1.3 matches upstream on 144/144 (identical prompt hashes, max prob gap 0.049), is deterministic, fails the negative control as it should (14/144), and shared matches direct on 72/72. The v0.1.2 table below is the reference-kernel baseline.

Acceptance (2026-09-27, v0.1.2)

Raw: services/semif-serve/acceptance/result-2026-09-27-v0.1.2.json.

check result
parity with SemIf's committed torch predictions (authored144) 142/144 same top choice; 144/144 identical prompt SHA-256; max prob gap 0.093
noise floor (same 144 twice) 144/144, gap 0.0: deterministic
negative control (option descriptions rotated) 14/144: the check catches a wrong answer
shared vs direct (36 shared states, 72 rows) 72/72, max gap 0.045
21 binary criteria over one state, from nh3-dev shared 159 ms (3 runs, 159–160) vs 21 sequential calls 981 ms

The two parity misses are exact bf16 ties in our output (top-2 margin 0.000), where upstream's prefix-cache path gave margins of 0.054 and 0.185; one of the two now matches the label. So the misses come from the numeric path, not the wrapper. The speed figure includes one network round trip (~27 ms). The burst release costs ~16 ms on it (0.1.1 measured 143 ms without the release).

Building

# from nh3-dev
tar -C services/semif-serve -cf - --exclude=.venv --exclude=.pytest_cache --exclude=__pycache__ --exclude=acceptance . \
  | ssh infra-ops@10.251.50.54 'mkdir -p /opt/docker/src/semif-serve-X.Y.Z && tar -x -C /opt/docker/src/semif-serve-X.Y.Z'
# on fv-ml1
cd /opt/docker/src/semif-serve-X.Y.Z && docker build -t semif-serve:X.Y.Z .

To move SemIf forward: bump the commit in services/semif-serve/pyproject.toml (and SEMIF_COMMIT in config.py), run uv lock, rebuild, and re-run the acceptance. Upstream is research code that changes weekly, which is why it is pinned.