Files
esh-pfi-infrastructure/stacks/semif/README.md
T

13 KiB
Raw Blame History

semif

⚠ REPLACED by intern-decision (Prime, 2026-09-30). Prime ruled that morning: "replace semif with intern-decision now". The replacement is stacks/intern-decision (Intern-Decision-4B, live since 0941 PT, http://10.251.50.54:8033, intern-decision.fv.internal). It keeps this service's HTTP surface (/decide, /decide/shared, /health), so callers only change the URL and the token (secret get intern-decision/api-token). Its README lists every deliberate difference.

The semif container was stopped at 0135 PT that day, to give scriberr its GPU 1 headroom back. It was removed at 0949 PT (docker compose down) so that Homepage stops listing a dead tile. The image (semif-serve:0.1.4), the weights, this stack's files and the token are all kept. Port 8032 and semif.fv.internal stay reserved for it.

Rollback, only on Prime's word.

  1. Stop intern-decision first. The two services do not fit GPU 1 together next to scriberr. semif held 9.2 GB at rest and peaked at 12.9 GB. cd /opt/docker/compose/intern-decision && docker compose stop.
  2. Recreate semif: cd /opt/docker/compose/semif && docker compose up -d (the container is removed, so start will not work).
  3. Check nvidia-smi Free on GPU 1 against semif's 12.9 GB peak plus scriberr's 5.5 GB (its 120 s slices) before calling it done. If it does not cover both, semif collides with scriberr.

Everything below describes the service as it was deployed.

SemIf option-logit decisions on fv-ml1 GPU 1, the utility card beside vllm-coder, the erp/meromero seats and scriberr. Deployed 2026-09-27 at Prime's request. Hosting assessment: infra-hermes, 2026-09-25.

SemIf (TheoLeeCJ/SemIf-OpenJev, MIT) asks a small model a typed question and reads the answer straight from the logits of the option letters, after one forward pass with no decoding. Upstream ships only a batch CLI, so services/semif-serve/ wraps its two torch scorers in a small FastAPI service. The contract is services/semif-serve/semif-serve.contract.md.

URL http://10.251.50.54:8032 (/health is open; POSTs need Authorization: Bearer $(secret get semif/api-token))
Model Qwen/Qwen3.5-4B @ 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a, BF16, from /tank/aimodels/huggingface (read-only, offline)
SemIf commit 23cf1f39fc9534fe81437200959b6dfc7106e45a; torch 2.10.0+cu128, transformers 5.17.0, the same stack SemIf's committed predictions were made on
Image semif-serve:<version>, built on fv-ml1 from services/semif-serve/
State none. If the weights are ever lost, re-pull them by hand at the pinned revision (the service itself never downloads: HF_HUB_OFFLINE=1, read-only mount). No backup needed beyond the host's /opt/docker restic.

API

T=$(secret get semif/api-token)
curl -s -H "Authorization: Bearer $T" http://10.251.50.54:8032/decide -d '{
  "id": "q1", "state": "Health checks passed in all three zones.",
  "question": "Did the deployment succeed?",
  "options": [{"id":"yes","description":"It succeeded."},{"id":"no","description":"It failed."}]}'
  • POST /decide: one decision, returning SemIf's result dict unchanged (option_ids, probabilities, option_logits, prompt_sha256, model, …).
  • POST /decide/shared: {state, decisions: [{id, question, options}]}. It prefills the state once and scores every criterion in one batch. Use it when many questions share one long state.
  • GET /health: the pins, limits and calibrated workloads.
  • Order averaging (0.1.3): add "orderings": "rotations" to a decision, or "all" for ≤ 4 options. The option list is asked in every rotation inside ONE shared batch. The reply keeps each ordering's native result under orderings and adds combined: {method, orderings, probabilities, top, agreement, spread}. Use it for anything real: a small model leans toward the first-listed option on ambiguous inputs, and averaging cancels that. Through the service on SemIf's labelled sets, accuracy goes from 78.6% to 88.1% (group-bootstrap 95% CI +5.1..+14.3 pts, 252 rows). agreement is the cheap confidence signal: unanimous rows are 94.5% accurate, split rows 76.4%. The orderings count toward the decision cap. workload calibration is not available together with orderings yet (422).
  • Past MAX_QUEUE (32) requests in progress, new POSTs get 429 busy before their body is read.
  • ⚠ Rotations cost options², not options. The options live in each row's suffix, and the shared prefix is only the state. So rotations over n options sends n rows each carrying all n options. Measured 2026-09-27 on a short state, over 16 options of about 40 tokens each: 850 ms with rotations against 109 ms for one ordering (10,304 against 644 suffix tokens, 5 calls after warm-up). One cold call at that size returned 503. The VRAM table below covers binary decisions only. For many options, use one ordering or shorter option text.
  • Object states ending in ), ; or } work as of 0.1.4. Until then, if the state was an object whose last value ended in one of those characters, the service returned 422 "The fixed state prefix does not match every full prompt". SemIf trims one token at the state boundary, and the JSON that follows re-merged two tokens back. The fix (contract INV-7) moves only where the shared prefix ends; every row still scores the same tokens. Measured with the real tokenizer: of 154 states (22 endings × 7 shapes) SemIf alone refused 23, and with the fix none. None of the 131 ordinary states, nor any of the authored144 states, changed its prefix. Startup proves the fix is live by scoring {"person_said": "ok :)"} through the shared path.

⚠ Probabilities are uncalibrated

SemIf labels its output "conditional option score; uncalibrated as decision confidence", and means it: on WANLI the model is right ~64% of the time while reporting far higher confidence. Before a caller thresholds on probabilities, it brings labelled rows (≥ ~150) for its workload. We fit one temperature T with SemIf's benchmarks/calibrate.py, add {"<workload>": T} to /opt/docker/conf/semif/calibration.json (stacks/semif/conf/), and restart. The caller then passes "workload": "<name>" and gets a calibrated block beside the native scores. The argmax never changes.

VRAM: a hard cap, released after every burst

VRAM_CAP_GIB=12 becomes torch.cuda.set_per_process_memory_fraction before the weights load. At rest the process holds ~8.7 GB (nvidia-smi); the weights are 7.84 GiB. After any call that grows torch's reserved memory past the post-warm-up baseline

  • 512 MiB, the engine calls empty_cache(), so a burst returns to the card and does not squeeze scriberr, which shares GPU 1. A request that would exceed the cap gets 503 out_of_memory, memory returns to baseline, and the service stays up. Both behaviours were verified on the card (0.1.0 held 11.9 GiB after an OOM, and 12.6 GB after a large request; 0.1.2 returns to 7.85 GiB in both cases).

What fits under 12 GiB (measured on 0.1.3, /decide/shared, binary decisions; 0.1.2 figures in brackets, before the fast kernels):

state size (prefix tokens) max rows in one request
~140 63 (52)
~520 51 (43)
~1,960 26 (19)
~3,900 16 (13)

Rows = decisions × orderings, so rotations over 3 options uses 3 rows per decision.

/decide fits at the full 4,096-token limit. Past the table you get a 503, so split the decisions across requests.

Fast kernels (0.1.3)

The image ships Qwen3.5's fast kernels, flash-linear-attention 0.5.2 and causal-conv1d 1.7.0 (a prebuilt cu12/torch2.10 wheel). Without them transformers logs that it falls back to "much slower" reference PyTorch paths. They are adopted because an A/B on the empty GPU 3 showed:

  • parity improved: 144/144 vs upstream (142/144 without), so upstream evidently ran with them;
  • long inputs got much faster: a ~2,000-token /decide went from 169 to 92 ms server-side.

Short 3-rotation batches cost ~3–6 ms more; everything else is equal or faster. ⚠ triton builds a C shim at runtime, so the image carries gcc. Without it the warm-up fails, and startup fails closed. Build without the kernels: --build-arg EXTRAS="--extra model".

Latency (0.1.3, from nh3-dev, 3 runs × 20, network floor ~33 ms)

request end to end server
/decide, short (~130 tok) 69 ms 35 ms
/decide, ~2,000-token state 131 ms 92 ms
3 rotations, short 115 ms 78 ms
6 orderings, short 118 ms 81 ms
3 rotations, ~2,000-token state 200 ms 158 ms

Acceptance (2026-09-27, v0.1.4)

Raw: services/semif-serve/acceptance/result-2026-09-27-v0.1.4.json. Parity with upstream 144/144 (identical prompt hashes, max prob gap 0.060). Deterministic within the process (A-vs-A gap 0.0). The negative control fails as it should (14/144). Shared vs direct 71/72. The one miss is an exact bf16 tie in the shared result (0.444/0.444). INV-7 did not move that row's prefix. After a plain docker restart of the same image, the row read 0.369/0.537 and agreed.

⚠ So "deterministic" holds within one process, not across restarts. Logits come in bf16 steps (0.125 here), and a near-tie can land differently after a restart, most likely because the fast kernels autotune at startup. That is n=1 row across one restart, and the cross-restart floor is otherwise unmeasured. When comparing two versions, compare them against that floor and not against zero.

Update 2026-09-30 (the Jev bench on GPU 3, docs/pfi/jev-candidates-bench-2026-09-30.md): across 4 restarts of the same image, SemIf changed 0 labels, and its probabilities were bit-identical on every set, the 144 authored rows included. So the near-tie flip described above did not reproduce. The cross-restart floor is now measured at 0 for n=4 restarts on an idle card. Treat the one flip on 2026-09-27 as possible but rare, not as the norm.

Acceptance (2026-09-27, v0.1.3)

Raw: services/semif-serve/acceptance/result-2026-09-27-v0.1.3.json, averaging-2026-09-27-v0.1.3.json. v0.1.3 matches upstream on 144/144 (identical prompt hashes, max prob gap 0.049), is deterministic, fails the negative control as it should (14/144), and shared matches direct on 72/72. The v0.1.2 table below is the reference-kernel baseline.

Acceptance (2026-09-27, v0.1.2)

Raw: services/semif-serve/acceptance/result-2026-09-27-v0.1.2.json.

check result
parity with SemIf's committed torch predictions (authored144) 142/144 same top choice; 144/144 identical prompt SHA-256; max prob gap 0.093
noise floor (same 144 twice) 144/144, gap 0.0: deterministic
negative control (option descriptions rotated) 14/144: the check catches a wrong answer
shared vs direct (36 shared states, 72 rows) 72/72, max gap 0.045
21 binary criteria over one state, from nh3-dev shared 159 ms (3 runs, 159–160) vs 21 sequential calls 981 ms

The two parity misses are exact bf16 ties in our output (top-2 margin 0.000), where upstream's prefix-cache path gave margins of 0.054 and 0.185; one of the two now matches the label. So the misses come from the numeric path, not the wrapper. The speed figure includes one network round trip (~27 ms). The burst release costs ~16 ms on it (0.1.1 measured 143 ms without the release).

Building and deploying

# from nh3-dev. /opt/docker/src is root-owned, so create the version dir with sudo first.
ssh infra-ops@10.251.50.54 'sudo -n install -d -o infra-ops -g infra-ops /opt/docker/src/semif-serve-X.Y.Z'
tar -C services/semif-serve -cf - --exclude=.venv --exclude=.pytest_cache --exclude=__pycache__ \
    --exclude=acceptance --exclude=spike --exclude='*.egg-info' . \
  | ssh infra-ops@10.251.50.54 'tar -x -C /opt/docker/src/semif-serve-X.Y.Z'
# on fv-ml1
cd /opt/docker/src/semif-serve-X.Y.Z && docker build -t semif-serve:X.Y.Z .
# ⚠ /opt/docker/compose/semif is root-owned, so `sed -i` cannot write its temp file there. .env
#   itself is infra-ops's: rewrite it IN PLACE, which keeps its owner and mode (0600).
cd /opt/docker/compose/semif && new=$(sed 's/^IMAGE=.*/IMAGE=semif-serve:X.Y.Z/' .env) \
  && printf '%s\n' "$new" > .env && docker compose up -d

Startup fails closed (INV-3, INV-7), so a container that does not reach healthy did not pass its own checks: read docker logs semif. Record the deploy with scripts/ops-log.

To move SemIf forward: bump the commit in services/semif-serve/pyproject.toml (and SEMIF_COMMIT in config.py), run uv lock, rebuild, and re-run the acceptance. Upstream is research code that changes weekly, which is why it is pinned.