feat(semif): SemIf option-logit decisions on fv-ml1 GPU 1 (Prime)
services/semif-serve is a FastAPI wrapper around SemIf's direct and shared torch scorers (SemIf-OpenJev @ 23cf1f39, MIT). Upstream ships only a batch CLI. The wrapper loads the pinned Qwen3.5-4B (851bf6e8, BF16) once from the offline HF cache and returns SemIf's result dicts unchanged, with an optional per-workload temperature-calibrated view. Contract: semif-serve.contract.md. Built with a short contract, TDD (39 tests, fake engine and fake torch, no GPU) and a heid bug-hunt panel (pending). On the card: - torch 2.10.0+cu128 with sm_120 kernels, which is SemIf's own stack; - a hard 12 GiB VRAM cap. Two defects surfaced only on the card, and each fix is covered by a test: - 0.1.1: an OOM raised as a chained exception kept the failed request's tensors alive (11.9 GiB after the 503). It is now raised unchained, after gc. - 0.1.2: a large request left 12.6 GB reserved on the shared card. After each call, reserved memory over the baseline + 512 MiB is now released. Acceptance against SemIf's committed torch predictions (authored144): - 142/144 same top choice; both misses are exact bf16 ties; - 144/144 identical prompt hashes; - deterministic A-vs-A; - negative control 14/144; - shared vs direct 72/72. 21 binary criteria over one state take 159 ms. The shared-mode capacity table under the cap is in stacks/semif/README.md. The Dockerfile installs dependencies from a manifest with the project version blanked, so a version bump reuses the ~4 GB torch layer. Verified: 41 s rebuild, dependency layer CACHED. DNS: semif.fv.internal. Token: vault semif/api-token.
This commit is contained in:
@@ -0,0 +1,11 @@
|
||||
# semif — copy to /opt/docker/compose/semif/.env on fv-ml1 (mode 0600).
|
||||
# Built on fv-ml1 from services/semif-serve (see README "Building").
|
||||
IMAGE=semif-serve:0.1.2
|
||||
PORT=8032
|
||||
HOST_IP=10.251.50.54
|
||||
# fv-ml1 GPU 1 = the utility card (vllm-coder, erp, meromero, scriberr).
|
||||
GPU_ID=1
|
||||
# HARD per-process cap, from the measured peak (README "VRAM").
|
||||
VRAM_CAP_GIB=12
|
||||
# >= 32 characters; source of truth: secret get semif/api-token
|
||||
SEMIF_API_TOKEN=
|
||||
@@ -0,0 +1,103 @@
|
||||
# semif
|
||||
|
||||
**SemIf option-logit decisions** on **fv-ml1 GPU 1**, the utility card beside
|
||||
`vllm-coder`, the erp/meromero seats and scriberr. Deployed 2026-09-27 at Prime's
|
||||
request. Hosting assessment: infra-hermes, 2026-09-25.
|
||||
|
||||
SemIf ([TheoLeeCJ/SemIf-OpenJev](https://github.com/TheoLeeCJ/SemIf-OpenJev), MIT)
|
||||
asks a small model a typed question and reads the answer straight from the logits of
|
||||
the option letters, after **one forward pass with no decoding**. Upstream ships only a
|
||||
batch CLI, so `services/semif-serve/` wraps its two torch scorers in a small FastAPI
|
||||
service. The contract is `services/semif-serve/semif-serve.contract.md`.
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| **URL** | `http://10.251.50.54:8032` (`/health` is open; POSTs need `Authorization: Bearer $(secret get semif/api-token)`) |
|
||||
| **Model** | `Qwen/Qwen3.5-4B` @ `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a`, BF16, from `/tank/aimodels/huggingface` (read-only, offline) |
|
||||
| **SemIf** | commit `23cf1f39fc9534fe81437200959b6dfc7106e45a`; torch `2.10.0+cu128`, transformers `5.17.0`, the same stack SemIf's committed predictions were made on |
|
||||
| **Image** | `semif-serve:<version>`, built on fv-ml1 from `services/semif-serve/` |
|
||||
| **State** | none. Weights re-download at the pinned revision. No backup needed beyond the host's `/opt/docker` restic. |
|
||||
|
||||
## API
|
||||
|
||||
```bash
|
||||
T=$(secret get semif/api-token)
|
||||
curl -s -H "Authorization: Bearer $T" http://10.251.50.54:8032/decide -d '{
|
||||
"id": "q1", "state": "Health checks passed in all three zones.",
|
||||
"question": "Did the deployment succeed?",
|
||||
"options": [{"id":"yes","description":"It succeeded."},{"id":"no","description":"It failed."}]}'
|
||||
```
|
||||
|
||||
- `POST /decide`: one decision, returning SemIf's result dict unchanged (`option_ids`,
|
||||
`probabilities`, `option_logits`, `prompt_sha256`, `model`, …).
|
||||
- `POST /decide/shared`: `{state, decisions: [{id, question, options}]}`. It prefills
|
||||
the state once and scores every criterion in one batch. Use it when many questions
|
||||
share one long state.
|
||||
- `GET /health`: the pins, limits and calibrated workloads.
|
||||
|
||||
## ⚠ Probabilities are uncalibrated
|
||||
|
||||
SemIf labels its output `"conditional option score; uncalibrated as decision
|
||||
confidence"`, and means it: on WANLI the model is right ~64% of the time while
|
||||
reporting far higher confidence. **Before a caller thresholds on `probabilities`, it
|
||||
brings labelled rows (≥ ~150) for its workload.** We fit one temperature `T` with
|
||||
SemIf's `benchmarks/calibrate.py`, add `{"<workload>": T}` to
|
||||
`/opt/docker/conf/semif/calibration.json` (`stacks/semif/conf/`), and restart. The
|
||||
caller then passes `"workload": "<name>"` and gets a `calibrated` block beside the
|
||||
native scores. The argmax never changes.
|
||||
|
||||
## VRAM: a hard cap, released after every burst
|
||||
|
||||
`VRAM_CAP_GIB=12` becomes `torch.cuda.set_per_process_memory_fraction` before the
|
||||
weights load. At rest the process holds **~8.7 GB** (nvidia-smi); the weights are 7.84
|
||||
GiB. After any call that grows torch's reserved memory past the post-warm-up baseline
|
||||
+ 512 MiB, the engine calls `empty_cache()`, so a burst returns to the card and does
|
||||
not squeeze scriberr, which shares GPU 1. A request that would exceed the cap gets
|
||||
`503 out_of_memory`, memory returns to baseline, and the service stays up. Both
|
||||
behaviours were verified on the card (0.1.0 held 11.9 GiB after an OOM, and 12.6 GB
|
||||
after a large request; 0.1.2 returns to 7.85 GiB in both cases).
|
||||
|
||||
**What fits under 12 GiB** (measured, `/decide/shared`, binary decisions):
|
||||
|
||||
| state size (prefix tokens) | max decisions in one request |
|
||||
|---|---|
|
||||
| ~140 | 52 |
|
||||
| ~520 | 43 |
|
||||
| ~1,960 | 19 |
|
||||
| ~3,900 | 13 |
|
||||
|
||||
`/decide` fits at the full 4,096-token limit. Past the table you get a 503, so split
|
||||
the decisions across requests.
|
||||
|
||||
## Acceptance (2026-09-27, v0.1.2)
|
||||
|
||||
Raw: `services/semif-serve/acceptance/result-2026-09-27-v0.1.2.json`.
|
||||
|
||||
| check | result |
|
||||
|---|---|
|
||||
| parity with SemIf's committed torch predictions (authored144) | **142/144** same top choice; **144/144** identical prompt SHA-256; max prob gap 0.093 |
|
||||
| noise floor (same 144 twice) | 144/144, gap 0.0: deterministic |
|
||||
| negative control (option descriptions rotated) | 14/144: the check catches a wrong answer |
|
||||
| shared vs direct (36 shared states, 72 rows) | 72/72, max gap 0.045 |
|
||||
| 21 binary criteria over one state, from nh3-dev | shared **159 ms** (3 runs, 159–160) vs 21 sequential calls 981 ms |
|
||||
|
||||
The two parity misses are **exact bf16 ties in our output** (top-2 margin 0.000),
|
||||
where upstream's prefix-cache path gave margins of 0.054 and 0.185; one of the two
|
||||
now matches the label. So the misses come from the numeric path, not the wrapper. The
|
||||
speed figure includes one network round trip (~27 ms). The burst release costs ~16 ms
|
||||
on it (0.1.1 measured 143 ms without the release).
|
||||
|
||||
## Building
|
||||
|
||||
```bash
|
||||
# from nh3-dev
|
||||
tar -C services/semif-serve -cf - --exclude=.venv --exclude=.pytest_cache --exclude=__pycache__ --exclude=acceptance . \
|
||||
| ssh infra-ops@10.251.50.54 'mkdir -p /opt/docker/src/semif-serve-X.Y.Z && tar -x -C /opt/docker/src/semif-serve-X.Y.Z'
|
||||
# on fv-ml1
|
||||
cd /opt/docker/src/semif-serve-X.Y.Z && docker build -t semif-serve:X.Y.Z .
|
||||
```
|
||||
|
||||
To move SemIf forward: bump the commit in `services/semif-serve/pyproject.toml`
|
||||
(and `SEMIF_COMMIT` in `config.py`), run `uv lock`, rebuild, and **re-run the
|
||||
acceptance**. Upstream is research code that changes weekly, which is why it is
|
||||
pinned.
|
||||
@@ -0,0 +1,64 @@
|
||||
# semif: SemIf option-logit decisions (github TheoLeeCJ/SemIf-OpenJev, MIT) behind semif-serve,
|
||||
# on fv-ml1 GPU 1 (the utility card, beside vllm-coder and scriberr). Prime, 2026-09-27.
|
||||
#
|
||||
# One forward pass of a pinned Qwen3.5-4B (BF16) per decision; the answer is read from the
|
||||
# option-letter logits, so there is no decoding. Service code + contract:
|
||||
# services/semif-serve/ (semif-serve.contract.md). Image built on fv-ml1 from that dir.
|
||||
#
|
||||
# ⚠ Scores are "conditional option score; uncalibrated as decision confidence". A caller
|
||||
# that needs thresholds brings labelled rows; we fit a per-workload temperature into
|
||||
# conf/calibration.json and the caller passes `workload`. See the README.
|
||||
# ⚠ SEMIF_VRAM_CAP_GIB is a HARD cap (torch per-process memory fraction), sized from a
|
||||
# measured peak, so SemIf cannot squeeze scriberr or the vLLM seats on this card. A
|
||||
# request that needs more gets 503 out_of_memory and the service stays up.
|
||||
#
|
||||
# .env (tunables): IMAGE, PORT, GPU_ID, VRAM_CAP_GIB, HOST_IP, SEMIF_API_TOKEN (vault
|
||||
# semif/api-token, >= 32 chars).
|
||||
|
||||
name: semif
|
||||
|
||||
services:
|
||||
semif:
|
||||
image: ${IMAGE:?set IMAGE}
|
||||
container_name: semif
|
||||
restart: unless-stopped
|
||||
ports:
|
||||
- "${PORT:-8032}:8000"
|
||||
environment:
|
||||
SEMIF_API_TOKEN: ${SEMIF_API_TOKEN:?set SEMIF_API_TOKEN}
|
||||
SEMIF_DEVICE: cuda
|
||||
SEMIF_VRAM_CAP_GIB: ${VRAM_CAP_GIB:?set VRAM_CAP_GIB}
|
||||
SEMIF_MAX_TOKENS: ${MAX_TOKENS:-4096}
|
||||
SEMIF_MAX_DECISIONS: ${MAX_DECISIONS:-64}
|
||||
SEMIF_CALIBRATION: /conf/calibration.json
|
||||
volumes:
|
||||
# Pinned weights, read offline (HF_HUB_OFFLINE=1 in the image). Never downloads.
|
||||
- /tank/aimodels/huggingface:/hf:ro
|
||||
- /opt/docker/conf/semif:/conf:ro
|
||||
deploy:
|
||||
resources:
|
||||
reservations:
|
||||
devices:
|
||||
- driver: nvidia
|
||||
device_ids: ["${GPU_ID:-1}"]
|
||||
capabilities: [gpu]
|
||||
healthcheck:
|
||||
test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=5)"]
|
||||
interval: 30s
|
||||
timeout: 10s
|
||||
retries: 3
|
||||
# Startup loads ~9 GB of weights and scores one warm-up decision before it serves.
|
||||
start_period: 300s
|
||||
networks:
|
||||
- tnet
|
||||
labels:
|
||||
- homepage.group=AI - Eval & Retrieval
|
||||
- homepage.name=SemIf — option-logit decisions
|
||||
- homepage.icon=mdi-scale-balance
|
||||
- homepage.description=Typed decisions from one forward pass (Qwen3.5-4B, fv-ml1 GPU1)
|
||||
- homepage.href=http://${HOST_IP:-10.251.50.54}:${PORT:-8032}/health
|
||||
|
||||
networks:
|
||||
tnet:
|
||||
name: traefik-net
|
||||
external: true
|
||||
@@ -0,0 +1 @@
|
||||
{}
|
||||
Reference in New Issue
Block a user