Order averaging (Prime, after the 739aa03 spike):
- A decision may set orderings: rotations|all (all only for <= 4 options). Every
ordering goes to the engine in one shared batch.
- The reply keeps each native result and adds combined {probabilities (log-mean),
top, agreement, spread}.
- Through the service on SemIf's labelled sets (252 rows): 78.6% -> 88.1%
(group-bootstrap 95% CI +5.1..+14.3). Unanimous agreement is 94.5% accurate.
Fast kernels: flash-linear-attention 0.5.2 and causal-conv1d 1.7.0 are now the
default build. A/B on the empty GPU 3:
- parity with upstream went from 142/144 to 144/144;
- a ~2k-token /decide went from 169 to 92 ms server-side;
- short 3-rotation batches cost ~3-6 ms more.
triton builds a C shim at runtime, so the image carries gcc. Without it the
warm-up failed and startup failed closed.
Heid bug-hunt panel (4/4 arms, thread 01M3H3F4RR7XBP90KQ3A39H4SX), folded:
- Startup validation: VRAM cap 0 no longer means uncapped (C1); limits must be
>= 1 (S1); the token must be visible ASCII (S2); the calibration file must
exist and parse, with T in [0.05, 20] (S8, and C3's NaN leg).
- The body limit is checked before a chunk is kept, and a Unicode-digit
Content-Length no longer crashes (C2, S3).
- Failures while building the response now get the 500 envelope (C3).
- 429 busy past SEMIF_MAX_QUEUE requests in progress (C6).
- The engine releases memory on every non-validation failure, unchained after
gc; an empty OOM message is handled; 'out of memory' RuntimeErrors map to 503
(C4, C5, S9).
- The entry point forces HF_HUB_OFFLINE (S10). README wording fixed (S5, S6).
- New guard tests close the gaps the arms' mutation grids exposed: early stop of
the body read, a shared-route lock, calibration pass-through, the gc cycle,
the exact caps, TorchEngine.load's arch and device checks, and the offline
entry point.
86 tests.
Deployed on fv-ml1 GPU 1: parity 144/144, OOM and burst release verified, shared
capacity 63/51/26/16 rows at ~140/520/1960/3900 prefix tokens.
67 lines
2.7 KiB
YAML
67 lines
2.7 KiB
YAML
# semif: SemIf option-logit decisions (github TheoLeeCJ/SemIf-OpenJev, MIT) behind semif-serve,
|
|
# on fv-ml1 GPU 1 (the utility card, beside vllm-coder and scriberr). Prime, 2026-09-27.
|
|
#
|
|
# One forward pass of a pinned Qwen3.5-4B (BF16) per decision; the answer is read from the
|
|
# option-letter logits, so there is no decoding. Service code + contract:
|
|
# services/semif-serve/ (semif-serve.contract.md). Image built on fv-ml1 from that dir.
|
|
#
|
|
# ⚠ Scores are "conditional option score; uncalibrated as decision confidence". A caller
|
|
# that needs thresholds brings labelled rows; we fit a per-workload temperature into
|
|
# conf/calibration.json and the caller passes `workload`. See the README.
|
|
# ⚠ SEMIF_VRAM_CAP_GIB is a HARD cap (torch per-process memory fraction), sized from a
|
|
# measured peak, so SemIf cannot squeeze scriberr or the vLLM seats on this card. A
|
|
# request that needs more gets 503 out_of_memory and the service stays up.
|
|
#
|
|
# .env (tunables): IMAGE, PORT, GPU_ID, VRAM_CAP_GIB, HOST_IP, SEMIF_API_TOKEN (vault
|
|
# semif/api-token, >= 32 chars).
|
|
|
|
name: semif
|
|
|
|
services:
|
|
semif:
|
|
image: ${IMAGE:?set IMAGE}
|
|
container_name: semif
|
|
restart: unless-stopped
|
|
ports:
|
|
- "${PORT:-8032}:8000"
|
|
environment:
|
|
SEMIF_API_TOKEN: ${SEMIF_API_TOKEN:?set SEMIF_API_TOKEN}
|
|
SEMIF_DEVICE: cuda
|
|
SEMIF_VRAM_CAP_GIB: ${VRAM_CAP_GIB:?set VRAM_CAP_GIB}
|
|
SEMIF_MAX_TOKENS: ${MAX_TOKENS:-4096}
|
|
SEMIF_MAX_DECISIONS: ${MAX_DECISIONS:-64}
|
|
# POSTs in progress (queued + scoring) before new ones get 429 busy.
|
|
SEMIF_MAX_QUEUE: ${MAX_QUEUE:-32}
|
|
SEMIF_CALIBRATION: /conf/calibration.json
|
|
volumes:
|
|
# Pinned weights, read offline (HF_HUB_OFFLINE=1 in the image). Never downloads.
|
|
- /tank/aimodels/huggingface:/hf:ro
|
|
- /opt/docker/conf/semif:/conf:ro
|
|
deploy:
|
|
resources:
|
|
reservations:
|
|
devices:
|
|
- driver: nvidia
|
|
device_ids: ["${GPU_ID:-1}"]
|
|
capabilities: [gpu]
|
|
healthcheck:
|
|
test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=5)"]
|
|
interval: 30s
|
|
timeout: 10s
|
|
retries: 3
|
|
# Startup loads ~9 GB of weights and scores one warm-up decision before it serves.
|
|
start_period: 300s
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=AI - Eval & Retrieval
|
|
- homepage.name=SemIf — option-logit decisions
|
|
- homepage.icon=mdi-scale-balance
|
|
- homepage.description=Typed decisions from one forward pass (Qwen3.5-4B, fv-ml1 GPU1)
|
|
- homepage.href=http://${HOST_IP:-10.251.50.54}:${PORT:-8032}/health
|
|
|
|
networks:
|
|
tnet:
|
|
name: traefik-net
|
|
external: true
|