feat(semif): SemIf option-logit decisions on fv-ml1 GPU 1 (Prime)
services/semif-serve is a FastAPI wrapper around SemIf's direct and shared torch scorers (SemIf-OpenJev @ 23cf1f39, MIT). Upstream ships only a batch CLI. The wrapper loads the pinned Qwen3.5-4B (851bf6e8, BF16) once from the offline HF cache and returns SemIf's result dicts unchanged, with an optional per-workload temperature-calibrated view. Contract: semif-serve.contract.md. Built with a short contract, TDD (39 tests, fake engine and fake torch, no GPU) and a heid bug-hunt panel (pending). On the card: - torch 2.10.0+cu128 with sm_120 kernels, which is SemIf's own stack; - a hard 12 GiB VRAM cap. Two defects surfaced only on the card, and each fix is covered by a test: - 0.1.1: an OOM raised as a chained exception kept the failed request's tensors alive (11.9 GiB after the 503). It is now raised unchained, after gc. - 0.1.2: a large request left 12.6 GB reserved on the shared card. After each call, reserved memory over the baseline + 512 MiB is now released. Acceptance against SemIf's committed torch predictions (authored144): - 142/144 same top choice; both misses are exact bf16 ties; - 144/144 identical prompt hashes; - deterministic A-vs-A; - negative control 14/144; - shared vs direct 72/72. 21 binary criteria over one state take 159 ms. The shared-mode capacity table under the cap is in stacks/semif/README.md. The Dockerfile installs dependencies from a manifest with the project version blanked, so a version bump reuses the ~4 GB torch layer. Verified: 41 s rebuild, dependency layer CACHED. DNS: semif.fv.internal. Token: vault semif/api-token.
This commit is contained in:
@@ -0,0 +1,64 @@
|
||||
# semif: SemIf option-logit decisions (github TheoLeeCJ/SemIf-OpenJev, MIT) behind semif-serve,
|
||||
# on fv-ml1 GPU 1 (the utility card, beside vllm-coder and scriberr). Prime, 2026-09-27.
|
||||
#
|
||||
# One forward pass of a pinned Qwen3.5-4B (BF16) per decision; the answer is read from the
|
||||
# option-letter logits, so there is no decoding. Service code + contract:
|
||||
# services/semif-serve/ (semif-serve.contract.md). Image built on fv-ml1 from that dir.
|
||||
#
|
||||
# ⚠ Scores are "conditional option score; uncalibrated as decision confidence". A caller
|
||||
# that needs thresholds brings labelled rows; we fit a per-workload temperature into
|
||||
# conf/calibration.json and the caller passes `workload`. See the README.
|
||||
# ⚠ SEMIF_VRAM_CAP_GIB is a HARD cap (torch per-process memory fraction), sized from a
|
||||
# measured peak, so SemIf cannot squeeze scriberr or the vLLM seats on this card. A
|
||||
# request that needs more gets 503 out_of_memory and the service stays up.
|
||||
#
|
||||
# .env (tunables): IMAGE, PORT, GPU_ID, VRAM_CAP_GIB, HOST_IP, SEMIF_API_TOKEN (vault
|
||||
# semif/api-token, >= 32 chars).
|
||||
|
||||
name: semif
|
||||
|
||||
services:
|
||||
semif:
|
||||
image: ${IMAGE:?set IMAGE}
|
||||
container_name: semif
|
||||
restart: unless-stopped
|
||||
ports:
|
||||
- "${PORT:-8032}:8000"
|
||||
environment:
|
||||
SEMIF_API_TOKEN: ${SEMIF_API_TOKEN:?set SEMIF_API_TOKEN}
|
||||
SEMIF_DEVICE: cuda
|
||||
SEMIF_VRAM_CAP_GIB: ${VRAM_CAP_GIB:?set VRAM_CAP_GIB}
|
||||
SEMIF_MAX_TOKENS: ${MAX_TOKENS:-4096}
|
||||
SEMIF_MAX_DECISIONS: ${MAX_DECISIONS:-64}
|
||||
SEMIF_CALIBRATION: /conf/calibration.json
|
||||
volumes:
|
||||
# Pinned weights, read offline (HF_HUB_OFFLINE=1 in the image). Never downloads.
|
||||
- /tank/aimodels/huggingface:/hf:ro
|
||||
- /opt/docker/conf/semif:/conf:ro
|
||||
deploy:
|
||||
resources:
|
||||
reservations:
|
||||
devices:
|
||||
- driver: nvidia
|
||||
device_ids: ["${GPU_ID:-1}"]
|
||||
capabilities: [gpu]
|
||||
healthcheck:
|
||||
test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=5)"]
|
||||
interval: 30s
|
||||
timeout: 10s
|
||||
retries: 3
|
||||
# Startup loads ~9 GB of weights and scores one warm-up decision before it serves.
|
||||
start_period: 300s
|
||||
networks:
|
||||
- tnet
|
||||
labels:
|
||||
- homepage.group=AI - Eval & Retrieval
|
||||
- homepage.name=SemIf — option-logit decisions
|
||||
- homepage.icon=mdi-scale-balance
|
||||
- homepage.description=Typed decisions from one forward pass (Qwen3.5-4B, fv-ml1 GPU1)
|
||||
- homepage.href=http://${HOST_IP:-10.251.50.54}:${PORT:-8032}/health
|
||||
|
||||
networks:
|
||||
tnet:
|
||||
name: traefik-net
|
||||
external: true
|
||||
Reference in New Issue
Block a user