Files
esh-pfi-infrastructure/stacks/semif/compose.yaml
T
vh 069725c4b3 feat(semif): SemIf option-logit decisions on fv-ml1 GPU 1 (Prime)
services/semif-serve is a FastAPI wrapper around SemIf's direct and shared torch
scorers (SemIf-OpenJev @ 23cf1f39, MIT). Upstream ships only a batch CLI. The
wrapper loads the pinned Qwen3.5-4B (851bf6e8, BF16) once from the offline HF
cache and returns SemIf's result dicts unchanged, with an optional per-workload
temperature-calibrated view. Contract: semif-serve.contract.md. Built with a
short contract, TDD (39 tests, fake engine and fake torch, no GPU) and a heid
bug-hunt panel (pending).

On the card:
- torch 2.10.0+cu128 with sm_120 kernels, which is SemIf's own stack;
- a hard 12 GiB VRAM cap.
Two defects surfaced only on the card, and each fix is covered by a test:
- 0.1.1: an OOM raised as a chained exception kept the failed request's tensors
  alive (11.9 GiB after the 503). It is now raised unchained, after gc.
- 0.1.2: a large request left 12.6 GB reserved on the shared card. After each
  call, reserved memory over the baseline + 512 MiB is now released.

Acceptance against SemIf's committed torch predictions (authored144):
- 142/144 same top choice; both misses are exact bf16 ties;
- 144/144 identical prompt hashes;
- deterministic A-vs-A;
- negative control 14/144;
- shared vs direct 72/72.
21 binary criteria over one state take 159 ms. The shared-mode capacity table
under the cap is in stacks/semif/README.md.

The Dockerfile installs dependencies from a manifest with the project version
blanked, so a version bump reuses the ~4 GB torch layer. Verified: 41 s rebuild,
dependency layer CACHED.

DNS: semif.fv.internal. Token: vault semif/api-token.
2026-09-27 02:36:56 -07:00

65 lines
2.6 KiB
YAML

# semif: SemIf option-logit decisions (github TheoLeeCJ/SemIf-OpenJev, MIT) behind semif-serve,
# on fv-ml1 GPU 1 (the utility card, beside vllm-coder and scriberr). Prime, 2026-09-27.
#
# One forward pass of a pinned Qwen3.5-4B (BF16) per decision; the answer is read from the
# option-letter logits, so there is no decoding. Service code + contract:
# services/semif-serve/ (semif-serve.contract.md). Image built on fv-ml1 from that dir.
#
# ⚠ Scores are "conditional option score; uncalibrated as decision confidence". A caller
# that needs thresholds brings labelled rows; we fit a per-workload temperature into
# conf/calibration.json and the caller passes `workload`. See the README.
# ⚠ SEMIF_VRAM_CAP_GIB is a HARD cap (torch per-process memory fraction), sized from a
# measured peak, so SemIf cannot squeeze scriberr or the vLLM seats on this card. A
# request that needs more gets 503 out_of_memory and the service stays up.
#
# .env (tunables): IMAGE, PORT, GPU_ID, VRAM_CAP_GIB, HOST_IP, SEMIF_API_TOKEN (vault
# semif/api-token, >= 32 chars).
name: semif
services:
semif:
image: ${IMAGE:?set IMAGE}
container_name: semif
restart: unless-stopped
ports:
- "${PORT:-8032}:8000"
environment:
SEMIF_API_TOKEN: ${SEMIF_API_TOKEN:?set SEMIF_API_TOKEN}
SEMIF_DEVICE: cuda
SEMIF_VRAM_CAP_GIB: ${VRAM_CAP_GIB:?set VRAM_CAP_GIB}
SEMIF_MAX_TOKENS: ${MAX_TOKENS:-4096}
SEMIF_MAX_DECISIONS: ${MAX_DECISIONS:-64}
SEMIF_CALIBRATION: /conf/calibration.json
volumes:
# Pinned weights, read offline (HF_HUB_OFFLINE=1 in the image). Never downloads.
- /tank/aimodels/huggingface:/hf:ro
- /opt/docker/conf/semif:/conf:ro
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids: ["${GPU_ID:-1}"]
capabilities: [gpu]
healthcheck:
test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=5)"]
interval: 30s
timeout: 10s
retries: 3
# Startup loads ~9 GB of weights and scores one warm-up decision before it serves.
start_period: 300s
networks:
- tnet
labels:
- homepage.group=AI - Eval & Retrieval
- homepage.name=SemIf — option-logit decisions
- homepage.icon=mdi-scale-balance
- homepage.description=Typed decisions from one forward pass (Qwen3.5-4B, fv-ml1 GPU1)
- homepage.href=http://${HOST_IP:-10.251.50.54}:${PORT:-8032}/health
networks:
tnet:
name: traefik-net
external: true