Files
esh-pfi-infrastructure/stacks/semif/compose.yaml
T
vh 77b8cb449c feat(semif): 0.1.3 — order averaging, fast kernels, bug-hunt hardening (Prime)
Order averaging (Prime, after the 739aa03 spike):
- A decision may set orderings: rotations|all (all only for <= 4 options). Every
  ordering goes to the engine in one shared batch.
- The reply keeps each native result and adds combined {probabilities (log-mean),
  top, agreement, spread}.
- Through the service on SemIf's labelled sets (252 rows): 78.6% -> 88.1%
  (group-bootstrap 95% CI +5.1..+14.3). Unanimous agreement is 94.5% accurate.

Fast kernels: flash-linear-attention 0.5.2 and causal-conv1d 1.7.0 are now the
default build. A/B on the empty GPU 3:
- parity with upstream went from 142/144 to 144/144;
- a ~2k-token /decide went from 169 to 92 ms server-side;
- short 3-rotation batches cost ~3-6 ms more.
triton builds a C shim at runtime, so the image carries gcc. Without it the
warm-up failed and startup failed closed.

Heid bug-hunt panel (4/4 arms, thread 01M3H3F4RR7XBP90KQ3A39H4SX), folded:
- Startup validation: VRAM cap 0 no longer means uncapped (C1); limits must be
  >= 1 (S1); the token must be visible ASCII (S2); the calibration file must
  exist and parse, with T in [0.05, 20] (S8, and C3's NaN leg).
- The body limit is checked before a chunk is kept, and a Unicode-digit
  Content-Length no longer crashes (C2, S3).
- Failures while building the response now get the 500 envelope (C3).
- 429 busy past SEMIF_MAX_QUEUE requests in progress (C6).
- The engine releases memory on every non-validation failure, unchained after
  gc; an empty OOM message is handled; 'out of memory' RuntimeErrors map to 503
  (C4, C5, S9).
- The entry point forces HF_HUB_OFFLINE (S10). README wording fixed (S5, S6).
- New guard tests close the gaps the arms' mutation grids exposed: early stop of
  the body read, a shared-route lock, calibration pass-through, the gc cycle,
  the exact caps, TorchEngine.load's arch and device checks, and the offline
  entry point.
86 tests.

Deployed on fv-ml1 GPU 1: parity 144/144, OOM and burst release verified, shared
capacity 63/51/26/16 rows at ~140/520/1960/3900 prefix tokens.
2026-09-27 03:27:15 -07:00

67 lines
2.7 KiB
YAML

# semif: SemIf option-logit decisions (github TheoLeeCJ/SemIf-OpenJev, MIT) behind semif-serve,
# on fv-ml1 GPU 1 (the utility card, beside vllm-coder and scriberr). Prime, 2026-09-27.
#
# One forward pass of a pinned Qwen3.5-4B (BF16) per decision; the answer is read from the
# option-letter logits, so there is no decoding. Service code + contract:
# services/semif-serve/ (semif-serve.contract.md). Image built on fv-ml1 from that dir.
#
# ⚠ Scores are "conditional option score; uncalibrated as decision confidence". A caller
# that needs thresholds brings labelled rows; we fit a per-workload temperature into
# conf/calibration.json and the caller passes `workload`. See the README.
# ⚠ SEMIF_VRAM_CAP_GIB is a HARD cap (torch per-process memory fraction), sized from a
# measured peak, so SemIf cannot squeeze scriberr or the vLLM seats on this card. A
# request that needs more gets 503 out_of_memory and the service stays up.
#
# .env (tunables): IMAGE, PORT, GPU_ID, VRAM_CAP_GIB, HOST_IP, SEMIF_API_TOKEN (vault
# semif/api-token, >= 32 chars).
name: semif
services:
semif:
image: ${IMAGE:?set IMAGE}
container_name: semif
restart: unless-stopped
ports:
- "${PORT:-8032}:8000"
environment:
SEMIF_API_TOKEN: ${SEMIF_API_TOKEN:?set SEMIF_API_TOKEN}
SEMIF_DEVICE: cuda
SEMIF_VRAM_CAP_GIB: ${VRAM_CAP_GIB:?set VRAM_CAP_GIB}
SEMIF_MAX_TOKENS: ${MAX_TOKENS:-4096}
SEMIF_MAX_DECISIONS: ${MAX_DECISIONS:-64}
# POSTs in progress (queued + scoring) before new ones get 429 busy.
SEMIF_MAX_QUEUE: ${MAX_QUEUE:-32}
SEMIF_CALIBRATION: /conf/calibration.json
volumes:
# Pinned weights, read offline (HF_HUB_OFFLINE=1 in the image). Never downloads.
- /tank/aimodels/huggingface:/hf:ro
- /opt/docker/conf/semif:/conf:ro
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids: ["${GPU_ID:-1}"]
capabilities: [gpu]
healthcheck:
test: ["CMD", "python", "-c", "import urllib.request; urllib.request.urlopen('http://127.0.0.1:8000/health', timeout=5)"]
interval: 30s
timeout: 10s
retries: 3
# Startup loads ~9 GB of weights and scores one warm-up decision before it serves.
start_period: 300s
networks:
- tnet
labels:
- homepage.group=AI - Eval & Retrieval
- homepage.name=SemIf — option-logit decisions
- homepage.icon=mdi-scale-balance
- homepage.description=Typed decisions from one forward pass (Qwen3.5-4B, fv-ml1 GPU1)
- homepage.href=http://${HOST_IP:-10.251.50.54}:${PORT:-8032}/health
networks:
tnet:
name: traefik-net
external: true