feat(semif): SemIf option-logit decisions on fv-ml1 GPU 1 (Prime)
services/semif-serve is a FastAPI wrapper around SemIf's direct and shared torch scorers (SemIf-OpenJev @ 23cf1f39, MIT). Upstream ships only a batch CLI. The wrapper loads the pinned Qwen3.5-4B (851bf6e8, BF16) once from the offline HF cache and returns SemIf's result dicts unchanged, with an optional per-workload temperature-calibrated view. Contract: semif-serve.contract.md. Built with a short contract, TDD (39 tests, fake engine and fake torch, no GPU) and a heid bug-hunt panel (pending). On the card: - torch 2.10.0+cu128 with sm_120 kernels, which is SemIf's own stack; - a hard 12 GiB VRAM cap. Two defects surfaced only on the card, and each fix is covered by a test: - 0.1.1: an OOM raised as a chained exception kept the failed request's tensors alive (11.9 GiB after the 503). It is now raised unchained, after gc. - 0.1.2: a large request left 12.6 GB reserved on the shared card. After each call, reserved memory over the baseline + 512 MiB is now released. Acceptance against SemIf's committed torch predictions (authored144): - 142/144 same top choice; both misses are exact bf16 ties; - 144/144 identical prompt hashes; - deterministic A-vs-A; - negative control 14/144; - shared vs direct 72/72. 21 binary criteria over one state take 159 ms. The shared-mode capacity table under the cap is in stacks/semif/README.md. The Dockerfile installs dependencies from a manifest with the project version blanked, so a version bump reuses the ~4 GB torch layer. Verified: 41 s rebuild, dependency layer CACHED. DNS: semif.fv.internal. Token: vault semif/api-token.
This commit is contained in:
@@ -180,7 +180,8 @@ embed/rerank/reward trio. GPUs are pinned per container via
|
||||
**GPU 1 — light / eval / retrieval + char-RP GGUF (~91/98 GB, on-demand):**
|
||||
|
||||
> ⚠ **This table is stale (checked 2026-09-27).** Live GPU 1 residents were `scriberr`,
|
||||
> `vllm-coder`, `vllm-erp-seat` and `vllm-meromero-rp`. Read the host
|
||||
> `vllm-coder`, `vllm-erp-seat` and `vllm-meromero-rp`, plus **`semif`** (:8032, ~8.7 GB resting,
|
||||
> hard-capped at 12 GiB; `stacks/semif`, since 2026-09-27). Read the host
|
||||
> (`docker inspect … DeviceRequests`), not this table. (A fixtures-only `augaman` instance ran
|
||||
> here for about an hour on 2026-09-27 for a speed bench, and was then removed on Prime's call.)
|
||||
|
||||
|
||||
Reference in New Issue
Block a user