Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-27-semif-live-on-fv-ml1-gpu1.md
T

36 lines
2.2 KiB
Markdown

# `[2026-09-27]` SemIf live on fv-ml1 GPU 1 (Prime)
**semif-serve 0.1.2 is live** at `http://10.251.50.54:8032` (`semif.fv.internal`), stack `stacks/semif`,
code and contract `services/semif-serve/` (`069725c`). It is a FastAPI wrapper around SemIf-OpenJev's
direct and shared torch scorers (MIT, pinned `23cf1f39fc9534fe81437200959b6dfc7106e45a`) on
Qwen3.5-4B `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a` (BF16, offline HF cache on `/tank`).
Hosting assessment: infra-hermes, `/tmp/semif.md` (2026-09-25).
Prime's calls (2026-09-27): fv-ml1 GPU 1 under a hard VRAM cap (not the reserved-empty GPU 3, not a Q4
GGUF on a 16 GB box); infra-ops builds it with a light process (short contract → TDD → heid bug hunt);
no consumer named yet, so every score is labelled uncalibrated.
**Two defects that only the card showed**, each fixed with a test:
- **0.1.1:** the OOM path raised `OutOfMemory(...) from exc` inside the `except` block. The chain kept the
torch exception's traceback, and with it the failed call's frames and tensors, so 11.9 GiB stayed
allocated after the 503. The fix raises after the block, with no chain, and runs `gc.collect()`
before `empty_cache()`.
- **0.1.2:** a large shared request left torch's cache holding 12.6 GB, which left scriberr 3.5 GB on GPU 1.
Now, after each call, reserved memory above the post-warm-up baseline + 512 MiB is released
(`empty_cache`). It costs ~16 ms on a big request.
**Acceptance** (`services/semif-serve/acceptance/result-2026-09-27-v0.1.2.json`), against SemIf's
committed torch predictions on authored144:
- 142/144 same top choice; both misses are exact bf16 ties in our output;
- 144/144 identical prompt SHA-256;
- A-vs-A gap 0.0;
- negative control (rotated option descriptions) 14/144;
- shared vs direct 72/72;
- 21 binary criteria over one state in 159 ms.
The shared-mode capacity under 12 GiB is 52 decisions at a ~140-token prefix and 13 at ~3,900.
**Build:** torch 2.10.0+cu128 from the pytorch index (SemIf's own stack; sm_120 present). The
Dockerfile installs dependencies from a manifest with the project version blanked, so a version bump
reuses the ~4 GB torch layer (41 s rebuild, layer CACHED). See [[2026-09-27-semif-order-averaging]].