Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-27-semif-live-on-fv-ml1-gpu1.md
T

2.2 KiB

[2026-09-27] SemIf live on fv-ml1 GPU 1 (Prime)

semif-serve 0.1.2 is live at http://10.251.50.54:8032 (semif.fv.internal), stack stacks/semif, code and contract services/semif-serve/ (069725c). It is a FastAPI wrapper around SemIf-OpenJev's direct and shared torch scorers (MIT, pinned 23cf1f39fc9534fe81437200959b6dfc7106e45a) on Qwen3.5-4B 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a (BF16, offline HF cache on /tank). Hosting assessment: infra-hermes, /tmp/semif.md (2026-09-25).

Prime's calls (2026-09-27): fv-ml1 GPU 1 under a hard VRAM cap (not the reserved-empty GPU 3, not a Q4 GGUF on a 16 GB box); infra-ops builds it with a light process (short contract → TDD → heid bug hunt); no consumer named yet, so every score is labelled uncalibrated.

Two defects that only the card showed, each fixed with a test:

  • 0.1.1: the OOM path raised OutOfMemory(...) from exc inside the except block. The chain kept the torch exception's traceback, and with it the failed call's frames and tensors, so 11.9 GiB stayed allocated after the 503. The fix raises after the block, with no chain, and runs gc.collect() before empty_cache().
  • 0.1.2: a large shared request left torch's cache holding 12.6 GB, which left scriberr 3.5 GB on GPU 1. Now, after each call, reserved memory above the post-warm-up baseline + 512 MiB is released (empty_cache). It costs ~16 ms on a big request.

Acceptance (services/semif-serve/acceptance/result-2026-09-27-v0.1.2.json), against SemIf's committed torch predictions on authored144:

  • 142/144 same top choice; both misses are exact bf16 ties in our output;
  • 144/144 identical prompt SHA-256;
  • A-vs-A gap 0.0;
  • negative control (rotated option descriptions) 14/144;
  • shared vs direct 72/72;
  • 21 binary criteria over one state in 159 ms.

The shared-mode capacity under 12 GiB is 52 decisions at a ~140-token prefix and 13 at ~3,900.

Build: torch 2.10.0+cu128 from the pytorch index (SemIf's own stack; sm_120 present). The Dockerfile installs dependencies from a manifest with the project version blanked, so a version bump reuses the ~4 GB torch layer (41 s rebuild, layer CACHED). See 2026-09-27-semif-order-averaging.