36 lines
2.2 KiB
Markdown
36 lines
2.2 KiB
Markdown
# `[2026-09-27]` SemIf live on fv-ml1 GPU 1 (Prime)
|
|
|
|
**semif-serve 0.1.2 is live** at `http://10.251.50.54:8032` (`semif.fv.internal`), stack `stacks/semif`,
|
|
code and contract `services/semif-serve/` (`069725c`). It is a FastAPI wrapper around SemIf-OpenJev's
|
|
direct and shared torch scorers (MIT, pinned `23cf1f39fc9534fe81437200959b6dfc7106e45a`) on
|
|
Qwen3.5-4B `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a` (BF16, offline HF cache on `/tank`).
|
|
Hosting assessment: infra-hermes, `/tmp/semif.md` (2026-09-25).
|
|
|
|
Prime's calls (2026-09-27): fv-ml1 GPU 1 under a hard VRAM cap (not the reserved-empty GPU 3, not a Q4
|
|
GGUF on a 16 GB box); infra-ops builds it with a light process (short contract → TDD → heid bug hunt);
|
|
no consumer named yet, so every score is labelled uncalibrated.
|
|
|
|
**Two defects that only the card showed**, each fixed with a test:
|
|
- **0.1.1:** the OOM path raised `OutOfMemory(...) from exc` inside the `except` block. The chain kept the
|
|
torch exception's traceback, and with it the failed call's frames and tensors, so 11.9 GiB stayed
|
|
allocated after the 503. The fix raises after the block, with no chain, and runs `gc.collect()`
|
|
before `empty_cache()`.
|
|
- **0.1.2:** a large shared request left torch's cache holding 12.6 GB, which left scriberr 3.5 GB on GPU 1.
|
|
Now, after each call, reserved memory above the post-warm-up baseline + 512 MiB is released
|
|
(`empty_cache`). It costs ~16 ms on a big request.
|
|
|
|
**Acceptance** (`services/semif-serve/acceptance/result-2026-09-27-v0.1.2.json`), against SemIf's
|
|
committed torch predictions on authored144:
|
|
- 142/144 same top choice; both misses are exact bf16 ties in our output;
|
|
- 144/144 identical prompt SHA-256;
|
|
- A-vs-A gap 0.0;
|
|
- negative control (rotated option descriptions) 14/144;
|
|
- shared vs direct 72/72;
|
|
- 21 binary criteria over one state in 159 ms.
|
|
|
|
The shared-mode capacity under 12 GiB is 52 decisions at a ~140-token prefix and 13 at ~3,900.
|
|
|
|
**Build:** torch 2.10.0+cu128 from the pytorch index (SemIf's own stack; sm_120 present). The
|
|
Dockerfile installs dependencies from a manifest with the project version blanked, so a version bump
|
|
reuses the ~4 GB torch layer (41 s rebuild, layer CACHED). See [[2026-09-27-semif-order-averaging]].
|