services/semif-serve is a FastAPI wrapper around SemIf's direct and shared torch scorers (SemIf-OpenJev @ 23cf1f39, MIT). Upstream ships only a batch CLI. The wrapper loads the pinned Qwen3.5-4B (851bf6e8, BF16) once from the offline HF cache and returns SemIf's result dicts unchanged, with an optional per-workload temperature-calibrated view. Contract: semif-serve.contract.md. Built with a short contract, TDD (39 tests, fake engine and fake torch, no GPU) and a heid bug-hunt panel (pending). On the card: - torch 2.10.0+cu128 with sm_120 kernels, which is SemIf's own stack; - a hard 12 GiB VRAM cap. Two defects surfaced only on the card, and each fix is covered by a test: - 0.1.1: an OOM raised as a chained exception kept the failed request's tensors alive (11.9 GiB after the 503). It is now raised unchained, after gc. - 0.1.2: a large request left 12.6 GB reserved on the shared card. After each call, reserved memory over the baseline + 512 MiB is now released. Acceptance against SemIf's committed torch predictions (authored144): - 142/144 same top choice; both misses are exact bf16 ties; - 144/144 identical prompt hashes; - deterministic A-vs-A; - negative control 14/144; - shared vs direct 72/72. 21 binary criteria over one state take 159 ms. The shared-mode capacity table under the cap is in stacks/semif/README.md. The Dockerfile installs dependencies from a manifest with the project version blanked, so a version bump reuses the ~4 GB torch layer. Verified: 41 s rebuild, dependency layer CACHED. DNS: semif.fv.internal. Token: vault semif/api-token.
37 lines
1.5 KiB
Python
37 lines
1.5 KiB
Python
"""Settings.from_env. Contract: semif-serve.contract.md § Configuration, INV-6."""
|
|
import json
|
|
|
|
import pytest
|
|
|
|
from semif_serve.config import DEFAULT_REVISION, Settings
|
|
|
|
TOKEN = "t" * 40
|
|
|
|
|
|
def test_defaults_from_a_minimal_env():
|
|
s = Settings.from_env({"SEMIF_API_TOKEN": TOKEN})
|
|
assert (s.api_token, s.revision, s.device, s.max_tokens, s.max_decisions) == (TOKEN, DEFAULT_REVISION, "cuda", 4096, 64)
|
|
assert s.vram_cap_gib is None and s.calibration == {}
|
|
|
|
|
|
@pytest.mark.parametrize("env", [{}, {"SEMIF_API_TOKEN": "short"}, {"SEMIF_API_TOKEN": "x" * 31}])
|
|
def test_a_missing_or_short_token_is_refused_at_startup(env):
|
|
with pytest.raises(ValueError, match="SEMIF_API_TOKEN"):
|
|
Settings.from_env(env)
|
|
|
|
|
|
def test_numbers_and_calibration_file_are_parsed(tmp_path):
|
|
cal = tmp_path / "cal.json"
|
|
cal.write_text(json.dumps({"triage": 2.5}))
|
|
s = Settings.from_env({"SEMIF_API_TOKEN": TOKEN, "SEMIF_VRAM_CAP_GIB": "12", "SEMIF_MAX_DECISIONS": "8",
|
|
"SEMIF_CALIBRATION": str(cal)})
|
|
assert (s.vram_cap_gib, s.max_decisions, s.calibration) == (12.0, 8, {"triage": 2.5})
|
|
|
|
|
|
@pytest.mark.parametrize("table", [{"w": 0}, {"w": -1.0}, {"w": "2"}, {"w": float("inf")}, ["w", 2.0]])
|
|
def test_a_calibration_table_needs_positive_finite_numbers(tmp_path, table):
|
|
cal = tmp_path / "cal.json"
|
|
cal.write_text(json.dumps(table))
|
|
with pytest.raises(ValueError, match="SEMIF_CALIBRATION"):
|
|
Settings.from_env({"SEMIF_API_TOKEN": TOKEN, "SEMIF_CALIBRATION": str(cal)})
|