services/semif-serve is a FastAPI wrapper around SemIf's direct and shared torch scorers (SemIf-OpenJev @ 23cf1f39, MIT). Upstream ships only a batch CLI. The wrapper loads the pinned Qwen3.5-4B (851bf6e8, BF16) once from the offline HF cache and returns SemIf's result dicts unchanged, with an optional per-workload temperature-calibrated view. Contract: semif-serve.contract.md. Built with a short contract, TDD (39 tests, fake engine and fake torch, no GPU) and a heid bug-hunt panel (pending). On the card: - torch 2.10.0+cu128 with sm_120 kernels, which is SemIf's own stack; - a hard 12 GiB VRAM cap. Two defects surfaced only on the card, and each fix is covered by a test: - 0.1.1: an OOM raised as a chained exception kept the failed request's tensors alive (11.9 GiB after the 503). It is now raised unchained, after gc. - 0.1.2: a large request left 12.6 GB reserved on the shared card. After each call, reserved memory over the baseline + 512 MiB is now released. Acceptance against SemIf's committed torch predictions (authored144): - 142/144 same top choice; both misses are exact bf16 ties; - 144/144 identical prompt hashes; - deterministic A-vs-A; - negative control 14/144; - shared vs direct 72/72. 21 binary criteria over one state take 159 ms. The shared-mode capacity table under the cap is in stacks/semif/README.md. The Dockerfile installs dependencies from a manifest with the project version blanked, so a version bump reuses the ~4 GB torch layer. Verified: 41 s rebuild, dependency layer CACHED. DNS: semif.fv.internal. Token: vault semif/api-token.
63 lines
1.4 KiB
JSON
63 lines
1.4 KiB
JSON
{
|
|
"url": "http://10.251.50.54:8032",
|
|
"health": {
|
|
"status": "ok",
|
|
"semif_commit": "23cf1f39fc9534fe81437200959b6dfc7106e45a",
|
|
"model": {
|
|
"source": "Qwen/Qwen3.5-4B",
|
|
"revision": "851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a",
|
|
"dtype": "bfloat16",
|
|
"device": "cuda:0",
|
|
"torch_version": "2.10.0+cu128",
|
|
"transformers_version": "5.17.0",
|
|
"device_name": "NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition",
|
|
"allocated_gib": 7.85,
|
|
"reserved_gib": 7.86
|
|
},
|
|
"vram_cap_gib": 12.0,
|
|
"max_tokens": 4096,
|
|
"max_decisions": 64,
|
|
"workloads": []
|
|
},
|
|
"1_parity_vs_upstream": {
|
|
"rows": 144,
|
|
"top_choice_agree": 142,
|
|
"max_abs_prob_gap": 0.09262644585476243
|
|
},
|
|
"1_prompt_sha256_equal": 144,
|
|
"2_noise_floor_a_vs_b": {
|
|
"rows": 144,
|
|
"top_choice_agree": 144,
|
|
"max_abs_prob_gap": 0.0
|
|
},
|
|
"3_negative_rotated_options": {
|
|
"rows": 144,
|
|
"top_choice_agree": 14,
|
|
"max_abs_prob_gap": 0.9987451392432198
|
|
},
|
|
"4_shared_vs_direct": {
|
|
"groups": 36,
|
|
"rows": 72,
|
|
"top_choice_agree": 72,
|
|
"max_abs_prob_gap": 0.044578713319407104
|
|
},
|
|
"5_speed_21_binary": {
|
|
"prefix_tokens": 62,
|
|
"shared_s": {
|
|
"runs": [
|
|
0.15981742000440136,
|
|
0.159110098000383,
|
|
0.15852549100236502
|
|
],
|
|
"median": 0.159110098000383
|
|
},
|
|
"sequential_decide_s": {
|
|
"runs": [
|
|
0.978545692996704,
|
|
0.9808067879930604,
|
|
0.9893363219889579
|
|
],
|
|
"median": 0.9808067879930604
|
|
}
|
|
}
|
|
} |