Order averaging (Prime, after the 739aa03 spike):
- A decision may set orderings: rotations|all (all only for <= 4 options). Every
ordering goes to the engine in one shared batch.
- The reply keeps each native result and adds combined {probabilities (log-mean),
top, agreement, spread}.
- Through the service on SemIf's labelled sets (252 rows): 78.6% -> 88.1%
(group-bootstrap 95% CI +5.1..+14.3). Unanimous agreement is 94.5% accurate.
Fast kernels: flash-linear-attention 0.5.2 and causal-conv1d 1.7.0 are now the
default build. A/B on the empty GPU 3:
- parity with upstream went from 142/144 to 144/144;
- a ~2k-token /decide went from 169 to 92 ms server-side;
- short 3-rotation batches cost ~3-6 ms more.
triton builds a C shim at runtime, so the image carries gcc. Without it the
warm-up failed and startup failed closed.
Heid bug-hunt panel (4/4 arms, thread 01M3H3F4RR7XBP90KQ3A39H4SX), folded:
- Startup validation: VRAM cap 0 no longer means uncapped (C1); limits must be
>= 1 (S1); the token must be visible ASCII (S2); the calibration file must
exist and parse, with T in [0.05, 20] (S8, and C3's NaN leg).
- The body limit is checked before a chunk is kept, and a Unicode-digit
Content-Length no longer crashes (C2, S3).
- Failures while building the response now get the 500 envelope (C3).
- 429 busy past SEMIF_MAX_QUEUE requests in progress (C6).
- The engine releases memory on every non-validation failure, unchained after
gc; an empty OOM message is handled; 'out of memory' RuntimeErrors map to 503
(C4, C5, S9).
- The entry point forces HF_HUB_OFFLINE (S10). README wording fixed (S5, S6).
- New guard tests close the gaps the arms' mutation grids exposed: early stop of
the body read, a shared-route lock, calibration pass-through, the gc cycle,
the exact caps, TorchEngine.load's arch and device checks, and the offline
entry point.
86 tests.
Deployed on fv-ml1 GPU 1: parity 144/144, OOM and burst release verified, shared
capacity 63/51/26/16 rows at ~140/520/1960/3900 prefix tokens.
47 lines
1.5 KiB
TOML
47 lines
1.5 KiB
TOML
[project]
|
|
name = "semif-serve"
|
|
version = "0.1.3"
|
|
description = "HTTP wrapper around SemIf's direct and shared option-logit scorers"
|
|
requires-python = ">=3.12"
|
|
dependencies = [
|
|
"fastapi==0.118.0",
|
|
"uvicorn==0.37.0",
|
|
]
|
|
|
|
[project.optional-dependencies]
|
|
# The real engine. Pulls torch 2.10.0 (cu128) and transformers 5.17.0 through SemIf's exact pins.
|
|
model = [
|
|
"semif-phase1 @ git+https://github.com/TheoLeeCJ/SemIf-OpenJev@23cf1f39fc9534fe81437200959b6dfc7106e45a",
|
|
# Same pin SemIf declares, taken from the cu128 index: SemIf's committed predictions report
|
|
# torch 2.10.0+cu128, and cu128 carries sm_120 kernels for the Blackwell cards.
|
|
"torch==2.10.0",
|
|
]
|
|
|
|
# Qwen3.5's fast kernels. Without them transformers runs its reference PyTorch paths
|
|
# ("correct but much slower"). Trialled 2026-09-27; adopted only if parity with upstream holds.
|
|
fast = [
|
|
"flash-linear-attention==0.5.2",
|
|
"causal-conv1d @ https://github.com/Dao-AILab/causal-conv1d/releases/download/v1.7.0/causal_conv1d-1.7.0+cu12torch2.10cxx11abiTRUE-cp312-cp312-linux_x86_64.whl ; sys_platform == 'linux' and platform_machine == 'x86_64'",
|
|
]
|
|
|
|
[dependency-groups]
|
|
dev = ["pytest==8.4.2", "httpx==0.28.1"]
|
|
|
|
[build-system]
|
|
requires = ["setuptools>=68"]
|
|
build-backend = "setuptools.build_meta"
|
|
|
|
[tool.setuptools.packages.find]
|
|
where = ["src"]
|
|
|
|
[tool.pytest.ini_options]
|
|
testpaths = ["tests"]
|
|
|
|
[[tool.uv.index]]
|
|
name = "pytorch-cu128"
|
|
url = "https://download.pytorch.org/whl/cu128"
|
|
explicit = true
|
|
|
|
[tool.uv.sources]
|
|
torch = { index = "pytorch-cu128" }
|