feat(semif): 0.1.3 — order averaging, fast kernels, bug-hunt hardening (Prime)
Order averaging (Prime, after the 739aa03 spike):
- A decision may set orderings: rotations|all (all only for <= 4 options). Every
ordering goes to the engine in one shared batch.
- The reply keeps each native result and adds combined {probabilities (log-mean),
top, agreement, spread}.
- Through the service on SemIf's labelled sets (252 rows): 78.6% -> 88.1%
(group-bootstrap 95% CI +5.1..+14.3). Unanimous agreement is 94.5% accurate.
Fast kernels: flash-linear-attention 0.5.2 and causal-conv1d 1.7.0 are now the
default build. A/B on the empty GPU 3:
- parity with upstream went from 142/144 to 144/144;
- a ~2k-token /decide went from 169 to 92 ms server-side;
- short 3-rotation batches cost ~3-6 ms more.
triton builds a C shim at runtime, so the image carries gcc. Without it the
warm-up failed and startup failed closed.
Heid bug-hunt panel (4/4 arms, thread 01M3H3F4RR7XBP90KQ3A39H4SX), folded:
- Startup validation: VRAM cap 0 no longer means uncapped (C1); limits must be
>= 1 (S1); the token must be visible ASCII (S2); the calibration file must
exist and parse, with T in [0.05, 20] (S8, and C3's NaN leg).
- The body limit is checked before a chunk is kept, and a Unicode-digit
Content-Length no longer crashes (C2, S3).
- Failures while building the response now get the 500 envelope (C3).
- 429 busy past SEMIF_MAX_QUEUE requests in progress (C6).
- The engine releases memory on every non-validation failure, unchained after
gc; an empty OOM message is handled; 'out of memory' RuntimeErrors map to 503
(C4, C5, S9).
- The entry point forces HF_HUB_OFFLINE (S10). README wording fixed (S5, S6).
- New guard tests close the gaps the arms' mutation grids exposed: early stop of
the body read, a shared-route lock, calibration pass-through, the gc cycle,
the exact caps, TorchEngine.load's arch and device checks, and the offline
entry point.
86 tests.
Deployed on fv-ml1 GPU 1: parity 144/144, OOM and burst release verified, shared
capacity 63/51/26/16 rows at ~140/520/1960/3900 prefix tokens.
This commit is contained in:
+56
-7
@@ -16,7 +16,7 @@ service. The contract is `services/semif-serve/semif-serve.contract.md`.
|
||||
| **Model** | `Qwen/Qwen3.5-4B` @ `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a`, BF16, from `/tank/aimodels/huggingface` (read-only, offline) |
|
||||
| **SemIf** | commit `23cf1f39fc9534fe81437200959b6dfc7106e45a`; torch `2.10.0+cu128`, transformers `5.17.0`, the same stack SemIf's committed predictions were made on |
|
||||
| **Image** | `semif-serve:<version>`, built on fv-ml1 from `services/semif-serve/` |
|
||||
| **State** | none. Weights re-download at the pinned revision. No backup needed beyond the host's `/opt/docker` restic. |
|
||||
| **State** | none. If the weights are ever lost, re-pull them by hand at the pinned revision (the service itself never downloads: `HF_HUB_OFFLINE=1`, read-only mount). No backup needed beyond the host's `/opt/docker` restic. |
|
||||
|
||||
## API
|
||||
|
||||
@@ -34,6 +34,18 @@ curl -s -H "Authorization: Bearer $T" http://10.251.50.54:8032/decide -d '{
|
||||
the state once and scores every criterion in one batch. Use it when many questions
|
||||
share one long state.
|
||||
- `GET /health`: the pins, limits and calibrated workloads.
|
||||
- **Order averaging (0.1.3):** add `"orderings": "rotations"` to a decision, or
|
||||
`"all"` for ≤ 4 options. The option list is asked in every rotation inside ONE
|
||||
shared batch. The reply keeps each ordering's native result under `orderings` and adds
|
||||
`combined: {method, orderings, probabilities, top, agreement, spread}`. **Use it for
|
||||
anything real:** a small model leans toward the first-listed option on ambiguous
|
||||
inputs, and averaging cancels that. Through the service on SemIf's labelled sets,
|
||||
accuracy goes from 78.6% to **88.1%** (group-bootstrap 95% CI +5.1..+14.3 pts, 252
|
||||
rows). **`agreement` is the cheap confidence signal**: unanimous rows are 94.5%
|
||||
accurate, split rows 76.4%. The orderings count toward the decision cap. `workload`
|
||||
calibration is not available together with `orderings` yet (422).
|
||||
- Past `MAX_QUEUE` (32) requests in progress, new POSTs get `429 busy` before their
|
||||
body is read.
|
||||
|
||||
## ⚠ Probabilities are uncalibrated
|
||||
|
||||
@@ -57,18 +69,55 @@ not squeeze scriberr, which shares GPU 1. A request that would exceed the cap ge
|
||||
behaviours were verified on the card (0.1.0 held 11.9 GiB after an OOM, and 12.6 GB
|
||||
after a large request; 0.1.2 returns to 7.85 GiB in both cases).
|
||||
|
||||
**What fits under 12 GiB** (measured, `/decide/shared`, binary decisions):
|
||||
**What fits under 12 GiB** (measured on 0.1.3, `/decide/shared`, binary decisions;
|
||||
0.1.2 figures in brackets, before the fast kernels):
|
||||
|
||||
| state size (prefix tokens) | max decisions in one request |
|
||||
| state size (prefix tokens) | max rows in one request |
|
||||
|---|---|
|
||||
| ~140 | 52 |
|
||||
| ~520 | 43 |
|
||||
| ~1,960 | 19 |
|
||||
| ~3,900 | 13 |
|
||||
| ~140 | 63 (52) |
|
||||
| ~520 | 51 (43) |
|
||||
| ~1,960 | 26 (19) |
|
||||
| ~3,900 | 16 (13) |
|
||||
|
||||
Rows = decisions × orderings, so `rotations` over 3 options uses 3 rows per decision.
|
||||
|
||||
`/decide` fits at the full 4,096-token limit. Past the table you get a 503, so split
|
||||
the decisions across requests.
|
||||
|
||||
## Fast kernels (0.1.3)
|
||||
|
||||
The image ships Qwen3.5's fast kernels, `flash-linear-attention` 0.5.2 and
|
||||
`causal-conv1d` 1.7.0 (a prebuilt cu12/torch2.10 wheel). Without them transformers
|
||||
logs that it falls back to "much slower" reference PyTorch paths. They are adopted
|
||||
because an A/B on the empty GPU 3 showed:
|
||||
- **parity improved**: 144/144 vs upstream (142/144 without), so upstream evidently
|
||||
ran with them;
|
||||
- **long inputs got much faster**: a ~2,000-token `/decide` went from 169 to 92 ms
|
||||
server-side.
|
||||
|
||||
Short 3-rotation batches cost ~3–6 ms more; everything else is equal or faster.
|
||||
⚠ triton builds a C shim at runtime, so the image carries `gcc`. Without it the
|
||||
warm-up fails, and startup fails closed. Build without the kernels:
|
||||
`--build-arg EXTRAS="--extra model"`.
|
||||
|
||||
## Latency (0.1.3, from nh3-dev, 3 runs × 20, network floor ~33 ms)
|
||||
|
||||
| request | end to end | server |
|
||||
|---|---|---|
|
||||
| `/decide`, short (~130 tok) | 69 ms | 35 ms |
|
||||
| `/decide`, ~2,000-token state | 131 ms | 92 ms |
|
||||
| 3 rotations, short | 115 ms | 78 ms |
|
||||
| 6 orderings, short | 118 ms | 81 ms |
|
||||
| 3 rotations, ~2,000-token state | 200 ms | 158 ms |
|
||||
|
||||
## Acceptance (2026-09-27, v0.1.3)
|
||||
|
||||
Raw: `services/semif-serve/acceptance/result-2026-09-27-v0.1.3.json`,
|
||||
`averaging-2026-09-27-v0.1.3.json`. v0.1.3 matches upstream on **144/144** (identical
|
||||
prompt hashes, max prob gap 0.049), is deterministic, fails the negative control as it
|
||||
should (14/144), and shared matches direct on 72/72. The v0.1.2 table below is the
|
||||
reference-kernel baseline.
|
||||
|
||||
## Acceptance (2026-09-27, v0.1.2)
|
||||
|
||||
Raw: `services/semif-serve/acceptance/result-2026-09-27-v0.1.2.json`.
|
||||
|
||||
Reference in New Issue
Block a user