feat(semif): 0.1.3 — order averaging, fast kernels, bug-hunt hardening (Prime)
Order averaging (Prime, after the 739aa03 spike):
- A decision may set orderings: rotations|all (all only for <= 4 options). Every
ordering goes to the engine in one shared batch.
- The reply keeps each native result and adds combined {probabilities (log-mean),
top, agreement, spread}.
- Through the service on SemIf's labelled sets (252 rows): 78.6% -> 88.1%
(group-bootstrap 95% CI +5.1..+14.3). Unanimous agreement is 94.5% accurate.
Fast kernels: flash-linear-attention 0.5.2 and causal-conv1d 1.7.0 are now the
default build. A/B on the empty GPU 3:
- parity with upstream went from 142/144 to 144/144;
- a ~2k-token /decide went from 169 to 92 ms server-side;
- short 3-rotation batches cost ~3-6 ms more.
triton builds a C shim at runtime, so the image carries gcc. Without it the
warm-up failed and startup failed closed.
Heid bug-hunt panel (4/4 arms, thread 01M3H3F4RR7XBP90KQ3A39H4SX), folded:
- Startup validation: VRAM cap 0 no longer means uncapped (C1); limits must be
>= 1 (S1); the token must be visible ASCII (S2); the calibration file must
exist and parse, with T in [0.05, 20] (S8, and C3's NaN leg).
- The body limit is checked before a chunk is kept, and a Unicode-digit
Content-Length no longer crashes (C2, S3).
- Failures while building the response now get the 500 envelope (C3).
- 429 busy past SEMIF_MAX_QUEUE requests in progress (C6).
- The engine releases memory on every non-validation failure, unchained after
gc; an empty OOM message is handled; 'out of memory' RuntimeErrors map to 503
(C4, C5, S9).
- The entry point forces HF_HUB_OFFLINE (S10). README wording fixed (S5, S6).
- New guard tests close the gaps the arms' mutation grids exposed: early stop of
the body read, a shared-route lock, calibration pass-through, the gc cycle,
the exact caps, TorchEngine.load's arch and device checks, and the offline
entry point.
86 tests.
Deployed on fv-ml1 GPU 1: parity 144/144, OOM and burst release verified, shared
capacity 63/51/26/16 rows at ~140/520/1960/3900 prefix tokens.
This commit is contained in:
@@ -45,6 +45,37 @@ probabilities}` with `softmax(option_logits / T)`. The argmax never changes. The
|
||||
native `probabilities` and `probability_status` stay untouched. An unknown
|
||||
`workload` is a 422. With no `workload`, no `calibrated` key appears.
|
||||
|
||||
**Order averaging (0.1.3, Prime 2026-09-27).** A decision (the `/decide` body, or an
|
||||
entry in `decisions`) may set `orderings`, whose default is `"none"`:
|
||||
|
||||
- `"rotations"`: the n cyclic shifts of the caller's option list, starting with
|
||||
the caller's order, so every option sits in every position exactly once.
|
||||
- `"all"`: every permutation (n!), the caller's order first. It is allowed only
|
||||
when n ≤ 4, else 422.
|
||||
|
||||
Every ordering of every decision in the request becomes its own SemIf row
|
||||
(`id` = `"<id>#o<k>"`, same state and question, reordered options). **All rows go
|
||||
to the engine in ONE `shared` call**, which includes `/decide`. The result for an
|
||||
averaged decision is:
|
||||
|
||||
```
|
||||
{id, option_ids (caller's order),
|
||||
combined: {method, orderings: n, probabilities, top, agreement, spread: {option_id: [min, max]}},
|
||||
orderings: [the SemIf result for each ordering, unchanged]}
|
||||
```
|
||||
|
||||
- `probabilities`: per ordering, log-softmax of `option_logits`; averaged per
|
||||
option id; renormalised; reported in the caller's order.
|
||||
- `agreement`: the fraction of orderings whose top option equals `combined.top`.
|
||||
- `spread`: each option's min and max native probability across orderings.
|
||||
|
||||
Every expanded row counts toward `SEMIF_MAX_DECISIONS`. `workload` together with
|
||||
`orderings` is a 422, because a temperature is fitted per method and none is
|
||||
fitted on combined scores yet. Decisions without `orderings` keep the exact
|
||||
pre-0.1.3 result shape. Averaging cancels any additive position bias exactly.
|
||||
Measured by the 2026-09-27 spike: 3 rotations take SemIf's labelled sets from
|
||||
78.6% to 87.7% accuracy, and unanimous agreement is 94.4% accurate.
|
||||
|
||||
## Invariants
|
||||
|
||||
- **INV-1 pass-through.** `option_ids`, `probabilities`, `option_logits`,
|
||||
@@ -64,10 +95,18 @@ native `probabilities` and `probability_status` stay untouched. An unknown
|
||||
baseline by more than 512 MiB, the engine calls `empty_cache()`. A burst must not
|
||||
keep the card's shared headroom: on 2026-09-27 a 64-decision request left the
|
||||
process holding 12.6 GB, leaving scriberr 3.5 GB.
|
||||
**Any other scorer failure except `ValueError`** also releases before it is
|
||||
reported: it is logged with its traceback, then re-raised unchained as
|
||||
`ScoringFailed` after `gc.collect()` + `empty_cache()` (bug hunt C5). A
|
||||
`RuntimeError` whose message says "out of memory" (cuBLAS/cuDNN allocation
|
||||
failures) counts as an OOM → 503 (S9). An OOM with an empty message is reported
|
||||
as "CUDA out of memory" rather than crashing the handler (C4).
|
||||
- **INV-5 no network at runtime.** Weights come from the mounted HF cache at the
|
||||
pinned revision (`HF_HUB_OFFLINE=1`).
|
||||
pinned revision. The entry point sets `HF_HUB_OFFLINE=1` itself before torch or
|
||||
transformers load, so this holds outside the image too (S10).
|
||||
- **INV-6 constant-time auth.** Token comparison uses `hmac.compare_digest`. The
|
||||
token is ≥ 32 characters, and startup refuses a shorter one.
|
||||
token is ≥ 32 characters of visible ASCII (33–126). Startup refuses anything
|
||||
else, because a CR, LF or NUL in the token can never arrive in a header (S2).
|
||||
|
||||
## Limits and errors
|
||||
|
||||
@@ -75,15 +114,21 @@ native `probabilities` and `probability_status` stay untouched. An unknown
|
||||
422, never truncated (SemIf raises).
|
||||
- `SEMIF_MAX_DECISIONS` (default 64) caps `decisions` per shared request. It must
|
||||
hold 1..max entries, else 422.
|
||||
- Request body ≤ `SEMIF_MAX_BODY_BYTES` (default 1 MiB), else 413.
|
||||
- Request body ≤ `SEMIF_MAX_BODY_BYTES` (default 1 MiB), else 413. The limit is
|
||||
checked **before** each chunk is kept, so no more than the limit is ever held
|
||||
(C2). A declared `Content-Length` is trusted only if it is ASCII digits (S3).
|
||||
- **Admission:** at most `SEMIF_MAX_QUEUE` (default 32) POSTs may be in progress,
|
||||
counting both queued and scoring. The next one is refused with 429 `busy`
|
||||
before its body is read (C6).
|
||||
|
||||
| status | code | when |
|
||||
|---|---|---|
|
||||
| 401 | `unauthorized` | missing or wrong bearer |
|
||||
| 413 | `request_too_large` | body over the limit |
|
||||
| 422 | `invalid_request` | bad JSON shape, a SemIf `ValueError` (validation, token limit, tokenisation), unknown workload, too many decisions |
|
||||
| 429 | `busy` | `SEMIF_MAX_QUEUE` requests already in progress |
|
||||
| 503 | `out_of_memory` | CUDA OOM during scoring |
|
||||
| 500 | `scoring_failed` | any other scorer exception |
|
||||
| 500 | `scoring_failed` | any other scorer exception, **or a failure while building the response** from a scorer result (calibration, averaging): always the envelope, never a bare 500 (C3) |
|
||||
|
||||
The error body is `{error: {code, message}}`.
|
||||
|
||||
@@ -92,7 +137,16 @@ The error body is `{error: {code, message}}`.
|
||||
`SEMIF_API_TOKEN` (required), `SEMIF_MODEL` (default `Qwen/Qwen3.5-4B`),
|
||||
`SEMIF_REVISION` (default the pinned SHA), `SEMIF_DEVICE` (default `cuda`),
|
||||
`SEMIF_VRAM_CAP_GIB`, `SEMIF_MAX_TOKENS`, `SEMIF_MAX_DECISIONS`,
|
||||
`SEMIF_MAX_BODY_BYTES`, `SEMIF_CALIBRATION` (path to a JSON `{workload: T}`; T > 0).
|
||||
`SEMIF_MAX_BODY_BYTES`, `SEMIF_MAX_QUEUE`, `SEMIF_CALIBRATION` (path to a JSON
|
||||
`{workload: T}`).
|
||||
|
||||
Startup validates every value and refuses a bad one with a `ValueError` naming
|
||||
the variable (C1, S1, S8):
|
||||
- the VRAM cap, when set, is finite and > 0 (`0` used to mean uncapped);
|
||||
- every limit is an integer ≥ 1;
|
||||
- each T is a finite number, not a bool, in [0.05, 20] (a tiny T overflowed to
|
||||
NaN, and the response then failed to render);
|
||||
- the calibration file must exist and parse.
|
||||
|
||||
## Tests (TDD, fake scorer: no torch, no model)
|
||||
|
||||
@@ -109,7 +163,13 @@ alone; any
|
||||
other exception → 500; malformed
|
||||
JSON or a wrong body shape → 422; too many decisions → 422; an oversized body → 413; requests
|
||||
are serialised (two concurrent calls never overlap inside the scorer); `/health`
|
||||
answers while a scorer call is blocked.
|
||||
answers while a scorer call is blocked. **Averaging:** `rotations` sends n rows
|
||||
in one shared call, each option once per position, with ids `<id>#o<k>`, and
|
||||
cancels a position bias exactly; `all` sends n! rows and is 422 above 4 options;
|
||||
agreement and spread are computed from the orderings; a mixed shared request
|
||||
(averaged + plain) is one engine call, with results in request order and plain
|
||||
results unchanged; expanded rows count toward the cap; `workload` + `orderings`
|
||||
→ 422.
|
||||
|
||||
## Acceptance (on fv-ml1, real model; not unit tests)
|
||||
|
||||
|
||||
Reference in New Issue
Block a user