feat(semif): 0.1.3 — order averaging, fast kernels, bug-hunt hardening (Prime)
Order averaging (Prime, after the 739aa03 spike):
- A decision may set orderings: rotations|all (all only for <= 4 options). Every
ordering goes to the engine in one shared batch.
- The reply keeps each native result and adds combined {probabilities (log-mean),
top, agreement, spread}.
- Through the service on SemIf's labelled sets (252 rows): 78.6% -> 88.1%
(group-bootstrap 95% CI +5.1..+14.3). Unanimous agreement is 94.5% accurate.
Fast kernels: flash-linear-attention 0.5.2 and causal-conv1d 1.7.0 are now the
default build. A/B on the empty GPU 3:
- parity with upstream went from 142/144 to 144/144;
- a ~2k-token /decide went from 169 to 92 ms server-side;
- short 3-rotation batches cost ~3-6 ms more.
triton builds a C shim at runtime, so the image carries gcc. Without it the
warm-up failed and startup failed closed.
Heid bug-hunt panel (4/4 arms, thread 01M3H3F4RR7XBP90KQ3A39H4SX), folded:
- Startup validation: VRAM cap 0 no longer means uncapped (C1); limits must be
>= 1 (S1); the token must be visible ASCII (S2); the calibration file must
exist and parse, with T in [0.05, 20] (S8, and C3's NaN leg).
- The body limit is checked before a chunk is kept, and a Unicode-digit
Content-Length no longer crashes (C2, S3).
- Failures while building the response now get the 500 envelope (C3).
- 429 busy past SEMIF_MAX_QUEUE requests in progress (C6).
- The engine releases memory on every non-validation failure, unchained after
gc; an empty OOM message is handled; 'out of memory' RuntimeErrors map to 503
(C4, C5, S9).
- The entry point forces HF_HUB_OFFLINE (S10). README wording fixed (S5, S6).
- New guard tests close the gaps the arms' mutation grids exposed: early stop of
the body read, a shared-route lock, calibration pass-through, the gc cycle,
the exact caps, TorchEngine.load's arch and device checks, and the offline
entry point.
86 tests.
Deployed on fv-ml1 GPU 1: parity 144/144, OOM and burst release verified, shared
capacity 63/51/26/16 rows at ~140/520/1960/3900 prefix tokens.
This commit is contained in:
+11
-27
@@ -144,33 +144,17 @@ _As of 2026-09-27 ~0255 PT._
|
||||
|
||||
### SemIf on fv-ml1 GPU 1 (2026-09-27, Prime)
|
||||
|
||||
- **LIVE: `semif-serve` 0.1.2** at `http://10.251.50.54:8032` (`semif.fv.internal`). It wraps the SemIf
|
||||
option-logit scorers (pinned `23cf1f39`, Qwen3.5-4B `851bf6e8`, BF16); token `semif/api-token`.
|
||||
Code + contract: `services/semif-serve/`; stack `stacks/semif`. **No consumer yet.**
|
||||
- Acceptance: 142/144 top choice vs upstream (both misses are exact bf16 ties), 144/144 prompt hashes,
|
||||
deterministic, negative control 14/144. It holds 12 GiB hard with release after a burst (~8.7 GB resting).
|
||||
- Latency from nh3-dev: `/decide` short 71 ms end to end (38 ms server); 3 rotations 113 ms; ~2k-token
|
||||
state 210 ms (`services/semif-serve/spike/latency-2026-09-27.txt`).
|
||||
- **NEXT, in this order (Prime 0250):**
|
||||
1. **Build order-averaging into semif-serve as 0.1.3.** Spike (`739aa03`): 3 rotations, log-mean,
|
||||
take accuracy 78.6% → 87.7% (group-bootstrap 95% CI +4.7..+13.8 pts); unanimous rows are 94.4%
|
||||
accurate, split rows 75.8%. Design: opt-in `orderings: rotations|all` (all only ≤ 4 options),
|
||||
every ordering in one shared batch, per-ordering SemIf results returned unchanged + a
|
||||
`combined` block (log-mean probabilities, top, agreement, per-option spread). Orderings count
|
||||
toward `max_decisions`. Contract amendment → TDD → deploy → acceptance (authored144 accuracy
|
||||
with/without; 2AM/2PM agreement as controls).
|
||||
2. **Trial the fast kernels** (`causal_conv1d`, `flash-linear-attention`). Today transformers runs
|
||||
Qwen3.5's reference PyTorch paths and logs that they are "much slower". Adopt only if the
|
||||
authored144 parity re-check holds.
|
||||
3. **heid bug-hunt panel** on 0.1.2 is pending (thread `01M3H3F4RR7XBP90KQ3A39H4SX`). Triage and fold
|
||||
its findings into 0.1.3.
|
||||
- NVFP4 is not worth it: the SemIf scorer runs on transformers, not vLLM; quant noise lands on the scores
|
||||
and would need its own calibration; BF16 already fits. SemIf's own MLX 4-bit run moved authored144
|
||||
0.813 → 0.789 (one run each, so indicative only).
|
||||
- Prime's probes (single scenarios, not benchmarks): at 2AM vs 2PM, booty-call probability averaged over
|
||||
all orderings goes 0.42 → 0.05; a baby dragon reads life-changing 0.65 / world-changing 0.25 with only
|
||||
2/3 rotations agreeing (one ordering said 0.925), the lottery ticket reads life-changing unanimously,
|
||||
and the paperclip control reads trivial 0.999.
|
||||
- **LIVE: `semif-serve` 0.1.3** at `http://10.251.50.54:8032` (`semif.fv.internal`), with order averaging
|
||||
and the fast kernels. SemIf `23cf1f39`, Qwen3.5-4B `851bf6e8`, BF16; token `semif/api-token`. Code +
|
||||
contract: `services/semif-serve/`; stack `stacks/semif`. **No consumer yet.**
|
||||
- 0.1.3 acceptance: 144/144 parity with upstream; averaging through the service 78.6% → 88.1%
|
||||
(95% CI +5.1..+14.3); unanimous rotations 94.5% accurate. Latency, envelope and the fast-kernel A/B
|
||||
are in `stacks/semif/README.md`.
|
||||
- The heid bug-hunt panel (thread `01M3H3F4RR7XBP90KQ3A39H4SX`) is triaged and folded into 0.1.3:
|
||||
C1–C6, S1–S3, S5, S6, S8–S10 fixed with tests; S4 settled; S7 (publish on all interfaces) accepted
|
||||
as LAN + auth.
|
||||
- NVFP4 is not worth it (see the SemIf detail files). Prime's probes (2AM/2PM, dragon/lottery) are
|
||||
recorded in [[2026-09-27-semif-order-averaging]].
|
||||
|
||||
### restic: credential leak fixed (2026-09-27, Prime)
|
||||
|
||||
|
||||
Reference in New Issue
Block a user