feat(semif): 0.1.3 — order averaging, fast kernels, bug-hunt hardening (Prime)

Order averaging (Prime, after the 739aa03 spike):
- A decision may set orderings: rotations|all (all only for <= 4 options). Every
  ordering goes to the engine in one shared batch.
- The reply keeps each native result and adds combined {probabilities (log-mean),
  top, agreement, spread}.
- Through the service on SemIf's labelled sets (252 rows): 78.6% -> 88.1%
  (group-bootstrap 95% CI +5.1..+14.3). Unanimous agreement is 94.5% accurate.

Fast kernels: flash-linear-attention 0.5.2 and causal-conv1d 1.7.0 are now the
default build. A/B on the empty GPU 3:
- parity with upstream went from 142/144 to 144/144;
- a ~2k-token /decide went from 169 to 92 ms server-side;
- short 3-rotation batches cost ~3-6 ms more.
triton builds a C shim at runtime, so the image carries gcc. Without it the
warm-up failed and startup failed closed.

Heid bug-hunt panel (4/4 arms, thread 01M3H3F4RR7XBP90KQ3A39H4SX), folded:
- Startup validation: VRAM cap 0 no longer means uncapped (C1); limits must be
  >= 1 (S1); the token must be visible ASCII (S2); the calibration file must
  exist and parse, with T in [0.05, 20] (S8, and C3's NaN leg).
- The body limit is checked before a chunk is kept, and a Unicode-digit
  Content-Length no longer crashes (C2, S3).
- Failures while building the response now get the 500 envelope (C3).
- 429 busy past SEMIF_MAX_QUEUE requests in progress (C6).
- The engine releases memory on every non-validation failure, unchained after
  gc; an empty OOM message is handled; 'out of memory' RuntimeErrors map to 503
  (C4, C5, S9).
- The entry point forces HF_HUB_OFFLINE (S10). README wording fixed (S5, S6).
- New guard tests close the gaps the arms' mutation grids exposed: early stop of
  the body read, a shared-route lock, calibration pass-through, the gc cycle,
  the exact caps, TorchEngine.load's arch and device checks, and the offline
  entry point.
86 tests.

Deployed on fv-ml1 GPU 1: parity 144/144, OOM and burst release verified, shared
capacity 63/51/26/16 rows at ~140/520/1960/3900 prefix tokens.
This commit is contained in:
vh
2026-09-27 03:27:15 -07:00
parent d7ad235365
commit 77b8cb449c
22 changed files with 1018 additions and 92 deletions
+11 -27
View File
@@ -144,33 +144,17 @@ _As of 2026-09-27 ~0255 PT._
### SemIf on fv-ml1 GPU 1 (2026-09-27, Prime)
- **LIVE: `semif-serve` 0.1.2** at `http://10.251.50.54:8032` (`semif.fv.internal`). It wraps the SemIf
option-logit scorers (pinned `23cf1f39`, Qwen3.5-4B `851bf6e8`, BF16); token `semif/api-token`.
Code + contract: `services/semif-serve/`; stack `stacks/semif`. **No consumer yet.**
- Acceptance: 142/144 top choice vs upstream (both misses are exact bf16 ties), 144/144 prompt hashes,
deterministic, negative control 14/144. It holds 12 GiB hard with release after a burst (~8.7 GB resting).
- Latency from nh3-dev: `/decide` short 71 ms end to end (38 ms server); 3 rotations 113 ms; ~2k-token
state 210 ms (`services/semif-serve/spike/latency-2026-09-27.txt`).
- **NEXT, in this order (Prime 0250):**
1. **Build order-averaging into semif-serve as 0.1.3.** Spike (`739aa03`): 3 rotations, log-mean,
take accuracy 78.6% → 87.7% (group-bootstrap 95% CI +4.7..+13.8 pts); unanimous rows are 94.4%
accurate, split rows 75.8%. Design: opt-in `orderings: rotations|all` (all only ≤ 4 options),
every ordering in one shared batch, per-ordering SemIf results returned unchanged + a
`combined` block (log-mean probabilities, top, agreement, per-option spread). Orderings count
toward `max_decisions`. Contract amendment → TDD → deploy → acceptance (authored144 accuracy
with/without; 2AM/2PM agreement as controls).
2. **Trial the fast kernels** (`causal_conv1d`, `flash-linear-attention`). Today transformers runs
Qwen3.5's reference PyTorch paths and logs that they are "much slower". Adopt only if the
authored144 parity re-check holds.
3. **heid bug-hunt panel** on 0.1.2 is pending (thread `01M3H3F4RR7XBP90KQ3A39H4SX`). Triage and fold
its findings into 0.1.3.
- NVFP4 is not worth it: the SemIf scorer runs on transformers, not vLLM; quant noise lands on the scores
and would need its own calibration; BF16 already fits. SemIf's own MLX 4-bit run moved authored144
0.813 → 0.789 (one run each, so indicative only).
- Prime's probes (single scenarios, not benchmarks): at 2AM vs 2PM, booty-call probability averaged over
all orderings goes 0.42 → 0.05; a baby dragon reads life-changing 0.65 / world-changing 0.25 with only
2/3 rotations agreeing (one ordering said 0.925), the lottery ticket reads life-changing unanimously,
and the paperclip control reads trivial 0.999.
- **LIVE: `semif-serve` 0.1.3** at `http://10.251.50.54:8032` (`semif.fv.internal`), with order averaging
and the fast kernels. SemIf `23cf1f39`, Qwen3.5-4B `851bf6e8`, BF16; token `semif/api-token`. Code +
contract: `services/semif-serve/`; stack `stacks/semif`. **No consumer yet.**
- 0.1.3 acceptance: 144/144 parity with upstream; averaging through the service 78.6% → 88.1%
(95% CI +5.1..+14.3); unanimous rotations 94.5% accurate. Latency, envelope and the fast-kernel A/B
are in `stacks/semif/README.md`.
- The heid bug-hunt panel (thread `01M3H3F4RR7XBP90KQ3A39H4SX`) is triaged and folded into 0.1.3:
C1–C6, S1–S3, S5, S6, S8–S10 fixed with tests; S4 settled; S7 (publish on all interfaces) accepted
as LAN + auth.
- NVFP4 is not worth it (see the SemIf detail files). Prime's probes (2AM/2PM, dragon/lottery) are
recorded in [[2026-09-27-semif-order-averaging]].
### restic: credential leak fixed (2026-09-27, Prime)