44 lines
2.4 KiB
Markdown
44 lines
2.4 KiB
Markdown
# `[2026-09-27]` SemIf order-averaging: spiked, Prime ruled build + fast-kernel trial
|
|
|
|
**Finding (Prime's probes):** SemIf's single-ordering answer leans toward whichever option is listed
|
|
first on ambiguous inputs. For "It's 2AM and I'm bored", the top pick flipped from casual 0.685 to booty
|
|
call 0.760 when the order was reversed.
|
|
|
|
**Spike** (`739aa03`, `services/semif-serve/spike/`, no service change). SemIf authored144 +
|
|
perturbations108, 252 rows in 72 groups, 3 options each, with all 6 orderings of every row sent in one
|
|
shared request:
|
|
|
|
| method | accuracy |
|
|
|---|---|
|
|
| single ordering, as sent | 78.6% |
|
|
| expected single ordering | 80.0% |
|
|
| 3 rotations, log-mean | **87.7%** (+9.1 pts, group-bootstrap 95% CI +4.7..+13.8) |
|
|
| all 6 orderings, log-mean | 88.1% |
|
|
|
|
Rows where the rotations agree unanimously: 161 rows, 94.4% accurate. Split rows: 91 rows, 75.8%.
|
|
The service is deterministic, so the CI measures item sampling, not run noise. This is one task family
|
|
(evidence interpretation, 3 options), not our workload.
|
|
|
|
**Latency** (`spike/latency-2026-09-27.txt`, from nh3-dev, network floor 31 ms):
|
|
- `/decide` short: 71 ms end to end (38 ms server);
|
|
- 3 rotations shared: 113 ms (79 ms);
|
|
- 6 orderings: 137 ms (99 ms);
|
|
- a ~2k-token state: 210 ms (169 ms).
|
|
|
|
**Prime ruled (2026-09-27 ~0250): build it in, and try the fast kernels too.**
|
|
- Averaging design: opt-in `orderings: rotations|all` (all only ≤ 4 options); all orderings in one shared
|
|
batch; per-ordering SemIf results returned unchanged; a `combined` block (log-mean, top, agreement,
|
|
spread); orderings count toward `max_decisions`.
|
|
- Fast kernels: `causal_conv1d` + `flash-linear-attention` are missing, so transformers runs Qwen3.5's
|
|
reference PyTorch paths. Adopt them only if the authored144 parity re-check holds.
|
|
|
|
Tracked in the in-flight SemIf section. See [[2026-09-27-semif-live-on-fv-ml1-gpu1]].
|
|
|
|
**Outcome (0.1.3, `77b8cb4`, ~0340 PT):** built and deployed. Through the service on the same 252 rows,
|
|
accuracy is 78.6% → 88.1% (95% CI +5.1..+14.3), and unanimous rows are 94.5% accurate. The fast kernels
|
|
won the A/B on the empty GPU 3:
|
|
- parity with upstream 144/144 (was 142/144);
|
|
- a ~2k-token `/decide` 169 → 92 ms;
|
|
- shared capacity 63/51/26/16 rows at ~140/520/1960/3900 prefix tokens (was 52/43/19/13).
|
|
All of heid's bug-hunt findings were folded into the same release (thread `01M3H3F4RR7XBP90KQ3A39H4SX`).
|