1.8 KiB
[2026-09-27] SemIf order-averaging: spiked, Prime ruled build + fast-kernel trial
Finding (Prime's probes): SemIf's single-ordering answer leans toward whichever option is listed first on ambiguous inputs. For "It's 2AM and I'm bored", the top pick flipped from casual 0.685 to booty call 0.760 when the order was reversed.
Spike (739aa03, services/semif-serve/spike/, no service change). SemIf authored144 +
perturbations108, 252 rows in 72 groups, 3 options each, with all 6 orderings of every row sent in one
shared request:
| method | accuracy |
|---|---|
| single ordering, as sent | 78.6% |
| expected single ordering | 80.0% |
| 3 rotations, log-mean | 87.7% (+9.1 pts, group-bootstrap 95% CI +4.7..+13.8) |
| all 6 orderings, log-mean | 88.1% |
Rows where the rotations agree unanimously: 161 rows, 94.4% accurate. Split rows: 91 rows, 75.8%. The service is deterministic, so the CI measures item sampling, not run noise. This is one task family (evidence interpretation, 3 options), not our workload.
Latency (spike/latency-2026-09-27.txt, from nh3-dev, network floor 31 ms):
/decideshort: 71 ms end to end (38 ms server);- 3 rotations shared: 113 ms (79 ms);
- 6 orderings: 137 ms (99 ms);
- a ~2k-token state: 210 ms (169 ms).
Prime ruled (2026-09-27 ~0250): build it in, and try the fast kernels too.
- Averaging design: opt-in
orderings: rotations|all(all only ≤ 4 options); all orderings in one shared batch; per-ordering SemIf results returned unchanged; acombinedblock (log-mean, top, agreement, spread); orderings count towardmax_decisions. - Fast kernels:
causal_conv1d+flash-linear-attentionare missing, so transformers runs Qwen3.5's reference PyTorch paths. Adopt them only if the authored144 parity re-check holds.
Tracked in the in-flight SemIf section. See 2026-09-27-semif-live-on-fv-ml1-gpu1.