Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-27-semif-order-averaging.md
T

2.4 KiB

[2026-09-27] SemIf order-averaging: spiked, Prime ruled build + fast-kernel trial

Finding (Prime's probes): SemIf's single-ordering answer leans toward whichever option is listed first on ambiguous inputs. For "It's 2AM and I'm bored", the top pick flipped from casual 0.685 to booty call 0.760 when the order was reversed.

Spike (739aa03, services/semif-serve/spike/, no service change). SemIf authored144 + perturbations108, 252 rows in 72 groups, 3 options each, with all 6 orderings of every row sent in one shared request:

method accuracy
single ordering, as sent 78.6%
expected single ordering 80.0%
3 rotations, log-mean 87.7% (+9.1 pts, group-bootstrap 95% CI +4.7..+13.8)
all 6 orderings, log-mean 88.1%

Rows where the rotations agree unanimously: 161 rows, 94.4% accurate. Split rows: 91 rows, 75.8%. The service is deterministic, so the CI measures item sampling, not run noise. This is one task family (evidence interpretation, 3 options), not our workload.

Latency (spike/latency-2026-09-27.txt, from nh3-dev, network floor 31 ms):

  • /decide short: 71 ms end to end (38 ms server);
  • 3 rotations shared: 113 ms (79 ms);
  • 6 orderings: 137 ms (99 ms);
  • a ~2k-token state: 210 ms (169 ms).

Prime ruled (2026-09-27 ~0250): build it in, and try the fast kernels too.

  • Averaging design: opt-in orderings: rotations|all (all only ≤ 4 options); all orderings in one shared batch; per-ordering SemIf results returned unchanged; a combined block (log-mean, top, agreement, spread); orderings count toward max_decisions.
  • Fast kernels: causal_conv1d + flash-linear-attention are missing, so transformers runs Qwen3.5's reference PyTorch paths. Adopt them only if the authored144 parity re-check holds.

Tracked in the in-flight SemIf section. See 2026-09-27-semif-live-on-fv-ml1-gpu1.

Outcome (0.1.3, 77b8cb4, ~0340 PT): built and deployed. Through the service on the same 252 rows, accuracy is 78.6% → 88.1% (95% CI +5.1..+14.3), and unanimous rows are 94.5% accurate. The fast kernels won the A/B on the empty GPU 3:

  • parity with upstream 144/144 (was 142/144);
  • a ~2k-token /decide 169 → 92 ms;
  • shared capacity 63/51/26/16 rows at ~140/520/1960/3900 prefix tokens (was 52/43/19/13). All of heid's bug-hunt findings were folded into the same release (thread 01M3H3F4RR7XBP90KQ3A39H4SX).