Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-27-semif-order-averaging.md
T

1.8 KiB

[2026-09-27] SemIf order-averaging: spiked, Prime ruled build + fast-kernel trial

Finding (Prime's probes): SemIf's single-ordering answer leans toward whichever option is listed first on ambiguous inputs. For "It's 2AM and I'm bored", the top pick flipped from casual 0.685 to booty call 0.760 when the order was reversed.

Spike (739aa03, services/semif-serve/spike/, no service change). SemIf authored144 + perturbations108, 252 rows in 72 groups, 3 options each, with all 6 orderings of every row sent in one shared request:

method accuracy
single ordering, as sent 78.6%
expected single ordering 80.0%
3 rotations, log-mean 87.7% (+9.1 pts, group-bootstrap 95% CI +4.7..+13.8)
all 6 orderings, log-mean 88.1%

Rows where the rotations agree unanimously: 161 rows, 94.4% accurate. Split rows: 91 rows, 75.8%. The service is deterministic, so the CI measures item sampling, not run noise. This is one task family (evidence interpretation, 3 options), not our workload.

Latency (spike/latency-2026-09-27.txt, from nh3-dev, network floor 31 ms):

  • /decide short: 71 ms end to end (38 ms server);
  • 3 rotations shared: 113 ms (79 ms);
  • 6 orderings: 137 ms (99 ms);
  • a ~2k-token state: 210 ms (169 ms).

Prime ruled (2026-09-27 ~0250): build it in, and try the fast kernels too.

  • Averaging design: opt-in orderings: rotations|all (all only ≤ 4 options); all orderings in one shared batch; per-ordering SemIf results returned unchanged; a combined block (log-mean, top, agreement, spread); orderings count toward max_decisions.
  • Fast kernels: causal_conv1d + flash-linear-attention are missing, so transformers runs Qwen3.5's reference PyTorch paths. Adopt them only if the authored144 parity re-check holds.

Tracked in the in-flight SemIf section. See 2026-09-27-semif-live-on-fv-ml1-gpu1.