spike(semif): latency profile and order-averaging measurement (no service change)
Latency, measured from nh3-dev (3 runs x 20 per condition; network floor 31 ms): - /decide short: 71 ms end to end, 38 ms server-side; - /decide with a ~2,000-token state: 210 / 169 ms; - shared, 3 rotations: 113 / 79 ms; - shared, 6 orderings: 137 / 99 ms. Qwen3.5's fast kernels (causal_conv1d, flash-linear-attention) are not installed, so transformers falls back to its reference PyTorch paths. That is a speed lever, and using it needs a parity re-check. Averaging over option orderings, on SemIf authored144 + perturbations108 (252 rows, 72 groups): - a single ordering scores 78.6%; - log-mean over the 3 rotations scores 87.7% (+9.1 pts, group-bootstrap 95% CI +4.7 to +13.8); - all 6 permutations score 88.1%. Rotations capture nearly all of the gain. Rows where the rotations agree unanimously (161) are 94.4% accurate; split rows (91) are 75.8%.
This commit is contained in: