Files
esh-pfi-infrastructure/services/semif-serve/spike/latency-2026-09-27.txt
T
vh 739aa03123 spike(semif): latency profile and order-averaging measurement (no service change)
Latency, measured from nh3-dev (3 runs x 20 per condition; network floor 31 ms):
- /decide short: 71 ms end to end, 38 ms server-side;
- /decide with a ~2,000-token state: 210 / 169 ms;
- shared, 3 rotations: 113 / 79 ms;
- shared, 6 orderings: 137 / 99 ms.
Qwen3.5's fast kernels (causal_conv1d, flash-linear-attention) are not installed,
so transformers falls back to its reference PyTorch paths. That is a speed lever,
and using it needs a parity re-check.

Averaging over option orderings, on SemIf authored144 + perturbations108 (252 rows,
72 groups):
- a single ordering scores 78.6%;
- log-mean over the 3 rotations scores 87.7% (+9.1 pts, group-bootstrap 95% CI +4.7
  to +13.8);
- all 6 permutations score 88.1%.
Rotations capture nearly all of the gain. Rows where the rotations agree unanimously
(161) are 94.4% accurate; split rows (91) are 75.8%.
2026-09-27 02:50:14 -07:00

8 lines
539 B
Plaintext

condition e2e p50 (runs) server p50 (runs)
health (floor) 31.1 [30.1-33.4] -
decide, short (~130 tok) 71.2 [70.6-71.4] 38.4 [38.4-38.4]
decide, long (~2,000 tok) 210.0 [209.5-211.0] 168.7 [168.3-168.9]
shared, 3 rotations (short) 112.6 [111.9-113.4] 78.8 [78.7-78.9]
shared, 6 orderings (short) 136.6 [135.2-139.5] 98.9 [98.6-99.0]
shared, 3 rotations (long) 272.1 [270.3-273.7] 228.1 [227.9-228.7]