7d0d991fcd
The probe finished. Median is 79.36 s/it across ten timed steps with a min/max of 79.30 to 79.45, and peak memory is 75.1 of 121.6 GiB by PyTorch's own max_memory_allocated -- 46 GiB spare rather than the 35 I estimated from a live free reading, which was counting fragmentation and the resident model rather than the allocation high-water mark. Resolved attention backend recorded as flex_attention, read off the loaded model rather than trusted from the request, which is the check I got wrong the first time.