5a24d77f12
The GX10 runs run 3c at about 79.3 s/it with 35 GB of headroom, which puts 604 steps at 13.3 hours against ana-ml2's 2.2 to 2.7. Six times slower where raw compute predicts 2.7, which points at memory bandwidth rather than FLOPs -- recorded as a hypothesis, since confirming it needs a bandwidth-bound microbenchmark nobody has run. That reverses the standing plan. Moving run 3c here was framed as the power answer, but the run did not die because ana-ml2 is unreliable. It died because save_steps was 100 and the breaker tripped at step 80, so no checkpoint existed. save_steps is now 50, which caps a power event at about eleven minutes. Trading 2.5 hours for 13.3 buys insurance against a risk already engineered out. Also records the five launch failures and their causes, and the one that matters most: the attention-backend trap was present and I first declared it absent. I checked whether flash-attn was installed, which is the wrong discriminator; the harness sets flex_attention explicitly in code. Absence of an alternative is not evidence of the default. The probe now reads the resolved backend back off the loaded model, and flex_attention does compile and run on sm_121.