dce261335e
The operator overruled my ana-ml2 recommendation and was right. I had step times in hand, so I priced a breaker trip as eleven minutes of lost training. That is the recompute cost and it is the cheapest component of the loss. A breaker trip at Anaheim is a forty-minute drive each way on the operator's time, whenever he happens to notice, with thirteen hosts dark until he arrives -- including the hypervisor most of them run on, the fleet's primary backup server, and three SureFire client machines that are a customer's production hosts under a hosting agreement. So save_steps 100 to 50 caps the recompute, not the outage, and it was never the mitigation I claimed. Thirteen hours unattended on a desk in NH3 beats two and a half hours that can put a client's hosts dark, especially when nothing is waiting on this run. Recorded the general form at length because it is the transferable part: when recommending between options, check whether you priced the failure mode in whatever units you happened to be measuring. A metric in hand will volunteer itself as the unit of risk. Also measured and dismissed the obvious third option: the RTX PRO 6000s have a 250 W floor against a 300 W default, so capping both saves 100 W on a box drawing about a kilowatt. Not enough to matter, and it costs throughput to buy.