b039aa19e8
Operator clarified the purpose, so the record now leads with it: this was a baseline for the box and a check that the tooling loads, not a decision about where run 3c runs. The placement reasoning stays because it is sound, but it is marked as a byproduct rather than the deliverable. Two things were actually delivered. The box trains: aarch64 and sm_121 run torch 2.14.0+cu130 with transformers, accelerate, peft, trl, datasets, safetensors and bitsandbytes, plus the harness's own flex_attention backend and chunked-loss path, and nothing beyond python3-dev was needed. And the baseline is 79.35 s/it median across seven timed steps with a 0.19% spread. Also recorded the port scope without executing it, so nobody re-derives it: about 2.5 GB of data, a venv rebuild on aarch64, no encode cache worth moving since the encode runs in 14 seconds, and copy the corpus rather than mounting NFS on a desk box that will be unattended for hours.