Files
esh-pfi-infrastructure/servers/pfi-gx10
vh dae77ee118 feat(pfi-gx10): stage ERP-seat SFT run 3c — verified, not launched
Rehome run 3c from ana-ml2 to pfi-gx10 unchanged: same corpus, base,
recipe and hyperparameters, different host. Slower (~13.3 h vs ~2.5 h)
and correct — an Anaheim breaker trip costs a 40-minute drive each way
and 13 hosts down, three of them SureFire client machines, while the
GX10 is a ~240 W appliance at NH3 that can take nothing else down.

Verified rather than assumed, because ana-ml2 ran transformers 5.15.1
on x86-64 and this box runs 5.16.1 on aarch64 — the silent
backend-delta class that has already voided conclusions here:

  - both 49 GB base shards sha256-match ana-ml2's (size equality is a
    weaker claim and was already true)
  - a full encode was run into a throwaway dir and the encoded corpus
    compared byte-for-byte: 197,360,233 B, sha256 c08bb1fe2ecb0be3,
    identical. Every aggregate matched too. That verified artifact is
    what the run will train on — it is seeded into run-03c/encode-cache
  - the harness's own suite: 122 passed on aarch64
  - the config generator asserts key-by-key that no non-path value
    differs from run-03c.json

The encode-cache filename differs by design (base_model_path is part of
the key) — an input hash, not an output hash. Documented so it is not
misread as drift, or "fixed" by faking /tank on this box.

Corpus is copied to local NVMe; the box mounts no NFS. nh3-nas is now on
the same subnet, which makes mounting it tempting and still wrong under
a 13 h unattended run.

The launcher refuses on a live pidfile rather than a pgrep: `pgrep -f
erp_sft_harness` invoked over ssh matches the invoking shell's own argv.
That self-match cost a shell during staging.

Not launched. 13.3 h is the operator's call.
2026-09-03 22:46:28 -07:00
..

pfi-gx10 — ASUS Ascent GX10 (NVIDIA GB10)

Grace-Blackwell desktop supercomputer. Registered 2026-09-01.

GPU NVIDIA GB10, driver 580.173.02, compute capability 12.1 (sm_121)
CPU 20 cores, aarch64
Memory 121 GB unified (CPU and GPU share it — not 121 GB plus VRAM)
Storage 916 GB NVMe, 6% used
Kernel 6.17.0-1031-nvidia
Hostname pfi-gx10 (shipped with static gx10-a745, corrected)

Network — racked, and single-path

Racked 2026-09-03. pfi-gx10.nh3.internal10.100.50.60, wired only on enP7s7, VLAN 50 (nh3-servers), UniFi switch port 22.

The address lives on the switch side, not the host — a DHCP reservation against the wired MAC 30:c5:99:3d:a7:45, with the host left on DHCP. Operator ruling: a reservation moves with the box, a netplan static goes stale the moment it does.

Wi-Fi is deliberately off and there is now exactly ONE path in. If the switch port or the reservation breaks, this is a rack visit. Correct end state for a racked server, but it is a posture change from the desk setup.

Full detail, including the order that made the move safe: docs/runbooks/gx10-rack-network.md.

The box mounts no NFS, on purpose. Working data is copied to local NVMe — see the training section below.

Access

infra-ops with NOPASSWD sudo (operator-bootstrapped). lkraven also has key auth but needs a password for sudo — automation must connect as infra-ops.

Headless conversion

playbooks/gx10-headless.yaml — run it with the infra-ops@ prefix, since elway's --sudo applies only to ad-hoc commands and playbook steps carry their own.

Ships booting to graphical.target with GDM and GNOME Remote Desktop. The playbook sets multi-user.target, stops the remote-desktop service, masks the sleep/suspend/hibernate targets, tells logind to ignore lid and idle, and adds sshd keepalives so a stalled link does not kill a long job.

GDM is static on Ubuntu — pulled in by display-manager.service, never "enabled". Guard and verify on is-active, not is-enabled; the latter passes trivially while the desktop is still running.

The playbook will not stop GDM while someone holds a seat session. Override with --var force_dm_stop=true, or just let the rack-install reboot handle it.

Training — run 3c is staged and ready

The ERP-seat SFT LoRA (run 3c) that died on ana-ml2 at step 24 of 604 to an Anaheim breaker trip is staged here, unchanged, and not launched — that call is the operator's.

ssh infra-ops@10.100.50.60 '~/erp-tune/launch-run-03c.sh'

~79.4 s/it measured on this geometry → 604 steps ≈ 13.3 h, peak 75.1 of 121.6 GiB. Slower than ana-ml2's ~2.5 h and still the right host: this is a ~240 W appliance at NH3 that cannot take a client's machines dark.

Base model and the encoded corpus were both sha256-verified identical to ana-ml2's, so the library delta (transformers 5.15.1 → 5.16.1, x86-64 → aarch64) is measured to be inert rather than assumed harmless. Runbook: docs/runbooks/gx10-run-03c.md; canonical config + launcher in scripts/erp-tune-gx10/.

Never pkill -f erp_sft_harness over SSH — the pattern is in your own ssh argv and you kill your shell with it. Kill by PID from ~/erp-tune/run-03c.pid.

Relevance to Flash-Next

sm_121, not sm_120. The SGLang fork evaluated for ana-ml2 (henge item 49) narrows to exact SM120 and explicitly excludes SM121/GB10 — it does not apply here. This chip has its own path: the DGX Spark recipe, which mmaps the ~48 GiB PLE n-gram table from NVMe rather than holding it in memory. 121 GB unified and 822 GB of free NVMe make that viable on this box in a way it is not on a 96 GB discrete card.