Rehome run 3c from ana-ml2 to pfi-gx10 unchanged: same corpus, base,
recipe and hyperparameters, different host. Slower (~13.3 h vs ~2.5 h)
and correct — an Anaheim breaker trip costs a 40-minute drive each way
and 13 hosts down, three of them SureFire client machines, while the
GX10 is a ~240 W appliance at NH3 that can take nothing else down.
Verified rather than assumed, because ana-ml2 ran transformers 5.15.1
on x86-64 and this box runs 5.16.1 on aarch64 — the silent
backend-delta class that has already voided conclusions here:
- both 49 GB base shards sha256-match ana-ml2's (size equality is a
weaker claim and was already true)
- a full encode was run into a throwaway dir and the encoded corpus
compared byte-for-byte: 197,360,233 B, sha256 c08bb1fe2ecb0be3,
identical. Every aggregate matched too. That verified artifact is
what the run will train on — it is seeded into run-03c/encode-cache
- the harness's own suite: 122 passed on aarch64
- the config generator asserts key-by-key that no non-path value
differs from run-03c.json
The encode-cache filename differs by design (base_model_path is part of
the key) — an input hash, not an output hash. Documented so it is not
misread as drift, or "fixed" by faking /tank on this box.
Corpus is copied to local NVMe; the box mounts no NFS. nh3-nas is now on
the same subnet, which makes mounting it tempting and still wrong under
a 13 h unattended run.
The launcher refuses on a live pidfile rather than a pgrep: `pgrep -f
erp_sft_harness` invoked over ssh matches the invoking shell's own argv.
That self-match cost a shell during staging.
Not launched. 13.3 h is the operator's call.
86 lines
3.8 KiB
Markdown
86 lines
3.8 KiB
Markdown
# pfi-gx10 — ASUS Ascent GX10 (NVIDIA GB10)
|
|
|
|
Grace-Blackwell desktop supercomputer. Registered 2026-09-01.
|
|
|
|
| | |
|
|
|---|---|
|
|
| GPU | **NVIDIA GB10**, driver 580.173.02, **compute capability 12.1 (`sm_121`)** |
|
|
| CPU | 20 cores, **aarch64** |
|
|
| Memory | **121 GB unified** (CPU and GPU share it — not 121 GB *plus* VRAM) |
|
|
| Storage | 916 GB NVMe, 6% used |
|
|
| Kernel | 6.17.0-1031-nvidia |
|
|
| Hostname | `pfi-gx10` (shipped with static `gx10-a745`, corrected) |
|
|
|
|
## Network — racked, and single-path
|
|
|
|
Racked 2026-09-03. `pfi-gx10.nh3.internal` → **10.100.50.60**, wired only on
|
|
`enP7s7`, VLAN 50 (`nh3-servers`), UniFi switch port 22.
|
|
|
|
**The address lives on the switch side, not the host** — a DHCP *reservation*
|
|
against the wired MAC `30:c5:99:3d:a7:45`, with the host left on DHCP. Operator
|
|
ruling: a reservation moves with the box, a netplan static goes stale the moment
|
|
it does.
|
|
|
|
⚠ **Wi-Fi is deliberately off and there is now exactly ONE path in.** If the
|
|
switch port or the reservation breaks, this is a rack visit. Correct end state
|
|
for a racked server, but it is a posture change from the desk setup.
|
|
|
|
Full detail, including the order that made the move safe:
|
|
[`docs/runbooks/gx10-rack-network.md`](../../docs/runbooks/gx10-rack-network.md).
|
|
|
|
⚠ **The box mounts no NFS, on purpose.** Working data is copied to local NVMe —
|
|
see the training section below.
|
|
|
|
## Access
|
|
|
|
`infra-ops` with NOPASSWD sudo (operator-bootstrapped). `lkraven` also has key
|
|
auth but needs a password for sudo — **automation must connect as `infra-ops`**.
|
|
|
|
## Headless conversion
|
|
|
|
`playbooks/gx10-headless.yaml` — run it with the `infra-ops@` prefix, since
|
|
elway's `--sudo` applies only to ad-hoc commands and playbook steps carry their
|
|
own.
|
|
|
|
Ships booting to `graphical.target` with GDM and GNOME Remote Desktop. The
|
|
playbook sets `multi-user.target`, stops the remote-desktop service, masks the
|
|
sleep/suspend/hibernate targets, tells logind to ignore lid and idle, and adds
|
|
sshd keepalives so a stalled link does not kill a long job.
|
|
|
|
⚠ **GDM is `static` on Ubuntu** — pulled in by `display-manager.service`, never
|
|
"enabled". Guard and verify on `is-active`, not `is-enabled`; the latter passes
|
|
trivially while the desktop is still running.
|
|
|
|
The playbook will not stop GDM while someone holds a seat session. Override with
|
|
`--var force_dm_stop=true`, or just let the rack-install reboot handle it.
|
|
|
|
## Training — run 3c is staged and ready
|
|
|
|
The ERP-seat SFT LoRA (run 3c) that died on ana-ml2 at step 24 of 604 to an
|
|
Anaheim breaker trip is staged here, unchanged, and **not launched** — that call
|
|
is the operator's.
|
|
|
|
ssh infra-ops@10.100.50.60 '~/erp-tune/launch-run-03c.sh'
|
|
|
|
~79.4 s/it measured on this geometry → 604 steps ≈ 13.3 h, peak 75.1 of
|
|
121.6 GiB. Slower than ana-ml2's ~2.5 h and still the right host: this is a
|
|
~240 W appliance at NH3 that cannot take a client's machines dark.
|
|
|
|
Base model and the encoded corpus were both sha256-verified identical to
|
|
ana-ml2's, so the library delta (transformers 5.15.1 → 5.16.1, x86-64 → aarch64)
|
|
is measured to be inert rather than assumed harmless. Runbook:
|
|
[`docs/runbooks/gx10-run-03c.md`](../../docs/runbooks/gx10-run-03c.md); canonical
|
|
config + launcher in [`scripts/erp-tune-gx10/`](../../scripts/erp-tune-gx10/).
|
|
|
|
⚠ **Never `pkill -f erp_sft_harness` over SSH** — the pattern is in your own ssh
|
|
argv and you kill your shell with it. Kill by PID from `~/erp-tune/run-03c.pid`.
|
|
|
|
## Relevance to Flash-Next
|
|
|
|
`sm_121`, not `sm_120`. The SGLang fork evaluated for ana-ml2 (henge item 49)
|
|
narrows to **exact SM120 and explicitly excludes SM121/GB10** — it does not apply
|
|
here. This chip has its own path: the DGX Spark recipe, which mmaps the ~48 GiB
|
|
PLE n-gram table from NVMe rather than holding it in memory. 121 GB unified and
|
|
822 GB of free NVMe make that viable on this box in a way it is not on a 96 GB
|
|
discrete card.
|