memory: snapshot — gx10 unracked and next up as an inference+training box
Run 3c's new intended home is the GX10 rather than a power triage on ana-ml2: a ~240 W appliance instead of the kilowatt-class box that tripped the breaker, and 121 GB unified holds the 49 GB bf16 base comfortably where ana-ml2 was tight. The box is bare, so the first move is a throughput probe rather than a harness port. Also banked: the Ada migration settling on zfs send with branch (b) ruled out by irv-ml1 keeping its eight services; the Synapse 39-release upgrade with its one-way schema migration, the appservice namespace opening and the admin API lockdown; the ratified room alias convention; and a named failure class — a correct check aimed at the wrong object — with six instances from one day across three sessions.
This commit is contained in:
@@ -0,0 +1,75 @@
|
||||
# `[2026-09-01]` Ada migration settled on `zfs send` — and branch (b) was never available
|
||||
|
||||
The Ada box (ComfyUI's new home, sm_89, x86-64) lands at **NH3**. irv-ml1 is in **Irvine**.
|
||||
comfy-dev asked whether `/storetank` rides along or the stack is rebuilt from source.
|
||||
|
||||
## The answer: (a) `zfs send`. Measured, not derived.
|
||||
|
||||
NH3 -> irv-ml1 11-26 ms, 0% loss
|
||||
throughput 99.0 MB/s (real 800 MB transfer over the WireGuard tunnel)
|
||||
payload 1.38 TB -> ~3.9 hours
|
||||
|
||||
`zfs send` is **incremental**: snapshot now, ship the base over ~4 hours while irv-ml1 keeps
|
||||
serving, then a small delta at cutover. Near-zero service interruption.
|
||||
|
||||
## Why (b) — physically moving the disks — was never on the table
|
||||
|
||||
`/storetank` is a two-disk **mirror** of Crucial MX500 2TB SATA SSDs, so it is genuinely
|
||||
portable hardware. That fact is true and **irrelevant**.
|
||||
|
||||
**irv-ml1 is NOT being decommissioned.** It runs arbo, comfyui, tts-gateway, dots-tts,
|
||||
voice-studio, waterland-studio, yt-voice-clipper and dockge — and **comfyui bind-mounts
|
||||
`/storetank/arbo/models`**. Pulling those disks does not inconvenience a source box that no
|
||||
longer needs them; it guts a live one.
|
||||
|
||||
⚠ **The lesson:** infra-ops asked whether the DATA could move and never asked whether the
|
||||
SOURCE still needed it. One `docker ps` away, at any point. Recommended (b) twice on the
|
||||
strength of a true-but-irrelevant fact. See [[2026-09-01-wrong-object-measurement]].
|
||||
|
||||
## (c) rebuild-from-source: rejected on reproducibility, not time
|
||||
|
||||
~6 hours of re-fetch at ~65 MB/s. The real objection is that **two of comfy-dev's pins are
|
||||
already paywalled** (Big Love went permanent-paid; Moody Krea 2 Mix's newest releases are
|
||||
gated while their pin is free). A rebuild today would **not reproduce today's stack**. A
|
||||
fallback that provably cannot restore what it exists to restore is not a fallback.
|
||||
|
||||
## ⚠ The two-boxes confusion — do not repeat it
|
||||
|
||||
There are **TWO new machines** and infra-ops collapsed them into one:
|
||||
|
||||
| | |
|
||||
|---|---|
|
||||
| **Ada box** | ComfyUI's target. sm_89, x86-64. Lands at NH3. Not yet arrived. |
|
||||
| **ASUS Ascent GX10** | Local inference + run 3c. **GB10, sm_121, aarch64.** On the operator's desk. → [[2026-09-01-pfi-gx10-onboarding]] |
|
||||
|
||||
infra-ops found no sm_89 part in the current inventory, saw the Ascent, and inferred the
|
||||
Ascent was the incoming box — then sent comfy-dev down an aarch64 re-platform investigation
|
||||
that was entirely void. **comfy-dev's original premise was correct throughout.**
|
||||
|
||||
Consequences of the retraction, all restored to their original state:
|
||||
- Their single-arch container image, 18 GB x86-64 venv and torch pin are **fine**.
|
||||
- Their nvfp4 pin is **RIGHT, not wrong** — sm_89 does not do native nvfp4 (that is
|
||||
Blackwell), so pinning away from the 7.74 GB nvfp4 build to the 12.84 GB int8 build was
|
||||
correct for the hardware they are actually getting.
|
||||
|
||||
**Not wasted:** comfy-dev's aarch64 runtime research (ecarmen16/SparkyUI — CUDA 13.0.2,
|
||||
torch 2.9.1+cu130 ARM64, SageAttention compiled with `TORCH_CUDA_ARCH_LIST="12.1"`, built on
|
||||
the box, ~10 min; verified from source by infra-ops) applies to the **GX10** if anything
|
||||
ComfyUI-shaped ever runs there.
|
||||
|
||||
## Their distinction, worth keeping
|
||||
|
||||
> **The weights port. The runtime does not.**
|
||||
|
||||
safetensors are architecture-neutral and travel anywhere. Container image, torch build, venv
|
||||
and attention kernels do not. "Migrate the model store" and "migrate ComfyUI" are different
|
||||
jobs, and the 1.38 TB transfer is the easy half.
|
||||
|
||||
## Open
|
||||
|
||||
- **Cutover window** — operator's, not yet set.
|
||||
- **comfy-dev's ~112 GB batch** — unheld by infra-ops, but they are deliberately NOT pulling
|
||||
until the operator approves putting that much onto his infrastructure. Their call, correct
|
||||
instinct. Manifest pinned and staged (`29324e9`).
|
||||
|
||||
Thread: `01M1EYBSYA4QRYK54PX0K1S8CS`.
|
||||
Reference in New Issue
Block a user