Checked the disk rather than waiting on comfy-dev. Both builds are there, pulled a minute apart on Aug 23: a 26 GB int8 and a 15 GB nvfp4-awq. I had asked whether their nvfp4 pin was set under a Blackwell assumption, since Ada has no native nvfp4, which would make the int8 file the right one on the new box. It cannot be. They pinned it on irv-ml1's A6000, which is Ampere sm_86 and has neither native nvfp4 nor native fp8. Ada sm_89 supports a strict superset, so a pin that was correct on the weaker card cannot be invalidated by moving to the stronger one. The migration is incapable of breaking it. The pin is about VRAM, not architecture. Eleven gigabytes on a 48 GB card that also holds a DiT and two VAEs decides whether a graph runs, and a text encoder runs once per prompt rather than once per diffusion step, so its throughput matters far less than the DiT's. That also explains why this pin went the opposite way from their other one without either being inconsistent. The RTX 6000 Ada is also 48 GB, so nothing relaxes. Reclassified the 26 GB int8 from orphan to spare: with the extra drives the destination lands near 14% full, so disk stops being the constraint and the pin-rot argument says keep it. Question withdrawn to comfy-dev.
332 lines
19 KiB
Markdown
332 lines
19 KiB
Markdown
# `[2026-09-01]` Ada migration settled on `zfs send` — and branch (b) was never available
|
|
|
|
The Ada box (ComfyUI's new home, sm_89, x86-64) lands at **NH3**. irv-ml1 is in **Irvine**.
|
|
comfy-dev asked whether `/storetank` rides along or the stack is rebuilt from source.
|
|
|
|
## The answer: (a) `zfs send`. Measured, not derived.
|
|
|
|
NH3 -> irv-ml1 11-26 ms, 0% loss
|
|
throughput 99.0 MB/s (real 800 MB transfer over the WireGuard tunnel)
|
|
payload 1.38 TB -> ~3.9 hours
|
|
|
|
`zfs send` is **incremental**: snapshot now, ship the base over ~4 hours while irv-ml1 keeps
|
|
serving, then a small delta at cutover. Near-zero service interruption.
|
|
|
|
## Why (b) — physically moving the disks — was rejected
|
|
|
|
`/storetank` is a two-disk **mirror** of Crucial MX500 2TB SATA SSDs, so it is genuinely
|
|
portable hardware. Rejected anyway: the send is non-disruptive at a measured 99 MB/s, and
|
|
physically handling disks across sites to save ~4 hours of background transfer is a bad
|
|
trade. Not reopened.
|
|
|
|
⚠ **CORRECTED 2026-09-01 by comfy-dev — the ORIGINAL reason given was wrong.** infra-ops
|
|
wrote that (b) "was never available" because *"irv-ml1 keeps running its eight services and
|
|
needs the pool."* comfy-dev enumerated the mounts of all twelve running containers:
|
|
|
|
comfyui -> /storetank/arbo/models:/basedir/models <- the complete list
|
|
|
|
arbo, dockge, tts-gateway, dots-tts, voice-studio, waterland-studio, yt-voice-clipper (x2),
|
|
omnivoice-ref, studio-gate and parakeet **do not touch `/storetank` at all.** Exactly one
|
|
container mounts it, and it is the one that moves. So (b) was unavailable **during the
|
|
transition** — while comfyui still serves from those disks — not structurally. After cutover
|
|
the constraint evaporates. Right conclusion, wrong reason, and the wrong reason mattered
|
|
because it mis-describes what irv-ml1 is holding afterwards (see the retain question below).
|
|
|
|
⚠ **The lesson, restated:** infra-ops reasoned about the BOX when the question was about the
|
|
MOUNT — the same wrong-object shape as asking whether the DATA could move without asking
|
|
whether the SOURCE still needed it. Both were one `docker ps` / `docker inspect` away.
|
|
See [[2026-09-01-wrong-object-measurement]].
|
|
|
|
## (c) rebuild-from-source: rejected on reproducibility, not time
|
|
|
|
~6 hours of re-fetch at ~65 MB/s. The real objection is that **two of comfy-dev's pins are
|
|
already paywalled** (Big Love went permanent-paid; Moody Krea 2 Mix's newest releases are
|
|
gated while their pin is free). A rebuild today would **not reproduce today's stack**. A
|
|
fallback that provably cannot restore what it exists to restore is not a fallback.
|
|
|
|
## ⚠ The two-boxes confusion — do not repeat it
|
|
|
|
There are **TWO new machines** and infra-ops collapsed them into one:
|
|
|
|
| | |
|
|
|---|---|
|
|
| **Ada box** | ComfyUI's target. sm_89, x86-64. Lands at NH3. Not yet arrived. |
|
|
| **ASUS Ascent GX10** | Local inference + run 3c. **GB10, sm_121, aarch64.** On the operator's desk. → [[2026-09-01-pfi-gx10-onboarding]] |
|
|
|
|
infra-ops found no sm_89 part in the current inventory, saw the Ascent, and inferred the
|
|
Ascent was the incoming box — then sent comfy-dev down an aarch64 re-platform investigation
|
|
that was entirely void. **comfy-dev's original premise was correct throughout.**
|
|
|
|
Consequences of the retraction, all restored to their original state:
|
|
- Their single-arch container image, 18 GB x86-64 venv and torch pin are **fine**.
|
|
- Their nvfp4 pin is **RIGHT, not wrong** — sm_89 does not do native nvfp4 (that is
|
|
Blackwell), so pinning away from the 7.74 GB nvfp4 build to the 12.84 GB int8 build was
|
|
correct for the hardware they are actually getting.
|
|
|
|
**Not wasted:** comfy-dev's aarch64 runtime research (ecarmen16/SparkyUI — CUDA 13.0.2,
|
|
torch 2.9.1+cu130 ARM64, SageAttention compiled with `TORCH_CUDA_ARCH_LIST="12.1"`, built on
|
|
the box, ~10 min; verified from source by infra-ops) applies to the **GX10** if anything
|
|
ComfyUI-shaped ever runs there.
|
|
|
|
## Their distinction, worth keeping
|
|
|
|
> **The weights port. The runtime does not.**
|
|
|
|
safetensors are architecture-neutral and travel anywhere. Container image, torch build, venv
|
|
and attention kernels do not. "Migrate the model store" and "migrate ComfyUI" are different
|
|
jobs, and the 1.38 TB transfer is the easy half.
|
|
|
|
## ⚠ DESTINATION SIZING — the as-bought Ada disks are SMALLER than the source pool
|
|
|
|
Measured 2026-09-01 (`zpool list` / `zfs list` on irv-ml1), not derived:
|
|
|
|
source pool storetank SIZE 1.81T ALLOC 1.45T FREE 367G CAP 80%
|
|
mirror-0: 2x ata-CT2000MX500SSD1 (2 TB class)
|
|
compression off, compressratio 1.00x, logicalused == used, 0 snapshots
|
|
|
|
Safetensors are **incompressible** (1.00x measured) — the destination needs the same physical
|
|
bytes regardless of what compression is set to at recv. There is no free win there.
|
|
|
|
**Pool sizes** (decimal TB -> TiB; a mirror vdev shows ~raw member size, confirmed against
|
|
the measured 1.81T for the 2x 2 TB source):
|
|
|
|
2x 1.92 TB mirrored 1.75 TiB <- SMALLER than the 1.82 TiB source pool
|
|
2x 1.92 TB striped 3.49 TiB (no redundancy)
|
|
4 drives, two mirror vdevs 3.57 TiB (1.92 pair + 2.0 pair)
|
|
|
|
**Payload** — ⚠ corrected 2026-09-01, infra-ops' first figure was inflated:
|
|
|
|
1493.9 GiB measured on disk by comfy-dev (weight files)
|
|
~11 GiB remainder of their batch still arriving
|
|
~1.47 TiB settled payload
|
|
~1.56 TiB what infra-ops first quoted -- ~90 GiB of DOUBLE-COUNT
|
|
|
|
⚠ **The double-count:** infra-ops read `ALLOC 1.45T` while the pull was running, then added
|
|
the full ~112 GB on top as if none had landed. Most of it was already inside that reading —
|
|
the 93.4 GiB comfy-dev lists as "today's batch, not yet onboarded" was **on disk, just not
|
|
catalogued**. **Onboarded is not landed.** Occupancy quoted as 90% is really **84%**.
|
|
|
|
config full (1.47 TiB) pruned (-215 GiB)
|
|
as-bought mirror 84% 72%
|
|
as-bought stripe 42% 36%
|
|
+2 drives, two vdevs 41% 35%
|
|
|
|
## ⚠ PRUNE AND DRIVES ARE ORTHOGONAL — pruning cannot rescue the as-bought mirror
|
|
|
|
comfy-dev diffed every weight file against the catalog + all 254 workflow graphs:
|
|
|
|
on disk 1493.9 GiB
|
|
referenced 1121.2 GiB
|
|
unreferenced 374.6 GiB (198 files)
|
|
- today's batch 93.4 GiB expected, catalogued at onboarding
|
|
- *ADA-PENDING* builds 24.2 GiB deliberate, correct on Ada
|
|
= UNEXPLAINED 257.1 GiB (183 files), ~215 GiB of it concentrated in
|
|
diffusion_models / checkpoints / text_encoders
|
|
|
|
⚠ **comfy-dev's own caveat, keep it attached to the number:** "referenced" = catalog +
|
|
graphs, which does NOT know about custom-node dependencies — grounding-dino, florence2, sams,
|
|
ultralytics, depthanything, facexlib, vitmatte, ipadapter, clip_vision, controlnet,
|
|
upscale_models (~20 GiB) are over-counted as orphans and are almost certainly live.
|
|
|
|
**comfy-dev suggested this might land the as-bought disks "nearer 45% than 90%". It does
|
|
not** — 45% is the *striped* figure. Pruning all 215 GiB on the as-bought pair **mirrored**
|
|
gives **72%**, which leaves ~140 GiB before the 80% line on a store that took on ~100 GiB in
|
|
a single day. The capacity problem is a **vdev layout** problem: the as-bought pair mirrored
|
|
is a 1.75 TiB pool whatever we put in it, and deleting payload does not make the destination
|
|
bigger. Only adding drives or dropping redundancy does.
|
|
|
|
**Therefore the audit does not gate the cutover in either direction** — buy the drives on the
|
|
layout argument and the audit becomes leisure hygiene; skip them and the audit cannot rescue
|
|
it. Two independent questions for the operator, not a trade. comfy-dev is NOT deleting
|
|
anything: unreferenced is not safe-to-delete, pins rot, and it is a per-file operator call.
|
|
|
|
**Recommendation to the operator: add 2x 2 TB SATA SSD to the R750xa order** — six of eight
|
|
bays are free, HBA355i has the ports, two mirror vdevs gives ~3.57 TiB at ~41% with
|
|
redundancy intact. Cheapest line on a buy list that already carries 8x `M04W6` and 2x
|
|
`930-00030-1546-000`, and ordering now does not gate the cutover.
|
|
→ [[2026-09-01-ada-inference-server-r750xa]]
|
|
|
|
⚠ **BUILD-TIME GOTCHA — pair like with like.** A ZFS mirror vdev is capped at its SMALLEST
|
|
member. Pairing each new 2 TB drive with an existing 1.92 TB caps both vdevs at 1.92 TB and
|
|
throws away ~150 GiB. Correct layout: the two 1.92s as one vdev, the two new drives as the
|
|
other (1.75 + 1.82 = 3.57 TiB).
|
|
|
|
The no-spend alternative — **striping the as-bought pair** for ~3.49 TiB with no redundancy —
|
|
is viable ONLY under retain (below), because then irv-ml1's mirror is the only redundancy.
|
|
|
|
## ⚠ OPERATOR LEANING 6 DRIVES (2026-09-01, "most likely I add 6 2TB SSDs") — NOT YET DECIDED
|
|
|
|
Six fills **all eight bays**, which turns a capacity top-up into a one-shot, hard-to-reverse
|
|
build decision. Three things change:
|
|
|
|
**1. Layout stops being obvious and becomes permanent.** raidz geometry is fixed at pool
|
|
creation — no conversion to mirrors, no widening, without destroy-and-rebuild.
|
|
|
|
layout (8 drives) usable* full survives
|
|
4x mirror vdevs 6.8 TiB 21.7% 1 per vdev (a both-halves loss kills the pool)
|
|
8-wide raidz2 10.1 TiB 14.6% ANY 2 drives
|
|
* derated ~3-4% for parity padding + ZFS slop
|
|
|
|
**Recommend raidz2.** The usual mirror arguments do not apply here: the workload is large
|
|
sequential reads of big safetensors (raidz's weakness is small random IOPS), 256 GB of ARC
|
|
fronts the metadata and hot checkpoints, SSD resilver has no seek penalty so the classic
|
|
raidz-resilver objection is void, and **"expand two drives at a time" is meaningless once all
|
|
eight bays are full.** raidz2 also survives ANY two failures where 4x mirrors dies to an
|
|
unlucky pair. Neither layout is throughput-limited — the HBA355i (PCIe 4.0 x8) and the
|
|
~6 GB/s backplane ceiling are both far above what checkpoint loading needs.
|
|
|
|
**2. Buy 8, not 6 — a raidz vdev caps every member at the SMALLEST.** The two as-bought
|
|
1.92 TB Hynix drives would cap all eight at 1.92 TB (~480 GB thrown away) AND put two used
|
|
drives of unknown remaining write-endurance, from the reseller who already stripped this box,
|
|
inside the parity set. Two more drives makes the set uniform and demotes the 1.92s to cold
|
|
spares. Cheap insurance on a pool that cannot be re-laid-out later.
|
|
|
|
**3. Drive SIZE is now the permanent ceiling.** The `321-BGKG` backplane is SAS/SATA only, so
|
|
with all bays full the ONLY later expansion is a PCIe NVMe add-in card — and the free-slot
|
|
inventory is still unpulled. Runway, measured in "days like the ~100 GiB acquisition day of
|
|
2026-09-01":
|
|
|
|
2x 2TB (the earlier rec) 3.46 TiB usable 42% ~13 days
|
|
6x 2TB, mixed, raidz2 10.06 TiB usable 14.6% ~67 days
|
|
8x 2TB, uniform, raidz2 10.48 TiB usable 14.0% ~71 days
|
|
8x 4TB, uniform, raidz2 20.95 TiB usable 7.0% ~157 days
|
|
|
|
⚠ The "days like today" rate is a **projection off a single deliberate acquisition batch**,
|
|
not a measured steady rate. Treat as an order-of-magnitude comparator only.
|
|
|
|
**Non-issues, closed:** 8 SATA SSDs draw ~40 W against a ~1 kW box; HBA355i and backplane are
|
|
nowhere near limiting.
|
|
|
|
**What 6-8 drives moots:** at ~14% the prune audit becomes pure hygiene with no capacity
|
|
stake at all. Retain on irv-ml1 stays recommended, but on **reproducibility alone** — the
|
|
redundancy argument for retain disappears under raidz2.
|
|
|
|
|
|
## Incidental — LTX 2.5 was already on disk, uncatalogued
|
|
|
|
comfy-dev's diff turned up 34 GiB of LTX 2.5 (`ltx-2.5-22b-distilled-transformer` 20.0 GiB +
|
|
`gemma4-12b-with-proj-ltx-2.5` 14.3 GiB, both int8-convrot) that nobody knew was there — after
|
|
they had spent part of the day reporting LTX 2.5 as the strategic acquisition of the civitai
|
|
scan and queuing a 15.9 GB REDGraft 2.5 checkpoint on top of a base already present. Their
|
|
domain, no infra-ops action.
|
|
|
|
## ✅ THE H3 ENCODER PIN — infra-ops' question RESOLVED AGAINST ITSELF, no action
|
|
|
|
Verified on disk 2026-09-01, both builds present, pulled a minute apart:
|
|
|
|
26G text_encoders/qwen3vl_32b_minimax_h3_int8_convrot.safetensors Aug 23 21:57
|
|
15G text_encoders/qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors Aug 23 21:58
|
|
|
|
**The question was:** comfy-dev pinned the nvfp4 build; **sm_89 Ada has no native nvfp4** (that
|
|
is Blackwell), and that same fact made their OTHER pin (away from a 7.74 GB nvfp4, toward a
|
|
12.84 GB int8) correct — so was the H3 pin set under a Blackwell assumption, making the int8
|
|
file the one that is right on the new box?
|
|
|
|
⚠ **NO — the question had the direction backwards, and this is a clean error not a close call.**
|
|
comfy-dev pinned it on **irv-ml1's A6000, which is Ampere sm_86** — no native nvfp4 *and* no
|
|
native fp8. **Ada sm_89 supports a strict superset** (adds fp8 e4m3, still no fp4). **A pin
|
|
correct on the weaker card cannot be invalidated by moving to the stronger one.** The migration
|
|
is incapable of breaking it.
|
|
|
|
**What the pin is actually about: VRAM, not architecture.** 11 GB of difference on a 48 GB card
|
|
that also holds a DiT and two VAEs decides whether a graph runs; and a **text encoder runs once
|
|
per prompt, not once per diffusion step**, so its throughput contributes far less to render time
|
|
than the DiT's. Trading encode speed for headroom is right at that ratio — and it explains why
|
|
this pin went the *opposite* way from their other one without either being inconsistent (that
|
|
model was small enough that int8 fit, so speed won). **The RTX 6000 Ada is also 48 GB, so the
|
|
constraint does not relax.** Pin stands.
|
|
|
|
⚠ **The int8 build is surplus but must NOT be read as a delete signal.** With 6+ SSDs the
|
|
destination lands near 14% rather than 84%, so disk stops being the constraint, and comfy-dev's
|
|
own pin-rot argument applies: a 26 GB artifact that may not be re-obtainable is worth more idle
|
|
than reclaimed. Reclassified from "orphan" to "spare". Question withdrawn to comfy-dev; nothing
|
|
owed by them.
|
|
|
|
## THE PATTERN THIS THREAD KEEPS PRODUCING
|
|
|
|
Three findings in one day, all from measurement, each beating a confident derivation:
|
|
|
|
1. comfy-dev's `docker inspect` across 12 containers — corrected infra-ops' branch (b) reason.
|
|
2. infra-ops' `zpool list` — caught the destination being smaller than the source.
|
|
3. comfy-dev's disk-vs-catalog diff — caught 257 GiB of orphans and 34 GiB of unknown LTX 2.5.
|
|
|
|
Same lesson as `docs/pfi/training-throughput-playbook.md` §4, on storage instead of training.
|
|
|
|
## Build-time: ARC sizing, and why NOT a ramfs/tmpfs model tier
|
|
|
|
Operator asked (2026-09-01) whether 512 GB justifies a ramfs for hot models, then self-
|
|
corrected to "or just let ZFS ARC handle it with a 128 GB pool or something." **ARC, yes.
|
|
128 GB, no — that is BELOW the Linux default.**
|
|
|
|
**Do not build a ramfs/tmpfs model tier.** ARC already is the RAM cache, and beats it:
|
|
|
|
- **ARC is elastic; tmpfs is a reservation.** A tmpfs holding a checkpoint costs that RAM
|
|
whether or not it is in use. ARC evicts cold data and hands the RAM back. On a box that
|
|
loads *different* models per workflow, adaptive wins outright.
|
|
- **tmpfs requires a curated hot-list that rots.** ARC decides from live access patterns and
|
|
is scan-resistant by construction (recency AND frequency).
|
|
- **`ramfs` specifically is dangerous** — no size limit, cannot swap, fills until OOM. If
|
|
ever used at all, `tmpfs`.
|
|
- **Not durable** — a 200 GB copy-in on every boot.
|
|
- **No speed to win.** Both are a memcpy; the floor is the host→VRAM PCIe 4.0 x16 hop
|
|
(~20 GB/s real) plus CPU-side deserialization, and neither tier can beat it:
|
|
|
|
cold, 2x SATA mirror 1.1 GB/s ~24 s for a 25 GB checkpoint
|
|
cold, 8-wide raidz2 3.3 GB/s ~9 s
|
|
warm in ARC ~12 GB/s ~3 s
|
|
tmpfs ~12 GB/s ~3 s <- identical, at the cost of curation
|
|
|
|
⚠ **The one honest argument FOR tmpfs:** safetensors loads via **mmap**, and on ZFS-on-Linux
|
|
mmap'd pages are double-buffered (ARC + page cache), so a 25 GB checkpoint can occupy ~50 GB.
|
|
tmpfs is single-copy. At 512 GB this is a capacity annoyance, not a reason to build a second
|
|
storage tier — size ARC for it instead.
|
|
|
|
**`zfs_arc_max`: recommend ~320 GB (62.5%).** OpenZFS on Linux already defaults to **50% of
|
|
RAM = 256 GB** here, so *doing nothing* is twice as good as the proposed 128 GB. A hot working
|
|
set of ~6 big checkpoints is ~150 GB, ~300 GB if mmap double-buffers — 128 GB would thrash
|
|
inside a single session. Do not go past ~75%: ARC shrink-under-pressure is laggy on Linux and
|
|
a memory-hungry ComfyUI on the same box can lose the race.
|
|
|
|
⚠ **This is NOT the [[idle VRAM is reserved, not waste]] case.** `gpu-memory-utilization`
|
|
pre-allocates a fixed pool; `zfs_arc_max` is a **ceiling on an elastic cache** that returns
|
|
memory under pressure. Raising the ARC cap reserves nothing. Do not apply the VRAM rule here.
|
|
|
|
⚠ **`recordsize=1M` cannot be won via `zfs send`.** Records are sized when written; `recv`
|
|
reproduces the source's block structure, so a 128K source (storetank is at the 128K default)
|
|
yields 128K records on the destination no matter what the property says. Getting 1M would mean
|
|
rsync instead of send — **not worth it**: the incremental-delta property at cutover is worth
|
|
far more, and on flash the 128K penalty is metadata overhead (~40 MB of ARC headers per 25 GB
|
|
file vs ~5 MB), not throughput. Set `recordsize=1M` only on datasets built fresh, not received.
|
|
Verify the recv behaviour at cutover rather than taking this on faith.
|
|
|
|
## Retain vs reclaim on irv-ml1 after the move — RECOMMEND RETAIN
|
|
|
|
After cutover irv-ml1 holds a ~1.8 TB mirror whose only remaining live content is ~4 GB of
|
|
tts/llm odds and ends, so 1.4 TB is either reclaimable capacity or a free full backup.
|
|
comfy-dev raised it; infra-ops concurs. **Operator's call.** Two independent reasons:
|
|
|
|
1. **Reproducibility (comfy-dev's).** Two of their pins have gone permanently paid, so a
|
|
from-source rebuild provably cannot reproduce today's stack. The irv-ml1 copy is the only
|
|
surviving reproduction path for some of those weights. Reclaiming early converts a
|
|
recoverable situation into an unrecoverable one for zero gain — nothing waits on the space.
|
|
2. **Redundancy (infra-ops').** If the operator declines the two extra drives and we stripe
|
|
the as-bought pair, retain stops being a recommendation and becomes a **requirement**.
|
|
|
|
⚠ Do not reclaim `/storetank` as routine post-migration tidying.
|
|
|
|
## Open
|
|
|
|
- **Cutover window** — operator's, not yet set.
|
|
- **Drives for the R750xa** — operator leaning **6x 2 TB** (fills all 8 bays). Open: 6-vs-8
|
|
(uniform set), 2 TB vs 4 TB (bays-full makes size the permanent ceiling), and raidz2 vs
|
|
mirrors (recommend raidz2). Not yet ordered.
|
|
- **Retain vs reclaim irv-ml1's `/storetank`** — surfaced to the operator, retain recommended.
|
|
- **comfy-dev's per-file prune audit** — offered, explicitly NOT gating the cutover, waiting
|
|
on the operator wanting it. Their judgment, not infra-ops'.
|
|
- **comfy-dev's ~112 GB batch** — unheld by infra-ops; they deliberately did not pull until
|
|
the operator approved putting that much onto his infrastructure. Manifest pinned and staged
|
|
(`29324e9`). **Landed 2026-09-01**, SHAs green; 93.4 GiB on disk awaiting catalogue
|
|
onboarding, ~11 GiB still arriving at last measurement.
|
|
|
|
Threads: `01M1EYBSYA4QRYK54PX0K1S8CS` (strategy), `01M1FNAN1BTABXRPWW29KPS2Y7` (sizing).
|