Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-01-ada-inference-server-r750xa.md
T
vh 2cf73fd556 memory: snapshot — 768 GB is not a population the R750xa takes
Operator asked why not move all 768 GB across. It does not fit the
board's shape: the R640 has 24 slots at 6 channels/socket x 2 DPC, the
R750xa has 16 at 8 channels/socket x 1 DPC.

768 GB is either 24x 32 GB (more DIMMs than slots) or 12x 64 GB (fits,
but populates 6 of 8 channels per socket and gives up ~25% of memory
bandwidth). The board wants 16 identical DIMMs. So the targets are
512 GB if they are 32s, or 1 TB if they are 64s — taking 12 from one
spare and 4 from the other. In the 64 GB case the answer beats the
question.

Also recorded: beyond ~512 GB the return is marginal for this workload,
so take 1 TB because it is free rather than because it is needed; 64 GB
LRDIMMs run ~50 W hotter in a chassis whose high-performance fans are
unaccounted for on the invoice; and DIMM slot count now needs to be on
the iDRAC pull, since the 16-slot figure is inferred from the factory
CSV and is load-bearing for a 512-vs-1024 decision.

Memory-only; no version bump per the SemVer SKIP list.
2026-09-01 17:35:21 -07:00

11 KiB
Raw Blame History

[2026-09-01] The Ada inference server is a Dell R750xa bought used — and the reseller stripped four things Dell shipped

Dell PowerEdge R750xa, service tag JPJ1ZP3, 8-bay 2.5" 2U. Bought from Entera Source (invoice INV-74422, 2026-08-25). 2× NVIDIA RTX 6000 Ada 48 GB to be fitted — confirmed by the operator, NOT L40/L40S. Destination NH3. This is ComfyUI's new home; the GX10 is a separate machine → 2026-09-01-pfi-gx10-onboarding.

As bought

CPU      2x Xeon Silver 4314   16C/32T each = 32C/64T, 135 W each
RAM      8x Micron 16 GB PC4-3200 = 128 GB
NIC      dual 25 GbE SFP28 (Broadcom 57414, OCP 3.0)
boot     BOSS-S2 + 2x 240 GB M.2 SATA, mirrored
storage  2x 1.92 TB SATA SSD (Hynix), 6 of 8 bays free
HBA      PERC HBA355i (LSI SAS3816)
PSU      2x 1400 W Platinum

THE DIFF — factory CSV (JPJ1ZP3.csv, booth 54-puffin) vs the invoice

Dell factory reseller invoice
RAM 16× 16 GB M04W6, "Performance Optimized" = 256 GB 8× = 128 GB
PSU 2400 W redundant (450-AJEV) 2× 1400 W (7DWXY)
GPUs 4× Tesla V100-32GB passive (490-BHUC) none
GPU risers 12XPY (right DW), 9TR6X (left DW) not listed
GPU cables 470-AEVS 4-GPU kit: 4RW1P W4K7M XC48N 6C77X not listed
Riser cables signal CP67W CXYF8 H4D7D N61TK; power HXJDR W1P56; pad C2JNP not listed
Fans 6× FD00R Very High Performance not listed
CPU 2× Xeon Platinum 8362 32C/64T, 265 W, DDR4-3200 (338-CCUS/W51V5) 2× Xeon Silver 4314 16C/32T, 135 W, DDR4-2666
Controller PERC H345 HBA355i — an UPGRADE for us
Drives 2× 480 GB 2× 1.92 TB

"Not listed on the invoice" is not proof of removal. Operator will pull an iDRAC inventory to settle it. Diff that against the CSV rather than the invoice.

The CPU row was MISSING from this table until 2026-09-01 — the Silver 4314s were recorded under "As bought" but never diffed against the factory CSV, so the swap went unremarked. It has two live consequences:

  • Memory speed: Silver 4314 caps at DDR4-2666; Platinum 8362 ran 3200. The box cannot use the 3200 rating of the factory-spec M04W6 parts on the buy list. Any DDR4-2666 RDIMM performs identically in this chassis as configured. ⚠ Verify the 2666 figure on Intel ARK before spending on it.
  • Power, in our favour: 270 W of CPU instead of 530 W. The ~1,020 W sustained figure below already assumes the Silvers and is correct as written.

Cores went 64C/128T → 32C/64T. For an inference/render box that is largely irrelevant; noted so nobody re-derives it as a surprise later.

The GPU power chain — resolved, and my first answer was wrong

RTX 6000 Ada: 300 W, ACTIVE cooling, one 16-pin 12VHPWR (CEM5) connector (nvidia.com).

The R750xa kit 470-AEVS is "W/CPU CONN" — CPU/EPS 8-pin, because V100 takes that. I initially flagged this as a mismatch requiring the PCIe-type cable RCCWC / Dell 470-BBXH. That was wrong and is withdrawnRCCWC is for H100/L40/L40S.

NVIDIA sells the bridge for exactly this case:

930-00030-1546-000 — NVIDIA Power Cable, RTX 6000 Ada Generation, 1× CPU 8-Pin → 1× CEM5 16-Pin

So the chain is 470-AEVS CPU 8-pin → 930-00030-1546-000 → card. The chassis' existing cabling is the correct source type. Two adapters needed, one per card.

⚠ Factory brackets are VFYGM — the V100S extended bracket. RTX 6000 Ada is a different physical card; bracket fit is UNKNOWN, check at fitting.

Active cooling in a passive-card chassis. Six high-performance fans expect to push air through passive cards; the RTX 6000 Ada has its own blower. It works in practice, but Dell never validated this combination for this box.

Storage — NVMe in the drive bays is CLOSED

Factory backplane is 321-BGKG : 2.5" Chassis with up to 8 **SAS/SATA** Drives. Not NVMe-capable; bays would need a backplane swap.

Arithmetic if filling bays with SATA:

2x SATA (today)     ~1.1 GB/s
6x striped          ~3.3 GB/s      <- roughly ONE NVMe drive
8x striped          ~4.4 GB/s
backplane ceiling   ~6 GB/s (8 lanes x SATA III)

The HBA355i (PCIe 4.0 x8) is nowhere near limiting. Drives are the constraint.

ARC probably beats spindles here. ComfyUI reloads the same checkpoints; with the factory 256 GB restored, ZFS ARC caches most of the working set at RAM speed. The RAM upgrade is both cheaper and likely faster than four more SATA drives.

CAPACITY, not throughput, is the binding constraint — the as-bought drives are SMALLER than the pool they receive from (measured 2026-09-01). ComfyUI's /storetank on irv-ml1 is a 2x 2 TB mirror: SIZE 1.81 TiB, ALLOC 1.45 TiB, CAP 80%. Settled payload ~1.47 TiB. The R750xa's 2x 1.92 TB mirrored is only ~1.75 TiB — smaller than the source pool — so the migration arrives at ~84% full with no growth room, past the ~80% line where ZFS allocation degrades. Compression buys nothing: safetensors measured compressratio 1.00x, logicalused == used.

-> Add 2x 2 TB SATA SSD to the buy list. Six of eight bays free, HBA355i has the ports; two mirror vdevs = ~3.57 TiB at ~41%, redundancy intact. Cheapest line on this buy list and it does not gate the cutover window. The no-spend alternative is striping the as-bought pair (~3.49 TiB, no redundancy) — only acceptable if irv-ml1 retains its copy, which makes retain a requirement rather than a recommendation.

Pruning does NOT substitute for the drives. comfy-dev found ~215 GiB of unreferenced weights; deleting every byte still lands the as-bought mirror at 72%. The constraint is vdev layout, not payload — a 1.75 TiB pool stays 1.75 TiB whatever goes in it. The prune audit and the drive purchase are independent decisions.

Build-time: pair like with like. A mirror vdev caps at its smallest member; pairing each new 2 TB with an existing 1.92 TB caps both vdevs at 1.92 TB and wastes ~150 GiB. Correct: the two 1.92s as one vdev, the two new drives as the other. → 2026-09-01-ada-migration-branch-a

UNCHECKED, and it may moot the whole bay question: free PCIe slots. Riser Config 0 is 330-BBUM : Riser Config 0, 6x16 + 2x8 slots; two double-wide GPUs will not consume all of it. An NVMe add-in card (U.2 HBA or bifurcating M.2 carrier) bypasses the SAS/SATA backplane entirely. Add slot inventory to the iDRAC pull.

Power — the ceiling is deliberate now, not discovered later

2x RTX 6000 Ada   600 W
2x Xeon 4314      270 W
rest             ~150 W
               ~1,020 W sustained

One 1400 W PSU carries that, so redundancy holds. But factory specced 2400 W for four V100s: there is no headroom for a third or fourth GPU. And ~1 kW is the same class as ana-ml2, the box whose training run tripped a breaker on 2026-08-26 — decide which NH3 circuit this lands on before racking, not after.

Buy list

THE RAM LINE MAY BE FREE — operator has 2x Dell R640 with 768 GB each (2026-09-01). Both are DDR4 RDIMM platforms (factory M04W6 = DIMM,16GB,3200,2RX8,8G,DDR4,**R**), so the part class is the same 288-pin 1.2 V DDR4 RDIMM and Dell does not vendor-lock DIMMs. Three things decide it, in order:

  1. RDIMM vs LRDIMM. 768 GB in an R640's 24 slots is almost certainly 24x 32 GB 2Rx4 RDIMM — the safe case, drops straight in. If it is instead 12x 64 GB LRDIMM, verify Ice Lake-SP / R750xa LRDIMM support first; do NOT assume it.
  2. Speed is a non-issue as configured. R640 DIMMs are 2666 or 2933; the Silver 4314s cap the box at 2666 regardless (see the CPU note above), so 2666 parts lose nothing against the 3200-rated M04W6. ⚠ It WOULD become a cap if the Platinum 8362s were ever restored — the factory CPUs ran 3200.
  3. Slot arithmetic. R750xa has 16 slots (8 channels/socket, 1 DPC); R640 has 24. 16x 32 GB = 512 GB — double the factory 256 GB spec, and it kills the 8x M04W6 line entirely. Cleanest harvest: strip ONE R640 of all 24, use 16, keep 8 as spares, leaving the second R640 whole — rather than half-emptying both into unbalanced populations.

"Just move all 768 GB across" is not a shape this board takes. R640 = 24 slots (6 ch/socket x 2 DPC); R750xa = 16 slots (8 ch/socket x 1 DPC). 768 GB is either 24x 32 GB (more DIMMs than there are slots) or 12x 64 GB (fits, but populates only 6 of 8 channels per socket, leaving ~25% of memory bandwidth unused). The board wants 16 identical DIMMs, 8 per socket, all channels. So the real targets are:

if 32 GB parts   16x 32 =   512 GB     (16 of the 24 in one box)
if 64 GB parts   16x 64 = 1,024 GB     (12 from one box + 4 from the other)

Neither is 768. And in the 64 GB case the answer is better than the question — 1 TB, not 768 GB, because the two spares hold 24 such DIMMs between them.

⚠ Beyond ~512 GB the return is marginal for this workload (hot set ~150-300 GB; ARC at ~320 GB already covers it). Take 1 TB because it is free, not because it is needed. Two costs to weigh if it lands there: 64 GB LRDIMMs run hotter (~6-8 W each vs ~3-5 W, so ~+50 W over the 32 GB case) in a chassis whose 6 high-performance fans FD00R are unaccounted for on the invoice — and this box already sits at ~1 kW on a site where a training run tripped a breaker on 2026-08-26.

Confirm the actual DIMM part numbers from iDRAC or the DIMM labels before ordering anything. R640s confirmed spares, not in service (operator, 2026-09-01). Also add DIMM SLOT COUNT to the iDRAC pull — the 16-slot figure is inferred from the factory CSV (qty 16 M04W6, "Performance Optimized") and is now load-bearing for a 512-vs-1024 decision.

Certain (pending the R640 harvest above, which may delete the RAM line):

item part qty
RDIMM 16 GB 3200 2Rx8 M04W6 8 → restores 256 GB. Buy only if the R640 harvest fails; 16x 32 GB from an R640 gives 512 GB free
NVIDIA 12VHPWR adapter 930-00030-1546-000 2
2 TB SATA SSD (any; match/exceed the MX500s) 2 → pool ~3.57 TiB at ~41%; without it the migration lands at ~84%. Pair the two NEW drives together, not one-each with a 1.92

Only if the iDRAC inventory shows them absent: 12XPY, 9TR6X, 4RW1P, W4K7M, XC48N, 6C77X, CP67W, CXYF8, H4D7D, N61TK, HXJDR, W1P56, C2JNP, FD00R ×6.

Spend order: check the R640 DIMM part numbers FIRST (may zero the RAM line and double the RAM) → RAM only if that fails → SATA SSDs (certain, capacity-driven) → check PCIe slots → only then consider filling the remaining bays. (The original order deprioritised drives on a throughput argument, which still holds; the drives are now on the certain list for capacity, which is a different axis.)

⚠ Dell's Installation & Service Manual publishes no part numbers at all for GPU kit hardware — names and quantities only. The factory CSV keyed to the service tag is the only authoritative source of DPNs for this machine. Do not accept a five-character DPN from anywhere else without corroboration.