Local GPU infrastructure overview#

Status: Current onboarding summary Last consolidated: 2026-09-29 Scope: The three owned RTX 5090 ComfyUI service nodes (PC1, PC2, PC3): what was ordered, why, and how each node is accepted into production Record of the ordered builds: rtx5090-ram-and-cost-performance-decision.md. Where this summary and that record disagree, the record wins.

One-minute summary#

ComfyICU is moving its predictable base load from Modal onto owned, single-GPU RTX 5090 machines, with Modal kept for overflow, rollouts and anything that needs more than 32 GB of VRAM (strategy). Three nodes have been ordered. They share one platform: Ryzen 9 9950X, MSI MAG X870E Tomahawk Max, a Samsung 9100 Pro 2 TB model cache on its own filesystem (no RAID), a separate Samsung OS drive, a Seasonic Vertex PX-1200, an Arctic Liquid Freezer III Pro 360 and a Fractal North XL.

The fleet's main return is demand, not Modal savings (decided 2026-09-26): fastest-in-category cold starts from always-warm, model-resident nodes, and unmetered idle time given away as short, preemptible free, relaxed and win-back work. Long or >32 GB jobs go to rented capacity so they never block the fast lane (capacity lanes).

The benchmarks showed that warm runs are GPU-bound, while cold model changes depend on the whole storage, filesystem, container and host-to-GPU path. Once a single Gen5 NVMe is attached directly and working, RAID 0 and higher synthetic bandwidth did not repeatably save more than a second or two. The biggest free latency lever is routing: keep each model family on a fixed node so its model stays loaded.

The fleet#

PC1 PC2 PC3
Source IdealTech J0926118, 2026-09-21 IdealTech J0926134, 2026-09-23 C-Zone, deposit 2026-09-24, spec confirmed 2026-09-25
GPU GIGABYTE RTX 5090 GAMING OC (600 W) Zotac RTX 5090 (575 W) Zotac RTX 5090 Solid OC (575 W)
RAM T-Create Expert 2×32 GB, EXPO 6000 T-Create Expert 2×32 GB, EXPO 6000 Kingston Fury Beast 2×64 GB, EXPO 5600
M2_1, CPU Gen5 x4 9100 Pro 2 TB, model cache 9100 Pro 2 TB (OS + cache partitions for now) 9100 Pro 2 TB, model cache
M2_2, CPU Gen5 x4 9100 Pro 2 TB, OS/Docker/outputs 9100 Pro 1 TB OS drive, to be bought 9100 Pro 1 TB, OS
Price RM36,748 RM34,999 RM43,324

Common to all three: 9950X, Tomahawk Max, Vertex PX-1200, Arctic LF III Pro 360, North XL, Linux, onboard 5GbE. RM115,071 for the fleet; how each was bought is in procurement.

PC3 is the only 128 GB node; PC1/PC2 stay at 64 GB (the upgrade was closed for them on 2026-09-24). PC3 was ordered ahead of the strategy doc's 30-day production gate; that gate now applies to a fourth node.

Why storage is central to the decision#

The MiniMax H3 input and model set is 42,471,968,783 bytes (42.47 GB, 39.56 GiB); Krea 2 is 18.6 GB. The benchmark does not trust an advertised SSD number or a single fio result. It measures the normal buffered read path after explicitly evicting the files from cache, and schema v4 records counters at the process, loop/RAID and physical-NVMe layers, because the older process counters could not prove that reads passed through an outer buffered loop-backing-file cache on Vast hosts.

Evidence Result What it establishes
Single-versus-RAID dual-workload matrix No repeated RAID win above the predeclared 2 s threshold (0.7–1.3 s) Array bandwidth did not translate into ComfyUI latency
Single physical Corsair MP700 control 4.536 GiB/s reader; 7.302 s cold penalty One Gen5 drive served every cold byte and completed MiniMax without RAID
Single physical Samsung 9100 Pro 4 TB control 5.699 GiB/s reader; 11.776 s penalty on a slower 500 W/Intel host The exact drive family served the workload physically; not a drive A/B
Single 9100 Pro 2 TB hosts 143669 and 148771 9.6–9.8 GiB/s reader; 6.2–7.1 s penalty A single 9100 Pro reaches the ~10 GiB/s tier and a small penalty
Same-machine SN850X RAID replication 11.380 GiB/s reader; 5.419 s penalty RAID can be fast, but no same-machine single-to-RAID test prices its extra value

The six one-hop PCIe 5.0 x4 single-drive hosts in the first cohort averaged a 6.804 GiB/s reader (4.536–10.098). Beyond about 5 GiB/s the application gains little: ComfyUI's loader interleaves reads with model construction and never realised more than about 4.6 GiB/s during a run. Hence one 2 TB 9100 Pro per node, no RAID, and the node's own acceptance run as the deciding measurement. The earlier 10 GiB/s purchase threshold and the automatic two-drive recommendation are withdrawn. The OS drive sits separately in M2_2 so container, log and output writes never contend with cold model reads; on PC1 it is a second 9100 Pro 2 TB (TLC endurance, and one spare part fits either slot). Detail: single vs RAID, RAID analysis, single-path analysis, node record.

Why the platform choices stay#

GPU and PCIe width#

The RTX 5090's 32 GB of VRAM and measured ComfyUI throughput are the reason to build the node; a 16 GB RTX 5080 is not an equivalent substitute. Two independent 5090s give 60–80% more aggregate cold MiniMax throughput than one RTX Pro 6000 for less money, at the cost of a 32 GB per-job ceiling (strategy). Power limit matters: 450 W was about 12% slower warm, while 575 and 600 W were equal, so every node runs at 575 W or more.

PCIe 5.0 x8 does not halve the GPU's compute or VRAM, but it halves host-link bandwidth from about 63.0 to 31.5 GB/s in each direction, which carries model upload, CPU offload and spilled tensors. The nodes run at x16 because the Tomahawk provides it for free. "Keep x16" was the founder's instruction during purchasing, not a measured requirement: three Gen5 x8 hosts in the benchmarks showed no loss, so x8 is acceptable when it buys something real. One GPU per node is chosen for RAM, cooling, failure isolation and how the cards were sold, not lanes (node record).

Motherboard and lane budget#

The 9950X exposes 24 usable device lanes: x16 for the GPU and two x4 paths. The MSI MAG X870E Tomahawk Max gives both x4 paths to independent CPU Gen5 M.2 slots while the GPU keeps x16; the ordered builds use them for the model drive (M2_1) and the OS drive (M2_2). The included PRO X870E-S EVO has only one CPU Gen5 M.2 slot. M2_2 shares its lanes with rear USB4, so it runs x2 until set to x4 in the BIOS.

A GPU at x16 plus three CPU-direct x4 drives would need 28 lanes, which AM5 cannot supply. Boards with three CPU M.2 sockets take the extra lanes from the GPU, and another AM5 CPU cannot change the budget. Threadripper can supply more direct lanes but costs RM5,800 more for the CPU and board alone, for a slot the design does not need. The Tomahawk is also a widely deployed board with far more validated memory kits than the newer PRO X870E-S EVO, and one board across the fleet means one BIOS setup and one spare. Detail: motherboard and PCIe lanes.

Network and motherboard choice#

Use onboard 5GbE. A healthy link moves roughly 500–550 MB/s after overhead, but one NAS hard drive fills a cache at roughly 190–220 MB/s, and prefetch hides fills from jobs. Buying 10GbE before measuring the NAS would accelerate a link that is not the bottleneck.

If sustained cache-fill traffic ever justifies 10GbE, keep the board and add a replaceable NIC (a RM499 TP-Link TX401 in the chipset PCI_E3 slot, after a dry-fit against the GPU). The onboard-10GbE boards cost more for integration, not throughput: the RM2,690 ASUS ProArt X870E-Creator drops the GPU to x8 when its second M.2 is used, the RM4,120 Dark Hero costs RM1,922 more than the Tomahawk-plus-NIC path, and the RM4,999 Godlike RM2,801 more. Detail: motherboard and PCIe lanes, dual-worker storage and NFS.

CPU and memory#

The Ryzen 7 9850X3D in the promotional package was already sufficient: one ComfyUI job averages about 1.8 cores during a cold run and 1.2 during a warm one. The 9950X was a RM700 upgrade that doubles the core count for LLM and CPU-side work, background prefetch hashing and concurrent services, and it is a high-volume part whose failure modes get found and fixed quickly. It is an operational-margin choice, not a warm-inference requirement.

PC1 and PC2 keep 64 GB (2×32 GB at EXPO 6000). A 64 GB host completed every measured workload, and the production replay put 128 GB's benefit at about 0.3 s per run for RM12,598 more across the pair, so that upgrade is closed for them. PC3 was bought with 128 GB (2×64 GB at EXPO 5600) and is the natural WAN 2.2 node. Every node uses two DIMMs only: AMD rates the 9950X at DDR5-5600 with two DIMMs but 3600 with four, so a future upgrade replaces the kit rather than adding sticks. Monitor RSS, swap, major faults and page-cache churn in production. Detail: node record.

Power supply, site power, and UPS#

Every node uses a 1200 W ATX 3.1 PSU with its own supplied 12V-2x6 GPU cable. The worst case, a 575–600 W GPU plus a fully loaded 9950X at its 230 W stock PPT and the rest of the system, is about 875–930 W: up to 93% of a 1000 W unit against 78% of a 1200 W one. The vendors' 1000 W recommendation is a floor; 1200 W keeps a 24/7 PSU cooler and slower-ageing. Larger units add no throughput or protection. The package's Antec GSK1000 was rejected for its ripple and short ride-through.

The Seasonic Vertex PX-1200 is the fleet standard (decided 2026-09-25): PC1, PC2 and PC3 all have it, which gives one spare pool and one native 16-to-16 cable type. The Corsair HX1200i is the fallback if the Vertex is unavailable or more than about RM500 dearer. The 91-incident 16-pin melt record shows no PSU-brand effect; the robust signals are GPU-bundled adapters, especially MSI's, and the need for monitoring. So: use only the PSU's own cable, never the GPU's adapter; seat it once; never mix modular cables across PSUs; and fit the Pi Pico connector thermal watchdog (designed and bench-simulated, not yet built). Detail: PSU decision, melt evidence.

A premium PSU does not condition the building supply or bridge an outage. The immediate site priorities are one dedicated protected circuit per node, verified earthing and RCD/RCBO operation, a correctly selected switchboard surge-protection device, and branch power-quality logging, all installed by a Malaysian licensed electrician. Do not buy UPS units merely to "regulate voltage": record at least 30 days of voltage, sag, swell and outage data first. If interruption cost or customer commitments justify it, start around 2200 VA with at least 1800 W continuous pure-sine output per node, protect the network path too, and test automatic job draining and shutdown. Detail: PSU and site power.

Electricity (2026-09-27). The site stays on single-phase supply and the domestic tariff; no solar for now.

Detail, including the verified bill formula: electricity cost and supply headroom.

Shared model origin and network#

Each node serves active models from its local NVMe. One shared HDD-backed NFS origin holds the canonical model repository and stages missing models to a node before a job is admitted, hash-verified and published atomically. Normal jobs never read model weights over NFS. The balanced start is a 2×4 TB IronWolf CMR mirror, moving to 2×8 TB only if the measured catalogue plus a twelve-month forecast exceeds about 3.2 TB. Do not add an HDD to each node. Background prefetch rules (idle I/O priority, no page-cache eviction of the running job) are in the node record. Full design: dual-worker storage and NFS.

Build configuration#

Production acceptance#

The full list is in the node record. In short, for each node:

  1. Topology. Every 9100 Pro at 32 GT/s x4 on a CPU path (M2_2 set to x4 in the BIOS); the GPU at Gen5 x16 under load; power limit ≥575 W. See motherboard and PCIe lanes.
  2. Direct filesystem. Models on plain XFS or ext4 on M2_1, mounted into the container, not a buffered loop file; confirm the Samsung namespace counters rise by about the resolved bytes on each cold read. Process counters alone are insufficient.
  3. Latency. Five cold/warm pairs per workload with the benchmark protocol: every cold sample verified at the process and NVMe layers, every warm sample reading zero model bytes, and MiniMax H3 cold-minus-warm ≤6.5 s and Krea 2 ≤5 s. Try read_ahead_kb and max_sectors_kb settings and keep the winner.
  4. Memory. EXPO with UCLK = MEMCLK and a multi-hour stress test: 6000 on PC1/PC2 (5600 fallback), 5600 on PC3.
  5. Electrical and connector. Record idle, ComfyUI and combined CPU/GPU stress input power at the branch meter; check for voltage sag, nuisance trips and OCP trips. Commission the watchdog and IR-check both 16-pin ends at full load. After initial thermal cycling, isolate mains and inspect the 12V-2x6 connection for full insertion, discolouration or deformation; do not reseat a healthy connector without cause.
  6. Thermals. Read at least 200 GB continuously while logging drive temperatures and link states; correct any configuration that passes briefly and then throttles.
  7. NAS. 5GbE full duplex; iperf3 about 4.5 Gbit/s per node to a RAM- or SSD-backed endpoint; a real model-set fill from the HDD origin, alone and with nodes filling together, recording whether the disk, NAS uplink or node link is the bottleneck; NAS loss must leave cached jobs running while misses wait or fail cleanly.

Open items#

Assumptions and boundaries#

Where to go deeper#