RTX 5090 nodes: RAM, cold loads and cost-performance#

Status: Decision record for the three ordered nodes; PC1/PC2 decisions settled by the founder on 2026-09-24, PC3 ordered on 2026-09-25 Date: 2026-09-24; PC3 and the PSU standard updated 2026-09-26 Scope: The RTX 5090 ComfyUI nodes, before assembly: PC1 (GIGABYTE RTX 5090 GAMING OC, IdealTech invoice J0926118, 2026-09-21), PC2 (Zotac RTX 5090, IdealTech invoice J0926134, 2026-09-23) and PC3 (Zotac RTX 5090 Solid OC, C-Zone, deposit 2026-09-24, spec confirmed 2026-09-25; see PC3). The "already built first worker" in the older IdealTech decision and overview is PC1: its invoice matches that parts list except the GPU brand. Goal: Best price-performance. The target is a cold-minus-warm penalty of about 5 seconds; money not spent on marginal latency goes to the next RTX 5090 instead. Supersedes: Where overview.md or the IdealTech decision describe other configurations (O11 Vision, HydroShift, G.Skill, RAID or three-drive layouts), this doc is the record.

Decisions#

Component Decision Why, in one line
RAM PC1/PC2: keep 64 GB (T-Create Expert 2×32 GB); 128 GB upgrade closed for them. PC3: 128 GB (Kingston 2×64 GB 5600) 128 GB saves about 0.3 s per run for RM12,598; the money is better spent on throughput. PC3 was bought with 128 GB (see PC3)
Memory speed EXPO 6000, with 5600 as the stress-test fallback ComfyUI does not notice memory speed; LLM offload would; 5600 is no more in-spec
Model storage Samsung 9100 Pro 2 TB, own filesystem, no RAID The loader caps at about 4.6 GiB/s; 2 TB holds the production working set
PC1 second drive Second 9100 Pro for OS, Docker, outputs and future LLM weights Price hedge and TLC endurance, not speed
PC2 OS drive Buy a Samsung 9100 Pro 1 TB later (RM1,399 on the 22 Sep list), not the package's Gen4 Kioxia PC2 ships with one drive; see Open items for the interim layout
Motherboard Keep MAG X870E Tomahawk Max Widely used and easier to maintain; two CPU-connected Gen5 M.2 slots
CPU Keep Ryzen 9 9950X For future LLM and CPU work, and part popularity; ComfyUI uses about two cores
PSU Seasonic Vertex PX-1200 on every node, the fleet standard since 2026-09-25 (PC1, PC2, PC3) Worst case 875–930 W, so 1200 W for 24/7 margin; one PSU and cable type fleet-wide (PSU doc)
GPU One GPU per node; power limit at least 575 W 450 W was about 12% slower warm; 575 and 600 W were equal
Case, cooler Fractal North XL, Arctic Liquid Freezer III Pro 360 Free swaps; the North XL fits the 342 mm GIGABYTE card and cools Gen5 drives better

The biggest latency lever is free: route each model family to a fixed node so its model stays loaded. In the production replay, 68–75% of runs needed no model load at all.

Throughput over marginal latency. Money not spent on memory upgrades was earmarked for another RTX 5090 node. That node, PC3, was ordered on 2026-09-25, ahead of the 30-day production gate in the owned-GPU strategy; see PC3. A fourth node remains subject to that gate: the nodes serving production for at least 30 days, with demand, spill and VRAM evidence.

Ordered configuration (from the invoices)#

Channels, package mechanics and the price gates for the next node are in procurement.

IdealTech sells a promotional package, whose price appears on the GPU line of the invoice. Each upgrade over a package part adds the difference between the two parts' prices on the IdealTech list of 2026-09-22; like-for-like swaps are free. The package rose by RM1,000 on 2026-09-22, from RM29,999 to RM30,999. The package's base parts are:

Line PC1, 2026-09-21 PC2, 2026-09-23 Check against the price list
Package (GPU line) RM29,999, GIGABYTE GAMING OC RM30,999, Zotac Package rose RM1,000 on 22 Sep
Ryzen 9 9950X +RM700 +RM700 Over the 9850X3D
MAG X870E Tomahawk Max WiFi +RM450 +RM450 RM1,699 − RM1,249
Samsung 9100 Pro 2 TB (replaces the Kioxia) +RM1,950 +RM1,950 RM2,849 − RM899
Second Samsung 9100 Pro 2 TB +RM2,749 none Full extra drive, RM100 under list
Seasonic Vertex PX-1200 +RM900 +RM900 RM1,700 − RM799
T-Create Expert 2×32 GB DDR5-6000 CL34 free free Same RM4,999 as the G.Skill kit
Arctic Liquid Freezer III Pro 360 free free Replaces the HydroShift 360
Fractal Design North XL TG free free Replaces the O11 Vision
Invoice total RM36,748 RM34,999 RM71,747 for both

Options considered and closed on 2026-09-24, per node:

Change Price effect Outcome
Seasonic PX-1200 → included Antec GSK1000 −RM900 Declined; the Seasonic stays
T-Create 2×32 GB → Kingston Fury Beast 2×64 GB 5600 CL36, 1.25 V (KF556C36BBEAK2-128) +RM6,299 (back-order) Declined. Sold out; supplier price RM11,298 (the RM8,999 list was the last batch; retail RM11,405). Technically sound: EXPO 5600 is AMD's rating (datasheet)
T-Create 2×32 GB → Corsair Vengeance RGB 2×64 GB 6400 CL42 (CMH128GX5M2B6400C42) +RM4,000 Rejected: Intel XMP-only, "optimized for Intel 700/800 Series" (Corsair); 6400 is above what AM5 usually runs 1:1; not validated on this board

The 128 GB upgrade was closed after a step back from the individual quotes:

Reopen 128 GB only if, after delivery, WAN at 64 GB re-reads experts on warm runs and the fp8 expert is unacceptable, and then only for the WAN node. The preferred kit is then the T-Create Expert 2×64 GB 6000 CL40 at 1.25 V (CTCED5128G6000HC40FDC01), which is validated on the Tomahawk Max.

Decision details#

RAM: 64 GB, two DIMMs, EXPO 6000#

Capacity. Every measured workload fits in 64 GB. The production replay below shows 128 GB would move about 11% of runs from an NVMe load to a RAM hit, saving about 2–3.5 s each: roughly 0.3 s per run on average. The one unmeasured risk is WAN 2.2 (see Open items).

Two DIMMs only. AMD's Ryzen 9 9950X specification rates two DIMMs at DDR5-5600 and four DIMMs at only DDR5-3600, and four DIMMs at 6000 work only on hand-tuned combinations with long memory training (Level1Techs). A future upgrade therefore replaces the kit rather than adding sticks.

Speed: EXPO 6000. The ordered kit (CTCED564G6000HC34BDC01) defaults to 4800 and runs EXPO 6000 at 34-44-44-84 and 1.35 V. DDR5-6000 is above AMD's 5600 rating, but:

Storage: 9100 Pro, separate filesystems, no RAID#

Past about 5 GiB/s the drive stops mattering, because ComfyUI's loader never read faster than 4.3–4.6 GiB/s. RAID 0 therefore showed no application gain. The exact drive being bought, a Samsung 9100 Pro 2 TB, measured 6.2 s cold minus warm on MiniMax and 4.8 s on Krea (machine 148771). The NVMe replay shows 2 TB holds the working set; 1 TB would save RM1,450 per node for 0.2–0.6% more NAS fetches.

PC1's second 9100 Pro adds no ComfyUI speed. It buys:

Streaming LLM weights from disk would use its raw bandwidth.

Slot Drive Use
M2_1, CPU Gen5 x4 9100 Pro 2 TB Model cache only: XFS or ext4, noatime, 15–20% free. Disposable, rebuilt from the NAS.
M2_2, CPU Gen5 x4 PC1: 9100 Pro 2 TB; PC2: 1 TB drive, bought later Linux, Docker, outputs, staging; later an LLM-weights partition on PC1.

If the model cache outgrows 2 TB, split model families across drives rather than striping. Further drives go in the chipset Gen4 slots, which are fast enough for the loader.

Motherboard: Tomahawk Max#

The included PRO X870E-S EVO is a new board. As of 2026-09-24, G.Skill's validation list does not include it, and TeamGroup has validated only two kits on it, against nine on the Tomahawk Max. The Tomahawk line is widely deployed, so BIOS issues and failure modes are known and later debugging is easier. It also has two independent CPU-connected Gen5 x4 M.2 slots while the GPU keeps x16. The lane budget, slot layout, the M2_2/USB4 BIOS setting and the rejected alternatives are in motherboard and PCIe lanes.

CPU: Ryzen 9 9950X#

ComfyUI does not need it: one job averaged 1.8 cores during MiniMax cold runs and 1.2 during warm runs. The 9950X is for later work:

All of this costs RM700 per node.

PSU: 1200 W class#

Zotac specifies 575 W and a 1000 W system supply. The GIGABYTE GAMING OC ships at a 600 W default (TechPowerUp), and GIGABYTE also recommends 1000 W. The worst case, with the 9950X at its 230 W stock PPT, is about 875–905 W with the Zotac and 900–930 W with the GIGABYTE. That is up to 93% of a 1000 W unit and 78% of a 1200 W unit. For 24/7 service the 1200 W class keeps the PSU cooler and slower-ageing, so the Antec GSK1000 revert was declined. PC1 and PC2 keep the Seasonic: IdealTech could not swap it for the HX1200i. On 2026-09-24 the HX1200i was preferred for future nodes for its Linux telemetry; on 2026-09-25 the Vertex PX-1200 became the fleet standard instead, and PC3 has one (PSU doc). Use only the PSU's own supplied 12V-2x6 cable, and fit the connector thermal watchdog.

One GPU per node#

It is also a bet that models will need less VRAM. That is already visible for images but not yet for video. Production runs from May to August 2026, by the largest model file loaded:

Model Size Share of runs, May → Aug
Krea 2 Turbo, fp8 13.1 GB 0% → 18.1%
Flux 2 Klein 9B, fp8 9.4 GB 0.1% → 5.2%
Z-Image Turbo, bf16 12.3 GB 27.9% → 15.4%
WAN 2.2 image-to-video 49.9 GB set 26.9% → 22.3% (range 17–27%)

Runs with a 16–24 GB model set rose from 43% to 52%. WAN 2.2 still drives most of the 28–36% of runs that load more than 32 GB. The production WAN workflow loads its low-noise expert in fp16 (28.6 GB) while the high-noise expert is already fp8 (14.3 GB). An fp8 low-noise expert would cut the set to about 36 GB; that is an unevaluated quality decision. Track the share of runs over 32 GB before buying a third node.

Evidence#

Why RAM matters when models end up in VRAM#

There is no direct disk-to-GPU path: NVMe ──► RAM ──► PCIe ──► VRAM. RAM plays three roles:

  1. Staging. Every byte passes through it.
  2. Overflow. A model set larger than 32 GB of VRAM keeps its remainder in RAM. MiniMax H3 (42.5 GB) held 46 GB of process RAM throughout a run. WAN 2.2's two 14B experts (42.9 GB) take turns in VRAM, so one waits in RAM.
  3. Cache. Linux keeps recently read files in unused RAM.

Roles 1 and 2 are capacity: the job must fit. Only role 3 buys speed with more RAM.

Benchmarks#

The RTX 5090 cohort in ../vast-*-cache-evicted-v1.json covered 36 MiniMax and 30 Krea hosts as of 2026-09-24. rtx5090_cold_penalty_table.py tabulates every host.

Validity. Every schema-v4 cold sample read its full model set from the block layer after eviction: 42.4–46.7 GB for MiniMax (151 samples) and 17.9–20.7 GB for Krea (130 samples). Of 281 warm samples, 273 read nothing and 8 read at most 2.2 MB. Excluded:

"Cold" here is the worst case: a new process with files evicted from RAM.

Why disk speed varied from 0.7 to 11.5 GiB/s.

MiniMax on desktop hosts, by drive read speed:

Drive read speed Hosts Cold − warm
under 3 GiB/s 3 7.8–28.7 s
3–5 GiB/s 4 7.3–13.5 s
5–8 GiB/s 6 5.6–21.3 s (high end: excluded 150656)
8 GiB/s and up 4 5.4–8.8 s

A good Gen4 drive is close to Gen5: the 990 Pro hosts reached 7.8 and 9.2 s even on 8-core, 450 W machines.

The floor. The older *-resident-paired.json runs kept the page cache, so their first cold sample started a fresh process with files still in RAM:

Host Files in RAM Files on disk
Ryzen 9950X, 123 GB 3.5 s 7.1 s
RTX 5090, 10.7 GB/s disk 3.3 s 5.6 s
RTX Pro 6000 WS, two hosts 4.3 / 4.6 s 8.5 / 14.9 s

About 3.5 s is process start and model construction; the NVMe adds 2–4 s. Only RAM can remove that 2–4 s.

GPU. Warm MiniMax took 60.0–65.4 s at 575–600 W versus 69.0–71.3 s at 450 W. GPU PCIe width did not matter: three of the best hosts ran at Gen5 x8 and measured 5.6, 7.1 and 7.3 s.

Production replay#

production_ram_nvme_replay.py replays every completed July and August 2026 run in order (15,010 and 14,162 runs), using the model files each run resolved and their sizes (../../docs/cost-analysis/data/months/*/model_stream.json). The model:

It ignores concurrency and spill to Modal, so the "already loaded" share is optimistic against production's current 21% residency hits. The comparisons between sizes are what it measures. Values are July / August.

July August
Model GB per run, p50 / p90 / p99 22.4 / 52.9 / 53.2 21.0 / 52.9 / 67.0
Runs over 56 GB 0.2% 1.4% (LTX and Flux 2 Dev)
WAN 2.2 runs (49.9 GB set) 3,975 (26%) 3,162 (22%)
RAM per node Already loaded Files in RAM Load from NVMe NVMe GB per run
64 GB 75.0% / 68.2% 8.7% / 10.9% 16.3% / 20.8% 3.6 / 5.5
96 GB 75.0% / 68.2% 17.1% / 18.7% 7.9% / 13.1% 1.7 / 3.5
128 GB 75.0% / 68.2% 19.6% / 22.3% 5.4% / 9.5% 1.1 / 2.4
Usable model cache per node Runs needing a NAS fetch GB fetched per month
800 GB (1 TB drive) 2.37% / 3.0% 2,996 / 3,121
1,600 GB (2 TB drive) 2.15% / 2.41% 2,542 / 2,571

With background prefetch, the next job's model set fits in RAM beside the running job for 57% / 58% of model switches at 64 GB and 100% / 98% at 128 GB. That extra coverage at 128 GB is worth about 1–1.3 hours of total latency per month. That is an upper bound, since prefetch needs a queued next job, which is often missing at current traffic. It does not change the RAM decision.

Third node: two versus three (2026-09-24)#

The same script now takes --nodes and adds a capacity replay. Every run that held a GPU (16,047 in July and 15,617 in August) arrives at its created_at and holds a node for its recorded Modal run_time. Nodes serve in FIFO order, and a run spills to Modal if it would wait longer than the limit. The model ignores model affinity and switch cost, and it assumes every run fits a 32 GB card. ×1.3 is a sensitivity case for a 5090 slower than the Modal card that ran the job.

Share of GPU-hours served locally, July / August:

Case 1 node 2 nodes 3 nodes Third node adds
No waiting, ×1.0 59.5 / 67.0% 78.8 / 85.2% 88.1 / 91.6% +9.3 / +6.4 pts
No waiting, ×1.3 55.9 / 62.7% 75.9 / 82.8% 86.2 / 90.2% +10.3 / +7.4 pts
Wait up to 60 s, ×1.0 73.6 / 80.7% 89.3 / 92.3% 95.5 / 96.0% +6.2 / +3.7 pts
Wait up to 60 s, ×1.3 68.3 / 76.0% 85.4 / 90.1% 92.8 / 94.4% +7.4 / +4.3 pts
Node occupancy, no waiting ×1.0 15.4 / 16.9% 10.2 / 10.7% 7.6 / 7.7%

What the third node does and does not buy at July–August demand:

The case for PC3 is therefore growth headroom, redundancy and model residency, not current Modal spend. Real spill and occupancy data from PC1/PC2, which the 30-day gate asks for, replaces this estimate.

PC3 was ordered on 2026-09-25, before that gate, so this replay is now the planning estimate for a three-node fleet rather than a purchase test.

Open items#

None of these block assembly.

  1. PC2's interim drive layout. PC2 ships with one 9100 Pro 2 TB, and a 1 TB OS drive will be bought later (a Samsung 9100 Pro 1 TB, which goes in M2_2).
    • Until then, give the 2 TB drive separate partitions: a small OS partition and a model-cache partition.
    • When the 1 TB drive arrives, reprovision the OS onto it and grow the model-cache partition over the whole 2 TB drive.
    • The cache is disposable, so this needs no data migration.
  2. Check WAN 2.2 at 64 GB, on PC1 or PC2 (PC3 has 128 GB, so it cannot answer this). Run the production WAN workflow (umt5 fp8, high-noise 14B fp8, low-noise 14B fp16) cold once and then warm three times. Record peak RSS, minimum available memory, swap, and NVMe bytes read during the warm runs.
    • Warm runs should read close to 0 GB.
    • If they re-read an expert (14–29 GB), evaluate the fp8 low-noise expert first.
  3. Ask IdealTech whether they have built the Tomahawk Max with the T-Create 2×32 GB kit at EXPO 6000. The kit is not on TeamGroup's validation list for this board.

PC3 ordered (2026-09-25)#

PC3 was bought from C-Zone (czone.my) on price. IdealTech's package rose RM1,000 on Tuesday 2026-09-22 and a further RM2,000 on Thursday 2026-09-24. C-Zone's Tuesday quotation was cheaper for the same build: at IdealTech's price it would have covered a whole extra 9100 Pro 2 TB. A deposit on Thursday locked that quotation. On Friday 2026-09-25 the spec was confirmed with the RAM upgraded to 128 GB and the PSU kept as the Vertex PX-1200, the same as PC1/PC2. The final quote is dated that day (../pc3.png); the full week is in the procurement timeline. PC3 keeps the PC1/PC2 platform, so one BIOS setup, OS image, acceptance suite and spare-parts set covers the CPU, board, drives, cooler, case and PSU.

Part Ordered item RM Against PC1/PC2
CPU AMD Ryzen 9 9950X 2,549 Same
Motherboard MSI MAG X870E Tomahawk Max WiFi 1,600 Same
GPU Zotac RTX 5090 32 GB Solid OC 22,799 Same card as PC2 (575 W)
RAM Kingston Fury Beast 128 GB (2×64 GB) DDR5-5600 CL36 (KF556C36BBEAK2-128) 9,899 Differs: PC1/PC2 have 64 GB T-Create 6000
OS drive, M2_2 Samsung 9100 Pro 1 TB 1,199 Same as PC2's planned OS drive
Model drive, M2_1 Samsung 9100 Pro 2 TB 2,399 Same
Cooler Arctic Liquid Freezer III Pro 360 499 Same
Case Fractal Design North XL Dark TG 780 Same
PSU Seasonic Vertex PX-1200 ATX 3.1 1,600 Same (fleet standard)
Total Sum of the quote lines 43,324 PC1 RM36,748, PC2 RM34,999

What PC3 changes:

Buying outside the package still requires:

Acceptance targets for the delivered nodes#

Background prefetch rules#

Nodes can fetch the next job's missing files and warm them into RAM while the current job runs. No hardware change is needed.

  1. Stage on the model drive. Download into a staging directory on the M2_1 filesystem, verify the hash, then rename. The atomic rename means ComfyUI never sees a partial file.
  2. The loader has priority. Run prefetch at idle I/O priority (ionice -c3 or a low cgroup io.weight). NAS fills run at about 200 MB/s from the HDD mirror.
  3. Protect the page cache. Write prefetched files with O_DIRECT, or call posix_fadvise(POSIX_FADV_DONTNEED) after writing, so the running job's model pages are not evicted.
  4. Warm RAM only when it fits. Pre-read the next set only when it fits beside the running set and MemAvailable confirms the headroom.

Prefetch hides NAS and network speed, which further weakens the case for 10GbE.

Reproducing#

cd comfyicu/orchestrator/benchmarks/analysis
python3 rtx5090_cold_penalty_table.py             # benchmark cohort + page-cache control
python3 production_ram_nvme_replay.py 2026-07 --nodes 2,3   # capacity, cache and prefetch replay (also 2026-08)

The replay reads gitignored runs.csv files; regenerate them with npx tsx fetch_runs.ts --months … from orchestrator/ (see ../../docs/cost-analysis/data/months/README.md).