RTX 5090 nodes: RAM, cold loads and cost-performance#
Status: Decision record for the three ordered nodes; PC1/PC2 decisions settled by the founder on 2026-09-24, PC3 ordered on 2026-09-25 Date: 2026-09-24; PC3 and the PSU standard updated 2026-09-26 Scope: The RTX 5090 ComfyUI nodes, before assembly: PC1 (GIGABYTE RTX 5090 GAMING OC, IdealTech invoice J0926118, 2026-09-21), PC2 (Zotac RTX 5090, IdealTech invoice J0926134, 2026-09-23) and PC3 (Zotac RTX 5090 Solid OC, C-Zone, deposit 2026-09-24, spec confirmed 2026-09-25; see PC3). The "already built first worker" in the older IdealTech decision and overview is PC1: its invoice matches that parts list except the GPU brand. Goal: Best price-performance. The target is a cold-minus-warm penalty of about 5 seconds; money not spent on marginal latency goes to the next RTX 5090 instead. Supersedes: Where overview.md or the IdealTech decision describe other configurations (O11 Vision, HydroShift, G.Skill, RAID or three-drive layouts), this doc is the record.
Decisions#
| Component | Decision | Why, in one line |
|---|---|---|
| RAM | PC1/PC2: keep 64 GB (T-Create Expert 2×32 GB); 128 GB upgrade closed for them. PC3: 128 GB (Kingston 2×64 GB 5600) | 128 GB saves about 0.3 s per run for RM12,598; the money is better spent on throughput. PC3 was bought with 128 GB (see PC3) |
| Memory speed | EXPO 6000, with 5600 as the stress-test fallback | ComfyUI does not notice memory speed; LLM offload would; 5600 is no more in-spec |
| Model storage | Samsung 9100 Pro 2 TB, own filesystem, no RAID | The loader caps at about 4.6 GiB/s; 2 TB holds the production working set |
| PC1 second drive | Second 9100 Pro for OS, Docker, outputs and future LLM weights | Price hedge and TLC endurance, not speed |
| PC2 OS drive | Buy a Samsung 9100 Pro 1 TB later (RM1,399 on the 22 Sep list), not the package's Gen4 Kioxia | PC2 ships with one drive; see Open items for the interim layout |
| Motherboard | Keep MAG X870E Tomahawk Max | Widely used and easier to maintain; two CPU-connected Gen5 M.2 slots |
| CPU | Keep Ryzen 9 9950X | For future LLM and CPU work, and part popularity; ComfyUI uses about two cores |
| PSU | Seasonic Vertex PX-1200 on every node, the fleet standard since 2026-09-25 (PC1, PC2, PC3) | Worst case 875–930 W, so 1200 W for 24/7 margin; one PSU and cable type fleet-wide (PSU doc) |
| GPU | One GPU per node; power limit at least 575 W | 450 W was about 12% slower warm; 575 and 600 W were equal |
| Case, cooler | Fractal North XL, Arctic Liquid Freezer III Pro 360 | Free swaps; the North XL fits the 342 mm GIGABYTE card and cools Gen5 drives better |
The biggest latency lever is free: route each model family to a fixed node so its model stays loaded. In the production replay, 68–75% of runs needed no model load at all.
Throughput over marginal latency. Money not spent on memory upgrades was earmarked for another RTX 5090 node. That node, PC3, was ordered on 2026-09-25, ahead of the 30-day production gate in the owned-GPU strategy; see PC3. A fourth node remains subject to that gate: the nodes serving production for at least 30 days, with demand, spill and VRAM evidence.
Ordered configuration (from the invoices)#
Channels, package mechanics and the price gates for the next node are in procurement.
IdealTech sells a promotional package, whose price appears on the GPU line of the invoice. Each upgrade over a package part adds the difference between the two parts' prices on the IdealTech list of 2026-09-22; like-for-like swaps are free. The package rose by RM1,000 on 2026-09-22, from RM29,999 to RM30,999. The package's base parts are:
- Ryzen 7 9850X3D
- MSI PRO X870E-S EVO WiFi
- G.Skill 64 GB DDR5-6000
- Kioxia Exceria Basic 1 TB (Gen4 QLC, RM899)
- Antec GSK1000
- Lian Li HydroShift 360 and O11 Vision
- the RTX 5090
| Line | PC1, 2026-09-21 | PC2, 2026-09-23 | Check against the price list |
|---|---|---|---|
| Package (GPU line) | RM29,999, GIGABYTE GAMING OC | RM30,999, Zotac | Package rose RM1,000 on 22 Sep |
| Ryzen 9 9950X | +RM700 | +RM700 | Over the 9850X3D |
| MAG X870E Tomahawk Max WiFi | +RM450 | +RM450 | RM1,699 − RM1,249 |
| Samsung 9100 Pro 2 TB (replaces the Kioxia) | +RM1,950 | +RM1,950 | RM2,849 − RM899 |
| Second Samsung 9100 Pro 2 TB | +RM2,749 | none | Full extra drive, RM100 under list |
| Seasonic Vertex PX-1200 | +RM900 | +RM900 | RM1,700 − RM799 |
| T-Create Expert 2×32 GB DDR5-6000 CL34 | free | free | Same RM4,999 as the G.Skill kit |
| Arctic Liquid Freezer III Pro 360 | free | free | Replaces the HydroShift 360 |
| Fractal Design North XL TG | free | free | Replaces the O11 Vision |
| Invoice total | RM36,748 | RM34,999 | RM71,747 for both |
Options considered and closed on 2026-09-24, per node:
| Change | Price effect | Outcome |
|---|---|---|
| Seasonic PX-1200 → included Antec GSK1000 | −RM900 | Declined; the Seasonic stays |
| T-Create 2×32 GB → Kingston Fury Beast 2×64 GB 5600 CL36, 1.25 V (KF556C36BBEAK2-128) | +RM6,299 (back-order) | Declined. Sold out; supplier price RM11,298 (the RM8,999 list was the last batch; retail RM11,405). Technically sound: EXPO 5600 is AMD's rating (datasheet) |
| T-Create 2×32 GB → Corsair Vengeance RGB 2×64 GB 6400 CL42 (CMH128GX5M2B6400C42) | +RM4,000 | Rejected: Intel XMP-only, "optimized for Intel 700/800 Series" (Corsair); 6400 is above what AM5 usually runs 1:1; not validated on this board |
The 128 GB upgrade was closed after a step back from the individual quotes:
- The latency benefit is small. About 0.3 s per run on average, 1–1.3 hours of total latency per month even with prefetch, for RM12,598.
- Nothing is locked in. Swapping RAM is a ten-minute job at any time; only price argued for buying now.
- The WAN 2.2 risk has cheaper fixes first: an fp8 low-noise expert and model-family routing.
- The LLM case is weak. Offloaded LLM inference is too slow to sell, so paid LLM work would mostly fit in 32 GB of VRAM, where system RAM hardly matters.
Reopen 128 GB only if, after delivery, WAN at 64 GB re-reads experts on warm runs and the fp8 expert is unacceptable, and then only for the WAN node. The preferred kit is then the T-Create Expert 2×64 GB 6000 CL40 at 1.25 V (CTCED5128G6000HC40FDC01), which is validated on the Tomahawk Max.
Decision details#
RAM: 64 GB, two DIMMs, EXPO 6000#
Capacity. Every measured workload fits in 64 GB. The production replay below shows 128 GB would move about 11% of runs from an NVMe load to a RAM hit, saving about 2–3.5 s each: roughly 0.3 s per run on average. The one unmeasured risk is WAN 2.2 (see Open items).
Two DIMMs only. AMD's Ryzen 9 9950X specification rates two DIMMs at DDR5-5600 and four DIMMs at only DDR5-3600, and four DIMMs at 6000 work only on hand-tuned combinations with long memory training (Level1Techs). A future upgrade therefore replaces the kit rather than adding sticks.
Speed: EXPO 6000. The ordered kit (CTCED564G6000HC34BDC01) defaults to 4800 and runs EXPO 6000 at 34-44-44-84 and 1.35 V. DDR5-6000 is above AMD's 5600 rating, but:
- 5600 is no more in-spec. It needs EXPO voltages too, and AMD's warranty terms treat any EXPO use as overclocking (Windows Central). Only the 4800 default is fully in spec.
- Two DIMMs at EXPO 6000 is the most common AM5 setup, so it is the most community-tested.
- ComfyUI does not notice memory speed. Across Ryzen hosts, host memory copy speed ranged 13.7–20.8 GiB/s while warm MiniMax stayed at 60–65 s. CPU-offloaded LLM inference is memory-bandwidth-bound, and 6000 gives it about 7% more than 5600.
Storage: 9100 Pro, separate filesystems, no RAID#
Past about 5 GiB/s the drive stops mattering, because ComfyUI's loader never read faster than 4.3–4.6 GiB/s. RAID 0 therefore showed no application gain. The exact drive being bought, a Samsung 9100 Pro 2 TB, measured 6.2 s cold minus warm on MiniMax and 4.8 s on Krea (machine 148771). The NVMe replay shows 2 TB holds the working set; 1 TB would save RM1,450 per node for 0.2–0.6% more NAS fetches.
PC1's second 9100 Pro adds no ComfyUI speed. It buys:
- capacity at today's price, as NVMe prices keep rising;
- TLC endurance for OS write churn (container pulls, outputs, logs);
- a single spare part that covers either slot.
Streaming LLM weights from disk would use its raw bandwidth.
| Slot | Drive | Use |
|---|---|---|
| M2_1, CPU Gen5 x4 | 9100 Pro 2 TB | Model cache only: XFS or ext4, noatime, 15–20% free. Disposable, rebuilt from the NAS. |
| M2_2, CPU Gen5 x4 | PC1: 9100 Pro 2 TB; PC2: 1 TB drive, bought later | Linux, Docker, outputs, staging; later an LLM-weights partition on PC1. |
If the model cache outgrows 2 TB, split model families across drives rather than striping. Further drives go in the chipset Gen4 slots, which are fast enough for the loader.
Motherboard: Tomahawk Max#
The included PRO X870E-S EVO is a new board. As of 2026-09-24, G.Skill's validation list does not include it, and TeamGroup has validated only two kits on it, against nine on the Tomahawk Max. The Tomahawk line is widely deployed, so BIOS issues and failure modes are known and later debugging is easier. It also has two independent CPU-connected Gen5 x4 M.2 slots while the GPU keeps x16. The lane budget, slot layout, the M2_2/USB4 BIOS setting and the rejected alternatives are in motherboard and PCIe lanes.
CPU: Ryzen 9 9950X#
ComfyUI does not need it: one job averaged 1.8 cores during MiniMax cold runs and 1.2 during warm runs. The 9950X is for later work:
- LLMs. Once weights (for example MoE experts) are offloaded to system RAM, token generation is bound by dual-channel DDR5 bandwidth of roughly 80–90 GB/s, regardless of core count (llama.cpp MoE offload guide). Extra cores help prompt processing and concurrent jobs. RAM capacity is the tighter LLM limit.
- Background work. Prefetch hashing runs without touching the ComfyUI job.
- Popularity. High-volume parts get failure modes found and fixed sooner. The 2023 Ryzen 7000 EXPO SoC-voltage burnouts were fixed with a firmware cap within weeks (Tom's Hardware).
- Consistency. Every node (PC1, PC2, PC3) runs the same CPU.
All of this costs RM700 per node.
PSU: 1200 W class#
Zotac specifies 575 W and a 1000 W system supply. The GIGABYTE GAMING OC ships at a 600 W default (TechPowerUp), and GIGABYTE also recommends 1000 W. The worst case, with the 9950X at its 230 W stock PPT, is about 875–905 W with the Zotac and 900–930 W with the GIGABYTE. That is up to 93% of a 1000 W unit and 78% of a 1200 W unit. For 24/7 service the 1200 W class keeps the PSU cooler and slower-ageing, so the Antec GSK1000 revert was declined. PC1 and PC2 keep the Seasonic: IdealTech could not swap it for the HX1200i. On 2026-09-24 the HX1200i was preferred for future nodes for its Linux telemetry; on 2026-09-25 the Vertex PX-1200 became the fleet standard instead, and PC3 has one (PSU doc). Use only the PSU's own supplied 12V-2x6 cable, and fit the connector thermal watchdog.
One GPU per node#
- PCIe lanes are not the reason. Three of the best benchmark hosts ran their GPU at Gen5 x8 with no loss.
- CPU and NVMe sharing are minor.
- The real reasons:
- RAM: two concurrent jobs need up to about 2×53 GB.
- Cooling: two 575 W cards of about 3.5 slots each in one case.
- Failure isolation: one box failing would take out both GPUs.
- Economics: Malaysian retailers sold the RTX 5090 only inside a complete system.
It is also a bet that models will need less VRAM. That is already visible for images but not yet for video. Production runs from May to August 2026, by the largest model file loaded:
| Model | Size | Share of runs, May → Aug |
|---|---|---|
| Krea 2 Turbo, fp8 | 13.1 GB | 0% → 18.1% |
| Flux 2 Klein 9B, fp8 | 9.4 GB | 0.1% → 5.2% |
| Z-Image Turbo, bf16 | 12.3 GB | 27.9% → 15.4% |
| WAN 2.2 image-to-video | 49.9 GB set | 26.9% → 22.3% (range 17–27%) |
Runs with a 16–24 GB model set rose from 43% to 52%. WAN 2.2 still drives most of the 28–36% of runs that load more than 32 GB. The production WAN workflow loads its low-noise expert in fp16 (28.6 GB) while the high-noise expert is already fp8 (14.3 GB). An fp8 low-noise expert would cut the set to about 36 GB; that is an unevaluated quality decision. Track the share of runs over 32 GB before buying a third node.
Evidence#
Why RAM matters when models end up in VRAM#
There is no direct disk-to-GPU path: NVMe ──► RAM ──► PCIe ──► VRAM. RAM plays three roles:
- Staging. Every byte passes through it.
- Overflow. A model set larger than 32 GB of VRAM keeps its remainder in RAM. MiniMax H3 (42.5 GB) held 46 GB of process RAM throughout a run. WAN 2.2's two 14B experts (42.9 GB) take turns in VRAM, so one waits in RAM.
- Cache. Linux keeps recently read files in unused RAM.
Roles 1 and 2 are capacity: the job must fit. Only role 3 buys speed with more RAM.
Benchmarks#
The RTX 5090 cohort in ../vast-*-cache-evicted-v1.json covered 36 MiniMax and 30 Krea hosts as of
2026-09-24. rtx5090_cold_penalty_table.py tabulates every host.
Validity. Every schema-v4 cold sample read its full model set from the block layer after eviction: 42.4–46.7 GB for MiniMax (151 samples) and 17.9–20.7 GB for Krea (130 samples). Of 281 warm samples, 273 read nothing and 8 read at most 2.2 MB. Excluded:
- machine 45827, whose warm run alone takes 142 s;
- machine 132465, whose GPU link was stuck at Gen1;
- machine 150656, with one wild cold sample;
- single-pair runs.
"Cold" here is the worst case: a new process with files evicted from RAM.
Why disk speed varied from 0.7 to 11.5 GiB/s.
- Virtual disks. OpenStack/QEMU VMs read at 0.7–1.1 GiB/s and gave 55–57 s MiniMax penalties.
- Degraded links. A P310 ran at Gen4 x2 and a 990 EVO Plus at Gen3 x2.
- Budget drives. They read at 2–3.5 GiB/s.
- Different host paths for the same drive. The 9100 Pro 4 TB read at 5.3, 5.7 and 7.6 GiB/s on three hosts, depending on overlay or loop layers, read-ahead and request size.
- CPU, not disk. On 64–96-core EPYC hosts, much of the apparent disk variance was CPU: 1,300–2,000 CPU-seconds per cold run against 120–150 on a 9950X.
MiniMax on desktop hosts, by drive read speed:
| Drive read speed | Hosts | Cold − warm |
|---|---|---|
| under 3 GiB/s | 3 | 7.8–28.7 s |
| 3–5 GiB/s | 4 | 7.3–13.5 s |
| 5–8 GiB/s | 6 | 5.6–21.3 s (high end: excluded 150656) |
| 8 GiB/s and up | 4 | 5.4–8.8 s |
A good Gen4 drive is close to Gen5: the 990 Pro hosts reached 7.8 and 9.2 s even on 8-core, 450 W machines.
The floor. The older *-resident-paired.json runs kept the page cache, so their first cold sample
started a fresh process with files still in RAM:
| Host | Files in RAM | Files on disk |
|---|---|---|
| Ryzen 9950X, 123 GB | 3.5 s | 7.1 s |
| RTX 5090, 10.7 GB/s disk | 3.3 s | 5.6 s |
| RTX Pro 6000 WS, two hosts | 4.3 / 4.6 s | 8.5 / 14.9 s |
About 3.5 s is process start and model construction; the NVMe adds 2–4 s. Only RAM can remove that 2–4 s.
GPU. Warm MiniMax took 60.0–65.4 s at 575–600 W versus 69.0–71.3 s at 450 W. GPU PCIe width did not matter: three of the best hosts ran at Gen5 x8 and measured 5.6, 7.1 and 7.3 s.
Production replay#
production_ram_nvme_replay.py replays every completed July and
August 2026 run in order (15,010 and 14,162 runs), using the model files each run resolved and their sizes
(../../docs/cost-analysis/data/months/*/model_stream.json). The model:
- two nodes, each base-model set pinned to one node and balanced by run count;
- RAM as a least-recently-used file cache of RAM minus 12 GB;
- NVMe as a least-recently-used cache, starting empty each month;
- about 7% of file references unsized (almost all LoRAs), counted as 0.5 GB.
It ignores concurrency and spill to Modal, so the "already loaded" share is optimistic against production's current 21% residency hits. The comparisons between sizes are what it measures. Values are July / August.
| July | August | |
|---|---|---|
| Model GB per run, p50 / p90 / p99 | 22.4 / 52.9 / 53.2 | 21.0 / 52.9 / 67.0 |
| Runs over 56 GB | 0.2% | 1.4% (LTX and Flux 2 Dev) |
| WAN 2.2 runs (49.9 GB set) | 3,975 (26%) | 3,162 (22%) |
| RAM per node | Already loaded | Files in RAM | Load from NVMe | NVMe GB per run |
|---|---|---|---|---|
| 64 GB | 75.0% / 68.2% | 8.7% / 10.9% | 16.3% / 20.8% | 3.6 / 5.5 |
| 96 GB | 75.0% / 68.2% | 17.1% / 18.7% | 7.9% / 13.1% | 1.7 / 3.5 |
| 128 GB | 75.0% / 68.2% | 19.6% / 22.3% | 5.4% / 9.5% | 1.1 / 2.4 |
| Usable model cache per node | Runs needing a NAS fetch | GB fetched per month |
|---|---|---|
| 800 GB (1 TB drive) | 2.37% / 3.0% | 2,996 / 3,121 |
| 1,600 GB (2 TB drive) | 2.15% / 2.41% | 2,542 / 2,571 |
With background prefetch, the next job's model set fits in RAM beside the running job for 57% / 58% of model switches at 64 GB and 100% / 98% at 128 GB. That extra coverage at 128 GB is worth about 1–1.3 hours of total latency per month. That is an upper bound, since prefetch needs a queued next job, which is often missing at current traffic. It does not change the RAM decision.
Third node: two versus three (2026-09-24)#
The same script now takes --nodes and adds a capacity replay. Every run that held a GPU (16,047 in July
and 15,617 in August) arrives at its created_at and holds a node for its recorded Modal run_time. Nodes
serve in FIFO order, and a run spills to Modal if it would wait longer than the limit. The model ignores
model affinity and switch cost, and it assumes every run fits a 32 GB card. ×1.3 is a sensitivity case
for a 5090 slower than the Modal card that ran the job.
Share of GPU-hours served locally, July / August:
| Case | 1 node | 2 nodes | 3 nodes | Third node adds |
|---|---|---|---|---|
| No waiting, ×1.0 | 59.5 / 67.0% | 78.8 / 85.2% | 88.1 / 91.6% | +9.3 / +6.4 pts |
| No waiting, ×1.3 | 55.9 / 62.7% | 75.9 / 82.8% | 86.2 / 90.2% | +10.3 / +7.4 pts |
| Wait up to 60 s, ×1.0 | 73.6 / 80.7% | 89.3 / 92.3% | 95.5 / 96.0% | +6.2 / +3.7 pts |
| Wait up to 60 s, ×1.3 | 68.3 / 76.0% | 85.4 / 90.1% | 92.8 / 94.4% | +7.4 / +4.3 pts |
| Node occupancy, no waiting ×1.0 | 15.4 / 16.9% | 10.2 / 10.7% | 7.6 / 7.7% |
What the third node does and does not buy at July–August demand:
- Modal displacement is small. It moves another 4–10% of GPU-hours off Modal. The Modal expense in the dataplane was US$959 (July) and US$1,147 (August), so that is at most about US$40–115 a month, before PC3's own power. Payback on a node costing about RM35,000 takes many years at current demand.
- Throughput is not the constraint. Two nodes run at about 11% occupancy against the strategy doc's 50–60% demand gate. Runs spill only during concurrency peaks.
- Fewer model switches. With base sets spread over three nodes, the previous-run model is already loaded for 82.1% / 74.4% of runs, against 75.0% / 68.2% on two, and August model switches drop from 4,500 to 3,617. Per-node RAM and NVMe results barely change.
- Redundancy. With one node down for maintenance or failure, three nodes still give the two-node numbers above.
The case for PC3 is therefore growth headroom, redundancy and model residency, not current Modal spend. Real spill and occupancy data from PC1/PC2, which the 30-day gate asks for, replaces this estimate.
PC3 was ordered on 2026-09-25, before that gate, so this replay is now the planning estimate for a three-node fleet rather than a purchase test.
Open items#
None of these block assembly.
- PC2's interim drive layout. PC2 ships with one 9100 Pro 2 TB, and a 1 TB OS drive will be bought
later (a Samsung 9100 Pro 1 TB, which goes in M2_2).
- Until then, give the 2 TB drive separate partitions: a small OS partition and a model-cache partition.
- When the 1 TB drive arrives, reprovision the OS onto it and grow the model-cache partition over the whole 2 TB drive.
- The cache is disposable, so this needs no data migration.
- Check WAN 2.2 at 64 GB, on PC1 or PC2 (PC3 has 128 GB, so it cannot answer this). Run the production WAN
workflow (umt5 fp8, high-noise 14B fp8, low-noise 14B fp16) cold once and then warm three times. Record
peak RSS, minimum available memory, swap, and NVMe bytes read during the warm runs.
- Warm runs should read close to 0 GB.
- If they re-read an expert (14–29 GB), evaluate the fp8 low-noise expert first.
- Ask IdealTech whether they have built the Tomahawk Max with the T-Create 2×32 GB kit at EXPO 6000. The kit is not on TeamGroup's validation list for this board.
PC3 ordered (2026-09-25)#
PC3 was bought from C-Zone (czone.my) on price. IdealTech's package rose RM1,000 on Tuesday 2026-09-22 and a
further RM2,000 on Thursday 2026-09-24. C-Zone's Tuesday quotation was cheaper for the same build: at
IdealTech's price it would have covered a whole extra 9100 Pro 2 TB. A deposit on Thursday locked that
quotation. On Friday 2026-09-25 the spec was confirmed with
the RAM upgraded to 128 GB and the PSU kept as the Vertex PX-1200, the same as PC1/PC2. The final quote is
dated that day (../pc3.png); the full week is in the
procurement timeline. PC3 keeps the
PC1/PC2 platform, so one BIOS setup, OS image, acceptance suite and spare-parts set covers the CPU, board,
drives, cooler, case and PSU.
| Part | Ordered item | RM | Against PC1/PC2 |
|---|---|---|---|
| CPU | AMD Ryzen 9 9950X | 2,549 | Same |
| Motherboard | MSI MAG X870E Tomahawk Max WiFi | 1,600 | Same |
| GPU | Zotac RTX 5090 32 GB Solid OC | 22,799 | Same card as PC2 (575 W) |
| RAM | Kingston Fury Beast 128 GB (2×64 GB) DDR5-5600 CL36 (KF556C36BBEAK2-128) | 9,899 | Differs: PC1/PC2 have 64 GB T-Create 6000 |
| OS drive, M2_2 | Samsung 9100 Pro 1 TB | 1,199 | Same as PC2's planned OS drive |
| Model drive, M2_1 | Samsung 9100 Pro 2 TB | 2,399 | Same |
| Cooler | Arctic Liquid Freezer III Pro 360 | 499 | Same |
| Case | Fractal Design North XL Dark TG | 780 | Same |
| PSU | Seasonic Vertex PX-1200 ATX 3.1 | 1,600 | Same (fleet standard) |
| Total | Sum of the quote lines | 43,324 | PC1 RM36,748, PC2 RM34,999 |
What PC3 changes:
- RAM. PC3 is the only 128 GB node. The kit is the Kingston one declined for PC1/PC2 (it was sold out at IdealTech); its datasheet rates EXPO at 5600, AMD's two-DIMM rating for the 9950X. The upgrade was chosen when the spec was confirmed on 2026-09-25; the reasoning behind it is not recorded here. The reopen condition named a WAN-only node, so PC3 is the natural home for WAN 2.2 routing. Its RAM does not reopen 128 GB for PC1/PC2.
- Memory acceptance. Run PC3 at EXPO 5600, not 6000; see acceptance targets.
- Third-node gate. PC3 was bought before the strategy doc's 30-day production gate. The two-versus-three replay is the planning estimate for its value: growth headroom, redundancy and model residency, not Modal spend.
Buying outside the package still requires:
- each exact SKU on the invoice, with Malaysian warranty ownership for the GPU, PSU and drives;
- confirming the Tomahawk Max's shipped BIOS supports the 9950X (MAX boards ship with Ryzen 9000 support);
- assembly with only the PSU's own supplied 12V-2x6 cable, plus the connector thermal watchdog.
Acceptance targets for the delivered nodes#
- Latency: MiniMax H3 cold minus warm of 6.5 s or less, and Krea 2 of 5 s or less, with files on NVMe.
- Storage:
- Models on a plain XFS or ext4 filesystem on the M2_1 9100 Pro, mounted into the container, not in an
image layer: 1 MiB partition alignment,
noatime, periodicfstrimrather than continuousdiscard, 15–20% free. Confirm the Samsung namespace read counters rise by about the resolved bytes on each cold read; process counters alone do not prove physical reads. - M.2 heat spreaders fitted, with case airflow across the Gen5 drives.
- Every 9100 Pro at 32 GT/s x4. M2_2 shares lanes with rear USB4 and runs x2 until it is set to x4 in the BIOS (why).
- Drive temperatures logged during a sustained read of at least 200 GB. Put the model cache in the cooler slot.
- Models on a plain XFS or ext4 filesystem on the M2_1 9100 Pro, mounted into the container, not in an
image layer: 1 MiB partition alignment,
- Queue settings: benchmark cold loads with
read_ahead_kbat 128 (default), 2048 and 8192, and withmax_sectors_kbraised tomax_hw_sectors_kb. Keep the winner with a udev rule. Faster hosts hinted at larger values, but drive and host were confounded, so measure first. - GPU: power limit at least 575 W (
nvidia-smi -q -d POWER), with the link at Gen5 x16 under load. - Power: sustain full GPU load with concurrent CPU load and confirm no PSU OCP trip. Record idle,
ComfyUI and combined stress input power at the branch meter and check for voltage sag and nuisance trips.
Commission the connector thermal watchdog and IR-check both ends
of the 16-pin cable at full load. After initial thermal cycling, isolate mains and inspect the 12V-2x6
connection for full insertion, discolouration or deformation. (Only a fallback HX1200i node would also need the
corsair-psuhwmon check and, if multi-rail OCP trips,liquidctl initialize --single-12v-ocpat every boot.) - Memory:
- PC1/PC2: enable EXPO (6000) with UCLK = MEMCLK and pass a multi-hour memory stress test before taking jobs. If it fails, lower the frequency to 5600 with EXPO still on. Disabling EXPO falls back to 4800.
- PC3: enable the Kingston kit's EXPO 5600 profile with UCLK = MEMCLK and pass the same stress test. Two 64 GB DIMMs train more slowly than 32 GB ones; allow for a long first boot.
- Also measure a switch back to a recently used model in a long-running process. The benchmarks never covered it, and routing will make it common.
Background prefetch rules#
Nodes can fetch the next job's missing files and warm them into RAM while the current job runs. No hardware change is needed.
- Stage on the model drive. Download into a staging directory on the M2_1 filesystem, verify the hash, then rename. The atomic rename means ComfyUI never sees a partial file.
- The loader has priority. Run prefetch at idle I/O priority (
ionice -c3or a low cgroupio.weight). NAS fills run at about 200 MB/s from the HDD mirror. - Protect the page cache. Write prefetched files with O_DIRECT, or call
posix_fadvise(POSIX_FADV_DONTNEED)after writing, so the running job's model pages are not evicted. - Warm RAM only when it fits. Pre-read the next set only when it fits beside the running set and
MemAvailableconfirms the headroom.
Prefetch hides NAS and network speed, which further weakens the case for 10GbE.
Reproducing#
cd comfyicu/orchestrator/benchmarks/analysis
python3 rtx5090_cold_penalty_table.py # benchmark cohort + page-cache control
python3 production_ram_nvme_replay.py 2026-07 --nodes 2,3 # capacity, cache and prefetch replay (also 2026-08)
The replay reads gitignored runs.csv files; regenerate them with npx tsx fetch_runs.ts --months … from
orchestrator/ (see ../../docs/cost-analysis/data/months/README.md).