RTX 5090 NVMe RAID 0: broader evidence analysis#

Status: Decision Historical question: one 4 TB Samsung 9100 Pro versus two 2 TB NVMe devices in RAID 0 Evidence snapshot: 2026-09-23; tables below are intentionally frozen to that decision point Primary outcome: full ComfyUI cold-minus-warm latency, not fio or file-reader throughput Deployment application (2026-09-26): the "no RAID" finding holds, but the "one 4 TB drive" purchase advice does not. Available devices were limited to 2 TB, and every ordered node has one 9100 Pro 2 TB model drive plus a separate OS drive. Acceptance targets (MiniMax ≤6.5 s, Krea ≤5 s cold-minus-warm) replace the ~7 GiB/s reader gate below. Record: node decisions. Current variance analysis: single-path multivariable reference.

This document is the standalone decision summary. The chronological protocol, acquisition log, per-host topology and complete workload results remain in the full single-NVMe versus RAID 0 experiment journal. The multivariable investigation journal records later controls, metadata backfills and factor-audit decisions. The same evidence is available as a storage-path dashboard RAID evidence view, calculated from the manually refreshed normalized DuckDB table. The dashboard derives Pearson and Spearman correlations for multiple admission slices, the historical purchase-relevant 9100-Pro comparison, cross-workload agreement and reacquisition variance; none of those displayed statistics is copied from this document.

Executive conclusion#

The expanded evidence says two different things with different confidence. Storage-path quality clearly matters: latest-host MiniMax penalties span 5.419-34.445 seconds, and actual model-reader rate ranks the penalty far better than Vast's advertised disk score or high-queue-depth fio. But it does not follow that RAID 0 is required to reach the useful part of that curve.

The strongest balanced comparison now contains two single-drive and two RAID0 hosts, all using 575 W RTX 5090s and both complete workloads. RAID's mean model-reader rate is about 31% higher, while its observed paired advantage is 1.291 seconds on MiniMax H3 and 0.743 seconds on Krea 2. Both are below the predeclared two-second threshold. A single SN8100 host almost exactly matches the RAID means on both workloads: only 0.352 seconds slower on MiniMax and 0.050 seconds slower on Krea.

The broad full-width slice is not a layout estimate. Its eight single-drive hosts include ordinary fast paths, slow consumer drives and a 0.851 GiB/s virtual/block path with no exposed NVMe; its two RAID0 hosts are both fast paths. That mixture makes the single mean 6.021 seconds slower, while the median difference is 2.710 seconds. It establishes that avoiding a bad path matters, not that striping caused the entire gap.

The purchase-relevant observational slice in this frozen snapshot compares two 600 W single-9100-Pro hosts with the two 575 W RAID0 hosts. RAID increases mean model-reader throughput by 50.5% (7.114 to 10.705 GiB/s), but lowers the mean paired MiniMax penalty by only 0.996 seconds (7.926 to 6.930 seconds). The historical purchase implication was to start with one drive and preserve the second CPU-connected M.2 path. The ordered fleet applies that finding as one 2 TB model drive plus a separate OS drive, without model-cache RAID.

Scope and evidence rules#

This analysis excludes the older pre-fix results. It uses successful ComfyUI v0.37.0 runs with the fixed resident-cache protocol:

evict the resolved model-file page cache
  -> fresh-process cold run with verified process block reads
  -> different-seed warm run in the same process

The physical machine is the independent unit. Reacquisitions of the same machine demonstrate repeatability but do not increase the cross-host sample size. After deduplication, this 2026-09-23 snapshot has sixteen independent MiniMax hosts, eleven Krea hosts and ten machines with both workloads.

Evidence is separated into three levels:

  1. Cgroup-attributed physical: process bytes are corroborated through cgroup counters at the MD/NVMe layer. This is the strongest storage-path evidence.
  2. Host-wide physical support: /sys/block counters match process reads, but cgroup io.stat is unavailable. Background host I/O cannot be excluded.
  3. Process-block verified: the workload definitely misses its immediate file page cache, but an outer loop/container cache may still satisfy the read before physical NAND.

No cross-host comparison is treated as a causal SSD experiment. GPU power, H2D, CPU, kernel, PCIe width, filesystem and container path remain possible confounders.

1. Strongest balanced evidence: two workloads, four hosts#

These hosts ran both complete workloads under v0.37.0. All use 575 W GPUs, have visible NVMe metadata and measure at least 10 GiB/s H2D. They are still heterogeneous: one single host is Gen5 x8, one is a shared Threadripper Pro machine, and the RAID hosts are Gen5 x16. The point of the slice is balance and replication, not a causal layout claim.

Layout and host MiniMax cold / warm / penalty Krea cold / warm / penalty MiniMax / Krea model reader
Single T700, 30597 91.531 / 82.371 / 9.160 s 12.069 / 5.529 / 6.540 s 6.264 / 6.247 GiB/s
Single SN8100, 137184 71.084 / 63.802 / 7.282 s 9.417 / 4.263 / 5.154 s 10.098 / 10.578 GiB/s
RAID SN850X, 112410 70.823 / 65.404 / 5.419 s 8.664 / 4.265 / 4.399 s 11.504 / 11.488 GiB/s
RAID T705, 147512 72.176 / 63.736 / 8.440 s 10.081 / 4.272 / 5.809 s 9.906 / 10.390 GiB/s

The layout summaries are:

Workload RAID mean / range Single mean / range Single minus RAID mean
MiniMax H3 6.930 / 5.419-8.440 s 8.221 / 7.282-9.160 s +1.291 s
Krea 2 5.104 / 4.399-5.809 s 5.847 / 5.154-6.540 s +0.743 s

Raw MiniMax cold time differs by 9.808 seconds because machine 30597 has an 82.371-second warm floor. That is exactly why cold-minus-warm, rather than raw cold time, is the outcome. With only two hosts per layout, there are only six possible two-versus-two label allocations; the observed differences are not extreme. More importantly, the single SN8100 result essentially reaches the RAID plateau on both workloads despite its x8 GPU link.

2. Broader MiniMax slice after basic platform admission#

The broader comparison admits v0.37.0 hosts with:

Layout Hosts Model-reader mean Penalty mean Penalty median Penalty range
Two-drive MD RAID0 2 10.705 GiB/s 6.930 s 6.930 s 5.419-8.440 s
Single/container path 8 5.189 GiB/s 12.951 s 9.640 s 5.732-34.445 s

The 6.021-second mean gap is not robust evidence of a RAID effect. Machine 150337 alone has no exposed NVMe, a 0.851 GiB/s model reader and a 34.445-second penalty; omitting it reduces the gap to 2.951 seconds. The remaining singles still mix 3.0 GiB/s paths with 7.6-9.6 GiB/s paths. The fastest single result (5.732 seconds) is within 0.313 seconds of the fastest RAID result, while the slower RAID result (8.440 seconds) is slower than two singles.

A narrower purchase-relevant check is more informative:

Observational group Hosts Reader mean Penalty mean Warm mean Cold mean
600 W, single 9100 Pro 2 7.114 GiB/s 7.926 s 60.682 s 68.607 s
575 W, two-drive RAID0 2 10.705 GiB/s 6.930 s 64.570 s 71.500 s

The 50.5% reader-rate gain compresses to a 0.996-second paired gain. Raw RAID cold time is actually 2.893 seconds slower because its warm floor is 3.888 seconds slower; that is platform variance, not a storage result. This slice remains cross-host, but it is a better answer to the purchase question than the unadjusted eight-versus-two mean.

3. Physical-device evidence#

Only schema-v4 runs can distinguish immediate file-cache misses from reads that reached lower block layers.

Host and layout Evidence GPU/platform Model reader Penalty What it establishes
112410, 2x SN850X RAID Cgroup-attributed MD plus both NVMe members 575 W, Ryzen 9950X3D 11.504 GiB/s 5.419 s RAID really used both drives; same-host result replicated
147011, single MP700 Elite Cgroup-attributed physical NVMe 600 W, Ryzen 9950X 4.536 GiB/s 7.302 s One drive served every cold byte physically
62211, single 9100 Pro Cgroup-attributed physical NVMe 500 W, Core Ultra 7 5.699 GiB/s 11.776 s Exact planned drive served every cold byte physically
132465, single SN850X Host-wide NVMe support only 600 W, i9-12900K, H2D 3.359 GiB/s 5.300 GiB/s 20.354 s Physical propagation is strongly supported, but accelerator path is a major confounder

The closest fully physical single-versus-RAID comparison is still cross-host: machine 112410 RAID versus machine 147011 single MP700. RAID's reader is 2.54x faster, but its application penalty is only 1.883 seconds smaller. Raw cold time is 0.805 seconds slower on RAID because its warm floor is slower. This supports single-drive sufficiency but does not measure the incremental benefit of striping on one machine.

The physical 9100 Pro result is deliberately not compared numerically to RAID as a drive test: its 500 W GPU and Intel platform have a much slower 69.986-second warm floor. It proves the exact drive and eviction path work, not performance equality.

4. Why the slow single-SN850X result does not support RAID#

Machine 132465 recorded a 20.354-second penalty with one SN850X, but its pinned H2D was only 3.359 GiB/s. The other three 600 W hosts measured 46.747-51.220 GiB/s H2D and averaged a 7.718-second penalty while having nearly the same warm floor. Its direct QD1 storage rate was also only 3.549 GB/s.

The run shows that a machine can look like a 600 W, Gen5 x16 RTX 5090 on paper while having a badly restricted measured cold-loading path. Because both storage and accelerator transfer differ, assigning the extra time to “one drive instead of RAID” would be invalid.

5. What predicts the application result#

Across all sixteen latest MiniMax host observations, the useful storage metrics are the ones closest to the real access pattern:

Metric Pearson with penalty Spearman rank with penalty
Actual model-reader rate -0.707 -0.803
Buffered fio -0.752 not used for the primary ranking
Direct QD1 fio -0.597 -0.482
Direct QD32 fio -0.143 -0.344
Vast advertised disk speed -0.149 not decision-useful
Pinned H2D -0.173 +0.118

Full-workload read-active rate has a -0.912 Pearson and -0.994 rank relationship with penalty, but it is measured inside the timed cold run and is therefore mechanically coupled to the outcome. It is evidence about where time went, not an independent pre-run predictor.

Machine 30024 is the clearest counterexample to treating QD32 as the answer: its four-drive RAID10 reached 21.447 GB/s at QD32, but only 5.164 GiB/s in the model reader and a 30.302-second MiniMax penalty. It is classified as other MD and excluded from RAID0 summaries, but retained as evidence that array headline throughput can fail to reach the application.

The actual model reader exposes a clear knee in the ten-host, 575-600 W, H2D-at-least-10, x16 single/RAID0 slice:

Model-reader tier Hosts Mean reader Mean penalty Median penalty
Below 3.5 GiB/s 3 2.302 GiB/s 20.824 s 14.536 s
3.5 to below 7 GiB/s 3 5.814 GiB/s 8.861 s 9.160 s
At least 7 GiB/s 4 9.644 GiB/s 7.103 s 7.086 s

The slow tier includes machine 150337's non-NVMe 0.851 GiB/s path; without it, the other two slow hosts average a 14.014-second penalty. The important shape is diminishing returns: moving from roughly 3 to 6 GiB/s removes about five seconds, while moving from roughly 6 to 10 GiB/s removes only about 1.8 seconds on average. Above 5 GiB/s, a simple inverse-reader model explains only 28% of the remaining penalty variance. The application is approaching a non-storage floor.

The conclusion replicates across workloads. On the ten machines with both results, MiniMax and Krea paired penalties have 0.976 Pearson and Spearman correlations, while their independently measured model-reader rates correlate at 0.998. This is strong evidence that the host storage/container path is real and repeatable, rather than five-run noise. It is not evidence that RAID is the only way to build a fast path.

Within-run noise is much smaller than cross-host spread. Across the sixteen latest MiniMax results, the median standard deviation of the five paired penalties is 0.486 seconds; the host means span 29.026 seconds. Schema-v3-to-v4 reacquisitions changed paired means by only 0.046-1.291 seconds. The large host differences are therefore reproducible, but attributing them among drive, CPU, filesystem and container layers still requires the captured covariates.

6. Why a normal Vast rental cannot provide the missing crossover#

The proposed same-machine single drive -> RAID 0 -> single drive experiment is not available on any captured Vast host. Every host with a matched pair of benchmark-capable NVMe devices exposed that pair as an array the provider had already configured. The only captured multi-NVMe host without MD RAID, machine 150183, contained a 2 TB Samsung 9100 Pro and a 1 TB Seagate device rather than two matched, unused test devices.

Machine 112410 makes the ownership boundary explicit. It has a separate 1 TB system SN850X plus two 4 TB SN850X devices in a conventional Linux MD RAID 0 (md0, two members, 512 KiB chunks). The benchmark itself runs in a Docker OverlayFS root whose upper and work directories are below the host's /var/lib/docker. Schema-v4 cgroup counters prove that the workload's reads passed through md0 and split across both physical members. This is a real software RAID array, but it is host infrastructure underneath the rental rather than a spare array owned by the renter. Stopping or rebuilding it would remove the storage path serving the container and may affect other provider-managed data.

The normal Vast path is therefore:

physical NVMe devices
  -> optional host-configured MD RAID/filesystem
  -> host Docker storage
  -> the renter's unprivileged Docker container/OverlayFS
  -> ComfyUI

It is not Docker inside Docker, and container root is not host root. Vast's FAQ documents ordinary instances as unprivileged Docker containers, while its instance guide says Docker-in-Docker and renter bare-metal access are not supported. Linux block-device reconfiguration would require host-exposed raw devices plus elevated mount/device capabilities that ordinary offers intentionally do not grant; Docker documents that these require explicit device/capability or privileged-container access in its runtime privilege reference.

The MD RAID implementation is the same basic technology that could be used on the planned node, but the end-to-end storage systems are not equivalent. The local node can place models on a filesystem mounted directly from one NVMe or from md0, with full control over the filesystem, MD chunk/read-ahead, kernel and mount options. Vast adds provider-selected Docker/overlay, quota, caching, cgroup and contention behavior. Consequently the Vast results are valid observations of those complete container paths, not a controlled estimate of the incremental bare-metal benefit from adding the second SSD.

A valid pre-purchase crossover would require either a cooperative host to dedicate two identical unused drives and perform both configurations from the host side, or a true bare-metal rental with administrative control. Without that, more independent preconfigured Vast hosts can improve observational confidence but cannot turn the comparison into a causal same-machine intervention. Such a crossover is not treated as a prerequisite for the purchase decision. The feasible strategy is to use the heterogeneous-host distribution, then acceptance-test the purchased single drive and keep a second CPU-connected slot available if the real application path lands below the useful tier.

Decision and remaining uncertainty#

The evidence supports the following narrow statements:

It does not establish that RAID can never help or that a single 9100 Pro is intrinsically identical to a two-drive array. A normal Vast rental cannot isolate the incremental effect because the relevant host block devices and Docker backing store are provider-controlled. The analysis therefore treats the heterogeneous hosts probabilistically and reports overlap, repeatability and covariates instead of pretending they are a same-machine intervention.

For the ordered nodes, the decision has already been applied: one 2 TB 9100 Pro serves the model cache and a separate drive serves the OS/container/output path, with no model-cache RAID. Accept each node only after confirming the intended SSD and GPU links, normal pinned H2D, direct-filesystem physical NVMe reads and five complete cold/warm pairs. Apply the current MiniMax and Krea thresholds in the node decision, not the historical ~7 GiB/s reader or 5-9 second ranges in this snapshot. If a node misses, investigate firmware, thermals, filesystem, block layers, PCIe routing and actual in-workload delivery before assuming striping is the remedy.