RTX 5090 NVMe RAID 0: broader evidence analysis#
Status: Decision Historical question: one 4 TB Samsung 9100 Pro versus two 2 TB NVMe devices in RAID 0 Evidence snapshot: 2026-09-23; tables below are intentionally frozen to that decision point Primary outcome: full ComfyUI cold-minus-warm latency, not
fioor file-reader throughput Deployment application (2026-09-26): the "no RAID" finding holds, but the "one 4 TB drive" purchase advice does not. Available devices were limited to 2 TB, and every ordered node has one 9100 Pro 2 TB model drive plus a separate OS drive. Acceptance targets (MiniMax ≤6.5 s, Krea ≤5 s cold-minus-warm) replace the ~7 GiB/s reader gate below. Record: node decisions. Current variance analysis: single-path multivariable reference.
This document is the standalone decision summary. The chronological protocol, acquisition log, per-host topology and complete workload results remain in the full single-NVMe versus RAID 0 experiment journal. The multivariable investigation journal records later controls, metadata backfills and factor-audit decisions. The same evidence is available as a storage-path dashboard RAID evidence view, calculated from the manually refreshed normalized DuckDB table. The dashboard derives Pearson and Spearman correlations for multiple admission slices, the historical purchase-relevant 9100-Pro comparison, cross-workload agreement and reacquisition variance; none of those displayed statistics is copied from this document.
Executive conclusion#
The expanded evidence says two different things with different confidence. Storage-path quality clearly
matters: latest-host MiniMax penalties span 5.419-34.445 seconds, and actual model-reader rate ranks the
penalty far better than Vast's advertised disk score or high-queue-depth fio. But it does not follow
that RAID 0 is required to reach the useful part of that curve.
The strongest balanced comparison now contains two single-drive and two RAID0 hosts, all using 575 W RTX 5090s and both complete workloads. RAID's mean model-reader rate is about 31% higher, while its observed paired advantage is 1.291 seconds on MiniMax H3 and 0.743 seconds on Krea 2. Both are below the predeclared two-second threshold. A single SN8100 host almost exactly matches the RAID means on both workloads: only 0.352 seconds slower on MiniMax and 0.050 seconds slower on Krea.
The broad full-width slice is not a layout estimate. Its eight single-drive hosts include ordinary fast paths, slow consumer drives and a 0.851 GiB/s virtual/block path with no exposed NVMe; its two RAID0 hosts are both fast paths. That mixture makes the single mean 6.021 seconds slower, while the median difference is 2.710 seconds. It establishes that avoiding a bad path matters, not that striping caused the entire gap.
The purchase-relevant observational slice in this frozen snapshot compares two 600 W single-9100-Pro hosts with the two 575 W RAID0 hosts. RAID increases mean model-reader throughput by 50.5% (7.114 to 10.705 GiB/s), but lowers the mean paired MiniMax penalty by only 0.996 seconds (7.926 to 6.930 seconds). The historical purchase implication was to start with one drive and preserve the second CPU-connected M.2 path. The ordered fleet applies that finding as one 2 TB model drive plus a separate OS drive, without model-cache RAID.
Scope and evidence rules#
This analysis excludes the older pre-fix results. It uses successful ComfyUI v0.37.0 runs with the fixed resident-cache protocol:
evict the resolved model-file page cache
-> fresh-process cold run with verified process block reads
-> different-seed warm run in the same process
The physical machine is the independent unit. Reacquisitions of the same machine demonstrate repeatability but do not increase the cross-host sample size. After deduplication, this 2026-09-23 snapshot has sixteen independent MiniMax hosts, eleven Krea hosts and ten machines with both workloads.
Evidence is separated into three levels:
- Cgroup-attributed physical: process bytes are corroborated through cgroup counters at the MD/NVMe layer. This is the strongest storage-path evidence.
- Host-wide physical support:
/sys/blockcounters match process reads, but cgroupio.statis unavailable. Background host I/O cannot be excluded. - Process-block verified: the workload definitely misses its immediate file page cache, but an outer loop/container cache may still satisfy the read before physical NAND.
No cross-host comparison is treated as a causal SSD experiment. GPU power, H2D, CPU, kernel, PCIe width, filesystem and container path remain possible confounders.
1. Strongest balanced evidence: two workloads, four hosts#
These hosts ran both complete workloads under v0.37.0. All use 575 W GPUs, have visible NVMe metadata and measure at least 10 GiB/s H2D. They are still heterogeneous: one single host is Gen5 x8, one is a shared Threadripper Pro machine, and the RAID hosts are Gen5 x16. The point of the slice is balance and replication, not a causal layout claim.
| Layout and host | MiniMax cold / warm / penalty | Krea cold / warm / penalty | MiniMax / Krea model reader |
|---|---|---|---|
| Single T700, 30597 | 91.531 / 82.371 / 9.160 s | 12.069 / 5.529 / 6.540 s | 6.264 / 6.247 GiB/s |
| Single SN8100, 137184 | 71.084 / 63.802 / 7.282 s | 9.417 / 4.263 / 5.154 s | 10.098 / 10.578 GiB/s |
| RAID SN850X, 112410 | 70.823 / 65.404 / 5.419 s | 8.664 / 4.265 / 4.399 s | 11.504 / 11.488 GiB/s |
| RAID T705, 147512 | 72.176 / 63.736 / 8.440 s | 10.081 / 4.272 / 5.809 s | 9.906 / 10.390 GiB/s |
The layout summaries are:
| Workload | RAID mean / range | Single mean / range | Single minus RAID mean |
|---|---|---|---|
| MiniMax H3 | 6.930 / 5.419-8.440 s | 8.221 / 7.282-9.160 s | +1.291 s |
| Krea 2 | 5.104 / 4.399-5.809 s | 5.847 / 5.154-6.540 s | +0.743 s |
Raw MiniMax cold time differs by 9.808 seconds because machine 30597 has an 82.371-second warm floor. That is exactly why cold-minus-warm, rather than raw cold time, is the outcome. With only two hosts per layout, there are only six possible two-versus-two label allocations; the observed differences are not extreme. More importantly, the single SN8100 result essentially reaches the RAID plateau on both workloads despite its x8 GPU link.
2. Broader MiniMax slice after basic platform admission#
The broader comparison admits v0.37.0 hosts with:
- actual RTX 5090 power of 575-600 W;
- measured H2D of at least 10 GiB/s, excluding the 3.359 GiB/s transfer-path outlier;
- a full five-pair MiniMax result;
- and, for the stricter table below, a full-width x16 GPU link.
| Layout | Hosts | Model-reader mean | Penalty mean | Penalty median | Penalty range |
|---|---|---|---|---|---|
| Two-drive MD RAID0 | 2 | 10.705 GiB/s | 6.930 s | 6.930 s | 5.419-8.440 s |
| Single/container path | 8 | 5.189 GiB/s | 12.951 s | 9.640 s | 5.732-34.445 s |
The 6.021-second mean gap is not robust evidence of a RAID effect. Machine 150337 alone has no exposed NVMe, a 0.851 GiB/s model reader and a 34.445-second penalty; omitting it reduces the gap to 2.951 seconds. The remaining singles still mix 3.0 GiB/s paths with 7.6-9.6 GiB/s paths. The fastest single result (5.732 seconds) is within 0.313 seconds of the fastest RAID result, while the slower RAID result (8.440 seconds) is slower than two singles.
A narrower purchase-relevant check is more informative:
| Observational group | Hosts | Reader mean | Penalty mean | Warm mean | Cold mean |
|---|---|---|---|---|---|
| 600 W, single 9100 Pro | 2 | 7.114 GiB/s | 7.926 s | 60.682 s | 68.607 s |
| 575 W, two-drive RAID0 | 2 | 10.705 GiB/s | 6.930 s | 64.570 s | 71.500 s |
The 50.5% reader-rate gain compresses to a 0.996-second paired gain. Raw RAID cold time is actually 2.893 seconds slower because its warm floor is 3.888 seconds slower; that is platform variance, not a storage result. This slice remains cross-host, but it is a better answer to the purchase question than the unadjusted eight-versus-two mean.
3. Physical-device evidence#
Only schema-v4 runs can distinguish immediate file-cache misses from reads that reached lower block layers.
| Host and layout | Evidence | GPU/platform | Model reader | Penalty | What it establishes |
|---|---|---|---|---|---|
| 112410, 2x SN850X RAID | Cgroup-attributed MD plus both NVMe members | 575 W, Ryzen 9950X3D | 11.504 GiB/s | 5.419 s | RAID really used both drives; same-host result replicated |
| 147011, single MP700 Elite | Cgroup-attributed physical NVMe | 600 W, Ryzen 9950X | 4.536 GiB/s | 7.302 s | One drive served every cold byte physically |
| 62211, single 9100 Pro | Cgroup-attributed physical NVMe | 500 W, Core Ultra 7 | 5.699 GiB/s | 11.776 s | Exact planned drive served every cold byte physically |
| 132465, single SN850X | Host-wide NVMe support only | 600 W, i9-12900K, H2D 3.359 GiB/s | 5.300 GiB/s | 20.354 s | Physical propagation is strongly supported, but accelerator path is a major confounder |
The closest fully physical single-versus-RAID comparison is still cross-host: machine 112410 RAID versus machine 147011 single MP700. RAID's reader is 2.54x faster, but its application penalty is only 1.883 seconds smaller. Raw cold time is 0.805 seconds slower on RAID because its warm floor is slower. This supports single-drive sufficiency but does not measure the incremental benefit of striping on one machine.
The physical 9100 Pro result is deliberately not compared numerically to RAID as a drive test: its 500 W GPU and Intel platform have a much slower 69.986-second warm floor. It proves the exact drive and eviction path work, not performance equality.
4. Why the slow single-SN850X result does not support RAID#
Machine 132465 recorded a 20.354-second penalty with one SN850X, but its pinned H2D was only 3.359 GiB/s. The other three 600 W hosts measured 46.747-51.220 GiB/s H2D and averaged a 7.718-second penalty while having nearly the same warm floor. Its direct QD1 storage rate was also only 3.549 GB/s.
The run shows that a machine can look like a 600 W, Gen5 x16 RTX 5090 on paper while having a badly restricted measured cold-loading path. Because both storage and accelerator transfer differ, assigning the extra time to “one drive instead of RAID” would be invalid.
5. What predicts the application result#
Across all sixteen latest MiniMax host observations, the useful storage metrics are the ones closest to the real access pattern:
| Metric | Pearson with penalty | Spearman rank with penalty |
|---|---|---|
| Actual model-reader rate | -0.707 | -0.803 |
Buffered fio |
-0.752 | not used for the primary ranking |
Direct QD1 fio |
-0.597 | -0.482 |
Direct QD32 fio |
-0.143 | -0.344 |
| Vast advertised disk speed | -0.149 | not decision-useful |
| Pinned H2D | -0.173 | +0.118 |
Full-workload read-active rate has a -0.912 Pearson and -0.994 rank relationship with penalty, but it is measured inside the timed cold run and is therefore mechanically coupled to the outcome. It is evidence about where time went, not an independent pre-run predictor.
Machine 30024 is the clearest counterexample to treating QD32 as the answer: its four-drive RAID10 reached 21.447 GB/s at QD32, but only 5.164 GiB/s in the model reader and a 30.302-second MiniMax penalty. It is classified as other MD and excluded from RAID0 summaries, but retained as evidence that array headline throughput can fail to reach the application.
The actual model reader exposes a clear knee in the ten-host, 575-600 W, H2D-at-least-10, x16 single/RAID0 slice:
| Model-reader tier | Hosts | Mean reader | Mean penalty | Median penalty |
|---|---|---|---|---|
| Below 3.5 GiB/s | 3 | 2.302 GiB/s | 20.824 s | 14.536 s |
| 3.5 to below 7 GiB/s | 3 | 5.814 GiB/s | 8.861 s | 9.160 s |
| At least 7 GiB/s | 4 | 9.644 GiB/s | 7.103 s | 7.086 s |
The slow tier includes machine 150337's non-NVMe 0.851 GiB/s path; without it, the other two slow hosts average a 14.014-second penalty. The important shape is diminishing returns: moving from roughly 3 to 6 GiB/s removes about five seconds, while moving from roughly 6 to 10 GiB/s removes only about 1.8 seconds on average. Above 5 GiB/s, a simple inverse-reader model explains only 28% of the remaining penalty variance. The application is approaching a non-storage floor.
The conclusion replicates across workloads. On the ten machines with both results, MiniMax and Krea paired penalties have 0.976 Pearson and Spearman correlations, while their independently measured model-reader rates correlate at 0.998. This is strong evidence that the host storage/container path is real and repeatable, rather than five-run noise. It is not evidence that RAID is the only way to build a fast path.
Within-run noise is much smaller than cross-host spread. Across the sixteen latest MiniMax results, the median standard deviation of the five paired penalties is 0.486 seconds; the host means span 29.026 seconds. Schema-v3-to-v4 reacquisitions changed paired means by only 0.046-1.291 seconds. The large host differences are therefore reproducible, but attributing them among drive, CPU, filesystem and container layers still requires the captured covariates.
6. Why a normal Vast rental cannot provide the missing crossover#
The proposed same-machine single drive -> RAID 0 -> single drive experiment is not available on any
captured Vast host. Every host with a matched pair of benchmark-capable NVMe devices exposed that pair as an
array the provider had already configured. The only captured multi-NVMe host without MD RAID, machine
150183, contained a 2 TB Samsung 9100 Pro and a 1 TB Seagate device rather than two matched, unused test
devices.
Machine 112410 makes the ownership boundary explicit. It has a separate 1 TB system SN850X plus two 4 TB
SN850X devices in a conventional Linux MD RAID 0 (md0, two members, 512 KiB chunks). The benchmark itself
runs in a Docker OverlayFS root whose upper and work directories are below the host's /var/lib/docker.
Schema-v4 cgroup counters prove that the workload's reads passed through md0 and split across both physical
members. This is a real software RAID array, but it is host infrastructure underneath the rental rather than
a spare array owned by the renter. Stopping or rebuilding it would remove the storage path serving the
container and may affect other provider-managed data.
The normal Vast path is therefore:
physical NVMe devices
-> optional host-configured MD RAID/filesystem
-> host Docker storage
-> the renter's unprivileged Docker container/OverlayFS
-> ComfyUI
It is not Docker inside Docker, and container root is not host root. Vast's
FAQ documents ordinary instances as unprivileged Docker containers, while
its instance guide says
Docker-in-Docker and renter bare-metal access are not supported. Linux block-device reconfiguration would
require host-exposed raw devices plus elevated mount/device capabilities that ordinary offers intentionally
do not grant; Docker documents that these require explicit device/capability or privileged-container access
in its runtime privilege reference.
The MD RAID implementation is the same basic technology that could be used on the planned node, but the
end-to-end storage systems are not equivalent. The local node can place models on a filesystem mounted
directly from one NVMe or from md0, with full control over the filesystem, MD chunk/read-ahead, kernel and
mount options. Vast adds provider-selected Docker/overlay, quota, caching, cgroup and contention behavior.
Consequently the Vast results are valid observations of those complete container paths, not a controlled
estimate of the incremental bare-metal benefit from adding the second SSD.
A valid pre-purchase crossover would require either a cooperative host to dedicate two identical unused drives and perform both configurations from the host side, or a true bare-metal rental with administrative control. Without that, more independent preconfigured Vast hosts can improve observational confidence but cannot turn the comparison into a causal same-machine intervention. Such a crossover is not treated as a prerequisite for the purchase decision. The feasible strategy is to use the heterogeneous-host distribution, then acceptance-test the purchased single drive and keep a second CPU-connected slot available if the real application path lands below the useful tier.
Decision and remaining uncertainty#
The evidence supports the following narrow statements:
- RAID 0 can materially increase standalone sequential/model-reader throughput.
- In the balanced dual-workload cohort, that increase translates to 0.743-1.291 seconds, below the predeclared two-second threshold.
- A single fast Gen5 path can reach the same application plateau: machine 137184's single SN8100 is within 0.050 seconds of the RAID mean on Krea and 0.352 seconds on MiniMax.
- A single direct-attached Gen5 NVMe can serve every cold model byte physically and produce an acceptable application penalty.
- Vast's advertised speed and QD32
fiodo not predict the application result; the actual model reader and paired cold penalty must be measured. - Cross-host variance in GPU floor, H2D, CPU, kernel and container/storage path is comparable to or larger than the purchase-relevant observed layout difference.
It does not establish that RAID can never help or that a single 9100 Pro is intrinsically identical to a two-drive array. A normal Vast rental cannot isolate the incremental effect because the relevant host block devices and Docker backing store are provider-controlled. The analysis therefore treats the heterogeneous hosts probabilistically and reports overlap, repeatability and covariates instead of pretending they are a same-machine intervention.
For the ordered nodes, the decision has already been applied: one 2 TB 9100 Pro serves the model cache and a separate drive serves the OS/container/output path, with no model-cache RAID. Accept each node only after confirming the intended SSD and GPU links, normal pinned H2D, direct-filesystem physical NVMe reads and five complete cold/warm pairs. Apply the current MiniMax and Krea thresholds in the node decision, not the historical ~7 GiB/s reader or 5-9 second ranges in this snapshot. If a node misses, investigate firmware, thermals, filesystem, block layers, PCIe routing and actual in-workload delivery before assuming striping is the remedy.