RTX 5090 single-NVMe versus RAID 0 dual-workload experiment#
Status: Vast container-path matrix complete; initial one-drive purchase is supported, but direct physical single-versus-RAID performance remains a local bare-metal acceptance test Started: 2026-09-23 Decision: One 4 TB Gen5 NVMe versus two 2 TB Gen5 NVMe devices in RAID 0 for the planned RTX 5090 service node Procurement application: The deployment now permits only 2 TB NVMe devices. Start with one 2 TB Samsung model cache if the active catalogue fits with 15-20% headroom; add the second 2 TB device only for demonstrated capacity need or a same-machine application-level improvement As ordered (2026-09-26): no node uses RAID. Each has one 9100 Pro 2 TB model drive in M2_1 and a separate 9100 Pro OS drive in M2_2 (not the chipset Gen4 system SSD assumed below). Record and acceptance targets: node decisions.
Question#
Does two-drive RAID 0 materially reduce real ComfyUI cold-start latency compared with a sufficiently fast single Gen5 NVMe, or can one drive already feed model weights quickly enough that GPU loading and model initialization become the limiting path?
Measurement correction discovered after the matrix: on reacquired machine 147169, Vast stored the container filesystem in a buffered XFS loop-backing file. A cache-evicted 42.47 GB model read moved the full byte count through the process and
loop0, but zero read bytes through the Samsung NVMe. The oldphysical_read_bytesname therefore verifies an immediate model-file cache miss, not physical NAND I/O. Historical results remain valid end-to-end Vast/ComfyUI measurements, but they cannot by themselves prove the performance of a drive attached to a direct bare-metal filesystem. See the storage-path journal.
The evidence available at the start was not conclusive. MiniMax H3 runs associated two-drive RAID 0 with
much lower cold penalties, but those results came from different hosts and different container storage
implementations. The newer Krea 2 result demonstrated 10.80 GiB/s model-file reads and a 5.22-second cold
penalty from one WD Black SN8100. A controlled, replicated comparison must retain the full application
benchmark instead of treating fio or a sequential file reader as a proxy for ComfyUI.
Hypotheses#
- Single-drive-sufficient hypothesis: a correctly configured single Gen5 NVMe sustains approximately 10 GiB/s through the actual model path, and its end-to-end cold penalty is within one second or 10% of a comparable RAID 0 host on both workloads.
- RAID-beneficial hypothesis: two-drive RAID 0 repeatedly saves more than two seconds of end-to-end cold latency on both workloads after accounting for warm GPU performance and host differences.
- Synthetic throughput alone cannot decide between these hypotheses. The primary measurement is paired cold-versus-warm ComfyUI runtime with verified process block reads and, for new schema-v4 runs, corroborating counters at every loop, RAID and physical-NVMe layer.
Workloads#
Both workloads run during one provisioned-host session. ComfyUI, Python, PyTorch, CUDA dependencies and benchmark tooling are installed once. The union of both model sets is downloaded before either benchmark. Each workload nevertheless produces its own complete schema-v3 JSON result and ComfyUI log evidence.
| Workload | Persisted inputs and models | Current size |
|---|---|---|
| MiniMax H3 | 5 files | 42,471,968,783 bytes (39.56 GiB) |
| Krea 2 | 3 files | 18,638,004,998 bytes (17.36 GiB) |
| Combined session | 8 files before path de-duplication | 61,109,973,781 bytes (56.91 GiB) |
At least 160 GB of container storage will be allocated so both sets remain present while the sustained storage probes create their temporary test file.
Full per-workload benchmark#
Neither workload is reduced to a storage microbenchmark. Each independent result contains:
Physical GPU, CPU, RAM, motherboard, negotiated PCIe links, PCIe ancestry and storage topology.
CPU memory-copy and pinned CPU-to-GPU/GPU-to-CPU transfer measurements.
Sustained direct QD1, direct QD32 and cache-evicted buffered storage probes.
Five repetitions of:
fsync and evict every resolved workload file's immediate page cache -> launch a fresh ComfyUI process -> cold run, requiring process block reads >= 80% of resolved bytes -> warm run with a different seed in the same process -> stop ComfyUIThree normal buffered reads of the workload's real model/input set after page-cache eviction, including
/procprocess block-read verification. Historical schema-v3 files call thisphysical_read_*; that name does not account for an additional buffered loop-backing-file cache.Per-sample wall time, execution time, seeds, process resource traces and ComfyUI model-loading evidence.
The default resident ComfyUI cache is used. --cache-none is not used.
Experimental units and host selection#
The physical host, rather than each of its five workflow pairs, is the independent experimental unit. Five pairs quantify within-host noise; they do not substitute for replication across hosts.
The minimum primary matrix is:
| Layout | Required independent hosts | Required GPU/host properties |
|---|---|---|
| One Gen5 NVMe | 2 | Dedicated RTX 5090 host, Ryzen 9950X-class CPU, PCIe 5 GPU link, 575-600 W |
| Two-drive RAID 0 | 2 | Dedicated RTX 5090 host with the same preferred CPU/GPU properties |
Primary hosts must expose one allocated physical GPU (gpu_frac=1) negotiating PCIe 5.0 x16, matching the
intended build. Shared multi-GPU systems, GPUs reduced to x8 and four-drive arrays are excluded from the
primary comparison. Results from those machines may be retained as supporting evidence but cannot alone decide
the purchase.
Known candidates at the start of the experiment:
- Single 4 TB Samsung 9100 Pro: machines 147169 and 150656.
- Two 4 TB WD Black SN850X RAID 0: machine 112410.
- A second comparable dedicated RAID 0 RTX 5090 host must be acquired.
- Machine 137184's single SN8100 result is supporting evidence because the physical host has two GPUs.
All primary runs must use the same exact ComfyUI tag, Torch/CUDA build, cache policy and benchmark code. GPU power limit, PCIe width, CPU allocation, RAM, container runtime and filesystem path are captured rather than inferred from Vast's advertised disk-speed field.
PCIe topology controls#
Drive model and count do not fully specify a storage layout. Two nominally identical RAID members can each negotiate Gen5 x4 yet still share a chipset uplink, while two CPU-direct M.2 slots can each have an independent x4 path. Those configurations must not be treated as equivalent. The GPU link is another independent variable: a 5090 negotiating x8 is supporting evidence for a planned x16 machine, even when its warm runtime happens to be close.
Every new result therefore records, for the GPU and each NVMe endpoint:
- current and maximum PCIe generation/speed and link width;
- the resolved sysfs device path, PCI address and ordered PCIe ancestry;
- the number and identity of upstream bridges;
- shared upstream bridges between every pair of NVMe devices and between each NVMe and the GPU;
- the complete
lspcitree and path-orientedlspci -DPPnnlisting.
The ancestry is evidence, not an automatic CPU-direct/chipset label. Final classification should combine it with the motherboard or platform block diagram. A primary single-versus-RAID comparison must match GPU link width and, as far as possible, CPU socket/NUMA placement. A RAID result whose members share one chipset uplink answers a different question from a RAID result whose members use independent CPU root ports.
PCIe 5.0 x8 and PCIe 4.0 x16 have similar nominal one-direction link bandwidth, but they are not assumed to be application-equivalent. That claim needs a later controlled validation, preferably on the same 5090 host with x16 versus x8 selected through slot placement or BIOS bifurcation. The validation must include pinned host-to-device transfers and both full ComfyUI workloads, not only a PCIe microbenchmark. Until then, every x8 GPU result is supporting-only regardless of its observed warm runtime.
For the planned MSI MAG X870E Tomahawk Max build, the physical acceptance test must be performed with every intended device installed, and the 5090 must negotiate Gen5 x16 under load. In any two-drive comparison, the model drives in M2_1 and M2_2 must each negotiate Gen5 x4 on their CPU paths, the system drive must remain on its chipset path, and M2_2 must be configured for full x4 even though that disables the shared rear USB4 function. The recommended single-drive configuration uses M2_1 and leaves rear USB4 available; only a later justified RAID trial should consume M2_2 and disable USB4. If adding a second model drive reduces the GPU to x8 or puts either RAID member behind a shared chipset uplink, the machine is not the topology this experiment tested.
Session and order controls#
The suite runner must:
- provision and install once;
- download the de-duplicated union of resolved files before the first workload;
- run both complete benchmarks sequentially;
- fetch and persist the first result before starting the second;
- use distinct remote logs and local JSON filenames;
- attach a shared suite/session ID, workload order, setup versions and prepared-file totals to both results;
- refuse to overwrite an existing result.
Workload order is counterbalanced across hosts to expose heating or order effects:
| Host replicate | Order |
|---|---|
| Single A | MiniMax H3, then Krea 2 |
| Single B | Krea 2, then MiniMax H3 |
| RAID A | MiniMax H3, then Krea 2 |
| RAID B | Krea 2, then MiniMax H3 |
Page-cache eviction and process block-read verification remain mandatory for the second workload. Schema-v4 block-layer deltas must also show whether those reads reached the loop, RAID and physical NVMe devices. Storage temperatures and device counters should be captured around each workload so a warmed or throttled SSD is visible in the result.
Analysis#
Results remain separate by workload. For each host and workload calculate:
- median and mean paired
cold wall - warm wall; - cold and warm means, standard deviations and individual samples;
- actual model-file read mean, minimum and coefficient of variation;
- process block-read bytes and effective process block-read rate during each cold run;
- warm-run process block reads;
- for schema-v4 runs, per-layer loop/RAID/NVMe read deltas that determine whether bytes reached physical media;
- CPU-to-GPU bandwidth and warm runtime as host/GPU matching checks;
- negotiated GPU width, storage ancestry, shared bridges, NUMA placement and CPU-direct versus chipset path;
- device temperatures and throttling indicators when exposed.
Do not pool the five samples from different hosts as independent observations. Compare host-level summaries, show every host, and report layout ranges. A RAID claim requires the direction and practical magnitude to replicate across independent machines and both workloads.
Decision thresholds#
A single Gen5 NVMe is considered sufficient when both independent single-drive hosts satisfy all of the following for both workloads:
- all five cold samples pass process block-read verification;
- all warm samples show no material process block reads;
- underlying NVMe counters corroborate the cold read on any run used to make a physical-drive claim;
- the real model-file reader stays around 10 GiB/s or higher without large degradation;
- the median paired cold penalty is within one second or 10% of comparable RAID hosts;
- warm GPU runtime is close enough to rule out a materially different GPU floor.
RAID 0 is recommended only if it repeatedly saves more than approximately two seconds of end-to-end cold
latency on both workloads, or if single-drive throughput is unstable enough to miss the service target.
Higher fio bandwidth without an application-level improvement is not sufficient.
Validation and launch log#
Before parallel acquisition, one dual-workload session will be monitored manually. Validation requires:
- one ComfyUI/PyTorch installation;
- both model sets present before the first measured suite;
- two successful, separate JSON outputs with the same suite ID;
- five cold and five warm samples in each output;
- every cold sample physically verified;
- the first result safely persisted before the second begins;
- no existing benchmark file overwritten;
- expected instance retention or teardown behavior.
| Time | Machine | Layout | Workload order | Outcome |
|---|---|---|---|---|
| 2026-09-23 04:44 UTC | 147512 | 2x 4 TB Crucial T705 Gen5 x4, mdadm RAID 0 |
MiniMax H3, Krea 2 | Suite validation passed; instance 52170548 was destroyed after inspection |
| 2026-09-23 05:12 UTC | 137184 | 1x 2 TB WD Black SN8100 Gen5 x4 | Krea 2, MiniMax H3 | Supporting counterbalanced suite passed; instance 52173137 was destroyed automatically |
| 2026-09-23 05:49 UTC | 112410 | 2x 4 TB WD Black SN850X Gen4 x4, mdadm RAID 0 |
Krea 2, MiniMax H3 | Independent AM5 RAID suite passed; instance 52177233 was destroyed automatically |
| 2026-09-23 06:42 UTC | 55019 | 2x 2 TB WD Black SN8100 Gen5 x4, mdadm RAID 0 |
MiniMax H3, Krea 2 | Supporting same-media RAID suite passed; instance 52183753 was destroyed automatically |
| 2026-09-23 07:13 UTC | 150183 | 1x 2 TB Samsung 9100 Pro Gen5 x4 | MiniMax H3 | Five-pair exact-drive validation passed; verified host; instance 52186892 was destroyed automatically |
| 2026-09-23 07:43 UTC | 148782 | 1x 4 TB Samsung 9100 Pro Gen5 x4 | MiniMax H3 | Five-pair exact-capacity container-path validation passed; supporting unverified host; instance 52190732 was destroyed automatically |
| 2026-09-23 08:16 UTC | 147169 | Same 1x 4 TB Samsung 9100 Pro as the original slow run | MiniMax H3 | Five-pair same-machine repeat passed; block-layer probe proved the model pass was served by an outer loop-backing-file cache |
| 2026-09-23 09:11 UTC | 147011 | 1x Corsair MP700 Elite Gen5 x4 | MiniMax H3 | Five-pair schema-v4 physical-read control passed; every process read propagated 100% to the NVMe; instance 52203128 was destroyed automatically |
After the RAID replicate completed, machines 147169 and 150656—the two initially known dedicated Ryzen 9950X, Gen5 x16, single-4-TB-Samsung-9100-Pro hosts—remained unavailable through 60 consecutive checks over about 32 minutes. A later search deliberately relaxed country, advertised speed and verification filters, while still manually rejecting shared multi-GPU hosts. That found machines 150183 and 148782. The second is supporting rather than primary evidence because Vast marks it unverified; its result is retained with that caveat instead of being discarded.
Results#
Suite validation: RAID host 147512#
The first dual-workload session used one dedicated physical RTX 5090, a Threadripper 7960X, 128 GB RAM
and a PCIe 5.0 x16 GPU link. Physical inspection found a separate 1 TB WD Black SN7100 OS drive and two
4 TB Crucial T705 drives, each negotiating PCIe 5.0 x4. The T705 partitions formed a clean mdadm RAID 0
array with 512 KiB chunks and 2 MiB read-ahead. The PCI tree places the two T705 endpoints on separate root
ports (0000:00:01.1 and 0000:00:01.2) in PCI domain/NUMA side 0, while the GPU is under root port
0000:40:01.1 in domain/NUMA side 40. This is not a shared-chipset-uplink RAID, but the cross-domain storage
and GPU placement is another host-level difference that must remain visible in comparisons.
Both result files share suite ID vast-20260923T044419Z. The first workload downloaded the complete
61,108,590,469-byte de-duplicated file set; the second reported zero download time and found all eight files
ready. ComfyUI was installed once at v0.37.0 with Torch 2.9.1+cu130. The MiniMax result was persisted before
Krea began. The retained validation instance was explicitly destroyed after both results were inspected.
Results: MiniMax H3 JSON and Krea 2 JSON.
| Workload | Cold mean | Warm mean | Cold penalty | Model read | Buffered QD1 | Direct QD1 | Direct QD32 |
|---|---|---|---|---|---|---|---|
| MiniMax H3 | 72.176 s | 63.736 s | 8.440 s | 9.906 GiB/s | 10.957 GB/s | 12.509 GB/s | 27.962 GB/s |
| Krea 2 | 10.081 s | 4.272 s | 5.809 s | 10.390 GiB/s | 9.676 GB/s | 11.531 GB/s | 26.020 GB/s |
Validation invariants passed for both workloads:
- five cold and five warm samples were recorded;
- every cold sample passed the 80% physical-read requirement on its first attempt;
- every warm sample recorded zero physical reads;
- the files use separate names and neither replaced an older benchmark;
- suite metadata records one setup invocation, workload order and the shared prepared-file set.
NVMe SMART temperature data was unavailable because the container exposed NVMe sysfs topology but not the
corresponding /dev/nvme*n* device nodes. The failed command and error are retained in both JSON files
rather than silently treated as a temperature reading.
These values validate the runner but do not resolve the storage decision. Machine 147512 is one RAID host, and its Threadripper platform and cross-domain GPU/storage paths are not a matched replacement for the planned Ryzen 9950X system.
Supporting single-drive host 137184#
The counterbalanced session used a Ryzen 9950X and one WD Black SN8100 negotiating Gen5 x4. The SSD and
tested GPU occupy separate CPU root ports (0000:00:01.2 and 0000:00:01.3) with no chipset bridge in
either visible path. The physical host contains two 5090s and the tested GPU negotiated Gen5 x8, however,
so this host cannot be a primary match for the intended one-GPU x16 build. Its median pinned host-to-device
transfer was 26.89 GiB/s, versus 48.86-49.65 GiB/s on the x16 RAID host. This large transfer difference did
not produce a meaningful warm-runtime difference in either tested workflow, but that observation does not
replace a controlled x8-versus-x16 validation.
The single drive's 10.19-11.00 GiB/s real model reads also remained well below the measured 26.89 GiB/s host-to-device transfer rate. That makes the x8 link unlikely to be the immediate ceiling on its storage-fed load path and, directionally, an x16 production link should not make the single-drive result worse. Model loading includes CPU transforms, allocation and synchronization, however, so this is an inference rather than a substitute for the matched x16 run.
Both result files share suite ID vast-20260923T051235Z. Krea ran first and MiniMax second, the complete
model union was ready before Krea, MiniMax downloaded zero additional bytes, and all cold samples passed on
their first attempt. Every warm sample recorded zero physical reads.
Results: MiniMax H3 JSON and Krea 2 JSON.
| Workload | Cold mean | Warm mean | Cold penalty | Model read | Buffered QD1 | Direct QD1 | Direct QD32 |
|---|---|---|---|---|---|---|---|
| MiniMax H3 | 71.255 s | 63.762 s | 7.493 s | 10.190 GiB/s | 12.549 GB/s | 9.177 GB/s | 14.906 GB/s |
| Krea 2 | 9.463 s | 4.263 s | 5.200 s | 11.003 GiB/s | 13.110 GB/s | 9.196 GB/s | 14.913 GB/s |
First cross-host comparison#
| Workload | RAID minus single cold | RAID minus single warm | RAID minus single penalty | RAID minus single model read |
|---|---|---|---|---|
| MiniMax H3 | +0.921 s | -0.026 s | +0.947 s | -0.284 GiB/s |
| Krea 2 | +0.618 s | +0.009 s | +0.609 s | -0.613 GiB/s |
Positive latency values mean the two-drive RAID host was slower. In this first comparison the single drive
was faster on both cold means and actual model-file reads, despite the RAID host reaching 26.0-28.0 GB/s in
direct QD32 fio. Warm times were effectively identical. This is strong evidence that the RAID array's
headline parallel throughput did not help these application paths, but it is not yet a conclusive purchase
result: there is only one host per layout, their CPUs and PCIe paths differ, and the single host's GPU is x8.
Independent AM5 RAID host 112410#
The second RAID suite used a Ryzen 9950X-class CPU, 128 GB RAM and a dedicated 575 W 5090 negotiating
Gen5 x16. Its two 4 TB SN850X RAID members each negotiate Gen4 x4 on independent AM5 CPU root ports:
0000:00:01.2 and 0000:00:03.1. The new ancestry capture confirms that the members share no upstream PCI
bridge with each other or the GPU. A separate 1 TB SN850X system device traverses three upstream bridges
behind the chipset. This is the closest measured topology to the intended build, although its RAID media are
Gen4 rather than Samsung 9100 Pro Gen5 drives.
Both result files share suite ID vast-20260923T054900Z. The runner installed ComfyUI v0.37.0 and Torch
2.9.1+cu130 once, prepared the complete model union before Krea, persisted Krea before MiniMax, and downloaded
zero additional bytes for MiniMax. All ten cold samples passed on their first attempt and every warm sample
recorded zero physical reads. The automatically provisioned instance was destroyed after both JSON files
were saved.
Results: MiniMax H3 JSON and Krea 2 JSON.
| Workload | Cold mean | Warm mean | Cold penalty | Model read | Buffered QD1 | Direct QD1 | Direct QD32 |
|---|---|---|---|---|---|---|---|
| MiniMax H3 | 70.842 s | 65.458 s | 5.384 s | 11.380 GiB/s | 13.760 GB/s | 8.100 GB/s | 14.451 GB/s |
| Krea 2 | 8.664 s | 4.265 s | 4.399 s | 11.488 GiB/s | 13.678 GB/s | 8.526 GB/s | 14.445 GB/s |
The Krea means reproduce this same host's earlier independent v0.37.0 result to 15 ms cold and exactly at the reported millisecond warm mean. That repeatability makes the large host-to-host difference between the two RAID machines a real platform/path effect rather than five-sample noise.
Supporting same-media RAID host 55019#
Machine 55019 adds a useful negative control. Physical inspection found two 2 TB WD Black SN8100 drives,
each negotiating Gen5 x4 on a separate root port, combined into a clean two-member mdadm RAID 0 with
512 KiB chunks and 2 MiB read-ahead. The array reached 24.28 GiB/s direct QD32 on MiniMax H3. Actual model
reads nevertheless reached only 4.32 GiB/s and the cold penalty was 17.10 seconds. Krea 2 reproduced the
same 4.30 GiB/s model-read ceiling and an 8.84-second cold penalty.
Results: MiniMax H3 JSON and Krea 2 JSON.
| Workload | Cold mean | Warm mean | Cold penalty | Model read | Buffered QD1 | Direct QD1 | Direct QD32 |
|---|---|---|---|---|---|---|---|
| MiniMax H3 | 91.645 s | 74.545 s | 17.100 s | 4.323 GiB/s | 4.295 GiB/s | 8.459 GiB/s | 24.278 GiB/s |
| Krea 2 | 13.870 s | 5.027 s | 8.843 s | 4.296 GiB/s | 4.685 GiB/s | 9.394 GiB/s | 26.916 GiB/s |
All ten cold samples passed physical-read verification on their first attempt and all warm samples read zero model bytes. The host is supporting evidence only: it is a four-GPU EPYC 9135 system, the offer allocated one quarter of the host's GPU/CPU/RAM resources, and the selected GPU was limited to 450 W. Its two storage endpoints and selected GPU also reside in different PCI domains. Those differences prevent a causal single-versus-RAID comparison with machine 137184. They do, however, prove that two matching SN8100 drives and very high direct-QD32 throughput are not sufficient to produce fast buffered model reads or a small ComfyUI cold penalty.
Exact single Samsung 9100 Pro MiniMax validation#
Two later MiniMax H3 runs validate the exact SSD family without replacing any earlier host result. Machine 150183 is a verified, dedicated 575 W RTX 5090 host with a Ryzen 9950X3D, 64 GB RAM, Gen5 x16 GPU and one 2 TB 9100 Pro on a direct CPU Gen5 x4 root port. Machine 148782 is a separate, dedicated 600 W RTX 5090 host with a Ryzen 9950X3D, approximately 96 GB RAM, Gen5 x16 GPU and one 4 TB 9100 Pro on another direct CPU Gen5 x4 root port. Vast marks machine 148782 unverified but not VM-deverified, so it is supporting evidence.
Results: 2 TB machine 150183 JSON and 4 TB machine 148782 JSON.
| Host and drive | Cold mean | Warm mean | Cold penalty | Model read | Direct QD1 | Direct QD32 |
|---|---|---|---|---|---|---|
| 150183, single 2 TB 9100 Pro | 71.086 s | 62.266 s | 8.820 s | 9.581 GiB/s | 10.064 GB/s | 14.751 GB/s |
| 148782, single 4 TB 9100 Pro | 66.767 s | 61.035 s | 5.732 s | 7.585 GiB/s | 9.171 GB/s | 11.326 GB/s |
Both results contain five cold/five warm pairs with no rejected attempts. Every cold pass caused
42.72-42.98 GB of process block reads, and every warm pass caused zero model reads apart from an immaterial 1.39 MB in machine
150183's first warm sample. The drives have distinct serial numbers, the same 0B2QNXH7 firmware, and
different motherboard/kernel environments. A single 9100 Pro therefore does not have the approximately
5 GiB/s ceiling in the observed Vast container paths, and 4 TB capacity is not sufficient to explain that
anomaly. Because these schema-v3 runs did not corroborate the reads at the underlying NVMe layer, this is not
a direct physical-drive throughput claim.
The 4 TB host's separate sequential reader was slower than the 2 TB host's, yet its full-workload cold
penalty was 3.088 seconds smaller. This reinforces that fio and a plain sequential reader are diagnostics,
not substitutes for the verified ComfyUI cold/warm experiment. The source-level investigation of ComfyUI's
actual eight-worker, per-tensor file path and the unresolved old-host hypotheses are recorded in the
Samsung 9100 Pro investigation journal.
Schema-v4 physical single-drive control#
Machine 147011 adds the physical-layer control missing from the schema-v3 matrix. It used a dedicated 600 W RTX 5090, Ryzen 9950X, PCIe 5.0 x16 GPU and one Corsair MP700 Elite on a direct CPU PCIe 5.0 x4 root port. Unlike machine 147169, the Docker overlay path had no active loop-backing device; cgroup I/O attributed the path directly to the physical NVMe namespace.
| Cold mean | Warm mean | Cold penalty | Model read | Direct QD1 | Direct QD32 |
|---|---|---|---|---|---|
| 70.018 s | 62.716 s | 7.302 s | 4.536 GiB/s | 5.172 GB/s | 9.280 GB/s |
All five cold passes moved 42.69-42.85 GB at both the process and NVMe layers, a 100.0% propagation ratio. All three independent model-reader passes moved their full 42.472 GB through the NVMe as well. Warm passes recorded zero reads except one 4 KiB event. This proves the cache-eviction protocol creates a genuinely physical cold read when the model filesystem maps directly to the device.
The result is particularly relevant to the RAID decision because this single physical drive reached only 4.54 GiB/s in the real reader yet still landed at a 7.30-second application penalty, inside the existing single/RAID overlap. It does not substitute for a direct 9100 Pro result, but it disproves the assumption that a single drive must exceed 10 GiB/s to produce an acceptable MiniMax cold start.
Schema-v4 physical Samsung 9100 Pro control#
Machine 62211 closes the exact-drive gap with one physical 4 TB Samsung 9100 Pro. The drive negotiated PCIe 5.0 x4 on a direct root port, while the dedicated RTX 5090 negotiated PCIe 5.0 x16 on a separate root port. There was no MD array and no active loop device in the benchmark cgroup path.
| Cold mean | Warm mean | Cold penalty | Model read | Buffered | Direct QD1 | Direct QD32 |
|---|---|---|---|---|---|---|
| 81.762 s | 69.986 s | 11.776 s | 5.699 GiB/s | 7.137 GB/s | 7.811 GB/s | 9.868 GB/s |
All five cold passes and all three model readers propagated 100.0% of their process block bytes to the Samsung namespace. The paired penalty was stable (0.160-second sample SD), so this is a valid physical single-drive result. It is not a drive-only comparison with machine 147011: machine 62211 used a 500 W GPU, Intel Core Ultra 7 265K, Linux 6.8 and 44.743 GiB/s H2D, while machine 147011 used a 600 W GPU, Ryzen 9950X, Linux 6.17 and 51.220 GiB/s H2D. That slower platform had a larger penalty even though its separate model reader was faster than the MP700's. The result therefore proves physical sufficiency and cache eviction, not that the 9100 Pro is slower than the MP700.
Completed dual-workload v0.37.0 matrix#
| Layout and host | GPU link | MiniMax cold / warm / penalty | Krea cold / warm / penalty | MiniMax / Krea model read |
|---|---|---|---|---|
| Single SN8100, 137184 | Gen5 x8 | 71.255 / 63.762 / 7.493 s | 9.463 / 4.263 / 5.200 s | 10.190 / 11.003 GiB/s |
| RAID T705, 147512 | Gen5 x16 | 72.176 / 63.736 / 8.440 s | 10.081 / 4.272 / 5.809 s | 9.906 / 10.390 GiB/s |
| RAID SN850X, 112410 | Gen5 x16 | 70.842 / 65.458 / 5.384 s | 8.664 / 4.265 / 4.399 s | 11.380 / 11.488 GiB/s |
The single-drive result falls inside the two independent RAID hosts' range for cold time, cold penalty and real model-read rate on both workloads. Across the RAID hosts, Krea cold time varies by 1.417 s and MiniMax cold time by 1.334 s despite both using explicit two-drive arrays. The storage-layout effect, if any, is therefore smaller than ordinary cross-host platform/path variation in the completed matrix.
Within each host, the paired cold-minus-warm sample standard deviation was small relative to those cross-host differences:
| Layout and host | MiniMax penalty SD | Krea penalty SD |
|---|---|---|
| Single SN8100, 137184 | 0.427 s | 0.086 s |
| RAID T705, 147512 | 0.278 s | 0.097 s |
| RAID SN850X, 112410 | 0.380 s | 0.113 s |
Same-machine schema-v4 RAID replication#
Machine 112410 was later reacquired with the same GPU UUID, motherboard, CPU, kernel, drive serials, MD layout and 575 W power limit. The full MiniMax run reproduced the earlier schema-v3 result almost exactly:
| Capture | Cold mean | Warm mean | Cold penalty | Model read |
|---|---|---|---|---|
| Original schema v3 | 70.842 s | 65.458 s | 5.384 s | 11.380 GiB/s |
| Schema-v4 replication | 70.823 s | 65.404 s | 5.419 s | 11.504 GiB/s |
Schema v4 showed no active loop backing. Every cold process byte reached md0 and then split essentially
50/50 across the two physical SN850X members. Each of three independent 42.472 GB reader passes likewise
produced 42.472 GB at md0 and 21.236 GB at each member. This proves that the observed RAID path was
physical and the result is repeatable; it does not reveal what this same host would do with one member.
Adjacent single-SN850X host: accelerator-path confound#
Machine 132465 provided one physical 4 TB SN850X, but it did not provide a clean comparison with machine 112410's two-drive SN850X RAID. Its MiniMax result was 81.476 seconds cold, 61.122 seconds warm, a 20.354-second paired penalty and 5.300 GiB/s in the model reader. Direct QD1/QD32 were 3.549/6.956 GB/s.
The decisive exclusion is the accelerator path: pinned H2D measured only 3.359 GiB/s, agreeing with Vast's 3.3 GB/s PCIe score, versus 45.682 GiB/s on the RAID host. The single host's 61.122-second warm floor was still fast because the resident model avoided that cold transfer path. This shows why warm runtime, GPU power and a nominal x16 lane label are not enough to match cold-load hosts. Disk rate, H2D rate and CPU/ platform behavior can all bind the same cold phase.
The host lacked cgroup io.stat, so its block evidence is kept in a lower class. Host-wide SN850X counters
matched each approximately 39.8 GiB cold process read and each 39.555 GiB reader pass, while warm deltas were
zero apart from 0.0001 GiB once. That strongly supports physical propagation on this dedicated machine but
is not process-attributed proof. The result is retained as an adjacent multivariable observation and must
not be used to claim that RAID saves 14.935 seconds over one SN850X.
Feasibility correction: the missing crossover is not a normal Vast experiment#
The desired causal experiment remains a same-machine transition between one drive and two-drive RAID 0, but the captured Vast inventory does not expose a safe renter-controlled host on which to perform it. All captured hosts with a matched NVMe pair already had that pair assembled by the provider as MD RAID. The one multi-NVMe host without MD RAID, machine 150183, had mismatched 2 TB Samsung and 1 TB Seagate devices rather than two identical unused drives.
Machine 112410 illustrates the boundary. Its two 4 TB SN850X devices form a clean, conventional Linux MD
RAID 0 with 512 KiB chunks, while a separate 1 TB SN850X serves system duties. The benchmark container's root
is OverlayFS with its upper/work path under /var/lib/docker, and schema-v4 cgroup counters show workload
bytes flowing through md0 and splitting across both members. The array is therefore real and its MD layer
is comparable to the software RAID that could be chosen locally, but it is provider-owned infrastructure
underneath the container. Dismantling it would dismantle the storage path that serves the running instance,
not switch an isolated benchmark volume.
Vast's FAQ documents its normal instance model as an unprivileged Docker
container; its
instance guide says
Docker-in-Docker and renter bare-metal access are not supported. Seeing host topology, /sys entries, NVMe
metadata or block counters does not grant the raw-device and mount privileges required to stop, format or
rebuild the host array. Container root is not host root; Docker's
runtime privilege reference
shows that device access and system-administration capabilities must be granted explicitly.
This changes the experiment plan, not the observed results: future Vast runs can add independent preconfigured hosts and improve external confidence, but they cannot produce the missing causal crossover unless a cooperative provider dedicates two unused identical drives and performs both layouts from the host side. A true crossover instead needs administrative bare-metal control. The planned machine's eventual direct-filesystem before/after test remains the decisive test of the second SSD's incremental value.
Final finding and recommendation#
The completed fixed v0.37.0 Vast runs do not show a practical RAID 0 advantage in the end-to-end container paths tested. The single SN8100 lands inside the two RAID hosts' observed range on both full workloads, while the RAID layouts range from faster than the single host to slower than it. Neither workload shows the predeclared, repeated improvement of more than two seconds. The fastest synthetic array result—approximately 28 GB/s direct QD32—did not produce the fastest model reads or cold starts. The two later exact-9100 MiniMax runs independently show that one 9100 Pro can also produce a small container-path cold penalty without striping.
The loop-backing-file discovery narrows what the original matrix proves. It supports “RAID did not improve these Vast application runs,” but not “one physical 9100 Pro is proven equal to two direct-attached 9100 Pros.” The older collector evicted the inner model files but could not evict or measure the outer host cache. That cache may conceal physical-device differences in either direction.
The later same-machine schema-v4 replication removes that ambiguity for the SN850X RAID host specifically: its bytes reached both physical members and its 5.419-second penalty reproduced the original 5.384-second result. Against the unmatched physical single MP700 host, its reader was 2.54x faster but its penalty was only 1.883 seconds smaller and raw cold time was 0.805 seconds slower because the warm floor differed. That is useful supporting evidence, but it is still not the same-host single-to-RAID intervention required to price the second SSD's incremental benefit.
The two schema-v4 controls establish a narrower physical claim: one direct-attached Gen5 x4 NVMe can serve every cold byte from physical media and complete MiniMax repeatably without RAID. The MP700 did so with a 7.30-second penalty at a 4.54 GiB/s reader rate; the exact 4 TB 9100 Pro did so with an 11.78-second penalty at 5.70 GiB/s on a materially slower 500 W/Intel platform. This strengthens “start with one drive” and removes the need to treat 10 GiB/s as a purchase or acceptance threshold.
For the equal quoted total cost of one 4 TB Samsung 9100 Pro versus two 2 TB devices, the current purchase recommendation remains start with one 4 TB 9100 Pro, without RAID 0 initially. This is a reversible, evidence-disciplined initial configuration, not a claim that Vast conclusively measured the bare drive. It preserves the second CPU-connected M.2 path and rear USB4 function, avoids array/topology/thermal complexity, and has a smaller cache-loss surface. The evidence does not justify consuming those resources for synthetic throughput that has not translated into application latency.
The later deployment constraint permits only 2 TB NVMe devices, so the operational application is now start with one 2 TB 9100 Pro when the active cache fits. The experiment's performance conclusion is unchanged; only the capacity ceiling is lower. A second 2 TB Samsung should be treated first as a 4 TB cache capacity purchase and only secondarily as a RAID-performance purchase unless the local A/B test proves the latency benefit.
The later single-SN850X observation does not reverse that conclusion. Its slow cold result arrived with both a slower storage path and an anomalous 3.359 GiB/s H2D path, so it is evidence for stricter host matching—not evidence for a RAID effect. A valid storage comparison must now admit or stratify by measured H2D bandwidth as well as GPU power, warm floor, CPU and physical block topology.
For the later 2 TB-only deployment, install the initial single 2 TB device in M2_1. If a 4 TB model is ever allowed and selected instead, the same acceptance checks apply. Accept the physical node only after checking all of the following with every intended device present:
- The 9100 Pro negotiates Gen5 x4 on a direct CPU path.
- The RTX 5090 negotiates Gen5 x16 under load.
- The system SSD is on its intended chipset path and does not share the model-drive path.
- Models reside on a direct filesystem on the 9100 Pro rather than a buffered loop image.
- The real model-file reader repeatedly sustains the required rate after page-cache eviction, while the 9100 Pro's own namespace counters increase by approximately the resolved model byte count.
- Both full five-pair workloads reproduce acceptable cold times, with every cold run verified at both the process and NVMe layers and every warm run reading zero model bytes.
If the physical 9100 Pro build misses that storage target, diagnose slot placement, negotiated widths,
filesystem/container path, thermal throttling and CPU-to-GPU transfer first. Add a second drive only after a
same-host before/after test on the local bare-metal node demonstrates a repeatable application improvement—not
merely a higher fio number—and preferably clears the same two-second threshold on both workloads. That
test should copy the same model set and filesystem between single-drive and RAID configurations, preserve
the GPU/CPU/PCIe topology, and record physical NVMe counters so the comparison cannot be satisfied by an
outer cache.
The exact 4 TB 9100 Pro has now been observed with the heavier MiniMax workload on multiple physical hosts, including one schema-v4 direct-filesystem path where every model byte reached the physical Samsung namespace. This is still a containerized host rather than the planned local bare-metal node, so local acceptance remains necessary. An exact-drive Krea run was not collected because the later acquisitions intentionally focused on the open MiniMax anomaly rather than reinstalling the dual-workload suite. That narrower remaining gap does not change the RAID decision: RAID failed to show a repeated benefit across two complete application workloads; single 2 TB and 4 TB 9100 Pro container paths avoided the old failure mode; and a different single Gen5 drive has now delivered the complete workload with physical-NVMe reads corroborated.
This recommendation supersedes the earlier provisional two-drive recommendation in model-read-storage-layouts.md, which was based on pre-matrix results where storage layout and host effects were not separated as rigorously.