RTX 5090 multivariable investigation journal#
Status: Journal Started: 2026-09-23 Latest entry: 2026-09-27 Primary slice: MiniMax H3, ComfyUI v0.37.0, resident cache, five verified cold/warm pairs
This journal tracks the variables that can change an RTX 5090 result so cross-host differences are not
misattributed to storage. Raw factors are normalized by
rtx5090_minimax_factors.sql, and physical block propagation is
normalized by rtx5090_minimax_block_layers.sql. The
decision-level synthesis is maintained in
rtx5090-raid0-broader-evidence-analysis.md.
Analysis model#
The physical machine is the independent unit. A repeat of one machine is a longitudinal control, not another independent host. The end-to-end time is separated conceptually as:
cold wall time = warm GPU/platform floor + paired storage/container-path penalty
The saved factors include actual and advertised GPU power, GPU identity/VBIOS/driver/link, CPU and logical
allocation, motherboard, RAM/cgroup/swap, kernel, container implementation, exact NVMe identity/firmware/
link/PCI ancestry, MD layout, active loop backing, model reader, buffered/direct fio, H2D, read-active
duration/rate, CPU use and major faults. Schema v4 additionally measures whether process block reads reach
loop, MD and physical NVMe layers.
Hardware observations copied from a later capture of the same physical machine carry an explicit
hardware_metadata_backfill.provenance record. Stable identity/topology is kept separate from mutable
runtime configuration. Existing backfills cover old results from machines 112410, 147169 and 150656.
Prior exploratory checkpoint#
This checkpoint predates the seven-suite expansion documented at the end of this journal. After the single-SN850X control, there were eleven full v0.37.0 RTX 5090/MiniMax host results. Their warm runtime strata show why raw cold time is not a storage metric:
| Actual GPU power | Hosts | Warm mean | Range | Mean cold penalty |
|---|---|---|---|---|
| 420 W | 1 | 75.626 s | 75.626 s | 30.003 s |
| 450 W | 1 | 74.545 s | 74.545 s | 17.100 s |
| 500 W | 1 | 69.986 s | 69.986 s | 11.776 s |
| 575 W | 4 | 63.792 s | 62.266-65.404 s | 7.543 s |
| 600 W | 4 | 61.300 s | 60.328-62.716 s | 10.877 s |
Exploratory Pearson correlations across those eleven heterogeneous hosts are:
| Variables | Correlation |
|---|---|
| Warm runtime vs actual GPU power | -0.9840 |
| Cold penalty vs read-active duration | +0.9881 |
| Cold penalty vs model-reader rate | -0.7499 |
Cold penalty vs direct QD1 fio |
-0.6264 |
Cold penalty vs direct QD32 fio |
-0.3623 |
| Cold penalty vs pinned H2D rate | -0.1120 |
These associations are descriptive, not causal estimates. The sample is small and power, CPU, kernel, storage, accelerator transfer path, container path and motherboard are highly collinear. A many-coefficient regression would overfit. The high-value evidence is instead paired outcomes, explicit stratification and same-machine interventions.
The aggregate H2D correlation is not evidence that accelerator transfer is irrelevant. Ten hosts occupy a wide 26.894-53.896 GiB/s band, while one new host is isolated at 3.359 GiB/s; a linear coefficient across power strata is the wrong model for a possible low-bandwidth threshold. Among the seven 575-600 W hosts with H2D at least 10 GiB/s, two MD-array paths average a 6.930-second penalty and five single/container paths average 7.893 seconds. Their ranges overlap (5.419-8.440 versus 5.732-10.120 seconds), and the schema-v3 paths do not prove physical-media reads. This is not a RAID effect estimate.
Stratified patterns worth retaining#
Restricting the slice to the seven 575-600 W hosts with measured H2D at least 10 GiB/s removes the clear power-limited hosts and the new transfer-path outlier. Within that still-heterogeneous group:
| Variables | Pearson correlation |
|---|---|
| Cold penalty vs full-workload read-active duration | +0.9946 |
| Cold penalty vs separate model-reader rate | -0.2480 |
Cold penalty vs direct QD1 fio |
+0.6089 |
| Cold penalty vs pinned H2D | +0.0550 |
The read-active duration is part of the full workload trace, so its strong relationship is mechanically plausible rather than an independent causal estimate. More importantly, the synthetic and standalone throughput metrics stop predicting penalty reliably after basic platform admission. This is why the full paired workload remains the decision metric.
The closest observational pair currently available is useful despite not being a controlled intervention:
| Factor | RAID machine 112410 | Single machine 148782 |
|---|---|---|
| CPU family | Ryzen 9950X3D | Ryzen 9950X3D |
| Kernel | 6.8.0-138 | 6.8.0-138 |
| GPU topology / H2D | Gen5 x16 / 45.682 GiB/s | Gen5 x16 / 46.747 GiB/s |
| GPU power | 575 W | 600 W |
| Storage | 2x SN850X RAID 0 | 1x Samsung 9100 Pro |
| Model reader | 11.504 GiB/s | 7.585 GiB/s |
| Warm / penalty / cold | 65.404 / 5.419 / 70.823 s | 61.035 / 5.732 / 66.767 s |
The RAID reader is 1.52x faster, yet its paired penalty is only 0.313 seconds smaller—well below the predeclared two-second application threshold—and its raw cold time is 4.056 seconds slower because its warm floor is 4.369 seconds slower. The comparison remains cross-host and the single side is schema v3, so it is supporting evidence rather than a RAID effect estimate. It does show how a large reader-rate difference can collapse to a small full-application difference.
Other apparent categories are too sparse or collinear for attribution. Ryzen 9950X3D hosts average a 6.657-second penalty versus 8.305 seconds for Ryzen 9950X hosts in the admitted slice, but each group has only three machines and their storage/kernel mix differs. Only kernel 6.8.0-138 has two admitted hosts; every other kernel is represented once. There is one containerd v0.37 host, one x8 GPU host and one H2D under-10-GiB/s host. These are columns to preserve and future matching targets, not effects to claim.
The same-machine metadata audit found only machines 112410, 147169 and 150656 with later captures capable of informing older runs. Their provenance-labelled backfills already exist. No new old run has a safe same-machine source, so missing PCI ancestry remains null rather than being inferred from a similar board or drive model.
2026-09-23: exact physical single-drive result#
Machine 62211 used one 4 TB Samsung 9100 Pro on PCIe 5.0 x4 and a dedicated PCIe 5.0 x16 RTX 5090. Every one of five cold passes and three model readers propagated 100.0% of process block bytes to the Samsung namespace. Its primary results were 81.762 seconds cold, 69.986 seconds warm, 11.776 seconds paired penalty, 5.699 GiB/s model reader and 44.743 GiB/s H2D. This is an exact-drive physical storage control, not a warm-performance baseline: the GPU was limited to 500 W and the host used an Intel Core Ultra 7 265K.
2026-09-23: same-machine RAID replication#
Vast machine 112410 was reacquired as instance 52211104. This is the exact verified, dedicated Texas host used by the existing v0.37.0 RAID result: Ryzen 9950X3D2, 575 W RTX 5090, PCIe 5.0 x16 GPU and two 4 TB WD Black SN850X members in an MD array. The new run used schema v4 and the full five-pair protocol. Result: machine 112410 replication JSON.
It answered both questions without changing machines:
- The earlier 70.842 / 65.458 / 5.384-second cold/warm/penalty result repeated at 70.823 / 65.404 / 5.419 seconds. The differences were -0.019 / -0.054 / +0.035 seconds.
- Every cold and model-reader byte reached
md0, then divided essentially 50.00/50.00 between the two physical SN850X namespaces. No loop device appeared in the cgroup I/O path.
| Metric | Original schema v3 | Schema-v4 replication |
|---|---|---|
| Cold mean | 70.842 s | 70.823 s |
| Warm mean | 65.458 s | 65.404 s |
| Paired penalty | 5.384 s | 5.419 s |
| Model reader | 11.380 GiB/s | 11.504 GiB/s |
Buffered fio |
13.772 GB/s | 13.746 GB/s |
| Direct QD1 / QD32 | 8.187 / 14.451 GB/s | 8.374 / 14.461 GB/s |
| H2D | 45.291 GiB/s | 45.682 GiB/s |
The new paired-penalty sample SD was 0.481 seconds. Cold process reads were 46.329-46.680 GB in both the
old and new captures, not a new schema-v4 artifact. The independent reader remained exactly 42.472 GB.
For each reader pass, md0 read 42.472 GB and each physical member read 21.236 GB. The RAID path and cache
eviction are therefore physically corroborated.
This establishes that machine 112410 really uses both drives and that its strong result is repeatable. It still does not measure the incremental RAID benefit: there is no same-host, same-filesystem single-member run. Relative to the physical single MP700 control, the RAID reader was 2.54x faster (11.504 versus 4.536 GiB/s) but the paired penalty was only 1.883 seconds smaller (5.419 versus 7.302 seconds), while raw cold time was 0.805 seconds slower because the RAID host had a slower warm floor. This cross-host comparison is descriptive and falls just short of the experiment's predeclared two-second application threshold.
The single-SN8100 machine 137184 was also available, but only at gpu_frac=0.5; it was rejected rather
than weakening the dedicated-host rule. The two-drive T705 machine 147512 was unavailable. The remaining
missing PCI ancestry on their schema-v3 files cannot be reconstructed honestly without a same-machine
capture and is left null rather than guessed. Instance 52211104 was destroyed after result persistence.
2026-09-23: adjacent single-SN850X control#
Vast machine 132465 was acquired as instance 52214899 to add the missing single-drive side of the SN850X family. The verified dedicated Maryland host had one 4 TB SN850X on a direct PCIe 4.0 x4 root, a 600 W RTX 5090 in the CPU x16 slot, an i9-12900K, 64 GB RAM and no MD array. The full five-pair run completed without rejected attempts and the instance was confirmed destroyed. Result: machine 132465 JSON.
| Metric | Result |
|---|---|
| Cold mean | 81.476 s |
| Warm mean | 61.122 s |
| Paired cold penalty | 20.354 s |
| Model reader | 5.300 GiB/s |
Buffered fio |
4.922 GB/s |
Direct QD1 / QD32 fio |
3.549 / 6.956 GB/s |
| Pinned H2D | 3.359 GiB/s |
The H2D result reproduces Vast's 3.3 GB/s PCIe score despite the slot topology and maximum-capability metadata saying Gen5 x16. It is an actual transfer measurement, whereas the captured current Gen1 state is an idle power-management snapshot. Among the other three 600 W hosts, H2D is 46.747-51.220 GiB/s and the mean cold penalty is 7.718 seconds. Machine 132465 has essentially the same warm floor but a penalty 12.636 seconds larger. Its direct QD1 storage and model reader are also slower, so this run cannot isolate whether disk, host-to-device transfer, CPU/platform behavior or their interaction caused the long cold path. It is therefore not a valid single-SN850X versus dual-SN850X RAID estimate.
This host did not expose cgroup io.stat. Schema v4 did capture host-wide /sys/block counters. On the
otherwise dedicated host, each cold run's SN850X delta matched the process block-read count to within
0.0002 GiB (39.802-39.965 GiB), each 39.555 GiB reader produced a 39.556-39.557 GiB device delta, and warm
runs produced zero except 0.0001 GiB once. This is strong supporting evidence that bytes reached the only
visible NVMe, but remains labelled host-wide, unattributed because unrelated host I/O could enter those
counters. It is not promoted to the same evidence class as machine 112410's cgroup-attributed MD/member
deltas.
The important contribution is methodological: a nominal 600 W, Gen5 x16 RTX 5090 can retain a fast warm runtime while a severely restricted measured H2D path coincides with much slower cold loading. Future slices must retain H2D as an admission/matching variable rather than grouping machines only by GPU power and lane label.
2026-09-23: heterogeneous-host sampling replaces the matched-host search#
The experiment no longer treats close hardware matching as a prerequisite or represents cross-host pairs as
if they were controlled interventions. Normal Vast renters cannot reconfigure the host block devices, and a
search for nominally matched machines discards useful variation while still leaving unobserved differences.
The revised design treats each physical RTX 5090 host as an observational unit and uses its paired
cold wall - warm wall latency as the primary response.
All measured differences remain covariates rather than admission filters: MD layout and member count, exact NVMe identity, negotiated drive generation/width, PCI ancestry and shared bridges, GPU generation/width, actual power limit, H2D, CPU/RAM, kernel, overlay/loop path, filesystem behavior and physical-read evidence. The paired response subtracts much of the warm GPU floor but does not make GPU, CPU or transfer-path effects impossible; those fields remain available for stratification, threshold checks and interactions. The goal is to estimate distributions and look for patterns that reproduce across hosts, not to fit a large causal model to a small sample.
Same-machine reacquisitions have a separate role. They estimate session-to-session and mutable-runtime variance and upgrade older schema-v3 results with schema-v4 block-layer evidence. They are linked to the same machine ID and are not counted as additional independent hosts.
Seven dual-workload v0.37.0 suites were launched with five verified cold/warm pairs per workload:
| Role | Machine | Advertised distinguishing factors at acquisition | Final state |
|---|---|---|---|
| New host | 56897 | Dedicated, Ryzen 7900, 600 W, Gen5 x16, Netac 1 TB | Complete; instance destroyed |
| New host | 29706 | Dedicated, i5-12400F, 600 W, advertised Gen5 x16, SN850X 2 TB | Complete; instance destroyed |
| New host | 30597 | One GPU from two, Threadripper Pro 7975WX, 575 W, Gen5 x16, T700 4 TB | Complete; instance destroyed |
| New host | 30024 | One GPU from eight, EPYC 9654, advertised 600 W, Gen5 x16, Kioxia, 17.4 GB/s advertised disk | Complete; instance destroyed |
| Schema-v3 repeat | 108899 | Dedicated, previous MiniMax-only single/container-path result | Complete; instance destroyed |
| Schema-v3 repeat | 137184 | One GPU from two, previous dual-workload SN8100/x8 result | Complete; instance destroyed |
| Schema-v3 repeat | 150337 | Dedicated, previous Krea-only result | Complete; instance destroyed |
Schema-v3 RAID machine 33260 was also selected, but the runner rejected it before provisioning because Vast now marks the offer VM-deverified. That safety gate was not bypassed. The other previously sampled schema-v3 machines were unavailable at this check.
2026-09-23: seven-suite completion#
All seven suites exited successfully. Each produced separate Krea 2 and MiniMax H3 schema-v4 JSONs, all five cold attempts in every result passed the cache-eviction/read verification gate, and the runner destroyed every Vast instance after persisting the results. The files are timestamped and do not replace the older schema-v3 observations.
| Machine | Observed storage path | GPU path | MiniMax cold / warm / penalty | Krea cold / warm / penalty | MiniMax model reader |
|---|---|---|---|---|---|
| 56897 | Single Netac 1 TB, PCIe 4.0 x4; cgroup physical | 600 W, Gen5 x16 | 73.896 / 60.403 / 13.492 s | 11.984 / 4.015 / 7.969 s | 3.013 GiB/s |
| 29706 | Single SN850X 2 TB, PCIe 4.0 x4; host-wide physical | 600 W, Gen5 x16 | 77.688 / 63.151 / 14.536 s | 14.393 / 4.272 / 10.121 s | 3.040 GiB/s |
| 30597 | Single T700 4 TB, PCIe 5.0 x4; cgroup physical | 575 W, Gen5 x16 | 91.531 / 82.371 / 9.160 s | 12.069 / 5.529 / 6.540 s | 6.264 GiB/s |
| 30024 | Four Kioxia drives in RAID10 plus two RAID1 arrays; host-wide physical | 575 W measured, Gen5 x16 | 97.375 / 67.072 / 30.302 s | 17.264 / 4.289 / 12.976 s | 5.164 GiB/s |
| 108899 | P310 on PCIe 4.0 x2 behind three upstream hops and a buffered Docker loop; cgroup layers traced | 420 W, Gen5 x16 | 104.692 / 75.979 / 28.713 s | 18.083 / 5.021 / 13.062 s | 2.702 GiB/s |
| 137184 | Single SN8100 2 TB, PCIe 5.0 x4; host-wide physical | 575 W, Gen5 x8 | 71.084 / 63.802 / 7.282 s | 9.417 / 4.263 / 5.154 s | 10.098 GiB/s |
| 150337 | No NVMe exposed; cgroup block device 8:16 | 575 W, Gen4 x16 | 102.419 / 67.974 / 34.445 s | 18.613 / 4.314 / 14.299 s | 0.851 GiB/s |
The Krea JSON for each machine uses the same filename stem with vast-krea-2 in place of
vast-minimax-h3. The dashboard discovers all fourteen files through DuckDB without a hand-maintained run
list.
Longitudinal repeatability#
The reacquired machines are variance controls, not three new independent hosts. Their paired penalties were stable enough to support that separation:
| Machine / workload | Earlier schema-v3 penalty | New schema-v4 penalty | Change |
|---|---|---|---|
| 108899 / MiniMax | 30.004 s | 28.713 s | -1.291 s |
| 137184 / MiniMax | 7.493 s | 7.282 s | -0.211 s |
| 137184 / Krea | 5.200 s | 5.154 s | -0.046 s |
| 150337 / Krea | 13.857 s | 14.299 s | +0.442 s |
Machine 137184 is particularly strong repeatability evidence: its model reader moved only from 10.190 to
10.098 GiB/s for MiniMax and its warm floor moved from 63.762 to 63.802 seconds. Machine 108899's new
schema-v4 evidence also explains why its large synthetic headline is misleading: QD32 direct fio reached
10.384 GB/s, but the application reader reached only 2.702 GiB/s through a buffered loop and a PCIe 4.0 x2
drive path.
What the four new hosts add#
Machine 30597 demonstrates why the paired response is necessary. Its MiniMax warm floor was a slow 82.371 seconds, but its cold penalty was only 9.160 seconds; comparing raw cold time would incorrectly assign the GPU/platform floor to storage. Machines 56897 and 29706 add ordinary single-drive paths with 13.492-14.536 second MiniMax penalties.
Machine 30024 is deliberately classified as other MD, not RAID0. Its QD32 fio reached 21.447 GB/s,
while QD1 reached 3.249 GB/s, the model reader reached 5.164 GiB/s, and the MiniMax penalty was 30.302
seconds. This is useful adjacent evidence that a very high advertised or high-queue-depth array throughput
does not guarantee a small application penalty. It does not estimate the benefit of the two-drive RAID0
planned for the dedicated machine. The dynamic RAID dashboard now recognizes only an explicit raid0
topology as RAID0; RAID10 and RAID1 observations stay available but are excluded from the single-versus-
RAID0 summaries.
2026-09-24: replace the remaining current schema-v3 runs#
The longitudinal controls above show that full schema-v3 and schema-v4 reruns on the same physical machine produce consistent application results. Nevertheless, schema v3 cannot be converted into schema v4 after the fact: it did not record per-pass loop, MD and physical-device counters. Later static topology is useful as a provenance-labelled supplement, but it cannot prove where the bytes from an earlier timed pass went.
The remaining latest schema-v3 RTX 5090 results will therefore be replaced in the current analysis only by full five-pair schema-v4 reruns on the exact same physical machine IDs. Existing schema-v3 JSONs remain as longitudinal controls and are never overwritten. Each reacquisition runs both MiniMax H3 and Krea 2 after a single setup so a machine that previously had only one workload gains the other at the same time.
| Machine | Current schema-v3 coverage | 2026-09-24 availability check |
|---|---|---|
| 45379 | Krea 2 | Unavailable |
| 55019 | MiniMax H3 and Krea 2 | Unavailable |
| 112410 | Krea 2; MiniMax already has a full schema-v4 repeat | Unavailable |
| 147169 | MiniMax H3; one-pair schema-v4 validation is not a full replacement | Complete; instance 52257980 destroyed |
| 147512 | MiniMax H3 and Krea 2 | Unavailable |
| 148782 | MiniMax H3 | Unavailable |
| 150183 | MiniMax H3 | Unavailable |
The acquired machine 147169 was the exact prior Oklahoma host: Ryzen 9950X, ROG Crosshair X870E Hero, 600 W RTX 5090, advertised PCIe 5.0 x16 and the previously identified 4 TB Samsung 9100 Pro. The suite pins ComfyUI v0.37.0, PyTorch 2.9.1+cu130, resident cache policy and five cold-to-warm pairs per workload. The runner persisted new timestamped JSON files and destroyed the instance after both workloads completed.
Machine 147169 completion and block-layer result#
Both workloads completed with five accepted cold-to-warm pairs and zero rejected cold attempts:
| Workload | Cold mean | Warm mean | Paired penalty | Model reader |
|---|---|---|---|---|
| MiniMax H3 | 73.786 s | 60.554 s | 13.231 s | 5.257 GiB/s |
| Krea 2 | 10.046 s | 4.015 s | 6.031 s | 7.498 GiB/s |
The full MiniMax schema-v3 result was 70.448 / 60.328 / 10.120 seconds with a 6.643 GiB/s reader. The new warm floor changed by only +0.226 seconds, but cold penalty increased by 3.112 seconds and the reader fell by 20.9%. The intervening one-pair schema-v4 validation measured an 11.362-second penalty and 6.199 GiB/s. This preserves the conclusion that the GPU/platform floor is stable while showing meaningful session and pass variation in the storage/container path; schema conversion alone was not responsible for the earlier number.
Schema v4 also resolves what “cold” means on this Vast host. The active model path is a buffered loop0
backed by /var/lib/docker-loop.xfs, above the one-hop PCIe 5.0 x4 Samsung namespace. Across five MiniMax
cold passes, the loop layer served 215.236 GB while the physical NVMe recorded only 59.891 GB, or 27.8% of
the loop bytes. The three independent MiniMax readers served 127.416 GB at the loop but only 13.868 GB at
the NVMe, or 10.9%. One reader reached 7.149 GiB/s while recording zero NVMe reads.
Krea makes the outer-cache behavior especially clear. Its first cold pass sent 17.847 of 18.338 GB to the NVMe. Cold passes two through five each missed the inner model-file cache and recorded about 18.3 GB at the loop, but recorded zero or 0.0008 GB at the NVMe. Thus schema v4 does not merely add metadata: it proves that the existing eviction policy does not produce repeated physical-media-cold samples through Vast's buffered loop. These results remain valid application/container-path measurements, but they must not be labelled physical-NVMe cold throughput.
2026-09-24: variance inside the PCIe 5.0 x4 SSD class#
The dashboard's six-host PCIe 5.0 x4 group spans 4.536-10.098 GiB/s in the independent model reader and 5.732-13.231 seconds in MiniMax paired cold penalty. The negotiated link is therefore a capacity ceiling, not a performance class. All six hosts expose one NVMe namespace through one upstream PCI hop, but they differ in drive/controller, block-queue limits, read-ahead, CPU/platform and container backing path.
| Machine | Drive | Reader | Active workload read | Active read duration | Overlapped by other startup | Paired penalty | Evidence note |
|---|---|---|---|---|---|---|---|
| 30597 | Crucial T700 | 6.264 GiB/s | 2.804 GiB/s | 14.161 s | 5.001 s | 9.160 s | Cgroup-attributed NVMe |
| 62211 | Samsung 9100 Pro | 5.699 GiB/s | 2.464 GiB/s | 16.154 s | 4.378 s | 11.776 s | Cgroup-attributed NVMe |
| 137184 | WD Black SN8100 | 10.098 GiB/s | 3.913 GiB/s | 11.035 s | 3.753 s | 7.282 s | Host-wide NVMe counters |
| 147011 | Corsair MP700 Elite | 4.536 GiB/s | 3.460 GiB/s | 11.506 s | 4.204 s | 7.302 s | Cgroup-attributed NVMe |
| 147169 | Samsung 9100 Pro | 5.257 GiB/s | 2.394 GiB/s | 16.919 s | 3.687 s | 13.231 s | Buffered loop; not media-cold |
| 148782 | Samsung 9100 Pro | 7.585 GiB/s | 4.187 GiB/s | 9.540 s | 3.809 s | 5.732 s | Schema v3; lower layer unproven |
The realized full-workload read rate explains the end-to-end spread much better than the standalone reader. Across the six host means, active read rate versus penalty has Pearson -0.942 and Spearman -1.000. Across all 30 cold/warm pairs the correlations are -0.926 and -0.951, and the within-host demeaned Pearson correlation is -0.805. The active read duration exceeds the paired penalty by only 3.687-5.001 seconds per host: roughly four seconds of model I/O overlaps CPU/CUDA startup, while the rest appears on the critical path. This quantity is outcome-adjacent and partly mechanical, so it diagnoses the penalty rather than independently predicting an untested machine.
The separate reader still validates storage-path capability. Buffered fio versus reader throughput has
Pearson 0.931 and Spearman 0.829 across all six hosts. After excluding the buffered-loop machine it rises to
Pearson 0.988 and Spearman 1.000. Reader versus paired penalty is much weaker at Pearson -0.478 and Spearman
-0.600 because ComfyUI/PyTorch interleaves reads with CPU tensor work and hides part of the I/O behind other
startup. Machine 147011 is the clearest example: its standalone MP700 reader is the slowest in the group,
but the application realizes 3.460 GiB/s while active and lands at the same 7.3-second penalty as the much
faster SN8100 reader.
Block-queue configuration is a plausible contributor, not yet a causal result. The SN8100 host exposes a 2,048 KiB read-ahead and 1,024 KiB maximum request, while the three directly observed 9100 Pro hosts expose 128 KiB read-ahead and maximum requests of 128, 128 and 512 KiB respectively. The 512 KiB 9100 host is the fastest of those three, but it is also the schema-v3 host and has a different CPU, motherboard and runtime. The setting cannot be credited in isolation. GPU link width is not the explanation either: the fastest reader uses a Gen5 x8 GPU/H2D path, while a Gen5 x16 host has essentially the same paired penalty despite a reader less than half as fast.
Machine 147169 must remain a separate backing-path class. Its loop layer read about 40 GiB per cold pass, but only 6-17 GiB reached the NVMe; one independent reader reached 7.149 GiB/s with zero NVMe reads. Its reader and penalty variance measure the buffered container path and outer-cache state, not repeated physical SSD performance. Machine 148782 remains lower-confidence until an exact schema-v4 rerun is available.
The operational conclusion is to retain SSD generation and lane width as covariates, never as grouping shortcuts. The useful chain is: queue/controller and backing path -> standalone reader capability -> realized active workload read duration -> paired cold penalty, with CPU/platform overlap between the last two steps. The dynamic single-path explorer now exposes the NVMe maximum-request class, in-workload active rate and duration, overlapped startup time and additional cold process CPU so this within-link variance remains visible when new result JSONs are discovered.
2026-09-24: cost-aware global RTX 5090 sampling#
The next inventory search is global rather than US-restricted. GPU power, allocated fraction, reliability, location and download speed are preferences rather than admission gates; physical machine identity remains the independent unit. Network tariff is weighted heavily because the previous suite downloaded 65.9 GB and paid $2.57 at $0.0391/GB, versus only $0.67 for GPU time.
Three new hosts were acquired with a combined expected 65.9 GB transfer charge of about $0.34: dedicated 600 W/QEMU-storage machine 90896 at $0/GB, dedicated 600 W/Intel-NVMe machine 142728 at $0.002604/GB and dedicated 575 W/Kioxia-NVMe machine 150265 at $0.002604/GB. All three completed the full MiniMax H3 and Krea 2 schema-v4 suite. The results are:
| Machine | Storage label | Workload | Cold mean | Warm mean | Paired penalty |
|---|---|---|---|---|---|
| 90896 | QEMU | MiniMax H3 | 120.180 s | 64.229 s | 55.951 s |
| 90896 | QEMU | Krea 2 | 30.115 s | 4.284 s | 25.831 s |
| 142728 | Intel SSDPF2KX076T1 | MiniMax H3 | 105.083 s | 68.162 s | 36.921 s |
| 142728 | Intel SSDPF2KX076T1 | Krea 2 | 23.404 s | 4.789 s | 18.615 s |
| 150265 | Kioxia KXG80ZN84T09 | MiniMax H3 | 79.254 s | 64.660 s | 14.594 s |
| 150265 | Kioxia KXG80ZN84T09 | Krea 2 | 12.995 s | 4.300 s | 8.695 s |
The 600 W fractional host 58084 advertised 24 GB/s storage at the same low tariff, but it never accepted the attached SSH key within ten minutes and the runner destroyed instance 52265805 without running a workload. Samsung 9100 Pro machine 42736 and 990 Pro machine 138147 disappeared before acquisition, so neither created an instance or incurred a charge.
The runner now persists Vast's download and upload tariffs, the sum of prepared model bytes and an estimated prepared-model transfer charge in each future result. The estimate deliberately excludes ComfyUI, Python and CUDA package setup traffic; provider billing remains the authority for total transferred bytes.
Full-fraction inventory expansion#
A subsequent live inventory snapshot returned 23 rentable gpu_frac=1.0 RTX 5090 physical machines. After
excluding machines already completed or active, 15 untested candidates charged no more than $0.006511/GB,
equivalent to at most about $0.40 for the suite's 61.109 GB of prepared models. Seven untested candidates
above that threshold were deliberately excluded; their prepared-model charge alone ranged from about $0.61
to $2.44 per run.
Eleven cost-safe machines were acquired for the same complete two-workload, five-pair schema-v4 suite: 30121, 139086, 151196, 150457, 45827, 148036, 148215, 129526, 110783, 129256 and 144671. Four more candidates (36518, 55535, 151410 and 31538) disappeared between inventory lookup and acquisition and created no billable instance. The acquired cohort spans 400-600 W cards, PCIe 4.0/5.0 GPU links, physical SN850X/990-series and other NVMe devices, plus repeated QEMU-backed storage paths. Results are written to new timestamped JSONs; existing observations are never overwritten and the dynamic dashboard discovers new JSONs automatically.
The expanded search found no rentable full-fraction RTX 5090 with a Samsung 9100 PRO. Machine 148771 is the
only visible exact full-fraction match (545 W, Threadripper 9960X, 9100 PRO 2 TB), but it was unavailable.
Two rentable exact-drive alternatives were rejected from this cohort because machine 4129 reports
gpu_frac=0.5 and machine 42736 reports gpu_frac=0.125.
Official drive-spec covariates#
Manufacturer specifications are now stored as normalized, joinable JSON tables in
storage-device-specs.json and
storage-device-aliases.json. The current version contains 22 exact product/capacity rows
and 29 raw-label aliases. It covers the identified Samsung, Western Digital, Crucial, Corsair, Solidigm/Intel, KIOXIA,
Kingston and HIKSEMI devices in the sampled or candidate cohort. QEMU, generic NVMe, ambiguous Netac, Fanxiang and
unpublished Phison OEM labels remain explicitly unresolved rather than being assigned guessed specifications.
This adds rated interface, NAND/cache design, sequential and random maxima, and endurance as analysis covariates. It does
not replace observed throughput: official maxima use vendor-specific queue depths and fresh-drive conditions, whereas the
ComfyUI path also includes filesystem/container layers, block-queue limits, PCIe topology, thermals and contention. The
schema and a DuckDB join example are documented in
storage-device-spec-database.md.
First completed full-fraction expansion results#
Seven more cost-screened, full-fraction machines completed both five-pair schema-v4 workloads. These files are additional observations; they do not replace earlier captures. The dynamic dashboard selects the latest full result for each physical machine and workload.
| Machine | Observed storage | Workload | Cold mean | Warm mean | Paired penalty |
|---|---|---|---|---|---|
| 45827 | WD Black SN850X 4 TB | MiniMax H3 | 146.915 s | 142.064 s | 4.851 s |
| 45827 | WD Black SN850X 4 TB | Krea 2 | 13.694 s | 8.629 s | 5.065 s |
| 129256 | Samsung 990 EVO Plus 1 TB | MiniMax H3 | 87.207 s | 64.724 s | 22.483 s |
| 129256 | Samsung 990 EVO Plus 1 TB | Krea 2 | 15.900 s | 4.271 s | 11.629 s |
| 139086 | Lexar NM1090 PRO 2 TB | MiniMax H3 | 71.562 s | 60.033 s | 11.530 s |
| 139086 | Lexar NM1090 PRO 2 TB | Krea 2 | 11.089 s | 4.165 s | 6.924 s |
| 147380 | Samsung 990 EVO Plus 1 TB | MiniMax H3 | 108.463 s | 66.231 s | 42.231 s |
| 147380 | Samsung 990 EVO Plus 1 TB | Krea 2 | 25.155 s | 4.546 s | 20.609 s |
| 148214 | QEMU block device; no exposed NVMe | MiniMax H3 | 122.613 s | 67.555 s | 55.057 s |
| 148214 | QEMU block device; no exposed NVMe | Krea 2 | 29.994 s | 4.301 s | 25.693 s |
| 150215 | Phison OEM 2 TB | MiniMax H3 | 88.815 s | 71.274 s | 17.541 s |
| 150215 | Phison OEM 2 TB | Krea 2 | 11.853 s | 4.778 s | 7.075 s |
| 151196 | Fanxiang S880 2 TB path | MiniMax H3 | 89.682 s | 66.179 s | 23.503 s |
| 151196 | Fanxiang S880 2 TB path | Krea 2 | 18.004 s | 4.534 s | 13.470 s |
The new observations expose three useful controls. Machine 147380 negotiated its 990 EVO Plus at only PCIe 3.0 x2 and realized 1.069 GiB/s, versus 3.331 GiB/s for the same model on machine 129256 at PCIe 4.0 x4; their MiniMax penalties were 42.231 and 22.483 seconds respectively. Machine 148214 repeats the slow virtual-storage pattern: no physical NVMe is exposed, the model reader reaches only 0.810 GiB/s, and the penalty is 55.057 seconds. Machine 45827 demonstrates why raw cold time is not the storage outcome: its 400 W GPU has a 142.064-second warm floor, but cold-minus-warm is only 4.851 seconds.
Machine 45827: a small penalty caused by a slow floor, not a faster RTX 5090#
The MiniMax result for machine 45827 can look unusually fast if it is ranked only by cold-minus-warm. It is the opposite when ranked by end-to-end service time:
| Measure | Machine 45827 | Cohort context |
|---|---|---|
| MiniMax cold mean | 146.915 s | slow, not fast |
| MiniMax warm mean | 142.064 s | slowest of the 30 latest RTX 5090 hosts; cohort median 66.231 s |
| MiniMax paired penalty | 4.851 s | small because cold work overlaps the 142 s execution |
| Krea cold / warm / penalty | 13.694 / 8.629 / 5.065 s | its warm Krea floor is also the cohort's slow tail |
| Independent model reader | 5.617 GiB/s | healthy PCIe 4.0 SN850X result, but not the fastest reader |
| Workload active-read throughput | 3.013 GiB/s | about 13.189 s of active reads per cold run |
The cache protocol did not accidentally turn the cold samples warm. All five accepted cold processes report
42.56-42.73 GB of block reads after eviction. Four warm processes report zero block reads and one reports
only 1.38 MB. Because DynamicVRAM accesses the staged weights lazily, those reads are distributed across
nearly the complete cold execution window. The approximately 13.189 seconds of active reading add only
4.851 seconds to wall time; the remaining approximately 8.337 seconds are hidden behind work already on the
critical path. active read duration - paired penalty is only an overlap heuristic, not a measured stage
boundary, but it explains why the penalty alone ranks this host misleadingly.
The SSD and host-to-device path do not explain the slow warm floor. The physical WD Black SN850X 4 TB
negotiates PCIe 4.0 x4, has no MD RAID, reads the prepared model set at 5.617 GiB/s, and the GPU link
negotiates PCIe 5.0 x16 with 47.555 GiB/s pinned H2D. There was no process swap activity. The warm MiniMax
process instead consumes about 152.986 CPU-seconds over 142.064 wall seconds, compared with roughly
88.7 CPU-seconds over 76.0 wall seconds on machine 108899. That control has the same Ryzen 9950X, a nearby 420 W
power limit, GPU subsystem ID 53061462 and VBIOS 98.02.2E.40.AF. Two other machines with that exact GPU
subsystem/VBIOS complete warm MiniMax in 67.974-71.274 seconds. The 400-420 W range and board identity are
therefore insufficient explanations for a twofold slowdown.
The remaining evidence localizes the anomaly to the runtime/loaded accelerator path but does not identify
one physical cause. Machine 45827 is the only current observation with NVIDIA driver 590.48.01 and kernel
6.8.0-58-lowlatency; those are confounded host-level differences, not proof of a driver or kernel defect.
The schema did not sample loaded GPU clocks, actual power draw, temperature, performance state, throttle
reasons, CPU frequency/governor or competing GPU processes. A reacquisition would need that time-series
telemetry during both warm workloads to distinguish an effective sub-400 W power cap, clock/thermal
throttling, host contention and a driver/runtime path. Until then the defensible conclusion is:
- machine 45827 is a compute/runtime outlier, not evidence that its single SN850X makes MiniMax faster;
- its small cold penalty is real for that run, but is an overlap artifact relative to an unacceptable warm floor and must not be used as evidence against faster storage; and
- cold wall time and warm wall time must remain visible beside the paired penalty in every comparison.
Targeted reacquisition and loaded-GPU telemetry#
Machine 45827 was reacquired on 2026-09-24 with the same offer-visible hardware, driver and power limit. A new one-pair schema-v4 run reproduced the anomaly:
| Observation | Original five-pair result | Reacquired one-pair diagnostic |
|---|---|---|
| Cold wall | 146.915 s mean | 140.030 s |
| Warm wall | 142.064 s mean | 137.865 s |
| Cold-minus-warm | 4.851 s mean | 2.165 s |
| Independent model reader | 5.617 GiB/s | 5.624 GiB/s |
| Pinned H2D | 47.555 GiB/s | 47.640 GiB/s |
The repeat rules out a one-off seed, transient cache state and the earlier benchmark process. During the
new cold and warm passes, nvidia-smi dmon recorded the following loaded state. This follows NVIDIA's own
guidance to run dmon alongside an inference workload for power, utilization and clock evidence
(NVIDIA TensorRT performance measurement guidance).
| Loaded telemetry | Cold | Warm |
|---|---|---|
| Samples with at least 50% reported SM activity | 133 | 123 |
| Mean reported SM utilization | 98.93% | 99.98% |
| Mean graphics clock | 2,693 MHz | 2,644 MHz |
| Mean board power | 303.65 W | 302.83 W |
| Mean GPU temperature | 67.35 C | 67.13 C |
| Thermal-violation flag | 0% | 0% |
Process telemetry showed only the benchmark Python process during the two passes. A live nvidia-smi -q
snapshot reported P1, no active idle/application-clock/hardware/thermal slowdown reason, 2,662 MHz graphics
clock, approximately 307 W average power, and 100% utilization. The device exposes all 170 SMs, 32 GB VRAM,
default compute mode, no MIG/vGPU mode and no MPS process or active-thread-percentage environment limit.
CPU frequency telemetry showed cores boosting above 5.3 GHz rather than remaining at the 600 MHz idle
floor. This rules out the simple versions of thermal throttling, a stuck low clock, another GPU process,
MPS/MIG partitioning and a non-boosting CPU.
A committed wall-clock CUDA probe adds direct evidence that the abnormal warm floor is an accelerator-path problem. It uses the identical PyTorch 2.9.1/CUDA 13.0 userspace on machine 45827 and on an otherwise idle, second RTX 5090 allocated with machine 56934. The control is a 600 W card on driver 595.84, so power and driver remain explicit confounders rather than being silently treated as matched:
| Probe | Machine 45827, 400 W / driver 590.48.01 | Machine 56934 GPU 1, 600 W / driver 595.84 | Difference |
|---|---|---|---|
| BF16 16,384-square GEMM | 180.30 TFLOP/s | 239.14 TFLOP/s | -24.6% |
| FP32 with TF32 enabled, 12,288-square GEMM | 50.81 TFLOP/s | 124.18 TFLOP/s | -59.1% |
| 2 GiB device-to-device copy | 561.41 GB/s | 767.40 GB/s | -26.8% |
The control is not sufficient to call the 590 driver causal: its higher power limit explains some of the BF16 and memory difference, and the driver, board and host also differ. It is nevertheless decisive for the storage question. Machine 45827 is slow and unusually variable in device-local operations that do not touch the SN850X, host page cache or PCIe H2D path. The nearby 420 W machine with the same GPU subsystem ID and VBIOS, 108899 produced a 75.979-second warm MiniMax mean, so nominal power alone still cannot explain 137.865 seconds. The remaining credible boundary is a machine-specific accelerator/runtime condition, with the 590.48.01 driver plus low-latency kernel and latent board/firmware behavior still confounded.
The raw diagnostic artifacts live under ../diagnostics/: the complete one-pair result,
per-second GPU/process telemetry and both compute-probe JSONs are retained separately from canonical
five-pair dashboard inputs. Machine 45827 should be flagged as a compute outlier and excluded from claims
about SSD or RAID effects; it remains useful as evidence that a small cold-minus-warm penalty can be
manufactured by a pathological warm floor.
Four further full-fraction suites completed while the first analysis pass was running:
| Machine | Observed storage | Workload | Cold mean | Warm mean | Paired penalty |
|---|---|---|---|---|---|
| 30121 | QEMU block device; no exposed NVMe | MiniMax H3 | 140.618 s | 107.040 s | 33.578 s |
| 30121 | QEMU block device; no exposed NVMe | Krea 2 | 19.761 s | 6.787 s | 12.974 s |
| 148036 | Samsung 990 PRO 2 TB | MiniMax H3 | 77.856 s | 70.085 s | 7.771 s |
| 148036 | Samsung 990 PRO 2 TB | Krea 2 | 10.316 s | 4.768 s | 5.548 s |
| 148215 | QEMU block device; no exposed NVMe | MiniMax H3 | 125.043 s | 67.919 s | 57.124 s |
| 148215 | QEMU block device; no exposed NVMe | Krea 2 | 29.279 s | 4.294 s | 24.985 s |
| 150457 | Samsung 990 PRO 4 TB | MiniMax H3 | 78.221 s | 68.975 s | 9.246 s |
| 150457 | Samsung 990 PRO 4 TB | Krea 2 | 10.444 s | 4.766 s | 5.678 s |
Machines 129526 and 110783 never progressed beyond Vast's provider-side loading state and were destroyed by
the runner before any workload. Machine 144671 downloaded the first 19 GiB model at only about 10 MiB/s; its
remaining instance 52270786 was explicitly destroyed after the benchmark process disappeared. The Vast
account was then checked and had zero active instances. No result JSON was produced for those three machines.
Expanded-cohort analysis#
With the completed files discovered dynamically, the latest-host dataset contains 30 independent MiniMax
machines, 24 Krea machines, 23 machines with both workloads and 26 non-MD-RAID MiniMax paths. On those 26
single/container paths, buffered fio predicts the independent model reader at Pearson 0.946 and Spearman
0.928; the model reader predicts paired MiniMax penalty at -0.696 and -0.752. On the narrower 19-host
575-600 W, H2D-at-least-10 GiB/s, GPU-x16 slice, the reader-to-penalty Spearman relationship is -0.932.
The official-spec join identifies 18 hosts without guessing virtual, generic or ambiguous multi-device paths. Manufacturer sequential rating predicts model-reader throughput at Pearson 0.743 and Spearman 0.769, but is weaker against final penalty (-0.481 and -0.658). Realized model reads range from 16% to 83% of the official maximum. This supports using the data-sheet rating as a prior and the measured reader as the acceptance test.
The two additional exact Linux identities are Crucial P310 1 TB (CT1000P310SSD8) and Lexar NM1090 PRO
2 TB. Official manufacturer sources were added to the JSON spec table. The Lexar host is especially useful:
the drive is rated for PCIe 5.0 x4 and 14.0 GB/s, but negotiates PCIe 4.0 x4 on machine 139086 and realizes
3.298 GiB/s. This makes the difference between product capability and installed-path realization explicit.
The result also replicates between applications: across 23 dual-workload hosts, MiniMax and Krea paired
penalties correlate at Pearson 0.985 and Spearman 0.975, while their model readers correlate at 0.978 and
0.962. The expanded interpretation, outlier explanations and owned-node acceptance gates are recorded in
rtx5090-single-path-multivariable-analysis.md.
Disk × setup × performance dashboard view#
The dynamic RAID analysis dashboard now includes a row-level view over the latest deduplicated results for
both MiniMax H3 and Krea 2. It keeps the advertised disk identity and all visible NVMe devices beside the
negotiated topology/link, upstream-hop inference, buffered-loop evidence, CPU/RAM, GPU PCIe/H2D/power path,
kernel/runtime and observed workload performance. Filters select workload, exact disk identity and layout;
the focus can be changed between paired cold penalty, cold/warm wall time, model-reader throughput and the
three fio views. This is deliberately a host-level comparison rather than an SSD-product leaderboard:
installed-path covariates remain visible instead of being averaged away.
Why the negotiated PCIe 5.0 x4 group still varies#
The latest independent MiniMax cohort contains seven non-RAID hosts whose exposed SSD negotiated PCIe 5.0 x4. That label is a link-width ceiling, not an end-to-end performance class:
| Measure | Observed range | Fastest / slowest |
|---|---|---|
Buffered fio after invalidation |
4.876–13.027 GB/s | 2.67× |
| Actual model reader | 4.536–10.098 GiB/s | 2.23× |
| Paired cold penalty | 5.732–17.541 s | 3.06× |
Within these seven hosts, buffered fio versus the independent model reader has Pearson 0.900 and Spearman
0.679; direct QD32 fio versus the model reader is 0.831 and 0.536. The immediate explanation for most of
the reader spread is therefore the realized mounted/container filesystem path, not a difference in the
nominal PCIe generation. The remaining physical causes are not isolated by this cross-host sample. They
include different SSD/controller implementations, queue limits, filesystem/runtime behavior and contention.
Queue configuration is a concrete clue, but not yet causal evidence. The fastest host in this class exposes 2 MiB read-ahead and a 1 MiB maximum request, while the four 128 KiB-maximum-request hosts average only 5.898 GiB/s. Those groups also use different drives and machines, so this cannot be called a queue-setting effect. The three Samsung 9100 Pro 4 TB observations further demonstrate installed-path variance: their model readers are 5.257, 5.699 and 7.585 GiB/s. The slowest has an observed active buffered loop; the fastest has a 512 KiB maximum request but only schema-v3 loop evidence.
Application penalty has more remaining variance than the model reader. Within the seven-host class, reader-versus-penalty correlation is only Pearson -0.336 and Spearman -0.393. Once gross storage pathology is removed, CPU deserialization, many-file access, GPU transfer/overlap and other host activity can still affect the final cold-minus-warm result. PCIe 5.0 x4 is neither a sufficient performance guarantee nor the variable that explains the class internally.
Vast pre-rental filtering boundary#
A read-only live query on 2026-09-24 confirmed that Vast's /bundles/ endpoint returns disk_name and accepts
exact eq and multi-value in filters for it. It also accepts the documented disk_bw filter. It does not
expose an SSD negotiated-generation field. Vast's pci_gen and pcie_bw offer fields describe the CPU-to-GPU
PCIe path, not the disk path, as documented in the
official CLI source. disk_name is not
listed among the CLI's documented offer fields even though the live backend accepted it, so this project
should use the raw API/client-side selection rather than depend on it as a stable CLI contract.
The practical pre-rental strategy is therefore:
- Filter normal constraints plus
disk_bw, download cost and full-GPU preference. - Match
disk_nameagainst the official SSD-spec alias database and prefer known PCIe 5.0 x4 models. - Treat generic labels such as
nvmeas unknown rather than PCIe 5. - After acquisition, verify
32.0 GT/s PCIe x4, visible devices, topology, upstream hops and buffered-loop evidence before admitting the host to the PCIe 5.0 x4 analysis.
This remains a capability prefilter, not proof of the installed link: an earlier Lexar NM1090 Pro host is a PCIe 5.0-capable product that actually negotiated PCIe 4.0 x4.
2026-09-24: additional PCIe 5 SSD sampling#
Four more RTX 5090 hosts completed the full five-pair MiniMax H3 and Krea 2 suite. Every accepted cold sample passed the process-block-read threshold and every warm pass remained in the same ComfyUI process with a different seed. The dashboard discovered all eight result files from the benchmark directory without a hand-maintained row.
| Machine | Mounted model-storage path | GPU / allocation caveat | Workload | Cold | Warm | Penalty | Model reader | Buffered fio |
Direct QD32 fio |
|---|---|---|---|---|---|---|---|---|---|
| 143669 | Single Samsung 9100 Pro 2 TB, PCIe 5.0 x4; separate 1 TB OS SSD | 600 W; GPU at PCIe 5.0 x8; allocation exposed two GPUs and the suite used GPU 0 | MiniMax H3 | 68.216 s | 61.126 s | 7.090 s | 9.755 GiB/s | 10.976 GB/s | 14.717 GB/s |
| 143669 | same | same | Krea 2 | 9.439 s | 4.015 s | 5.424 s | 9.827 GiB/s | 10.707 GB/s | 14.721 GB/s |
| 56934 | Single Crucial T710 2 TB, PCIe 5.0 x4 | 600 W; GPU at PCIe 5.0 x8; allocation exposed two GPUs and the suite used GPU 0 | MiniMax H3 | 66.691 s | 61.088 s | 5.603 s | 5.522 GiB/s | 10.623 GB/s | 14.844 GB/s |
| 56934 | same | same | Krea 2 | 8.458 s | 4.065 s | 4.393 s | 5.565 GiB/s | 10.597 GB/s | 14.815 GB/s |
| 148771 | Single Samsung 9100 Pro 2 TB, PCIe 5.0 x4 | one 545 W GPU; GPU at PCIe 5.0 x16 | MiniMax H3 | 70.577 s | 64.330 s | 6.248 s | 9.596 GiB/s | 10.006 GB/s | 14.717 GB/s |
| 148771 | same | same | Krea 2 | 9.082 s | 4.267 s | 4.814 s | 9.730 GiB/s | 10.793 GB/s | 14.700 GB/s |
| 145570 | Linux MD RAID0 over two WD Black SN8100 2 TB drives, both PCIe 5.0 x4, 512 KiB chunk | one GPU allocated from a four-GPU host; 450 W; GPU at PCIe 5.0 x16 | MiniMax H3 | 103.782 s | 72.646 s | 31.136 s | 6.005 GiB/s | 6.283 GB/s | 22.346 GB/s |
| 145570 | same | same | Krea 2 | 17.737 s | 4.782 s | 12.954 s | 5.864 GiB/s | 6.730 GB/s | 24.285 GB/s |
This expansion changes three parts of the interpretation:
- A single SSD can reach the approximately 10 GiB/s application-shaped tier. Machine 143669's mounted 9100 Pro reads the real prepared model set at 9.755-9.827 GiB/s, so RAID0 is not a prerequisite for that throughput. Machine 148771 independently repeats the result at 9.596-9.730 GiB/s on another 9100 Pro, a different Threadripper host and an x16 GPU link.
- RAID0 is not a guarantee that the application will receive aggregate sequential bandwidth. Machine 145570 exposes a real, clean, two-member MD RAID0 and reaches 22.346-24.285 GB/s at direct QD32, yet its buffered path and model reader remain near 6 GiB/s. Its 450 W GPU, EPYC host and much larger additional cold-process CPU time make it a multivariable observation, not a clean estimate of RAID's causal effect.
- Headline
fiobandwidth is not interchangeable with the application-shaped reader. The single T710 sustains 10.623 GB/s in bufferedfioand 14.844 GB/s at direct QD32, but the multi-file reader reaches only 5.522 GiB/s. Even so, its MiniMax paired penalty is only 5.603 seconds because about 3.781 seconds of its 9.383-second active-read interval overlaps other work. This is another reason to retain the complete cold/warm pair rather than predict service time from one storage metric.
The direct storage decision remains unchanged: buy one 4 TB 9100 Pro first and accept it using both the independent model-reader test and the complete cold/warm suite. Do not buy a second drive solely to obtain RAID0; add capacity or RAID later only if the owned-node measurements miss the service target. The new samples strengthen the sufficiency claim for a high-performing single Gen5 SSD, but they do not provide a same-machine estimate of the incremental benefit of RAID0.
Across the now eleven non-RAID MiniMax paths that expose at least one SSD negotiated at PCIe 5.0 x4, the
application-shaped reader spans 4.536-10.098 GiB/s and the paired penalty spans 5.603-17.541 seconds. Buffered
fio still tracks the reader directionally (Pearson 0.784, Spearman 0.700), while reader versus penalty is
weak inside this already-fast link class (Pearson -0.355, Spearman -0.300). The broad cohort establishes that
gross storage performance matters; the within-Gen5 slice establishes that once the worst storage paths are
removed, overlap, CPU-side loading and platform/runtime variance become large relative to the remaining
storage difference.
One other exact PCIe 5 offer, machine 46165 with a 4 TB 9100 Pro, never accepted the provisioned SSH key and was destroyed without producing a result. A replacement candidate became provider-deverified before provisioning and was not rented. The runner now rejects explicit offers that allocate more than one GPU or advertise less than 48 GB of CPU RAM, preventing those offer-shape mistakes in later samples.
2026-09-27: focused multivariable reanalysis of single-path variance#
The current dataset has grown to 34 independent MiniMax machines, 30 Krea machines, 29 machines with both
workloads and 29 MiniMax single/no-MD-RAID paths. The canonical interpretation has been rewritten in
rtx5090-single-path-multivariable-analysis.md. This entry
records the changes in reasoning without rewriting the earlier chronological observations.
Correction: “startup overlap” is a derived residual#
The API field startup_overlap_s is calculated as active_read_duration_s - cold_penalty_s. It is not an
independently observed stage duration. Earlier entries used “overlap” as shorthand and sometimes gave that
residual too much explanatory weight. A positive residual is consistent with concurrent reading and other
startup work, but cannot prove the amount of overlap or distinguish concurrency from instrumentation and
stage-boundary effects.
The read accounting is otherwise unusually tight. Single-path cold processes read 39.53-43.17 GiB, with a
39.84 GiB median. active_read_duration_s agrees with process bytes divided by active rate to a median
absolute difference of 0.015 seconds. Across 29 hosts, the descriptive fit
penalty ~= -4.92 + 41.848 / active_read_GiB_s has in-sample R² 0.994. Because the rate is measured inside
the timed run and shares its read volume with the response, this is a diagnostic localization, not a causal
SSD-speed model.
Aggregate, restricted and within-class relationships#
On all 29 single paths, Spearman correlations with MiniMax penalty are -0.604 for Vast's advertised disk,
-0.761 for direct QD32 fio, -0.837 for buffered fio, -0.779 for the standalone model reader and -0.980
for in-workload active rate. Additional cold CPU is +0.800; warm runtime and H2D are weak at +0.264 and
-0.130. Ten thousand host-level bootstrap resamples produce intervals that exclude zero for the first five
storage/active metrics and additional CPU, but include zero for warm runtime and H2D.
Removing the five virtual/no-exposed-NVMe paths weakens the advertised, QD32, buffered and model-reader rank relationships to -0.337, -0.601, -0.718 and -0.635 respectively. Active rate remains -0.966. The broad relationship is therefore partly a contrast between pathological virtual paths and physical NVMe, not a clean marginal-speed curve among healthy SSDs.
Within exact single-device link classes, independent-probe relationships weaken further. The nine PCIe 5.0 x4 hosts span 5.60-17.54 seconds and have model-reader-versus-penalty Spearman -0.250. The eleven PCIe 4.0 x4 hosts span 4.85-36.92 seconds and have Spearman -0.091. Active rate remains almost perfectly ordered in both groups (-1.000 and -0.991). This supports treating link class as a ceiling and path descriptor rather than an explanation of within-class variance.
Predictive stress test#
Nested leave-one-host-out validation standardizes predictors within each training fold and selects ridge regularization inside that fold. Independent storage probes together reach MAE 5.434 seconds, RMSE 6.761 seconds and predictive R² 0.818. The captured platform variables alone reach RMSE 16.221 and R² -0.047. Adding them to storage improves MAE to 4.844 but does not improve RMSE or R² (6.855 and 0.813). An inverse active-rate diagnostic reaches RMSE 1.319 and R² 0.993, as expected from the read accounting above.
Categorical hardware labels do not isolate the missing causes. Path categories alone reach predictive R² 0.358, drive family 0.449, storage probes plus path categories 0.717, and storage plus path plus drive labels 0.760. These high-dimensional fits have only 29 machines and are sensitivity checks, not stable effect estimates.
Resolution of the apparent slow-disk/fast-penalty cases#
Machine 148036 is the cleanest apparent contradiction. Its standalone 990 PRO reader is only 2.001 GiB/s, but ComfyUI realizes 3.251 GiB/s during the workload and its 12.182-second read-active interval accompanies a 7.771-second penalty with only 0.120 seconds of paired SD. The active-rate fit predicts about 7.95 seconds. The standalone probe understated the loaded path; the data does not identify which filesystem or access- pattern detail caused that difference.
Machine 150215 shows the reverse: 6.374 GiB/s standalone but only 1.796 GiB/s during ComfyUI, yielding a 17.541-second penalty. Machine 45827 remains the largest residual: its active rate predicts about 8.97 seconds but the observed mean is 4.851 seconds. Its five-pair SD is the cohort maximum at 2.560 seconds, its warm floor is pathological, and device-local diagnostics show an accelerator/runtime outlier. Its small penalty must not be promoted into evidence that the SN850X hid a measured amount of storage time.
The reanalysis therefore narrows rather than reverses the storage conclusion: cold-penalty variance is
mostly storage-path variance expressed through actual workload delivery, but it is not SSD-spec variance.
Buffered fio is the strongest independent screening proxy in the current sample. The full paired workload
and lower-layer read evidence remain the acceptance test.