RTX 5090 multivariable investigation journal#

Status: Journal Started: 2026-09-23 Latest entry: 2026-09-27 Primary slice: MiniMax H3, ComfyUI v0.37.0, resident cache, five verified cold/warm pairs

This journal tracks the variables that can change an RTX 5090 result so cross-host differences are not misattributed to storage. Raw factors are normalized by rtx5090_minimax_factors.sql, and physical block propagation is normalized by rtx5090_minimax_block_layers.sql. The decision-level synthesis is maintained in rtx5090-raid0-broader-evidence-analysis.md.

Analysis model#

The physical machine is the independent unit. A repeat of one machine is a longitudinal control, not another independent host. The end-to-end time is separated conceptually as:

cold wall time = warm GPU/platform floor + paired storage/container-path penalty

The saved factors include actual and advertised GPU power, GPU identity/VBIOS/driver/link, CPU and logical allocation, motherboard, RAM/cgroup/swap, kernel, container implementation, exact NVMe identity/firmware/ link/PCI ancestry, MD layout, active loop backing, model reader, buffered/direct fio, H2D, read-active duration/rate, CPU use and major faults. Schema v4 additionally measures whether process block reads reach loop, MD and physical NVMe layers.

Hardware observations copied from a later capture of the same physical machine carry an explicit hardware_metadata_backfill.provenance record. Stable identity/topology is kept separate from mutable runtime configuration. Existing backfills cover old results from machines 112410, 147169 and 150656.

Prior exploratory checkpoint#

This checkpoint predates the seven-suite expansion documented at the end of this journal. After the single-SN850X control, there were eleven full v0.37.0 RTX 5090/MiniMax host results. Their warm runtime strata show why raw cold time is not a storage metric:

Actual GPU power Hosts Warm mean Range Mean cold penalty
420 W 1 75.626 s 75.626 s 30.003 s
450 W 1 74.545 s 74.545 s 17.100 s
500 W 1 69.986 s 69.986 s 11.776 s
575 W 4 63.792 s 62.266-65.404 s 7.543 s
600 W 4 61.300 s 60.328-62.716 s 10.877 s

Exploratory Pearson correlations across those eleven heterogeneous hosts are:

Variables Correlation
Warm runtime vs actual GPU power -0.9840
Cold penalty vs read-active duration +0.9881
Cold penalty vs model-reader rate -0.7499
Cold penalty vs direct QD1 fio -0.6264
Cold penalty vs direct QD32 fio -0.3623
Cold penalty vs pinned H2D rate -0.1120

These associations are descriptive, not causal estimates. The sample is small and power, CPU, kernel, storage, accelerator transfer path, container path and motherboard are highly collinear. A many-coefficient regression would overfit. The high-value evidence is instead paired outcomes, explicit stratification and same-machine interventions.

The aggregate H2D correlation is not evidence that accelerator transfer is irrelevant. Ten hosts occupy a wide 26.894-53.896 GiB/s band, while one new host is isolated at 3.359 GiB/s; a linear coefficient across power strata is the wrong model for a possible low-bandwidth threshold. Among the seven 575-600 W hosts with H2D at least 10 GiB/s, two MD-array paths average a 6.930-second penalty and five single/container paths average 7.893 seconds. Their ranges overlap (5.419-8.440 versus 5.732-10.120 seconds), and the schema-v3 paths do not prove physical-media reads. This is not a RAID effect estimate.

Stratified patterns worth retaining#

Restricting the slice to the seven 575-600 W hosts with measured H2D at least 10 GiB/s removes the clear power-limited hosts and the new transfer-path outlier. Within that still-heterogeneous group:

Variables Pearson correlation
Cold penalty vs full-workload read-active duration +0.9946
Cold penalty vs separate model-reader rate -0.2480
Cold penalty vs direct QD1 fio +0.6089
Cold penalty vs pinned H2D +0.0550

The read-active duration is part of the full workload trace, so its strong relationship is mechanically plausible rather than an independent causal estimate. More importantly, the synthetic and standalone throughput metrics stop predicting penalty reliably after basic platform admission. This is why the full paired workload remains the decision metric.

The closest observational pair currently available is useful despite not being a controlled intervention:

Factor RAID machine 112410 Single machine 148782
CPU family Ryzen 9950X3D Ryzen 9950X3D
Kernel 6.8.0-138 6.8.0-138
GPU topology / H2D Gen5 x16 / 45.682 GiB/s Gen5 x16 / 46.747 GiB/s
GPU power 575 W 600 W
Storage 2x SN850X RAID 0 1x Samsung 9100 Pro
Model reader 11.504 GiB/s 7.585 GiB/s
Warm / penalty / cold 65.404 / 5.419 / 70.823 s 61.035 / 5.732 / 66.767 s

The RAID reader is 1.52x faster, yet its paired penalty is only 0.313 seconds smaller—well below the predeclared two-second application threshold—and its raw cold time is 4.056 seconds slower because its warm floor is 4.369 seconds slower. The comparison remains cross-host and the single side is schema v3, so it is supporting evidence rather than a RAID effect estimate. It does show how a large reader-rate difference can collapse to a small full-application difference.

Other apparent categories are too sparse or collinear for attribution. Ryzen 9950X3D hosts average a 6.657-second penalty versus 8.305 seconds for Ryzen 9950X hosts in the admitted slice, but each group has only three machines and their storage/kernel mix differs. Only kernel 6.8.0-138 has two admitted hosts; every other kernel is represented once. There is one containerd v0.37 host, one x8 GPU host and one H2D under-10-GiB/s host. These are columns to preserve and future matching targets, not effects to claim.

The same-machine metadata audit found only machines 112410, 147169 and 150656 with later captures capable of informing older runs. Their provenance-labelled backfills already exist. No new old run has a safe same-machine source, so missing PCI ancestry remains null rather than being inferred from a similar board or drive model.

2026-09-23: exact physical single-drive result#

Machine 62211 used one 4 TB Samsung 9100 Pro on PCIe 5.0 x4 and a dedicated PCIe 5.0 x16 RTX 5090. Every one of five cold passes and three model readers propagated 100.0% of process block bytes to the Samsung namespace. Its primary results were 81.762 seconds cold, 69.986 seconds warm, 11.776 seconds paired penalty, 5.699 GiB/s model reader and 44.743 GiB/s H2D. This is an exact-drive physical storage control, not a warm-performance baseline: the GPU was limited to 500 W and the host used an Intel Core Ultra 7 265K.

2026-09-23: same-machine RAID replication#

Vast machine 112410 was reacquired as instance 52211104. This is the exact verified, dedicated Texas host used by the existing v0.37.0 RAID result: Ryzen 9950X3D2, 575 W RTX 5090, PCIe 5.0 x16 GPU and two 4 TB WD Black SN850X members in an MD array. The new run used schema v4 and the full five-pair protocol. Result: machine 112410 replication JSON.

It answered both questions without changing machines:

  1. The earlier 70.842 / 65.458 / 5.384-second cold/warm/penalty result repeated at 70.823 / 65.404 / 5.419 seconds. The differences were -0.019 / -0.054 / +0.035 seconds.
  2. Every cold and model-reader byte reached md0, then divided essentially 50.00/50.00 between the two physical SN850X namespaces. No loop device appeared in the cgroup I/O path.
Metric Original schema v3 Schema-v4 replication
Cold mean 70.842 s 70.823 s
Warm mean 65.458 s 65.404 s
Paired penalty 5.384 s 5.419 s
Model reader 11.380 GiB/s 11.504 GiB/s
Buffered fio 13.772 GB/s 13.746 GB/s
Direct QD1 / QD32 8.187 / 14.451 GB/s 8.374 / 14.461 GB/s
H2D 45.291 GiB/s 45.682 GiB/s

The new paired-penalty sample SD was 0.481 seconds. Cold process reads were 46.329-46.680 GB in both the old and new captures, not a new schema-v4 artifact. The independent reader remained exactly 42.472 GB. For each reader pass, md0 read 42.472 GB and each physical member read 21.236 GB. The RAID path and cache eviction are therefore physically corroborated.

This establishes that machine 112410 really uses both drives and that its strong result is repeatable. It still does not measure the incremental RAID benefit: there is no same-host, same-filesystem single-member run. Relative to the physical single MP700 control, the RAID reader was 2.54x faster (11.504 versus 4.536 GiB/s) but the paired penalty was only 1.883 seconds smaller (5.419 versus 7.302 seconds), while raw cold time was 0.805 seconds slower because the RAID host had a slower warm floor. This cross-host comparison is descriptive and falls just short of the experiment's predeclared two-second application threshold.

The single-SN8100 machine 137184 was also available, but only at gpu_frac=0.5; it was rejected rather than weakening the dedicated-host rule. The two-drive T705 machine 147512 was unavailable. The remaining missing PCI ancestry on their schema-v3 files cannot be reconstructed honestly without a same-machine capture and is left null rather than guessed. Instance 52211104 was destroyed after result persistence.

2026-09-23: adjacent single-SN850X control#

Vast machine 132465 was acquired as instance 52214899 to add the missing single-drive side of the SN850X family. The verified dedicated Maryland host had one 4 TB SN850X on a direct PCIe 4.0 x4 root, a 600 W RTX 5090 in the CPU x16 slot, an i9-12900K, 64 GB RAM and no MD array. The full five-pair run completed without rejected attempts and the instance was confirmed destroyed. Result: machine 132465 JSON.

Metric Result
Cold mean 81.476 s
Warm mean 61.122 s
Paired cold penalty 20.354 s
Model reader 5.300 GiB/s
Buffered fio 4.922 GB/s
Direct QD1 / QD32 fio 3.549 / 6.956 GB/s
Pinned H2D 3.359 GiB/s

The H2D result reproduces Vast's 3.3 GB/s PCIe score despite the slot topology and maximum-capability metadata saying Gen5 x16. It is an actual transfer measurement, whereas the captured current Gen1 state is an idle power-management snapshot. Among the other three 600 W hosts, H2D is 46.747-51.220 GiB/s and the mean cold penalty is 7.718 seconds. Machine 132465 has essentially the same warm floor but a penalty 12.636 seconds larger. Its direct QD1 storage and model reader are also slower, so this run cannot isolate whether disk, host-to-device transfer, CPU/platform behavior or their interaction caused the long cold path. It is therefore not a valid single-SN850X versus dual-SN850X RAID estimate.

This host did not expose cgroup io.stat. Schema v4 did capture host-wide /sys/block counters. On the otherwise dedicated host, each cold run's SN850X delta matched the process block-read count to within 0.0002 GiB (39.802-39.965 GiB), each 39.555 GiB reader produced a 39.556-39.557 GiB device delta, and warm runs produced zero except 0.0001 GiB once. This is strong supporting evidence that bytes reached the only visible NVMe, but remains labelled host-wide, unattributed because unrelated host I/O could enter those counters. It is not promoted to the same evidence class as machine 112410's cgroup-attributed MD/member deltas.

The important contribution is methodological: a nominal 600 W, Gen5 x16 RTX 5090 can retain a fast warm runtime while a severely restricted measured H2D path coincides with much slower cold loading. Future slices must retain H2D as an admission/matching variable rather than grouping machines only by GPU power and lane label.

The experiment no longer treats close hardware matching as a prerequisite or represents cross-host pairs as if they were controlled interventions. Normal Vast renters cannot reconfigure the host block devices, and a search for nominally matched machines discards useful variation while still leaving unobserved differences. The revised design treats each physical RTX 5090 host as an observational unit and uses its paired cold wall - warm wall latency as the primary response.

All measured differences remain covariates rather than admission filters: MD layout and member count, exact NVMe identity, negotiated drive generation/width, PCI ancestry and shared bridges, GPU generation/width, actual power limit, H2D, CPU/RAM, kernel, overlay/loop path, filesystem behavior and physical-read evidence. The paired response subtracts much of the warm GPU floor but does not make GPU, CPU or transfer-path effects impossible; those fields remain available for stratification, threshold checks and interactions. The goal is to estimate distributions and look for patterns that reproduce across hosts, not to fit a large causal model to a small sample.

Same-machine reacquisitions have a separate role. They estimate session-to-session and mutable-runtime variance and upgrade older schema-v3 results with schema-v4 block-layer evidence. They are linked to the same machine ID and are not counted as additional independent hosts.

Seven dual-workload v0.37.0 suites were launched with five verified cold/warm pairs per workload:

Role Machine Advertised distinguishing factors at acquisition Final state
New host 56897 Dedicated, Ryzen 7900, 600 W, Gen5 x16, Netac 1 TB Complete; instance destroyed
New host 29706 Dedicated, i5-12400F, 600 W, advertised Gen5 x16, SN850X 2 TB Complete; instance destroyed
New host 30597 One GPU from two, Threadripper Pro 7975WX, 575 W, Gen5 x16, T700 4 TB Complete; instance destroyed
New host 30024 One GPU from eight, EPYC 9654, advertised 600 W, Gen5 x16, Kioxia, 17.4 GB/s advertised disk Complete; instance destroyed
Schema-v3 repeat 108899 Dedicated, previous MiniMax-only single/container-path result Complete; instance destroyed
Schema-v3 repeat 137184 One GPU from two, previous dual-workload SN8100/x8 result Complete; instance destroyed
Schema-v3 repeat 150337 Dedicated, previous Krea-only result Complete; instance destroyed

Schema-v3 RAID machine 33260 was also selected, but the runner rejected it before provisioning because Vast now marks the offer VM-deverified. That safety gate was not bypassed. The other previously sampled schema-v3 machines were unavailable at this check.

2026-09-23: seven-suite completion#

All seven suites exited successfully. Each produced separate Krea 2 and MiniMax H3 schema-v4 JSONs, all five cold attempts in every result passed the cache-eviction/read verification gate, and the runner destroyed every Vast instance after persisting the results. The files are timestamped and do not replace the older schema-v3 observations.

Machine Observed storage path GPU path MiniMax cold / warm / penalty Krea cold / warm / penalty MiniMax model reader
56897 Single Netac 1 TB, PCIe 4.0 x4; cgroup physical 600 W, Gen5 x16 73.896 / 60.403 / 13.492 s 11.984 / 4.015 / 7.969 s 3.013 GiB/s
29706 Single SN850X 2 TB, PCIe 4.0 x4; host-wide physical 600 W, Gen5 x16 77.688 / 63.151 / 14.536 s 14.393 / 4.272 / 10.121 s 3.040 GiB/s
30597 Single T700 4 TB, PCIe 5.0 x4; cgroup physical 575 W, Gen5 x16 91.531 / 82.371 / 9.160 s 12.069 / 5.529 / 6.540 s 6.264 GiB/s
30024 Four Kioxia drives in RAID10 plus two RAID1 arrays; host-wide physical 575 W measured, Gen5 x16 97.375 / 67.072 / 30.302 s 17.264 / 4.289 / 12.976 s 5.164 GiB/s
108899 P310 on PCIe 4.0 x2 behind three upstream hops and a buffered Docker loop; cgroup layers traced 420 W, Gen5 x16 104.692 / 75.979 / 28.713 s 18.083 / 5.021 / 13.062 s 2.702 GiB/s
137184 Single SN8100 2 TB, PCIe 5.0 x4; host-wide physical 575 W, Gen5 x8 71.084 / 63.802 / 7.282 s 9.417 / 4.263 / 5.154 s 10.098 GiB/s
150337 No NVMe exposed; cgroup block device 8:16 575 W, Gen4 x16 102.419 / 67.974 / 34.445 s 18.613 / 4.314 / 14.299 s 0.851 GiB/s

The Krea JSON for each machine uses the same filename stem with vast-krea-2 in place of vast-minimax-h3. The dashboard discovers all fourteen files through DuckDB without a hand-maintained run list.

Longitudinal repeatability#

The reacquired machines are variance controls, not three new independent hosts. Their paired penalties were stable enough to support that separation:

Machine / workload Earlier schema-v3 penalty New schema-v4 penalty Change
108899 / MiniMax 30.004 s 28.713 s -1.291 s
137184 / MiniMax 7.493 s 7.282 s -0.211 s
137184 / Krea 5.200 s 5.154 s -0.046 s
150337 / Krea 13.857 s 14.299 s +0.442 s

Machine 137184 is particularly strong repeatability evidence: its model reader moved only from 10.190 to 10.098 GiB/s for MiniMax and its warm floor moved from 63.762 to 63.802 seconds. Machine 108899's new schema-v4 evidence also explains why its large synthetic headline is misleading: QD32 direct fio reached 10.384 GB/s, but the application reader reached only 2.702 GiB/s through a buffered loop and a PCIe 4.0 x2 drive path.

What the four new hosts add#

Machine 30597 demonstrates why the paired response is necessary. Its MiniMax warm floor was a slow 82.371 seconds, but its cold penalty was only 9.160 seconds; comparing raw cold time would incorrectly assign the GPU/platform floor to storage. Machines 56897 and 29706 add ordinary single-drive paths with 13.492-14.536 second MiniMax penalties.

Machine 30024 is deliberately classified as other MD, not RAID0. Its QD32 fio reached 21.447 GB/s, while QD1 reached 3.249 GB/s, the model reader reached 5.164 GiB/s, and the MiniMax penalty was 30.302 seconds. This is useful adjacent evidence that a very high advertised or high-queue-depth array throughput does not guarantee a small application penalty. It does not estimate the benefit of the two-drive RAID0 planned for the dedicated machine. The dynamic RAID dashboard now recognizes only an explicit raid0 topology as RAID0; RAID10 and RAID1 observations stay available but are excluded from the single-versus- RAID0 summaries.

2026-09-24: replace the remaining current schema-v3 runs#

The longitudinal controls above show that full schema-v3 and schema-v4 reruns on the same physical machine produce consistent application results. Nevertheless, schema v3 cannot be converted into schema v4 after the fact: it did not record per-pass loop, MD and physical-device counters. Later static topology is useful as a provenance-labelled supplement, but it cannot prove where the bytes from an earlier timed pass went.

The remaining latest schema-v3 RTX 5090 results will therefore be replaced in the current analysis only by full five-pair schema-v4 reruns on the exact same physical machine IDs. Existing schema-v3 JSONs remain as longitudinal controls and are never overwritten. Each reacquisition runs both MiniMax H3 and Krea 2 after a single setup so a machine that previously had only one workload gains the other at the same time.

Machine Current schema-v3 coverage 2026-09-24 availability check
45379 Krea 2 Unavailable
55019 MiniMax H3 and Krea 2 Unavailable
112410 Krea 2; MiniMax already has a full schema-v4 repeat Unavailable
147169 MiniMax H3; one-pair schema-v4 validation is not a full replacement Complete; instance 52257980 destroyed
147512 MiniMax H3 and Krea 2 Unavailable
148782 MiniMax H3 Unavailable
150183 MiniMax H3 Unavailable

The acquired machine 147169 was the exact prior Oklahoma host: Ryzen 9950X, ROG Crosshair X870E Hero, 600 W RTX 5090, advertised PCIe 5.0 x16 and the previously identified 4 TB Samsung 9100 Pro. The suite pins ComfyUI v0.37.0, PyTorch 2.9.1+cu130, resident cache policy and five cold-to-warm pairs per workload. The runner persisted new timestamped JSON files and destroyed the instance after both workloads completed.

Machine 147169 completion and block-layer result#

Both workloads completed with five accepted cold-to-warm pairs and zero rejected cold attempts:

Workload Cold mean Warm mean Paired penalty Model reader
MiniMax H3 73.786 s 60.554 s 13.231 s 5.257 GiB/s
Krea 2 10.046 s 4.015 s 6.031 s 7.498 GiB/s

The full MiniMax schema-v3 result was 70.448 / 60.328 / 10.120 seconds with a 6.643 GiB/s reader. The new warm floor changed by only +0.226 seconds, but cold penalty increased by 3.112 seconds and the reader fell by 20.9%. The intervening one-pair schema-v4 validation measured an 11.362-second penalty and 6.199 GiB/s. This preserves the conclusion that the GPU/platform floor is stable while showing meaningful session and pass variation in the storage/container path; schema conversion alone was not responsible for the earlier number.

Schema v4 also resolves what “cold” means on this Vast host. The active model path is a buffered loop0 backed by /var/lib/docker-loop.xfs, above the one-hop PCIe 5.0 x4 Samsung namespace. Across five MiniMax cold passes, the loop layer served 215.236 GB while the physical NVMe recorded only 59.891 GB, or 27.8% of the loop bytes. The three independent MiniMax readers served 127.416 GB at the loop but only 13.868 GB at the NVMe, or 10.9%. One reader reached 7.149 GiB/s while recording zero NVMe reads.

Krea makes the outer-cache behavior especially clear. Its first cold pass sent 17.847 of 18.338 GB to the NVMe. Cold passes two through five each missed the inner model-file cache and recorded about 18.3 GB at the loop, but recorded zero or 0.0008 GB at the NVMe. Thus schema v4 does not merely add metadata: it proves that the existing eviction policy does not produce repeated physical-media-cold samples through Vast's buffered loop. These results remain valid application/container-path measurements, but they must not be labelled physical-NVMe cold throughput.

2026-09-24: variance inside the PCIe 5.0 x4 SSD class#

The dashboard's six-host PCIe 5.0 x4 group spans 4.536-10.098 GiB/s in the independent model reader and 5.732-13.231 seconds in MiniMax paired cold penalty. The negotiated link is therefore a capacity ceiling, not a performance class. All six hosts expose one NVMe namespace through one upstream PCI hop, but they differ in drive/controller, block-queue limits, read-ahead, CPU/platform and container backing path.

Machine Drive Reader Active workload read Active read duration Overlapped by other startup Paired penalty Evidence note
30597 Crucial T700 6.264 GiB/s 2.804 GiB/s 14.161 s 5.001 s 9.160 s Cgroup-attributed NVMe
62211 Samsung 9100 Pro 5.699 GiB/s 2.464 GiB/s 16.154 s 4.378 s 11.776 s Cgroup-attributed NVMe
137184 WD Black SN8100 10.098 GiB/s 3.913 GiB/s 11.035 s 3.753 s 7.282 s Host-wide NVMe counters
147011 Corsair MP700 Elite 4.536 GiB/s 3.460 GiB/s 11.506 s 4.204 s 7.302 s Cgroup-attributed NVMe
147169 Samsung 9100 Pro 5.257 GiB/s 2.394 GiB/s 16.919 s 3.687 s 13.231 s Buffered loop; not media-cold
148782 Samsung 9100 Pro 7.585 GiB/s 4.187 GiB/s 9.540 s 3.809 s 5.732 s Schema v3; lower layer unproven

The realized full-workload read rate explains the end-to-end spread much better than the standalone reader. Across the six host means, active read rate versus penalty has Pearson -0.942 and Spearman -1.000. Across all 30 cold/warm pairs the correlations are -0.926 and -0.951, and the within-host demeaned Pearson correlation is -0.805. The active read duration exceeds the paired penalty by only 3.687-5.001 seconds per host: roughly four seconds of model I/O overlaps CPU/CUDA startup, while the rest appears on the critical path. This quantity is outcome-adjacent and partly mechanical, so it diagnoses the penalty rather than independently predicting an untested machine.

The separate reader still validates storage-path capability. Buffered fio versus reader throughput has Pearson 0.931 and Spearman 0.829 across all six hosts. After excluding the buffered-loop machine it rises to Pearson 0.988 and Spearman 1.000. Reader versus paired penalty is much weaker at Pearson -0.478 and Spearman -0.600 because ComfyUI/PyTorch interleaves reads with CPU tensor work and hides part of the I/O behind other startup. Machine 147011 is the clearest example: its standalone MP700 reader is the slowest in the group, but the application realizes 3.460 GiB/s while active and lands at the same 7.3-second penalty as the much faster SN8100 reader.

Block-queue configuration is a plausible contributor, not yet a causal result. The SN8100 host exposes a 2,048 KiB read-ahead and 1,024 KiB maximum request, while the three directly observed 9100 Pro hosts expose 128 KiB read-ahead and maximum requests of 128, 128 and 512 KiB respectively. The 512 KiB 9100 host is the fastest of those three, but it is also the schema-v3 host and has a different CPU, motherboard and runtime. The setting cannot be credited in isolation. GPU link width is not the explanation either: the fastest reader uses a Gen5 x8 GPU/H2D path, while a Gen5 x16 host has essentially the same paired penalty despite a reader less than half as fast.

Machine 147169 must remain a separate backing-path class. Its loop layer read about 40 GiB per cold pass, but only 6-17 GiB reached the NVMe; one independent reader reached 7.149 GiB/s with zero NVMe reads. Its reader and penalty variance measure the buffered container path and outer-cache state, not repeated physical SSD performance. Machine 148782 remains lower-confidence until an exact schema-v4 rerun is available.

The operational conclusion is to retain SSD generation and lane width as covariates, never as grouping shortcuts. The useful chain is: queue/controller and backing path -> standalone reader capability -> realized active workload read duration -> paired cold penalty, with CPU/platform overlap between the last two steps. The dynamic single-path explorer now exposes the NVMe maximum-request class, in-workload active rate and duration, overlapped startup time and additional cold process CPU so this within-link variance remains visible when new result JSONs are discovered.

2026-09-24: cost-aware global RTX 5090 sampling#

The next inventory search is global rather than US-restricted. GPU power, allocated fraction, reliability, location and download speed are preferences rather than admission gates; physical machine identity remains the independent unit. Network tariff is weighted heavily because the previous suite downloaded 65.9 GB and paid $2.57 at $0.0391/GB, versus only $0.67 for GPU time.

Three new hosts were acquired with a combined expected 65.9 GB transfer charge of about $0.34: dedicated 600 W/QEMU-storage machine 90896 at $0/GB, dedicated 600 W/Intel-NVMe machine 142728 at $0.002604/GB and dedicated 575 W/Kioxia-NVMe machine 150265 at $0.002604/GB. All three completed the full MiniMax H3 and Krea 2 schema-v4 suite. The results are:

Machine Storage label Workload Cold mean Warm mean Paired penalty
90896 QEMU MiniMax H3 120.180 s 64.229 s 55.951 s
90896 QEMU Krea 2 30.115 s 4.284 s 25.831 s
142728 Intel SSDPF2KX076T1 MiniMax H3 105.083 s 68.162 s 36.921 s
142728 Intel SSDPF2KX076T1 Krea 2 23.404 s 4.789 s 18.615 s
150265 Kioxia KXG80ZN84T09 MiniMax H3 79.254 s 64.660 s 14.594 s
150265 Kioxia KXG80ZN84T09 Krea 2 12.995 s 4.300 s 8.695 s

The 600 W fractional host 58084 advertised 24 GB/s storage at the same low tariff, but it never accepted the attached SSH key within ten minutes and the runner destroyed instance 52265805 without running a workload. Samsung 9100 Pro machine 42736 and 990 Pro machine 138147 disappeared before acquisition, so neither created an instance or incurred a charge.

The runner now persists Vast's download and upload tariffs, the sum of prepared model bytes and an estimated prepared-model transfer charge in each future result. The estimate deliberately excludes ComfyUI, Python and CUDA package setup traffic; provider billing remains the authority for total transferred bytes.

Full-fraction inventory expansion#

A subsequent live inventory snapshot returned 23 rentable gpu_frac=1.0 RTX 5090 physical machines. After excluding machines already completed or active, 15 untested candidates charged no more than $0.006511/GB, equivalent to at most about $0.40 for the suite's 61.109 GB of prepared models. Seven untested candidates above that threshold were deliberately excluded; their prepared-model charge alone ranged from about $0.61 to $2.44 per run.

Eleven cost-safe machines were acquired for the same complete two-workload, five-pair schema-v4 suite: 30121, 139086, 151196, 150457, 45827, 148036, 148215, 129526, 110783, 129256 and 144671. Four more candidates (36518, 55535, 151410 and 31538) disappeared between inventory lookup and acquisition and created no billable instance. The acquired cohort spans 400-600 W cards, PCIe 4.0/5.0 GPU links, physical SN850X/990-series and other NVMe devices, plus repeated QEMU-backed storage paths. Results are written to new timestamped JSONs; existing observations are never overwritten and the dynamic dashboard discovers new JSONs automatically.

The expanded search found no rentable full-fraction RTX 5090 with a Samsung 9100 PRO. Machine 148771 is the only visible exact full-fraction match (545 W, Threadripper 9960X, 9100 PRO 2 TB), but it was unavailable. Two rentable exact-drive alternatives were rejected from this cohort because machine 4129 reports gpu_frac=0.5 and machine 42736 reports gpu_frac=0.125.

Official drive-spec covariates#

Manufacturer specifications are now stored as normalized, joinable JSON tables in storage-device-specs.json and storage-device-aliases.json. The current version contains 22 exact product/capacity rows and 29 raw-label aliases. It covers the identified Samsung, Western Digital, Crucial, Corsair, Solidigm/Intel, KIOXIA, Kingston and HIKSEMI devices in the sampled or candidate cohort. QEMU, generic NVMe, ambiguous Netac, Fanxiang and unpublished Phison OEM labels remain explicitly unresolved rather than being assigned guessed specifications.

This adds rated interface, NAND/cache design, sequential and random maxima, and endurance as analysis covariates. It does not replace observed throughput: official maxima use vendor-specific queue depths and fresh-drive conditions, whereas the ComfyUI path also includes filesystem/container layers, block-queue limits, PCIe topology, thermals and contention. The schema and a DuckDB join example are documented in storage-device-spec-database.md.

First completed full-fraction expansion results#

Seven more cost-screened, full-fraction machines completed both five-pair schema-v4 workloads. These files are additional observations; they do not replace earlier captures. The dynamic dashboard selects the latest full result for each physical machine and workload.

Machine Observed storage Workload Cold mean Warm mean Paired penalty
45827 WD Black SN850X 4 TB MiniMax H3 146.915 s 142.064 s 4.851 s
45827 WD Black SN850X 4 TB Krea 2 13.694 s 8.629 s 5.065 s
129256 Samsung 990 EVO Plus 1 TB MiniMax H3 87.207 s 64.724 s 22.483 s
129256 Samsung 990 EVO Plus 1 TB Krea 2 15.900 s 4.271 s 11.629 s
139086 Lexar NM1090 PRO 2 TB MiniMax H3 71.562 s 60.033 s 11.530 s
139086 Lexar NM1090 PRO 2 TB Krea 2 11.089 s 4.165 s 6.924 s
147380 Samsung 990 EVO Plus 1 TB MiniMax H3 108.463 s 66.231 s 42.231 s
147380 Samsung 990 EVO Plus 1 TB Krea 2 25.155 s 4.546 s 20.609 s
148214 QEMU block device; no exposed NVMe MiniMax H3 122.613 s 67.555 s 55.057 s
148214 QEMU block device; no exposed NVMe Krea 2 29.994 s 4.301 s 25.693 s
150215 Phison OEM 2 TB MiniMax H3 88.815 s 71.274 s 17.541 s
150215 Phison OEM 2 TB Krea 2 11.853 s 4.778 s 7.075 s
151196 Fanxiang S880 2 TB path MiniMax H3 89.682 s 66.179 s 23.503 s
151196 Fanxiang S880 2 TB path Krea 2 18.004 s 4.534 s 13.470 s

The new observations expose three useful controls. Machine 147380 negotiated its 990 EVO Plus at only PCIe 3.0 x2 and realized 1.069 GiB/s, versus 3.331 GiB/s for the same model on machine 129256 at PCIe 4.0 x4; their MiniMax penalties were 42.231 and 22.483 seconds respectively. Machine 148214 repeats the slow virtual-storage pattern: no physical NVMe is exposed, the model reader reaches only 0.810 GiB/s, and the penalty is 55.057 seconds. Machine 45827 demonstrates why raw cold time is not the storage outcome: its 400 W GPU has a 142.064-second warm floor, but cold-minus-warm is only 4.851 seconds.

Machine 45827: a small penalty caused by a slow floor, not a faster RTX 5090#

The MiniMax result for machine 45827 can look unusually fast if it is ranked only by cold-minus-warm. It is the opposite when ranked by end-to-end service time:

Measure Machine 45827 Cohort context
MiniMax cold mean 146.915 s slow, not fast
MiniMax warm mean 142.064 s slowest of the 30 latest RTX 5090 hosts; cohort median 66.231 s
MiniMax paired penalty 4.851 s small because cold work overlaps the 142 s execution
Krea cold / warm / penalty 13.694 / 8.629 / 5.065 s its warm Krea floor is also the cohort's slow tail
Independent model reader 5.617 GiB/s healthy PCIe 4.0 SN850X result, but not the fastest reader
Workload active-read throughput 3.013 GiB/s about 13.189 s of active reads per cold run

The cache protocol did not accidentally turn the cold samples warm. All five accepted cold processes report 42.56-42.73 GB of block reads after eviction. Four warm processes report zero block reads and one reports only 1.38 MB. Because DynamicVRAM accesses the staged weights lazily, those reads are distributed across nearly the complete cold execution window. The approximately 13.189 seconds of active reading add only 4.851 seconds to wall time; the remaining approximately 8.337 seconds are hidden behind work already on the critical path. active read duration - paired penalty is only an overlap heuristic, not a measured stage boundary, but it explains why the penalty alone ranks this host misleadingly.

The SSD and host-to-device path do not explain the slow warm floor. The physical WD Black SN850X 4 TB negotiates PCIe 4.0 x4, has no MD RAID, reads the prepared model set at 5.617 GiB/s, and the GPU link negotiates PCIe 5.0 x16 with 47.555 GiB/s pinned H2D. There was no process swap activity. The warm MiniMax process instead consumes about 152.986 CPU-seconds over 142.064 wall seconds, compared with roughly 88.7 CPU-seconds over 76.0 wall seconds on machine 108899. That control has the same Ryzen 9950X, a nearby 420 W power limit, GPU subsystem ID 53061462 and VBIOS 98.02.2E.40.AF. Two other machines with that exact GPU subsystem/VBIOS complete warm MiniMax in 67.974-71.274 seconds. The 400-420 W range and board identity are therefore insufficient explanations for a twofold slowdown.

The remaining evidence localizes the anomaly to the runtime/loaded accelerator path but does not identify one physical cause. Machine 45827 is the only current observation with NVIDIA driver 590.48.01 and kernel 6.8.0-58-lowlatency; those are confounded host-level differences, not proof of a driver or kernel defect. The schema did not sample loaded GPU clocks, actual power draw, temperature, performance state, throttle reasons, CPU frequency/governor or competing GPU processes. A reacquisition would need that time-series telemetry during both warm workloads to distinguish an effective sub-400 W power cap, clock/thermal throttling, host contention and a driver/runtime path. Until then the defensible conclusion is:

Targeted reacquisition and loaded-GPU telemetry#

Machine 45827 was reacquired on 2026-09-24 with the same offer-visible hardware, driver and power limit. A new one-pair schema-v4 run reproduced the anomaly:

Observation Original five-pair result Reacquired one-pair diagnostic
Cold wall 146.915 s mean 140.030 s
Warm wall 142.064 s mean 137.865 s
Cold-minus-warm 4.851 s mean 2.165 s
Independent model reader 5.617 GiB/s 5.624 GiB/s
Pinned H2D 47.555 GiB/s 47.640 GiB/s

The repeat rules out a one-off seed, transient cache state and the earlier benchmark process. During the new cold and warm passes, nvidia-smi dmon recorded the following loaded state. This follows NVIDIA's own guidance to run dmon alongside an inference workload for power, utilization and clock evidence (NVIDIA TensorRT performance measurement guidance).

Loaded telemetry Cold Warm
Samples with at least 50% reported SM activity 133 123
Mean reported SM utilization 98.93% 99.98%
Mean graphics clock 2,693 MHz 2,644 MHz
Mean board power 303.65 W 302.83 W
Mean GPU temperature 67.35 C 67.13 C
Thermal-violation flag 0% 0%

Process telemetry showed only the benchmark Python process during the two passes. A live nvidia-smi -q snapshot reported P1, no active idle/application-clock/hardware/thermal slowdown reason, 2,662 MHz graphics clock, approximately 307 W average power, and 100% utilization. The device exposes all 170 SMs, 32 GB VRAM, default compute mode, no MIG/vGPU mode and no MPS process or active-thread-percentage environment limit. CPU frequency telemetry showed cores boosting above 5.3 GHz rather than remaining at the 600 MHz idle floor. This rules out the simple versions of thermal throttling, a stuck low clock, another GPU process, MPS/MIG partitioning and a non-boosting CPU.

A committed wall-clock CUDA probe adds direct evidence that the abnormal warm floor is an accelerator-path problem. It uses the identical PyTorch 2.9.1/CUDA 13.0 userspace on machine 45827 and on an otherwise idle, second RTX 5090 allocated with machine 56934. The control is a 600 W card on driver 595.84, so power and driver remain explicit confounders rather than being silently treated as matched:

Probe Machine 45827, 400 W / driver 590.48.01 Machine 56934 GPU 1, 600 W / driver 595.84 Difference
BF16 16,384-square GEMM 180.30 TFLOP/s 239.14 TFLOP/s -24.6%
FP32 with TF32 enabled, 12,288-square GEMM 50.81 TFLOP/s 124.18 TFLOP/s -59.1%
2 GiB device-to-device copy 561.41 GB/s 767.40 GB/s -26.8%

The control is not sufficient to call the 590 driver causal: its higher power limit explains some of the BF16 and memory difference, and the driver, board and host also differ. It is nevertheless decisive for the storage question. Machine 45827 is slow and unusually variable in device-local operations that do not touch the SN850X, host page cache or PCIe H2D path. The nearby 420 W machine with the same GPU subsystem ID and VBIOS, 108899 produced a 75.979-second warm MiniMax mean, so nominal power alone still cannot explain 137.865 seconds. The remaining credible boundary is a machine-specific accelerator/runtime condition, with the 590.48.01 driver plus low-latency kernel and latent board/firmware behavior still confounded.

The raw diagnostic artifacts live under ../diagnostics/: the complete one-pair result, per-second GPU/process telemetry and both compute-probe JSONs are retained separately from canonical five-pair dashboard inputs. Machine 45827 should be flagged as a compute outlier and excluded from claims about SSD or RAID effects; it remains useful as evidence that a small cold-minus-warm penalty can be manufactured by a pathological warm floor.

Four further full-fraction suites completed while the first analysis pass was running:

Machine Observed storage Workload Cold mean Warm mean Paired penalty
30121 QEMU block device; no exposed NVMe MiniMax H3 140.618 s 107.040 s 33.578 s
30121 QEMU block device; no exposed NVMe Krea 2 19.761 s 6.787 s 12.974 s
148036 Samsung 990 PRO 2 TB MiniMax H3 77.856 s 70.085 s 7.771 s
148036 Samsung 990 PRO 2 TB Krea 2 10.316 s 4.768 s 5.548 s
148215 QEMU block device; no exposed NVMe MiniMax H3 125.043 s 67.919 s 57.124 s
148215 QEMU block device; no exposed NVMe Krea 2 29.279 s 4.294 s 24.985 s
150457 Samsung 990 PRO 4 TB MiniMax H3 78.221 s 68.975 s 9.246 s
150457 Samsung 990 PRO 4 TB Krea 2 10.444 s 4.766 s 5.678 s

Machines 129526 and 110783 never progressed beyond Vast's provider-side loading state and were destroyed by the runner before any workload. Machine 144671 downloaded the first 19 GiB model at only about 10 MiB/s; its remaining instance 52270786 was explicitly destroyed after the benchmark process disappeared. The Vast account was then checked and had zero active instances. No result JSON was produced for those three machines.

Expanded-cohort analysis#

With the completed files discovered dynamically, the latest-host dataset contains 30 independent MiniMax machines, 24 Krea machines, 23 machines with both workloads and 26 non-MD-RAID MiniMax paths. On those 26 single/container paths, buffered fio predicts the independent model reader at Pearson 0.946 and Spearman 0.928; the model reader predicts paired MiniMax penalty at -0.696 and -0.752. On the narrower 19-host 575-600 W, H2D-at-least-10 GiB/s, GPU-x16 slice, the reader-to-penalty Spearman relationship is -0.932.

The official-spec join identifies 18 hosts without guessing virtual, generic or ambiguous multi-device paths. Manufacturer sequential rating predicts model-reader throughput at Pearson 0.743 and Spearman 0.769, but is weaker against final penalty (-0.481 and -0.658). Realized model reads range from 16% to 83% of the official maximum. This supports using the data-sheet rating as a prior and the measured reader as the acceptance test.

The two additional exact Linux identities are Crucial P310 1 TB (CT1000P310SSD8) and Lexar NM1090 PRO 2 TB. Official manufacturer sources were added to the JSON spec table. The Lexar host is especially useful: the drive is rated for PCIe 5.0 x4 and 14.0 GB/s, but negotiates PCIe 4.0 x4 on machine 139086 and realizes 3.298 GiB/s. This makes the difference between product capability and installed-path realization explicit.

The result also replicates between applications: across 23 dual-workload hosts, MiniMax and Krea paired penalties correlate at Pearson 0.985 and Spearman 0.975, while their model readers correlate at 0.978 and 0.962. The expanded interpretation, outlier explanations and owned-node acceptance gates are recorded in rtx5090-single-path-multivariable-analysis.md.

Disk × setup × performance dashboard view#

The dynamic RAID analysis dashboard now includes a row-level view over the latest deduplicated results for both MiniMax H3 and Krea 2. It keeps the advertised disk identity and all visible NVMe devices beside the negotiated topology/link, upstream-hop inference, buffered-loop evidence, CPU/RAM, GPU PCIe/H2D/power path, kernel/runtime and observed workload performance. Filters select workload, exact disk identity and layout; the focus can be changed between paired cold penalty, cold/warm wall time, model-reader throughput and the three fio views. This is deliberately a host-level comparison rather than an SSD-product leaderboard: installed-path covariates remain visible instead of being averaged away.

Why the negotiated PCIe 5.0 x4 group still varies#

The latest independent MiniMax cohort contains seven non-RAID hosts whose exposed SSD negotiated PCIe 5.0 x4. That label is a link-width ceiling, not an end-to-end performance class:

Measure Observed range Fastest / slowest
Buffered fio after invalidation 4.876–13.027 GB/s 2.67×
Actual model reader 4.536–10.098 GiB/s 2.23×
Paired cold penalty 5.732–17.541 s 3.06×

Within these seven hosts, buffered fio versus the independent model reader has Pearson 0.900 and Spearman 0.679; direct QD32 fio versus the model reader is 0.831 and 0.536. The immediate explanation for most of the reader spread is therefore the realized mounted/container filesystem path, not a difference in the nominal PCIe generation. The remaining physical causes are not isolated by this cross-host sample. They include different SSD/controller implementations, queue limits, filesystem/runtime behavior and contention.

Queue configuration is a concrete clue, but not yet causal evidence. The fastest host in this class exposes 2 MiB read-ahead and a 1 MiB maximum request, while the four 128 KiB-maximum-request hosts average only 5.898 GiB/s. Those groups also use different drives and machines, so this cannot be called a queue-setting effect. The three Samsung 9100 Pro 4 TB observations further demonstrate installed-path variance: their model readers are 5.257, 5.699 and 7.585 GiB/s. The slowest has an observed active buffered loop; the fastest has a 512 KiB maximum request but only schema-v3 loop evidence.

Application penalty has more remaining variance than the model reader. Within the seven-host class, reader-versus-penalty correlation is only Pearson -0.336 and Spearman -0.393. Once gross storage pathology is removed, CPU deserialization, many-file access, GPU transfer/overlap and other host activity can still affect the final cold-minus-warm result. PCIe 5.0 x4 is neither a sufficient performance guarantee nor the variable that explains the class internally.

Vast pre-rental filtering boundary#

A read-only live query on 2026-09-24 confirmed that Vast's /bundles/ endpoint returns disk_name and accepts exact eq and multi-value in filters for it. It also accepts the documented disk_bw filter. It does not expose an SSD negotiated-generation field. Vast's pci_gen and pcie_bw offer fields describe the CPU-to-GPU PCIe path, not the disk path, as documented in the official CLI source. disk_name is not listed among the CLI's documented offer fields even though the live backend accepted it, so this project should use the raw API/client-side selection rather than depend on it as a stable CLI contract.

The practical pre-rental strategy is therefore:

  1. Filter normal constraints plus disk_bw, download cost and full-GPU preference.
  2. Match disk_name against the official SSD-spec alias database and prefer known PCIe 5.0 x4 models.
  3. Treat generic labels such as nvme as unknown rather than PCIe 5.
  4. After acquisition, verify 32.0 GT/s PCIe x4, visible devices, topology, upstream hops and buffered-loop evidence before admitting the host to the PCIe 5.0 x4 analysis.

This remains a capability prefilter, not proof of the installed link: an earlier Lexar NM1090 Pro host is a PCIe 5.0-capable product that actually negotiated PCIe 4.0 x4.

2026-09-24: additional PCIe 5 SSD sampling#

Four more RTX 5090 hosts completed the full five-pair MiniMax H3 and Krea 2 suite. Every accepted cold sample passed the process-block-read threshold and every warm pass remained in the same ComfyUI process with a different seed. The dashboard discovered all eight result files from the benchmark directory without a hand-maintained row.

Machine Mounted model-storage path GPU / allocation caveat Workload Cold Warm Penalty Model reader Buffered fio Direct QD32 fio
143669 Single Samsung 9100 Pro 2 TB, PCIe 5.0 x4; separate 1 TB OS SSD 600 W; GPU at PCIe 5.0 x8; allocation exposed two GPUs and the suite used GPU 0 MiniMax H3 68.216 s 61.126 s 7.090 s 9.755 GiB/s 10.976 GB/s 14.717 GB/s
143669 same same Krea 2 9.439 s 4.015 s 5.424 s 9.827 GiB/s 10.707 GB/s 14.721 GB/s
56934 Single Crucial T710 2 TB, PCIe 5.0 x4 600 W; GPU at PCIe 5.0 x8; allocation exposed two GPUs and the suite used GPU 0 MiniMax H3 66.691 s 61.088 s 5.603 s 5.522 GiB/s 10.623 GB/s 14.844 GB/s
56934 same same Krea 2 8.458 s 4.065 s 4.393 s 5.565 GiB/s 10.597 GB/s 14.815 GB/s
148771 Single Samsung 9100 Pro 2 TB, PCIe 5.0 x4 one 545 W GPU; GPU at PCIe 5.0 x16 MiniMax H3 70.577 s 64.330 s 6.248 s 9.596 GiB/s 10.006 GB/s 14.717 GB/s
148771 same same Krea 2 9.082 s 4.267 s 4.814 s 9.730 GiB/s 10.793 GB/s 14.700 GB/s
145570 Linux MD RAID0 over two WD Black SN8100 2 TB drives, both PCIe 5.0 x4, 512 KiB chunk one GPU allocated from a four-GPU host; 450 W; GPU at PCIe 5.0 x16 MiniMax H3 103.782 s 72.646 s 31.136 s 6.005 GiB/s 6.283 GB/s 22.346 GB/s
145570 same same Krea 2 17.737 s 4.782 s 12.954 s 5.864 GiB/s 6.730 GB/s 24.285 GB/s

This expansion changes three parts of the interpretation:

The direct storage decision remains unchanged: buy one 4 TB 9100 Pro first and accept it using both the independent model-reader test and the complete cold/warm suite. Do not buy a second drive solely to obtain RAID0; add capacity or RAID later only if the owned-node measurements miss the service target. The new samples strengthen the sufficiency claim for a high-performing single Gen5 SSD, but they do not provide a same-machine estimate of the incremental benefit of RAID0.

Across the now eleven non-RAID MiniMax paths that expose at least one SSD negotiated at PCIe 5.0 x4, the application-shaped reader spans 4.536-10.098 GiB/s and the paired penalty spans 5.603-17.541 seconds. Buffered fio still tracks the reader directionally (Pearson 0.784, Spearman 0.700), while reader versus penalty is weak inside this already-fast link class (Pearson -0.355, Spearman -0.300). The broad cohort establishes that gross storage performance matters; the within-Gen5 slice establishes that once the worst storage paths are removed, overlap, CPU-side loading and platform/runtime variance become large relative to the remaining storage difference.

One other exact PCIe 5 offer, machine 46165 with a 4 TB 9100 Pro, never accepted the provisioned SSH key and was destroyed without producing a result. A replacement candidate became provider-deverified before provisioning and was not rented. The runner now rejects explicit offers that allocate more than one GPU or advertise less than 48 GB of CPU RAM, preventing those offer-shape mistakes in later samples.

2026-09-27: focused multivariable reanalysis of single-path variance#

The current dataset has grown to 34 independent MiniMax machines, 30 Krea machines, 29 machines with both workloads and 29 MiniMax single/no-MD-RAID paths. The canonical interpretation has been rewritten in rtx5090-single-path-multivariable-analysis.md. This entry records the changes in reasoning without rewriting the earlier chronological observations.

Correction: “startup overlap” is a derived residual#

The API field startup_overlap_s is calculated as active_read_duration_s - cold_penalty_s. It is not an independently observed stage duration. Earlier entries used “overlap” as shorthand and sometimes gave that residual too much explanatory weight. A positive residual is consistent with concurrent reading and other startup work, but cannot prove the amount of overlap or distinguish concurrency from instrumentation and stage-boundary effects.

The read accounting is otherwise unusually tight. Single-path cold processes read 39.53-43.17 GiB, with a 39.84 GiB median. active_read_duration_s agrees with process bytes divided by active rate to a median absolute difference of 0.015 seconds. Across 29 hosts, the descriptive fit penalty ~= -4.92 + 41.848 / active_read_GiB_s has in-sample R² 0.994. Because the rate is measured inside the timed run and shares its read volume with the response, this is a diagnostic localization, not a causal SSD-speed model.

Aggregate, restricted and within-class relationships#

On all 29 single paths, Spearman correlations with MiniMax penalty are -0.604 for Vast's advertised disk, -0.761 for direct QD32 fio, -0.837 for buffered fio, -0.779 for the standalone model reader and -0.980 for in-workload active rate. Additional cold CPU is +0.800; warm runtime and H2D are weak at +0.264 and -0.130. Ten thousand host-level bootstrap resamples produce intervals that exclude zero for the first five storage/active metrics and additional CPU, but include zero for warm runtime and H2D.

Removing the five virtual/no-exposed-NVMe paths weakens the advertised, QD32, buffered and model-reader rank relationships to -0.337, -0.601, -0.718 and -0.635 respectively. Active rate remains -0.966. The broad relationship is therefore partly a contrast between pathological virtual paths and physical NVMe, not a clean marginal-speed curve among healthy SSDs.

Within exact single-device link classes, independent-probe relationships weaken further. The nine PCIe 5.0 x4 hosts span 5.60-17.54 seconds and have model-reader-versus-penalty Spearman -0.250. The eleven PCIe 4.0 x4 hosts span 4.85-36.92 seconds and have Spearman -0.091. Active rate remains almost perfectly ordered in both groups (-1.000 and -0.991). This supports treating link class as a ceiling and path descriptor rather than an explanation of within-class variance.

Predictive stress test#

Nested leave-one-host-out validation standardizes predictors within each training fold and selects ridge regularization inside that fold. Independent storage probes together reach MAE 5.434 seconds, RMSE 6.761 seconds and predictive R² 0.818. The captured platform variables alone reach RMSE 16.221 and R² -0.047. Adding them to storage improves MAE to 4.844 but does not improve RMSE or R² (6.855 and 0.813). An inverse active-rate diagnostic reaches RMSE 1.319 and R² 0.993, as expected from the read accounting above.

Categorical hardware labels do not isolate the missing causes. Path categories alone reach predictive R² 0.358, drive family 0.449, storage probes plus path categories 0.717, and storage plus path plus drive labels 0.760. These high-dimensional fits have only 29 machines and are sensitivity checks, not stable effect estimates.

Resolution of the apparent slow-disk/fast-penalty cases#

Machine 148036 is the cleanest apparent contradiction. Its standalone 990 PRO reader is only 2.001 GiB/s, but ComfyUI realizes 3.251 GiB/s during the workload and its 12.182-second read-active interval accompanies a 7.771-second penalty with only 0.120 seconds of paired SD. The active-rate fit predicts about 7.95 seconds. The standalone probe understated the loaded path; the data does not identify which filesystem or access- pattern detail caused that difference.

Machine 150215 shows the reverse: 6.374 GiB/s standalone but only 1.796 GiB/s during ComfyUI, yielding a 17.541-second penalty. Machine 45827 remains the largest residual: its active rate predicts about 8.97 seconds but the observed mean is 4.851 seconds. Its five-pair SD is the cohort maximum at 2.560 seconds, its warm floor is pathological, and device-local diagnostics show an accelerator/runtime outlier. Its small penalty must not be promoted into evidence that the SN850X hid a measured amount of storage time.

The reanalysis therefore narrows rather than reverses the storage conclusion: cold-penalty variance is mostly storage-path variance expressed through actual workload delivery, but it is not SSD-spec variance. Buffered fio is the strongest independent screening proxy in the current sample. The full paired workload and lower-layer read evidence remain the acceptance test.