Why cold-start penalties vary across single-storage RTX 5090 hosts#

Status: Reference Last verified: 2026-09-27 Canonical for: multivariable interpretation of RTX 5090 single-storage cold-start variance Population: latest fixed-protocol MiniMax H3 result per physical RTX 5090 machine, ComfyUI v0.37.0 Independent unit: physical machine Dashboard: interactive storage-path explorer Experiment record: RTX 5090 multivariable investigation journal Owned-node acceptance targets: RTX 5090 RAM and cost/performance decision

Executive answer#

The large spread in cold-start penalty is real, repeatable and mostly associated with how quickly the complete host path delivers model bytes to ComfyUI. It is not explained reliably by SSD model, advertised bandwidth, PCIe generation or RAID label alone.

That distinction resolves the apparent contradiction in the data. A slow standalone disk test can coexist with a small penalty when ComfyUI receives the same files faster than that probe did. A fast standalone test can coexist with a large penalty when the loaded application path delivers them slowly. The useful unit of analysis is therefore not “the SSD”; it is the installed path:

drive + negotiated link + filesystem/container + host state + loader access pattern
    -> delivery while ComfyUI is reading
    -> read-active duration
    -> incremental cold-start penalty

Storage is still important. Independent storage probes predict the penalty well across the broad cohort, and buffered fio is the strongest current screening probe. But their relationships weaken within healthy PCIe classes, where the remaining host differences are bundled together and cannot be assigned to one component from this observational sample.

The fleet implication is narrow: the evidence supports one correctly attached and validated Gen5 model drive; it does not support adding RAID merely to improve a synthetic number. The node itself must pass the full ComfyUI workload and physical-read checks. A small cold - warm penalty also does not necessarily mean a fast service—the cold and warm wall times must remain visible beside it.

What is being measured#

The current dataset contains 34 independent MiniMax hosts and 30 Krea hosts. Twenty-nine machines ran both workloads. This analysis uses the latest complete MiniMax result from each of the 29 machines classified as single/no-MD-RAID storage paths.

Each benchmark runs five resident-cache pairs:

evict resolved model-file pages
  -> fresh-process cold run with verified process block reads
  -> different-seed warm run in the same process
  -> stop ComfyUI

The primary response is the mean paired cold wall - warm wall time. It estimates the incremental cost of the cold process relative to the resident run on that host. It is not raw user-perceived cold latency, and pairing does not make CPU, GPU transfer, driver or host-state effects impossible.

The measurements form an evidence hierarchy:

Measurement What it observes Proper use
SSD specification and Vast disk score Product capability or provider headline Pre-rental prior
Direct fio Block path under a synthetic access pattern Media/controller diagnostic
Buffered fio Mounted buffered path after invalidation Best current independent screening probe
Standalone model-file reader Real model files, normal buffered reads Application-shaped independent probe
In-workload active read rate and duration Reads inside the timed ComfyUI cold process Outcome diagnosis, not independent prediction
Paired cold penalty Incremental end-to-end application time Primary response and admission result

Schema v4 additionally records whether process reads reached a loop, MD or physical-NVMe layer. Schema-v3 observations remain valid process-cache-cold application results, but cannot prove physical-media coldness through an outer buffered loop. All comparisons here describe the observed complete path, not intrinsic NAND performance.

What the data establishes#

The host-to-host spread is larger than measurement noise#

The 29 single-path MiniMax penalties span 4.851-57.124 s, a 52.273 s range. Across all 34 current MiniMax results, the median within-run SD of the five paired penalties is 0.525 s; the mean is 0.761 s and the maximum is 2.560 s. Across seven machine reacquisitions, the median absolute change is 0.211 s and the maximum is 3.112 s.

The ordering also transfers to another workload. Across the 29 machines with both results, MiniMax and Krea penalties correlate at Pearson 0.983 and Spearman 0.975; their standalone model-reader rates correlate at 0.983 and 0.972. Individual runs can be noisy, but the overall spread is persistent host-path behavior rather than seed noise.

Storage predicts the broad cohort better than it explains healthy subclasses#

The table shows Spearman correlations with MiniMax cold penalty. The first two columns progressively remove the slow virtual-storage extreme; the last two compare hosts only within one negotiated physical-link class.

Measurement All single paths, n=29 Physical NVMe, n=24 PCIe 5.0 x4, n=9 PCIe 4.0 x4, n=11
Vast advertised disk -0.604 -0.337 +0.033 -0.073
Direct fio QD32 -0.761 -0.601 -0.533 -0.355
Buffered fio -0.837 -0.718 -0.683 -0.364
Standalone model reader -0.779 -0.635 -0.250 -0.091
In-workload active rate -0.980 -0.966 -1.000 -0.991

In the full cohort, 10,000 host-level bootstrap resamples (seed 20260927) keep the buffered-fio interval (-0.937 to -0.632) and standalone-reader interval (-0.909 to -0.547) below zero. Those are strong associations in this sample, not causal effects.

Five hosts expose no NVMe and present a virtual block path. Removing them weakens every independent storage relationship. Restricting further to one negotiated link class weakens the standalone-reader relationship again, even though penalties still span 5.60-17.54 s on PCIe 5.0 x4 and 4.85-36.92 s on PCIe 4.0 x4.

The pattern is coherent rather than contradictory:

  1. avoiding a pathological storage path matters greatly;
  2. independent path measurements remain useful among physical drives; and
  3. within an already healthy link class, product and topology labels are too coarse to explain the remaining application variance.

The six single Samsung 9100 Pro observations make the last point concrete. Despite sharing a drive family, they span 5.26-9.76 GiB/s in the standalone reader, 2.39-4.19 GiB/s during ComfyUI and 5.73-13.23 s of penalty. One has an active buffered loop; another uses an Intel platform without that observed loop. The sample does not isolate firmware, queue configuration, filesystem behavior, CPU work or contention.

Actual workload delivery bridges storage and latency#

MiniMax process block-read volume is almost fixed: minimum 39.53 GiB, median 39.84 GiB, maximum 43.17 GiB. The recorded read-active duration agrees with process bytes divided by active rate to a median absolute difference of 0.015 s and a maximum of 0.169 s.

Consequently, a descriptive fit across the 29 hosts is extremely tight:

cold penalty seconds ~= -4.92 + 41.848 / active_read_GiB_per_second

Its in-sample R² is 0.994. This localizes the elapsed time; it does not show what hardware change will produce a faster rate. The predictor is measured inside the timed cold process and shares its nearly fixed read volume with the response, so it cannot be used as an independent pre-rental predictor.

Independent probes approximate that workload rate imperfectly. Their Spearman correlations with active rate are 0.843 for buffered fio, 0.791 for direct QD32, 0.779 for the standalone model reader and 0.592 for Vast's advertised score. That ordering is why buffered fio is a useful screen while the application run remains decisive.

One derived field requires special care. The API calculates startup_overlap_s as:

active_read_duration_s - (cold_mean_s - warm_mean_s)

It is a residual, not an independently observed overlap interval. A positive value is consistent with read activity coexisting with other startup work, but it cannot measure the amount of CPU/GPU overlap or identify a stage boundary. Earlier journal entries used “overlap” as shorthand; this definition is authoritative.

Why slower probes sometimes accompany smaller penalties#

The apparent exceptions are mainly disagreements between an independent probe and delivery inside the real loader, plus one abnormal compute host:

Machine Standalone reader In-workload rate Penalty Interpretation
148036, Samsung 990 PRO 2.001 GiB/s 3.251 GiB/s 7.771 s ComfyUI receives bytes much faster than the standalone probe
150215, Gen5 Phison 6.374 GiB/s 1.796 GiB/s 17.541 s The standalone probe overstates delivery in ComfyUI
45827, WD SN850X 5.617 GiB/s 3.013 GiB/s 4.851 s Noisy accelerator/runtime outlier; not a clean storage result

Machine 148036 is the clearest resolution. Its read-active interval is 12.182 s, and the inverse-active- rate fit predicts about 7.95 s, close to the observed 7.77 s. Its paired SD is only 0.120 s. The data does not identify whether readahead, access ordering, cache state or another path detail caused the probe/workload difference, but no special SSD exception is needed to explain the application result.

Machine 45827 needs the opposite caution. Its active-rate fit predicts about 8.97 s, leaving the cohort's largest residual at -4.12 s. Its five penalties range from about 1.67 to 8.43 s, with the cohort's largest paired SD (2.560 s). Its MiniMax cold and warm means are both extremely slow (146.915 s and 142.064 s), and separate device-local probes identify an abnormal accelerator/runtime path. Krea also records a small incremental penalty, but this machine is not evidence that an SN850X eliminates storage time. It is evidence that a small paired penalty can coexist with unacceptable end-to-end service time.

Multivariable check#

Pairwise correlations can reflect the virtual-storage tail or collinear hardware. To test whether the captured platform variables add predictive information, the analysis uses leave-one-host-out validation. Predictors are standardized inside each training fold; ridge regularization is selected inside that fold.

Predictor set LOOCV MAE LOOCV RMSE Predictive R²
Intercept only 13.430 s 16.417 s -0.073
Buffered fio alone, inverse 5.625 s 6.829 s 0.814
All independent storage probes 5.434 s 6.761 s 0.818
Captured platform variables only 13.008 s 16.221 s -0.047
Storage plus captured platform variables 4.844 s 6.855 s 0.813
In-workload active rate, inverse; diagnostic 0.903 s 1.319 s 0.993

The captured platform set contains warm runtime, H2D, GPU power, CPU-memory bandwidth and cold-process core usage. It improves MAE slightly when added to storage, but not RMSE or predictive R². A separate high-dimensional stress test reaches R² 0.358 from path categories, 0.449 from drive family and 0.760 from storage probes plus path and drive labels—still below the 0.818 from continuous storage probes.

These results do not prove that CPU, kernel, GPU or filesystem state is irrelevant. They say that the captured platform fields and sparse categorical labels do not explain additional hosts reliably in this 29-machine sample. Important loaded state and contention remain incompletely observed, and the categories are highly collinear.

Interpretation#

Taken together, the evidence supports three claims. Complete storage-path quality materially affects cold penalty; buffered fio is the best current independent screening proxy; and actual in-workload rate diagnoses where the timed cold process spent its time. A healthy single Gen5 path has reached the useful latency tier repeatedly, so RAID is not a prerequisite for the fleet design.

The evidence does not show that SSD model or advertised speed determines the penalty, that a synthetic probe can accept a node, or that in-workload rate independently predicts an untested host. It also does not show that RAID can never help or isolate a causal effect for each CPU, kernel, PCIe and container category.

The central conclusion is therefore narrow: the observed cold-penalty spread is best described as complete-path delivery variance, not simply SSD-specification variance.

Operational use#

For rental screening, prefer a known physical NVMe, full negotiated links, a low transfer tariff and a healthy provider record. Treat the SSD specification and Vast score as priors. After acquisition, record the block topology and run buffered fio plus the model reader. Admit the host only from the repeated application result, with raw cold and warm times visible beside the paired penalty.

For each owned node:

  1. Confirm the model 9100 Pro negotiates PCIe 5.0 x4 on the intended CPU-connected path.
  2. Confirm the RTX 5090 link and pinned H2D measurement under the production BIOS configuration.
  3. Put models on a direct filesystem and corroborate cold reads at the physical namespace.
  4. Run both complete five-pair workloads and apply the MiniMax and Krea thresholds from the owned-node decision.
  5. If the node misses, inspect in-workload delivery, block/loop layers, thermals, loaded GPU state and CPU loading before treating a second SSD or RAID as the remedy.

There is deliberately no universal reader-throughput gate. A node should not fail because one independent probe is below a historical threshold if its repeated, physically corroborated application result passes. Conversely, fast fio or a fast data sheet cannot compensate for a failing application result.

Interpretation limits#

This is an observational cross-host study, not a same-machine intervention. Bootstrap intervals quantify sampling uncertainty but do not remove omitted-variable bias. The sample supports path-level associations, not isolated causal effects for every CPU, SSD, kernel, loop or PCIe category. The result also applies to the current MiniMax and Krea loaders and their roughly 40/17 GiB model sets; a different loader or access pattern can move the useful storage tier.

The RAID-specific evidence and fleet decision remain in the broader RAID analysis. Raw acquisitions, metadata backfills and the chronological changes in interpretation remain in the experiment journal.