Dual-worker model storage and NFS/NAS decision#

Status: Current deployment decision; local RAID 0 remains optional pending a same-machine test Decision date: 2026-09-23 Scope: Two local RTX 5090 ComfyUI workers connected initially through their onboard 5GbE ports Primary decision: Keep active models on a local NVMe cache in each worker; use one shared HDD-backed NFS/NAS as the model origin; do not add a local HDD to each worker and do not serve the normal customer workload directly from NFS

Decision in one page#

Use this hierarchy:

Tier 0: GPU VRAM and Linux RAM/page cache
   │
Tier 1: local Samsung 9100 Pro cache in each worker
        one 2 TB drive initially on the new worker; retain an existing two-drive cache
        active model files use a direct filesystem on CPU-connected Gen5 x4 storage
   ▲
   │ staged and verified before a job is admitted
   │
Tier 2: one shared NFS/NAS model origin
        large CMR HDD capacity; preferably a two-drive mirror
   ▲
   │ backup/synchronization, not the live model path
   │
Tier 3: independent off-machine copy or reproducible upstream sources
        manifests, private models, configuration and irreplaceable data

The initial layout for a newly configured worker should be:

CPU Gen5 x4 M.2               ── Samsung 9100 Pro 2 TB ── /srv/model-cache
chipset Gen4 x4 M.2           ── separate 2 TB OS SSD   ── Linux, containers, logs, outputs
second CPU Gen5 x4 M.2        ── empty initially; Tomahawk only, reserved for conditional expansion
onboard 5GbE                  ── NFS model origin and worker-to-worker/site traffic

Both the promotional MSI PRO X870E-S EVO and the upgraded Tomahawk provide onboard 5GbE, one CPU-connected Gen5 x4 M.2 path, chipset Gen4 M.2 storage and a CPU-connected Gen5 x16 GPU slot. The RM450 Tomahawk upgrade is therefore not required for the initial one-model-drive layout. It buys the independent second CPU Gen5 x4 M.2 path needed for a later two-drive experiment or 4 TB striped cache.

Keep the Tomahawk while that expansion is genuinely undecided: RM450 is inexpensive option value compared with replacing the motherboard later. If the local test is completed before ordering, RAID is rejected and a 2 TB cache with headroom is demonstrated to be sufficient, retain the included PRO X870E-S EVO and save RM450. MSI's PRO X870E-S EVO specification confirms its 5GbE controller, single CPU Gen5 x4 M.2 socket, chipset M.2 sockets and Gen5 x16 GPU slot.

Do not install an HDD in each worker as an intermediate cache. It duplicates capacity, adds another copy and eviction policy, increases heat and failure monitoring, and is still much slower than the local NVMe. The shared NAS is already the correct next tier beneath local NVMe.

Do not point the production ComfyUI model directories directly at NFS. A scheduler or cache agent should stage every required model bundle to local NVMe, verify it, atomically publish it, and only then admit the job. A NAS outage must prevent a cache miss from being filled; it must not stop a worker from serving models that are already local.

The ordered nodes (updated 2026-09-26; the record is the node decision doc) use no RAID and put the OS on a separate Samsung drive in the second CPU slot, not on a chipset Gen4 SSD as drawn above:

Node M2_1, CPU Gen5 x4 M2_2, CPU Gen5 x4
PC1 9100 Pro 2 TB, model cache 9100 Pro 2 TB: OS, Docker, outputs, later LLM weights
PC2 9100 Pro 2 TB, OS and model-cache partitions until the OS drive arrives 9100 Pro 1 TB OS drive, bought later
PC3 9100 Pro 2 TB, model cache 9100 Pro 1 TB, OS

Each node therefore has one 2 TB model cache. Advertise available models and free cache capacity to the scheduler so it can route each job to the better-prepared worker. The NAS, NFS and staging design below is unchanged by the order, apart from serving three workers instead of two.

Why direct NFS is not the normal serving path#

The current MiniMax model and input set is 42,471,968,783 bytes. The lower-bound transfer times are:

Source path Planning throughput Ideal time for 42.47 GB Appropriate role
Local Samsung 9100 Pro, measured physical reader 5.699 GiB/s about 7 seconds of raw reading Active model cache
5GbE raw line rate 625 MB/s 68 seconds Mathematical ceiling only
Healthy 5GbE NFS planning rate 500-550 MB/s 77-85 seconds Cache fill, not request-time model loading
Seagate BarraCuda 2 TB ST2000DM008 maximum 220 MB/s about 3.2 minutes Cheap staging/origin media only
Seagate BarraCuda 4 TB ST4000DM004 maximum 190 MB/s about 3.7 minutes Cheap staging/origin media only
Selected Seagate IronWolf 4 TB ST4000VN006 maximum 202 MB/s about 3.5 minutes NAS capacity/origin media

The 5GbE figures are decimal: 5 Gbit/s is 625 MB/s before Ethernet, TCP and NFS overhead. They are not 5 GB/s. A single HDD therefore will not saturate 5GbE. This is acceptable when copying is asynchronous; it is unacceptable if every cold customer request waits for that copy.

Seagate specifies 220 MB/s for the 2 TB BarraCuda and 190 MB/s for the 4 TB model. It also identifies both listed models as SMR drives with a desktop-class 55 TB/year workload limit and 2,400 power-on hours per year. They are inexpensive but are not the balanced foundation for a mirrored 24/7 NAS. Seagate's IronWolf line uses CMR, is rated for 8,760 power-on hours and a 180 TB/year workload, and is the more appropriate class for the shared origin. Relevant manufacturer references are the 2 TB BarraCuda manual, BarraCuda recording-technology sheet, and IronWolf specification.

What the latest RAID evidence means#

The earlier purchase record selected two Samsung 9100 Pro drives in RAID 0. That recommendation is superseded by the completed single-NVMe versus RAID 0 experiment.

The fixed dual-workload matrix did not show the required repeated application-level RAID advantage. A later schema-v4 physical control showed that one direct-attached Gen5 NVMe could serve every cold byte from the physical device and complete MiniMax with a 7.30-second cold penalty. A separate physical 4 TB Samsung 9100 Pro control also completed every run without RAID. RAID greatly improved some synthetic results but did not reliably improve the complete ComfyUI workloads.

The current 2 TB-only purchasing constraint changes the capacity choice, not the performance conclusion:

Worker option Usable model-cache capacity Incremental economics Decision
One 2 TB Samsung 9100 Pro + separate 2 TB OS SSD About 2 TB before formatting/headroom Baseline Start here if the active model set fits
Two 2 TB Samsung 9100 Pros in RAID 0 + separate 2 TB OS SSD About 4 TB, no redundancy Adds RM2,749 under the prior package quote or RM2,849 at list price per worker Buy only for measured latency benefit or demonstrated cache-capacity need

One 2 TB drive can hold approximately 47 copies of the current 42.47 GB set before filesystem and free-space allowance, or roughly 37 copies after reserving 20%. Actual capacity planning must de-duplicate shared files and measure the complete model catalogue rather than multiplying one workflow blindly.

If the cache must exceed 2 TB, the second Samsung can be added later. RAID 0 is acceptable for a disposable cache, but its purchase should be justified primarily as a 4 TB capacity expansion unless the local same-machine experiment proves a repeatable latency win. Keep M2_2 empty initially so the option remains available and rear USB4 remains enabled.

Price/performance interpretation of the retailer list#

The 2026-09-22 IdealTech extraction captured these relevant listed prices:

Item Price Useful interpretation
Samsung 9100 Pro 2 TB RM2,849 list; prior package add-on RM2,749 Proven model-cache family; do not buy the second device solely for synthetic RAID speed
MSI Spatium M480 2 TB, Gen4 TLC RM1,499 Balanced 2 TB OS candidate; the OS path does not need Gen5
Kioxia Exceria Basic 2 TB, Gen4 QLC RM1,399 RM100 cheaper OS option; take the TLC M480 if its installed quote remains this close
MSI Spatium M480 Pro 4 TB, Gen4 TLC RM2,199 Can saturate 5GbE as NAS flash, but duplicates local NVMe and is not initially necessary
Seagate BarraCuda 2 TB ST2000DM008 RM599 Poor capacity value and desktop/SMR characteristics
Seagate BarraCuda 4 TB ST4000DM004 RM799 Cheapest temporary shared origin, but desktop SMR and not a robust NAS mirror choice
Seagate IronWolf 4 TB RM815 dealer / RM725 bundle promotion in M-Link's 8 September 2026 list Two-drive mirror costs RM1,630 / RM1,450 and yields 4 TB usable before formatting
Seagate IronWolf 8 TB RM1,615 dealer / RM1,435 bundle promotion Two-drive mirror costs RM3,230 / RM2,870 and yields 8 TB usable before formatting
Seagate IronWolf 12 TB RM2,530 dealer / RM2,250 bundle promotion Two-drive mirror costs RM5,060 / RM4,500 and yields 12 TB usable before formatting

At standalone list prices, one Samsung model drive plus the preferred M480 2 TB OS drive is RM4,348 per worker. Adding the second Samsung raises that to RM7,197, an extra RM2,849 per worker or RM5,698 across two newly configured workers. The previous promotional package quoted the second Samsung RM100 below list, so the final comparison must use the builder's written installed price.

The current CMR-NAS comparison is unusually clear:

Mirrored NAS capacity Bundle-promotion price Price per usable TB Decision
2x 4 TB IronWolf RM1,450 RM362.50/TB Balanced starting point while the repository plus forecast fits below about 3.2 TB
2x 8 TB IronWolf RM2,870 RM358.75/TB Capacity-triggered choice; only about 1% cheaper per usable TB, so do not pre-buy unused space
2x 12 TB IronWolf RM4,500 RM375.00/TB More capacity, but not the current price/performance sweet spot

These are planning prices from M-Link's 8 September 2026 dealer and promotional list, not a written end-user quote. Confirm stock, tax, delivery and warranty eligibility before purchase. The 8 TB mirror costs RM1,420 more than the 4 TB mirror while improving promotional cost per usable TB by only RM3.75, or about 1%. That tiny unit-price difference does not justify doubling capital expenditure for unused capacity. The currently measured two-workload union is only about 56.91 GiB. Start with the 4 TB mirror if the measured canonical repository plus a realistic twelve-month forecast is no more than about 3.2 TB, leaving 20% free. Select 2x 8 TB before purchase only if that capacity test fails; do not infer NAS capacity by adding the two worker-cache sizes because their contents should overlap heavily.

The retailer's two 4 TB BarraCuda drives would cost RM1,598, RM148 more than the promotional 4 TB IronWolf mirror while providing the less suitable SMR desktop duty profile. There is therefore no economic reason to build the production mirror from those BarraCuda drives at these prices.

That second-drive money is better retained until either:

Balanced NAS design#

Where the NFS origin should run#

The logical design needs one shared origin; it does not require buying a third computer on day one:

Placement Capital cost Operational trade-off Use
Existing reliable, always-on Linux/NAS host with native SATA bays HDDs and networking only Independent of either worker Best choice when already available
One two-disk mirror inside one GPU worker, exported to both workers Lowest bootstrap cost Rebooting or servicing that worker temporarily blocks cache misses on both workers Acceptable first phase; it is the one central origin, not a per-worker HDD cache
New dedicated NAS chassis/host Highest Best fault isolation and maintenance independence Buy when origin availability or capacity earns the added host cost

If there is no suitable existing NAS, co-locate the mirror in one worker initially rather than purchasing a third host merely for architectural neatness. Use native SATA bays and cooling, not an unproven consumer USB enclosure. Because jobs use local NVMe, loss of the origin-hosting worker still leaves already-cached jobs on the other worker serviceable. Move the same HDDs into an independent NAS later if maintenance coupling or cache-miss availability becomes economically important.

Initial production recommendation#

Use one Linux NFS origin—independent if a suitable host already exists, otherwise temporarily co-located in one worker—with:

A two-disk mirror is selected for availability and operational simplicity, not to promise 5GbE from one stream. It can distribute independent reads in some implementations, but do not budget performance from that assumption. If the NAS has only one 5GbE uplink, it also cannot deliver 5GbE to both workers simultaneously. Serialize initial fills and let cache-aware scheduling avoid duplicate emergency downloads.

Do not buy a NAS SSD tier initially. Local NVMe already performs that job. Add flash to the NAS only if telemetry shows that HDD-backed promotion is repeatedly on the critical path. The listed single 4 TB M480 Pro would reduce the ideal 42.47 GB fill from roughly 3.5 minutes on one HDD to roughly 77-85 seconds over 5GbE, but it costs RM2,199: RM749 more than the complete 4 TB IronWolf mirror, with neither the mirror's redundancy nor the 8 TB mirror's capacity. Saving two to three minutes on a background cache fill has value only when it happens frequently or blocks paid jobs.

Lean temporary variant#

If all model binaries are reproducible from the internet or object storage, one large CMR HDD plus an independent copy of manifests and private data is an acceptable bootstrap phase. Use one 4 TB IronWolf in that case, not the RM799 4 TB BarraCuda: the IronWolf is RM725 under the conditional promotion or RM815 at the ordinary dealer price, so the desktop SMR drive saves at most RM16 and may cost RM74 more. The second matching IronWolf costs little enough that the mirror remains the balanced production choice. Do not buy a staging HDD for each worker.

NFS and cache operation#

Export the canonical model repository read-only to the workers. Keep it at a path distinct from the active ComfyUI model directory:

NAS export:                    /srv/model-origin
Worker NFS mount:              /mnt/model-origin
Worker local cache:            /srv/model-cache
ComfyUI model paths/symlinks:  point only to verified local cache objects

Use NFSv4.2 over TCP with a hard mount. Modern Linux negotiates the largest read and write size supported by both sides; Red Hat documents a 1,048,576-byte maximum on current systems. nconnect=4 can be tested where both the client kernel and mount tooling support it, but it is not a substitute for measuring one large-file copy. See the Red Hat NFS mount documentation.

An indicative read-only worker mount is:

nas:/srv/model-origin /mnt/model-origin nfs4 ro,hard,_netdev,noatime,nconnect=4 0 0

Keep the mount out of the live serving path so a hard-mounted but unavailable NAS cannot stall an already cached job. A dedicated cache-fill process should:

  1. Resolve a workflow to a versioned manifest of model files and cryptographic hashes.
  2. Check which objects already exist on that worker's NVMe cache.
  3. Copy missing objects into a staging directory without exposing partial files to ComfyUI.
  4. Verify size and hash, then atomically rename the completed object into the cache.
  5. Publish or update local symlinks only after every required object is present.
  6. Admit the job, and update last-access metadata for cache-aware eviction.
  7. Evict whole verified objects or bundles when free space falls below 15-20%.

The scheduler should prefer the worker that already has the required model bundle. Prefetch likely next bundles while the GPU is busy. These two software policies save more customer-visible time than forcing the NAS itself to behave like a local Gen5 drive.

Use MTU 1500 initially. Jumbo frames have little economic value at two nodes and introduce an end-to-end configuration dependency. Enable them only after every NIC, switch port and NAS interface is validated.

Acceptance measurements and upgrade gates#

Record these measurements before buying more storage or networking:

  1. ethtool must show 5,000 Mbit/s full duplex on both workers and the NAS link at its intended rate.
  2. iperf3 to a RAM- or SSD-backed NAS endpoint should sustain at least about 4.5 Gbit/s per worker when tested alone. This isolates the network from HDD performance.
  3. Copy the real 42.47 GB model set from the NAS to an empty local cache and record elapsed time, NAS disk throughput, worker network throughput and hash-verification time.
  4. Repeat with both workers filling simultaneously. Record whether the bottleneck is the NAS disk, NAS uplink, switch, or worker link.
  5. Run both full ComfyUI workloads from local NVMe and retain process and physical-NVMe counters.
  6. Simulate NAS loss. Already-cached jobs must continue; only cache misses should wait or fail cleanly.
  7. Simulate deletion of the local cache and prove it can be reconstructed from the NAS.

Upgrade in this order:

  1. Improve cache affinity, prefetching and job admission.
  2. Add the second local 2 TB Samsung when capacity or a controlled A/B result justifies it.
  3. Add NAS flash only when HDD refill latency repeatedly blocks work.
  4. Add 10GbE to workers only after 5GbE is measurably saturated often enough to repay the adapters and switch ports.

Boundaries#