Dual-worker model storage and NFS/NAS decision#
Status: Current deployment decision; local RAID 0 remains optional pending a same-machine test Decision date: 2026-09-23 Scope: Two local RTX 5090 ComfyUI workers connected initially through their onboard 5GbE ports Primary decision: Keep active models on a local NVMe cache in each worker; use one shared HDD-backed NFS/NAS as the model origin; do not add a local HDD to each worker and do not serve the normal customer workload directly from NFS
Decision in one page#
Use this hierarchy:
Tier 0: GPU VRAM and Linux RAM/page cache
│
Tier 1: local Samsung 9100 Pro cache in each worker
one 2 TB drive initially on the new worker; retain an existing two-drive cache
active model files use a direct filesystem on CPU-connected Gen5 x4 storage
▲
│ staged and verified before a job is admitted
│
Tier 2: one shared NFS/NAS model origin
large CMR HDD capacity; preferably a two-drive mirror
▲
│ backup/synchronization, not the live model path
│
Tier 3: independent off-machine copy or reproducible upstream sources
manifests, private models, configuration and irreplaceable data
The initial layout for a newly configured worker should be:
CPU Gen5 x4 M.2 ── Samsung 9100 Pro 2 TB ── /srv/model-cache
chipset Gen4 x4 M.2 ── separate 2 TB OS SSD ── Linux, containers, logs, outputs
second CPU Gen5 x4 M.2 ── empty initially; Tomahawk only, reserved for conditional expansion
onboard 5GbE ── NFS model origin and worker-to-worker/site traffic
Both the promotional MSI PRO X870E-S EVO and the upgraded Tomahawk provide onboard 5GbE, one CPU-connected Gen5 x4 M.2 path, chipset Gen4 M.2 storage and a CPU-connected Gen5 x16 GPU slot. The RM450 Tomahawk upgrade is therefore not required for the initial one-model-drive layout. It buys the independent second CPU Gen5 x4 M.2 path needed for a later two-drive experiment or 4 TB striped cache.
Keep the Tomahawk while that expansion is genuinely undecided: RM450 is inexpensive option value compared with replacing the motherboard later. If the local test is completed before ordering, RAID is rejected and a 2 TB cache with headroom is demonstrated to be sufficient, retain the included PRO X870E-S EVO and save RM450. MSI's PRO X870E-S EVO specification confirms its 5GbE controller, single CPU Gen5 x4 M.2 socket, chipset M.2 sockets and Gen5 x16 GPU slot.
Do not install an HDD in each worker as an intermediate cache. It duplicates capacity, adds another copy and eviction policy, increases heat and failure monitoring, and is still much slower than the local NVMe. The shared NAS is already the correct next tier beneath local NVMe.
Do not point the production ComfyUI model directories directly at NFS. A scheduler or cache agent should stage every required model bundle to local NVMe, verify it, atomically publish it, and only then admit the job. A NAS outage must prevent a cache miss from being filled; it must not stop a worker from serving models that are already local.
The ordered nodes (updated 2026-09-26; the record is the node decision doc) use no RAID and put the OS on a separate Samsung drive in the second CPU slot, not on a chipset Gen4 SSD as drawn above:
| Node | M2_1, CPU Gen5 x4 | M2_2, CPU Gen5 x4 |
|---|---|---|
| PC1 | 9100 Pro 2 TB, model cache | 9100 Pro 2 TB: OS, Docker, outputs, later LLM weights |
| PC2 | 9100 Pro 2 TB, OS and model-cache partitions until the OS drive arrives | 9100 Pro 1 TB OS drive, bought later |
| PC3 | 9100 Pro 2 TB, model cache | 9100 Pro 1 TB, OS |
Each node therefore has one 2 TB model cache. Advertise available models and free cache capacity to the scheduler so it can route each job to the better-prepared worker. The NAS, NFS and staging design below is unchanged by the order, apart from serving three workers instead of two.
Why direct NFS is not the normal serving path#
The current MiniMax model and input set is 42,471,968,783 bytes. The lower-bound transfer times are:
| Source path | Planning throughput | Ideal time for 42.47 GB | Appropriate role |
|---|---|---|---|
| Local Samsung 9100 Pro, measured physical reader | 5.699 GiB/s | about 7 seconds of raw reading | Active model cache |
| 5GbE raw line rate | 625 MB/s | 68 seconds | Mathematical ceiling only |
| Healthy 5GbE NFS planning rate | 500-550 MB/s | 77-85 seconds | Cache fill, not request-time model loading |
Seagate BarraCuda 2 TB ST2000DM008 maximum |
220 MB/s | about 3.2 minutes | Cheap staging/origin media only |
Seagate BarraCuda 4 TB ST4000DM004 maximum |
190 MB/s | about 3.7 minutes | Cheap staging/origin media only |
Selected Seagate IronWolf 4 TB ST4000VN006 maximum |
202 MB/s | about 3.5 minutes | NAS capacity/origin media |
The 5GbE figures are decimal: 5 Gbit/s is 625 MB/s before Ethernet, TCP and NFS overhead. They are not 5 GB/s. A single HDD therefore will not saturate 5GbE. This is acceptable when copying is asynchronous; it is unacceptable if every cold customer request waits for that copy.
Seagate specifies 220 MB/s for the 2 TB BarraCuda and 190 MB/s for the 4 TB model. It also identifies both listed models as SMR drives with a desktop-class 55 TB/year workload limit and 2,400 power-on hours per year. They are inexpensive but are not the balanced foundation for a mirrored 24/7 NAS. Seagate's IronWolf line uses CMR, is rated for 8,760 power-on hours and a 180 TB/year workload, and is the more appropriate class for the shared origin. Relevant manufacturer references are the 2 TB BarraCuda manual, BarraCuda recording-technology sheet, and IronWolf specification.
What the latest RAID evidence means#
The earlier purchase record selected two Samsung 9100 Pro drives in RAID 0. That recommendation is superseded by the completed single-NVMe versus RAID 0 experiment.
The fixed dual-workload matrix did not show the required repeated application-level RAID advantage. A later schema-v4 physical control showed that one direct-attached Gen5 NVMe could serve every cold byte from the physical device and complete MiniMax with a 7.30-second cold penalty. A separate physical 4 TB Samsung 9100 Pro control also completed every run without RAID. RAID greatly improved some synthetic results but did not reliably improve the complete ComfyUI workloads.
The current 2 TB-only purchasing constraint changes the capacity choice, not the performance conclusion:
| Worker option | Usable model-cache capacity | Incremental economics | Decision |
|---|---|---|---|
| One 2 TB Samsung 9100 Pro + separate 2 TB OS SSD | About 2 TB before formatting/headroom | Baseline | Start here if the active model set fits |
| Two 2 TB Samsung 9100 Pros in RAID 0 + separate 2 TB OS SSD | About 4 TB, no redundancy | Adds RM2,749 under the prior package quote or RM2,849 at list price per worker | Buy only for measured latency benefit or demonstrated cache-capacity need |
One 2 TB drive can hold approximately 47 copies of the current 42.47 GB set before filesystem and free-space allowance, or roughly 37 copies after reserving 20%. Actual capacity planning must de-duplicate shared files and measure the complete model catalogue rather than multiplying one workflow blindly.
If the cache must exceed 2 TB, the second Samsung can be added later. RAID 0 is acceptable for a disposable cache, but its purchase should be justified primarily as a 4 TB capacity expansion unless the local same-machine experiment proves a repeatable latency win. Keep M2_2 empty initially so the option remains available and rear USB4 remains enabled.
Price/performance interpretation of the retailer list#
The 2026-09-22 IdealTech extraction captured these relevant listed prices:
| Item | Price | Useful interpretation |
|---|---|---|
| Samsung 9100 Pro 2 TB | RM2,849 list; prior package add-on RM2,749 | Proven model-cache family; do not buy the second device solely for synthetic RAID speed |
| MSI Spatium M480 2 TB, Gen4 TLC | RM1,499 | Balanced 2 TB OS candidate; the OS path does not need Gen5 |
| Kioxia Exceria Basic 2 TB, Gen4 QLC | RM1,399 | RM100 cheaper OS option; take the TLC M480 if its installed quote remains this close |
| MSI Spatium M480 Pro 4 TB, Gen4 TLC | RM2,199 | Can saturate 5GbE as NAS flash, but duplicates local NVMe and is not initially necessary |
Seagate BarraCuda 2 TB ST2000DM008 |
RM599 | Poor capacity value and desktop/SMR characteristics |
Seagate BarraCuda 4 TB ST4000DM004 |
RM799 | Cheapest temporary shared origin, but desktop SMR and not a robust NAS mirror choice |
| Seagate IronWolf 4 TB | RM815 dealer / RM725 bundle promotion in M-Link's 8 September 2026 list | Two-drive mirror costs RM1,630 / RM1,450 and yields 4 TB usable before formatting |
| Seagate IronWolf 8 TB | RM1,615 dealer / RM1,435 bundle promotion | Two-drive mirror costs RM3,230 / RM2,870 and yields 8 TB usable before formatting |
| Seagate IronWolf 12 TB | RM2,530 dealer / RM2,250 bundle promotion | Two-drive mirror costs RM5,060 / RM4,500 and yields 12 TB usable before formatting |
At standalone list prices, one Samsung model drive plus the preferred M480 2 TB OS drive is RM4,348 per worker. Adding the second Samsung raises that to RM7,197, an extra RM2,849 per worker or RM5,698 across two newly configured workers. The previous promotional package quoted the second Samsung RM100 below list, so the final comparison must use the builder's written installed price.
The current CMR-NAS comparison is unusually clear:
| Mirrored NAS capacity | Bundle-promotion price | Price per usable TB | Decision |
|---|---|---|---|
| 2x 4 TB IronWolf | RM1,450 | RM362.50/TB | Balanced starting point while the repository plus forecast fits below about 3.2 TB |
| 2x 8 TB IronWolf | RM2,870 | RM358.75/TB | Capacity-triggered choice; only about 1% cheaper per usable TB, so do not pre-buy unused space |
| 2x 12 TB IronWolf | RM4,500 | RM375.00/TB | More capacity, but not the current price/performance sweet spot |
These are planning prices from M-Link's 8 September 2026 dealer and promotional list, not a written end-user quote. Confirm stock, tax, delivery and warranty eligibility before purchase. The 8 TB mirror costs RM1,420 more than the 4 TB mirror while improving promotional cost per usable TB by only RM3.75, or about 1%. That tiny unit-price difference does not justify doubling capital expenditure for unused capacity. The currently measured two-workload union is only about 56.91 GiB. Start with the 4 TB mirror if the measured canonical repository plus a realistic twelve-month forecast is no more than about 3.2 TB, leaving 20% free. Select 2x 8 TB before purchase only if that capacity test fails; do not infer NAS capacity by adding the two worker-cache sizes because their contents should overlap heavily.
The retailer's two 4 TB BarraCuda drives would cost RM1,598, RM148 more than the promotional 4 TB IronWolf mirror while providing the less suitable SMR desktop duty profile. There is therefore no economic reason to build the production mirror from those BarraCuda drives at these prices.
That second-drive money is better retained until either:
- the local cache catalogue, with 15-20% free-space allowance, no longer fits in 2 TB; or
- a direct-filesystem, same-worker A/B test shows RAID 0 saving more than approximately two seconds on both production workloads.
Balanced NAS design#
Where the NFS origin should run#
The logical design needs one shared origin; it does not require buying a third computer on day one:
| Placement | Capital cost | Operational trade-off | Use |
|---|---|---|---|
| Existing reliable, always-on Linux/NAS host with native SATA bays | HDDs and networking only | Independent of either worker | Best choice when already available |
| One two-disk mirror inside one GPU worker, exported to both workers | Lowest bootstrap cost | Rebooting or servicing that worker temporarily blocks cache misses on both workers | Acceptable first phase; it is the one central origin, not a per-worker HDD cache |
| New dedicated NAS chassis/host | Highest | Best fault isolation and maintenance independence | Buy when origin availability or capacity earns the added host cost |
If there is no suitable existing NAS, co-locate the mirror in one worker initially rather than purchasing a third host merely for architectural neatness. Use native SATA bays and cooling, not an unproven consumer USB enclosure. Because jobs use local NVMe, loss of the origin-hosting worker still leaves already-cached jobs on the other worker serviceable. Move the same HDDs into an independent NAS later if maintenance coupling or cache-miss availability becomes economically important.
Initial production recommendation#
Use one Linux NFS origin—independent if a suitable host already exists, otherwise temporarily co-located in one worker—with:
- 2x 4 TB IronWolf CMR drives as the balanced starting point while the measured catalogue and realistic
twelve-month forecast fit below about 3.2 TB; otherwise choose 2x 8 TB before purchase; use
mdadmRAID 1 plus ext4, or an equivalent mirror in an already-familiar storage stack; - a small independent boot SSD when using a dedicated host, or the worker's existing OS SSD when co-located;
- at least a 5GbE uplink initially;
- a 10GbE uplink only when both 5GbE workers must refill simultaneously at close to line rate; and
- an independent backup or reproducible upstream source for private/irreplaceable content and metadata.
A two-disk mirror is selected for availability and operational simplicity, not to promise 5GbE from one stream. It can distribute independent reads in some implementations, but do not budget performance from that assumption. If the NAS has only one 5GbE uplink, it also cannot deliver 5GbE to both workers simultaneously. Serialize initial fills and let cache-aware scheduling avoid duplicate emergency downloads.
Do not buy a NAS SSD tier initially. Local NVMe already performs that job. Add flash to the NAS only if telemetry shows that HDD-backed promotion is repeatedly on the critical path. The listed single 4 TB M480 Pro would reduce the ideal 42.47 GB fill from roughly 3.5 minutes on one HDD to roughly 77-85 seconds over 5GbE, but it costs RM2,199: RM749 more than the complete 4 TB IronWolf mirror, with neither the mirror's redundancy nor the 8 TB mirror's capacity. Saving two to three minutes on a background cache fill has value only when it happens frequently or blocks paid jobs.
Lean temporary variant#
If all model binaries are reproducible from the internet or object storage, one large CMR HDD plus an independent copy of manifests and private data is an acceptable bootstrap phase. Use one 4 TB IronWolf in that case, not the RM799 4 TB BarraCuda: the IronWolf is RM725 under the conditional promotion or RM815 at the ordinary dealer price, so the desktop SMR drive saves at most RM16 and may cost RM74 more. The second matching IronWolf costs little enough that the mirror remains the balanced production choice. Do not buy a staging HDD for each worker.
NFS and cache operation#
Export the canonical model repository read-only to the workers. Keep it at a path distinct from the active ComfyUI model directory:
NAS export: /srv/model-origin
Worker NFS mount: /mnt/model-origin
Worker local cache: /srv/model-cache
ComfyUI model paths/symlinks: point only to verified local cache objects
Use NFSv4.2 over TCP with a hard mount. Modern Linux negotiates the largest read and write size supported by
both sides; Red Hat documents a 1,048,576-byte maximum on current systems. nconnect=4 can be tested where
both the client kernel and mount tooling support it, but it is not a substitute for measuring one large-file
copy. See the Red Hat NFS mount documentation.
An indicative read-only worker mount is:
nas:/srv/model-origin /mnt/model-origin nfs4 ro,hard,_netdev,noatime,nconnect=4 0 0
Keep the mount out of the live serving path so a hard-mounted but unavailable NAS cannot stall an already cached job. A dedicated cache-fill process should:
- Resolve a workflow to a versioned manifest of model files and cryptographic hashes.
- Check which objects already exist on that worker's NVMe cache.
- Copy missing objects into a staging directory without exposing partial files to ComfyUI.
- Verify size and hash, then atomically rename the completed object into the cache.
- Publish or update local symlinks only after every required object is present.
- Admit the job, and update last-access metadata for cache-aware eviction.
- Evict whole verified objects or bundles when free space falls below 15-20%.
The scheduler should prefer the worker that already has the required model bundle. Prefetch likely next bundles while the GPU is busy. These two software policies save more customer-visible time than forcing the NAS itself to behave like a local Gen5 drive.
Use MTU 1500 initially. Jumbo frames have little economic value at two nodes and introduce an end-to-end configuration dependency. Enable them only after every NIC, switch port and NAS interface is validated.
Acceptance measurements and upgrade gates#
Record these measurements before buying more storage or networking:
ethtoolmust show 5,000 Mbit/s full duplex on both workers and the NAS link at its intended rate.iperf3to a RAM- or SSD-backed NAS endpoint should sustain at least about 4.5 Gbit/s per worker when tested alone. This isolates the network from HDD performance.- Copy the real 42.47 GB model set from the NAS to an empty local cache and record elapsed time, NAS disk throughput, worker network throughput and hash-verification time.
- Repeat with both workers filling simultaneously. Record whether the bottleneck is the NAS disk, NAS uplink, switch, or worker link.
- Run both full ComfyUI workloads from local NVMe and retain process and physical-NVMe counters.
- Simulate NAS loss. Already-cached jobs must continue; only cache misses should wait or fail cleanly.
- Simulate deletion of the local cache and prove it can be reconstructed from the NAS.
Upgrade in this order:
- Improve cache affinity, prefetching and job admission.
- Add the second local 2 TB Samsung when capacity or a controlled A/B result justifies it.
- Add NAS flash only when HDD refill latency repeatedly blocks work.
- Add 10GbE to workers only after 5GbE is measurably saturated often enough to repay the adapters and switch ports.
Boundaries#
- NFS, RAID and snapshots are not backups. Private models, manifests, configuration and customer data need an independent copy.
- The local 2 TB model NVMe is a cache. No irreplaceable data belongs on it, whether single or RAID 0.
- The 42.47 GB timing examples describe the current MiniMax set. Catalogue size and model-switch frequency determine the real economics.
- The M-Link IronWolf figures distinguish ordinary dealer prices from conditional bundle-promotion prices; neither is a confirmed end-user offer. Obtain a written all-in quote before locking capacity. Even at the non-promotion dealer price, the 4 TB IronWolf mirror is only RM32 more than two IdealTech 4 TB BarraCuda drives, which is a strong trade for CMR NAS drives and their higher duty rating.
- The exact 2 TB OS SSD remains a quote decision. The M480 TLC is the value default only while its installed price remains within roughly RM100 of the Kioxia QLC option.