Owned GPU strategy and architecture decision#
Status: Reference Last verified: 2026-10-03 (competitive premises; fleet, telemetry and financial figures retain their stated evidence dates) Canonical for: why ComfyICU owns its GPU compute, how the owned fleet is used, what it must prove, and the evidence behind it. Founder decisions are recorded in
second-brain/DECISIONS.md(2026-09-26 and its 2026-09-27 updates); open execution questions insecond-brain/NOTEBOOK.md#8. Fleet: three single-GPU RTX 5090 nodes (PC1–PC3), bought 21–25 September 2026 for RM115,071 (procurement, ordered builds). Deployment status recorded in the September evidence: not serving production traffic; owned-node scheduling was not built. The October competitive review did not recheck live deployment state.
Position#
- Owning GPU compute is a permanent goal. It is not an investment with an exit, and not a bet to be reversed if renting gets cheaper. Resale value never enters a decision.
- Owning the compute is the foundation of the moat. It is why we're fast (always-warm workers and resident models on local NVMe) and why our base capacity has no per-second GPU rent. Residency and caching are architectural choices, not exclusive to ownership; competitor-owned fleets and committed leases can also use already-paid idle time. A market cost or latency lead must be measured (competitor evidence).
- The fleet's return is demand, not Modal savings. Fastest-in-category cold starts, and idle time given away as short, preemptible free, relaxed and win-back work. Displacing Modal spend is the floor, not the case (the demand lever).
- Capacity runs in three lanes: protected fast paid work and gap-fill work on owned nodes, and long, bursty or >32 GB work on rented capacity (capacity lanes).
- Renting is a bridge, for work the fleet can't do yet. It shrinks as the fleet grows. We rent our own compute out only if there is no application for it.
- Buy now, not gradually. Node prices are projected to rise before they fall, so any node put to work in the next two to three years is cheapest now. Capital limits the next purchase to one or two nodes (capacity expansion).
- Integration is the gate. None of this earns until the scheduler can place work on the nodes (required production integration).
How to reason about the fleet#
The founder's framing (DECISIONS 2026-09-26 and 2026-09-27), in his words where quoted:
- "Doing hard things is where the value is. Anyone could rent GPUs ... What's our advantage, we're fast, we're cheap? Why? We own the compute."
- We stay in business. "We won't shutdown the company." The company is not VC-backed and needs about US$1k (RM4k) a month to cover living costs. Reason from that, not from a startup's runway-and-exit logic.
- Nodes are long-lived productive assets. A 3090 from 2020 still runs AI models; a 5090 is RTX PRO 6000 compute with less VRAM and no ECC. The next fleet is added once the current one shows some ROI, and it "doesn't even have to break even because it's a productive asset."
- Idle capacity is something to put to work, not a sunk loss.
- Malaysia is an advantage to use fully. Owning our power (solar) is a long-term goal with no immediate plan.
Framings that were made and corrected, and should not come back:
- rent versus own, regret triggers, exit or shutdown scenarios, and resale value;
- gradual buying, cost-averaging, or invented buying processes, funds, bands or cadences;
- demand treated as fixed, and owned node time priced by electricity rather than by what paid work loses;
- site upgrades (three-phase supply, commercial tariff, solar) proposed as the next step.
Owned capacity as a demand lever#
Founder decision, second-brain/DECISIONS.md 2026-09-26.
The reframe#
The fleet is three nodes (PC1–PC3, RM115,071), and its main return is demand, not Modal savings. At current load, displacement alone is roughly break-even over a three-year life (economics): August Modal+Verda spend was about US$1.2k (dataplane, inputs about 28 days stale), down from about US$3.1k in March through Modal optimisation, and the two-versus-three replay puts each node at about 8–11% occupancy. The capabilities we are building to turn ownership into demand are:
- The fastest cold starts in the category. Always-warm workers, resident model families and local NVMe remove GPU acquisition and container restore. Production Krea cold runs had a 22.1 s p50 and Modal snapshot workers took a 15.7 s median (132.5 s maximum) just to become ready, against about 9 s for a complete cold Krea run on a fast-storage 5090 (evidence below). This is a positioning claim to prove locally, then sell.
- Unmetered idle time. Owned capacity is not billed per second, so idle time can be given away to pull users in: a cheaper relaxed tier, free credits to win back churned users, and free-tier usage. A competitor on metered rented capacity incurs provider charges for that work. This is not exclusive to us: owned fleets and committed leases can also exploit already-paid idle time, and Spot/negotiated rates are not hyperscaler list prices (GPU-supply research).
Capacity lanes#
Owned node time is priced by its opportunity cost to paid work, not by electricity. With one GPU per node, a job on an owned node makes arriving paid jobs wait or spill to Modal (per-second billing plus cold restores) and evicts the node's resident model. So:
| Lane | Runs on | Work | Rules |
|---|---|---|---|
| 1. Fast (protected) | Owned nodes | Short, latency-sensitive paid work on each node's resident model families | Always first; nothing else may degrade its latency or residency |
| 2. Gap-fill | Owned nodes | Free, relaxed-price and win-back work | Short preemptible jobs only (no longer than the 300 s default); prefer the node's resident family; never the last free node; never spills to Modal |
| 3. Rented | Modal, or a rented Verda/DataCrunch block | Long (e.g. 30-minute video), bursty, or >32 GB work, and peaks beyond the fleet | Priced at its cost; a rented block is cheaper than blocking a third of the owned fleet |
Lane 2 is also the safest first production traffic for the nodes, because it carries no latency promise.
Conditions#
- First-run reliability comes with win-back. 43% of first executions fail, mostly on graph configuration; credits that bring users back to the same failure are wasted.
- Quotas start small and grow as idle capacity is proven (raising a free allowance is welcome; cutting one breaks trust).
- Integration is the gate. None of this earns until the scheduler can use the nodes.
Open execution questions, signals and the review date are tracked in second-brain/NOTEBOOK.md #8.
Evidence language#
When this document distinguishes certainty, the following terms apply:
- Measured: produced by a benchmark, production query, or provider billing export, dated where it is used.
- Reported: supplied by the founder or a vendor but not independently reconciled here.
- Assumption: a planning input used to make a decision; it must be replaced by telemetry when available.
- Decision: the operating choice made from the current evidence and assumptions.
The purchase is known; full workload compatibility and the share of demand the fleet can absorb remain operating assumptions until production integration. Worker-lifecycle and residency telemetry now establish that cold state dominates, but only a matched production A/B can measure how much of its penalty the owned nodes remove.
How we reached this point#
Phase 1: rented GPU hosts and remote model storage#
ComfyICU first depended heavily on rented infrastructure, including Verda. The difficult part was not only GPU execution. Workers repeatedly paid latency and operational complexity for downloading large, changing model sets to remote machines. Availability varied, download performance was inconsistent, and the fast NVMe storage needed to hide those limitations was itself expensive.
Engineering effort went into improving download paths and working around provider behavior. Some of that work created reusable knowledge about the workload, but it did not remove the root condition: every ephemeral or poorly cached worker had to reconstruct a large local working set before useful inference could begin.
The June workload analysis found:
- 12,601 runs in the month;
- median execution time of approximately 26 seconds, p90 of 100 seconds, and p99 of 293 seconds;
- a median of approximately 21 GB of model files referenced per run;
- median cold download time of approximately 259 seconds;
- cold download time greater than execution time for about 79% of runs;
- a model corpus of roughly 1.47 TB across about 350 models; and
- a relatively concentrated hot set: the top 500 GB covered approximately 96.7% of observed accesses.
The full methodology and limitations are in the workload analysis.
The important conclusion was not that downloads needed another round of optimization. It was that repeatedly reconstructing the same hot model set on rented machines was the wrong steady-state architecture for predictable demand.
Phase 2: Modal and memory snapshots#
The workload moved to Modal to reduce operational burden and improve elasticity. This was initially successful: serverless execution removed much of the direct VM management, Modal's network-backed persistent storage was much faster than repeated origin downloads, and memory snapshots substantially improved framework and PyTorch startup. The models still traversed a provider-controlled network/distributed filesystem and container lifecycle rather than a direct local NVMe path.
The next experiment attempted to include loaded models in snapshots. It improved the time from a restored worker becoming ready to completing inference, but it did not improve total economics. Restoring the larger memory state took longer, failed intermittently, and consumed billable GPU lifecycle time before useful execution.
The clearest H100 snapshot incident recorded:
- approximately 1,261 seconds between application readiness and job completion;
- approximately 10,938 billable GPU seconds;
- twelve restoration failures of roughly six minutes each; and
- approximately US$13.73 of application cost for the incident.
See the H100 snapshot restore billing incident.
The broader August reconciliation reached the same conclusion. A model based on completed execution plus two seconds per run predicted 165.5 GPU-hours and US$490.88. The production Modal resource bill recorded 312.6 GPU-hours and US$925.52. Including failed runs and all recorded method-side GPU time raised the estimate only to 180.2 hours. Container lifecycle, cold restore work, and resident-loop tail explained most of the remaining gap. The detailed reconstruction is in GPU economics and the billing reconciliation experiment.
This means the miss was not merely an inaccurate per-inference estimate. The billable system boundary was wider than the useful-compute boundary. A cloud worker can be operationally convenient while still charging for provisioning, restore, failure, idle tail, CPU, memory, and storage around the inference itself.
Phase 3: delete the recurring problem#
The owned-node strategy changes the problem instead of continuing to tune it:
- download the hot model corpus once and retain it locally;
- pay for hardware and electricity rather than every container lifecycle transition;
- keep the worker processes alive and selected model families resident without paying a cloud-GPU idle tail;
- load non-resident models from direct-attached NVMe rather than a mounted provider Volume or a remote origin;
- control the worker software, cache, lifecycle, and observability directly; and
- rent only the portion of demand for which elasticity or large memory is actually valuable.
The prior cloud work was therefore not entirely wasted. It established workload shape, concurrency, model diversity, failure modes, and the real billing boundary. The stopping rule is that further cloud lifecycle optimization should not delay deployment of an architecture that removes that lifecycle from the predictable base load.
The selected architecture#
ComfyICU scheduler
|
classify by lane, compatibility, model,
residency, capacity, and SLA
|
+-----------------+-------------+---+-----------------+
| | | |
v v v v
PC1, 64 GB RAM PC2, 64 GB RAM PC3, 128 GB RAM Rented lane
RTX 5090 32 GB RTX 5090 32 GB RTX 5090 32 GB Modal / Verda block
local cache local cache local cache long, bursty, >32 GB
PCIe 5.0 x16 PCIe 5.0 x16 PCIe 5.0 x16
| | |
+-----------------+--------+--------+
|
durable metadata and artifacts
remain outside disposable model caches
Each owned GPU is placed in a separate complete machine. The cards are deliberately not concentrated in one multi-GPU host.
The separate-node design provides:
- one full CPU and host-memory domain per GPU;
- an independent PCIe x16 path and local NVMe model cache per GPU;
- one concurrent service slot per node;
- independent maintenance, restart, and failure domains;
- incremental growth and replacement in one-GPU cells; and
- the ability to keep different model families warm on each node.
Each local model cache is disposable. Missing models are staged from the shared authoritative model repository before a job is admitted; normal inference never reads model weights live over the NAS or provider Volume. Durable metadata and customer artifacts remain outside the worker caches.
This is a throughput and operability design, not a distributed-memory design. A job must fit on one card unless the application explicitly supports a tested multi-node execution path. No such path is assumed here.
The ordered node builds and acceptance targets are in the node decision record, summarised in the local infrastructure overview. The pre-order lane-topology and networking analysis is in the IdealTech decision, and site power is covered by the power and site decision.
Why 32 GB RTX 5090 cells rather than one RTX Pro 6000#
The first purchase, on 2026-09-23, compared two complete RTX 5090 machines (about RM70,000) with one RTX Pro 6000 Workstation Edition GPU (about RM85,000 before the rest of its host). The evidence below decided it, and it is why the fleet is built from 32 GB cells.
Production evidence: cold state dominates, but “cold” has three meanings#
“Cold” previously bundled together different costs. The production implementation and telemetry allow them to be separated:
- Worker-cold: the request is the first job handled by that worker/container.
- Model-residency miss: the required base-model signature does not match the prior resident signature, so the request cannot reuse the already-constructed model state.
- Origin-cold: model files are absent from the persistent model Volume and must be fetched from Hugging Face, CivitAI, R2, or another origin.
A read-only query of production runs created from September 1 through the fixed September 23, 15:17 UTC capture read worker session IDs and the executor's explicit residency_hit, model_signature, and model_download_bytes metrics. Drafts and insufficient-credit rows had no positive runtime and were excluded. Failed attempts were retained for operational accounting but excluded from customer-latency comparisons. It found:
| Production condition | Runs | Share | Interpretation |
|---|---|---|---|
| All positive-runtime attempts | 12,264 | 100% | Operational/cost cohort; includes 668 failed, timed-out, canceled, or crashed attempts |
| Completed positive-runtime runs | 11,596 | 94.6% of attempts | Customer-latency cohort |
| Completed runs with residency telemetry | 6,949 | 59.9% of completed | Denominator for successful-job residency results; older/uninstrumented runs are excluded |
| Completed model-residency miss | 5,465 | 78.6% of instrumented completed | Required base model was not already resident |
| Completed model-residency hit | 1,484 | 21.4% of instrumented completed | Warm reuse existed, mainly in repeated sequences |
| Completed run with positive origin-download bytes | 282 | 4.1% of instrumented completed | Most successful measured runs did not fetch model files from origin |
| All-attempt residency miss | 6,075 of 7,575 | 80.2% | Operational view including failed attempts |
| First job on its worker, all attempts | 10,779 | 87.9% | These first jobs accounted for 90.4% of all-attempt execution time |
| First job on its worker, completed only | 10,127 | 87.3% | These first jobs accounted for 89.9% of completed execution time |
The first-job result and the residency result measure different things. A first job restored from a model-warmed snapshot can be a residency hit, while a reused worker can miss after switching model families. Both measures independently show that the request path is usually not a simple warm inference. The low origin-cold rate shows that the remaining delay is not primarily slow internet download: it is usually container/worker lifecycle plus reading cached weights from the mounted Volume, deserializing them, applying any transforms, and moving/constructing the executable model state.
The phase timings reinforce that interpretation. For instrumented completed runs, the model-preparation phase—including cache checks, cleanup, links, and any actual origin fetch—had a p50 of 1.53 seconds and p90 of 3.67 seconds. The executor then opens mount-direct model files during ComfyUI execution, so those Volume reads appear inside run_time, not as model_download_ms. Looking only at “download time” would therefore miss the recurring network-backed model-read cost that local NVMe is intended to remove.
The completed-run latency split is large enough to matter to the product, not just to a microbenchmark:
| Production group or base-model signature | Miss runs | Hit runs | Miss p50 | Hit p50 | Observed p50 gap |
|---|---|---|---|---|---|
| Dedicated Krea 2 pool | 406 | 67 | 22.13 s | 5.07 s | 17.06 s |
| Dedicated WAN 2.2 pool | 148 | 168 | 67.43 s | 25.88 s | 41.54 s |
| Dedicated Z-Image pool | 132 | 53 | 20.59 s | 3.28 s | 17.31 s |
| MiniMax H3 INT8 signature, generic workers | 276 | 266 | 127.97 s | 86.60 s | 41.37 s |
| Flux 1 Dev FP8 signature, generic workers | 1,300 | 227 | 37.41 s | 14.27 s | 23.14 s |
| Flux 2 Klein 9B FP8 signature, generic workers | 575 | 184 | 24.31 s | 4.71 s | 19.60 s |
These gaps are observational, not controlled load-time measurements. Runs sharing a base-model signature can still differ in resolution, steps, LoRAs, graph shape, output count, and input media; one less common LTX signature even showed the opposite ordering because of that workload mix. The table therefore does not prove that every second of the gap is removable by storage. Its value is that several high-volume groups independently show the same direction and tens-of-seconds scale. A matched route-level A/B on the owned nodes is still required.
Warm compute is a floor, not the primary service metric#
The original GPU ranking used the same approximately 42.47 GB MiniMax H3 model set and five cold/warm pairs. Warm runs kept the constructed model objects resident and performed zero model-file reads. They isolate the GPU and in-memory execution floor reasonably well, but they do not represent the common ComfyICU request: only 21.4% of instrumented completed production runs were residency hits. The all-attempt rate was 19.8%.
The warm comparison remains useful as a boundary:
| Configuration | Mean warm run | Warm throughput | VRAM available to one job |
|---|---|---|---|
| 1x RTX Pro 6000 Workstation Edition | 48.858 s | 73.68 runs/hour | 96 GB |
| 1x RTX 5090 in a 9950X-class host | 63.344 s | 56.83 runs/hour | 32 GB |
| 2x RTX 5090, independent hosts | Parallel jobs | 113.66 aggregate runs/hour | 32 GB per job |
One Pro 6000 has about 30% more warm MiniMax throughput than one 5090. Two 5090s have about 54% more aggregate warm throughput than the single Pro. This is the lower-bound compute comparison; it is not the end-to-end user experience.
Cold MiniMax H3 performance#
The fresh-process MiniMax results are more relevant to production. Each cold run starts a new ComfyUI process without resident model objects. The newer schema-v4 controls also evict model-file pages and trace reads through the lower storage layers.
| GPU and host path | Cold mean | Warm mean | Cold penalty | Interpretation |
|---|---|---|---|---|
| Pro 6000, fast host A | 56.540 s | 48.858 s | 7.682 s | Fastest replicated Pro result; older protocol retained OS cache |
| Pro 6000, fast host B | 61.064 s | 48.217 s | 12.847 s | Independent fast-Pro replication; older protocol retained OS cache |
| 5090, single physical Gen5 MP700 | 70.018 s | 62.716 s | 7.302 s | Schema-v4 proved every cold model byte reached one physical NVMe |
| 5090, 2 TB Samsung 9100-backed host | 71.086 s | 62.266 s | 8.820 s | 575 W, Ryzen 9950X3D; provider container path |
| 5090, 4 TB Samsung 9100-backed host | 66.767 s | 61.035 s | 5.732 s | 600 W, Ryzen 9950X3D; provider container path, host marked unverified |
| 5090, physical Samsung 9100 control | 81.762 s | 69.986 s | 11.776 s | Exact drive reached physically, but 500 W GPU and slower Intel platform prevent a speed comparison |
The 5090 is slower on a single MiniMax job. Using the two high-power Samsung-backed results against the two fast Pro results, the observed cold gap ranges from roughly 9% to 26%, with a midpoint near 17%. That is the relevant meaning of “close”: it is not performance parity, but the per-job cold deficit is modest relative to the capital and concurrency difference.
At those observed means:
- one Pro 6000 supplies approximately 59-64 cold MiniMax runs per hour;
- one 5090 supplies approximately 51-54 cold runs per hour; and
- two independent 5090s supply approximately 101-108 cold runs per hour in aggregate.
The pair therefore provides roughly 60-80% more aggregate cold MiniMax throughput than one Pro across the observed range. This assumes independent queued work and does not reduce the latency of one job.
Krea 2 illustrates why the cold path matters#
Krea is shorter, so model loading is a larger share of the user-visible runtime:
| GPU and storage path | Cold mean | Warm mean | Cold penalty | Model-reader rate |
|---|---|---|---|---|
| Pro 6000 rented host | 13.108 s | 3.023 s | 10.085 s | 3.324 GiB/s |
| 5090, single Gen5 SN8100 | 9.463 s | 4.263 s | 5.200 s | 11.003 GiB/s |
| 5090, dual SN850X physical RAID path | 8.664 s | 4.265 s | 4.399 s | 11.488 GiB/s |
| 5090, dual T705 RAID path | 10.081 s | 4.272 s | 5.809 s | 10.390 GiB/s |
The Pro is roughly 41% faster in warm Krea throughput, yet the fast-storage 5090 hosts were approximately 23-34% faster end to end on a cold run. In these cross-host results, the combined storage and model-construction path outweighed the Pro's warm-compute advantage.
The individual result files are Pro 6000 MiniMax host A, Pro 6000 MiniMax host B, physical single-Gen5 5090 control, physical Samsung 9100 control, Krea Pro 6000, and the dual-workload 5090 matrix.
These are cross-host observations, not a matched GPU-or-storage A/B. CPU, power, filesystem, kernel, host-to-device path, and provider cache differ. They support the direction and bound plausible performance; the purchased machines' direct-filesystem acceptance run is the deciding measurement.
Expected local-NVMe UX advantage#
Owned hardware improves cold UX through more than raw SSD bandwidth:
- The worker already exists; there is no GPU acquisition or container provisioning before it can accept a job.
- The complete current 1.47 TB model catalogue fits on a 2 TB local drive, with roughly 0.53 TB of raw decimal headroom before formatting and reserve, so normal jobs need neither an origin download nor live reads from a mounted provider Volume.
- A direct filesystem removes provider-selected overlay, loop-backing-file, network-filesystem, and contention layers from the model path.
- Because idle GPU time is no longer billed by the second, worker processes and the most valuable model pools can remain resident between bursts.
- Jobs that are still cold load from a predictable local device rather than a provider lifecycle whose queue, restore, and storage behavior are outside ComfyICU's control.
The storage evidence is directionally consistent with that design. A committed 27.1 GiB file read back from a Modal Volume at 2,226 MiB/s, or about 2.17 GiB/s. The MP700 and Samsung 9100 Pro physical controls measured 4.536 and 5.699 GiB/s, but those are the two slowest observations among six one-hop PCIe 5.0 x4 single-storage hosts. That group averaged 6.804 GiB/s, had a median of approximately 6.454 GiB/s, and ranged from 4.536 to 10.098 GiB/s. The three attributable single-path 9100 Pro observations ranged from 5.699 to 7.585 GiB/s and averaged 6.642 GiB/s; a fourth 9.581 GiB/s 9100 host exposed another NVMe, so the active device could not be assigned. Relative to the Modal readback, the relevant groups center near 3 times the throughput rather than 2.1-2.6 times.
The two slower controls remain valuable because their counters prove that the physical SSD served the cold bytes; they should not be presented as representative speed estimates. The broader sample is also cross-host and uses a different reader from the Modal test, so even its ratio is not a promised application speedup. Its 10.098 GiB/s ceiling came from an SN8100 host with 2 MiB read-ahead and is not an expected 9100 Pro baseline. The purchased nodes' same-machine acceptance measurements remain authoritative.
Application timings provide the actual service hypothesis. Across those six Gen5 single-storage hosts, the MiniMax cold penalty averaged 8.56 seconds and ranged from 5.73 to 11.78 seconds, versus a 41.4-second completed-production residency-miss gap for the same broad model family. Fast-storage Krea benchmark penalties were 4.4-5.8 seconds, versus 17.1 seconds in the dedicated production pool. If a matched owned-node deployment moves production toward those local bands, the plausible improvement is tens of seconds for MiniMax and roughly ten seconds for Krea. This remains a hypothesis: production graphs differ, and residency gaps include more than storage.
Historical Modal snapshot evidence shows the size of the removable platform layer. Krea 2 RTX Pro snapshot workers took a median 15.7 seconds and a maximum 132.5 seconds merely to move from enqueue to container-ready, before request execution. The failed H100 snapshot path took a median 109.3 seconds and a maximum 622.6 seconds. In comparison, the complete fast-storage 5090 Krea cold runs above took 8.7-10.1 seconds after the local process accepted the prompt.
This is not a direct Modal-Volume-versus-local-NVMe A/B: the timings begin at different boundaries and the current Modal route no longer uses all of those snapshot pools. It nevertheless explains why a locally available worker can improve customer latency much more than the 5090-versus-Pro warm compute table suggests. The production comparison must measure request enqueue to completed artifact on both routes.
Local NVMe will not make every miss equal to a warm hit. Safetensors still must be opened and deserialized; quantization or other CPU transforms may run; weights still move through host memory and PCIe; and ComfyUI still constructs its model objects. The architectural claim is narrower: persistent workers should create more residency hits, and direct NVMe should reduce and stabilize the unavoidable miss path. Both effects must be measured separately.
Gen5 branding alone is not the claim. The broader storage experiment found that after a sound direct-attached path is established, RAID and higher synthetic bandwidth do not consistently save more than one or two seconds. One well-configured local Gen5 NVMe is the starting point; a second drive is justified by capacity or a same-machine application improvement, not by fio numbers.
What the professional-card premium buys#
The RTX Pro 6000 has 96 GB of GDDR7 ECC memory and professional workstation features. That is valuable when a single production workload cannot fit in 32 GB, when ECC is contractually or operationally required, or when professional partitioning and support features have proven value.
Two 32 GB cards do not form one 64 GB or 96 GB memory pool. The founder reports that all current ComfyICU workloads are 5090-eligible, including MiniMax H3, the most demanding current workflow, so the professional card's memory premium solves no demonstrated current production constraint.
Those benefits do not create more service slots. With no demonstrated current 32 GB constraint, they did not justify an RM85,000 bare-GPU purchase: the Pro was faster per job, but its observed MiniMax cold advantage was approximately 9-26%, while two 5090s delivered substantially more aggregate throughput.
Therefore:
- Decision: buy throughput and concurrency as 32 GB RTX 5090 service cells.
- Boundary: >32 GB work runs in the rented lane.
- Future option: a 48 GB, 72 GB, or 96 GB cell follows the larger-memory ladder.
Workload and capacity interpretation#
The operating assumption is that 100% of the current supported workload is technically RTX 5090-eligible. This is founder-reported and supported by the successful MiniMax H3 benchmark, which exercised the most demanding current workflow; it is not yet an exhaustive production compatibility result. The first production month must validate it per workflow.
Eligibility is different from capacity capture. A compatible job may still spill to Modal because every owned node is busy, its model is not cached or resident, a node is draining, or a deployment is being canaried. The production evidence above also makes cold queue-to-result latency—not only warm compute—the primary service metric.
Production evidence queried on 2026-09-23 showed:
- 14,525 created September-to-date records, of which 12,264 recorded positive execution time totaling approximately 174.8 GPU-hours;
- 11,596 completed records accounting for approximately 164.8 of those hours;
- a last-30-day peak observed concurrency of 12;
- executed-attempt arrivals during active hours averaging approximately 22.9, with p95 of 58 and a maximum of 123; and
- a diverse workload rather than a single dominant model family.
The dedicated executor pools represented 46.2% of positive-runtime attempts: Krea 2 at 19.8%, WAN 2.2 at 17.1%, and Z-Image at 9.2%. The remaining 53.8% used the generic queue and spanned many model signatures. This concentration is enough to make model-affine placement useful, but the generic majority and observed peak concurrency still require general overflow.
The 174.8 execution hours equal about 32% of one node's elapsed time over the reporting window, or about 11% across three nodes, before adjusting for GPU performance or scheduling losses. Applying the observed 9-26% MiniMax cold slowdown to every workload would raise that mechanical range to 190-220 local hours, or 35-41% of one node and 12-14% of the fleet. This is a stress illustration, not a forecast, because:
- arrivals are bursty rather than uniform;
- different workflows have different GPU performance ratios;
- model transitions and cold loads consume time;
- future workflows may introduce requirements not represented by the current benchmark set; and
- maintenance and failures reduce usable capacity.
The two-versus-three replay of July and August demand puts three nodes at about 7.6–7.7% occupancy, serving 88–96% of GPU-hours locally, with the rest spilling during concurrency peaks. At today's demand the fleet is justified by concurrency, model residency and failure isolation, and above all by the demand it is meant to create; not by average utilisation. The peak concurrency of 12 is why the rented lane remains part of the architecture.
Economics#
The return case is demand (the reframe). Modal displacement is the floor: what the fleet earns if demand does not grow at all. The figures below are dated and labelled; replace them with the dataplane and the first electricity bill once the nodes carry traffic.
Observed cloud spend#
The latest provider ledger available on 2026-09-23 contained:
| Month | Verda | Modal | Combined GPU-related provider spend |
|---|---|---|---|
| 2026-06 | US$941.58 | US$788.26 | US$1,729.84 |
| 2026-07 | US$410.10 | US$958.63 | US$1,368.73 |
| 2026-08 through August 29 | US$53.32 | US$1,146.71 | US$1,200.03 |
The ledger had last been refreshed on August 29 and was approximately 25 days old when these figures were captured. August is therefore incomplete. These figures establish the recent cost range, not a final invoice or current-month forecast.
The August Modal subtotal was distributed across both GPUs and their supporting resources:
| Modal resource, through August 29 | Cost |
|---|---|
| RTX Pro 6000 GPU | US$660.52 |
| H100 GPU | US$168.48 |
| L40S GPU | US$111.03 |
| Persistent volume | US$104.31 |
| Memory | US$65.19 |
| CPU | US$35.07 |
| Other small GPU/resource lines and rounding | Approximately US$2.11 |
| Modal subtotal | US$1,146.71 |
This composition is why displacement is modelled as a share rather than the whole bill disappearing. Modal overflow, deployment safety, local downtime, and some supporting costs continue after local routing is active.
Displacement floor (ours, derived)#
Inputs:
- the incomplete August baseline, US$1,200.03 at RM4.20 per US dollar: about RM5,040 a month;
- the replay's August share of GPU-hours served locally by three nodes: 91.6% without waiting and 96.0% allowing up to 60 s of queueing (applying the GPU-hour share to the whole bill overstates displacement, since volume, CPU and memory lines don't all move);
- electricity for three nodes: RM1,020–1,320 a month (electricity cost). That range assumes 50–70% utilisation; at today's modelled 8% occupancy the bill is lower, but idle draw is not yet metered.
avoided cloud spend = 91.6–96.0% × RM5,040 = RM4,617–4,838 a month
net of electricity = RM4,617 − 1,320 to 4,838 − 1,020 = RM3,297–3,818 a month
simple payback = RM115,071 / RM3,297–3,818 = about 30–35 months
That is the "roughly break-even over a three-year life" behind the demand lever: on August's demand alone, displacement recovers the fleet's cost in about three years. The nodes are long-lived productive assets, so this scores how well the fleet is bought and run; it does not decide whether to own.
Electricity is a larger share of owned cost than a round per-kWh rate suggests. The household uses about 740 kWh a month, and any node load pushes the premises over TNB's 1,500 kWh cliff, after which every kWh costs about RM0.636 all-in, including the household's own use. Each node beyond three adds RM304–402 a month. Per-node branch metering and the first bill with the nodes running replace these estimates with observed kWh per completed job.
Where the advantage comes from#
Owning the compute is the foundation: persistent local workers remove our ephemeral GPU acquisition and container restore path, while ownership removes per-second GPU rent from the base load. These mechanisms support the fast and gap-fill lanes; they do not establish a unique capability or measured lead against TensorArt's self-operated fleet, SeaArt's optimized cloud deployments, or Haima's GPUaaS platform behind RunningHub (dated competitor findings). Paid-work opportunity cost still governs gap-fill admission. The system around the hardware is how that advantage turns into demand:
- ComfyICU already has paid, varied GPU demand;
- historical runs reveal which models and workflows deserve residency;
- the hot model set can be stored once on fast local media;
- the scheduler can route by lane, compatibility, model affinity, queue depth, and cost;
- local lifecycle and cache behavior are directly observable and controllable;
- the base load is not exposed to provider restore, tail, and storage billing boundaries; and
- rented capacity absorbs peaks and memory tiers the fleet can't serve yet.
Rapid progress in open-weight models, including capable models in the roughly 27B class, makes commodity 32 GB cells more useful over time: more useful software fits on broadly available hardware, and the software layer can be replaced without replacing the service.
Alternatives considered#
| Alternative | Decision and reason | Revisit when |
|---|---|---|
| Remain all-Modal | Rejected for predictable base load; lifecycle, restore, tail, and mounted-storage exposure remain even after individual optimizations | Not a standing option: renting is a bridge |
| Return to conventional rented VMs | Rejected as the default; it restores the Verda-era availability, storage-price, transfer, and cache-reconstruction problems | A specific GPU or a long job is needed tactically (lane 3) |
| Buy one RTX Pro 6000 | Deferred; it offers 96 GB and faster single-job compute but only one service slot and failure domain at a higher bare-GPU price | The larger-memory ladder names a workload |
| Build one dual-5090 workstation | Rejected at this fleet size; density does not outweigh lane, cooling, power, host-contention, and failure-domain concentration, and VRAM is not transparently pooled | Space, circuits, or fleet-management overhead becomes the constraint |
| Buy nodes gradually as demand is proven | Rejected 2026-09-26/27; node prices are projected to rise first, so waiting costs more | The price mechanism in the buy-now analysis changes |
| Gate ownership on rental prices or resale value | Rejected 2026-09-26; owning is permanent | Not revisited |
| Rent out idle owned capacity | Last resort, only if there is no application for it | Owned capacity has no lane-1 or lane-2 use |
Operating boundaries and failure model#
The owned nodes are service capacity, not the durable system of record.
- Any owned node may be drained or lost without data loss.
- The rented lane remains available when the owned nodes are busy or unhealthy.
- A local site or network outage reduces capacity; it should not corrupt durable state.
- Availability and durability are treated separately. The fleet does not require datacenter-style local high availability.
Each 5090 is rated for very high board power, so one protected electrical circuit per worker, correct supplied 12V-2x6 cabling, verified earthing, surge protection, thermal monitoring, and branch power telemetry are part of acceptance. Large UPS purchases remain evidence-gated rather than assumed necessary. The site stays on a single-phase domestic supply (founder, 2026-09-27); confirm the main breaker and supply rating before adding nodes 4 and 5.
Required production integration#
The hardware produces no return until the scheduler can use it. The hardware and benchmark documentation are ahead of the production runtime: the active orchestrator is centered on Modal, with older VM-provider paths dormant or commented out. Production integration is therefore the highest-priority work after physical acceptance. The first question (NOTEBOOK #8) is whether the dormant DataCrunch VM worker path (session pre-create, frontend registration and heartbeat, the getJob claim) can serve an always-on owned node, or whether a separate lane is cleaner.
The minimum owned-worker control plane must provide:
- authenticated registration, health heartbeats, and capability reporting for GPU, VRAM, software image, cached models, and free space;
- compatibility rules that keep incompatible or >32 GB jobs off local nodes;
- model-affinity placement that avoids unnecessary cache and residency churn;
- safe drain, retry, and spill-to-Modal behavior;
- per-run phase timings and explicit spill reasons such as
local_busy,model_missing,vram_ineligible, andworker_unhealthy; - node-level energy and cost attribution against avoided cloud spend; and
- the three capacity lanes: lane-1 paid work always claims ahead of lane-2 gap-fill work; lane-2 jobs are capped at the default timeout, never take the last free node and never spill to Modal; lane-3 long or >32 GB work routes to rented capacity.
The first deployment should favor correctness and observability over sophisticated bin-packing. The suggested first traffic is lane 2 (no latency promise), before paid work moves across. Static assignment of the most valuable model pools plus safe cloud spill is sufficient to prove the architecture. Dynamic residency can follow after real traces show that it will reduce cost or latency.
Acceptance and review plan#
Before paid traffic#
For each node:
- verify the exact component inventory, supplied GPU cabling, PCIe 5.0 x16 link under load, and CPU-connected model-drive path;
- place the model cache on a direct local filesystem rather than a network mount or buffered container loop image;
- reproduce the application-level cold and warm MiniMax and Krea benchmarks with physical-NVMe counters;
- stress CPU and GPU concurrently while logging power, temperature, clock, and errors;
- test registration, drain, forced failure, retry, Modal spill, and rejection of a synthetic incompatible or oversized job; and
- preserve benchmark and acceptance artifacts for later regression comparison.
First 30 production days#
Record at least:
- compatibility by workflow, treating any exception as a contradiction of the current 100% assumption;
- local placement and Modal spill rates by paid runs, GPU time, cloud cost, and explicit spill reason;
- occupancy and queue delay per node;
- worker-first, origin-cold, residency-miss, residency-hit, and model-transition rates as separate states;
- completed jobs, revenue, and p50/p90/p99 end-to-end latency by workflow, model pool, and route, including a matched cold local-versus-Modal comparison;
- lane-2 usage: free, relaxed and win-back runs served, and the users they reactivate;
- local failures, retries, and customer-visible failures;
- kWh, cooling impact, energy cost, and weekly founder operations time; and
- actual Modal invoice reduction after normalizing for demand changes.
Success criteria#
The founder's review signals (NOTEBOOK #8, for example at 60 days of live traffic) are the share of GPU-hours served locally, the cold-latency delta on Krea and MiniMax, node occupancy, reactivated users who pay within 30 days, any drop in fast-mode revenue per payer, and founder hours per week on the nodes.
The direction is working if, after normalization for demand:
- demand grows: gap-fill and win-back work reactivates users who go on to pay, without cutting fast-mode revenue per payer;
- the cold-start claim holds: local cold queue-to-result latency materially beats Modal for the same workflows, so it can be said publicly;
- customer-visible reliability is no worse than the prior cloud-only path;
- the rented lane absorbs local peaks and faults without pathological retry behavior;
- local nodes require no more than one or two hours of routine founder attention per week;
- direct-attached NVMe makes cold model loading predictable, while model-affine placement converts some burst traffic to warm reuse; and
- production validation finds no current profitable workload segment excluded by the 32 GB limit.
Modal displacement is reported alongside, as the floor, not as the test.
If these conditions fail, first determine whether the cause is routing, software compatibility, demand shape, memory capacity, or site operation. Do not treat every failure as evidence that another GPU is required.
Capacity expansion policy#
Current position: owning is permanent; bulk buy now#
Founder decisions, second-brain/DECISIONS.md 2026-09-26 ("Owning compute is a permanent goal") and its
2026-09-27 updates.
- Owning GPU compute is a permanent goal. Avoid renting where possible; it is a bridge, not the alternative. Cheaper rentals never pause buying; they only make lane 3 cheaper in the meantime.
- Resale value never enters a decision.
- Bulk buy now rather than cost-average: "we're just starting, so we better to bulk buy now ... because it will be the cheapest now."
- Quantity is set by capital: "I might only have capital for 1 or 2 more nodes that's it," for a fleet of four or five.
- Add the next fleet once the current one shows some ROI; it "doesn't even have to break even because it's a productive asset."
Why now: should we buy more now? A first-principles price analysis reasons from how each part's price is set and applies that to our node's actual parts:
- Next 12 months: a node like ours is projected at +3% to +20% (+RM1.0k to +7.1k). Memory (RAM plus the GPU's GDDR7, about 70% of the node) keeps rising through 2027. Nvidia has every incentive to starve the 5090 in favour of the RTX PRO 6000, which is cut from the same die. The next consumer generation has reportedly slipped to 2028.
- 24 months: about −5% at the centre (range −21% to +9%). Memory stays expensive through 2028, and RTX 60 launches into it. Waiting two years saves little and costs two years of the node's work.
- 36 months: clearly down, likely −20% or more, if the 2028 memory supply lands on time. The first real discount is about three years out.
- The risk to this call: an AI spending correction in 2027. Its early warning is H100 rental prices falling while capex still rises.
So any node we'll use before roughly late 2028 is cheapest bought now. Which GPU and build the next nodes use is worked in GPU theory, market price and workload performance; how to buy them is in procurement.
Site limits (founder, 2026-09-27): single-phase supply, domestic tariff, no immediate plan for solar. Five nodes at full load is about 6 kW including cooling on a single-phase supply shared with the household: it fits, with little margin (supply headroom).
A larger-memory node#
Prefer the smallest memory tier that clears a measured production constraint. Before buying a 96 GB RTX Pro 6000, collect:
- the number and revenue of otherwise-valid jobs rejected or spilled for VRAM;
- their actual peak allocated and reserved VRAM;
- their cloud cost and latency;
- whether a 48 GB or 72 GB card is sufficient; and
- whether quantization or offload preserves the required output and latency.
The decision ladder is therefore 32 GB commodity cells first, 48/72 GB when measured demand requires them, and 96 GB only when the larger pool has a named workload and economic case. A glut of used data-centre cards (about 35% odds, mid-2027 to mid-2028) is the likely buying window (NOTEBOOK #8). Meanwhile >32 GB work runs in the rented lane.
Stop-doing list#
- Do not benchmark for its own sake. Each benchmark round serves a named decision: the next node's GPU and build, or an owned node's acceptance and integration.
- Do not optimize Modal snapshots for large loaded models unless a bounded experiment targets a specific remaining cloud workload.
- Do not buy network, UPS, memory, or platform upgrades without production telemetry showing the bottleneck.
- Do not confuse synthetic SSD bandwidth with end-to-end model-load performance.
- Do not stop caring about cold storage performance merely because warm GPU benchmarks look good; 78.6% of explicitly instrumented completed September runs missed model residency.
- Do not keep refining local storage after direct-filesystem cold latency is stable and acceptable. The evidence does not show a repeatable RAID advantage large enough to justify optimization by synthetic bandwidth alone.
- Do not reopen rent versus own, resale value, or gradual buying (how to reason).
The next unit of engineering effort belongs in production routing, observability, and cost attribution.
Consequences#
Expected benefits#
- A target of the fastest cold starts in the category, from always-warm workers and resident models on local NVMe; competitor-relative performance is not yet measured.
- Idle capacity that costs only electricity, available as a demand lever.
- More aggregate throughput, concurrency, and failure isolation per ringgit than a single RTX Pro 6000.
- Direct control of worker lifecycle, model residency, cache, software images, and telemetry.
- Lower, more predictable base-load cost, and incremental growth in replaceable one-GPU cells.
Accepted costs and risks#
- RM115,071 of capital is committed before the demand it is meant to create is proven.
- Each local job is limited to one 32 GB VRAM pool.
- The site now owns power, heat, hardware failure, security, and maintenance responsibilities, on a single-phase domestic supply with little margin at five nodes.
- Consumer cards lack the Pro card's ECC and professional feature set.
- PC1 and PC2 have 64 GB of host memory: sufficient for the tested MiniMax process but not generous (the physical MP700 control reached about 45.7 GiB of process RSS), so RSS, swap, page-cache churn, and model-transition latency must be observed. PC3's 128 GB makes it the natural WAN node.
- Residential or office power and internet are weaker failure domains than a datacenter.
- The fleet adds scheduler and observability work that Modal previously abstracted.
These are accepted because each node is independently replaceable, durable state stays elsewhere, and the rented lane remains available while the fleet grows.
What changed#
- 2026-09-23: the first decision. Two RTX 5090 nodes (RM71,747) chosen over one RTX Pro 6000, with Modal displacement as the return and a third node gated on 30 days of production evidence.
- 2026-09-25: PC3 (RM43,324, 128 GB RAM) ordered ahead of that gate.
- 2026-09-26: owned capacity is a demand lever, with three lanes; owning compute is permanent and resale is excluded; the expansion gates and a briefly proposed rental-price "regret trigger" were withdrawn.
- 2026-09-27: bulk buy now, quantity set by capital (1–2 more nodes); site limits; electricity recomputed from the verified TNB bill formula.
- 2026-09-29: this document rewritten as one current position. The earlier wording is in git history;
the dated decisions are in
second-brain/DECISIONS.md. - 2026-10-03: competitor GPU-supply research qualified the blanket claims about renters, cold starts and exclusive idle-capacity economics. The ownership goal, buying decisions and three lanes are unchanged; production state and financial inputs were not refreshed by this research.
Evidence index#
- Chinese AI platform GPU supply: dated, sourced TensorArt, SeaArt and RunningHub procurement findings, operating-versus-ownership distinctions and competitive limits.
- Local GPU infrastructure overview: current topology, storage evidence, component choices, and acceptance gates.
- Node decision record: the ordered builds, RAM and storage choices, and the two-versus-three capacity replay.
- Procurement: what each node cost and how to buy the next ones.
- Buy-now analysis: part-by-part price projections behind bulk buying now.
- GPU theory, market price and workload performance: the evidence model for choosing the next node's GPU.
- Owned GPU infrastructure proposal: original owned-base-load and cloud-overflow thesis.
- GPU economics: cloud cost history and the August Modal billing correction.
- Modal billing reconciliation experiment: reconstruction of billed versus method-side GPU time.
- H100 snapshot restore billing incident: evidence that model-warmed restore time and failures were billable.
- Pool snapshot startup health: measured enqueue-to-ready latency for the Krea 2 RTX Pro and H100 snapshot routes.
- Workload analysis: run-time distribution, model working set, concurrency, and download burden.
- September workload snapshot: frozen September 1-17 production extract and analysis; the September 23 residency figures in this document were refreshed directly from the production database using the same read-only usage/worker source.
- Production evidence SQL: fixed-window, read-only queries for the status audit, worker reuse, residency, origin-download, pool, and per-signature latency figures in this document.
- Modal model-Volume throughput: controlled origin-transfer and committed-Volume readback measurements, including the 2,226 MiB/s readback result.
- Resident executor implementation: definitions and emission of model signature, prior resident signature, residency hit, preparation time, and cold-origin byte metrics.
- Volume-path model: documents that mount-direct Modal model files are read during execution rather than copied during preparation.
- GPU comparison data: measured RTX 5090 and RTX Pro 6000 results used in the first purchase.
- Single-path multivariable analysis: six-host PCIe 5.0 x4 reader distribution, host-level explanations, and separation of storage-path from GPU-transfer effects.
- Single-NVMe versus RAID dual-workload experiment: current MiniMax and Krea cold/warm matrix, physical-read controls, and the decision not to assume RAID is required.
- IdealTech pre-order decision (historical): bill of materials, Malaysian procurement context, and price gates for the first nodes.
- Power and site decision: PSU, circuits, surge protection, metering, UPS criteria, and the electricity cost model.
- NVIDIA RTX 5090 specifications: 32 GB VRAM and published board-power characteristics.
- NVIDIA RTX Pro 6000 specifications: 96 GB ECC VRAM and professional-card characteristics.
- Modal pricing: provider list prices; actual invoiced cost depends on the complete container lifecycle and supporting resources.