Owned GPU strategy and architecture decision#

Status: Reference Last verified: 2026-10-03 (competitive premises; fleet, telemetry and financial figures retain their stated evidence dates) Canonical for: why ComfyICU owns its GPU compute, how the owned fleet is used, what it must prove, and the evidence behind it. Founder decisions are recorded in second-brain/DECISIONS.md (2026-09-26 and its 2026-09-27 updates); open execution questions in second-brain/NOTEBOOK.md #8. Fleet: three single-GPU RTX 5090 nodes (PC1–PC3), bought 21–25 September 2026 for RM115,071 (procurement, ordered builds). Deployment status recorded in the September evidence: not serving production traffic; owned-node scheduling was not built. The October competitive review did not recheck live deployment state.

Position#

How to reason about the fleet#

The founder's framing (DECISIONS 2026-09-26 and 2026-09-27), in his words where quoted:

Framings that were made and corrected, and should not come back:

Owned capacity as a demand lever#

Founder decision, second-brain/DECISIONS.md 2026-09-26.

The reframe#

The fleet is three nodes (PC1–PC3, RM115,071), and its main return is demand, not Modal savings. At current load, displacement alone is roughly break-even over a three-year life (economics): August Modal+Verda spend was about US$1.2k (dataplane, inputs about 28 days stale), down from about US$3.1k in March through Modal optimisation, and the two-versus-three replay puts each node at about 8–11% occupancy. The capabilities we are building to turn ownership into demand are:

Capacity lanes#

Owned node time is priced by its opportunity cost to paid work, not by electricity. With one GPU per node, a job on an owned node makes arriving paid jobs wait or spill to Modal (per-second billing plus cold restores) and evicts the node's resident model. So:

Lane Runs on Work Rules
1. Fast (protected) Owned nodes Short, latency-sensitive paid work on each node's resident model families Always first; nothing else may degrade its latency or residency
2. Gap-fill Owned nodes Free, relaxed-price and win-back work Short preemptible jobs only (no longer than the 300 s default); prefer the node's resident family; never the last free node; never spills to Modal
3. Rented Modal, or a rented Verda/DataCrunch block Long (e.g. 30-minute video), bursty, or >32 GB work, and peaks beyond the fleet Priced at its cost; a rented block is cheaper than blocking a third of the owned fleet

Lane 2 is also the safest first production traffic for the nodes, because it carries no latency promise.

Conditions#

Open execution questions, signals and the review date are tracked in second-brain/NOTEBOOK.md #8.

Evidence language#

When this document distinguishes certainty, the following terms apply:

The purchase is known; full workload compatibility and the share of demand the fleet can absorb remain operating assumptions until production integration. Worker-lifecycle and residency telemetry now establish that cold state dominates, but only a matched production A/B can measure how much of its penalty the owned nodes remove.

How we reached this point#

Phase 1: rented GPU hosts and remote model storage#

ComfyICU first depended heavily on rented infrastructure, including Verda. The difficult part was not only GPU execution. Workers repeatedly paid latency and operational complexity for downloading large, changing model sets to remote machines. Availability varied, download performance was inconsistent, and the fast NVMe storage needed to hide those limitations was itself expensive.

Engineering effort went into improving download paths and working around provider behavior. Some of that work created reusable knowledge about the workload, but it did not remove the root condition: every ephemeral or poorly cached worker had to reconstruct a large local working set before useful inference could begin.

The June workload analysis found:

The full methodology and limitations are in the workload analysis.

The important conclusion was not that downloads needed another round of optimization. It was that repeatedly reconstructing the same hot model set on rented machines was the wrong steady-state architecture for predictable demand.

Phase 2: Modal and memory snapshots#

The workload moved to Modal to reduce operational burden and improve elasticity. This was initially successful: serverless execution removed much of the direct VM management, Modal's network-backed persistent storage was much faster than repeated origin downloads, and memory snapshots substantially improved framework and PyTorch startup. The models still traversed a provider-controlled network/distributed filesystem and container lifecycle rather than a direct local NVMe path.

The next experiment attempted to include loaded models in snapshots. It improved the time from a restored worker becoming ready to completing inference, but it did not improve total economics. Restoring the larger memory state took longer, failed intermittently, and consumed billable GPU lifecycle time before useful execution.

The clearest H100 snapshot incident recorded:

See the H100 snapshot restore billing incident.

The broader August reconciliation reached the same conclusion. A model based on completed execution plus two seconds per run predicted 165.5 GPU-hours and US$490.88. The production Modal resource bill recorded 312.6 GPU-hours and US$925.52. Including failed runs and all recorded method-side GPU time raised the estimate only to 180.2 hours. Container lifecycle, cold restore work, and resident-loop tail explained most of the remaining gap. The detailed reconstruction is in GPU economics and the billing reconciliation experiment.

This means the miss was not merely an inaccurate per-inference estimate. The billable system boundary was wider than the useful-compute boundary. A cloud worker can be operationally convenient while still charging for provisioning, restore, failure, idle tail, CPU, memory, and storage around the inference itself.

Phase 3: delete the recurring problem#

The owned-node strategy changes the problem instead of continuing to tune it:

The prior cloud work was therefore not entirely wasted. It established workload shape, concurrency, model diversity, failure modes, and the real billing boundary. The stopping rule is that further cloud lifecycle optimization should not delay deployment of an architecture that removes that lifecycle from the predictable base load.

The selected architecture#

                                ComfyICU scheduler
                                        |
                  classify by lane, compatibility, model,
                        residency, capacity, and SLA
                                        |
        +-----------------+-------------+---+-----------------+
        |                 |                 |                 |
        v                 v                 v                 v
   PC1, 64 GB RAM    PC2, 64 GB RAM    PC3, 128 GB RAM    Rented lane
   RTX 5090 32 GB    RTX 5090 32 GB    RTX 5090 32 GB     Modal / Verda block
   local cache       local cache       local cache        long, bursty, >32 GB
   PCIe 5.0 x16      PCIe 5.0 x16      PCIe 5.0 x16
        |                 |                 |
        +-----------------+--------+--------+
                                   |
                    durable metadata and artifacts
                 remain outside disposable model caches

Each owned GPU is placed in a separate complete machine. The cards are deliberately not concentrated in one multi-GPU host.

The separate-node design provides:

Each local model cache is disposable. Missing models are staged from the shared authoritative model repository before a job is admitted; normal inference never reads model weights live over the NAS or provider Volume. Durable metadata and customer artifacts remain outside the worker caches.

This is a throughput and operability design, not a distributed-memory design. A job must fit on one card unless the application explicitly supports a tested multi-node execution path. No such path is assumed here.

The ordered node builds and acceptance targets are in the node decision record, summarised in the local infrastructure overview. The pre-order lane-topology and networking analysis is in the IdealTech decision, and site power is covered by the power and site decision.

Why 32 GB RTX 5090 cells rather than one RTX Pro 6000#

The first purchase, on 2026-09-23, compared two complete RTX 5090 machines (about RM70,000) with one RTX Pro 6000 Workstation Edition GPU (about RM85,000 before the rest of its host). The evidence below decided it, and it is why the fleet is built from 32 GB cells.

Production evidence: cold state dominates, but “cold” has three meanings#

“Cold” previously bundled together different costs. The production implementation and telemetry allow them to be separated:

A read-only query of production runs created from September 1 through the fixed September 23, 15:17 UTC capture read worker session IDs and the executor's explicit residency_hit, model_signature, and model_download_bytes metrics. Drafts and insufficient-credit rows had no positive runtime and were excluded. Failed attempts were retained for operational accounting but excluded from customer-latency comparisons. It found:

Production condition Runs Share Interpretation
All positive-runtime attempts 12,264 100% Operational/cost cohort; includes 668 failed, timed-out, canceled, or crashed attempts
Completed positive-runtime runs 11,596 94.6% of attempts Customer-latency cohort
Completed runs with residency telemetry 6,949 59.9% of completed Denominator for successful-job residency results; older/uninstrumented runs are excluded
Completed model-residency miss 5,465 78.6% of instrumented completed Required base model was not already resident
Completed model-residency hit 1,484 21.4% of instrumented completed Warm reuse existed, mainly in repeated sequences
Completed run with positive origin-download bytes 282 4.1% of instrumented completed Most successful measured runs did not fetch model files from origin
All-attempt residency miss 6,075 of 7,575 80.2% Operational view including failed attempts
First job on its worker, all attempts 10,779 87.9% These first jobs accounted for 90.4% of all-attempt execution time
First job on its worker, completed only 10,127 87.3% These first jobs accounted for 89.9% of completed execution time

The first-job result and the residency result measure different things. A first job restored from a model-warmed snapshot can be a residency hit, while a reused worker can miss after switching model families. Both measures independently show that the request path is usually not a simple warm inference. The low origin-cold rate shows that the remaining delay is not primarily slow internet download: it is usually container/worker lifecycle plus reading cached weights from the mounted Volume, deserializing them, applying any transforms, and moving/constructing the executable model state.

The phase timings reinforce that interpretation. For instrumented completed runs, the model-preparation phase—including cache checks, cleanup, links, and any actual origin fetch—had a p50 of 1.53 seconds and p90 of 3.67 seconds. The executor then opens mount-direct model files during ComfyUI execution, so those Volume reads appear inside run_time, not as model_download_ms. Looking only at “download time” would therefore miss the recurring network-backed model-read cost that local NVMe is intended to remove.

The completed-run latency split is large enough to matter to the product, not just to a microbenchmark:

Production group or base-model signature Miss runs Hit runs Miss p50 Hit p50 Observed p50 gap
Dedicated Krea 2 pool 406 67 22.13 s 5.07 s 17.06 s
Dedicated WAN 2.2 pool 148 168 67.43 s 25.88 s 41.54 s
Dedicated Z-Image pool 132 53 20.59 s 3.28 s 17.31 s
MiniMax H3 INT8 signature, generic workers 276 266 127.97 s 86.60 s 41.37 s
Flux 1 Dev FP8 signature, generic workers 1,300 227 37.41 s 14.27 s 23.14 s
Flux 2 Klein 9B FP8 signature, generic workers 575 184 24.31 s 4.71 s 19.60 s

These gaps are observational, not controlled load-time measurements. Runs sharing a base-model signature can still differ in resolution, steps, LoRAs, graph shape, output count, and input media; one less common LTX signature even showed the opposite ordering because of that workload mix. The table therefore does not prove that every second of the gap is removable by storage. Its value is that several high-volume groups independently show the same direction and tens-of-seconds scale. A matched route-level A/B on the owned nodes is still required.

Warm compute is a floor, not the primary service metric#

The original GPU ranking used the same approximately 42.47 GB MiniMax H3 model set and five cold/warm pairs. Warm runs kept the constructed model objects resident and performed zero model-file reads. They isolate the GPU and in-memory execution floor reasonably well, but they do not represent the common ComfyICU request: only 21.4% of instrumented completed production runs were residency hits. The all-attempt rate was 19.8%.

The warm comparison remains useful as a boundary:

Configuration Mean warm run Warm throughput VRAM available to one job
1x RTX Pro 6000 Workstation Edition 48.858 s 73.68 runs/hour 96 GB
1x RTX 5090 in a 9950X-class host 63.344 s 56.83 runs/hour 32 GB
2x RTX 5090, independent hosts Parallel jobs 113.66 aggregate runs/hour 32 GB per job

One Pro 6000 has about 30% more warm MiniMax throughput than one 5090. Two 5090s have about 54% more aggregate warm throughput than the single Pro. This is the lower-bound compute comparison; it is not the end-to-end user experience.

Cold MiniMax H3 performance#

The fresh-process MiniMax results are more relevant to production. Each cold run starts a new ComfyUI process without resident model objects. The newer schema-v4 controls also evict model-file pages and trace reads through the lower storage layers.

GPU and host path Cold mean Warm mean Cold penalty Interpretation
Pro 6000, fast host A 56.540 s 48.858 s 7.682 s Fastest replicated Pro result; older protocol retained OS cache
Pro 6000, fast host B 61.064 s 48.217 s 12.847 s Independent fast-Pro replication; older protocol retained OS cache
5090, single physical Gen5 MP700 70.018 s 62.716 s 7.302 s Schema-v4 proved every cold model byte reached one physical NVMe
5090, 2 TB Samsung 9100-backed host 71.086 s 62.266 s 8.820 s 575 W, Ryzen 9950X3D; provider container path
5090, 4 TB Samsung 9100-backed host 66.767 s 61.035 s 5.732 s 600 W, Ryzen 9950X3D; provider container path, host marked unverified
5090, physical Samsung 9100 control 81.762 s 69.986 s 11.776 s Exact drive reached physically, but 500 W GPU and slower Intel platform prevent a speed comparison

The 5090 is slower on a single MiniMax job. Using the two high-power Samsung-backed results against the two fast Pro results, the observed cold gap ranges from roughly 9% to 26%, with a midpoint near 17%. That is the relevant meaning of “close”: it is not performance parity, but the per-job cold deficit is modest relative to the capital and concurrency difference.

At those observed means:

The pair therefore provides roughly 60-80% more aggregate cold MiniMax throughput than one Pro across the observed range. This assumes independent queued work and does not reduce the latency of one job.

Krea 2 illustrates why the cold path matters#

Krea is shorter, so model loading is a larger share of the user-visible runtime:

GPU and storage path Cold mean Warm mean Cold penalty Model-reader rate
Pro 6000 rented host 13.108 s 3.023 s 10.085 s 3.324 GiB/s
5090, single Gen5 SN8100 9.463 s 4.263 s 5.200 s 11.003 GiB/s
5090, dual SN850X physical RAID path 8.664 s 4.265 s 4.399 s 11.488 GiB/s
5090, dual T705 RAID path 10.081 s 4.272 s 5.809 s 10.390 GiB/s

The Pro is roughly 41% faster in warm Krea throughput, yet the fast-storage 5090 hosts were approximately 23-34% faster end to end on a cold run. In these cross-host results, the combined storage and model-construction path outweighed the Pro's warm-compute advantage.

The individual result files are Pro 6000 MiniMax host A, Pro 6000 MiniMax host B, physical single-Gen5 5090 control, physical Samsung 9100 control, Krea Pro 6000, and the dual-workload 5090 matrix.

These are cross-host observations, not a matched GPU-or-storage A/B. CPU, power, filesystem, kernel, host-to-device path, and provider cache differ. They support the direction and bound plausible performance; the purchased machines' direct-filesystem acceptance run is the deciding measurement.

Expected local-NVMe UX advantage#

Owned hardware improves cold UX through more than raw SSD bandwidth:

  1. The worker already exists; there is no GPU acquisition or container provisioning before it can accept a job.
  2. The complete current 1.47 TB model catalogue fits on a 2 TB local drive, with roughly 0.53 TB of raw decimal headroom before formatting and reserve, so normal jobs need neither an origin download nor live reads from a mounted provider Volume.
  3. A direct filesystem removes provider-selected overlay, loop-backing-file, network-filesystem, and contention layers from the model path.
  4. Because idle GPU time is no longer billed by the second, worker processes and the most valuable model pools can remain resident between bursts.
  5. Jobs that are still cold load from a predictable local device rather than a provider lifecycle whose queue, restore, and storage behavior are outside ComfyICU's control.

The storage evidence is directionally consistent with that design. A committed 27.1 GiB file read back from a Modal Volume at 2,226 MiB/s, or about 2.17 GiB/s. The MP700 and Samsung 9100 Pro physical controls measured 4.536 and 5.699 GiB/s, but those are the two slowest observations among six one-hop PCIe 5.0 x4 single-storage hosts. That group averaged 6.804 GiB/s, had a median of approximately 6.454 GiB/s, and ranged from 4.536 to 10.098 GiB/s. The three attributable single-path 9100 Pro observations ranged from 5.699 to 7.585 GiB/s and averaged 6.642 GiB/s; a fourth 9.581 GiB/s 9100 host exposed another NVMe, so the active device could not be assigned. Relative to the Modal readback, the relevant groups center near 3 times the throughput rather than 2.1-2.6 times.

The two slower controls remain valuable because their counters prove that the physical SSD served the cold bytes; they should not be presented as representative speed estimates. The broader sample is also cross-host and uses a different reader from the Modal test, so even its ratio is not a promised application speedup. Its 10.098 GiB/s ceiling came from an SN8100 host with 2 MiB read-ahead and is not an expected 9100 Pro baseline. The purchased nodes' same-machine acceptance measurements remain authoritative.

Application timings provide the actual service hypothesis. Across those six Gen5 single-storage hosts, the MiniMax cold penalty averaged 8.56 seconds and ranged from 5.73 to 11.78 seconds, versus a 41.4-second completed-production residency-miss gap for the same broad model family. Fast-storage Krea benchmark penalties were 4.4-5.8 seconds, versus 17.1 seconds in the dedicated production pool. If a matched owned-node deployment moves production toward those local bands, the plausible improvement is tens of seconds for MiniMax and roughly ten seconds for Krea. This remains a hypothesis: production graphs differ, and residency gaps include more than storage.

Historical Modal snapshot evidence shows the size of the removable platform layer. Krea 2 RTX Pro snapshot workers took a median 15.7 seconds and a maximum 132.5 seconds merely to move from enqueue to container-ready, before request execution. The failed H100 snapshot path took a median 109.3 seconds and a maximum 622.6 seconds. In comparison, the complete fast-storage 5090 Krea cold runs above took 8.7-10.1 seconds after the local process accepted the prompt.

This is not a direct Modal-Volume-versus-local-NVMe A/B: the timings begin at different boundaries and the current Modal route no longer uses all of those snapshot pools. It nevertheless explains why a locally available worker can improve customer latency much more than the 5090-versus-Pro warm compute table suggests. The production comparison must measure request enqueue to completed artifact on both routes.

Local NVMe will not make every miss equal to a warm hit. Safetensors still must be opened and deserialized; quantization or other CPU transforms may run; weights still move through host memory and PCIe; and ComfyUI still constructs its model objects. The architectural claim is narrower: persistent workers should create more residency hits, and direct NVMe should reduce and stabilize the unavoidable miss path. Both effects must be measured separately.

Gen5 branding alone is not the claim. The broader storage experiment found that after a sound direct-attached path is established, RAID and higher synthetic bandwidth do not consistently save more than one or two seconds. One well-configured local Gen5 NVMe is the starting point; a second drive is justified by capacity or a same-machine application improvement, not by fio numbers.

What the professional-card premium buys#

The RTX Pro 6000 has 96 GB of GDDR7 ECC memory and professional workstation features. That is valuable when a single production workload cannot fit in 32 GB, when ECC is contractually or operationally required, or when professional partitioning and support features have proven value.

Two 32 GB cards do not form one 64 GB or 96 GB memory pool. The founder reports that all current ComfyICU workloads are 5090-eligible, including MiniMax H3, the most demanding current workflow, so the professional card's memory premium solves no demonstrated current production constraint.

Those benefits do not create more service slots. With no demonstrated current 32 GB constraint, they did not justify an RM85,000 bare-GPU purchase: the Pro was faster per job, but its observed MiniMax cold advantage was approximately 9-26%, while two 5090s delivered substantially more aggregate throughput.

Therefore:

Workload and capacity interpretation#

The operating assumption is that 100% of the current supported workload is technically RTX 5090-eligible. This is founder-reported and supported by the successful MiniMax H3 benchmark, which exercised the most demanding current workflow; it is not yet an exhaustive production compatibility result. The first production month must validate it per workflow.

Eligibility is different from capacity capture. A compatible job may still spill to Modal because every owned node is busy, its model is not cached or resident, a node is draining, or a deployment is being canaried. The production evidence above also makes cold queue-to-result latency—not only warm compute—the primary service metric.

Production evidence queried on 2026-09-23 showed:

The dedicated executor pools represented 46.2% of positive-runtime attempts: Krea 2 at 19.8%, WAN 2.2 at 17.1%, and Z-Image at 9.2%. The remaining 53.8% used the generic queue and spanned many model signatures. This concentration is enough to make model-affine placement useful, but the generic majority and observed peak concurrency still require general overflow.

The 174.8 execution hours equal about 32% of one node's elapsed time over the reporting window, or about 11% across three nodes, before adjusting for GPU performance or scheduling losses. Applying the observed 9-26% MiniMax cold slowdown to every workload would raise that mechanical range to 190-220 local hours, or 35-41% of one node and 12-14% of the fleet. This is a stress illustration, not a forecast, because:

The two-versus-three replay of July and August demand puts three nodes at about 7.6–7.7% occupancy, serving 88–96% of GPU-hours locally, with the rest spilling during concurrency peaks. At today's demand the fleet is justified by concurrency, model residency and failure isolation, and above all by the demand it is meant to create; not by average utilisation. The peak concurrency of 12 is why the rented lane remains part of the architecture.

Economics#

The return case is demand (the reframe). Modal displacement is the floor: what the fleet earns if demand does not grow at all. The figures below are dated and labelled; replace them with the dataplane and the first electricity bill once the nodes carry traffic.

Observed cloud spend#

The latest provider ledger available on 2026-09-23 contained:

Month Verda Modal Combined GPU-related provider spend
2026-06 US$941.58 US$788.26 US$1,729.84
2026-07 US$410.10 US$958.63 US$1,368.73
2026-08 through August 29 US$53.32 US$1,146.71 US$1,200.03

The ledger had last been refreshed on August 29 and was approximately 25 days old when these figures were captured. August is therefore incomplete. These figures establish the recent cost range, not a final invoice or current-month forecast.

The August Modal subtotal was distributed across both GPUs and their supporting resources:

Modal resource, through August 29 Cost
RTX Pro 6000 GPU US$660.52
H100 GPU US$168.48
L40S GPU US$111.03
Persistent volume US$104.31
Memory US$65.19
CPU US$35.07
Other small GPU/resource lines and rounding Approximately US$2.11
Modal subtotal US$1,146.71

This composition is why displacement is modelled as a share rather than the whole bill disappearing. Modal overflow, deployment safety, local downtime, and some supporting costs continue after local routing is active.

Displacement floor (ours, derived)#

Inputs:

avoided cloud spend   = 91.6–96.0% × RM5,040        = RM4,617–4,838 a month
net of electricity    = RM4,617 − 1,320 to 4,838 − 1,020 = RM3,297–3,818 a month
simple payback        = RM115,071 / RM3,297–3,818    = about 30–35 months

That is the "roughly break-even over a three-year life" behind the demand lever: on August's demand alone, displacement recovers the fleet's cost in about three years. The nodes are long-lived productive assets, so this scores how well the fleet is bought and run; it does not decide whether to own.

Electricity is a larger share of owned cost than a round per-kWh rate suggests. The household uses about 740 kWh a month, and any node load pushes the premises over TNB's 1,500 kWh cliff, after which every kWh costs about RM0.636 all-in, including the household's own use. Each node beyond three adds RM304–402 a month. Per-node branch metering and the first bill with the nodes running replace these estimates with observed kWh per completed job.

Where the advantage comes from#

Owning the compute is the foundation: persistent local workers remove our ephemeral GPU acquisition and container restore path, while ownership removes per-second GPU rent from the base load. These mechanisms support the fast and gap-fill lanes; they do not establish a unique capability or measured lead against TensorArt's self-operated fleet, SeaArt's optimized cloud deployments, or Haima's GPUaaS platform behind RunningHub (dated competitor findings). Paid-work opportunity cost still governs gap-fill admission. The system around the hardware is how that advantage turns into demand:

Rapid progress in open-weight models, including capable models in the roughly 27B class, makes commodity 32 GB cells more useful over time: more useful software fits on broadly available hardware, and the software layer can be replaced without replacing the service.

Alternatives considered#

Alternative Decision and reason Revisit when
Remain all-Modal Rejected for predictable base load; lifecycle, restore, tail, and mounted-storage exposure remain even after individual optimizations Not a standing option: renting is a bridge
Return to conventional rented VMs Rejected as the default; it restores the Verda-era availability, storage-price, transfer, and cache-reconstruction problems A specific GPU or a long job is needed tactically (lane 3)
Buy one RTX Pro 6000 Deferred; it offers 96 GB and faster single-job compute but only one service slot and failure domain at a higher bare-GPU price The larger-memory ladder names a workload
Build one dual-5090 workstation Rejected at this fleet size; density does not outweigh lane, cooling, power, host-contention, and failure-domain concentration, and VRAM is not transparently pooled Space, circuits, or fleet-management overhead becomes the constraint
Buy nodes gradually as demand is proven Rejected 2026-09-26/27; node prices are projected to rise first, so waiting costs more The price mechanism in the buy-now analysis changes
Gate ownership on rental prices or resale value Rejected 2026-09-26; owning is permanent Not revisited
Rent out idle owned capacity Last resort, only if there is no application for it Owned capacity has no lane-1 or lane-2 use

Operating boundaries and failure model#

The owned nodes are service capacity, not the durable system of record.

Each 5090 is rated for very high board power, so one protected electrical circuit per worker, correct supplied 12V-2x6 cabling, verified earthing, surge protection, thermal monitoring, and branch power telemetry are part of acceptance. Large UPS purchases remain evidence-gated rather than assumed necessary. The site stays on a single-phase domestic supply (founder, 2026-09-27); confirm the main breaker and supply rating before adding nodes 4 and 5.

Required production integration#

The hardware produces no return until the scheduler can use it. The hardware and benchmark documentation are ahead of the production runtime: the active orchestrator is centered on Modal, with older VM-provider paths dormant or commented out. Production integration is therefore the highest-priority work after physical acceptance. The first question (NOTEBOOK #8) is whether the dormant DataCrunch VM worker path (session pre-create, frontend registration and heartbeat, the getJob claim) can serve an always-on owned node, or whether a separate lane is cleaner.

The minimum owned-worker control plane must provide:

  1. authenticated registration, health heartbeats, and capability reporting for GPU, VRAM, software image, cached models, and free space;
  2. compatibility rules that keep incompatible or >32 GB jobs off local nodes;
  3. model-affinity placement that avoids unnecessary cache and residency churn;
  4. safe drain, retry, and spill-to-Modal behavior;
  5. per-run phase timings and explicit spill reasons such as local_busy, model_missing, vram_ineligible, and worker_unhealthy;
  6. node-level energy and cost attribution against avoided cloud spend; and
  7. the three capacity lanes: lane-1 paid work always claims ahead of lane-2 gap-fill work; lane-2 jobs are capped at the default timeout, never take the last free node and never spill to Modal; lane-3 long or >32 GB work routes to rented capacity.

The first deployment should favor correctness and observability over sophisticated bin-packing. The suggested first traffic is lane 2 (no latency promise), before paid work moves across. Static assignment of the most valuable model pools plus safe cloud spill is sufficient to prove the architecture. Dynamic residency can follow after real traces show that it will reduce cost or latency.

Acceptance and review plan#

Before paid traffic#

For each node:

First 30 production days#

Record at least:

Success criteria#

The founder's review signals (NOTEBOOK #8, for example at 60 days of live traffic) are the share of GPU-hours served locally, the cold-latency delta on Krea and MiniMax, node occupancy, reactivated users who pay within 30 days, any drop in fast-mode revenue per payer, and founder hours per week on the nodes.

The direction is working if, after normalization for demand:

Modal displacement is reported alongside, as the floor, not as the test.

If these conditions fail, first determine whether the cause is routing, software compatibility, demand shape, memory capacity, or site operation. Do not treat every failure as evidence that another GPU is required.

Capacity expansion policy#

Current position: owning is permanent; bulk buy now#

Founder decisions, second-brain/DECISIONS.md 2026-09-26 ("Owning compute is a permanent goal") and its 2026-09-27 updates.

Why now: should we buy more now? A first-principles price analysis reasons from how each part's price is set and applies that to our node's actual parts:

So any node we'll use before roughly late 2028 is cheapest bought now. Which GPU and build the next nodes use is worked in GPU theory, market price and workload performance; how to buy them is in procurement.

Site limits (founder, 2026-09-27): single-phase supply, domestic tariff, no immediate plan for solar. Five nodes at full load is about 6 kW including cooling on a single-phase supply shared with the household: it fits, with little margin (supply headroom).

A larger-memory node#

Prefer the smallest memory tier that clears a measured production constraint. Before buying a 96 GB RTX Pro 6000, collect:

The decision ladder is therefore 32 GB commodity cells first, 48/72 GB when measured demand requires them, and 96 GB only when the larger pool has a named workload and economic case. A glut of used data-centre cards (about 35% odds, mid-2027 to mid-2028) is the likely buying window (NOTEBOOK #8). Meanwhile >32 GB work runs in the rented lane.

Stop-doing list#

The next unit of engineering effort belongs in production routing, observability, and cost attribution.

Consequences#

Expected benefits#

Accepted costs and risks#

These are accepted because each node is independently replaceable, durable state stays elsewhere, and the rented lane remains available while the fleet grows.

What changed#

Evidence index#