GPU theory, market price and workload performance#

Status: Working reference
Last verified: 2026-09-29 Canonical for: the current GPU procurement evidence model, tier-2 GPU-class findings, and the benchmark and live-price work still required before the next purchase Complete tier-2 host and price record: tier2-amd-node-procurement.md

Decision frame#

The purchase question has three independent evidence layers:

  1. Theoretical capability from gpu-device-specs.json: VRAM, CUDA cores, FP32, memory bandwidth, board power, ECC, MIG and physical configuration.
  2. Current obtainable price from a dated, exact-SKU, seller- and stock-attributed market observation.
  3. Use-case performance from the fixed-protocol benchmark: warm throughput, cold latency, model reads, host memory, CPU use and power for a named workload.

No one layer substitutes for another. Specifications screen candidates but do not predict ComfyUI throughput. Card price alone omits the fixed cost of the host. A workload result from one rented machine remains confounded by its CPU, storage, power limit and provider path. The useful joined record is therefore:

GPU specification + exact market observation + fixed-protocol workload result + complete node BOM

The current dashboard joins the first three only partially: specification and benchmark identities are materialized, but purchase pricing is still a single manually captured value rather than an append-only live market history, and the purchase frontier is card-level rather than complete-node-level.

Market movement observed by the founder#

These are same-model Malaysian observations supplied by the founder, not inferred MSRP or resale values:

Exact model 2026-08-10 2026-09-27 Change Change over 48 days
Zotac RTX 5070 Ti 16 GB Solid OC SFF RM5,299 RM6,499 +RM1,200 +22.65%
Zotac RTX 5090 32 GB Solid OC RM20,999 RM22,799 +RM1,800 +8.57%

The 5090 therefore became cheaper relative to the 5070 Ti: its price multiple fell from 3.96x to 3.51x while the fixed-protocol Krea result shows 2.16x the warm throughput and twice the VRAM. Waiting for a production measurement before purchasing is not the current policy. Procurement and owned-scheduler integration proceed in parallel; a written quotation and deposit are the price-lock mechanism.

The older gpu-pricing.json entries and the complete 2026-09-22 Ideal Tech capture are immutable observations and must not be overwritten or silently replace this history. For example, its RM5,499 RTX 5070 Ti value is a Colorful Battle AX component in the Ideal Tech list captured 2026-09-22, not the Zotac history above. Exact SKU, seller, capture time, stock state, quote scope and quote validity are part of the price fact.

The saved September 22 Next.js Flight response is now also represented as an append-only flat snapshot under idealtech-price-snapshots/2026-09-22T140313Z/. It retains all 3,372 source rows: 2,902 purchasable products and 470 grouping rows. The originals remain untouched. Capture and DuckDB details are in idealtech-price-snapshots.md.

A live root Flight fetch on 2026-09-28 added a second immutable snapshot with 3,376 source rows and the site's own update label 26 Sept 2026, 1:46 pm. Its exact Zotac entries list the RTX 5070 Ti Solid SFF OC at RM5,899, RTX 5080 Solid Core OC at RM7,069, and RTX 5090 Solid OC at RM22,999. These retailer facts do not replace the founder-supplied observations above: their amounts, dates, seller context and stock/quote status differ.

Current workload evidence#

Warm runs/hour is the performance term used in the Pareto calculations because the intended service reuses a loaded process and cold disk/model loading is a separate operational cost. Here, resident means resident in the ComfyUI process, not necessarily wholly resident in GPU VRAM. If a workload exceeds VRAM, its warm result can include recurring host-memory/PCIe offload and remains valid as achieved steady-state throughput for that complete node, but it must not be described as a pure GPU-compute floor. Linux swap staying at zero does not establish VRAM residency. Pareto comparisons must keep workload, cache policy, VRAM/offload behavior and host class visible rather than mixing resident and host-assisted cohorts without qualification.

Krea 2 warm throughput#

GPU class VRAM Representative warm mean Warm runs/hour Independent observations Interpretation
RTX 5060 Ti 16 GB 18.034 s 199.6 1 Too slow for the current tier-2 target
RTX 5070 12 GB 12.796 s 281.3 1 unverified host Slower and lower-capacity than the tier-2 finalists
RTX 5070 Ti 16 GB 8.849 s 406.8 3 (1 verified, 2 unverified) Card-only value candidate
RTX 5080 16 GB 7.782 s 462.6 3 verified Current tier-2 front-runner
RTX 5090 32 GB 4.065 s 885.7 31 General/heavy owned tier
RTX PRO 4000 24 GB 9.469 s 380.2 1 Capacity/ECC, not Krea value
RTX PRO 5000 48 GB 4.867 s 739.6 1 Buy for 48/72 GB capacity
RTX PRO 6000 Max-Q 96 GB 3.765 s 956.1 1 96 GB and power-density option
RTX PRO 6000 Workstation 96 GB 3.016 s 1,193.6 1 Faster but large price premium
RTX PRO 6000 Server 96 GB 2.767 s 1,301.1 1 Specialized server configuration

The consumer representatives above use the mean of the independent host means when a class has multiple hosts. The RTX 5070 Ti and RTX 5080 now each have three independent physical-host observations, although two of the RTX 5070 Ti providers were unverified. These remain screening measurements rather than procurement-grade estimates of population performance. Host and storage effects are especially visible on the cold path.

What the first consumer-host measurements establish#

The RTX 5070 Ti Krea host used an Intel i3-12100. Its warm samples averaged about 1.04 process CPU cores, reached about 19.2 GiB process RSS and still delivered 409.5 runs/hour. The RTX 5080 host used an i7-11700 and PCIe 4.0 x16, reached 25.1 GiB/s median host-to-device bandwidth and delivered 462.7 warm runs/hour. This is strong evidence that Krea's resident path does not require the Ryzen 9 9950X or an X870E platform.

It does not prove that storage can be reduced indiscriminately. The observed Krea cold-minus-warm penalty was 14.7 seconds on the RTX 5070 Ti host's Crucial P3 Plus path, 24.9 seconds on the RTX 5080 host's ORICO path, and 4.4 seconds on a fast-storage RTX 5090 host. CPU, GPU and storage differ between those machines, so the direction is useful but the size of each component's effect requires a same-machine experiment.

Source results:

2026-09-28 fast-host repeat#

The repeat used the canonical schema-v4 resident protocol: pinned ComfyUI v0.37.0, five verified cold fresh-process runs and five warm different-seed runs per workload, 120 GB container disks and full-host GPU allocations. All six accepted files passed the structural gate, used zero process swap and had zero rejected cold attempts.

GPU CPU and storage Model reader Krea cold / warm Krea warm runs/hour MiniMax cold / warm MiniMax warm runs/hour
RTX 5060 Ti 16 GB Core Ultra 7 265K; HS-SSD-FUTURE 4096G 3.90 GiB/s 24.691 / 18.034 s 199.6 298.579 / 285.861 s 12.6
RTX 5080 16 GB i7-14700F; Corsair MP600 PRO LPX path 4.04 GiB/s 14.999 / 7.789 s 462.2 128.648 / 119.493 s 30.1
RTX 5090 32 GB Ryzen 7 9800X3D; Kingston Fury Renegade 8.13 GiB/s 8.948 / 4.262 s 844.7 68.855 / 60.541 s 59.5

The RTX 5080 repeat is the cleanest host-path comparison. Against the first i7-11700/ORICO observation, model-reader throughput increased from 1.25 to 4.04 GiB/s and median host-to-device bandwidth increased from 25.12 to 41.25 GiB/s. Krea cold latency fell from 32.672 to 14.999 seconds and its cold-minus-warm penalty fell from 24.892 to 7.210 seconds. Warm latency did not improve: 7.780 versus 7.789 seconds. The combined faster CPU, storage and PCIe path therefore removed cold-start delay but did not change resident image throughput. This experiment does not isolate CPU from storage individually.

The MiniMax result exposes a different frontier. On these strong hosts the RTX 5090 delivered 59.5 warm runs/hour, 1.97x the RTX 5080 and 4.72x the RTX 5060 Ti. The RTX 5060 Ti also delivered only 43.2% of the RTX 5080's warm Krea throughput. Its 16 GB capacity does not make it an efficient substitute for the upper cards unless its complete-node price is exceptionally low.

2026-09-28 expanded consumer cohort#

A later inventory round added two independent RTX 5070 Ti hosts, a third RTX 5080 host, and an RTX 5070 baseline. The RTX 5080 host was provider-verified; the RTX 5070 and both RTX 5070 Ti hosts were unverified, which was explicitly accepted for benchmark screening. Verification status remains part of every source record and must stay visible when filtering evidence. All eight consumer result files passed the same schema-v4 structural gate with zero rejected cold attempts.

GPU and host evidence CPU and advertised disk Krea cold / warm Krea warm runs/hour MiniMax cold / warm MiniMax warm runs/hour
RTX 5070 12 GB, machine 152222, unverified i9-14900KF; 5,468 MB/s 20.928 / 12.796 s 281.3 223.170 / 202.683 s 17.8
RTX 5070 Ti 16 GB, machine 152371, unverified Ryzen 9 9900X; 6,536 MB/s 13.945 / 8.925 s 403.4 148.039 / 141.562 s 25.4
RTX 5070 Ti 16 GB, machine 152326, unverified Ryzen 9 5900XT; 6,768 MB/s 16.809 / 8.831 s 407.7 151.445 / 141.562 s 25.4
RTX 5080 16 GB, machine 45749, verified Ryzen 7 7700; 4,843 MB/s 15.007 / 7.776 s 463.0 129.661 / 120.716 s 29.8

Across all three independent hosts per class, RTX 5070 Ti versus RTX 5080 warm means are 8.849 versus 7.782 seconds for Krea and 141.091 versus 119.897 seconds for MiniMax. The RTX 5080 therefore delivers 13.72% more Krea throughput and 17.68% more MiniMax throughput. The two new RTX 5070 Ti hosts have very different CPUs and storage yet converge on exactly 141.562 seconds warm MiniMax; the RTX 5080 hosts also remain tightly clustered. Faster host paths materially reduce cold latency but do not erase the steady-state GPU gap.

The same round screened a verified CMP 170HX: 7.739 seconds warm Krea (465.2 runs/hour) but 140.099 seconds warm MiniMax (25.7 runs/hour). Its HBM bandwidth makes it interesting for Krea, but the MiniMax result shows why it cannot be ranked from bandwidth alone. It is benchmark evidence, not a current consumer purchase candidate. Three additional RTX 5060 Ti acquisition attempts produced no measurements: two failed during dependency download and one unverified host never accepted SSH; every rented instance was automatically destroyed.

Provisional two-tier fleet#

The existing three RTX 5090 nodes cover the higher-capability lane. A lower-cost tier makes sense only when it is deliberately specialized rather than being the 5090 BOM with a cheaper GPU:

Tier Hardware role Workload role
Tier 1 RTX 5090, 32 GB, existing full platform >16 GB jobs, short heavy video/image work, MiniMax/WAN, protected paid latency and tier-2 overflow
Tier 2 RTX 5080, 16 GB, six-core CPU, modest board, 64 GB candidate, one good TLC NVMe, quality 850 W PSU Proven <=16 GB resident image families such as Krea, plus relaxed/free gap-fill
Rented bridge Provider capacity Long video jobs that would block an owned node, bursts, and >32 GB jobs

The RTX 5080 remains the tier-2 front-runner at complete-system pricing, not at every standalone price. The current AMD host analysis selects a CPU-direct PCIe 5.0 x16 GPU slot and Gen5 x4 model SSD path for RM6,263 with 32 GB or RM9,063 with 64 GB. The 32 GB figure is a Krea-only hypothesis pending a controlled memory-limit test; the observed 45.94 GiB MiniMax process peak requires the 64 GB build.

With the captured builder prices, the Colorful RTX 5080 at RM6,799 produces RM13,062/15,862 nodes and the ordinary Zotac RTX 5080 at RM7,069 produces RM13,332/16,132 nodes at 32/64 GB. Both beat the corresponding RTX 5070 Ti nodes on measured Krea capex/throughput. The founder's standalone Zotac/Gigabyte/MSI offers at RM8,089/RM8,399/RM8,790 do not: for a Krea-led node, use the cheaper RTX 5070 Ti if the RTX 5080 cannot be obtained inside the complete-system price gate. Exact calculations, component choices and quote conditions are in the tier-2 AMD procurement record.

Complete-node Pareto boundaries#

Use the measured workload throughput and a real complete-node price. For candidate c and reference r:

maximum_candidate_node_price = reference_node_price
                               * candidate_warm_runs_per_hour
                               / reference_warm_runs_per_hour

Against PC2's paid RM34,999 and the current Krea representatives:

Candidate Maximum complete-node price matching PC2's Krea capex/throughput
RTX 5060 Ti RM7,888
RTX 5070 RM11,116
RTX 5070 Ti RM16,076
RTX 5080 RM18,282

This is why a tier-2 host can be rational even though two smaller cards duplicate CPU, memory and storage. It is also why card-only price/performance is insufficient: every ringgit of common host overhead favors the faster GPU. The frontier must additionally expose single-job VRAM, node slots and energy. Two RTX 5080 cards provide about 4.5% more aggregate warm Krea throughput than one RTX 5090, but use two nodes and 720 W combined GPU TGP rather than one node and 575 W.

RTX PRO interpretation#

RTX PRO changes the capacity frontier, not the current Krea value conclusion. The catalog covers 16–96 GB, including distinct RTX PRO 4500 workstation/server boards, RTX PRO 5000 48/72 GB, and RTX PRO 6000 Max-Q/workstation/server configurations. Buy the smallest professional memory tier only after a real workload requires more than 32 GB, ECC or MIG. Current 48/72/96 GB rental measurements remain the bridge and evidence source.

Required next evidence#

Current conclusion#

The evidence supports buying now while developing the scheduler in parallel. Three independent observations per finalist establish that the RTX 5080 provides 13.72% more Krea and 17.68% more MiniMax warm throughput than the RTX 5070 Ti. On the lean AMD host, the RM6,799 Colorful and RM7,069 ordinary Zotac builder cards make the RTX 5080 the better complete-node value; the RM8,089–RM8,790 standalone offers make the RTX 5070 Ti the better Krea value. The RTX 5070 and RTX 5060 Ti remain outside the current complete-node frontier. More RTX 5090 nodes remain the density and capability choice; RTX PRO remains a measured-capacity purchase rather than the default throughput purchase.