GPU theory, market price and workload performance#
Status: Working reference
Last verified: 2026-09-29 Canonical for: the current GPU procurement evidence model, tier-2 GPU-class findings, and the benchmark and live-price work still required before the next purchase Complete tier-2 host and price record:tier2-amd-node-procurement.md
Decision frame#
The purchase question has three independent evidence layers:
- Theoretical capability from
gpu-device-specs.json: VRAM, CUDA cores, FP32, memory bandwidth, board power, ECC, MIG and physical configuration. - Current obtainable price from a dated, exact-SKU, seller- and stock-attributed market observation.
- Use-case performance from the fixed-protocol benchmark: warm throughput, cold latency, model reads, host memory, CPU use and power for a named workload.
No one layer substitutes for another. Specifications screen candidates but do not predict ComfyUI throughput. Card price alone omits the fixed cost of the host. A workload result from one rented machine remains confounded by its CPU, storage, power limit and provider path. The useful joined record is therefore:
GPU specification + exact market observation + fixed-protocol workload result + complete node BOM
The current dashboard joins the first three only partially: specification and benchmark identities are materialized, but purchase pricing is still a single manually captured value rather than an append-only live market history, and the purchase frontier is card-level rather than complete-node-level.
Market movement observed by the founder#
These are same-model Malaysian observations supplied by the founder, not inferred MSRP or resale values:
| Exact model | 2026-08-10 | 2026-09-27 | Change | Change over 48 days |
|---|---|---|---|---|
| Zotac RTX 5070 Ti 16 GB Solid OC SFF | RM5,299 | RM6,499 | +RM1,200 | +22.65% |
| Zotac RTX 5090 32 GB Solid OC | RM20,999 | RM22,799 | +RM1,800 | +8.57% |
The 5090 therefore became cheaper relative to the 5070 Ti: its price multiple fell from 3.96x to 3.51x while the fixed-protocol Krea result shows 2.16x the warm throughput and twice the VRAM. Waiting for a production measurement before purchasing is not the current policy. Procurement and owned-scheduler integration proceed in parallel; a written quotation and deposit are the price-lock mechanism.
The older gpu-pricing.json entries and the complete 2026-09-22 Ideal Tech capture are immutable observations
and must not be overwritten or silently replace this history.
For example, its RM5,499 RTX 5070 Ti value is a Colorful Battle AX component in the Ideal Tech list captured
2026-09-22, not the Zotac history above. Exact SKU, seller, capture time, stock state, quote scope and quote
validity are part of the price fact.
The saved September 22 Next.js Flight response is now also represented as an append-only flat snapshot under
idealtech-price-snapshots/2026-09-22T140313Z/. It retains all 3,372 source rows: 2,902 purchasable products
and 470 grouping rows. The originals remain untouched. Capture and DuckDB details are in
idealtech-price-snapshots.md.
A live root Flight fetch on 2026-09-28 added a second immutable snapshot with 3,376 source rows and the site's
own update label 26 Sept 2026, 1:46 pm. Its exact Zotac entries list the RTX 5070 Ti Solid SFF OC at RM5,899,
RTX 5080 Solid Core OC at RM7,069, and RTX 5090 Solid OC at RM22,999. These retailer facts do not replace the
founder-supplied observations above: their amounts, dates, seller context and stock/quote status differ.
Current workload evidence#
Warm runs/hour is the performance term used in the Pareto calculations because the intended service reuses a
loaded process and cold disk/model loading is a separate operational cost. Here, resident means resident in
the ComfyUI process, not necessarily wholly resident in GPU VRAM. If a workload exceeds VRAM, its warm result
can include recurring host-memory/PCIe offload and remains valid as achieved steady-state throughput for that
complete node, but it must not be described as a pure GPU-compute floor. Linux swap staying at zero does not
establish VRAM residency. Pareto comparisons must keep workload, cache policy, VRAM/offload behavior and host
class visible rather than mixing resident and host-assisted cohorts without qualification.
Krea 2 warm throughput#
| GPU class | VRAM | Representative warm mean | Warm runs/hour | Independent observations | Interpretation |
|---|---|---|---|---|---|
| RTX 5060 Ti | 16 GB | 18.034 s | 199.6 | 1 | Too slow for the current tier-2 target |
| RTX 5070 | 12 GB | 12.796 s | 281.3 | 1 unverified host | Slower and lower-capacity than the tier-2 finalists |
| RTX 5070 Ti | 16 GB | 8.849 s | 406.8 | 3 (1 verified, 2 unverified) | Card-only value candidate |
| RTX 5080 | 16 GB | 7.782 s | 462.6 | 3 verified | Current tier-2 front-runner |
| RTX 5090 | 32 GB | 4.065 s | 885.7 | 31 | General/heavy owned tier |
| RTX PRO 4000 | 24 GB | 9.469 s | 380.2 | 1 | Capacity/ECC, not Krea value |
| RTX PRO 5000 | 48 GB | 4.867 s | 739.6 | 1 | Buy for 48/72 GB capacity |
| RTX PRO 6000 Max-Q | 96 GB | 3.765 s | 956.1 | 1 | 96 GB and power-density option |
| RTX PRO 6000 Workstation | 96 GB | 3.016 s | 1,193.6 | 1 | Faster but large price premium |
| RTX PRO 6000 Server | 96 GB | 2.767 s | 1,301.1 | 1 | Specialized server configuration |
The consumer representatives above use the mean of the independent host means when a class has multiple hosts. The RTX 5070 Ti and RTX 5080 now each have three independent physical-host observations, although two of the RTX 5070 Ti providers were unverified. These remain screening measurements rather than procurement-grade estimates of population performance. Host and storage effects are especially visible on the cold path.
What the first consumer-host measurements establish#
The RTX 5070 Ti Krea host used an Intel i3-12100. Its warm samples averaged about 1.04 process CPU cores, reached about 19.2 GiB process RSS and still delivered 409.5 runs/hour. The RTX 5080 host used an i7-11700 and PCIe 4.0 x16, reached 25.1 GiB/s median host-to-device bandwidth and delivered 462.7 warm runs/hour. This is strong evidence that Krea's resident path does not require the Ryzen 9 9950X or an X870E platform.
It does not prove that storage can be reduced indiscriminately. The observed Krea cold-minus-warm penalty was 14.7 seconds on the RTX 5070 Ti host's Crucial P3 Plus path, 24.9 seconds on the RTX 5080 host's ORICO path, and 4.4 seconds on a fast-storage RTX 5090 host. CPU, GPU and storage differ between those machines, so the direction is useful but the size of each component's effect requires a same-machine experiment.
Source results:
vast-krea-2-rtx-5070-ti-machine-149605-suite-vast-20260927T141752Z-cache-evicted-v1.jsonvast-krea-2-rtx-5080-machine-151209-suite-vast-20260927T143428Z-cache-evicted-v1.jsonvast-krea-2-rtx-5060-ti-machine-117553-suite-vast-20260928T071323Z-cache-evicted-v1.jsonvast-krea-2-rtx-5080-machine-141717-suite-vast-20260928T072523Z-cache-evicted-v1.jsonvast-krea-2-rtx-5090-machine-129520-suite-vast-20260928T071323Z-cache-evicted-v1.jsonvast-krea-2-rtx-5090-raid0-machine-112410-cache-evicted-v1.json
2026-09-28 fast-host repeat#
The repeat used the canonical schema-v4 resident protocol: pinned ComfyUI v0.37.0, five verified cold
fresh-process runs and five warm different-seed runs per workload, 120 GB container disks and full-host GPU
allocations. All six accepted files passed the structural gate, used zero process swap and had zero rejected
cold attempts.
| GPU | CPU and storage | Model reader | Krea cold / warm | Krea warm runs/hour | MiniMax cold / warm | MiniMax warm runs/hour |
|---|---|---|---|---|---|---|
| RTX 5060 Ti 16 GB | Core Ultra 7 265K; HS-SSD-FUTURE 4096G | 3.90 GiB/s | 24.691 / 18.034 s | 199.6 | 298.579 / 285.861 s | 12.6 |
| RTX 5080 16 GB | i7-14700F; Corsair MP600 PRO LPX path | 4.04 GiB/s | 14.999 / 7.789 s | 462.2 | 128.648 / 119.493 s | 30.1 |
| RTX 5090 32 GB | Ryzen 7 9800X3D; Kingston Fury Renegade | 8.13 GiB/s | 8.948 / 4.262 s | 844.7 | 68.855 / 60.541 s | 59.5 |
The RTX 5080 repeat is the cleanest host-path comparison. Against the first i7-11700/ORICO observation, model-reader throughput increased from 1.25 to 4.04 GiB/s and median host-to-device bandwidth increased from 25.12 to 41.25 GiB/s. Krea cold latency fell from 32.672 to 14.999 seconds and its cold-minus-warm penalty fell from 24.892 to 7.210 seconds. Warm latency did not improve: 7.780 versus 7.789 seconds. The combined faster CPU, storage and PCIe path therefore removed cold-start delay but did not change resident image throughput. This experiment does not isolate CPU from storage individually.
The MiniMax result exposes a different frontier. On these strong hosts the RTX 5090 delivered 59.5 warm runs/hour, 1.97x the RTX 5080 and 4.72x the RTX 5060 Ti. The RTX 5060 Ti also delivered only 43.2% of the RTX 5080's warm Krea throughput. Its 16 GB capacity does not make it an efficient substitute for the upper cards unless its complete-node price is exceptionally low.
2026-09-28 expanded consumer cohort#
A later inventory round added two independent RTX 5070 Ti hosts, a third RTX 5080 host, and an RTX 5070 baseline. The RTX 5080 host was provider-verified; the RTX 5070 and both RTX 5070 Ti hosts were unverified, which was explicitly accepted for benchmark screening. Verification status remains part of every source record and must stay visible when filtering evidence. All eight consumer result files passed the same schema-v4 structural gate with zero rejected cold attempts.
| GPU and host evidence | CPU and advertised disk | Krea cold / warm | Krea warm runs/hour | MiniMax cold / warm | MiniMax warm runs/hour |
|---|---|---|---|---|---|
| RTX 5070 12 GB, machine 152222, unverified | i9-14900KF; 5,468 MB/s | 20.928 / 12.796 s | 281.3 | 223.170 / 202.683 s | 17.8 |
| RTX 5070 Ti 16 GB, machine 152371, unverified | Ryzen 9 9900X; 6,536 MB/s | 13.945 / 8.925 s | 403.4 | 148.039 / 141.562 s | 25.4 |
| RTX 5070 Ti 16 GB, machine 152326, unverified | Ryzen 9 5900XT; 6,768 MB/s | 16.809 / 8.831 s | 407.7 | 151.445 / 141.562 s | 25.4 |
| RTX 5080 16 GB, machine 45749, verified | Ryzen 7 7700; 4,843 MB/s | 15.007 / 7.776 s | 463.0 | 129.661 / 120.716 s | 29.8 |
Across all three independent hosts per class, RTX 5070 Ti versus RTX 5080 warm means are 8.849 versus 7.782 seconds for Krea and 141.091 versus 119.897 seconds for MiniMax. The RTX 5080 therefore delivers 13.72% more Krea throughput and 17.68% more MiniMax throughput. The two new RTX 5070 Ti hosts have very different CPUs and storage yet converge on exactly 141.562 seconds warm MiniMax; the RTX 5080 hosts also remain tightly clustered. Faster host paths materially reduce cold latency but do not erase the steady-state GPU gap.
The same round screened a verified CMP 170HX: 7.739 seconds warm Krea (465.2 runs/hour) but 140.099 seconds warm MiniMax (25.7 runs/hour). Its HBM bandwidth makes it interesting for Krea, but the MiniMax result shows why it cannot be ranked from bandwidth alone. It is benchmark evidence, not a current consumer purchase candidate. Three additional RTX 5060 Ti acquisition attempts produced no measurements: two failed during dependency download and one unverified host never accepted SSH; every rented instance was automatically destroyed.
Provisional two-tier fleet#
The existing three RTX 5090 nodes cover the higher-capability lane. A lower-cost tier makes sense only when it is deliberately specialized rather than being the 5090 BOM with a cheaper GPU:
| Tier | Hardware role | Workload role |
|---|---|---|
| Tier 1 | RTX 5090, 32 GB, existing full platform | >16 GB jobs, short heavy video/image work, MiniMax/WAN, protected paid latency and tier-2 overflow |
| Tier 2 | RTX 5080, 16 GB, six-core CPU, modest board, 64 GB candidate, one good TLC NVMe, quality 850 W PSU | Proven <=16 GB resident image families such as Krea, plus relaxed/free gap-fill |
| Rented bridge | Provider capacity | Long video jobs that would block an owned node, bursts, and >32 GB jobs |
The RTX 5080 remains the tier-2 front-runner at complete-system pricing, not at every standalone price. The current AMD host analysis selects a CPU-direct PCIe 5.0 x16 GPU slot and Gen5 x4 model SSD path for RM6,263 with 32 GB or RM9,063 with 64 GB. The 32 GB figure is a Krea-only hypothesis pending a controlled memory-limit test; the observed 45.94 GiB MiniMax process peak requires the 64 GB build.
With the captured builder prices, the Colorful RTX 5080 at RM6,799 produces RM13,062/15,862 nodes and the ordinary Zotac RTX 5080 at RM7,069 produces RM13,332/16,132 nodes at 32/64 GB. Both beat the corresponding RTX 5070 Ti nodes on measured Krea capex/throughput. The founder's standalone Zotac/Gigabyte/MSI offers at RM8,089/RM8,399/RM8,790 do not: for a Krea-led node, use the cheaper RTX 5070 Ti if the RTX 5080 cannot be obtained inside the complete-system price gate. Exact calculations, component choices and quote conditions are in the tier-2 AMD procurement record.
Complete-node Pareto boundaries#
Use the measured workload throughput and a real complete-node price. For candidate c and reference r:
maximum_candidate_node_price = reference_node_price
* candidate_warm_runs_per_hour
/ reference_warm_runs_per_hour
Against PC2's paid RM34,999 and the current Krea representatives:
| Candidate | Maximum complete-node price matching PC2's Krea capex/throughput |
|---|---|
| RTX 5060 Ti | RM7,888 |
| RTX 5070 | RM11,116 |
| RTX 5070 Ti | RM16,076 |
| RTX 5080 | RM18,282 |
This is why a tier-2 host can be rational even though two smaller cards duplicate CPU, memory and storage. It is also why card-only price/performance is insufficient: every ringgit of common host overhead favors the faster GPU. The frontier must additionally expose single-job VRAM, node slots and energy. Two RTX 5080 cards provide about 4.5% more aggregate warm Krea throughput than one RTX 5090, but use two nodes and 720 W combined GPU TGP rather than one node and 575 W.
RTX PRO interpretation#
RTX PRO changes the capacity frontier, not the current Krea value conclusion. The catalog covers 16–96 GB, including distinct RTX PRO 4500 workstation/server boards, RTX PRO 5000 48/72 GB, and RTX PRO 6000 Max-Q/workstation/server configurations. Buy the smallest professional memory tier only after a real workload requires more than 32 GB, ECC or MIG. Current 48/72/96 GB rental measurements remain the bridge and evidence source.
Required next evidence#
- Add one more verified RTX 5070 Ti host. Three independent fixed-protocol observations now exist, but two came from unverified providers. Keep the verification distinction visible and replace screening uncertainty with another verified physical host when inventory appears.
- Repeat RTX 5060 Ti only on a different ready host. The one accepted observation is sufficient to keep it outside the current frontier; three later provisioning/setup failures are not GPU evidence.
- Run one same-machine tier-2 matrix. Hold the RTX 5080 fixed; compare 32/48/64 GB cgroup limits, an affordable TLC Gen4 drive against the Samsung 9100 Pro, and 80/90/100% GPU power. Record warm throughput, cold p50/p95, model-read rate, RSS, swap, wall power and failures.
- Continue the verified Ideal Tech capture.
extract_idealtech_prices.tsreads the root Next.js Flight catalogue, can reproduce quote-to-edit navigation, preserves each raw response as a timestamped snapshot, and emits flat JSON for DuckDB. Keep the retailer'sproduct_nameexact. Do not normalize, rewrite or merge part/component names in the extraction layer. Any GPU identity or component-family mapping belongs in a separate attributed join table. - Materialize append-only market observations. Never overwrite an older price. Derive the current purchase view from the newest available in-stock observation and show staleness explicitly.
- Add complete-node options to the economics view. Join component BOM observations, GPU specifications and fixed-protocol performance. Maintain separate frontiers for <=16 GB image work, <=32 GB general work and larger-memory capacity; VRAM is not pooled.
Current conclusion#
The evidence supports buying now while developing the scheduler in parallel. Three independent observations per finalist establish that the RTX 5080 provides 13.72% more Krea and 17.68% more MiniMax warm throughput than the RTX 5070 Ti. On the lean AMD host, the RM6,799 Colorful and RM7,069 ordinary Zotac builder cards make the RTX 5080 the better complete-node value; the RM8,089–RM8,790 standalone offers make the RTX 5070 Ti the better Krea value. The RTX 5070 and RTX 5060 Ti remain outside the current complete-node frontier. More RTX 5090 nodes remain the density and capability choice; RTX PRO remains a measured-capacity purchase rather than the default throughput purchase.