C. GPU rental market#

Status: Plan Last verified: 2026-09-26 Canonical for: the Gemini Deep Research prompt for the GPU-rental workstream

Feeds decision 2 (regret trigger). Paste everything in the block below.

CONTEXT
I run a small AI inference business in Malaysia (image and video generation). In September 2026 I bought three consumer PCs, each with one NVIDIA RTX 5090 (32 GB), a Ryzen 9 9950X, 64–128 GB DDR5 and a PCIe 5.0 NVMe drive, for about RM115,000 in total (roughly USD 27,000). They serve my baseline load; bursts and jobs that need more than 32 GB of VRAM go to serverless cloud GPUs. I do not need data-centre redundancy; brief downtime is acceptable because traffic fails over to the cloud.

Amortised over 36 months with local electricity, the owned cost works out to roughly USD 0.40 per GPU-hour when the nodes are busy. My planning horizon is October 2026 to September 2029.

The decisions this research informs:
1. Whether and when to buy a fourth node (now, after memory prices fall, or never).
2. The rental price at which owning becomes the wrong call. The relevant alternative is a dedicated GPU server with fast local NVMe, not per-second serverless.
3. What a used RTX 5090 will be worth in 18–36 months.
4. Malaysia-specific risks to the cost of owning.

TASK: THE GPU RENTAL MARKET
My workload loads large models (10–40 GB each) from a long tail of hundreds of models, so most runs start cold. What matters is a GPU with fast local NVMe storage that keeps the models resident, reliably available, near Southeast Asia. Cheap GPU-hours with slow network storage or unreliable spot availability do not solve my problem. Answer these questions:

1. What do on-demand, spot and reserved GPU-hours cost now for the RTX 5090, RTX 4090, L40S, RTX PRO 6000 (Blackwell), H100 and B200? Cover marketplaces (for example Vast.ai, RunPod, TensorDock, SF Compute) and neoclouds (for example CoreWeave, Lambda, Nebius, Crusoe, Verda/DataCrunch). Show the 12–24 month trend and any published GPU rental price index (for example SemiAnalysis's).
2. What do dedicated monthly GPU servers with local NVMe cost? Examples: Hetzner GPU servers, OVHcloud, Leaseweb, and providers in Singapore, Malaysia, Japan or Hong Kong. Give per-month and implied per-GPU-hour prices, NVMe capacity and speed, and availability and lead time.
3. How reliable is capacity? Evidence of spot availability problems, waitlists or preemption rates for consumer and mid-range GPUs.
4. How healthy are the neoclouds financially: debt loads, GPU-backed loans and their interest rates, utilisation, contract lengths, depreciation assumptions (useful-life debates), customer concentration? Have any failed or been distressed, and what happened to their GPUs?
5. If AI spending slows or a neocloud fails, how fast do rental prices fall, and for which GPU classes? Does that flood reach consumer-class GPUs and local-NVMe servers, or mostly H100/B200 clusters?
6. Serverless GPU platforms (for example Modal, RunPod Serverless, Replicate, fal): current per-second pricing, how model loading and storage are handled, and any trend in cold-start performance.
7. Where is the break-even? At what dedicated-server price per GPU-hour, with comparable local NVMe and reliability, does renting beat owning a USD 9,000 node that costs about USD 0.40 per GPU-hour all-in?

Tell me: (a) how likely a comparable dedicated local-NVMe GPU server near Southeast Asia is to fall below USD 0.40 per GPU-hour within 12, 24 and 36 months, and (b) the specific price indices or listings to watch quarterly as my regret trigger.

SOURCES
- Today is late September 2026. Prioritise sources published in the last 90 days. Use older material only as a historical baseline, and label it as such.
- Prefer leading and primary sources: independent semiconductor analysts (for example SemiAnalysis), market trackers (for example TrendForce), company earnings calls and transcripts, regulatory filings, price indices, marketplace price histories and sold listings, supply-chain reporting.
- Investment banks, consultancies and auditors are acceptable but usually lag. When you cite them, say so.
- Where credible sources disagree, show both positions with their reasoning. Do not average them.
- Give a URL for every sourced claim.

OUTPUT FORMAT
1. Answer: at most 10 bullets answering the questions above directly, with numbers.
2. Evidence table with columns: claim | number | as-of date | source (URL) | observed / forecast / rumour | leading / lagging.
3. Scenarios for October 2026 to September 2029: base, upside and downside for my position. For each: a rough probability, what happens to the prices that matter to me, and the indicator that would tell me early that this scenario is unfolding.
4. What would change the call: concrete thresholds (indicator, level, where to watch it, how often it updates).
5. Unknowns: what you could not find or could not verify.

Do not give generic investment advice or stock recommendations. Keep the focus on the decisions above.