Vast benchmark operator handbook#
Status: Reference Last verified: 2026-09-28 Canonical for: launching, validating, preserving, and publishing ComfyICU GPU benchmarks on Vast.ai
Scope#
This is the start-to-finish runbook for adding a machine to the current benchmark dataset. It covers selecting a Vast offer, running the fixed cold/warm protocol, checking the resulting schema-v4 JSON, confirming that the rented instance was destroyed, and viewing the result in the dynamic dashboard.
The normal deliverable is one immutable JSON result per workload under orchestrator/benchmarks/. A
canonical full-suite launch produces separate MiniMax H3 and Krea 2 files from one provisioned machine and
one software setup.
This handbook does not decide which GPU or storage layout to buy. The analysis documents linked under
Related interpret the evidence. It also does not cover production GPU provisioning, which lives
in orchestrator/docs/gpu-operations.md.
Model#
The units being measured#
- A source run is a persisted ComfyICU workflow, prompt, and resolved file list read from the production
database through the read-only
warehouse_roaccount. - A workload gives that source run a stable analytical identity. The built-in workloads are
minimax-h3andkrea-2. - A suite runs two or more built-in workloads after one machine setup. All suite files are downloaded before the first workload, so the setup and network cost are shared.
- A pair is one cold run followed immediately by one warm run in the same ComfyUI process. The warm run changes the seed. ComfyUI is stopped after the pair.
- A benchmark result is the complete JSON evidence for one workload on one machine. It includes raw samples, actual and advertised hardware, storage topology, PCIe links, resource traces, storage probes, accelerator-path probes, software versions, source identity, and provider pricing.
- The independent sampling unit is the physical Vast machine, identified by
machine_id. Reacquiring the same machine is useful for repeatability, but it is not another independent host.
The canonical protocol#
New comparative results use schema v4, --cache-mode resident, and five pairs:
prepare all suite files once
repeat 5 times per workload:
evict every resolved model/input file with fsync + POSIX_FADV_DONTNEED
start a fresh ComfyUI process
cold run
require process block reads >= 80% of resolved bytes
warm run in the same process with a different seed
stop ComfyUI, clearing resident model state
capture model-reader, fio, CPU-memory, H2D/D2H, block-layer, PCIe, CPU,
RAM, GPU, storage-topology, software-version, and pricing evidence
This separates three useful quantities:
- Warm time is the same-process steady-state time after cold file loading has been removed. It approximates
the GPU/compute floor only when the workload's active working set fits in VRAM without recurring host-RAM or
PCIe offload.
residentdescribes ComfyUI's process cache; it does not prove that every weight remains in VRAM. - Cold time is the real fresh-process experience after file-page eviction.
- Cold penalty is paired cold wall time minus paired warm wall time on the same machine. It is the useful application-level storage/loading outcome, but it still includes CPU deserialization, memory copies, and other cold-only work.
The cold-read verifier proves that the ComfyUI process incurred block I/O after page eviction. On a nested container/loop/device-mapper path, it does not by itself prove that every byte reached physical flash. Use the captured cgroup, block-layer, NVMe, filesystem, and model-reader evidence together.
Software baseline#
The runner currently installs:
nvidia/cuda:13.0.3-cudnn-devel-ubuntu24.04- PyTorch
2.9.1+cu130, torchvision0.24.1, and torchaudio2.9.1 - the exact ComfyUI release supplied with
--comfyui-tag
The current RTX 5090 storage cohort uses ComfyUI v0.37.0. Use that tag when adding a directly comparable
host. When starting a new cohort, resolve the latest release once, record it, and then pass that exact tag
to every run; do not let “latest” move between machines:
COMFYUI_TAG="$(curl -fsS \
-H 'Accept: application/vnd.github+json' \
https://api.github.com/repos/Comfy-Org/ComfyUI/releases/latest | jq -r .tag_name)"
test -n "$COMFYUI_TAG" && test "$COMFYUI_TAG" != null
printf 'Pinned ComfyUI tag: %s\n' "$COMFYUI_TAG"
How it works#
Golden path#
For a normal new-machine sample:
- Load
.envand pass the local preflight. - Freeze the experiment's ComfyUI tag and five-pair resident-cache protocol.
- Inspect a verified Vast offer, especially its download price, host RAM, GPU power/fraction, and machine ID.
- Launch
minimax-h3,krea-2by exact offer ID inside tmux. - Let the runner provision, set up once, run both workloads, write two JSON files, and tear down.
- Confirm that Vast has no leftover benchmark instance.
- Run the structural gate and human metadata summary on both JSON files.
- Open the dashboard, inspect both run-detail views, and commit only the two immutable result files.
1. Prepare the local operator environment#
Run benchmark commands from the comfyicu repository root. The repository dependencies must already be
installed. You need:
- Bun, pnpm, PostgreSQL's
psql, OpenSSH,curl,jq, and optionallytmux; VASTAI_API_KEYand the read-onlyPGPASSWORDin the repository.env;- an Ed25519 SSH private key and matching
.pubfile. The default is~/.ssh/id_ed25519; or pass--ssh-keyexplicitly; - network access to the production read replica and to Vast.ai.
Load the secrets without printing them, then run the preflight:
set -a
source .env
set +a
for command_name in bun pnpm psql ssh curl jq; do
command -v "$command_name" >/dev/null || echo "missing: $command_name"
done
test -n "$VASTAI_API_KEY" && echo "Vast API key: present"
test -n "$PGPASSWORD" && echo "read-only DB password: present"
test -f ~/.ssh/id_ed25519 && test -f ~/.ssh/id_ed25519.pub && echo "SSH key pair: present"
psql -X -qAt \
-h gke-pgbouncer.tail6bb0f.ts.net \
-U warehouse_ro \
-d comfyicu \
-c 'select 1'
bun run orchestrator/benchmark_vast.ts --help >/dev/null
Expected database output is 1. The benchmark lookup is read-only; never substitute a write-capable
production account.
2. Define the comparison before renting anything#
Write down the question and the controlled dimensions. At minimum, freeze:
- workload or full suite;
- ComfyUI tag;
- cache mode (
residentfor the current dataset); - sample count (five for a canonical result);
- required GPU name and VRAM range;
- whether a one-GPU physical host is required;
- any power, GPU PCIe-width, storage, or host-memory preference;
- whether the run is a new independent machine or a reacquisition for repeatability.
Prefer a full suite when the machine is new. It amortizes the expensive environment and model downloads
while producing two independent workload result files. A one-sample run is only a paid diagnostic and must
go under orchestrator/benchmarks/diagnostics/; it is not evidence for the main dashboard.
3. Inspect and choose a Vast offer#
For a first run, inspect offers in the Vast console and pass the exact offer ID. That prevents automatic selection from surprising you with expensive bandwidth. Use these checks:
| Check | Canonical default | Why |
|---|---|---|
| Verification | Verified | Reduces unexplained host risk |
| Allocation | Exactly one GPU | The runner executes on one GPU |
| Physical host | Prefer gpu_frac = 1 |
Avoids shared CPU, RAM, and storage for baseline runs |
| Host RAM | 64 GB preferred; 48 GB hard minimum | MiniMax cold loading has approached the minimum |
| Container disk | 120 GB | Holds the environment, both workloads, probes, and output |
| Reliability | At least 0.97; 0.98 preferred | Avoids wasting the large download |
| Download rate | Inspect price per GB, not only Mbps | Egress can cost more than GPU time |
| CUDA compatibility | cuda_max_good >= 13 |
The pinned CUDA 13 container cannot initialize on older host drivers |
| GPU power | Prefer the target class; record rather than hide variants | Power affects the warm compute floor |
| GPU PCIe | Prefer full width for the target setup | Host-to-device loading can be link-sensitive |
| Disk headline | Treat only as a search hint | Vast's number is not the application read path |
The built-in suite currently prepares 61,108,590,469 bytes, or about 61.1 decimal GB. Before renting, estimate model traffic as:
download cost = advertised download USD/GB × 61.1
At $0.04/GB, models alone cost about $2.44, before CUDA, Python, ComfyUI, or package downloads. Prefer a
cheap download rate even if the GPU hourly price is slightly higher. Region is a preference, not a hard
requirement; use --prefer-country '' when searching globally.
Offer ID and machine ID are different:
--offer-idacquires one specific currently rentable offer.--machine-idfinds a new offer on a previously measured physical machine.- with neither, the runner searches matching offers itself. This is convenient but does not impose a network-price ceiling, so inspect the selected-offer log immediately.
--instance-idreuses an instance that is already running.
The default rejects a GPU allocated from a multi-GPU physical host. --allow-multi-gpu-host is appropriate
only when shared-host behavior is part of the experiment. A reported fraction such as 0.25 can mean one
whole allocated GPU on a four-GPU host; the main concern is shared host resources, not necessarily a
quarter of the GPU's compute.
4. Launch the canonical full suite#
Use a durable terminal so an SSH disconnect from your workstation does not kill the local runner:
tmux new -s vast-benchmark
Inside that session, load .env, set the offer and cohort constants, and launch:
set -a
source .env
set +a
OFFER_ID=12345678
GPU_NAME='RTX 5090'
COMFYUI_TAG='v0.37.0'
bun run orchestrator/benchmark_vast.ts \
--offer-id "$OFFER_ID" \
--gpu "$GPU_NAME" \
--min-vram 30 \
--max-vram 34 \
--min-reliability 0.97 \
--min-internet 0 \
--prefer-country '' \
--disk 120 \
--samples 5 \
--cache-mode resident \
--comfyui-tag "$COMFYUI_TAG" \
--suite minimax-h3,krea-2 \
--output-dir orchestrator/benchmarks
Replace the example offer ID. The GPU string must exactly match Vast's gpu_name; change the VRAM bounds
for other GPU classes. Add --min-power only when power is an admission requirement, not merely a
preference.
To reacquire a known physical host, replace --offer-id with:
--machine-id 45827
To let the runner search globally, omit both selectors. It queries verified, rentable, one-GPU offers,
applies the supplied constraints, and normally requires gpu_frac = 1.
To reuse an already-running instance, replace the selector with:
--instance-id 52270704 --destroy-after
Without --destroy-after, a reused instance remains running. A newly provisioned instance is destroyed by
default. --keep-instance overrides that default and should be rare.
Detach from tmux with Ctrl-b d and return with:
tmux attach -t vast-benchmark
5. Understand the lifecycle output#
The runner:
- reads the source run and resolved files through
warehouse_ro; - selects and acquires the offer, then logs the Vast instance ID;
- waits up to ten minutes for the instance and another ten minutes for SSH-key login;
- installs the pinned CUDA/PyTorch stack and ComfyUI tag;
- downloads the union of the suite's resolved files once;
- executes each workload's five cold/warm pairs and hardware/storage probes;
- fetches and writes one JSON file per workload; and
- destroys an instance it created, including after ordinary setup or benchmark failures.
BENCHMARK_RESULT lines are per-pass progress. BENCHMARK_REJECTED means a cold attempt failed the block-
read threshold and was retried rather than silently accepted. BENCHMARK_SUMMARY is the remote aggregate.
The final local log states each output path and whether the instance was destroyed.
The outer SSH command allows up to four hours so a slow but valid model download can still finish. Each individual workflow sample retains its separate 15-minute timeout; extending the transport ceiling does not permit a hung inference sample to run indefinitely.
Do not constantly poll Vast while a healthy benchmark is running. Watch the runner's lifecycle output and check the provider when a stage exceeds its expected window or after the process exits.
6. Confirm teardown every time#
The finally cleanup handles normal completion and ordinary thrown errors. It cannot guarantee cleanup
after local power loss, kill -9, or every abrupt interruption. After every run, inspect the Vast Instances
page or list your active instances without printing the API key:
bun -e '
const response = await fetch("https://console.vast.ai/api/v0/instances/?owner=me", {
headers: { Authorization: `Bearer ${process.env.VASTAI_API_KEY}` }
});
if (!response.ok) throw new Error(`${response.status}: ${await response.text()}`);
const { instances = [] } = await response.json();
console.table(instances.map(instance => ({
id: instance.id,
status: instance.actual_status,
label: instance.label,
gpu: instance.gpu_name,
hourly_usd: instance.dph_total
})));
'
If a benchmark instance remains, verify its exact numeric ID and destroy it in the Vast console. Do not assume that closing the terminal stopped billing.
7. Validate each result before using it#
Suite output names include workload, GPU slug, physical machine ID, suite timestamp, and the
cache-evicted-v1.json suffix. The runner refuses to overwrite an existing suite result.
Set RESULT to each emitted file and run the structural gate:
RESULT='orchestrator/benchmarks/vast-minimax-h3-rtx-5090-machine-45827-suite-vast-20260923T174517Z-cache-evicted-v1.json'
jq -e '
.status == "success"
and .benchmark_schema_version >= 4
and .cache_mode == "resident"
and (.samples | length) == 5
and (.warm_samples | length) == 5
and all(.samples[];
.phase == "cold_fresh_process"
and .process_block_read_verification.verified == true)
and all(.warm_samples[];
.phase == "warm_resident_different_seed")
' "$RESULT"
Expected output is true with exit code zero. Then print the human review summary:
jq '{
file: input_filename,
schema: .benchmark_schema_version,
workload: .workload.id,
comfyui: .comfyui_tag,
torch: .torch,
machine_id: .vast.advertised_hardware.machine_id,
gpu: .hardware.gpu.product_name,
gpu_power_w: .hardware.gpu.power_limit_w,
gpu_pcie: [
.hardware.gpu.pcie_current_link_gen,
.hardware.gpu.pcie_current_link_width
],
cpu: .hardware.host.cpu_model,
memory_gib: (.hardware.host.cgroup_memory_limit_bytes / 1073741824),
nvme: [.hardware.storage.nvme_devices[]? | {
model,
link_speed: .pcie_current_link_speed,
link_width: .pcie_current_link_width,
upstream_hops: .upstream_pci_hops
}],
md_arrays: .hardware.storage.md_arrays,
cold_mean_s: (.mean_wall_ms / 1000),
warm_mean_s: (.mean_warm_wall_ms / 1000),
cold_penalty_s: ((.mean_wall_ms - .mean_warm_wall_ms) / 1000),
model_reader_mean_gib_s: (
([.model_file_read_benchmark.samples[].throughput_mib_s] | add / length) / 1024
),
h2d_median_gib_s: .accelerator_path_benchmarks.h2d_median_gib_s,
rejected_cold_attempts: (.rejected_cold_attempts | length),
hourly_usd: .vast.hourly_usd,
estimated_model_download_usd: .vast.network_pricing.estimated_prepared_model_download_usd
}' "$RESULT"
Review both suite files. Check that they have the same suite ID, machine ID, GPU identity, ComfyUI tag, and prepared-byte total. Investigate before accepting a result when:
- any cold attempt was rejected;
- actual GPU power or PCIe width differs from the intended class;
- physical storage topology is missing or contradicts the advertised disk;
- host memory is close to exhaustion or swap was used;
- the warm floor is an outlier for the GPU class;
- cold and warm are unexpectedly identical but warm samples show substantial block reads;
- the result came from the same physical machine as an alleged independent replicate.
Use warm throughput for repeated-service Pareto calculations because it removes fresh-process and storage startup from the numerator. This remains the correct observed steady-state throughput when ComfyUI must offload weights, but it is then a host-assisted result rather than a pure GPU-compute result. Keep the workload, cache policy, VRAM class and offload state visible; do not merge resident and offloaded cohorts as if their warm floors measured the same thing. Zero Linux process swap is necessary host-health evidence, but does not prove VRAM residency or the absence of CPU-RAM-to-GPU transfers.
Never hand-edit a completed result. If metadata is wrong, fix the collector and rerun. Put one-off probes
and noncanonical diagnostics under orchestrator/benchmarks/diagnostics/ so the dashboard's root JSON glob
does not treat them as benchmark runs.
8. View the result in the dynamic dashboard#
Start the server from the orchestrator directory:
cd orchestrator
HOST=0.0.0.0 PORT=3011 pnpm benchmark:verified:serve
Open http://127.0.0.1:3011/ locally, /storage-analysis for the RTX 5090 storage-path analysis, or /docs/ for the
rendered benchmark documentation. For a remote hostname, replace 127.0.0.1 with that reachable host; the
server already listens on all interfaces, but the host firewall and overlay network must still permit the
port.
Health and data checks:
curl -fsS http://127.0.0.1:3011/healthz | jq
curl -fsS http://127.0.0.1:3011/api/runs | jq '.count'
At startup the server reads orchestrator/benchmarks/*-cache-evicted-v1.json through the existing DuckDB
normalization query and materializes its output once as the normalized_runs table. All aggregate APIs use
that table rather than independently reparsing the result corpus. After adding or replacing a result,
explicitly rebuild the table without restarting the server:
curl -fsS -X POST http://127.0.0.1:3011/api/normalized-runs/refresh | jq
Refreshes are serialized and API requests continue using the previous table while the replacement is built. The health response reports the last successful refresh, row and source-file counts, duration, and any refresh error. Clicking a run reveals its full details, while the raw JSON remains the source of truth.
For a persistent local session:
tmux new -s benchmark-dashboard
cd orchestrator
HOST=0.0.0.0 PORT=3011 pnpm benchmark:verified:serve
For the persistent K3s service, always name the homelab context explicitly because the default kubectl
context may be production GKE:
kubectl --context k3s apply -f orchestrator/benchmarks/k3s.yaml
kubectl --context k3s -n benchmark-dashboard rollout status \
deployment/benchmark-dashboard --timeout=180s
curl -fsS http://192.168.1.3:3011/healthz | jq
curl -fsS http://192.168.1.3:3011/api/runs | jq '.count'
The deployment mounts this benchmark directory read-only at /app and deliberately runs from writable
/tmp. DuckDB spills large JSON scans into .tmp relative to the working directory; running from /app
causes Failed to create directory ".tmp": Read-only file system and can end in an OOM restart as the
dataset grows. The Bun entry point must therefore remain the absolute /app/verified_server.ts path. The
deployment also caps DuckDB itself at 1 GB inside the 6 GiB pod limit so startup and manual
materialization can spill predictably instead of letting DuckDB size itself from host RAM and get killed by
the cgroup. Materialization can take roughly a minute at the current corpus size; once complete, dashboard
queries read the compact table and do not repeat that work.
9. Preserve the evidence#
Before committing, inspect only the files from your run. This repository often has unrelated work in
progress, so never use a blanket git add .:
git status --short -- \
orchestrator/benchmarks/<minimax-result>.json \
orchestrator/benchmarks/<krea-result>.json
git add -- \
orchestrator/benchmarks/<minimax-result>.json \
orchestrator/benchmarks/<krea-result>.json
git diff --cached --stat
git commit -m 'benchmarks: add RTX 5090 machine <machine-id> suite'
Commit result files without reformatting or trimming their rich metadata. Put interpretation in a separate analysis document or dashboard query so evidence and conclusions remain distinguishable.
Invariants#
- Canonical new results are schema v4, resident-cache, five-pair results.
- One pair is always cold fresh process, then warm same process with a different seed, then process stop.
- Every accepted cold sample must pass the process block-read threshold. Rejected attempts remain recorded.
- Compare machines only within the same workload, source run, ComfyUI tag, PyTorch/CUDA stack, cache policy, and sample protocol.
- Preserve the full device string, power limit, VBIOS/part identity, VRAM, PCIe link, CPU, RAM limit, kernel, storage devices, topology, and container/block path. “RTX 5090” or “9100 Pro” alone is not a configuration.
- Treat a physical machine as the independent unit. Multiple offer or instance IDs can still refer to it.
- Treat advertised
disk_bwas provider metadata, not measured application throughput. - Treat process block reads as page-cache evidence, not automatic proof of physical-media reads through every hidden host layer.
- Keep result JSON immutable and retain raw samples. Derived tables and correlations can change; evidence must not.
- Calculate bandwidth cost before provisioning and verify that no paid instance remains afterward.
- Do not mix legacy
cache-noneresults into the resident-cache cohort.
Interfaces#
Required external interfaces#
| Interface | Purpose | Failure symptom |
|---|---|---|
Vast API via VASTAI_API_KEY |
Offer search, acquisition, instance state, teardown | API or acquisition error |
Read-only production Postgres via PGPASSWORD |
Loads the exact persisted prompt and resolved files | run lookup failed |
| GitHub | Resolves latest ComfyUI only when no tag is supplied | latest-release lookup fails |
| Resolved model URLs | Downloads source-run inputs and models | remote setup/download failure |
| Local SSH key pair | Attaches the public key and controls the machine | ten-minute SSH login timeout |
Important runner flags#
| Flag | Meaning |
|---|---|
--offer-id |
Acquire this exact rentable offer |
--machine-id |
Reacquire an offer on this exact physical machine |
--instance-id |
Reuse a running instance; it remains unless --destroy-after is set |
--gpu |
Exact Vast GPU-name contract |
--min-vram, --max-vram |
Reject the wrong memory variant |
--min-power |
Hard lower bound, not a preference |
--min-reliability, --min-internet |
Hard offer bounds |
--prefer-country |
Ordering preference with global fallback; empty disables preference |
--suite |
Comma-separated built-in workloads; currently minimax-h3,krea-2 |
--samples |
Number of cold/warm pairs |
--cache-mode |
Use resident for current canonical work |
--comfyui-tag |
Exact release for reproducibility |
--output-dir |
Suite output directory; use orchestrator/benchmarks for dashboard data |
--keep-instance |
Keep a newly created instance and continue billing |
--destroy-after |
Destroy a reused --instance-id after the run |
--allow-multi-gpu-host |
Allow shared physical-host CPU/RAM/storage conditions |
--cuda-device |
Select the physical GPU index exposed to ComfyUI |
For a single custom persisted run, use --workflow-id and --run-id. A workflow not known to the runner
also requires --workload-id, --workload-name, and --workload-kind. Adding a workload to a reusable
suite requires adding its fixed source IDs to known_suite_runs in orchestrator/benchmark_vast.ts.
Result and dashboard interfaces#
- Collector:
orchestrator/benchmark_vast.ts - Canonical data:
orchestrator/benchmarks/*-cache-evicted-v1.json - Diagnostic data:
orchestrator/benchmarks/diagnostics/ - Dashboard server:
orchestrator/benchmarks/verified_server.ts - Site navigation:
orchestrator/benchmarks/site_nav.ts, the one list of pages (/,/storage-analysis,/power-cost,/docs/). The server swaps each page's<!-- site-nav -->marker for the bar, and the docs layout renders it too. Add a new page there instead of hand-linking it from another page. HTML edits go live on reload from the hostPath mount; changes to.tsfiles need a pod restart (kubectl --context k3s -n benchmark-dashboard rollout restart deploy/benchmark-dashboard). - Run list/detail API:
/api/runsand/api/runs/:id - Storage analysis API:
/api/storage-analysis - Legacy compatibility aliases:
/raid-analysisand/api/raid-analysis - Economics API:
/api/economics - Source GPU specification catalog:
gpu-device-specs.json - GPU specification API:
/api/gpu-specs - Append-only Ideal Tech source:
idealtech-price-snapshots/*/products.json - Materialized Ideal Tech table:
idealtech_price_observations - Health check:
/healthz
The economics API discovers measured GPU classes directly from the materialized normalized_runs table.
Classes are keyed by normalized GPU identity plus VRAM, so a modified 48 GB RTX 4090 remains separate from a
standard 24 GB card. It first keeps the latest admitted result per physical machine, then chooses a class
representative by zero rejected cold attempts, full-host placement, captured rental price, maximum measured
GPU power, schema version, and capture time. A newly admitted GPU therefore appears in rental economics after
the normal materialization refresh without an allowlist edit.
gpu-pricing.json contains only external purchase facts and exceptional derived options. The server
materializes these as pricing_config, purchase_prices, and configured_purchase_options DuckDB tables at
startup and on /api/normalized-runs/refresh. Every measured class also appears in the purchase table; a class
without a joined MYR price is labelled price required and excluded from the purchase frontier until a quote
is configured.
gpu-device-specs.json is the source-attributed hardware catalog. Its gpu_key values use the same
GPU-plus-VRAM identity as normalized_runs, so specification, benchmark, rental, and purchase data can be
joined without model-name guessing. The server materializes the file as gpu_device_specs at startup and on
the same manual refresh endpoint. /api/gpu-specs returns the raw materialized rows separately from derived
ratios and theoretical per-watt figures. It defaults to the RTX 5090 as the ratio baseline; use, for example,
/api/gpu-specs?baseline=rtx_5080_16gb&minimum_vram_gb=16. Pass include_modified=false to exclude the
non-reference RTX 4090 48 GB profile.
Ideal Tech market observations are separate from gpu-pricing.json. The Next.js Flight extractor writes one
immutable raw response, flat products.json, and metadata file per timestamp. The server materializes all
snapshots as idealtech_price_observations; use WHERE NOT is_label for purchasable components. Retailer
product_name and category_name values remain exact source fields. See
idealtech-price-snapshots.md for capture, overwrite protection, and direct
DuckDB query instructions.
Specification ratios are screening evidence, not workload predictions. NVIDIA's advertised AI TOPS can use different numeric precision and sparsity modes between generations. Compare like-for-like fields where possible, and use the fixed-protocol benchmark for purchase decisions. The modified RTX 4090 profile inherits only AD102 compute-silicon fields; board-specific memory bandwidth, clocks, and power remain null until they are measured on the actual card.
The catalog also covers the RTX PRO Blackwell capacity ladder: RTX PRO 2000 16 GB, 4000 24 GB, 4500 32 GB, 5000 48/72 GB, and 6000 Workstation/Max-Q/Server 96 GB. ECC, MIG capacity, board power, thermal design, and form factor are recorded where NVIDIA publishes them. Treat editions as different boards: for example, the active RTX PRO 4500 Workstation Edition is 200 W with 896 GB/s memory bandwidth, while its passive Server Edition is 165 W with 800 GB/s. Likewise, the 300 W RTX PRO 6000 Max-Q and 600 W Workstation Edition share core and memory capacity but not theoretical or measured throughput.
Footguns#
- Bandwidth can dominate cost. The runner filters minimum download speed, not download price. Exact offer selection is safer for a 61.1 GB suite.
- Latest is not a reproducible version. Omitting
--comfyui-tagresolves the current GitHub release at launch time. Pin the resolved tag across a cohort. - A provider headline is not topology.
disk_bw, disk marketing name, and “20 GB/s” do not establish single drive versus RAID, CPU-direct versus chipset, or the container's real read path. - Warm means same-process model residency.
--cache-nonechanges ComfyUI's caching policy and belongs to a separate experiment; it is not the current baseline. Same-process residency can still use ComfyUI DynamicVRAM or other host-memory offload when the active model does not fit in VRAM, so warm is not automatically a pure GPU-compute measurement. - Cold does not mean power-cycle. It means a fresh ComfyUI process with resolved file pages evicted and verified process block reads.
gpu_fracis easy to misread. On a one-GPU allocation it can encode the fraction of a multi-GPU host, not a MIG-like compute fraction. Shared host resources still make it a different cohort.- An idle PCIe link can downtrain. Interpret current and maximum link fields with the workload context; investigate unexpected width or generation instead of silently relabelling it.
- Vast offers disappear. A specific offer can be rented or deverified between inspection and acquisition. Select the next qualifying offer rather than weakening constraints unnoticed.
- Interruptions need a billing audit. Normal exceptions enter cleanup; abrupt process death might not.
- Output location controls inclusion. Only root-level
*-cache-evicted-v1.jsonfiles feed the dashboard. Keep smoke tests and probes indiagnostics/. - A repeat is not necessarily a replicate. Group by physical
machine_id, not instance or offer ID. - Cold penalty alone can mislead. Inspect absolute cold and warm times. A slow warm floor can make a weak storage path look like it has a small penalty.
- Do not reconstruct missing v4 fields. Old results cannot honestly recover run-time physical-read or resource-trace counters. Rerun when those fields matter.
Testing#
There is no cost-free end-to-end test because provisioning and model download are the system under test. Use these layers:
bun run orchestrator/benchmark_vast.ts --helpchecks that the local TypeScript entry point and CLI dependencies load without renting anything.- The preflight read-only
select 1checks database reachability and credentials. - A one-pair paid diagnostic checks a materially changed collector or a new GPU class. Keep it outside the canonical result glob.
- The five-pair structural
jqgate checks protocol completion and cold-read verification. - The human summary checks hardware identity, actual links, topology, memory, outliers, and cost.
/healthz,/api/runs, and the clicked run-detail view check dashboard ingestion.- The provider instance list checks the billing boundary.
The honest coverage gap is provider opacity: even schema v4 cannot always observe a host cache or storage layer outside the container. That is why the result preserves multiple counters and why cross-host storage claims remain probabilistic rather than causal same-machine A/B claims.
Operations#
Choose the right launch mode#
| Goal | Selector | Teardown behavior |
|---|---|---|
| New independent host | --offer-id after inspecting the offer |
Destroyed by default |
| Repeat known hardware | --machine-id |
Destroyed by default |
| Broad sampling | no selector; use explicit bounds | Destroyed by default; inspect selected price immediately |
| Continue an existing box | --instance-id |
Kept by default; add --destroy-after when finished |
| Manual investigation after suite | --keep-instance |
Kept and billing until explicitly destroyed |
Failure triage#
| Symptom | Likely cause | Action |
|---|---|---|
no rentable Vast offer found |
Constraints currently have no match | Wait or broaden one documented preference; do not silently change the experiment |
| Offer unavailable during acquire | Another renter won the offer | Choose another inspected offer; automatic mode tries its next candidate |
| Instance/SSH timeout | Provider boot or key attachment failure | Let normal cleanup run, then verify the instance list |
| Download is unexpectedly slow or costly | Provider network path or expensive per-GB rate | Stop only if justified, then manually verify teardown; choose a cheaper inspected offer |
| CUDA error 804 during setup | Host driver advertises cuda_max_good < 13 |
Reject the offer; do not change the pinned CUDA/PyTorch cohort |
| Cold attempt rejected | Insufficient process block reads after eviction | Keep the rejection as evidence; the runner retries automatically |
| Remote workload failed | ComfyUI/model/runtime error | Inspect the logged workload and any failed JSON that was written before changing anything |
| Result absent from dashboard | Wrong directory/name, failed status, or unsupported schema | Check the root glob, JSON status/schema, API response, and server log |
| Physical layout fields empty | Host/container hid lower-level devices | Classify the topology as unobserved; do not infer RAID from advertised speed |
When a result looks suspicious#
Start with the paired raw samples and captured metadata, not a cross-GPU comparison:
- compare cold and warm block-read bytes;
- inspect absolute warm time against machines in the same GPU/power class;
- inspect the model reader, buffered fio, direct fio QD1/QD32, and H2D probes;
- inspect NVMe count/model/link, MD arrays, PCI ancestry, filesystem/overlay path, and cgroup evidence;
- inspect CPU, memory limit, swap, kernel, GPU link, and GPU power;
- compare another workload from the same suite;
- reacquire the physical machine if repeatability is the question;
- sample another physical machine if generality is the question.
Document a new interpretation in the experiment journal before changing the dashboard's conclusions. Keep the raw result untouched.
Related#
benchmark_vast.ts— executable source of truth for selection, setup, measurement, result writing, and cleanup.overview.md— local GPU infrastructure and current purchase summary.rtx5090-multivariable-investigation-journal.md— full experiment journal and evolving evidence.rtx5090-single-path-multivariable-analysis.md— current multivariable interpretation of single-path RTX 5090 hosts.model-read-storage-layouts.md— how captured storage layouts map to model read behavior.rtx5090-single-nvme-vs-raid0-dual-workload-experiment.md— the dual-workload storage experiment design and results.storage-device-spec-database.md— official device-spec enrichment used alongside observed results.gpu-theory-price-performance-findings.md— current specification, market-price, workload-performance, and tier-2 findings.idealtech-price-snapshots.md— append-only live component-price capture and DuckDB contract.