Vast benchmark operator handbook#

Status: Reference Last verified: 2026-09-28 Canonical for: launching, validating, preserving, and publishing ComfyICU GPU benchmarks on Vast.ai

Scope#

This is the start-to-finish runbook for adding a machine to the current benchmark dataset. It covers selecting a Vast offer, running the fixed cold/warm protocol, checking the resulting schema-v4 JSON, confirming that the rented instance was destroyed, and viewing the result in the dynamic dashboard.

The normal deliverable is one immutable JSON result per workload under orchestrator/benchmarks/. A canonical full-suite launch produces separate MiniMax H3 and Krea 2 files from one provisioned machine and one software setup.

This handbook does not decide which GPU or storage layout to buy. The analysis documents linked under Related interpret the evidence. It also does not cover production GPU provisioning, which lives in orchestrator/docs/gpu-operations.md.

Model#

The units being measured#

The canonical protocol#

New comparative results use schema v4, --cache-mode resident, and five pairs:

prepare all suite files once

repeat 5 times per workload:
  evict every resolved model/input file with fsync + POSIX_FADV_DONTNEED
  start a fresh ComfyUI process
  cold run
  require process block reads >= 80% of resolved bytes
  warm run in the same process with a different seed
  stop ComfyUI, clearing resident model state

capture model-reader, fio, CPU-memory, H2D/D2H, block-layer, PCIe, CPU,
RAM, GPU, storage-topology, software-version, and pricing evidence

This separates three useful quantities:

The cold-read verifier proves that the ComfyUI process incurred block I/O after page eviction. On a nested container/loop/device-mapper path, it does not by itself prove that every byte reached physical flash. Use the captured cgroup, block-layer, NVMe, filesystem, and model-reader evidence together.

Software baseline#

The runner currently installs:

The current RTX 5090 storage cohort uses ComfyUI v0.37.0. Use that tag when adding a directly comparable host. When starting a new cohort, resolve the latest release once, record it, and then pass that exact tag to every run; do not let “latest” move between machines:

COMFYUI_TAG="$(curl -fsS \
  -H 'Accept: application/vnd.github+json' \
  https://api.github.com/repos/Comfy-Org/ComfyUI/releases/latest | jq -r .tag_name)"
test -n "$COMFYUI_TAG" && test "$COMFYUI_TAG" != null
printf 'Pinned ComfyUI tag: %s\n' "$COMFYUI_TAG"

How it works#

Golden path#

For a normal new-machine sample:

  1. Load .env and pass the local preflight.
  2. Freeze the experiment's ComfyUI tag and five-pair resident-cache protocol.
  3. Inspect a verified Vast offer, especially its download price, host RAM, GPU power/fraction, and machine ID.
  4. Launch minimax-h3,krea-2 by exact offer ID inside tmux.
  5. Let the runner provision, set up once, run both workloads, write two JSON files, and tear down.
  6. Confirm that Vast has no leftover benchmark instance.
  7. Run the structural gate and human metadata summary on both JSON files.
  8. Open the dashboard, inspect both run-detail views, and commit only the two immutable result files.

1. Prepare the local operator environment#

Run benchmark commands from the comfyicu repository root. The repository dependencies must already be installed. You need:

Load the secrets without printing them, then run the preflight:

set -a
source .env
set +a

for command_name in bun pnpm psql ssh curl jq; do
  command -v "$command_name" >/dev/null || echo "missing: $command_name"
done

test -n "$VASTAI_API_KEY" && echo "Vast API key: present"
test -n "$PGPASSWORD" && echo "read-only DB password: present"
test -f ~/.ssh/id_ed25519 && test -f ~/.ssh/id_ed25519.pub && echo "SSH key pair: present"

psql -X -qAt \
  -h gke-pgbouncer.tail6bb0f.ts.net \
  -U warehouse_ro \
  -d comfyicu \
  -c 'select 1'

bun run orchestrator/benchmark_vast.ts --help >/dev/null

Expected database output is 1. The benchmark lookup is read-only; never substitute a write-capable production account.

2. Define the comparison before renting anything#

Write down the question and the controlled dimensions. At minimum, freeze:

Prefer a full suite when the machine is new. It amortizes the expensive environment and model downloads while producing two independent workload result files. A one-sample run is only a paid diagnostic and must go under orchestrator/benchmarks/diagnostics/; it is not evidence for the main dashboard.

3. Inspect and choose a Vast offer#

For a first run, inspect offers in the Vast console and pass the exact offer ID. That prevents automatic selection from surprising you with expensive bandwidth. Use these checks:

Check Canonical default Why
Verification Verified Reduces unexplained host risk
Allocation Exactly one GPU The runner executes on one GPU
Physical host Prefer gpu_frac = 1 Avoids shared CPU, RAM, and storage for baseline runs
Host RAM 64 GB preferred; 48 GB hard minimum MiniMax cold loading has approached the minimum
Container disk 120 GB Holds the environment, both workloads, probes, and output
Reliability At least 0.97; 0.98 preferred Avoids wasting the large download
Download rate Inspect price per GB, not only Mbps Egress can cost more than GPU time
CUDA compatibility cuda_max_good >= 13 The pinned CUDA 13 container cannot initialize on older host drivers
GPU power Prefer the target class; record rather than hide variants Power affects the warm compute floor
GPU PCIe Prefer full width for the target setup Host-to-device loading can be link-sensitive
Disk headline Treat only as a search hint Vast's number is not the application read path

The built-in suite currently prepares 61,108,590,469 bytes, or about 61.1 decimal GB. Before renting, estimate model traffic as:

download cost = advertised download USD/GB × 61.1

At $0.04/GB, models alone cost about $2.44, before CUDA, Python, ComfyUI, or package downloads. Prefer a cheap download rate even if the GPU hourly price is slightly higher. Region is a preference, not a hard requirement; use --prefer-country '' when searching globally.

Offer ID and machine ID are different:

The default rejects a GPU allocated from a multi-GPU physical host. --allow-multi-gpu-host is appropriate only when shared-host behavior is part of the experiment. A reported fraction such as 0.25 can mean one whole allocated GPU on a four-GPU host; the main concern is shared host resources, not necessarily a quarter of the GPU's compute.

4. Launch the canonical full suite#

Use a durable terminal so an SSH disconnect from your workstation does not kill the local runner:

tmux new -s vast-benchmark

Inside that session, load .env, set the offer and cohort constants, and launch:

set -a
source .env
set +a

OFFER_ID=12345678
GPU_NAME='RTX 5090'
COMFYUI_TAG='v0.37.0'

bun run orchestrator/benchmark_vast.ts \
  --offer-id "$OFFER_ID" \
  --gpu "$GPU_NAME" \
  --min-vram 30 \
  --max-vram 34 \
  --min-reliability 0.97 \
  --min-internet 0 \
  --prefer-country '' \
  --disk 120 \
  --samples 5 \
  --cache-mode resident \
  --comfyui-tag "$COMFYUI_TAG" \
  --suite minimax-h3,krea-2 \
  --output-dir orchestrator/benchmarks

Replace the example offer ID. The GPU string must exactly match Vast's gpu_name; change the VRAM bounds for other GPU classes. Add --min-power only when power is an admission requirement, not merely a preference.

To reacquire a known physical host, replace --offer-id with:

--machine-id 45827

To let the runner search globally, omit both selectors. It queries verified, rentable, one-GPU offers, applies the supplied constraints, and normally requires gpu_frac = 1.

To reuse an already-running instance, replace the selector with:

--instance-id 52270704 --destroy-after

Without --destroy-after, a reused instance remains running. A newly provisioned instance is destroyed by default. --keep-instance overrides that default and should be rare.

Detach from tmux with Ctrl-b d and return with:

tmux attach -t vast-benchmark

5. Understand the lifecycle output#

The runner:

  1. reads the source run and resolved files through warehouse_ro;
  2. selects and acquires the offer, then logs the Vast instance ID;
  3. waits up to ten minutes for the instance and another ten minutes for SSH-key login;
  4. installs the pinned CUDA/PyTorch stack and ComfyUI tag;
  5. downloads the union of the suite's resolved files once;
  6. executes each workload's five cold/warm pairs and hardware/storage probes;
  7. fetches and writes one JSON file per workload; and
  8. destroys an instance it created, including after ordinary setup or benchmark failures.

BENCHMARK_RESULT lines are per-pass progress. BENCHMARK_REJECTED means a cold attempt failed the block- read threshold and was retried rather than silently accepted. BENCHMARK_SUMMARY is the remote aggregate. The final local log states each output path and whether the instance was destroyed.

The outer SSH command allows up to four hours so a slow but valid model download can still finish. Each individual workflow sample retains its separate 15-minute timeout; extending the transport ceiling does not permit a hung inference sample to run indefinitely.

Do not constantly poll Vast while a healthy benchmark is running. Watch the runner's lifecycle output and check the provider when a stage exceeds its expected window or after the process exits.

6. Confirm teardown every time#

The finally cleanup handles normal completion and ordinary thrown errors. It cannot guarantee cleanup after local power loss, kill -9, or every abrupt interruption. After every run, inspect the Vast Instances page or list your active instances without printing the API key:

bun -e '
const response = await fetch("https://console.vast.ai/api/v0/instances/?owner=me", {
  headers: { Authorization: `Bearer ${process.env.VASTAI_API_KEY}` }
});
if (!response.ok) throw new Error(`${response.status}: ${await response.text()}`);
const { instances = [] } = await response.json();
console.table(instances.map(instance => ({
  id: instance.id,
  status: instance.actual_status,
  label: instance.label,
  gpu: instance.gpu_name,
  hourly_usd: instance.dph_total
})));
'

If a benchmark instance remains, verify its exact numeric ID and destroy it in the Vast console. Do not assume that closing the terminal stopped billing.

7. Validate each result before using it#

Suite output names include workload, GPU slug, physical machine ID, suite timestamp, and the cache-evicted-v1.json suffix. The runner refuses to overwrite an existing suite result.

Set RESULT to each emitted file and run the structural gate:

RESULT='orchestrator/benchmarks/vast-minimax-h3-rtx-5090-machine-45827-suite-vast-20260923T174517Z-cache-evicted-v1.json'

jq -e '
  .status == "success"
  and .benchmark_schema_version >= 4
  and .cache_mode == "resident"
  and (.samples | length) == 5
  and (.warm_samples | length) == 5
  and all(.samples[];
    .phase == "cold_fresh_process"
    and .process_block_read_verification.verified == true)
  and all(.warm_samples[];
    .phase == "warm_resident_different_seed")
' "$RESULT"

Expected output is true with exit code zero. Then print the human review summary:

jq '{
  file: input_filename,
  schema: .benchmark_schema_version,
  workload: .workload.id,
  comfyui: .comfyui_tag,
  torch: .torch,
  machine_id: .vast.advertised_hardware.machine_id,
  gpu: .hardware.gpu.product_name,
  gpu_power_w: .hardware.gpu.power_limit_w,
  gpu_pcie: [
    .hardware.gpu.pcie_current_link_gen,
    .hardware.gpu.pcie_current_link_width
  ],
  cpu: .hardware.host.cpu_model,
  memory_gib: (.hardware.host.cgroup_memory_limit_bytes / 1073741824),
  nvme: [.hardware.storage.nvme_devices[]? | {
    model,
    link_speed: .pcie_current_link_speed,
    link_width: .pcie_current_link_width,
    upstream_hops: .upstream_pci_hops
  }],
  md_arrays: .hardware.storage.md_arrays,
  cold_mean_s: (.mean_wall_ms / 1000),
  warm_mean_s: (.mean_warm_wall_ms / 1000),
  cold_penalty_s: ((.mean_wall_ms - .mean_warm_wall_ms) / 1000),
  model_reader_mean_gib_s: (
    ([.model_file_read_benchmark.samples[].throughput_mib_s] | add / length) / 1024
  ),
  h2d_median_gib_s: .accelerator_path_benchmarks.h2d_median_gib_s,
  rejected_cold_attempts: (.rejected_cold_attempts | length),
  hourly_usd: .vast.hourly_usd,
  estimated_model_download_usd: .vast.network_pricing.estimated_prepared_model_download_usd
}' "$RESULT"

Review both suite files. Check that they have the same suite ID, machine ID, GPU identity, ComfyUI tag, and prepared-byte total. Investigate before accepting a result when:

Use warm throughput for repeated-service Pareto calculations because it removes fresh-process and storage startup from the numerator. This remains the correct observed steady-state throughput when ComfyUI must offload weights, but it is then a host-assisted result rather than a pure GPU-compute result. Keep the workload, cache policy, VRAM class and offload state visible; do not merge resident and offloaded cohorts as if their warm floors measured the same thing. Zero Linux process swap is necessary host-health evidence, but does not prove VRAM residency or the absence of CPU-RAM-to-GPU transfers.

Never hand-edit a completed result. If metadata is wrong, fix the collector and rerun. Put one-off probes and noncanonical diagnostics under orchestrator/benchmarks/diagnostics/ so the dashboard's root JSON glob does not treat them as benchmark runs.

8. View the result in the dynamic dashboard#

Start the server from the orchestrator directory:

cd orchestrator
HOST=0.0.0.0 PORT=3011 pnpm benchmark:verified:serve

Open http://127.0.0.1:3011/ locally, /storage-analysis for the RTX 5090 storage-path analysis, or /docs/ for the rendered benchmark documentation. For a remote hostname, replace 127.0.0.1 with that reachable host; the server already listens on all interfaces, but the host firewall and overlay network must still permit the port.

Health and data checks:

curl -fsS http://127.0.0.1:3011/healthz | jq
curl -fsS http://127.0.0.1:3011/api/runs | jq '.count'

At startup the server reads orchestrator/benchmarks/*-cache-evicted-v1.json through the existing DuckDB normalization query and materializes its output once as the normalized_runs table. All aggregate APIs use that table rather than independently reparsing the result corpus. After adding or replacing a result, explicitly rebuild the table without restarting the server:

curl -fsS -X POST http://127.0.0.1:3011/api/normalized-runs/refresh | jq

Refreshes are serialized and API requests continue using the previous table while the replacement is built. The health response reports the last successful refresh, row and source-file counts, duration, and any refresh error. Clicking a run reveals its full details, while the raw JSON remains the source of truth.

For a persistent local session:

tmux new -s benchmark-dashboard
cd orchestrator
HOST=0.0.0.0 PORT=3011 pnpm benchmark:verified:serve

For the persistent K3s service, always name the homelab context explicitly because the default kubectl context may be production GKE:

kubectl --context k3s apply -f orchestrator/benchmarks/k3s.yaml
kubectl --context k3s -n benchmark-dashboard rollout status \
  deployment/benchmark-dashboard --timeout=180s
curl -fsS http://192.168.1.3:3011/healthz | jq
curl -fsS http://192.168.1.3:3011/api/runs | jq '.count'

The deployment mounts this benchmark directory read-only at /app and deliberately runs from writable /tmp. DuckDB spills large JSON scans into .tmp relative to the working directory; running from /app causes Failed to create directory ".tmp": Read-only file system and can end in an OOM restart as the dataset grows. The Bun entry point must therefore remain the absolute /app/verified_server.ts path. The deployment also caps DuckDB itself at 1 GB inside the 6 GiB pod limit so startup and manual materialization can spill predictably instead of letting DuckDB size itself from host RAM and get killed by the cgroup. Materialization can take roughly a minute at the current corpus size; once complete, dashboard queries read the compact table and do not repeat that work.

9. Preserve the evidence#

Before committing, inspect only the files from your run. This repository often has unrelated work in progress, so never use a blanket git add .:

git status --short -- \
  orchestrator/benchmarks/<minimax-result>.json \
  orchestrator/benchmarks/<krea-result>.json

git add -- \
  orchestrator/benchmarks/<minimax-result>.json \
  orchestrator/benchmarks/<krea-result>.json

git diff --cached --stat
git commit -m 'benchmarks: add RTX 5090 machine <machine-id> suite'

Commit result files without reformatting or trimming their rich metadata. Put interpretation in a separate analysis document or dashboard query so evidence and conclusions remain distinguishable.

Invariants#

Interfaces#

Required external interfaces#

Interface Purpose Failure symptom
Vast API via VASTAI_API_KEY Offer search, acquisition, instance state, teardown API or acquisition error
Read-only production Postgres via PGPASSWORD Loads the exact persisted prompt and resolved files run lookup failed
GitHub Resolves latest ComfyUI only when no tag is supplied latest-release lookup fails
Resolved model URLs Downloads source-run inputs and models remote setup/download failure
Local SSH key pair Attaches the public key and controls the machine ten-minute SSH login timeout

Important runner flags#

Flag Meaning
--offer-id Acquire this exact rentable offer
--machine-id Reacquire an offer on this exact physical machine
--instance-id Reuse a running instance; it remains unless --destroy-after is set
--gpu Exact Vast GPU-name contract
--min-vram, --max-vram Reject the wrong memory variant
--min-power Hard lower bound, not a preference
--min-reliability, --min-internet Hard offer bounds
--prefer-country Ordering preference with global fallback; empty disables preference
--suite Comma-separated built-in workloads; currently minimax-h3,krea-2
--samples Number of cold/warm pairs
--cache-mode Use resident for current canonical work
--comfyui-tag Exact release for reproducibility
--output-dir Suite output directory; use orchestrator/benchmarks for dashboard data
--keep-instance Keep a newly created instance and continue billing
--destroy-after Destroy a reused --instance-id after the run
--allow-multi-gpu-host Allow shared physical-host CPU/RAM/storage conditions
--cuda-device Select the physical GPU index exposed to ComfyUI

For a single custom persisted run, use --workflow-id and --run-id. A workflow not known to the runner also requires --workload-id, --workload-name, and --workload-kind. Adding a workload to a reusable suite requires adding its fixed source IDs to known_suite_runs in orchestrator/benchmark_vast.ts.

Result and dashboard interfaces#

The economics API discovers measured GPU classes directly from the materialized normalized_runs table. Classes are keyed by normalized GPU identity plus VRAM, so a modified 48 GB RTX 4090 remains separate from a standard 24 GB card. It first keeps the latest admitted result per physical machine, then chooses a class representative by zero rejected cold attempts, full-host placement, captured rental price, maximum measured GPU power, schema version, and capture time. A newly admitted GPU therefore appears in rental economics after the normal materialization refresh without an allowlist edit.

gpu-pricing.json contains only external purchase facts and exceptional derived options. The server materializes these as pricing_config, purchase_prices, and configured_purchase_options DuckDB tables at startup and on /api/normalized-runs/refresh. Every measured class also appears in the purchase table; a class without a joined MYR price is labelled price required and excluded from the purchase frontier until a quote is configured.

gpu-device-specs.json is the source-attributed hardware catalog. Its gpu_key values use the same GPU-plus-VRAM identity as normalized_runs, so specification, benchmark, rental, and purchase data can be joined without model-name guessing. The server materializes the file as gpu_device_specs at startup and on the same manual refresh endpoint. /api/gpu-specs returns the raw materialized rows separately from derived ratios and theoretical per-watt figures. It defaults to the RTX 5090 as the ratio baseline; use, for example, /api/gpu-specs?baseline=rtx_5080_16gb&minimum_vram_gb=16. Pass include_modified=false to exclude the non-reference RTX 4090 48 GB profile.

Ideal Tech market observations are separate from gpu-pricing.json. The Next.js Flight extractor writes one immutable raw response, flat products.json, and metadata file per timestamp. The server materializes all snapshots as idealtech_price_observations; use WHERE NOT is_label for purchasable components. Retailer product_name and category_name values remain exact source fields. See idealtech-price-snapshots.md for capture, overwrite protection, and direct DuckDB query instructions.

Specification ratios are screening evidence, not workload predictions. NVIDIA's advertised AI TOPS can use different numeric precision and sparsity modes between generations. Compare like-for-like fields where possible, and use the fixed-protocol benchmark for purchase decisions. The modified RTX 4090 profile inherits only AD102 compute-silicon fields; board-specific memory bandwidth, clocks, and power remain null until they are measured on the actual card.

The catalog also covers the RTX PRO Blackwell capacity ladder: RTX PRO 2000 16 GB, 4000 24 GB, 4500 32 GB, 5000 48/72 GB, and 6000 Workstation/Max-Q/Server 96 GB. ECC, MIG capacity, board power, thermal design, and form factor are recorded where NVIDIA publishes them. Treat editions as different boards: for example, the active RTX PRO 4500 Workstation Edition is 200 W with 896 GB/s memory bandwidth, while its passive Server Edition is 165 W with 800 GB/s. Likewise, the 300 W RTX PRO 6000 Max-Q and 600 W Workstation Edition share core and memory capacity but not theoretical or measured throughput.

Footguns#

Testing#

There is no cost-free end-to-end test because provisioning and model download are the system under test. Use these layers:

  1. bun run orchestrator/benchmark_vast.ts --help checks that the local TypeScript entry point and CLI dependencies load without renting anything.
  2. The preflight read-only select 1 checks database reachability and credentials.
  3. A one-pair paid diagnostic checks a materially changed collector or a new GPU class. Keep it outside the canonical result glob.
  4. The five-pair structural jq gate checks protocol completion and cold-read verification.
  5. The human summary checks hardware identity, actual links, topology, memory, outliers, and cost.
  6. /healthz, /api/runs, and the clicked run-detail view check dashboard ingestion.
  7. The provider instance list checks the billing boundary.

The honest coverage gap is provider opacity: even schema v4 cannot always observe a host cache or storage layer outside the container. That is why the result preserves multiple counters and why cross-host storage claims remain probabilistic rather than causal same-machine A/B claims.

Operations#

Choose the right launch mode#

Goal Selector Teardown behavior
New independent host --offer-id after inspecting the offer Destroyed by default
Repeat known hardware --machine-id Destroyed by default
Broad sampling no selector; use explicit bounds Destroyed by default; inspect selected price immediately
Continue an existing box --instance-id Kept by default; add --destroy-after when finished
Manual investigation after suite --keep-instance Kept and billing until explicitly destroyed

Failure triage#

Symptom Likely cause Action
no rentable Vast offer found Constraints currently have no match Wait or broaden one documented preference; do not silently change the experiment
Offer unavailable during acquire Another renter won the offer Choose another inspected offer; automatic mode tries its next candidate
Instance/SSH timeout Provider boot or key attachment failure Let normal cleanup run, then verify the instance list
Download is unexpectedly slow or costly Provider network path or expensive per-GB rate Stop only if justified, then manually verify teardown; choose a cheaper inspected offer
CUDA error 804 during setup Host driver advertises cuda_max_good < 13 Reject the offer; do not change the pinned CUDA/PyTorch cohort
Cold attempt rejected Insufficient process block reads after eviction Keep the rejection as evidence; the runner retries automatically
Remote workload failed ComfyUI/model/runtime error Inspect the logged workload and any failed JSON that was written before changing anything
Result absent from dashboard Wrong directory/name, failed status, or unsupported schema Check the root glob, JSON status/schema, API response, and server log
Physical layout fields empty Host/container hid lower-level devices Classify the topology as unobserved; do not infer RAID from advertised speed

When a result looks suspicious#

Start with the paired raw samples and captured metadata, not a cross-GPU comparison:

  1. compare cold and warm block-read bytes;
  2. inspect absolute warm time against machines in the same GPU/power class;
  3. inspect the model reader, buffered fio, direct fio QD1/QD32, and H2D probes;
  4. inspect NVMe count/model/link, MD arrays, PCI ancestry, filesystem/overlay path, and cgroup evidence;
  5. inspect CPU, memory limit, swap, kernel, GPU link, and GPU power;
  6. compare another workload from the same suite;
  7. reacquire the physical machine if repeatability is the question;
  8. sample another physical machine if generality is the question.

Document a new interpretation in the experiment journal before changing the dashboard's conclusions. Keep the raw result untouched.