C-Zone price snapshots#
Status: Working reference
Last verified: 2026-09-29
Canonical for: capturing and querying append-only C-Zone DIY price-list data
Scope and design#
C-Zone publishes its current DIY price list as a two-page Excel-generated PDF at a stable Google Drive
file ID. The page contains several tables side-by-side, so ordinary pdftotext output interleaves unrelated
products. extract_czone_prices.ts uses Poppler's word bounding boxes and explicit C-Zone panel/column rules
instead.
This is intentionally a C-Zone extractor, not a universal PDF parser. It captures the component classes needed for complete-node comparisons:
- CPU
- motherboard
- GPU
- performance RAM kits
- NVMe SSD
- PSU
- air and AIO cooler
- case
It does not capture monitors, networking, audio, peripherals, software, portable storage, SATA SSDs or hard disks. Another retailer or another PDF layout needs its own small extraction profile. Sentinel products and category checks make layout drift fail before a snapshot is written; do not weaken those checks to make a changed PDF appear to work.
Capture the current list#
From orchestrator/benchmarks/:
bun run prices:czone
The command downloads the canonical Drive file, verifies that it is a PDF, computes its SHA-256, extracts and validates the supported tables, then writes:
czone-price-snapshots/<UTC timestamp>/
metadata.json
products.json
source.pdf
Snapshots are append-only. If the downloaded PDF has the same SHA-256 as an existing snapshot, the command reports that it is unchanged and does not create a duplicate directory.
An already downloaded PDF can be imported with an explicit observation time:
bun run prices:czone -- \
--input /path/to/C-ZoneDIY.pdf \
--captured-at 2026-09-29T08:28:22Z
Poppler's pdftotext and pdfinfo executables must be installed. Run the parser tests with:
bun run test:prices:czone
Data contract#
products.json is a flat JSON array. Every row carries the capture time, canonical source URL, source PDF
hash and the PDF's own effective-date label, followed by:
category_idandcategory_name- deterministic per-snapshot
product_id - exact extracted
product_name - typed
current_price_myr - source page, source panel and reconstructed
raw_line - category-specific attributes such as SSD read/write speed, RAM kit capacity or PSU wattage
The exact category_name + product_name pair is the catalog identity. Build candidates must use it exactly;
fuzzy matching is forbidden because a plausible near-match can silently price the wrong capacity or SKU.
The current extractor retains 674 build-relevant products from the list effective 28 SEPT 2026:
| Category | Products |
|---|---|
| CPU | 33 |
| Motherboard | 129 |
| GPU | 92 |
| RAM | 25 |
| SSD | 42 |
| PSU | 104 |
| Cooler | 123 |
| Case | 126 |
Its source PDF SHA-256 is
af3b0c96def81053d8b2fa5b6a8080e3c05f864d7a93a35fd454b33ab6c5e782.
Extractor version 1 retained 627 products because 47 AIO rows used slightly different word coordinates for their lighting column. That snapshot remains immutable evidence. Version 2 widened only that bounded panel, added an exact Arctic AIO sentinel, and re-extracted the same source hash as a new snapshot. Materialization uses products from the newest complete snapshot rather than carrying absent identities forward from older extractor versions.
DuckDB and build comparison#
The benchmark server materializes all snapshots into czone_price_observations at startup and on the manual
POST /api/normalized-runs/refresh endpoint. The health response exposes capture/product counters under
czone_prices.
latest_czone_prices selects the newest exact product identity. A component in build-candidates.json uses
it like this:
{
"product_name": "SAMSUNG 9100 PRO 2TB (GEN 5.0)",
"price_source": "czone_latest",
"catalog_vendor": "czone",
"category_name": "SSD"
}
The materialized build refresh fails if that exact product cannot be resolved. This is the intended signal to inspect a catalog change and update the BOM explicitly.