GPU plotter for Chia v2 proofs of space (CHIP-48). Produces .plot2 files
byte-identical to the pinned
pos2-chip CPU reference.
This is the main branch, using SYCL/AdaptiveCpp with a CUB fast path
on NVIDIA. The cuda-only branch
provides the native CUDA implementation for NVIDIA.
This is a work in progress. Future changes to the plot format, including grouping, may require replotting.
Quick start · Hardware · Build · Commands · Performance · Documentation
Install the build dependencies first. For containers or Windows, follow INSTALL.md.
cargo install --git https://github.com/Jsewill/xchplot2 --locked
xchplot2 devices
# Replace the key, contract address, and output directory with your values.
xchplot2 plot -k 28 -n 10 \
-f <farmer-pk-hex> \
-c <pool-contract-xch1-or-txch1> \
-o /mnt/plotsEach completed output path is printed to stdout. Check one output with:
xchplot2 verify /mnt/plots/NAME.plot2 --full --trials 100verify --full samples challenges and validates the resulting full proofs;
it does not check every byte. See validation
for CPU-reference comparisons.
| Resource | Requirement or tested scope |
|---|---|
| NVIDIA | Maxwell or newer via CUDA/CUB. Pre-Turing GPUs require a CUDA 12.x build. RTX 4090 is in the current hardware benchmark set. |
| AMD | AdaptiveCpp HIP; RX 6700 XT (gfx1031) is in the current hardware benchmark set. RDNA1 needs the installation-path guidance. |
| Intel | Arc B580 with AdaptiveCpp Level Zero; see the runtime workaround. |
| VRAM | Tiny's base k=28 floor is 1,228 MiB free after context creation, including the default 128 MiB buffer. Backend sort scratch can raise it. |
| Host RAM | Depends on tier and worker count. Lower VRAM tiers generally use more host RAM; see memory requirements. |
| CPU plotting | Opt in with --devices cpu, --devices all, or --cpu; uses pos2-chip's CPU plotter. |
| OS | Linux is tested. WSL2 requires support from the GPU vendor. Native Windows SYCL and macOS are unsupported by this build. |
The benchmark report records tested hardware and toolchains. Build checks and GPU hardware checks are described separately in CONTRIBUTING.md. Physical low-capacity cards are outside the latest benchmark set; a software VRAM cap does not certify another card.
See INSTALL.md for dependencies, containers, Cargo, CMake, architecture selection, and Windows/WSL2. CMake also builds the parity and host test binaries.
| Command | Purpose and guide |
|---|---|
| xchplot2 plot | Create plots from farmer and pool keys |
| xchplot2 batch | Run or resume a saved plot manifest |
| xchplot2 bench | Measure throughput and estimate time to fill storage |
| xchplot2 devices | List GPUs and CPU NUMA nodes |
| xchplot2 verify | Check an existing plot, including full proofs with --full |
| xchplot2 test | Build a test plot from a raw plot ID and memo |
| xchplot2 parity-check | Run the built parity and host tests |
| xchplot2 completions | Generate Bash, zsh, or fish completions |
See configuration and argument files
for reusable options. xchplot2 --help prints command syntax.
Ordinary plotting uses one GPU. Select all GPUs with --devices gpu; add
CPU workers with --cpu, or select both with --devices all. Each GPU
chooses a tier from its own free VRAM. See the
device reference.
Before starting, plot saves an xchplot2-job-*.tsv manifest in the output
directory. It contains private plot keys; keep it private and retain it to
recover the job. Repeat the original plot command with --resume, or
resume directly from the saved manifest:
xchplot2 batch /path/to/job.tsv --resumeResume validates existing files before skipping them. See plotting and recovery for identity matching, manifest selection, progress, stdout, and exit status.
If host RAM is short, use --temp-dir to select a real disk for automatic
spill, or --max-host-ram to set a budget. Available reductions differ by
tier and branch; see host RAM and disk-offload.
Measured September 9, 2026, at k=28, strength=2, using real file writes, FSE compression, and durability barriers. Times are mean completion intervals over ten measured plots after two warmups, using one GPU per host. Each non-auto tier was forced.
Tier (main) |
RTX 4090, CUDA/CUB | RX 6700 XT, AdaptiveCpp HIP | Arc B580, AdaptiveCpp Level Zero |
|---|---|---|---|
| Auto (pool) | 2.50 s | 9.63 s | 13.75 s |
| Plain | 3.91 s; 2.66 s repeat | 9.76 s | 14.17 s |
| Compact | 3.84 s | 10.51 s | — |
| Minimal | 16.40 s | 22.19 s | — |
| Tiny | 27.74 s | 29.77 s | — |
| Pinned | 27.42 s | 29.73 s | — |
— means not benchmarked. See the benchmark configurations.
The two RTX 4090 Plain timings differ for an undetermined reason.
The native cuda-only auto path measured 2.20 s/plot on the same RTX
4090, with its optional D2H/Xs overlap enabled. These runs do not isolate the
cause of the difference between the native and SYCL runtimes.
The benchmarks include configurations, variability, and memory use. Forced tiers can use additional match scratch on these roomy GPUs; their timings do not predict performance on a card restricted to a tier's minimum VRAM.
Multi-GPU throughput also depends on shared PCIe bandwidth, CPU compression, and storage. This benchmark set uses one GPU per host.
| Guide | Contents |
|---|---|
| Installation | Dependencies, containers, Cargo/CMake, Windows and WSL2 |
| Command reference | All commands, configuration, devices, memory, environment variables, troubleshooting |
| Benchmark results | Dated measurements, methodology, and memory use |
| Contributing | Architecture, local tests, CI, and contribution conventions |
| Security | Private key and manifest handling; vulnerability reporting |
MIT — see LICENSE and NOTICE for third-party attributions. Built collaboratively with Claude.
If you appreciate this, and want to give back, feel free.
xch1d80tfje65xy97fpxg7kl89wugnd6svlv5uag2qays0um5ay5sn0qz8vph8