Benchmarks
Benchmark methodology
How the StateSync-GKR measurements were controlled, which statistics and timer boundaries apply, which hardware ran them, how to replay the reproducible parts and what the measurements do not establish.
This page explains how the benchmark results were produced. Read it to judge whether a result applies to your workload, to compare your own measurements fairly and to replay the parts that can be replayed.
Measured subject#
Every study measured the same immutable kernel: the historical 1.1.0 source, commit 1b8d2f829792172b347dedbfde969016c2c05789. That commit is identified in the papers' artifact records; the current public repository history begins after it. The benchmark callers and observers have their own source and binary identities, and no study edits the production source. The memory observer runs as a test child inside an isolated copy of the source.
The results describe that source on the recorded hosts. Treat a different source revision, compiler, configuration or host as a new subject and measure it yourself.
Controlled comparisons#
Each ratio changes one thing inside StateSync-GKR and holds the rest fixed.
| Comparison | What changes | What stays fixed |
|---|---|---|
| Memory | Prover oracle: test-only Dense reference with six quadratic factor tables against the production Libra-style two-phase sparse oracle | Circuit, field, transcript order, prover driver and fixture; proof observations are checked for equality |
| Prepared verification | Wiring oracle: TableWiring against DerivedRegularWiring | Circuit, request, proof, commitment and verifier body; call order is randomized |
| Preparation reuse | Fresh compilation and commitment for every request against reused preparation | Operation, depth and configuration |
| Output policy | Serial against parallel encoding and encoded verification | Parallel witness generation and proving, batch size, worker count and fixtures; proof bytes are checked for equality |
StateSync-GKR implements its own sumcheck, GKR assembly and sparse reduction. Plonky3 0.4.3 supplies the field arithmetic, Poseidon2 hash and challenger primitives to both sides of every comparison. Plonky3 is a component provider in these studies, not a comparison baseline.
Study units and sample sizes#
| Study | Execution units | Timed and warmup observations |
|---|---|---|
| Earlier CPU campaign | 81 conditions, three processes each: 243 | 63,540 timed samples |
| CPU supplement | 35 conditions, five processes each: 175 | 156,100 timed, 3,925 warmups |
| Generic memory | Four circuit families, eleven widths from 4 to 4,096, two oracle modes, three processes: 264 | 792 timed, 264 warmups |
| Memory limits and concurrency | 17 attempts: 14 successes, one address-space abort, two out-of-memory kills | 38 successful concurrent jobs; 12 timed and 4 warmup scale proofs |
| 48-core arrivals | 36 screens, one pilot, one confirmation | 730,000 timed, 5,888 warmups |
| Equal resources | 12 frontend and 18 worker processes | 180,000 timed, 3,072 warmups |
| 48-core fine study | 19 conditions, three process pairs each: 57 | 456,750 timed, 14,592 warmups |
| 192-core VM | 47 frontend and worker process pairs | 1,102,000 timed, 12,032 warmups |
These counts are different experimental units and are never added into one sample population. Process repeats share a host and a fixed corpus, so they measure run-to-run variation on one machine. A request cycled from the fixed corpus counts as a repeat of its fixture. Different hosts, profiles and confirmation periods are kept separate.
In the 192-core study, the 1,114,032 offers equal 1,079,666 submitted and accepted requests plus 34,366 explicit pre-send frontend_max_pending outcomes. Submitted requests had no server rejection, loss, duplicate or unresolved response.
Statistics#
Each result is summarized per process first, and process distributions are never pooled into one sample.
| Result | Summary |
|---|---|
| Memory | Median of three timed observations per process; the reported value is the median of the three process medians |
| CPU throughput | Median of three process-level rates; each rate uses the median of three measured batch intervals |
| Complete local batch throughput | Each process rate is the batch size divided by the median of ten complete-pipeline intervals after one warmup batch; the reported value is the median of five process rates |
| Prepared verification ratio | Each process summarizes its 300 same-proof paired ratios by their median; the reported value is the median of five process summaries |
| Output policy ratio | Each process summarizes ten paired serial/parallel interval ratios by their median; the range spans the three operation medians across five processes |
| Preparation reuse ratio | Ratios of run-matched process summaries, not of individual requests |
The studies keep their original quantile conventions rather than unifying them after the fact:
| Study | Quantile convention |
|---|---|
| Arrival and serving studies | Observed nearest rank at sorted index ceil(q*N)-1, integer nanoseconds, no interpolation |
| Earlier CPU campaign | Sorted index floor((N-1)*q) |
| CPU supplement | Type-7 linear interpolation at (N-1)*q; p50 equals the ordinary sample median |
Ranges across processes describe the observed spread and are not confidence intervals. No outlier is removed, and no interpolated operating point is reported as measured. In the supplement, p99 from 300 or 128 requests is a descriptive sample statistic only.
For serving studies, success fractions divide by all timed offers, including refusals and pre-send failures, and accepted-only latency distributions are reported separately. A completion-window rate includes the drain after input stops. Dividing an accepted count by the scheduled duration can hide queue growth, so it is not used as a capacity measure.
Timer boundaries#
| Result | Inside the timer | Outside the timer |
|---|---|---|
Requested allocation and Vec capacity | The prover interval, with allocator counters and stage observation enabled and not subtracted | Allocator metadata and slack, stacks and transient reallocation overlap |
| Whole-process RSS | The entire process, including fixtures, logging and the JSON, hashing and verification that follow the prover timer | Nothing is subtracted |
| Prepared verification | The complete verify_sync_op_with call, including request validation and native leaf and path hashing | Circuit, wiring and commitment preparation, proof generation, encoding and transport |
| Preparation reuse | The per-request proving path | Encoding, transport and acceptance |
| Complete local batch | Witness generation, proving, encoding and encoded verification | Preparation, transport and post-timer hashing and logging |
| CPU throughput | Proving, plus witness creation in the witness columns | Encoding, hash and control checks, and delivery |
| Serving latency | From each scheduled arrival to strict encoded-proof acceptance at the frontend, including queueing and transport | Database commits, root adoption, consensus and settlement |
Serving latency is measured from the intended arrival time, so a slow service cannot lower the offered demand by delaying its replies. The frontend uses its own monotonic clock, and original offered timestamps are preserved. Server queue and compute times use the server's clock. Clocks are never subtracted across hosts to produce a one-way network latency.
Batch timers are shared by the jobs in a batch and counted once. CPU usage comes from tick counts at HZ 100 over actual telemetry intervals; missing telemetry stays missing. GNU time peaks and sampled /proc RSS are reported separately when they differ. Logging stays inside the measured system and is not subtracted to claim a higher throughput.
Hardware and software#
| Study | Hosts |
|---|---|
| Earlier CPU campaign and supplement | Separate C8a bare-metal hosts with two AMD EPYC 9R45 sockets, 192 physical cores, SMT off, two NUMA nodes and 384 GiB installed memory |
| Serving studies | C8a virtual machines with 48, 96 and 192 cores and advertised 96, 192 and 384 GiB of RAM, plus a separate 16-core, 32-GiB frontend connected over private TCP |
| Generic memory contrast | One prover job on CPU 0 of the 48-core virtual machine |
| Memory limits | An isolated development host |
All hosts ran Linux. The CPU studies used Rust 1.96.1 with LLVM 22.1.2 and -Ctarget-cpu=native. The reported CPU model was AMD EPYC 9R45 with one thread per reported core. CPU frequency and the placement of individual worker threads were not fixed. Virtual machine sizes change CPU and RAM together, so these studies do not isolate a RAM-capacity effect. Runs with fewer workers on a large machine are labelled as such and say nothing about a smaller machine.
The serving workload cycles 192 fixed fixtures per operation kind with periodic arrivals. It supports repeated comparisons, and it does not represent random customer traffic, bursts or every circuit.
Reproduce the replayable parts#
Absolute values are comparable across machines only when hardware, firmware, operating system, compiler flags, thread placement, frequency control, background load and repetition policy are fixed. A small replay checks functional behavior. It does not reproduce an archived performance rate.
Development harness#
The repository includes a development harness for single-path latency, component timing, proof size before wrapping, sparse oracle memory and independent-job batch throughput at depths 24, 28 and 32:
RUSTFLAGS="-Ctarget-cpu=native" cargo run --release --locked --bin measureIts output is a development observation for your machine.
Small caller check#
On Linux with Rust 1.96.1, from the repository root:
cargo build --release --locked --manifest-path benches/controlled-capacity-2026-10-07/Load/Cargo.toml
python3 benches/controlled-capacity-2026-10-07/smoke.py --binary benches/controlled-capacity-2026-10-07/Load/target/release/ssgkr-load-driver --out /tmp/ssgkr-caller-small-checkThis runs six small control proofs, two fixtures for each operation kind. It checks sequential and parallel equality and honest encoded acceptance, and it confirms rejection of a mutated typed proof, root, value digest and trailing bytes. The output directory must be new and empty. The script records the built binary hash and the source commit. It does not rerun the archived campaigns.
Supplement data check#
The two earlier CPU studies keep their notes and data in the repository under benches/controlled-cpu-2026-10-06 and benches/controlled-cpu-2026-10-06-supplement. From the supplement folder, recalculate its compressed public data without compiling or proving:
python3 data.py check-package .Full dataset#
The complete raw observations, detailed tables, figure sources and replay drivers form a separate measurement dataset (dataset identifier to follow). The dataset has been prepared for review. Its public identifier and its data and figure license are still pending, and those terms are decided separately from the software license. When the dataset is available, verify the extracted payload against the contract hash in its publication metadata:
python3 benches/controlled-capacity-2026-10-07/verify-dataset.py --contract PACKAGE_CONTRACT.json --contract-sha256 THE_PUBLISHED_CONTRACT_SHA256 --manifest PUBLIC_MANIFEST.json --public PublicThe verifier rejects empty or partial manifests, missing or unmanifested files, changed bytes and schema mismatches. Its result is file-contract verification. It does not reproduce a benchmark result.
Scope of the measurements#
- Comparisons: Every ratio compares StateSync-GKR with its own reference implementation under the same conditions, so each gain belongs to one design choice.
- System outcomes: Measurements end at inner-proof acceptance. Database commits, root adoption, consensus, settlement and finality belong to the surrounding system.
- Service capacity: CPU throughput is a computation rate; only the serving studies include queueing and transport. Serving results describe the named host, mix, batch policy, period and latency budget.
- Causal attribution: The engine's nine design mechanisms (sparse reduction, derived wiring, cube and affine fusion, parallel path constraints, preparation reuse, independent jobs, three-state leaves, strict request binding and the reusable core) are supported by a mix of measurements, source structure, rejection controls and formal models, and each one names the kind of evidence behind it.
- Security: Security properties are covered in Formal verification and Trust boundaries.
- Development notes: The repository's development measurement report was produced on a development laptop and records development history, separate from the controlled studies on this page.