Benchmarks
Benchmark results
Controlled measurements of StateSync-GKR memory, prepared verification, preparation reuse, batch output, proof size, CPU throughput and full-proof serving, each with its comparison target, conditions, sample and exclusions.
This page summarizes the controlled measurements behind the design and evaluation paper (arXiv link to follow) and the repository's benchmark notes. Use it to estimate what StateSync-GKR changes for memory, repeated requests and serving, and read each result together with its comparison table.
Results at a glance#
| Result | Measured value | Reference or offered load | Scope |
|---|---|---|---|
| Peak process memory | 89.46× lower whole-process RSS | Test-only Dense prover reference in the same engine | Two-layer mixed circuit, width 4,096, one job |
| Prover allocation | 3,242.41× lower additional requested allocation | Same Dense reference | Same circuit and job |
| Prepared verification | 10.80–23.68× faster | Table wiring on the same proof and request | Complete prepared typed verification call, nine operation and depth profiles, preparation excluded |
| Preparation reuse | 3.46–4.45× lower latency | Fresh setup for every request | Per-request proving path at depth 24 |
| Complete local batch | 9.71–10.71× lower time | Serial encoding and encoded verification | Depth 24, batch 768, 192 workers |
| Proof size across depth | About 2.93% larger | Membership proof at depth 24 | Membership proof at depth 32 |
| Sustained serving | 450,000 requests, every one accepted within 500 ms | Offered load of 500 requests/s | 48 cores, 15 minutes |
| Delivery at higher load | 69,000 requests, every one accepted within 2 s, 1,123.89 accepted requests/s including drain | Offered load of 1,150 requests/s | 192 cores, 60-second input window |
For a visual overview, see StateSync-GKR performance.
Memory with sparse layer reduction#
The production prover uses a Libra-style two-phase sparse oracle. It keeps input-sized buffers and a sparse list of multiplication gates instead of quadratic input-pair tables. What you gain is room for a wider circuit or more concurrent proof jobs within the same RAM budget.
| Item | Condition |
|---|---|
| Comparison | Production SparseLayerOracle against the test-only DenseLayerOracle, which builds six quadratic factor tables |
| Held fixed | Circuit, field, transcript order and prover driver |
| Workload | Two-layer mixed linear and product circuit, width 4,096 |
| Host | One prover job pinned to CPU 0 of a 48-core VM |
| Sample | Three processes per mode, each with one warmup and three timed proofs; each value is the median of the three process medians |
| Result check | Canonical proof observations match for cases where both modes complete, and verification accepts every completed pair |
| Excluded | Allocator metadata and slack, stacks and transient reallocation overlap are outside requested bytes |
| Boundary | Dense reference | Sparse (production) |
|---|---|---|
| Additional requested allocation peak | 1,879,247,872 B | 579,584 B |
Exposed oracle Vec capacity | 1,610,612,736 B | 508,096 B |
| Whole-process peak RSS (GNU time) | 1,855,016 KiB | 20,736 KiB |
| Observed prover interval | 11,610,216,436 ns | 4,630,590 ns |
Requested allocation is 3,242.41 times lower and whole-process RSS is 89.46 times lower. The two ratios cover different ownership intervals and cannot be substituted for each other. Allocator counters and stage observation stay inside both prover timers and are not subtracted. JSON construction, hashing and verification run after the timer but remain inside process RSS.
The wider the circuit, the larger the saving.
Same engine, circuit, field, transcript and prover driver; only the layer representation changes. Two-layer mixed circuit, three processes per point, one job pinned to one core. Below width 256 the two are about the same; the gap opens as layers widen. The dense reference is StateSync-GKR's own implementation of the conventional dense layout.
The Dense reference is a test-only part of StateSync-GKR, so the result measures a representation change inside one engine, on a generic two-layer circuit, for one job on CPU 0. Plonky3 supplies field, hash and challenger code to both modes and is not a comparison target, and no external GKR implementation is involved. The ratio does not predict the memory reduction for a production sparse-Merkle circuit, and a single pinned job says nothing about 48-core throughput.
Memory limits and tested scale#
A separate study on an isolated development host ran Sparse and Dense under explicit memory limits. It preserved 17 attempts: 14 successes, one address-space allocation abort and two native out-of-memory kills. Concurrent jobs share one generic circuit and witness while each keeps its own oracle and proof state.
| Constraint | Width | Sparse (production) | Dense reference |
|---|---|---|---|
| 4-GiB address-space guard, concurrent jobs | 2,048 | 1, 2, 4 and 8 jobs completed | 1, 2 and 4 jobs completed; allocation abort at 8 |
| 3-GiB cgroup with swap disabled, concurrent jobs | 2,048 | 4 and 8 jobs completed | 4 jobs completed; killed by the out-of-memory mechanism at 8 |
| 3-GiB cgroup with swap disabled, one job | 8,192 to 65,536 | Each width returned one warmup and three timed proofs | Killed at width 8,192 before returning a proof |
Width 65,536 is the largest width tested. It bounds the evidence without defining a supported maximum.
Prepared verification with derived wiring#
Derived wiring evaluates regular gate families in closed form and keeps an exact sparse residue for the rest. It lowers the repeated work of checking a prepared circuit's connections while the proof stays the same.
| Item | Condition |
|---|---|
| Comparison | TableWiring, a sparse gate-list oracle, against DerivedRegularWiring |
| Held fixed | Circuit, request, proof, commitment and verifier body; call order randomized |
| Timed | The complete prepared typed verify_sync_op_with call, including request validation and native leaf and path hashing |
| Excluded | Circuit, wiring and commitment preparation, proof generation, encoding and transport |
| Host | 192-core host, single worker |
| Sample | Five processes per profile, 300 measured requests and five warmups per process; each process is summarized by the median of its 300 same-proof paired ratios, and the reported value is the median of the five process summaries |
Table/Derived time ratio (higher means derived wiring is faster):
| Operation | Depth 24 | Depth 28 | Depth 32 |
|---|---|---|---|
| Membership | 11.11× | 12.08× | 14.05× |
| Non-membership | 10.80× | 11.95× | 13.86× |
| Update | 18.95× | 19.96× | 23.68× |
The advantage grows as the tree gets deeper.
Same proof and request, verified with derived wiring instead of explicit table wiring, which is already sparse. Each bar is the median of five process-level medians of 300 paired calls. The timer covers the whole typed verification call, including path hashing; preparation, proving, encoding and transport are outside it.
TableWiring already iterates a sparse gate list and is unrelated to the quadratic Dense prover reference in the memory comparison. Both wiring evaluators are part of StateSync-GKR. The verifier still recomputes native leaf and path hashes from the private witness, and that work stays inside both timers.
Preparation reuse#
PreparedSync holds the compiled circuit, its full circuit commitment and the derived wiring for one operation kind and configuration. Reusing it removes per-request compilation and commitment work. Each request still has its own witness, transcript and proof.
| Item | Condition |
|---|---|
| Comparison | Fresh path, which compiles the circuit and computes its commitment for every request, against the prepared path, which reuses that material |
| Timed | Per-request proving path |
| Excluded | Encoding, transport and acceptance |
| Sample | Depth 24; three processes per operation and mode, 1,000 timed requests per process; ratios compare run-matched process summaries, not individual requests |
| Reuse condition | Same operation kind, SMT parameters, strategy, configuration and circuit |
| Operation | Fresh/prepared ratio |
|---|---|
| Membership | About 3.69× |
| Non-membership | About 3.46× |
| Update | About 4.45× |
Prepare the circuit once. Pay less for every request.
Depth 24, one request per call on a 192-core bare-metal host, three processes of 1,000 timed requests per operation and mode. Prepared material is reusable only for the same operation kind and configuration. Encoding, transport and acceptance are excluded.
Direct calls show where the setup cost sits. These are medians of five run-specific request p50 values for depth-24 membership, with 300 measured requests and five warmups per process. They are separate direct measurements; for the end-to-end comparison, use the fresh/prepared ratios above.
| Direct call | p50 (ms) |
|---|---|
| Compile with hints | 7.9021 |
| Derive verifier wiring | 4.4190 |
| Circuit commitment | 105.0273 |
| Witness generation | 0.5544 |
| Prove on circuit | 41.5387 |
| Encode | 0.0864 |
| Derived verifier | 5.8663 |
| Table verifier | 65.1941 |
Complete local batch with parallel output#
Parallelizing the output stages, encoding and encoded verification, lowered the time of the complete prepared local batch. Both policies already parallelize witness generation and proving.
| Item | Condition |
|---|---|
| Comparison | Parallel output policy against serial output policy; only encoding and encoded verification change mode |
| Timed | The complete prepared local batch: witness generation, proving, encoding and encoded verification |
| Excluded | Preparation, transport, post-timer proof hashing and logging, and state commits |
| Workload | Depth 24, batch 768, 192 workers, two-field occupied payloads |
| Sample | Five processes per operation; each process is summarized by the median of ten paired serial/parallel interval ratios, and the range spans the three operation medians across processes |
| Result check | 327,360 serial/parallel proof-byte pairs preserve the canonical result |
| Operation | Serial/parallel time ratio | Parallel output (proofs/s) | Serial output (proofs/s) |
|---|---|---|---|
| Membership | 10.71× | 1,593.17 | 148.98 |
| Non-membership | 9.97× | 1,431.66 | 144.01 |
| Update | 9.71× | 1,167.49 | 121.01 |
Parallel output keeps a full batch moving.
Depth 24, batch of 768, 192 workers on a 192-core bare-metal host, five processes with ten paired batches each. Witness generation and proving are parallel in both modes; only encoding and encoded verification change. Each request keeps its own proof, nothing is aggregated, and transport is excluded.
The ratio comes from paired interval ratios, so it can differ from the quotient of the two throughput columns. Each throughput value is the median of five process rates, and each process rate is 768 divided by the median of ten complete-pipeline intervals after one warmup batch. In the same study, smaller batches and runs with 48 workers produced smaller ratios.
Proof size across tree depth#
Cube gates with affine fusion represent a Poseidon2 round in one layer, and Merkle path compressions are checked as parallel constraints. As a result, the compiler produces 118 layers in all nine measured profiles: membership, non-membership and update at depths 24, 28 and 32.
| Item | Condition |
|---|---|
| Comparison | Each operation's proof at depth 32 against the same operation's proof at depth 24 |
| Measured | Length of the complete inner-proof-v1 encoded proof, before any external wrapping |
| Profiles | Membership, non-membership and update at depths 24, 28 and 32 |
| Sample | The encoded layout is fixed by the compiled circuit, so each profile has one proof length |
| Excluded | The private witness, which is never part of the encoding |
| Profile | Depth 24 | Depth 32 | Change |
|---|---|---|---|
| Membership encoded proof | 176,948 bytes | 182,132 bytes | About +2.93% |
| Non-membership encoded proof | 181,646 bytes | 182,132 bytes | About +0.27% |
| Single-leaf update encoded proof | 186,992 bytes | 201,086 bytes | About +7.54% |
Deeper trees add only a few percent to the proof.
Complete encoded inner proofs, 768 fixtures per operation and depth, all with the same 118-layer schedule. Membership grows 2.93%, non-membership 0.27% and update 7.54% from depth 24 to 32; circuit width and total work still grow with depth.
Circuit width, witness work and total computation still grow with depth, so the compact schedule limits proof growth without making the work constant. Cube and affine fusion has no separately measured speed multiplier.
CPU throughput#
These rates measure inner proof jobs on one large host. They show how much prepared work the engine completes in parallel.
| Item | Condition |
|---|---|
| Host | Two AMD EPYC 9R45 sockets, 192 physical cores, SMT off, 384 GiB installed RAM, Linux, Rust 1.96.1 with -Ctarget-cpu=native |
| Workload | Depth 24, batch 768, 192 workers, two-field occupied payloads, fixtures with independent roots |
| Sample | Each value is the median of three process-level rates; each rate uses the median of three measured batch intervals |
| Excluded | Encoding, hash and control checks, and delivery |
| Operation | Proving only (proofs/s) | Serial witness and proving (proofs/s) | Parallel witness and proving (proofs/s) |
|---|---|---|---|
| Membership | 3,207.74 | 1,060.35 | 2,881.55 |
| Non-membership | 2,930.85 | 1,021.87 | 2,654.03 |
| Update | 2,092.03 | 594.56 | 1,984.00 |
Independent proofs scale with the cores you give them.
Prepared proving only at depth 24, batch 768, three processes per point on a two-socket bare-metal host with SMT off. Witness generation, encoding and transport are excluded. Worker counts are thread limits on one machine, not smaller machines.
The parallel-witness column is orchestration in the benchmark caller around the public APIs; the facade has no built-in witness-parallel method. These are inner-job computation rates. For end-to-end serving, see the serving results below, which add queueing and transport but still exclude state commits.
Serving through full-proof acceptance#
The serving studies send complete encoded proofs over private TCP to a separate 16-core frontend. A request counts as served only when the frontend's strict encoded verifier accepts it against the original request, so queueing and transport are inside the measured latency.
| Item | Condition |
|---|---|
| Arrivals | Periodic and scheduled independently of responses |
| Workload | Membership, non-membership and update in a 1:1:1 mix at depth 24; 192 fixed fixtures per operation kind, cycled |
| Worker policy | One FIFO per operation kind, oldest ready batch first, one compute batch at a time, 5-ms maximum batch wait |
| Endpoint | Frontend checks framing, hash and original-request binding, then strict encoded verification accepts |
| Latency | From the scheduled arrival to acceptance, on the frontend's monotonic clock |
| Excluded | Database commits, root adoption, consensus and settlement |
| Deployment | Offered requests/s | Workers / batch cap | Confirmation |
|---|---|---|---|
| 48 cores, 96 GiB | 500 | 48 / 48 | One 900-second run |
| Two 48-core, 96-GiB workers | 1,000 total, 500 each | 48 / 48 each | Three 20-second runs |
| 192 cores, 384 GiB | 1,000 | 192 / 192 | Three 60-second runs |
| 192 cores, 384 GiB | 1,150 | 192 / 192 | One 60-second run |
48 cores for 15 minutes#
| Metric | Value |
|---|---|
| Timed requests | 450,000 at 500 requests/s over 900 s |
| Accepted within 500 ms | All |
| Accepted within 250 ms | 95.601333% |
| p50 / p95 / p99 / maximum | 165.114 / 248.255 / 267.025 / 307.096 ms |
| Mean CPU use | 27.108 core equivalents on the worker, 3.479 on the frontend |
Fifteen minutes at a steady 500 requests a second.
One 900-second run on a 48-core server: 48 workers, batch cap 48, a 1:1:1 mix of membership, non-membership and update at depth 24, offered at 500 per second, with a separate 16-core frontend. Latency ends when the complete proof is cryptographically accepted against its original request.
Minute-level mean pending counts stayed around 48 to 49 during the run. The result covers this 15-minute window and makes no availability guarantee beyond it.
Two 48-core workers against one 96-core worker#
This comparison holds total worker resources at 96 cores and 192 GiB and uses the same frontend and 5-ms wait. Each policy ran three 20-second processes at 1,000 requests/s in total.
| Policy | Result |
|---|---|
| Two 48-core workers, 48 workers and batch cap 48 each | Every timed request accepted within 500 ms; process p99 282.108–329.785 ms |
| One 96-core worker, 96 workers and batch cap 96 | Every submitted request eventually accepted while the queue grew; process p99 2.545–2.752 s; 27.075–32.31% accepted within one second |
Each policy has its own batch cap, so this compares deployment policies as a whole. The result supports the two-worker policy at this load and period. It says nothing about other 96-core settings or about stability over 15 minutes.
192 cores at 1,000 and 1,150 requests per second#
At 1,000 requests/s, three 60-second runs accepted all 180,000 timed requests within one second. Process p99 values were 426.032, 438.775 and 443.471 ms, and 99.9183%, 99.7167% and 99.6217% of requests were accepted within 500 ms. Mean worker use was about 64.57 to 65.24 of the 192 available core equivalents.
At 1,150 requests/s, one 60-second input window produced these results:
| Metric | Value |
|---|---|
| Timed requests | 69,000 (256 warmups excluded) |
| Accepted within 2 s | All 69,000 |
| Accepted within 1 s | 48,553 (70.3667%) |
| p99 / maximum latency | 1,471.284 / 1,598.489 ms |
| Pending requests | Grew from 562 at 5 s to 1,559 at 60 s |
| Final acceptance | At 61.394 s |
| Completion-window rate | 1,123.89 accepted requests/s, including drain |
Above 1,000 requests a second, still inside two seconds.
One 60-second window at 1,150 per second on a 192-core server (192 workers, batch cap 192), all 69,000 requests accepted within 2 seconds, slowest 1.598 s, finishing at 1,123.89 per second including the drain. The queue grew during the window and 70.37% met 1 second. Three runs at 1,000 per second met 1 second for all 180,000 requests.
Three earlier 15-second screens at 1,150 requests/s met one second for every timed request; the 60-second confirmation did not. The 20,447 requests above one second were all accepted later; none was rejected or lost. Because the queue grew during the window, the run establishes the finite-window result above and leaves sustained capacity at 1,150 requests/s unmeasured. Sixty seconds is the length of the input window, and the prover remains usable after it.
Most of the added delay occurred before computation began. Server pre-compute wait had p99 1,367.973 ms at 1,150 requests/s, against 343.967–354.871 ms in the 1,000 requests/s confirmations, while mean worker use was 78.09 core equivalents. The serving caller processes one batch at a time and finishes hashing, framing, writing and logging each result before it starts the next batch, so the caller's return path matters alongside the prover.
Tuning observations on 48 cores#
A fine-grained study on the 48-core VM tested 19 conditions with three processes each across worker count, batch cap, wait, rate and operation mix.
- Of 456,750 timed offers, 451,429 were accepted and 5,321 received explicit server-run-limit refusals. The refusals occurred at batch cap 16 with 500 requests/s and at batch cap 48 with 1,000 requests/s, when the 30-second run limit left jobs uncomputed, and they stay in the success denominators.
- At a balanced 500 requests/s, 32 workers had p99 latency around 4.56 to 4.66 s.
- Worker counts of 36 to 40 and a batch cap of 40 kept every timed request below 500 ms in 15-second repetitions. Treat them as short-run candidates measured on the 48-core VM; a smaller VM needs its own measurement.
- An update-only load at 500 requests/s did not meet one second in every process, so the balanced result does not transfer to every operation mix.
To apply these observations to your own deployment, see Performance tuning.