Benchmarks

Benchmark results

Controlled measurements of StateSync-GKR memory, prepared verification, preparation reuse, batch output, proof size, CPU throughput and full-proof serving, each with its comparison target, conditions, sample and exclusions.

This page summarizes the controlled measurements behind the design and evaluation paper (arXiv link to follow) and the repository's benchmark notes. Use it to estimate what StateSync-GKR changes for memory, repeated requests and serving, and read each result together with its comparison table.

Results at a glance#

ResultMeasured valueReference or offered loadScope
Peak process memory89.46× lower whole-process RSSTest-only Dense prover reference in the same engineTwo-layer mixed circuit, width 4,096, one job
Prover allocation3,242.41× lower additional requested allocationSame Dense referenceSame circuit and job
Prepared verification10.80–23.68× fasterTable wiring on the same proof and requestComplete prepared typed verification call, nine operation and depth profiles, preparation excluded
Preparation reuse3.46–4.45× lower latencyFresh setup for every requestPer-request proving path at depth 24
Complete local batch9.71–10.71× lower timeSerial encoding and encoded verificationDepth 24, batch 768, 192 workers
Proof size across depthAbout 2.93% largerMembership proof at depth 24Membership proof at depth 32
Sustained serving450,000 requests, every one accepted within 500 msOffered load of 500 requests/s48 cores, 15 minutes
Delivery at higher load69,000 requests, every one accepted within 2 s, 1,123.89 accepted requests/s including drainOffered load of 1,150 requests/s192 cores, 60-second input window

For a visual overview, see StateSync-GKR performance.

Memory with sparse layer reduction#

The production prover uses a Libra-style two-phase sparse oracle. It keeps input-sized buffers and a sparse list of multiplication gates instead of quadratic input-pair tables. What you gain is room for a wider circuit or more concurrent proof jobs within the same RAM budget.

ItemCondition
ComparisonProduction SparseLayerOracle against the test-only DenseLayerOracle, which builds six quadratic factor tables
Held fixedCircuit, field, transcript order and prover driver
WorkloadTwo-layer mixed linear and product circuit, width 4,096
HostOne prover job pinned to CPU 0 of a 48-core VM
SampleThree processes per mode, each with one warmup and three timed proofs; each value is the median of the three process medians
Result checkCanonical proof observations match for cases where both modes complete, and verification accepts every completed pair
ExcludedAllocator metadata and slack, stacks and transient reallocation overlap are outside requested bytes
BoundaryDense referenceSparse (production)
Additional requested allocation peak1,879,247,872 B579,584 B
Exposed oracle Vec capacity1,610,612,736 B508,096 B
Whole-process peak RSS (GNU time)1,855,016 KiB20,736 KiB
Observed prover interval11,610,216,436 ns4,630,590 ns

Requested allocation is 3,242.41 times lower and whole-process RSS is 89.46 times lower. The two ratios cover different ownership intervals and cannot be substituted for each other. Allocator counters and stage observation stay inside both prover timers and are not subtracted. JSON construction, hashing and verification run after the timer but remain inside process RSS.

Memory

The wider the circuit, the larger the saving.

89.46×lower peak process RAM at width 4,096
2101001K416642561K4KWidth of one circuit layer (wires)Peak process RAM, MiB (log scale)1,811.5 MiB20.25 MiB

Same engine, circuit, field, transcript and prover driver; only the layer representation changes. Two-layer mixed circuit, three processes per point, one job pinned to one core. Below width 256 the two are about the same; the gap opens as layers widen. The dense reference is StateSync-GKR's own implementation of the conventional dense layout.

The Dense reference is a test-only part of StateSync-GKR, so the result measures a representation change inside one engine, on a generic two-layer circuit, for one job on CPU 0. Plonky3 supplies field, hash and challenger code to both modes and is not a comparison target, and no external GKR implementation is involved. The ratio does not predict the memory reduction for a production sparse-Merkle circuit, and a single pinned job says nothing about 48-core throughput.

Memory limits and tested scale#

A separate study on an isolated development host ran Sparse and Dense under explicit memory limits. It preserved 17 attempts: 14 successes, one address-space allocation abort and two native out-of-memory kills. Concurrent jobs share one generic circuit and witness while each keeps its own oracle and proof state.

ConstraintWidthSparse (production)Dense reference
4-GiB address-space guard, concurrent jobs2,0481, 2, 4 and 8 jobs completed1, 2 and 4 jobs completed; allocation abort at 8
3-GiB cgroup with swap disabled, concurrent jobs2,0484 and 8 jobs completed4 jobs completed; killed by the out-of-memory mechanism at 8
3-GiB cgroup with swap disabled, one job8,192 to 65,536Each width returned one warmup and three timed proofsKilled at width 8,192 before returning a proof

Width 65,536 is the largest width tested. It bounds the evidence without defining a supported maximum.

Prepared verification with derived wiring#

Derived wiring evaluates regular gate families in closed form and keeps an exact sparse residue for the rest. It lowers the repeated work of checking a prepared circuit's connections while the proof stays the same.

ItemCondition
ComparisonTableWiring, a sparse gate-list oracle, against DerivedRegularWiring
Held fixedCircuit, request, proof, commitment and verifier body; call order randomized
TimedThe complete prepared typed verify_sync_op_with call, including request validation and native leaf and path hashing
ExcludedCircuit, wiring and commitment preparation, proof generation, encoding and transport
Host192-core host, single worker
SampleFive processes per profile, 300 measured requests and five warmups per process; each process is summarized by the median of its 300 same-proof paired ratios, and the reported value is the median of the five process summaries

Table/Derived time ratio (higher means derived wiring is faster):

OperationDepth 24Depth 28Depth 32
Membership11.11×12.08×14.05×
Non-membership10.80×11.95×13.86×
Update18.95×19.96×23.68×
Prepared verification

The advantage grows as the tree gets deeper.

23.68×peak, single-leaf update at depth 32
0×5×10×15×20×25×11.11×12.08×14.05×Membership10.80×11.95×13.86×Non-membership18.95×19.96×23.68×Single-leaf update

Same proof and request, verified with derived wiring instead of explicit table wiring, which is already sparse. Each bar is the median of five process-level medians of 300 paired calls. The timer covers the whole typed verification call, including path hashing; preparation, proving, encoding and transport are outside it.

TableWiring already iterates a sparse gate list and is unrelated to the quadratic Dense prover reference in the memory comparison. Both wiring evaluators are part of StateSync-GKR. The verifier still recomputes native leaf and path hashes from the private witness, and that work stays inside both timers.

Preparation reuse#

PreparedSync holds the compiled circuit, its full circuit commitment and the derived wiring for one operation kind and configuration. Reusing it removes per-request compilation and commitment work. Each request still has its own witness, transcript and proof.

ItemCondition
ComparisonFresh path, which compiles the circuit and computes its commitment for every request, against the prepared path, which reuses that material
TimedPer-request proving path
ExcludedEncoding, transport and acceptance
SampleDepth 24; three processes per operation and mode, 1,000 timed requests per process; ratios compare run-matched process summaries, not individual requests
Reuse conditionSame operation kind, SMT parameters, strategy, configuration and circuit
OperationFresh/prepared ratio
MembershipAbout 3.69×
Non-membershipAbout 3.46×
UpdateAbout 4.45×
Repeated requests

Prepare the circuit once. Pay less for every request.

3.46–4.45×lower per-request proving time at depth 24
0 ms80 ms160 ms240 ms320 ms155.6 ms42.1 msMembership3.69× lower159.5 ms46.1 msNon-membership3.46× lower289.1 ms64.9 msSingle-leaf update4.45× lower

Depth 24, one request per call on a 192-core bare-metal host, three processes of 1,000 timed requests per operation and mode. Prepared material is reusable only for the same operation kind and configuration. Encoding, transport and acceptance are excluded.

Direct calls show where the setup cost sits. These are medians of five run-specific request p50 values for depth-24 membership, with 300 measured requests and five warmups per process. They are separate direct measurements; for the end-to-end comparison, use the fresh/prepared ratios above.

Direct callp50 (ms)
Compile with hints7.9021
Derive verifier wiring4.4190
Circuit commitment105.0273
Witness generation0.5544
Prove on circuit41.5387
Encode0.0864
Derived verifier5.8663
Table verifier65.1941

Complete local batch with parallel output#

Parallelizing the output stages, encoding and encoded verification, lowered the time of the complete prepared local batch. Both policies already parallelize witness generation and proving.

ItemCondition
ComparisonParallel output policy against serial output policy; only encoding and encoded verification change mode
TimedThe complete prepared local batch: witness generation, proving, encoding and encoded verification
ExcludedPreparation, transport, post-timer proof hashing and logging, and state commits
WorkloadDepth 24, batch 768, 192 workers, two-field occupied payloads
SampleFive processes per operation; each process is summarized by the median of ten paired serial/parallel interval ratios, and the range spans the three operation medians across processes
Result check327,360 serial/parallel proof-byte pairs preserve the canonical result
OperationSerial/parallel time ratioParallel output (proofs/s)Serial output (proofs/s)
Membership10.71×1,593.17148.98
Non-membership9.97×1,431.66144.01
Update9.71×1,167.49121.01
Local batches

Parallel output keeps a full batch moving.

9.71–10.71×faster complete batch of 768 proofs
0 s1 s2 s3 s4 s5 s6 s7 s5.16 s0.48 sMembership10.71× faster5.33 s0.54 sNon-membership9.97× faster6.35 s0.66 sSingle-leaf update9.71× faster

Depth 24, batch of 768, 192 workers on a 192-core bare-metal host, five processes with ten paired batches each. Witness generation and proving are parallel in both modes; only encoding and encoded verification change. Each request keeps its own proof, nothing is aggregated, and transport is excluded.

The ratio comes from paired interval ratios, so it can differ from the quotient of the two throughput columns. Each throughput value is the median of five process rates, and each process rate is 768 divided by the median of ten complete-pipeline intervals after one warmup batch. In the same study, smaller batches and runs with 48 workers produced smaller ratios.

Proof size across tree depth#

Cube gates with affine fusion represent a Poseidon2 round in one layer, and Merkle path compressions are checked as parallel constraints. As a result, the compiler produces 118 layers in all nine measured profiles: membership, non-membership and update at depths 24, 28 and 32.

ItemCondition
ComparisonEach operation's proof at depth 32 against the same operation's proof at depth 24
MeasuredLength of the complete inner-proof-v1 encoded proof, before any external wrapping
ProfilesMembership, non-membership and update at depths 24, 28 and 32
SampleThe encoded layout is fixed by the compiled circuit, so each profile has one proof length
ExcludedThe private witness, which is never part of the encoding
ProfileDepth 24Depth 32Change
Membership encoded proof176,948 bytes182,132 bytesAbout +2.93%
Non-membership encoded proof181,646 bytes182,132 bytesAbout +0.27%
Single-leaf update encoded proof186,992 bytes201,086 bytesAbout +7.54%
Proof size

Deeper trees add only a few percent to the proof.

+0.27–7.54%encoded proof bytes, depth 24 to 32, by operation
0 KiB50 KiB100 KiB150 KiB200 KiB173178178Membership177178178Non-membership183196196Single-leaf update

Complete encoded inner proofs, 768 fixtures per operation and depth, all with the same 118-layer schedule. Membership grows 2.93%, non-membership 0.27% and update 7.54% from depth 24 to 32; circuit width and total work still grow with depth.

Circuit width, witness work and total computation still grow with depth, so the compact schedule limits proof growth without making the work constant. Cube and affine fusion has no separately measured speed multiplier.

CPU throughput#

These rates measure inner proof jobs on one large host. They show how much prepared work the engine completes in parallel.

ItemCondition
HostTwo AMD EPYC 9R45 sockets, 192 physical cores, SMT off, 384 GiB installed RAM, Linux, Rust 1.96.1 with -Ctarget-cpu=native
WorkloadDepth 24, batch 768, 192 workers, two-field occupied payloads, fixtures with independent roots
SampleEach value is the median of three process-level rates; each rate uses the median of three measured batch intervals
ExcludedEncoding, hash and control checks, and delivery
OperationProving only (proofs/s)Serial witness and proving (proofs/s)Parallel witness and proving (proofs/s)
Membership3,207.741,060.352,881.55
Non-membership2,930.851,021.872,654.03
Update2,092.03594.561,984.00
Parallel proving

Independent proofs scale with the cores you give them.

3,207.74membership proofs per second on 192 workers
05001,0001,5002,0002,5003,0003,500124816324896192Worker threads on one 192-core hostProofs per second

Prepared proving only at depth 24, batch 768, three processes per point on a two-socket bare-metal host with SMT off. Witness generation, encoding and transport are excluded. Worker counts are thread limits on one machine, not smaller machines.

The parallel-witness column is orchestration in the benchmark caller around the public APIs; the facade has no built-in witness-parallel method. These are inner-job computation rates. For end-to-end serving, see the serving results below, which add queueing and transport but still exclude state commits.

Serving through full-proof acceptance#

The serving studies send complete encoded proofs over private TCP to a separate 16-core frontend. A request counts as served only when the frontend's strict encoded verifier accepts it against the original request, so queueing and transport are inside the measured latency.

ItemCondition
ArrivalsPeriodic and scheduled independently of responses
WorkloadMembership, non-membership and update in a 1:1:1 mix at depth 24; 192 fixed fixtures per operation kind, cycled
Worker policyOne FIFO per operation kind, oldest ready batch first, one compute batch at a time, 5-ms maximum batch wait
EndpointFrontend checks framing, hash and original-request binding, then strict encoded verification accepts
LatencyFrom the scheduled arrival to acceptance, on the frontend's monotonic clock
ExcludedDatabase commits, root adoption, consensus and settlement
DeploymentOffered requests/sWorkers / batch capConfirmation
48 cores, 96 GiB50048 / 48One 900-second run
Two 48-core, 96-GiB workers1,000 total, 500 each48 / 48 eachThree 20-second runs
192 cores, 384 GiB1,000192 / 192Three 60-second runs
192 cores, 384 GiB1,150192 / 192One 60-second run

48 cores for 15 minutes#

MetricValue
Timed requests450,000 at 500 requests/s over 900 s
Accepted within 500 msAll
Accepted within 250 ms95.601333%
p50 / p95 / p99 / maximum165.114 / 248.255 / 267.025 / 307.096 ms
Mean CPU use27.108 core equivalents on the worker, 3.479 on the frontend
Sustained serving

Fifteen minutes at a steady 500 requests a second.

450,000requests in 15 minutes, every one within 500 ms
0 ms100 ms200 ms300 ms400 ms500 ms02468101214Minute of the runRequest to accepted proof500 ms budget

One 900-second run on a 48-core server: 48 workers, batch cap 48, a 1:1:1 mix of membership, non-membership and update at depth 24, offered at 500 per second, with a separate 16-core frontend. Latency ends when the complete proof is cryptographically accepted against its original request.

Minute-level mean pending counts stayed around 48 to 49 during the run. The result covers this 15-minute window and makes no availability guarantee beyond it.

Two 48-core workers against one 96-core worker#

This comparison holds total worker resources at 96 cores and 192 GiB and uses the same frontend and 5-ms wait. Each policy ran three 20-second processes at 1,000 requests/s in total.

PolicyResult
Two 48-core workers, 48 workers and batch cap 48 eachEvery timed request accepted within 500 ms; process p99 282.108–329.785 ms
One 96-core worker, 96 workers and batch cap 96Every submitted request eventually accepted while the queue grew; process p99 2.545–2.752 s; 27.075–32.31% accepted within one second

Each policy has its own batch cap, so this compares deployment policies as a whole. The result supports the two-worker policy at this load and period. It says nothing about other 96-core settings or about stability over 15 minutes.

192 cores at 1,000 and 1,150 requests per second#

At 1,000 requests/s, three 60-second runs accepted all 180,000 timed requests within one second. Process p99 values were 426.032, 438.775 and 443.471 ms, and 99.9183%, 99.7167% and 99.6217% of requests were accepted within 500 ms. Mean worker use was about 64.57 to 65.24 of the 192 available core equivalents.

At 1,150 requests/s, one 60-second input window produced these results:

MetricValue
Timed requests69,000 (256 warmups excluded)
Accepted within 2 sAll 69,000
Accepted within 1 s48,553 (70.3667%)
p99 / maximum latency1,471.284 / 1,598.489 ms
Pending requestsGrew from 562 at 5 s to 1,559 at 60 s
Final acceptanceAt 61.394 s
Completion-window rate1,123.89 accepted requests/s, including drain
Peak delivery

Above 1,000 requests a second, still inside two seconds.

69,000requests at 1,150 per second, all accepted within 2 s
0 ms500 ms1 s1.5 s2 s0102030405060Second of the input window99th percentile, per arrival second2 s budget1 s

One 60-second window at 1,150 per second on a 192-core server (192 workers, batch cap 192), all 69,000 requests accepted within 2 seconds, slowest 1.598 s, finishing at 1,123.89 per second including the drain. The queue grew during the window and 70.37% met 1 second. Three runs at 1,000 per second met 1 second for all 180,000 requests.

Three earlier 15-second screens at 1,150 requests/s met one second for every timed request; the 60-second confirmation did not. The 20,447 requests above one second were all accepted later; none was rejected or lost. Because the queue grew during the window, the run establishes the finite-window result above and leaves sustained capacity at 1,150 requests/s unmeasured. Sixty seconds is the length of the input window, and the prover remains usable after it.

Most of the added delay occurred before computation began. Server pre-compute wait had p99 1,367.973 ms at 1,150 requests/s, against 343.967–354.871 ms in the 1,000 requests/s confirmations, while mean worker use was 78.09 core equivalents. The serving caller processes one batch at a time and finishes hashing, framing, writing and logging each result before it starts the next batch, so the caller's return path matters alongside the prover.

Tuning observations on 48 cores#

A fine-grained study on the 48-core VM tested 19 conditions with three processes each across worker count, batch cap, wait, rate and operation mix.

  • Of 456,750 timed offers, 451,429 were accepted and 5,321 received explicit server-run-limit refusals. The refusals occurred at batch cap 16 with 500 requests/s and at batch cap 48 with 1,000 requests/s, when the 30-second run limit left jobs uncomputed, and they stay in the success denominators.
  • At a balanced 500 requests/s, 32 workers had p99 latency around 4.56 to 4.66 s.
  • Worker counts of 36 to 40 and a batch cap of 40 kept every timed request below 500 ms in 15-second repetitions. Treat them as short-run candidates measured on the 48-core VM; a smaller VM needs its own measurement.
  • An update-only load at 500 requests/s did not meet one second in every process, so the balanced result does not transfer to every operation mix.

To apply these observations to your own deployment, see Performance tuning.