Performance

Measured, run by run,
with every condition shown.

Every figure on this page comes from a controlled experiment reported in the engine paper. Each one is measured against StateSync-GKR's own reference and states its own conditions.

~89×less peak process memory than the dense reference, width 4,096
23.68×faster prepared verification at peak, 10.80× at minimum
450,000requests in 15 minutes on 48 cores, all accepted within 500 ms
1,123.89/scompletion rate at 1,150 offered per second, all within 2 s

Memory

Fit more work in the same machine.

The sparse prover never builds the square tables that the dense reference needs, so memory stays small as circuits widen.

Memory

The wider the circuit, the larger the saving.

89.46×lower peak process RAM at width 4,096
2101001K416642561K4KWidth of one circuit layer (wires)Peak process RAM, MiB (log scale)1,811.5 MiB20.25 MiB

Same engine, circuit, field, transcript and prover driver; only the layer representation changes. Two-layer mixed circuit, three processes per point, one job pinned to one core. Below width 256 the two are about the same; the gap opens as layers widen. The dense reference is StateSync-GKR's own implementation of the conventional dense layout.

Memory headroom

Under a 3 GiB cap, the sparse prover keeps going.

65,536widest layer proved under a 3 GiB cap
Layer widthDense referenceStateSync-GKR sparse
8,192Stopped: out of memoryProved · peak RAM 36 MiB
16,384Not run past the failureProved · peak RAM 67.4 MiB
32,768Not run past the failureProved · peak RAM 132 MiB
65,536Not run past the failureProved · peak RAM 259.9 MiB

Bounded tests on an isolated development host with a 3 GiB memory limit and swap disabled. The dense reference was stopped by the out-of-memory mechanism at width 8,192; the sparse prover returned proofs at every tested width. Peak RAM is the peak resident memory of the proving process. Width 65,536 is the largest point tested, not a supported maximum.

Verification

Make checking cheap for whoever receives the proof.

Derived wiring replaces lookups with structure. The verifying party, whether a counterparty, an auditor or a settlement service, does less work for the same proof.

Prepared verification

The advantage grows as the tree gets deeper.

23.68×peak, single-leaf update at depth 32
0×5×10×15×20×25×11.11×12.08×14.05×Membership10.80×11.95×13.86×Non-membership18.95×19.96×23.68×Single-leaf update

Same proof and request, verified with derived wiring instead of explicit table wiring, which is already sparse. Each bar is the median of five process-level medians of 300 paired calls. The timer covers the whole typed verification call, including path hashing; preparation, proving, encoding and transport are outside it.

Proof size

Deeper trees add only a few percent to the proof.

+0.27–7.54%encoded proof bytes, depth 24 to 32, by operation
0 KiB50 KiB100 KiB150 KiB200 KiB173178178Membership177178178Non-membership183196196Single-leaf update

Complete encoded inner proofs, 768 fixtures per operation and depth, all with the same 118-layer schedule. Membership grows 2.93%, non-membership 0.27% and update 7.54% from depth 24 to 32; circuit width and total work still grow with depth.

Repeated requests

Pay for setup once, then keep proofs independent.

Preparation is shared across requests of the same kind; witnesses, transcripts and proofs never are. Parallel output keeps a full batch from stalling at the end.

Repeated requests

Prepare the circuit once. Pay less for every request.

3.46–4.45×lower per-request proving time at depth 24
0 ms80 ms160 ms240 ms320 ms155.6 ms42.1 msMembership3.69× lower159.5 ms46.1 msNon-membership3.46× lower289.1 ms64.9 msSingle-leaf update4.45× lower

Depth 24, one request per call on a 192-core bare-metal host, three processes of 1,000 timed requests per operation and mode. Prepared material is reusable only for the same operation kind and configuration. Encoding, transport and acceptance are excluded.

Local batches

Parallel output keeps a full batch moving.

9.71–10.71×faster complete batch of 768 proofs
0 s1 s2 s3 s4 s5 s6 s7 s5.16 s0.48 sMembership10.71× faster5.33 s0.54 sNon-membership9.97× faster6.35 s0.66 sSingle-leaf update9.71× faster

Depth 24, batch of 768, 192 workers on a 192-core bare-metal host, five processes with ten paired batches each. Witness generation and proving are parallel in both modes; only encoding and encoded verification change. Each request keeps its own proof, nothing is aggregated, and transport is excluded.

Parallel proving

Independent proofs scale with the cores you give them.

3,207.74membership proofs per second on 192 workers
05001,0001,5002,0002,5003,0003,500124816324896192Worker threads on one 192-core hostProofs per second

Prepared proving only at depth 24, batch 768, three processes per point on a two-socket bare-metal host with SMT off. Witness generation, encoding and transport are excluded. Worker counts are thread limits on one machine, not smaller machines.

Serving

Hold a latency budget under load.

Serving runs measure the whole path: a request is offered, proved, encoded, sent to a separate verifying frontend and accepted against its original request.

Sustained serving

Fifteen minutes at a steady 500 requests a second.

450,000requests in 15 minutes, every one within 500 ms
0 ms100 ms200 ms300 ms400 ms500 ms02468101214Minute of the runRequest to accepted proof500 ms budget

One 900-second run on a 48-core server: 48 workers, batch cap 48, a 1:1:1 mix of membership, non-membership and update at depth 24, offered at 500 per second, with a separate 16-core frontend. Latency ends when the complete proof is cryptographically accepted against its original request.

Peak delivery

Above 1,000 requests a second, still inside two seconds.

69,000requests at 1,150 per second, all accepted within 2 s
0 ms500 ms1 s1.5 s2 s0102030405060Second of the input window99th percentile, per arrival second2 s budget1 s

One 60-second window at 1,150 per second on a 192-core server (192 workers, batch cap 192), all 69,000 requests accepted within 2 seconds, slowest 1.598 s, finishing at 1,123.89 per second including the drain. The queue grew during the window and 70.37% met 1 second. Three runs at 1,000 per second met 1 second for all 180,000 requests.

Deployment

How you split the cores matters.

2 × 48beat 1 × 96 at the same total resources
0 ms500 ms1 s1.5 s2 s2.5 s3 s284 ms282 ms330 msTwo 48-core workers2.75 s2.55 s2.74 sOne 96-core worker

Three 20-second runs per policy at 1,000 requests per second with the same total worker resources (96 cores, 192 GiB) and a separate 16-core frontend. Two workers with batch cap 48 each accepted every request within 500 ms; one worker with batch cap 96 built a queue. This compares deployment policies, not topology alone.

Reading the numbers

Separate results, never one multiplier.

Each ratio has its own reference: the engine's dense representation for memory, explicit table wiring for verification, fresh preparation for repeated requests and serial output for batches. They describe different parts of the engine, so they are never multiplied or added into a single figure.

Each comparison changes one design choice inside StateSync-GKR and keeps everything else fixed, so the gain belongs to that choice. Plonky3 supplies the field, hash and challenger components, and the sparse reduction follows the Libra approach.

Methodology and reproductionDesign and evaluation paperarXiv link to follow