Performance
Measured, run by run,
with every condition shown.
Every figure on this page comes from a controlled experiment reported in the engine paper. Each one is measured against StateSync-GKR's own reference and states its own conditions.
Memory
Fit more work in the same machine.
The sparse prover never builds the square tables that the dense reference needs, so memory stays small as circuits widen.
The wider the circuit, the larger the saving.
Same engine, circuit, field, transcript and prover driver; only the layer representation changes. Two-layer mixed circuit, three processes per point, one job pinned to one core. Below width 256 the two are about the same; the gap opens as layers widen. The dense reference is StateSync-GKR's own implementation of the conventional dense layout.
Under a 3 GiB cap, the sparse prover keeps going.
Bounded tests on an isolated development host with a 3 GiB memory limit and swap disabled. The dense reference was stopped by the out-of-memory mechanism at width 8,192; the sparse prover returned proofs at every tested width. Peak RAM is the peak resident memory of the proving process. Width 65,536 is the largest point tested, not a supported maximum.
Verification
Make checking cheap for whoever receives the proof.
Derived wiring replaces lookups with structure. The verifying party, whether a counterparty, an auditor or a settlement service, does less work for the same proof.
The advantage grows as the tree gets deeper.
Same proof and request, verified with derived wiring instead of explicit table wiring, which is already sparse. Each bar is the median of five process-level medians of 300 paired calls. The timer covers the whole typed verification call, including path hashing; preparation, proving, encoding and transport are outside it.
Deeper trees add only a few percent to the proof.
Complete encoded inner proofs, 768 fixtures per operation and depth, all with the same 118-layer schedule. Membership grows 2.93%, non-membership 0.27% and update 7.54% from depth 24 to 32; circuit width and total work still grow with depth.
Repeated requests
Pay for setup once, then keep proofs independent.
Preparation is shared across requests of the same kind; witnesses, transcripts and proofs never are. Parallel output keeps a full batch from stalling at the end.
Prepare the circuit once. Pay less for every request.
Depth 24, one request per call on a 192-core bare-metal host, three processes of 1,000 timed requests per operation and mode. Prepared material is reusable only for the same operation kind and configuration. Encoding, transport and acceptance are excluded.
Parallel output keeps a full batch moving.
Depth 24, batch of 768, 192 workers on a 192-core bare-metal host, five processes with ten paired batches each. Witness generation and proving are parallel in both modes; only encoding and encoded verification change. Each request keeps its own proof, nothing is aggregated, and transport is excluded.
Independent proofs scale with the cores you give them.
Prepared proving only at depth 24, batch 768, three processes per point on a two-socket bare-metal host with SMT off. Witness generation, encoding and transport are excluded. Worker counts are thread limits on one machine, not smaller machines.
Serving
Hold a latency budget under load.
Serving runs measure the whole path: a request is offered, proved, encoded, sent to a separate verifying frontend and accepted against its original request.
Fifteen minutes at a steady 500 requests a second.
One 900-second run on a 48-core server: 48 workers, batch cap 48, a 1:1:1 mix of membership, non-membership and update at depth 24, offered at 500 per second, with a separate 16-core frontend. Latency ends when the complete proof is cryptographically accepted against its original request.
Above 1,000 requests a second, still inside two seconds.
One 60-second window at 1,150 per second on a 192-core server (192 workers, batch cap 192), all 69,000 requests accepted within 2 seconds, slowest 1.598 s, finishing at 1,123.89 per second including the drain. The queue grew during the window and 70.37% met 1 second. Three runs at 1,000 per second met 1 second for all 180,000 requests.
How you split the cores matters.
Three 20-second runs per policy at 1,000 requests per second with the same total worker resources (96 cores, 192 GiB) and a separate 16-core frontend. Two workers with batch cap 48 each accepted every request within 500 ms; one worker with batch cap 96 built a queue. This compares deployment policies, not topology alone.
Reading the numbers
Separate results, never one multiplier.
Each ratio has its own reference: the engine's dense representation for memory, explicit table wiring for verification, fresh preparation for repeated requests and serial output for batches. They describe different parts of the engine, so they are never multiplied or added into a single figure.
Each comparison changes one design choice inside StateSync-GKR and keeps everything else fixed, so the gain belongs to that choice. Plonky3 supplies the field, hash and challenger components, and the sparse reduction follows the Libra approach.