Guides

Performance tuning

Choose tree depth, leaf bound, memory budget, worker count, batch cap and deployment shape for StateSync-GKR, with the measured evidence behind each choice and general guidance where no measurement exists.

This guide helps you choose the settings that decide how fast StateSync-GKR runs and how much memory it uses. Each section separates what was measured, and under which conditions, from guidance for the cases the measurements do not cover.

Before you begin#

  • Build with --release and use the prepared path described in Prepared execution. The fresh calls compile the circuit and compute its commitment on every call.
  • Parallelize witness generation and output processing in your own code where it helps; Batching and parallelism shows how and what that gained.
  • Read every measured number with its conditions. The results on this page and in Benchmark results come from separate experiments, and their ratios do not combine.

What each setting changes#

SettingWhere you set itChanges the circuit and proofsWhat it trades
Tree depth, smt.depthStateSyncGkrConfigYesKey space against witness, proving and verification work
Leaf bound, smt.leaf_max_fieldsStateSyncGkrConfigYesPayload capacity against the cost of every leaf hash and the depth of the circuit
Worker countYour Rayon poolNoParallel work against memory for the jobs in flight
Batch cap and maximum waitYour schedulerNoFuller batches against time spent waiting
Deployment shapeYour infrastructureNoLatency and queueing for a given total resource
Build flags such as -Ctarget-cpu=nativeYour buildNoCPU-specific code generation

A setting that changes the circuit gives it a new commitment and a new circuit identity, so provers and verifiers must change together and existing trees keep their own configuration. The other settings leave proof bytes unchanged, because each proof depends only on its own request.

Tree depth and leaf bound#

What was measured#

Depths 24, 28 and 32 were measured. All nine operation and depth profiles compiled to 118 layers, and the complete encoded proofs grew slowly with depth. Each profile had 768 accepted fixtures, all with the same proof length:

OperationDepth 24 (bytes)Depth 28 (bytes)Depth 32 (bytes)Change from 24 to 32
Membership176,948181,970182,132+2.93%
Non-membership181,646181,970182,132+0.27%
Update186,992200,762201,086+7.54%

Prepared verification time changed little with depth when the verifier used derived wiring. The table gives the median time of the complete prepared verification call, including request validation and the native hashing of the leaf and its path, on one worker of the 192-core host. Each value is the median of five process medians, each over 300 calls:

OperationDepth 24 (ms)Depth 28 (ms)Depth 32 (ms)
Membership5.876.185.94
Non-membership6.096.246.02
Update7.247.987.56

Witness generation, proving and total computation still grow with depth, because the circuit becomes wider.

How to choose#

  • Pick the smallest depth whose 2^depth slots cover your key space with room to grow. The default depth, 24, gives about 16.7 million slots.
  • Treat the depth and the leaf bound as fixed for the life of a tree. Changing either produces a different circuit, needs new preparations on every prover and verifier, and for the leaf bound can change every leaf digest and root (Configuration).
  • Expect to measure for yourself outside the measured range. Depths other than 24, 28 and 32 have no measurements, and the studies did not vary the leaf bound. The leaf pre-image is the bound plus one, rounded up to a multiple of eight field elements, and the leaf hash spends one Poseidon2 permutation per eight elements. Raising the bound past one of those steps therefore adds a permutation to every leaf hash, both in the prover's witness and in the input the verifier rebuilds. Each extra permutation also adds 29 layers to the circuit, which lengthens every proof and adds proving and verification work. The default bound of 31 fills four permutations exactly, so raising it to 32 already adds a fifth. This follows from the circuit layout; the studies did not measure it.
  • Before you choose a depth of 31 or more, read the note on key absorption in Configuration.

Memory#

What was measured#

The production prover uses Libra-style two-phase sparse booking. For a layer with n_in input wires, n_out output wires and m multiplication gates, its auxiliary memory is O(n_out + n_in + m), and it allocates no table over pairs of input wires.

A separate study on an isolated development host ran the production sparse prover and the engine's own test-only dense reference on generic two-layer circuits under explicit memory limits:

Memory headroom

Under a 3 GiB cap, the sparse prover keeps going.

65,536widest layer proved under a 3 GiB cap
Layer widthDense referenceStateSync-GKR sparse
8,192Stopped: out of memoryProved · peak RAM 36 MiB
16,384Not run past the failureProved · peak RAM 67.4 MiB
32,768Not run past the failureProved · peak RAM 132 MiB
65,536Not run past the failureProved · peak RAM 259.9 MiB

Bounded tests on an isolated development host with a 3 GiB memory limit and swap disabled. The dense reference was stopped by the out-of-memory mechanism at width 8,192; the sparse prover returned proofs at every tested width. Peak RAM is the peak resident memory of the proving process. Width 65,536 is the largest point tested, not a supported maximum.

  • Under a 3-GiB cgroup with swap disabled, the sparse prover returned proofs at widths 8,192, 16,384, 32,768 and 65,536, with peak process memory between 36.0 and 259.9 MiB. The dense reference was killed by the out-of-memory mechanism at width 8,192 before it returned a proof.
  • At width 2,048, the sparse prover completed 1, 2, 4 and 8 concurrent jobs under a 4-GiB address-space guard, and 4 and 8 concurrent jobs under the 3-GiB cgroup. The dense reference completed 1, 2 and 4 jobs under the address-space guard and 4 under the cgroup, and failed at 8 under both.
  • Width 65,536 is the largest width tested, not a supported maximum. The study used generic circuits and does not predict the memory of the sparse-Merkle circuit, and the dense reference is part of StateSync-GKR rather than another GKR implementation.

How to choose#

  • Budget memory per job in flight. Concurrent jobs share one prepared circuit, but each keeps its own witness, prover state, transcript and proof, so memory grows with the number of jobs running at once.
  • Measure your own circuit at your depth and concurrency under the memory limit you will deploy with, and keep the failed attempts. An address-space allocation failure and an out-of-memory kill are different outcomes, and a killed process leaves no peak memory to report.
  • Share one PreparedSync per operation kind instead of cloning it, because a clone copies the compiled circuit.

Worker count#

What was measured#

On one 192-core bare-metal host, with two AMD EPYC 9R45 sockets and SMT off, prepared proving-only throughput at depth 24 with batches of 768 grew with the worker count:

Parallel proving

Independent proofs scale with the cores you give them.

3,207.74membership proofs per second on 192 workers
05001,0001,5002,0002,5003,0003,500124816324896192Worker threads on one 192-core hostProofs per second

Prepared proving only at depth 24, batch 768, three processes per point on a two-socket bare-metal host with SMT off. Witness generation, encoding and transport are excluded. Worker counts are thread limits on one machine, not smaller machines.

WorkersMembership (proofs/s)Non-membership (proofs/s)Update (proofs/s)
124.0121.9515.66
8190.92174.30124.52
481,116.711,019.66735.33
962,058.711,917.651,383.05
1923,207.742,930.852,092.03

Each value is the median of three processes. These are proving-only rates, without witness generation, encoding or transport. The worker counts were thread limits on one large host, so a lower count does not show how a smaller machine would perform.

In serving, every confirmed setting used one worker per available core: 48 workers on the 48-core VM and 192 on the 192-core VM. On the 48-core VM at 500 requests/s, 32 workers let the queue grow, and accepted-request p99 latency reached 4.56 to 4.66 seconds.

How to choose#

  • Size the Rayon pool explicitly and run batches inside pool.install, as described in Batching and parallelism.
  • Start from one worker per available core, then check that memory suffices for that many jobs in flight.
  • Measure on the machine size you will deploy. Lowering the thread count on a larger machine does not stand in for a smaller one.

Batch cap and maximum wait#

What was measured#

A fine-grained study on the 48-core VM at a balanced 500 requests/s varied the worker count with batch cap 48, and the batch cap with 48 workers, in 15-second runs repeated three times each. The p99 column is over accepted requests:

SettingAccepted-request p99Notes
32 workers, batch cap 484,557.771–4,663.165 ms15.93–16.83% of requests within one second
36, 40 or 44 workers, batch cap 48266.386–277.384 msEvery request within one second
48 workers, batch cap 1611,231.230–11,347.565 ms847, 912 and 896 explicit refusals in the three runs
48 workers, batch cap 244,402.638–4,789.998 ms
48 workers, batch cap 32433.563–474.808 ms
48 workers, batch cap 40270.028–278.601 ms
48 workers, batch cap 48262.789–263.385 msThree earlier screening runs, the reference for both comparisons

The refusals at batch cap 16 came from the benchmark server's run limit, which ended runs while jobs were still uncomputed, and they remain in the success denominators. Workers 36 to 40 with batch cap 48, and batch cap 40 with 48 workers, kept every timed request below 500 ms in these short runs. The study also tested maximum waits of 1, 5 and 10 ms, and every confirmed serving setting uses 5 ms. The maximum wait decides when a batch is ready to run; it does not limit how long a ready batch waits behind one that is already running.

How to choose#

  • Start with the batch cap equal to the worker count, as in every confirmed setting, and a short maximum wait.
  • Small caps leave workers idle and let queues grow. Lower the cap only when measurements show the queue staying flat.
  • Tune for the operation mix you expect. An update-only load at 500 requests/s did not meet one second in every short run, although the balanced mix did.
  • Judge a setting by the fraction of offered requests accepted within your latency budget, the queue trend and memory use, not by accepted requests divided by the offered duration.

Deployment shape#

What was measured#

With the same total worker resources, 96 cores and 192 GiB, and the same frontend, two policies were compared at 1,000 requests/s in three 20-second runs each:

Deployment

How you split the cores matters.

2 × 48beat 1 × 96 at the same total resources
0 ms500 ms1 s1.5 s2 s2.5 s3 s284 ms282 ms330 msTwo 48-core workers2.75 s2.55 s2.74 sOne 96-core worker

Three 20-second runs per policy at 1,000 requests per second with the same total worker resources (96 cores, 192 GiB) and a separate 16-core frontend. Two workers with batch cap 48 each accepted every request within 500 ms; one worker with batch cap 96 built a queue. This compares deployment policies, not topology alone.

Policyp99 across the three runsAccepted within one second
Two 48-core workers, each with 48 workers and batch cap 48282.11–329.79 msAll requests
One 96-core worker with 96 workers and batch cap 962,545.17–2,751.96 ms27.075–32.31%; every request was eventually accepted while the queue grew

Worker CPU use was similar in all six runs, between 55.94 and 57.08 core equivalents. The comparison covers each policy as a whole, including its batch cap. It does not isolate the effect of topology, does not show that every 96-core setting is worse and does not confirm two workers over a longer period.

How to choose#

  • Compare candidate deployments at equal total resources and under the same offered load.
  • Expect similar CPU use to coexist with very different latency when batching and the serial part of the serving loop differ.
  • Confirm the shape you choose over a longer window before you rely on it.

Confirm a setting before relying on it#

The evaluation followed these steps, and they apply to your own service as well:

  1. Offer requests at a fixed rate on a schedule that does not wait for replies, and measure each latency from the request's scheduled time. A slow service then cannot reduce the load it receives by answering slowly.
  2. Count every offered request, and keep accepted, late, rejected and refused requests apart.
  3. Screen settings with short runs, then confirm the best one with a longer run. On the 192-core VM at 1,150 requests/s, three 15-second screens met one second for every request, and the 60-second confirmation met it for 70.37% of requests.
  4. Watch the queue length over time. A queue that keeps growing at a constant offered rate means the setting does not sustain that rate, whatever its average throughput.
  5. Report throughput over the completion window including the drain, together with the length of the window. Production integration lists what else to record.

Build settings#

The controlled measurements used Rust 1.96.1 with -Ctarget-cpu=native. Proof bytes do not depend on CPU-specific code generation: CI compares proof digests from a baseline x86-64 build with those from a native build, as described in Installation. The published measurements do not compare the speed of the two builds.

Checklist#

  • The depth covers the key space with room to grow, and depth and leaf bound stay fixed for each tree.
  • Memory is budgeted per job in flight and measured under the deployed limit.
  • The worker pool is sized explicitly, starting from one worker per available core.
  • The batch cap starts equal to the worker count, with a short maximum wait.
  • Candidate deployments are compared at equal total resources.
  • The final setting is confirmed with a long run on your own operation mix, with every offered request counted.

Next steps#