Guides

Batching and parallelism

Prove many independent requests on a Rayon pool that you size, keep every proof matched to its own request, and know exactly what a batch does and does not mean in StateSync-GKR.

StateSync-GKR speeds up large request volumes by running independent proving jobs in parallel over one prepared circuit. This guide shows the batch API, a complete pipeline that builds witnesses, proves, encodes and verifies on a worker pool you control, and the rules that keep each proof tied to the request it belongs to.

Before you begin#

  • Read Prepared execution. Every batch function on this page except make_job and prove_batch takes a PreparedSync.
  • Add rayon next to statesync-gkr so that you can create and size a thread pool:
Cargo.tomlTOML
[package]
name = "my-batch-prover"
version = "0.1.0"
edition = "2024"
publish = false

[dependencies]
statesync-gkr = { git = "https://github.com/Oraclizer/statesync-gkr.git", branch = "main" }
rayon = "1.10"

The facade declares rayon = "1.10" for its parallel batch API. Cargo links only one rayon-core into a build, so your pool and the facade's parallel calls share one Rayon runtime. A rayon requirement that cannot unify with the facade's 1.10 fails during dependency resolution.

What a batch is#

In StateSync-GKR, a batch is a set of independent jobs of one operation kind that share one prepared circuit. Each job keeps its own witness, transcript and proof, and the proof for job i is byte for byte the proof that the single-request path would produce for that request. That definition has four consequences:

  • One proof per request. A batch does not merge statements into an aggregate proof, and verification still runs once per proof.
  • One kind per batch. Jobs share a compiled circuit, so membership, non-membership and update jobs are batched separately.
  • Independent outcomes. A malformed witness fails its own request, and the other jobs proceed.
  • No ordering between jobs. Two updates in one batch are proved independently. The engine neither applies one before the other nor commits either one to a store, and an atomic change to several leaves is not provided. To prove a sequence of updates, build each request against the root that the previous one produced.

The batch API#

FunctionCircuit setup on each callRuns onReturns
make_job(&request)Compiles the circuitThe calling threadResult<ProveJob, SyncError>
make_job_prepared(&prepared, &request)NoneThe calling thread; parallelize it yourselfResult<ProveJob, SyncError>
prove_batch(kind, &jobs)Compiles the circuit and computes its commitment onceThe calling thread, one job after anotherVec<GkrProof>, empty if compilation fails
prove_batch_prepared(&prepared, &jobs)NoneThe calling thread, one job after anotherVec<GkrProof> in job order
prove_batch_parallel(&prepared, &jobs)NoneThe current Rayon pool; requires the default host featureVec<GkrProof> in job order

All of these are methods of StateSyncProver. A ProveJob (from statesync_gkr::batching) holds the operation kind, the request's public inputs and the complete circuit witness. prove_batch_parallel returns exactly what prove_batch_prepared returns for the same jobs, element for element, whatever the worker count.

Three checks belong in your code, because these functions do not make them:

  • make_job_prepared and prove_sync_op_prepared check that the request's kind matches the preparation only in debug builds. Compare request.operation.kind() with prepared.kind() before you build jobs.
  • prove_batch_prepared and prove_batch_parallel do not compare each job's kind with the preparation. Pass only jobs built by make_job_prepared from that same PreparedSync.
  • prove_batch returns an empty vector when the prover's configuration cannot be compiled. Compare the number of proofs with the number of jobs.

Size the worker pool#

prove_batch_parallel runs on the Rayon pool that is current when you call it. Inside pool.install, that is your pool; anywhere else it is Rayon's global pool, whose size Rayon chooses. Create the pool yourself, so that the worker count is a decision you make and record:

  • Build the pool once and reuse it for every batch. Building a pool starts its threads.
  • Each running job owns its own prover state, so memory use grows with the number of jobs in flight.
  • Small batches can leave workers idle. The serving configurations confirmed in the controlled studies set the worker count equal to the batch cap (48 and 48 on a 48-core worker, 192 and 192 on a 192-core VM) with a 5-ms maximum batch wait. Those settings belong to the benchmark's own serving caller, not to a built-in service; Performance tuning explains how to choose your own.
Parallel proving

Independent proofs scale with the cores you give them.

3,207.74membership proofs per second on 192 workers
05001,0001,5002,0002,5003,0003,500124816324896192Worker threads on one 192-core hostProofs per second

Prepared proving only at depth 24, batch 768, three processes per point on a two-socket bare-metal host with SMT off. Witness generation, encoding and transport are excluded. Worker counts are thread limits on one machine, not smaller machines.

The figure shows prepared proving-only throughput on one host as the worker count grows. Witness creation, encoding and transport are outside those measurements.

A complete batch pipeline#

The program below proves 16 membership requests in three parallel stages on a four-worker pool: witness construction, proving, and encoding with verification of the encoded bytes. Each request carries its own fixture root, like the benchmark fixtures; a batch does not require its requests to share a root.

src/main.rsRust
use std::process::ExitCode;

use rayon::prelude::*;
use statesync_gkr::batching::ProveJob;
use statesync_gkr::compiler::{
    AssetId, LayerStrategy, LeafPayload, LeafState, MerklePath, PublicInputs, SmtOpKind,
    SmtOperation, SmtParams, SmtWitness,
};
use statesync_gkr::primitives::field::{BaseField, PrimeCharacteristicRing};
use statesync_gkr::primitives::hash::{Digest, HashGadget, Poseidon2Gadget};
use statesync_gkr::{StateSyncGkrConfig, StateSyncProver, SyncError, SyncRequest, SyncResult};

const DEPTH: u32 = 4;
const REQUESTS: u32 = 16;
const WORKERS: usize = 4;

fn field(value: u32) -> BaseField {
    BaseField::from_u32(value)
}

/// A self-contained membership request with its own fixture root.
fn membership_request(params: &SmtParams, index: u32) -> Result<SyncRequest, String> {
    let hasher = Poseidon2Gadget::new(params.leaf_max_fields as usize);
    let key = AssetId(u64::from(index));
    let payload = LeafPayload {
        sync_state: vec![field(100 + index), field(7)],
        identity_digest: [9_u8; 32],
    };
    let leaf = LeafState::Occupied(payload.clone());
    let path = MerklePath {
        siblings: (0..params.depth)
            .map(|level| Digest([field(level * 13 + index + 1); 8]))
            .collect(),
    };
    let root = path
        .compute_root(&hasher, params, key, &leaf)
        .map_err(|error| format!("root construction failed: {error:?}"))?;
    let value_digest = hasher
        .hash_leaf(&leaf.encode())
        .map_err(|error| format!("leaf hashing failed: {error:?}"))?;
    let operation = SmtOperation::Membership { key, payload };
    let op_kind_tag = PublicInputs::kind_tag(operation.kind());
    Ok(SyncRequest {
        operation,
        witness: SmtWitness { leaf, path },
        public_inputs: PublicInputs {
            old_root: root,
            new_root: root,
            op_kind_tag,
            asset_id: key,
            value_digest,
        },
    })
}

fn run() -> Result<(), String> {
    let params = SmtParams {
        depth: DEPTH,
        ..Default::default()
    };
    let prover = StateSyncProver::new(StateSyncGkrConfig {
        smt: params,
        layer_strategy: LayerStrategy::A,
        ..Default::default()
    });
    // One preparation serves every job of this kind.
    let prepared = prover
        .prepare(SmtOpKind::Membership)
        .map_err(|error| format!("preparation failed: {error:?}"))?;

    let requests = (0..REQUESTS)
        .map(|index| membership_request(&params, index))
        .collect::<Result<Vec<_>, _>>()?;

    // The prepared calls check the kind only in debug builds, so check it here.
    if let Some(index) = requests
        .iter()
        .position(|request| request.operation.kind() != prepared.kind())
    {
        return Err(format!("request {index} does not match the prepared kind"));
    }

    // Build the pool once and reuse it for every batch.
    let pool = rayon::ThreadPoolBuilder::new()
        .num_threads(WORKERS)
        .build()
        .map_err(|error| format!("thread pool construction failed: {error:?}"))?;

    // Stage 1: build witnesses in parallel. The collected vector keeps input order.
    let built = pool.install(|| {
        requests
            .par_iter()
            .map(|request| prover.make_job_prepared(&prepared, request))
            .collect::<Vec<Result<ProveJob, SyncError>>>()
    });

    // A failed witness affects only its own request. `owners[i]` is the index
    // of the request that job `i` belongs to.
    let mut owners = Vec::with_capacity(built.len());
    let mut jobs = Vec::with_capacity(built.len());
    for (index, outcome) in built.into_iter().enumerate() {
        match outcome {
            Ok(job) => {
                owners.push(index);
                jobs.push(job);
            }
            Err(error) => eprintln!("request {index}: witness construction failed: {error:?}"),
        }
    }

    // Stage 2: one independent proof per job, returned in job order.
    let proofs = pool.install(|| prover.prove_batch_parallel(&prepared, &jobs));
    if proofs.len() != jobs.len() {
        return Err("the prover returned a different number of proofs than jobs".to_owned());
    }

    // Stage 3: encode each proof and verify the bytes against the request it belongs to.
    let outcomes = pool.install(|| {
        owners
            .par_iter()
            .zip(proofs)
            .map(|(&index, proof)| -> Result<Vec<u8>, String> {
                let request = &requests[index];
                let result = SyncResult {
                    public_inputs: request.public_inputs.clone(),
                    proof,
                };
                let bytes = prover
                    .encode_sync_result(&prepared, &result)
                    .map_err(|error| format!("request {index}: encoding failed: {error:?}"))?;
                if !prover.verify_encoded_sync_op(&prepared, request, &bytes) {
                    return Err(format!("request {index}: encoded proof rejected"));
                }
                Ok(bytes)
            })
            .collect::<Vec<_>>()
    });

    let mut accepted = 0_usize;
    for outcome in outcomes {
        match outcome {
            Ok(_encoded) => accepted += 1,
            Err(error) => eprintln!("{error}"),
        }
    }
    println!("accepted={accepted} requests={}", requests.len());
    if accepted == requests.len() {
        Ok(())
    } else {
        Err("some requests were not accepted".to_owned())
    }
}

fn main() -> ExitCode {
    match run() {
        Ok(()) => ExitCode::SUCCESS,
        Err(error) => {
            eprintln!("batch=FAIL: {error}");
            ExitCode::FAILURE
        }
    }
}

With all 16 requests valid, the program prints accepted=16 requests=16.

The witness and output stages are caller code built on the public API. The facade has no built-in parallel witness builder or parallel encoder, which is why the program parallelizes those stages itself with par_iter. Stage 3 encodes and verifies in one process to keep the example self-contained; in a deployment, the prover encodes and a separate verifier decodes and checks the bytes. The encoded bytes carry the circuit identity, the public inputs and the proof, but not the private witness, which the verifier still needs; see Encoding and transport.

Keep results matched to requests#

  • prove_batch_prepared and prove_batch_parallel return proofs in job order: proofs[i] belongs to jobs[i]. If you skip failed requests while building jobs, keep a separate map from each job to its request, as owners does above.
  • A ProveJob records the operation kind and the public inputs, but not your request identifier. Keep your own mapping.
  • Pair each proof with the public inputs of its own request in a SyncResult, and verify it against that same request. The verifier rejects a result whose public inputs differ from the request's.
  • Count a request as accepted only after verify_sync_op_prepared or verify_encoded_sync_op returns true for that request. A hash of the proof bytes is not verification.
  • Between processes or machines, replies can arrive out of order. Give every request an explicit identifier and match replies by it. The measured serving caller bound each reply to its original request by request identifier, fixture identifier and operation kind.

Queue and scheduling building blocks#

statesync_gkr::batching also contains scheduling structure that you can use or replace:

  • WitnessQueue keeps one first-in, first-out lane per operation kind, with push, len, is_empty and drain.
  • DeadlineScheduler decides a batch size from a lane's length and the remaining deadline budget, using a BatchPolicy. Below min_batch_size it takes everything queued. Otherwise it caps the size at max_batch_size, and when less than the full deadline remains it scales the size down in proportion, but not below min_batch_size. The defaults (a cap of 32, a 200-ms deadline and a minimum of 2) are marked provisional in the source and are not tuned settings.
  • BatchProver::prove_round(&mut queue, remaining, prove_batch) visits each operation kind in turn, decides a size, drains that many jobs and passes them to your closure as prove_batch(kind, &jobs). Jobs it does not drain stay queued.

prove_round stores only the proofs your closure returns, in a BatchResult, without recording which jobs produced them. If you use it, select the preparation for the kind your closure receives and record the drained jobs inside the closure so that you can match proofs to requests. The batching field of StateSyncGkrConfig holds a BatchPolicy for your scheduler; the prover's own methods do not read it.

Measured effect#

At depth 24, with batches of 768 requests and 192 workers, processing the outputs in parallel instead of serially lowered the time of the complete prepared local batch by a factor of 9.71–10.71 across the three operations. The timed interval covers witness generation, proving, encoding and encoded verification. Both policies build witnesses and prove in parallel; only encoding and encoded verification change between serial and parallel. Preparation, transport and the hashing and logging after the timer are excluded. The range spans the three operation-specific medians across five processes, each process summarized by the median of ten paired serial and parallel interval ratios, so it is not a ratio of throughput medians. Smaller batches behave differently. Across the study, 327,360 pairs of serial and parallel proof bytes were compared and preserved the canonical result. The host had two AMD EPYC 9R45 sockets (192 CPUs, SMT off) and ran the 1.1.0 release source built with Rust 1.96.1 and -Ctarget-cpu=native.

Local batches

Parallel output keeps a full batch moving.

9.71–10.71×faster complete batch of 768 proofs
0 s1 s2 s3 s4 s5 s6 s7 s5.16 s0.48 sMembership10.71× faster5.33 s0.54 sNon-membership9.97× faster6.35 s0.66 sSingle-leaf update9.71× faster

Depth 24, batch of 768, 192 workers on a 192-core bare-metal host, five processes with ten paired batches each. Witness generation and proving are parallel in both modes; only encoding and encoded verification change. Each request keeps its own proof, nothing is aggregated, and transport is excluded.

Building witnesses in parallel matters as well. A separate, earlier campaign on a different host, at the same depth, batch size and worker count, measured membership at 1,060.35 inner proofs per second with witnesses built serially and 2,881.55 per second with the caller building them in parallel; both rates include proving. Non-membership and update show the same pattern. Benchmark results has the complete tables.

These are local computation rates. They do not include arrival, queueing, network transport or state commits, so they are not service throughput.

What the guarantees cover#

The formal model includes an independent-job batching result: each result in a batch agrees with its single-job counterpart, and acceptance and rejection are preserved. In the Rust code, every proving path funnels through one prove body, tests check that batch and parallel outputs agree with the single and sequential paths, and proof bytes do not depend on the worker count. None of this covers a single aggregate proof, state shared between jobs, scheduler liveness or a performance guarantee. Formal verification states the model result and its assumptions.

Common mistakes#

SymptomCauseFix
The worker count does not change throughputprove_batch_parallel runs outside pool.install and uses Rayon's global poolCall it inside install on a pool that you built
Proofs are verified against the wrong requestsFailed requests were skipped and the job index was used as the request indexKeep an explicit map from job to request
A batch returns no proofsprove_batch could not compile the configurationCompare proof and job counts; fix the configuration
A release build panics or produces rejected proofs for mixed requestsJobs of several kinds were built against one preparationGroup requests by kind and prepare each kind
Updates in one batch do not chainEach job was built against the same old rootBuild each update against the root produced by the previous one

Checklist#

  • One preparation per operation kind; each batch contains jobs of one kind only.
  • Request kinds are checked against prepared.kind() before jobs are built.
  • The Rayon pool is built once, sized deliberately and used through install.
  • Each proof is mapped back to its own request and verified against it.
  • Acceptance is counted only from the verifier's result.
  • Each update in a sequence is built against the root that the previous update produced.

Next steps#