Guides

Production integration

Design a service around StateSync-GKR. Learn what the library provides and what your service adds, how requests and results stay matched, how to handle failures without weakening checks, what to record and what the measured deployments show.

StateSync-GKR is a library. A service that proves or verifies state operations adds request intake, state access, scheduling, transport, storage and monitoring around it. This guide describes the boundary between the library and your service, and the design decisions on your side of that boundary.

Before you begin#

What the library provides and what your service adds#

ConcernStateSync-GKR providesYour service adds
Circuitsprepare and PreparedSyncOne preparation per operation kind and configuration, built at startup
Provingprove_sync_op_prepared, make_job_prepared and prove_batch_parallelWorker pools, the batching policy, admission control and timeouts
Verificationverify_sync_op_prepared and verify_encoded_sync_opThe list of configurations you accept and the handling of false
Transportencode_sync_result and the inner-proof-v1 formatFraming, request identifiers, size limits and retransmission
Statecompiler::smt_valid_native for a native check of one requestAuthenticated roots, witnesses, the order of updates and every change to stored state
Schedulingbatching::WitnessQueue, DeadlineScheduler and BatchProver as building blocksThe serving loop, queue bounds and overload handling
ObservabilityNothing: the library writes no logs and emits no metricsLogs, metrics and traces

The facade also defines OssCoreInterface, a minimal contract between a host system and the prover: next_request returns the next pending request and deliver hands back a result. MockOssCore implements it in memory for tests. The contract has no request identifiers, queue bounds, backpressure, priorities or retries, so a service adds those on its own side, and it is not one of the interfaces that the release line keeps fixed.

Follow one request through the service#

  1. Admit. Assign an identifier, check the request against your bounds, then queue it or refuse it. Record the admission outcome.
  2. Build. Authenticate the root, read the leaf and the sibling path from your state store and build the SyncRequest.
  3. Prove. Select the preparation by the request's operation kind, build the job and prove it on your worker pool.
  4. Deliver. Encode the result and send it with the identifier.
  5. Verify. Check the received bytes against the original request with verify_encoded_sync_op.
  6. Record and apply. Record the outcome and its timings. Apply any state change that depends on the proof, such as adopting the new root of an update, only after verification returns true.

Admit and validate requests#

  • Fix the tree depth, leaf bound and layer strategy at deployment. Requests carry data for a configuration; they never select one. Configuration values do not limit resource use, so choose them before you accept requests from untrusted sources (Configuration).
  • Bound every queue. When a bound is reached, refuse the request and record the refusal as an outcome of its own.
  • Check that each request's operation kind matches the preparation you will use. The prepared proving calls check it only in debug builds.
  • Authenticate every root against your trusted source before you build a request. A root recomputed from caller-supplied siblings shows only that the siblings are consistent with that root.
  • Turn away malformed requests before proving with compiler::smt_valid_native, as shown in Proving state operations. It does not replace verification.
  • Apply your own transition rules to updates before proving. The engine proves any pair of old and new leaf states.
  • Bound the size of the proof messages you receive (Encoding and transport).

Prepared state in a long-running service#

  • Build one PreparedSync per operation kind and configuration at startup, share it across workers, and rebuild it when the configuration or the source revision changes. Prepared execution has the details.
  • Log the circuit commitment of each preparation, which circuit_commitment() returns. It identifies the exact circuit behind every proof made or checked with that preparation. For a stable 32-byte form in logs, pass it to wrap::prepared::circuit_commitment_bytes, which returns its canonical little-endian bytes.
  • A separate verifier service builds its own preparations from the same configuration. Proofs made under any other configuration fail its identity comparison.

Keep requests and results matched#

  • ProveJob, SyncResult and the encoded bytes carry no request identifier. Keep your own mapping from admission until a verdict exists.
  • Within one process, map job positions back to requests, as in Batching and parallelism. Between processes, carry the identifier in your envelope and verify each reply against the request it names, as in Encoding and transport.
  • Proving is deterministic, so a second proof for the same request and configuration has the same bytes as the first. Treat a repeated delivery for an identifier you have already settled as a duplicate.

Errors and async runtimes#

  • Every error type implements Debug and Clone, and most do not implement std::error::Error (Errors). Map them into your service's own error type by matching on the variant, instead of relying on ? conversion into a generic error type.
  • Proving and verification are synchronous and CPU-bound. In an async service, run them on a dedicated pool, such as your Rayon pool or a blocking task, so they do not stall the async executor.

Handle failures without weakening checks#

FailureWhat the library reportsAction
The configuration does not compileSyncError::Compile from prepare, prove_sync_op or make_jobFail startup and fix the configuration; a retry returns the same error
A malformed request: a path of the wrong length or an oversized leaf encodingSyncError::WitnessReject that request; other requests are unaffected
A request that does not hold, a key out of range or a mismatched public inputProving still succeeds; verification returns falseReject the request. A successful proving call is never acceptance
Received bytes fail decoding or the identity comparisonfalse from verify_encoded_sync_opReject the bytes and check the sender's configuration and transport. Never change the expected identity to match
A proof is rejectedfalseFinal for those bytes. Do not weaken the verifier and do not retry the same input in a loop
Encoding failsEncodeErrorCountOverflow means a leaf bound above 65,535; the other variants mean the proof did not come unchanged from this prover
prove_batch returns fewer proofs than jobsA short vectorThe configuration does not compile; compare counts before pairing proofs with jobs
A tampered or mismatched proof is acceptedtrue where false was expectedStop and report it through the Security policy
A configuration value at or above the field order, or, in a debug build, a prepared call given a request of another operation kindA panicValidate the configuration at startup and operation kinds at admission. Isolate proving so that a panic ends one worker, not the service

Record the source revision, the configuration and the error class with every failure; Errors lists every variant. A type that places each offered request in exactly one outcome class keeps these rules visible in code and gives your success fractions a complete denominator:

src/outcome.rsRust
use statesync_gkr::SyncError;

/// The final outcome of one offered request. Every offered request ends in
/// exactly one outcome, so success fractions can count all offers.
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub enum Outcome {
    /// The verifier returned `true` within the latency budget.
    Accepted,
    /// The verifier returned `true` after the latency budget.
    AcceptedLate,
    /// The verifier returned `false`. Final for these bytes.
    Rejected,
    /// The service could not produce a proof.
    NotProved(Cause),
    /// Refused before any work started, for example because a bounded queue was full.
    Refused,
}

/// Why a request could not be proved.
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub enum Cause {
    /// The configuration does not compile. Fix it; a retry fails the same way.
    Configuration,
    /// The request is malformed. Reject it; other requests are unaffected.
    Request,
}

/// Map a proving error to its cause.
pub fn cause_of(error: &SyncError) -> Cause {
    match error {
        SyncError::Compile(_) => Cause::Configuration,
        SyncError::Witness(_) => Cause::Request,
    }
}

Record what happened#

RecordWhy it matters
Request identifier, operation kind and configuration: depth, leaf bound and strategyTies each outcome to one circuit
Source revision and the preparation's circuit commitmentIdentifies the code and the circuit that produced or checked a proof
Each root and the source that authenticated itRoot provenance is your responsibility
Scheduled or received time, admission, batch start and end, encoding, sending, receipt and verificationSeparates queueing from computation and transport
The outcome class of every offered requestSuccess fractions need every offered request in the denominator
Batch sizes and queue length per operation kind over timeA queue that keeps growing at a constant offered rate means the setting does not sustain that rate
Worker CPU use and process memoryCapacity planning. Spare average CPU is a reason to examine the serving path, not a proportional capacity reserve
Encoded proof sizeTransport and storage sizing

Measure each latency on one monotonic clock, from the request's scheduled or received time to the verifier's result, and do not subtract timestamps taken on different hosts to estimate network time. When you report throughput, compute rates over the completion window including the drain after input stops. Accepted requests divided by the scheduled duration can hide a growing queue.

Root provenance and state changes#

StateSync-GKR does not authenticate roots, commit state, order updates or reach consensus, and the measured serving profiles end at the acceptance of the inner proof. Your service owns these steps:

  • Record which trusted source supplied each root, and when.
  • Order the updates to one tree yourself, and build each update against the root that the previous update produced.
  • Adopt a new root, or make any other state change that depends on a proof, only after the verifier accepts that proof.

Choose a deployment shape#

The serving measurements cited on this page come from the benchmark serving harness built for the design and evaluation paper, which delivered every proof over TCP to a separate verifying frontend. StateSync-GKR is a library, so your service supplies the network layer. The four observations below all used a balanced mix of membership, non-membership and update requests at depth 24, periodic arrivals and a 5-ms maximum batch wait:

Worker deploymentOffered loadWorkers and batch capResultPeriod
One 48-core VM, 96 GiB500 requests/s48 and 48All 450,000 timed requests accepted within 500 msOne 15-minute run
Two 48-core VMs, 96 GiB each1,000 requests/s in total48 and 48 on eachEvery timed request accepted within 500 msThree 20-second runs
One 192-core VM, 384 GiB1,000 requests/s192 and 192Every timed request accepted within one secondThree 60-second runs
One 192-core VM, 384 GiB1,150 requests/s192 and 192All 69,000 timed requests accepted within two seconds; 70.37% within one second while the queue grewOne 60-second input window

Benchmark results gives the full distributions. What these runs show about deployment:

  • Compare deployments at equal total resources. At 1,000 requests/s, two 48-core workers kept every request within 500 ms, while one 96-core worker with batch cap 96 built up a queue. The comparison covers each policy as a whole; Performance tuning has the numbers and their limits.
  • Worker count and batch cap decide queueing. On the 48-core VM at 500 requests/s, too few workers or too small a batch cap turned sub-second latencies into p99 values of several seconds. Tune both on your own workload.
  • A short pass is not a confirmation. At 1,150 requests/s, three 15-second screens met one second for every request, and the 60-second run did not. Confirm a setting over a longer window before you rely on it.
  • A finite run is not sustained capacity. The 1,150 requests/s run completed 1,123.89 requests per second over its completion window, including the drain after input stopped, and its queue grew throughout, so it does not establish a rate that the deployment can sustain.
  • The return path matters. In the measured caller, each worker ran one compute batch at a time and hashed, framed, wrote and logged its results before it chose the next batch. At 1,150 requests/s, most of the added delay was waiting before computation began: p99 of 1,367.97 ms, against 343.97–354.87 ms at 1,000 requests/s, while the worker averaged 78.09 of its 192 core equivalents. The study did not separate how much each step of the return path contributed. Design your service so that delivering one batch's results does not hold back the next batch, and measure the effect yourself, because an overlapped design was not part of the measurements.
  • Your traffic is not the benchmark's. The benchmark cycled 192 fixed fixtures per operation kind with periodic arrivals. Random or bursty traffic and other operation mixes need their own measurements; an update-only load at 500 requests/s did not meet one second in every short run.

Checklist#

  • Roots are authenticated against a trusted source before requests are built.
  • Every queue is bounded, and refusals are recorded as their own outcome.
  • One preparation exists per operation kind and accepted configuration, and its circuit commitment is logged.
  • Every request keeps an identifier and its original form until a verdict exists.
  • Only a true verification result counts as acceptance, and state changes wait for it.
  • Configuration errors stop the service; request errors reject one request; false is final for the bytes.
  • Every offered request ends in exactly one recorded outcome, with timings on one clock.
  • Deployment settings are confirmed with a long run on your own traffic before you rely on them.

Next steps#