Guides
Production integration
Design a service around StateSync-GKR. Learn what the library provides and what your service adds, how requests and results stay matched, how to handle failures without weakening checks, what to record and what the measured deployments show.
StateSync-GKR is a library. A service that proves or verifies state operations adds request intake, state access, scheduling, transport, storage and monitoring around it. This guide describes the boundary between the library and your service, and the design decisions on your side of that boundary.
Before you begin#
- Read Trust boundaries. It states what a
trueverification result establishes and what your application must provide. - Work through Prepared execution, Batching and parallelism and Encoding and transport. This page builds on them and does not repeat their code.
- Production use that generates proofs, embeds or redistributes the proving functionality in a product or service, or offers proof generation to third parties requires a commercial license. A proving service needs one; see StateSync-GKR licensing.
What the library provides and what your service adds#
| Concern | StateSync-GKR provides | Your service adds |
|---|---|---|
| Circuits | prepare and PreparedSync | One preparation per operation kind and configuration, built at startup |
| Proving | prove_sync_op_prepared, make_job_prepared and prove_batch_parallel | Worker pools, the batching policy, admission control and timeouts |
| Verification | verify_sync_op_prepared and verify_encoded_sync_op | The list of configurations you accept and the handling of false |
| Transport | encode_sync_result and the inner-proof-v1 format | Framing, request identifiers, size limits and retransmission |
| State | compiler::smt_valid_native for a native check of one request | Authenticated roots, witnesses, the order of updates and every change to stored state |
| Scheduling | batching::WitnessQueue, DeadlineScheduler and BatchProver as building blocks | The serving loop, queue bounds and overload handling |
| Observability | Nothing: the library writes no logs and emits no metrics | Logs, metrics and traces |
The facade also defines OssCoreInterface, a minimal contract between a host system and the prover: next_request returns the next pending request and deliver hands back a result. MockOssCore implements it in memory for tests. The contract has no request identifiers, queue bounds, backpressure, priorities or retries, so a service adds those on its own side, and it is not one of the interfaces that the release line keeps fixed.
Follow one request through the service#
- Admit. Assign an identifier, check the request against your bounds, then queue it or refuse it. Record the admission outcome.
- Build. Authenticate the root, read the leaf and the sibling path from your state store and build the
SyncRequest. - Prove. Select the preparation by the request's operation kind, build the job and prove it on your worker pool.
- Deliver. Encode the result and send it with the identifier.
- Verify. Check the received bytes against the original request with
verify_encoded_sync_op. - Record and apply. Record the outcome and its timings. Apply any state change that depends on the proof, such as adopting the new root of an update, only after verification returns
true.
Admit and validate requests#
- Fix the tree depth, leaf bound and layer strategy at deployment. Requests carry data for a configuration; they never select one. Configuration values do not limit resource use, so choose them before you accept requests from untrusted sources (Configuration).
- Bound every queue. When a bound is reached, refuse the request and record the refusal as an outcome of its own.
- Check that each request's operation kind matches the preparation you will use. The prepared proving calls check it only in debug builds.
- Authenticate every root against your trusted source before you build a request. A root recomputed from caller-supplied siblings shows only that the siblings are consistent with that root.
- Turn away malformed requests before proving with
compiler::smt_valid_native, as shown in Proving state operations. It does not replace verification. - Apply your own transition rules to updates before proving. The engine proves any pair of old and new leaf states.
- Bound the size of the proof messages you receive (Encoding and transport).
Prepared state in a long-running service#
- Build one
PreparedSyncper operation kind and configuration at startup, share it across workers, and rebuild it when the configuration or the source revision changes. Prepared execution has the details. - Log the circuit commitment of each preparation, which
circuit_commitment()returns. It identifies the exact circuit behind every proof made or checked with that preparation. For a stable 32-byte form in logs, pass it towrap::prepared::circuit_commitment_bytes, which returns its canonical little-endian bytes. - A separate verifier service builds its own preparations from the same configuration. Proofs made under any other configuration fail its identity comparison.
Keep requests and results matched#
ProveJob,SyncResultand the encoded bytes carry no request identifier. Keep your own mapping from admission until a verdict exists.- Within one process, map job positions back to requests, as in Batching and parallelism. Between processes, carry the identifier in your envelope and verify each reply against the request it names, as in Encoding and transport.
- Proving is deterministic, so a second proof for the same request and configuration has the same bytes as the first. Treat a repeated delivery for an identifier you have already settled as a duplicate.
Errors and async runtimes#
- Every error type implements
DebugandClone, and most do not implementstd::error::Error(Errors). Map them into your service's own error type by matching on the variant, instead of relying on?conversion into a generic error type. - Proving and verification are synchronous and CPU-bound. In an async service, run them on a dedicated pool, such as your Rayon pool or a blocking task, so they do not stall the async executor.
Handle failures without weakening checks#
| Failure | What the library reports | Action |
|---|---|---|
| The configuration does not compile | SyncError::Compile from prepare, prove_sync_op or make_job | Fail startup and fix the configuration; a retry returns the same error |
| A malformed request: a path of the wrong length or an oversized leaf encoding | SyncError::Witness | Reject that request; other requests are unaffected |
| A request that does not hold, a key out of range or a mismatched public input | Proving still succeeds; verification returns false | Reject the request. A successful proving call is never acceptance |
| Received bytes fail decoding or the identity comparison | false from verify_encoded_sync_op | Reject the bytes and check the sender's configuration and transport. Never change the expected identity to match |
| A proof is rejected | false | Final for those bytes. Do not weaken the verifier and do not retry the same input in a loop |
| Encoding fails | EncodeError | CountOverflow means a leaf bound above 65,535; the other variants mean the proof did not come unchanged from this prover |
prove_batch returns fewer proofs than jobs | A short vector | The configuration does not compile; compare counts before pairing proofs with jobs |
| A tampered or mismatched proof is accepted | true where false was expected | Stop and report it through the Security policy |
| A configuration value at or above the field order, or, in a debug build, a prepared call given a request of another operation kind | A panic | Validate the configuration at startup and operation kinds at admission. Isolate proving so that a panic ends one worker, not the service |
Record the source revision, the configuration and the error class with every failure; Errors lists every variant. A type that places each offered request in exactly one outcome class keeps these rules visible in code and gives your success fractions a complete denominator:
use statesync_gkr::SyncError;
/// The final outcome of one offered request. Every offered request ends in
/// exactly one outcome, so success fractions can count all offers.
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub enum Outcome {
/// The verifier returned `true` within the latency budget.
Accepted,
/// The verifier returned `true` after the latency budget.
AcceptedLate,
/// The verifier returned `false`. Final for these bytes.
Rejected,
/// The service could not produce a proof.
NotProved(Cause),
/// Refused before any work started, for example because a bounded queue was full.
Refused,
}
/// Why a request could not be proved.
#[derive(Clone, Copy, Debug, PartialEq, Eq)]
pub enum Cause {
/// The configuration does not compile. Fix it; a retry fails the same way.
Configuration,
/// The request is malformed. Reject it; other requests are unaffected.
Request,
}
/// Map a proving error to its cause.
pub fn cause_of(error: &SyncError) -> Cause {
match error {
SyncError::Compile(_) => Cause::Configuration,
SyncError::Witness(_) => Cause::Request,
}
}Record what happened#
| Record | Why it matters |
|---|---|
| Request identifier, operation kind and configuration: depth, leaf bound and strategy | Ties each outcome to one circuit |
| Source revision and the preparation's circuit commitment | Identifies the code and the circuit that produced or checked a proof |
| Each root and the source that authenticated it | Root provenance is your responsibility |
| Scheduled or received time, admission, batch start and end, encoding, sending, receipt and verification | Separates queueing from computation and transport |
| The outcome class of every offered request | Success fractions need every offered request in the denominator |
| Batch sizes and queue length per operation kind over time | A queue that keeps growing at a constant offered rate means the setting does not sustain that rate |
| Worker CPU use and process memory | Capacity planning. Spare average CPU is a reason to examine the serving path, not a proportional capacity reserve |
| Encoded proof size | Transport and storage sizing |
Measure each latency on one monotonic clock, from the request's scheduled or received time to the verifier's result, and do not subtract timestamps taken on different hosts to estimate network time. When you report throughput, compute rates over the completion window including the drain after input stops. Accepted requests divided by the scheduled duration can hide a growing queue.
Root provenance and state changes#
StateSync-GKR does not authenticate roots, commit state, order updates or reach consensus, and the measured serving profiles end at the acceptance of the inner proof. Your service owns these steps:
- Record which trusted source supplied each root, and when.
- Order the updates to one tree yourself, and build each update against the root that the previous update produced.
- Adopt a new root, or make any other state change that depends on a proof, only after the verifier accepts that proof.
Choose a deployment shape#
The serving measurements cited on this page come from the benchmark serving harness built for the design and evaluation paper, which delivered every proof over TCP to a separate verifying frontend. StateSync-GKR is a library, so your service supplies the network layer. The four observations below all used a balanced mix of membership, non-membership and update requests at depth 24, periodic arrivals and a 5-ms maximum batch wait:
| Worker deployment | Offered load | Workers and batch cap | Result | Period |
|---|---|---|---|---|
| One 48-core VM, 96 GiB | 500 requests/s | 48 and 48 | All 450,000 timed requests accepted within 500 ms | One 15-minute run |
| Two 48-core VMs, 96 GiB each | 1,000 requests/s in total | 48 and 48 on each | Every timed request accepted within 500 ms | Three 20-second runs |
| One 192-core VM, 384 GiB | 1,000 requests/s | 192 and 192 | Every timed request accepted within one second | Three 60-second runs |
| One 192-core VM, 384 GiB | 1,150 requests/s | 192 and 192 | All 69,000 timed requests accepted within two seconds; 70.37% within one second while the queue grew | One 60-second input window |
Benchmark results gives the full distributions. What these runs show about deployment:
- Compare deployments at equal total resources. At 1,000 requests/s, two 48-core workers kept every request within 500 ms, while one 96-core worker with batch cap 96 built up a queue. The comparison covers each policy as a whole; Performance tuning has the numbers and their limits.
- Worker count and batch cap decide queueing. On the 48-core VM at 500 requests/s, too few workers or too small a batch cap turned sub-second latencies into p99 values of several seconds. Tune both on your own workload.
- A short pass is not a confirmation. At 1,150 requests/s, three 15-second screens met one second for every request, and the 60-second run did not. Confirm a setting over a longer window before you rely on it.
- A finite run is not sustained capacity. The 1,150 requests/s run completed 1,123.89 requests per second over its completion window, including the drain after input stopped, and its queue grew throughout, so it does not establish a rate that the deployment can sustain.
- The return path matters. In the measured caller, each worker ran one compute batch at a time and hashed, framed, wrote and logged its results before it chose the next batch. At 1,150 requests/s, most of the added delay was waiting before computation began: p99 of 1,367.97 ms, against 343.97–354.87 ms at 1,000 requests/s, while the worker averaged 78.09 of its 192 core equivalents. The study did not separate how much each step of the return path contributed. Design your service so that delivering one batch's results does not hold back the next batch, and measure the effect yourself, because an overlapped design was not part of the measurements.
- Your traffic is not the benchmark's. The benchmark cycled 192 fixed fixtures per operation kind with periodic arrivals. Random or bursty traffic and other operation mixes need their own measurements; an update-only load at 500 requests/s did not meet one second in every short run.
Checklist#
- Roots are authenticated against a trusted source before requests are built.
- Every queue is bounded, and refusals are recorded as their own outcome.
- One preparation exists per operation kind and accepted configuration, and its circuit commitment is logged.
- Every request keeps an identifier and its original form until a verdict exists.
- Only a
trueverification result counts as acceptance, and state changes wait for it. - Configuration errors stop the service; request errors reject one request;
falseis final for the bytes. - Every offered request ends in exactly one recorded outcome, with timings on one clock.
- Deployment settings are confirmed with a long run on your own traffic before you rely on them.
Next steps#
- Performance tuning: choose depth, memory budget, workers, batch caps and deployment shape.
- Troubleshooting: diagnose failed proving calls and rejected proofs.
- Trust boundaries: what a verification result does and does not establish.
- Benchmark methodology: how the serving studies were run and measured.