Randomized Benchmarking
Short Definition
Section titled “Short Definition”Randomized benchmarking (RB) estimates an ensemble-average error parameter for implemented quantum gates by measuring how the survival probability of random identity circuits decays with sequence length. In standard reference RB, one samples gates from a finite group, appends the group inverse, executes the compiled circuit, and asks whether the final state agrees with the intended return state.
Under a baseline model of stationary, Markovian, gate-independent noise and a unitary 2-design such as the Clifford group, the sequence-averaged survival probability is
The constants and contain fixed state-preparation, measurement, and edge effects. The decay parameter determines the average infidelity of the twirled noise on a -dimensional computational space:
This statement is powerful but conditional. RB estimates an average associated with the sampled group, its compiler, its physical context, and its time window. It does not reconstruct a noise channel, identify an error mechanism, return a worst-case error, or automatically yield an error per native gate.
This page is the canonical home for standard reference RB, its twirl derivation, sequence and shot statistics, fit design, interpretation, interleaved RB, and the main failure modes. Metrics for Quantum Hardware owns comparisons among hardware metrics. Process Tomography owns channel reconstruction, while Noise in Quantum Information owns the physical taxonomy of coherent, stochastic, leakage, correlated, and non-Markovian noise. Cycle-level randomized protocols belong to the separate Cycle Benchmarking page.
What Reference RB Estimates
Section titled “What Reference RB Estimates”Let be the declared reference group. In a one-qubit experiment it is commonly the 24-element Clifford group; for several qubits, sampled multi-qubit Cliffords must be compiled into available native operations. A reference RB result is attached to the complete implementation map
Changing changes the benchmark even when the abstract group is unchanged. Whole-circuit optimization, cancellation, virtual-frame handling, randomized compilation, pulse overlap, spectator activity, and calibration versions are therefore part of the estimand.
The primary object is the mean survival over random sequences:
not the survival of one representative sequence. Here and is the probability generated by the implemented circuit for that particular draw.
Under the standard model, is an error per sampled reference element. It is often called error per Clifford when is a Clifford group. It is not automatically error per pulse, error per primitive gate, error per layer, or error for the target application.
The Protocol
Section titled “The Protocol”For each selected length :
-
Draw independently and uniformly from .
-
Compute the ideal inverse
-
Compile all group elements under a frozen compiler and calibration policy.
-
Prepare a declared state , execute the sequence, and measure a declared survival effect .
-
Repeat shots for the same sequence to estimate its survival probability.
-
Repeat with independently sampled sequences at the same .
-
Repeat across lengths and fit the sequence-averaged data.
The implemented sequence channel is
and its survival probability is
The exact order convention must be fixed in software and tested. An inverse computed with the wrong multiplication order can still produce plausible decay data while benchmarking a different circuit family.
At each length, many random sequences and repeated shots estimate an ensemble mean. The fitted decay is distinct from the SPAM-dependent contrast and asymptote ; residuals and between-sequence variation remain part of the evidence.
Twirling Under the Baseline Model
Section titled “Twirling Under the Baseline Model”Represent an ideal group element by the unitary channel
Assume each physical implementation has the same Markovian error channel :
This gate-independent model is not exact hardware physics; it is the clean starting point from which the standard decay law follows.
The group twirl of is
If is a unitary 2-design and is trace preserving, the twirled channel acts as a depolarizing channel on the computational space:
This follows because conjugation leaves the identity direction fixed and acts irreducibly on traceless operators. By Schur’s lemma, a channel commuting with every group conjugation must multiply the entire traceless sector by one scalar .
After averaging independent group draws, the randomized frames convert each interior occurrence of into the same twirl. Preparation, measurement, and a final edge channel can be collected into effective objects and , giving
Since
one obtains
with
The decay is therefore insensitive to the leading constant SPAM amplitude, not independent of all SPAM behavior.
From Decay to Average Infidelity
Section titled “From Decay to Average Infidelity”The average gate fidelity of a channel relative to the identity is
Twirling preserves this average. For ,
Hence
For a single qubit,
The conversion is exact for the baseline channel associated with the RB decay. Under gate-dependent noise, modern RB theory can still justify a dominant exponential and an ensemble-level average error under suitable conditions, but the result is not generally the arithmetic mean of separately defined infidelities for every compiled gate.
Worked Example: A Depolarizing Qubit
Section titled “Worked Example: A Depolarizing Qubit”Suppose the effective one-qubit depolarizing parameter is
With ideal preparation and measurement,
so
The reported RB infidelity is
At length ,
At ,
Long sequences amplify a small per-element decay into observable contrast. They also increase exposure to drift, leakage, heating, and memory. Sequence lengths should span the informative part of the curve rather than merely reach the longest executable circuit.
If SPAM changes the curve to and while remaining stationary, the same and are recovered in the ideal model. The lower contrast increases statistical uncertainty even though it does not change the central decay parameter.
Sequences and Shots Are Different Samples
Section titled “Sequences and Shots Are Different Samples”Let index independently drawn sequences at length , and let index repeated shots. Conditional on sequence survival ,
The per-sequence estimate is
and the unweighted sequence-ensemble mean for equal shot counts is
Suppose each sequence uses shots and
The law of total variance gives
The first term is between-sequence variation; the second is shot noise. Once shot noise is small, spending the remaining budget on more shots of the same few sequences does not characterize the random-sequence ensemble. Conversely, one shot on many sequences can be inefficient when readout noise dominates.
Sequence is the experimental unit for the randomization. A bootstrap that resamples only individual shots while holding the sampled circuits fixed omits the first variance term.
Designing Sequence Lengths
Section titled “Designing Sequence Lengths”Lengths should identify , , and rather than cluster where the curve is nearly flat. A practical design includes:
- short lengths to resolve initial contrast;
- lengths near the characteristic decay scale;
- long lengths that constrain the asymptote without driving all signal below readout resolution;
- repeated lengths throughout the run to reveal drift;
- more sequence draws where between-sequence variance is large.
For near one,
so the characteristic scale is approximately
This estimate guides an initial design; a pilot run can then refine the length grid. Selecting lengths after seeing favorable random fluctuations without accounting for that adaptation biases the final analysis.
Fitting the Decay
Section titled “Fitting the Decay”The familiar three-parameter nonlinear least-squares fit is not the only statistical model. A defensible analysis keeps the data hierarchy visible.
Likelihood and weighting
Section titled “Likelihood and weighting”At minimum, account for heteroskedasticity: survival probabilities near have different binomial variance from probabilities near one, and between-sequence variation can change with . Weighted nonlinear least squares can be useful if weights are estimated honestly. A binomial or hierarchical likelihood retains more of the acquisition model.
Do not first average away sequence identifiers and then treat the resulting point means as equally precise. Preserve counts, sequence seeds, and shots per sequence.
Uncertainty
Section titled “Uncertainty”Useful approaches include:
- hierarchical parametric bootstrap under a declared RB model;
- nonparametric resampling of sequences within each length, with shot resampling nested inside;
- profile-likelihood or Bayesian intervals with transparent priors;
- rigorous concentration bounds when their assumptions and constants fit the experiment.
Report uncertainty in and propagate it through . Systematic uncertainty from model failure, compiler variation, drift, or interleaved-RB assumptions is separate from fit uncertainty.
Residuals and alternatives
Section titled “Residuals and alternatives”Inspect residuals against sequence length, acquisition time, random-sequence identity, compiler statistics, and leakage flags. Compare the single exponential with justified alternatives, such as an additional decaying mode for leakage or a transient term for gate dependence. An alternative model is evidence only if it improves predictive behavior on held-out or replicated data, not merely in-sample fit.
What SPAM Robustness Means
Section titled “What SPAM Robustness Means”Fixed preparation and measurement errors change and while leaving the interior decay unchanged in the baseline derivation. This is the sense in which reference RB is robust to SPAM.
SPAM can still matter in several ways:
- poor preparation or readout reduces and makes hard to identify;
- reset quality can depend on sequence length or previous outcomes;
- measurement response can depend on leakage or heating accumulated during the sequence;
- the final inversion has sequence-dependent compilation and can correlate edge error with the sampled circuit;
- drift can make , , , and vary during acquisition;
- postselection can change the estimand.
Calling RB “SPAM-free” is therefore incorrect. It is insensitive to a particular stationary nuisance structure under a specified model.
Gate-Dependent Noise
Section titled “Gate-Dependent Noise”Real implementations have
with error depending on , its native decomposition, preceding frames, and simultaneous activity. Gate dependence breaks the elementary replacement of every interior error by one identical twirl.
Perturbative and representation-theoretic analyses show why a dominant exponential often survives. For broad Markovian gate-dependent models, the decay can be written schematically as
where is a gate-dependent correction that can decay more rapidly than the dominant mode under suitable conditions. This does not license ignoring it. Short-length residuals, strong compilation dependence, or multiple long-lived modes can make a one-exponential summary inadequate.
There is also a gauge issue. Observable circuit probabilities are invariant under simultaneous similarity transformations of state, gates, and measurement. Under gate-dependent noise, relating to a mean fidelity requires a consistent representation or gauge of the noisy gate set. The RB decay itself is operational; a decomposition into individual gate errors need not be.
Time Dependence and Memory
Section titled “Time Dependence and Memory”If the error at step depends on time or history,
then random sequences sample both gates and an evolving environment. Slowly varying Markovian noise may appear as a mixture of nearby exponentials. Quasistatic or correlated noise can produce skewed sequence distributions, nonvanishing variance, shoulders, and acquisition-order dependence.
Defenses include:
- randomizing the order of sequence lengths;
- interleaving reference sequences and calibration checks;
- recording timestamps and control-system state;
- repeating independent epochs rather than pooling them immediately;
- plotting per-sequence and per-epoch distributions;
- testing whether a model fit in one block predicts later blocks.
A visually smooth average decay does not prove Markovianity. Averaging can hide the very correlations that matter to long algorithms.
Markovian and Non-Markovian Noise owns the broader fixed-step, CP-divisibility, causal-break, confounder, and model-escalation tests; this page retains RB sequence design, decay inference, correlated-noise signatures, uncertainty, and benchmark-specific failure modes.
Leakage Breaks the One-Mode Picture
Section titled “Leakage Breaks the One-Mode Picture”Standard RB assumes a trace-preserving channel on a fixed -dimensional computational space. Leakage violates this closure. Population can leave the space, remain outside it, or seep back, so survival may contain more than one decay mode and need not approach the standard asymptote.
Let project onto the computational subspace. Measuring both the target-state survival and total computational population,
helps separate logical randomization within the subspace from population loss. Leakage RB introduces models for average leakage and seepage rates. Coherent excursions that return during a gate can still generate logical phase error without leaving final population outside the subspace.
Do not force leaky data into and report as if the computational channel were trace preserving.
Average Error Is Not Worst-Case Error
Section titled “Average Error Is Not Worst-Case Error”Consider a coherent one-qubit overrotation
Its average infidelity is
whereas its diamond distance from the identity is
For small , average infidelity is quadratic while worst-case error is linear. A stochastic Pauli channel with the same has very different composition behavior. Standard RB alone does not distinguish these cases.
Unitarity benchmarking modifies the protocol to estimate how strongly a channel preserves the length of traceless operators, providing information about coherent content. It is a complementary statistic, not a conversion of average infidelity into a universal worst-case guarantee.
Error per Clifford Is Not Error per Native Gate
Section titled “Error per Clifford Is Not Error per Native Gate”Suppose, only for illustration, that every sampled Clifford consists of exactly identical native gates with the same depolarizing parameter . Then
For small errors, dividing an error per Clifford by approximates the native error. Actual compilations have variable lengths, unequal gate types, virtual operations, cancellation, idle errors, and gate-dependent coherent effects. The observed is then a weighted circuit property, and naive division need not recover any physical primitive’s fidelity.
At minimum, report the distribution of native gate counts and durations per sampled group element. To characterize one native gate, use a protocol whose theory and randomization directly support that target, such as carefully bounded interleaved RB or an appropriate direct benchmarking protocol.
Interleaved Randomized Benchmarking
Section titled “Interleaved Randomized Benchmarking”Interleaved RB alternates a target gate with random reference elements. Fit a reference decay and an interleaved decay .
If both reference and target errors act as independent depolarizing channels,
Thus
and the target estimate is
The approximation sign matters. The target error can interact coherently with reference errors; the interleaved circuit has a different duration and thermal history; and gate-dependent reference errors do not generally factor into one scalar. The original protocol supplies systematic bounds under stated conditions. A ratio with only statistical fit bars should not be presented as model-free gate fidelity.
Reference and interleaved data should be acquired close in time or interwoven so drift does not masquerade as target-gate error.
Other RB Variants
Section titled “Other RB Variants”Simultaneous RB
Section titled “Simultaneous RB”Run reference RB on subsystems separately and then simultaneously. A changed decay is evidence of context or addressability error under the compared schedules. It does not by itself localize the coupling mechanism or predict every parallel workload.
Unitarity RB
Section titled “Unitarity RB”Fit a purity-like decay of traceless observables to estimate the unitarity of the noise. Combining unitarity with average infidelity helps distinguish coherent from stochastic behavior, subject to the protocol’s state, measurement, and leakage assumptions.
Leakage RB
Section titled “Leakage RB”Track computational-subspace population as well as target survival and fit a model with leakage and seepage. The measurement must distinguish leaked population, and multiple leakage states or coherent return can require a richer model.
Direct and native-gate RB
Section titled “Direct and native-gate RB”Direct RB and related protocols sample customizable native-gate circuits rather than compiling every draw into a large Clifford. They improve scaling and make the sampled operation mix more explicit, but their correctness relies on their own scrambling and error assumptions. “RB” names a family of protocols, not one interchangeable fit.
Character and dihedral protocols
Section titled “Character and dihedral protocols”Representation-aware variants isolate selected decay modes or benchmark gate sets smaller than the full Clifford group. Their fitted parameters and conversion formulas must be taken from that protocol’s representation, not copied from standard Clifford RB.
A Defensible Experimental Workflow
Section titled “A Defensible Experimental Workflow”- Declare the estimand. Name the group or gate distribution, subsystem, compiler, native schedule, context, time window, and whether the result is per Clifford, per target gate, or another protocol-specific element.
- Freeze and validate circuit generation. Test group sampling, composition order, inverse correctness, seed replay, compiler determinism, and native gate accounting.
- Design lengths and replication. Use a pilot to span the decay, then allocate independent sequences and shots according to both sequence and binomial variance.
- Randomize acquisition order. Interleave lengths, reference checks, and when relevant interleaved-gate experiments.
- Preserve raw hierarchy. Store every seed, compiled circuit, native gate count, timestamp, calibration identifier, shot count, outcome, leakage flag, and exclusion.
- Fit a declared model. State the likelihood or weighting, parameter constraints, initialization, convergence checks, and treatment of overdispersion.
- Inspect model adequacy. Plot residuals and per-sequence distributions; compare epochs and justified alternative decay models.
- Propagate statistical uncertainty. Resample sequences, retain nested shot noise, and separate systematic protocol assumptions from fit intervals.
- Cross-check another protocol. Compare with direct calibration, unitarity, leakage, tomography, simultaneous tests, or workload circuits as appropriate.
- Publish the evidence bundle. Reproducible Notebooks defines the artifact and clean-execution requirements.
Minimum Reporting Record
Section titled “Minimum Reporting Record”Report at least:
- reference group, sampling distribution, subsystem dimension, and target state/effect;
- exact definition of sequence length and whether the inverse is counted;
- sequence lengths, independent sequences per length, shots per sequence, and acquisition order;
- random seeds, generator version, group-composition convention, and inverse tests;
- compiler version, optimization level, native gate counts, durations, virtual operations, idles, and pulse-schedule policy;
- calibration identifiers, timestamps, simultaneous activity, reset policy, leakage treatment, and exclusions;
- raw counts retaining sequence identity;
- fit model, objective or likelihood, weights, parameter constraints, residuals, and alternative-model checks;
- , , , converted , statistical intervals, and any systematic bounds;
- the precise language of the claim: per reference element, per interleaved target, per subsystem, and under which assumptions.
Common Mistakes
Section titled “Common Mistakes”Calling RB SPAM-free
Section titled “Calling RB SPAM-free”RB absorbs stationary leading SPAM factors into and . Length-dependent reset, leakage-sensitive readout, drift, and poor contrast still matter.
Reporting “gate fidelity” without naming the gate distribution
Section titled “Reporting “gate fidelity” without naming the gate distribution”Reference RB averages over implemented reference elements. The compiler and sampling measure define the average.
Dividing error per Clifford by the mean gate count
Section titled “Dividing error per Clifford by the mean gate count”This is only a low-error approximation under a restrictive homogeneous model. Variable compilations and gate-dependent errors invalidate the shortcut.
Using one random sequence per length
Section titled “Using one random sequence per length”Repeated shots then estimate that circuit’s survival, not the sequence-ensemble mean.
Bootstrapping shots but not sequences
Section titled “Bootstrapping shots but not sequences”This omits between-sequence uncertainty and can produce intervals that are far too narrow.
Fitting every curve to one exponential
Section titled “Fitting every curve to one exponential”Leakage, memory, gate-dependent transients, and drift can add modes or overdispersion. The fit form is a hypothesis to test.
Treating interleaved decay ratios as model-free
Section titled “Treating interleaved decay ratios as model-free”Reference and target errors need not factor. Report systematic bounds and acquisition timing.
Comparing RB numbers with different compilers
Section titled “Comparing RB numbers with different compilers”Two experiments over the same abstract Clifford group can implement different native workloads.
Equating average infidelity with worst-case error
Section titled “Equating average infidelity with worst-case error”Coherent and stochastic channels with the same RB infidelity can compose very differently.
Ignoring chronological order
Section titled “Ignoring chronological order”Pooling an improving or degrading calibration into one curve can produce a precise estimate of no stationary device state.
Further Connections
Section titled “Further Connections”- Logical Benchmarking lifts randomized sequence tests to encoded gates while retaining syndrome memory, decoder, acceptance, gadget-boundary, and resource caveats.
- Why Benchmarking Is Hard supplies the broader estimand, context, drift, verification, and reporting contract.
- Metrics for Quantum Hardware compares average infidelity with worst-case, leakage, logical, and workload metrics.
- Cycle Benchmarking keeps one scheduled layer fixed, uses local Pauli dressing, and estimates a dressed-cycle process fidelity from Pauli-orbit decays.
- Cross-Entropy Benchmarking scores random-circuit outputs with ideal probabilities, enabling non-Clifford system tests while adding scrambling, classical-reference, and spoofing assumptions absent from reference RB.
- Process Tomography reconstructs a detailed channel model rather than one randomized decay.
- Noise in Quantum Information develops the error mechanisms that can share one RB average.
- Stabilizer Formalism explains Clifford closure, Pauli conjugation, and efficient inverse computation.
- Calibration Loops owns drift detection, acceptance, publication, and rollback of the controls being benchmarked.
- Device Characterization uses RB as one diagnostic coordinate alongside unitarity, GST, mechanism-sensitive amplification, leakage, context, and drift tests.
References
Section titled “References”- Joseph Emerson, Robert Alicki, and Karol Życzkowski, “Scalable Noise Estimation with Random Unitary Operators”, Journal of Optics B 7, S347–S352 (2005) — random motion reversal and exponential noise estimation.
- E. Knill et al., “Randomized Benchmarking of Quantum Gates”, Physical Review A 77, 012307 (2008) — early experimental RB protocol and long-sequence gate-error estimate.
- Christoph Dankert, Richard Cleve, Joseph Emerson, and Etera Livine, “Exact and Approximate Unitary 2-Designs and Their Application to Fidelity Estimation”, Physical Review A 80, 012304 (2009) — unitary-design foundation for efficient twirling.
- Easwar Magesan, Jay M. Gambetta, and Joseph Emerson, “Scalable and Robust Randomized Benchmarking of Quantum Processes”, Physical Review Letters 106, 180504 (2011) — standard scalable RB decay and perturbative noise analysis.
- Easwar Magesan, Jay M. Gambetta, and Joseph Emerson, “Characterizing Quantum Gates via Randomized Benchmarking”, Physical Review A 85, 042311 (2012) — detailed derivation, SPAM treatment, and gate-dependence conditions.
- Easwar Magesan et al., “Efficient Measurement of Quantum Gate Error by Interleaved Randomized Benchmarking”, Physical Review Letters 109, 080505 (2012) — target-gate interleaving and systematic bounds.
- Jay M. Gambetta et al., “Characterization of Addressability by Simultaneous Randomized Benchmarking”, Physical Review Letters 109, 240504 (2012) — simultaneous RB for context and addressability.
- Jeffrey M. Epstein, Andrew W. Cross, Easwar Magesan, and Jay M. Gambetta, “Investigating the Limits of Randomized Benchmarking Protocols”, Physical Review A 89, 062321 (2014) — coherent, damping, leakage, and correlated-noise stress tests.
- Joel J. Wallman and Steven T. Flammia, “Randomized Benchmarking with Confidence”, New Journal of Physics 16, 103032 (2014) — sequence variance, experimental design, and confidence bounds.
- Harrison Ball, Thomas M. Stace, Steven T. Flammia, and Michael J. Biercuk, “The Effect of Noise Correlations on Randomized Benchmarking”, Physical Review A 93, 022303 (2016) — sequence distributions under correlated noise.
- Joel J. Wallman, Christopher Granade, Robin Harper, and Steven T. Flammia, “Estimating the Coherence of Noise”, New Journal of Physics 17, 113020 (2015) — unitarity and coherence-sensitive RB.
- Timothy Proctor, Kenneth Rudinger, Kevin Young, Mohan Sarovar, and Robin Blume-Kohout, “What Randomized Benchmarking Actually Measures”, Physical Review Letters 119, 130502 (2017) — operational interpretation and gauge issues under gate dependence.
- Joel J. Wallman, “Randomized Benchmarking with Gate-Dependent Noise”, Quantum 2, 47 (2018) — dominant exponential and decaying correction for general Markovian gate dependence.
- Christopher J. Wood and Jay M. Gambetta, “Quantification and Characterization of Leakage Errors”, Physical Review A 97, 032306 (2018) — leakage, seepage, and leakage RB.
- Robin Harper, Ian Hincks, Chris Ferrie, Steven T. Flammia, and Joel J. Wallman, “Statistical Analysis of Randomized Benchmarking”, Physical Review A 99, 052350 (2019) — resource allocation and principled RB inference.
- Timothy J. Proctor et al., “Direct Randomized Benchmarking for Multiqubit Devices”, Physical Review Letters 123, 030503 (2019) — native-circuit-oriented benchmarking beyond compiled full Cliffords.
- Jonas Helsen, Ingo Roth, Emilio Onorati, Albert H. Werner, and Jens Eisert, “A General Framework for Randomized Benchmarking”, PRX Quantum 3, 020357 (2022) — representation-theoretic multi-exponential framework and fitting analysis.
Exercises
Section titled “Exercises”1. Derive the RB decay
Section titled “1. Derive the RB decay”Assume the averaged interior noise is , with effective preparation and effect . Derive
and identify and .
Solution
Decompose the effective input into identity and traceless parts:
The depolarizing channel fixes and multiplies every traceless operator by . Therefore
Taking the expectation of gives
Thus
2. Convert decay to infidelity
Section titled “2. Convert decay to infidelity”An RB fit on a two-qubit computational space returns
Find and its propagated one-standard-deviation statistical uncertainty.
Solution
For two qubits, , so
The conversion is linear, hence
Thus the statistical result is
This does not include model, compiler, leakage, or drift uncertainty.
3. Allocate sequences and shots
Section titled “3. Allocate sequences and shots”At one length, suppose the between-sequence variance is and . Compare the variance of the estimated mean for:
- sequences and shots each;
- sequences and shots each.
Both designs use shots.
Solution
For equal ,
For the first design,
For the second,
More sequences reduce the first term, so the second design is better here. The improvement is modest because shot noise dominates. If were larger, the benefit would be greater.
4. Error per compiled element
Section titled “4. Error per compiled element”Assume every one-qubit Clifford contains exactly identical native gates with depolarizing parameter , and an RB fit gives . Find and the corresponding one-qubit native-gate infidelity. State why the calculation is usually only illustrative.
Solution
Under the stated homogeneous model,
For a qubit,
Real Clifford compilations have a distribution of native gates, unequal error channels, virtual operations, cancellations, idles, and coherent gate-dependent effects. One scalar does not generally factor into identical primitive channels.
5. Interleaved RB
Section titled “5. Interleaved RB”A one-qubit reference experiment gives , while the interleaved experiment gives . Compute the depolarizing-ratio estimate of the target-gate infidelity.
Solution
The inferred target depolarizing parameter is
For ,
This is the ratio-model estimate. A complete result also includes statistical propagation, reference gate dependence, target-reference interaction, drift, and the protocol’s systematic bounds.
6. Coherent overrotation
Section titled “6. Coherent overrotation”For radians, estimate the average infidelity and diamond distance of . Why can one RB number not distinguish this error from a stochastic channel with the same infidelity?
Solution
For small ,
The diamond distance is
Reference RB primarily estimates the average decay and therefore the average infidelity under its model. Twirling suppresses directional information. A stochastic channel can have the same but a much smaller worst-case-to- average separation and different composition behavior. A coherence-sensitive protocol such as unitarity benchmarking is needed for more information.
7. Diagnose a drifting experiment
Section titled “7. Diagnose a drifting experiment”An RB run acquires all short sequences in the morning and all long sequences in the afternoon. Calibration quality degrades monotonically during the day, but the pooled data fit one exponential with small residuals. Is the fitted a valid stationary gate-set error?
Solution
No. Sequence length is confounded with time. A lower afternoon survival can be explained by length, drift, or both. A smooth exponential does not identify the cause.
The experiment should randomize or interleave lengths, record timestamps and calibration identifiers, repeat reference lengths throughout the run, and compare time blocks. A hierarchical model may estimate time dependence, but it cannot recover information absent from the confounded acquisition. The pooled can be reported only as a property of that chronological campaign, not a stationary gate set.
8. Detect leakage-model failure
Section titled “8. Detect leakage-model failure”Suppose target-state survival decays below and total computational-space population also decreases with length. Explain why the standard RB conversion is invalid and name the additional observables needed for a leakage analysis.
Solution
The standard derivation assumes a trace-preserving channel on the -dimensional computational subspace. Decreasing computational population violates that assumption, and the asymptote and decay modes need not match . The fitted cannot be converted to
as though it described a closed depolarizing channel.
A leakage analysis should separately estimate target-state survival and total computational-subspace population. If the detector resolves leakage levels, their populations are also useful. A model with leakage and seepage rates, and possibly multiple or coherent leakage modes, should be fitted and validated.