Skip to content

Randomized Benchmarking

Randomized benchmarking (RB) estimates an ensemble-average error parameter for implemented quantum gates by measuring how the survival probability of random identity circuits decays with sequence length. In standard reference RB, one samples mm gates from a finite group, appends the group inverse, executes the compiled circuit, and asks whether the final state agrees with the intended return state.

Under a baseline model of stationary, Markovian, gate-independent noise and a unitary 2-design such as the Clifford group, the sequence-averaged survival probability is

F(m)=Apm+B.F(m) = A p^m+B.

The constants AA and BB contain fixed state-preparation, measurement, and edge effects. The decay parameter pp determines the average infidelity of the twirled noise on a dd-dimensional computational space:

rRB=1−Favg=d−1d(1−p).r_{\mathrm{RB}} = 1-F_{\mathrm{avg}} = \frac{d-1}{d}(1-p).

This statement is powerful but conditional. RB estimates an average associated with the sampled group, its compiler, its physical context, and its time window. It does not reconstruct a noise channel, identify an error mechanism, return a worst-case error, or automatically yield an error per native gate.

This page is the canonical home for standard reference RB, its twirl derivation, sequence and shot statistics, fit design, interpretation, interleaved RB, and the main failure modes. Metrics for Quantum Hardware owns comparisons among hardware metrics. Process Tomography owns channel reconstruction, while Noise in Quantum Information owns the physical taxonomy of coherent, stochastic, leakage, correlated, and non-Markovian noise. Cycle-level randomized protocols belong to the separate Cycle Benchmarking page.

Let G\mathbb G be the declared reference group. In a one-qubit experiment it is commonly the 24-element Clifford group; for several qubits, sampled multi-qubit Cliffords must be compiled into available native operations. A reference RB result is attached to the complete implementation map

ϕ:G∈G⟼scheduled pulses, frames, idles, and measurements.\phi: G\in\mathbb G \longmapsto \text{scheduled pulses, frames, idles, and measurements}.

Changing ϕ\phi changes the benchmark even when the abstract group is unchanged. Whole-circuit optimization, cancellation, virtual-frame handling, randomized compilation, pulse overlap, spectator activity, and calibration versions are therefore part of the estimand.

The primary object is the mean survival over random sequences:

F(m)=EG∼Gmq(m,G),F(m) = \mathbb E_{\boldsymbol G\sim\mathbb G^m} q(m,\boldsymbol G),

not the survival of one representative sequence. Here G=(G1,…,Gm)\boldsymbol G=(G_1,\ldots,G_m) and qq is the probability generated by the implemented circuit for that particular draw.

Under the standard model, rRBr_{\mathrm{RB}} is an error per sampled reference element. It is often called error per Clifford when G\mathbb G is a Clifford group. It is not automatically error per pulse, error per primitive gate, error per layer, or error for the target application.

For each selected length mm:

  1. Draw G1,…,GmG_1,\ldots,G_m independently and uniformly from G\mathbb G.

  2. Compute the ideal inverse

    Gm+1=(Gm⋯G2G1)−1.G_{m+1} = (G_m\cdots G_2G_1)^{-1}.
  3. Compile all m+1m+1 group elements under a frozen compiler and calibration policy.

  4. Prepare a declared state ρ\rho, execute the sequence, and measure a declared survival effect EE.

  5. Repeat shots for the same sequence to estimate its survival probability.

  6. Repeat with independently sampled sequences at the same mm.

  7. Repeat across lengths and fit the sequence-averaged data.

The implemented sequence channel is

S~G=G~m+1∘G~m∘⋯∘G~1,\widetilde{\mathcal S}_{\boldsymbol G} = \widetilde{\mathcal G}_{m+1} \circ\widetilde{\mathcal G}_{m} \circ\cdots\circ \widetilde{\mathcal G}_{1},

and its survival probability is

q(m,G)=Tr⁡ ⁣[E S~G(ρ)].q(m,\boldsymbol G) = \operatorname{Tr} \!\left[ E\, \widetilde{\mathcal S}_{\boldsymbol G}(\rho) \right].

The exact order convention must be fixed in software and tested. An inverse computed with the wrong multiplication order can still produce plausible decay data while benchmarking a different circuit family.

Reference randomized-benchmarking sequence followed by an inverse and a survival-decay fit

At each length, many random sequences and repeated shots estimate an ensemble mean. The fitted decay pp is distinct from the SPAM-dependent contrast AA and asymptote BB; residuals and between-sequence variation remain part of the evidence.

Represent an ideal group element by the unitary channel

G(ρ)=GρG†.\mathcal G(\rho) = G\rho G^\dagger.

Assume each physical implementation has the same Markovian error channel Λ\Lambda:

G~=Λ∘G.\widetilde{\mathcal G} = \Lambda\circ\mathcal G.

This gate-independent model is not exact hardware physics; it is the clean starting point from which the standard decay law follows.

The group twirl of Λ\Lambda is

TG(Λ)=1∣G∣∑G∈GG−1∘Λ∘G.\mathcal T_{\mathbb G}(\Lambda) = \frac{1}{|\mathbb G|} \sum_{G\in\mathbb G} \mathcal G^{-1} \circ\Lambda\circ\mathcal G.

If G\mathbb G is a unitary 2-design and Λ\Lambda is trace preserving, the twirled channel acts as a depolarizing channel on the computational space:

Dp(ρ)=pρ+(1−p)Id.\mathcal D_p(\rho) = p\rho +(1-p)\frac{I}{d}.

This follows because conjugation leaves the identity direction fixed and acts irreducibly on traceless operators. By Schur’s lemma, a channel commuting with every group conjugation must multiply the entire traceless sector by one scalar pp.

After averaging independent group draws, the randomized frames convert each interior occurrence of Λ\Lambda into the same twirl. Preparation, measurement, and a final edge channel can be collected into effective objects ρ′\rho' and E′E', giving

F(m)=Tr⁡ ⁣[E′Dpm(ρ′)].F(m) = \operatorname{Tr} \!\left[ E'\mathcal D_p^m(\rho') \right].

Since

Dpm(ρ′)=pm(ρ′−Id)+Id,\mathcal D_p^m(\rho') = p^m \left( \rho'-\frac{I}{d} \right) +\frac{I}{d},

one obtains

F(m)=Apm+B,F(m) = A p^m+B,

with

A=Tr⁡ ⁣[E′(ρ′−Id)],B=Tr⁡ ⁣(E′Id).\begin{aligned} A &= \operatorname{Tr} \!\left[ E' \left( \rho'-\frac{I}{d} \right) \right], \\ B &= \operatorname{Tr} \!\left( E'\frac{I}{d} \right). \end{aligned}

The decay is therefore insensitive to the leading constant SPAM amplitude, not independent of all SPAM behavior.

The average gate fidelity of a channel Λ\Lambda relative to the identity is

Favg(Λ)=∫dψ ⟨ψ∣Λ ⁣(∣ψ⟩⟨ψ∣)∣ψ⟩.F_{\mathrm{avg}}(\Lambda) = \int d\psi\, \langle\psi| \Lambda \!\left( |\psi\rangle\langle\psi| \right) |\psi\rangle.

Twirling preserves this average. For Dp\mathcal D_p,

Favg(Dp)=p+1−pd=1+(d−1)pd.F_{\mathrm{avg}}(\mathcal D_p) = p+\frac{1-p}{d} = \frac{1+(d-1)p}{d}.

Hence

rRB=1−Favg=d−1d(1−p).r_{\mathrm{RB}} = 1-F_{\mathrm{avg}} = \frac{d-1}{d}(1-p).

For a single qubit,

rRB=1−p2.r_{\mathrm{RB}} = \frac{1-p}{2}.

The conversion is exact for the baseline channel associated with the RB decay. Under gate-dependent noise, modern RB theory can still justify a dominant exponential and an ensemble-level average error under suitable conditions, but the result is not generally the arithmetic mean of separately defined infidelities for every compiled gate.

Suppose the effective one-qubit depolarizing parameter is

p=0.998.p=0.998.

With ideal preparation and measurement,

A=B=12,A=B=\frac12,

so

F(m)=12(0.998)m+12.F(m) = \frac12(0.998)^m+\frac12.

The reported RB infidelity is

rRB=1−0.9982=10−3.r_{\mathrm{RB}} = \frac{1-0.998}{2} = 10^{-3}.

At length m=100m=100,

F(100)=12(0.998)100+12≃0.9093.F(100) = \frac12(0.998)^{100}+\frac12 \simeq 0.9093.

At m=500m=500,

F(500)≃0.6838.F(500) \simeq 0.6838.

Long sequences amplify a small per-element decay into observable contrast. They also increase exposure to drift, leakage, heating, and memory. Sequence lengths should span the informative part of the curve rather than merely reach the longest executable circuit.

If SPAM changes the curve to A=0.42A=0.42 and B=0.47B=0.47 while remaining stationary, the same pp and rRBr_{\mathrm{RB}} are recovered in the ideal model. The lower contrast increases statistical uncertainty even though it does not change the central decay parameter.

Let k=1,…,Kmk=1,\ldots,K_m index independently drawn sequences at length mm, and let s=1,…,Sm,ks=1,\ldots,S_{m,k} index repeated shots. Conditional on sequence survival qm,kq_{m,k},

Ym,k,s∼Bernoulli⁡(qm,k).Y_{m,k,s} \sim \operatorname{Bernoulli}(q_{m,k}).

The per-sequence estimate is

q^m,k=1Sm,k∑sYm,k,s,\widehat q_{m,k} = \frac{1}{S_{m,k}} \sum_s Y_{m,k,s},

and the unweighted sequence-ensemble mean for equal shot counts is

F^(m)=1Km∑k=1Kmq^m,k.\widehat F(m) = \frac{1}{K_m} \sum_{k=1}^{K_m} \widehat q_{m,k}.

Suppose each sequence uses SS shots and

σseq2(m)=Var⁡G ⁣[q(m,G)].\sigma_{\mathrm{seq}}^2(m) = \operatorname{Var}_{\boldsymbol G} \!\left[ q(m,\boldsymbol G) \right].

The law of total variance gives

Var⁡ ⁣[F^(m)]=σseq2(m)Km+EG[q(1−q)]KmS.\operatorname{Var} \!\left[ \widehat F(m) \right] = \frac{\sigma_{\mathrm{seq}}^2(m)}{K_m} + \frac{ \mathbb E_{\boldsymbol G} [q(1-q)] }{K_m S}.

The first term is between-sequence variation; the second is shot noise. Once shot noise is small, spending the remaining budget on more shots of the same few sequences does not characterize the random-sequence ensemble. Conversely, one shot on many sequences can be inefficient when readout noise dominates.

Sequence is the experimental unit for the randomization. A bootstrap that resamples only individual shots while holding the sampled circuits fixed omits the first variance term.

Lengths should identify AA, pp, and BB rather than cluster where the curve is nearly flat. A practical design includes:

  • short lengths to resolve initial contrast;
  • lengths near the characteristic decay scale;
  • long lengths that constrain the asymptote without driving all signal below readout resolution;
  • repeated lengths throughout the run to reveal drift;
  • more sequence draws where between-sequence variance is large.

For pp near one,

pm≃exp⁡[−m(1−p)],p^m \simeq \exp[-m(1-p)],

so the characteristic scale is approximately

m⋆∼11−p.m_\star \sim \frac{1}{1-p}.

This estimate guides an initial design; a pilot run can then refine the length grid. Selecting lengths after seeing favorable random fluctuations without accounting for that adaptation biases the final analysis.

The familiar three-parameter nonlinear least-squares fit is not the only statistical model. A defensible analysis keeps the data hierarchy visible.

At minimum, account for heteroskedasticity: survival probabilities near 1/21/2 have different binomial variance from probabilities near one, and between-sequence variation can change with mm. Weighted nonlinear least squares can be useful if weights are estimated honestly. A binomial or hierarchical likelihood retains more of the acquisition model.

Do not first average away sequence identifiers and then treat the resulting point means as equally precise. Preserve counts, sequence seeds, and shots per sequence.

Useful approaches include:

  • hierarchical parametric bootstrap under a declared RB model;
  • nonparametric resampling of sequences within each length, with shot resampling nested inside;
  • profile-likelihood or Bayesian intervals with transparent priors;
  • rigorous concentration bounds when their assumptions and constants fit the experiment.

Report uncertainty in pp and propagate it through rRB=(d−1)(1−p)/dr_{\mathrm{RB}}=(d-1)(1-p)/d. Systematic uncertainty from model failure, compiler variation, drift, or interleaved-RB assumptions is separate from fit uncertainty.

Inspect residuals against sequence length, acquisition time, random-sequence identity, compiler statistics, and leakage flags. Compare the single exponential with justified alternatives, such as an additional decaying mode for leakage or a transient term for gate dependence. An alternative model is evidence only if it improves predictive behavior on held-out or replicated data, not merely in-sample fit.

Fixed preparation and measurement errors change AA and BB while leaving the interior decay pp unchanged in the baseline derivation. This is the sense in which reference RB is robust to SPAM.

SPAM can still matter in several ways:

  • poor preparation or readout reduces ∣A∣|A| and makes pp hard to identify;
  • reset quality can depend on sequence length or previous outcomes;
  • measurement response can depend on leakage or heating accumulated during the sequence;
  • the final inversion has sequence-dependent compilation and can correlate edge error with the sampled circuit;
  • drift can make ρ′\rho', E′E', AA, and BB vary during acquisition;
  • postselection can change the estimand.

Calling RB “SPAM-free” is therefore incorrect. It is insensitive to a particular stationary nuisance structure under a specified model.

Real implementations have

G~=ΛG∘G,\widetilde{\mathcal G} = \Lambda_G\circ\mathcal G,

with error depending on GG, its native decomposition, preceding frames, and simultaneous activity. Gate dependence breaks the elementary replacement of every interior error by one identical twirl.

Perturbative and representation-theoretic analyses show why a dominant exponential often survives. For broad Markovian gate-dependent models, the decay can be written schematically as

F(m)=Apm+B+Δ(m),F(m) = A p^m+B+\Delta(m),

where Δ(m)\Delta(m) is a gate-dependent correction that can decay more rapidly than the dominant mode under suitable conditions. This does not license ignoring it. Short-length residuals, strong compilation dependence, or multiple long-lived modes can make a one-exponential summary inadequate.

There is also a gauge issue. Observable circuit probabilities are invariant under simultaneous similarity transformations of state, gates, and measurement. Under gate-dependent noise, relating pp to a mean fidelity requires a consistent representation or gauge of the noisy gate set. The RB decay itself is operational; a decomposition into individual gate errors need not be.

If the error at step jj depends on time or history,

Λj=Λ(tj,Gj,G<j,hj),\Lambda_j = \Lambda \left( t_j,G_j,G_{<j},h_j \right),

then random sequences sample both gates and an evolving environment. Slowly varying Markovian noise may appear as a mixture of nearby exponentials. Quasistatic or correlated noise can produce skewed sequence distributions, nonvanishing variance, shoulders, and acquisition-order dependence.

Defenses include:

  • randomizing the order of sequence lengths;
  • interleaving reference sequences and calibration checks;
  • recording timestamps and control-system state;
  • repeating independent epochs rather than pooling them immediately;
  • plotting per-sequence and per-epoch distributions;
  • testing whether a model fit in one block predicts later blocks.

A visually smooth average decay does not prove Markovianity. Averaging can hide the very correlations that matter to long algorithms.

Markovian and Non-Markovian Noise owns the broader fixed-step, CP-divisibility, causal-break, confounder, and model-escalation tests; this page retains RB sequence design, decay inference, correlated-noise signatures, uncertainty, and benchmark-specific failure modes.

Standard RB assumes a trace-preserving channel on a fixed dd-dimensional computational space. Leakage violates this closure. Population can leave the space, remain outside it, or seep back, so survival may contain more than one decay mode and need not approach the standard 1/d1/d asymptote.

Let PCP_{\mathrm C} project onto the computational subspace. Measuring both the target-state survival and total computational population,

pC(m)=Tr⁡ ⁣[PCρm],p_{\mathrm C}(m) = \operatorname{Tr} \!\left[ P_{\mathrm C}\rho_m \right],

helps separate logical randomization within the subspace from population loss. Leakage RB introduces models for average leakage and seepage rates. Coherent excursions that return during a gate can still generate logical phase error without leaving final population outside the subspace.

Do not force leaky data into Apm+BA p^m+B and report r=(d−1)(1−p)/dr=(d-1)(1-p)/d as if the computational channel were trace preserving.

Consider a coherent one-qubit overrotation

Uθ=exp⁡(−iθZ2).U_\theta = \exp \left( -\frac{i\theta Z}{2} \right).

Its average infidelity is

r=1−cos⁡θ3≃θ26,r = \frac{1-\cos\theta}{3} \simeq \frac{\theta^2}{6},

whereas its diamond distance from the identity is

D⋄=∣sin⁡θ2∣≃∣θ∣2.D_\diamond = \left| \sin\frac{\theta}{2} \right| \simeq \frac{|\theta|}{2}.

For small θ\theta, average infidelity is quadratic while worst-case error is linear. A stochastic Pauli channel with the same rr has very different composition behavior. Standard RB alone does not distinguish these cases.

Unitarity benchmarking modifies the protocol to estimate how strongly a channel preserves the length of traceless operators, providing information about coherent content. It is a complementary statistic, not a conversion of average infidelity into a universal worst-case guarantee.

Error per Clifford Is Not Error per Native Gate

Section titled “Error per Clifford Is Not Error per Native Gate”

Suppose, only for illustration, that every sampled Clifford consists of exactly LL identical native gates with the same depolarizing parameter pgp_g. Then

pC=pgL,pg=pC1/L.p_{\mathrm C} = p_g^L, \qquad p_g = p_{\mathrm C}^{1/L}.

For small errors, dividing an error per Clifford by LL approximates the native error. Actual compilations have variable lengths, unequal gate types, virtual operations, cancellation, idle errors, and gate-dependent coherent effects. The observed pCp_{\mathrm C} is then a weighted circuit property, and naive division need not recover any physical primitive’s fidelity.

At minimum, report the distribution of native gate counts and durations per sampled group element. To characterize one native gate, use a protocol whose theory and randomization directly support that target, such as carefully bounded interleaved RB or an appropriate direct benchmarking protocol.

Interleaved RB alternates a target gate G⋆G_\star with random reference elements. Fit a reference decay prefp_{\mathrm{ref}} and an interleaved decay pintp_{\mathrm{int}}.

If both reference and target errors act as independent depolarizing channels,

pint=prefp⋆.p_{\mathrm{int}} = p_{\mathrm{ref}}p_\star.

Thus

p⋆=pintpref,p_\star = \frac{p_{\mathrm{int}}}{p_{\mathrm{ref}}},

and the target estimate is

r⋆≈d−1d(1−pintpref).r_\star \approx \frac{d-1}{d} \left( 1-\frac{p_{\mathrm{int}}}{p_{\mathrm{ref}}} \right).

The approximation sign matters. The target error can interact coherently with reference errors; the interleaved circuit has a different duration and thermal history; and gate-dependent reference errors do not generally factor into one scalar. The original protocol supplies systematic bounds under stated conditions. A ratio with only statistical fit bars should not be presented as model-free gate fidelity.

Reference and interleaved data should be acquired close in time or interwoven so drift does not masquerade as target-gate error.

Run reference RB on subsystems separately and then simultaneously. A changed decay is evidence of context or addressability error under the compared schedules. It does not by itself localize the coupling mechanism or predict every parallel workload.

Fit a purity-like decay of traceless observables to estimate the unitarity of the noise. Combining unitarity with average infidelity helps distinguish coherent from stochastic behavior, subject to the protocol’s state, measurement, and leakage assumptions.

Track computational-subspace population as well as target survival and fit a model with leakage and seepage. The measurement must distinguish leaked population, and multiple leakage states or coherent return can require a richer model.

Direct RB and related protocols sample customizable native-gate circuits rather than compiling every draw into a large Clifford. They improve scaling and make the sampled operation mix more explicit, but their correctness relies on their own scrambling and error assumptions. “RB” names a family of protocols, not one interchangeable fit.

Representation-aware variants isolate selected decay modes or benchmark gate sets smaller than the full Clifford group. Their fitted parameters and conversion formulas must be taken from that protocol’s representation, not copied from standard Clifford RB.

  1. Declare the estimand. Name the group or gate distribution, subsystem, compiler, native schedule, context, time window, and whether the result is per Clifford, per target gate, or another protocol-specific element.
  2. Freeze and validate circuit generation. Test group sampling, composition order, inverse correctness, seed replay, compiler determinism, and native gate accounting.
  3. Design lengths and replication. Use a pilot to span the decay, then allocate independent sequences and shots according to both sequence and binomial variance.
  4. Randomize acquisition order. Interleave lengths, reference checks, and when relevant interleaved-gate experiments.
  5. Preserve raw hierarchy. Store every seed, compiled circuit, native gate count, timestamp, calibration identifier, shot count, outcome, leakage flag, and exclusion.
  6. Fit a declared model. State the likelihood or weighting, parameter constraints, initialization, convergence checks, and treatment of overdispersion.
  7. Inspect model adequacy. Plot residuals and per-sequence distributions; compare epochs and justified alternative decay models.
  8. Propagate statistical uncertainty. Resample sequences, retain nested shot noise, and separate systematic protocol assumptions from fit intervals.
  9. Cross-check another protocol. Compare with direct calibration, unitarity, leakage, tomography, simultaneous tests, or workload circuits as appropriate.
  10. Publish the evidence bundle. Reproducible Notebooks defines the artifact and clean-execution requirements.

Report at least:

  • reference group, sampling distribution, subsystem dimension, and target state/effect;
  • exact definition of sequence length and whether the inverse is counted;
  • sequence lengths, independent sequences per length, shots per sequence, and acquisition order;
  • random seeds, generator version, group-composition convention, and inverse tests;
  • compiler version, optimization level, native gate counts, durations, virtual operations, idles, and pulse-schedule policy;
  • calibration identifiers, timestamps, simultaneous activity, reset policy, leakage treatment, and exclusions;
  • raw counts retaining sequence identity;
  • fit model, objective or likelihood, weights, parameter constraints, residuals, and alternative-model checks;
  • AA, BB, pp, converted rr, statistical intervals, and any systematic bounds;
  • the precise language of the claim: per reference element, per interleaved target, per subsystem, and under which assumptions.

RB absorbs stationary leading SPAM factors into AA and BB. Length-dependent reset, leakage-sensitive readout, drift, and poor contrast still matter.

Reporting “gate fidelity” without naming the gate distribution

Section titled “Reporting “gate fidelity” without naming the gate distribution”

Reference RB averages over implemented reference elements. The compiler and sampling measure define the average.

Dividing error per Clifford by the mean gate count

Section titled “Dividing error per Clifford by the mean gate count”

This is only a low-error approximation under a restrictive homogeneous model. Variable compilations and gate-dependent errors invalidate the shortcut.

Repeated shots then estimate that circuit’s survival, not the sequence-ensemble mean.

This omits between-sequence uncertainty and can produce intervals that are far too narrow.

Leakage, memory, gate-dependent transients, and drift can add modes or overdispersion. The fit form is a hypothesis to test.

Treating interleaved decay ratios as model-free

Section titled “Treating interleaved decay ratios as model-free”

Reference and target errors need not factor. Report systematic bounds and acquisition timing.

Comparing RB numbers with different compilers

Section titled “Comparing RB numbers with different compilers”

Two experiments over the same abstract Clifford group can implement different native workloads.

Equating average infidelity with worst-case error

Section titled “Equating average infidelity with worst-case error”

Coherent and stochastic channels with the same RB infidelity can compose very differently.

Pooling an improving or degrading calibration into one curve can produce a precise estimate of no stationary device state.

  • Logical Benchmarking lifts randomized sequence tests to encoded gates while retaining syndrome memory, decoder, acceptance, gadget-boundary, and resource caveats.
  • Why Benchmarking Is Hard supplies the broader estimand, context, drift, verification, and reporting contract.
  • Metrics for Quantum Hardware compares average infidelity with worst-case, leakage, logical, and workload metrics.
  • Cycle Benchmarking keeps one scheduled layer fixed, uses local Pauli dressing, and estimates a dressed-cycle process fidelity from Pauli-orbit decays.
  • Cross-Entropy Benchmarking scores random-circuit outputs with ideal probabilities, enabling non-Clifford system tests while adding scrambling, classical-reference, and spoofing assumptions absent from reference RB.
  • Process Tomography reconstructs a detailed channel model rather than one randomized decay.
  • Noise in Quantum Information develops the error mechanisms that can share one RB average.
  • Stabilizer Formalism explains Clifford closure, Pauli conjugation, and efficient inverse computation.
  • Calibration Loops owns drift detection, acceptance, publication, and rollback of the controls being benchmarked.
  • Device Characterization uses RB as one diagnostic coordinate alongside unitarity, GST, mechanism-sensitive amplification, leakage, context, and drift tests.

Assume the averaged interior noise is Dp\mathcal D_p, with effective preparation ρ′\rho' and effect E′E'. Derive

F(m)=Apm+BF(m)=A p^m+B

and identify AA and BB.

Solution

Decompose the effective input into identity and traceless parts:

ρ′=(ρ′−Id)+Id.\rho' = \left( \rho'-\frac{I}{d} \right) +\frac{I}{d}.

The depolarizing channel fixes I/dI/d and multiplies every traceless operator by pp. Therefore

Dpm(ρ′)=pm(ρ′−Id)+Id.\mathcal D_p^m(\rho') = p^m \left( \rho'-\frac{I}{d} \right) +\frac{I}{d}.

Taking the expectation of E′E' gives

F(m)=pmTr⁡ ⁣[E′(ρ′−Id)]+Tr⁡ ⁣(E′Id)=Apm+B.\begin{aligned} F(m) &= p^m \operatorname{Tr} \!\left[ E' \left( \rho'-\frac{I}{d} \right) \right] + \operatorname{Tr} \!\left( E'\frac{I}{d} \right) \\ &= A p^m+B. \end{aligned}

Thus

A=Tr⁡ ⁣[E′(ρ′−Id)],B=Tr⁡ ⁣(E′Id).A = \operatorname{Tr} \!\left[ E' \left( \rho'-\frac{I}{d} \right) \right], \qquad B = \operatorname{Tr} \!\left( E'\frac{I}{d} \right).

An RB fit on a two-qubit computational space returns

p=0.9920±0.0008.p=0.9920\pm0.0008.

Find rRBr_{\mathrm{RB}} and its propagated one-standard-deviation statistical uncertainty.

Solution

For two qubits, d=4d=4, so

rRB=34(1−p)=34(0.0080)=0.0060.r_{\mathrm{RB}} = \frac34(1-p) = \frac34(0.0080) = 0.0060.

The conversion is linear, hence

σr=34σp=34(0.0008)=0.0006.\sigma_r = \frac34\sigma_p = \frac34(0.0008) = 0.0006.

Thus the statistical result is

rRB=0.0060±0.0006.r_{\mathrm{RB}} = 0.0060\pm0.0006.

This does not include model, compiler, leakage, or drift uncertainty.

At one length, suppose the between-sequence variance is σseq2=4×10−4\sigma_{\mathrm{seq}}^2=4\times10^{-4} and E[q(1−q)]=0.16\mathbb E[q(1-q)]=0.16. Compare the variance of the estimated mean for:

  1. K=40K=40 sequences and S=100S=100 shots each;
  2. K=100K=100 sequences and S=40S=40 shots each.

Both designs use 40004000 shots.

Solution

For equal SS,

Var⁡(F^)=σseq2K+E[q(1−q)]KS.\operatorname{Var}(\widehat F) = \frac{\sigma_{\mathrm{seq}}^2}{K} + \frac{\mathbb E[q(1-q)]}{KS}.

For the first design,

Var⁡1=4×10−440+0.164000=10−5+4×10−5=5×10−5.\operatorname{Var}_1 = \frac{4\times10^{-4}}{40} + \frac{0.16}{4000} = 10^{-5}+4\times10^{-5} = 5\times10^{-5}.

For the second,

Var⁡2=4×10−4100+0.164000=4×10−6+4×10−5=4.4×10−5.\operatorname{Var}_2 = \frac{4\times10^{-4}}{100} + \frac{0.16}{4000} = 4\times10^{-6}+4\times10^{-5} = 4.4\times10^{-5}.

More sequences reduce the first term, so the second design is better here. The improvement is modest because shot noise dominates. If σseq2\sigma_{\mathrm{seq}}^2 were larger, the benefit would be greater.

Assume every one-qubit Clifford contains exactly L=2L=2 identical native gates with depolarizing parameter pgp_g, and an RB fit gives pC=0.996p_{\mathrm C}=0.996. Find pgp_g and the corresponding one-qubit native-gate infidelity. State why the calculation is usually only illustrative.

Solution

Under the stated homogeneous model,

pg=pC=0.996≃0.997998.p_g = \sqrt{p_{\mathrm C}} = \sqrt{0.996} \simeq 0.997998.

For a qubit,

rg=1−pg2≃1.001×10−3.r_g = \frac{1-p_g}{2} \simeq 1.001\times10^{-3}.

Real Clifford compilations have a distribution of native gates, unequal error channels, virtual operations, cancellations, idles, and coherent gate-dependent effects. One scalar pCp_{\mathrm C} does not generally factor into identical primitive channels.

A one-qubit reference experiment gives pref=0.9980p_{\mathrm{ref}}=0.9980, while the interleaved experiment gives pint=0.9940p_{\mathrm{int}}=0.9940. Compute the depolarizing-ratio estimate of the target-gate infidelity.

Solution

The inferred target depolarizing parameter is

p⋆=0.99400.9980≃0.995992.p_\star = \frac{0.9940}{0.9980} \simeq 0.995992.

For d=2d=2,

r⋆=12(1−p⋆)≃2.004×10−3.r_\star = \frac12(1-p_\star) \simeq 2.004\times10^{-3}.

This is the ratio-model estimate. A complete result also includes statistical propagation, reference gate dependence, target-reference interaction, drift, and the protocol’s systematic bounds.

For θ=0.02\theta=0.02 radians, estimate the average infidelity and diamond distance of Uθ=exp⁡(−iθZ/2)U_\theta=\exp(-i\theta Z/2). Why can one RB number not distinguish this error from a stochastic channel with the same infidelity?

Solution

For small θ\theta,

r≃θ26=(0.02)26≃6.67×10−5.r \simeq \frac{\theta^2}{6} = \frac{(0.02)^2}{6} \simeq 6.67\times10^{-5}.

The diamond distance is

D⋄≃∣θ∣2=0.01.D_\diamond \simeq \frac{|\theta|}{2} = 0.01.

Reference RB primarily estimates the average decay and therefore the average infidelity under its model. Twirling suppresses directional information. A stochastic channel can have the same rr but a much smaller worst-case-to- average separation and different composition behavior. A coherence-sensitive protocol such as unitarity benchmarking is needed for more information.

An RB run acquires all short sequences in the morning and all long sequences in the afternoon. Calibration quality degrades monotonically during the day, but the pooled data fit one exponential with small residuals. Is the fitted pp a valid stationary gate-set error?

Solution

No. Sequence length is confounded with time. A lower afternoon survival can be explained by length, drift, or both. A smooth exponential does not identify the cause.

The experiment should randomize or interleave lengths, record timestamps and calibration identifiers, repeat reference lengths throughout the run, and compare time blocks. A hierarchical model may estimate time dependence, but it cannot recover information absent from the confounded acquisition. The pooled pp can be reported only as a property of that chronological campaign, not a stationary gate set.

Suppose target-state survival decays below 1/d1/d and total computational-space population also decreases with length. Explain why the standard RB conversion is invalid and name the additional observables needed for a leakage analysis.

Solution

The standard derivation assumes a trace-preserving channel on the dd-dimensional computational subspace. Decreasing computational population violates that assumption, and the asymptote and decay modes need not match Apm+BA p^m+B. The fitted pp cannot be converted to

r=d−1d(1−p)r=\frac{d-1}{d}(1-p)

as though it described a closed depolarizing channel.

A leakage analysis should separately estimate target-state survival and total computational-subspace population. If the detector resolves leakage levels, their populations are also useful. A model with leakage and seepage rates, and possibly multiple or coherent leakage modes, should be fitted and validated.