Skip to content

Logical Benchmarking

A logical benchmark measures how encoded quantum information performs under a declared preparation, protected operation or memory interval, syndrome-processing rule, logical measurement, and resource boundary. It asks about the delivered logical channel, not merely whether physical errors were detected or whether one component had high fidelity.

The object being tested can be written schematically as

Lg=D∘Rg∘Ng∘E.\mathcal L_g = \mathcal D \circ \mathcal R_g \circ \mathcal N_g \circ \mathcal E.

E\mathcal E prepares or encodes the logical input, Ng\mathcal N_g is the implemented noisy memory or gate gadget, Rg\mathcal R_g denotes syndrome processing and recovery or Pauli-frame interpretation, and D\mathcal D decodes conceptually to the logical output space. The final decoding need not be a physical circuit; it can be the rule that maps physical measurements to a logical outcome.

This compact expression hides the benchmark’s most important choices. Which logical states are sampled? Does gg mean one syndrome round, one full error-correction cycle, one logical Clifford, or an entire circuit? Is the decoder offline or real time? Are flagged trials discarded? What physical reference defines break-even? A number called “logical error rate” is not interpretable until those questions are answered.

This page owns the design and interpretation of encoded-level benchmark protocols: logical memories, preparations, measurements, gates, random sequences, scaling with code size, break-even comparisons, decoder and postselection policies, rare-event statistics, and performance per unit time and resource.

Metrics for Quantum Hardware owns general metric definitions across the hardware stack. Randomized Benchmarking and Cycle Benchmarking own the underlying physical-level protocol theory. Decoders owns decoding algorithms and systems metrics. Error-Correction Case Studies owns dated experimental claims and comparisons. This page supplies the protocol contract with which those claims should be measured.

Quantum Error Correction and Fault Tolerance routes a finite logical claim here for its estimand, denominator, comparator, selection, uncertainty, decoder version, and resource checks; this page retains encoded benchmark design, break-even, below-threshold evidence, statistical reliability, and interpretation.

A reproducible logical benchmark can be represented by

BL=(C,S,G,X,N,R,M,A,B,Q).\mathfrak B_L = \left( \mathcal C, \mathcal S, \mathcal G, \mathcal X, \mathcal N, \mathcal R, \mathcal M, \mathcal A, \mathcal B, \mathcal Q \right).

The entries are:

  • C\mathcal C: the code, code size or distance, logical basis, and boundary conditions;
  • S\mathcal S: the input-state or logical-circuit sampling ensemble;
  • G\mathcal G: the protected memory interval, preparation, measurement, gate, or circuit family under test;
  • X\mathcal X: syndrome extraction, flags, resets, leakage handling, and timing;
  • N\mathcal N: the physical operating point, noise environment, drift window, and simultaneous activity;
  • R\mathcal R: the decoder, training data, recovery or frame rule, and latency mode;
  • M\mathcal M: the logical estimand and its denominator;
  • A\mathcal A: acceptance, heralding, postselection, and retry policy;
  • B\mathcal B: the comparison baseline; and
  • Q\mathcal Q: the physical, classical, temporal, and sampling resources.

Changing any entry creates a different benchmark. This is appropriate when studying a different question, but the results should not be plotted on the same curve as if only code distance changed.

Logical benchmark contract with a protected logical path, syndrome and decoder path, matched reference, and reported reliability and resource outputs

A logical benchmark compares a complete protected path with a declared reference. The state ensemble, timing, syndrome record, decoder, acceptance rule, and resources remain inside the boundary. Reliability without delivered rate or a matched denominator is only part of the result.

A logical-memory benchmark prepares an ensemble of encoded states, applies rr rounds of error correction or waits for time tt, decodes the record, and measures the logical output. It probes storage under a particular syndrome schedule and decoder. It does not establish the quality of logical gates or state factories.

For a one-logical-qubit Pauli memory, testing only ∣0L⟩|0_L\rangle probes logical bit flips but is insensitive to pure logical dephasing. A minimal cardinal-state ensemble includes eigenstates of XLX_L, YLY_L, and ZLZ_L. Different axes can have different failure probabilities:

pX(L),pY(L),pZ(L).p_X^{(L)}, \qquad p_Y^{(L)}, \qquad p_Z^{(L)}.

Noise bias may make that asymmetry intentional and useful, but averaging it away should be a declared choice. For multiple logical qubits, the benchmark must also detect correlated logical failures rather than report only marginal per-qubit rates.

Logical state preparation and logical measurement are operations in their own right. A preparation benchmark should say whether verification, repeated checks, or postselection is included. A measurement benchmark should report the logical confusion matrix

Cy∣x(L)=Pr⁡(x^=y∣x),C_{y|x}^{(L)} = \Pr(\widehat{x}=y\mid x),

including erasure or rejection outcomes when present. Averaging C0∣0(L)C_{0|0}^{(L)} and C1∣1(L)C_{1|1}^{(L)} can hide strong asymmetry.

If preparation and measurement surround every tested gate, their errors can be absorbed into fitted nuisance parameters only under a stable model. Direct logical process tomography exposes more structure but scales exponentially with the number of logical qubits and remains sensitive to its own SPAM model.

A logical gate benchmark must define the gadget boundary. Does it include leading and trailing error correction, ancilla preparation, repeated parity measurements, decoding, feed-forward, injection, state transport, and factory retries? The gate’s operational channel changes when the boundary changes.

For fault-tolerance claims, benchmark both nominal noise and controlled fault injections. High average fidelity does not show error containment. A fault-tolerant gadget should respond to every fault set within its design order without allowing an uncorrectable output error, subject to the stated code and input conditions. Fault-Tolerant Gates owns that structural contract.

An encoded circuit benchmark composes preparations, gates, measurements, factories, decoders, and controller decisions. It is more representative than an idle memory but harder to diagnose. The logical circuit family should span width, protected depth, non-Clifford demand, entanglement structure, and reaction depth rather than select one favorable demonstration.

Algorithmic Benchmarking owns the further step from a logical circuit to a task-level answer, including input access, output accuracy, classical comparison, and cost per accepted solution.

A logical failure probability is a ratio

p^L=kfailnopp,\widehat p_L = \frac{k_{\mathrm{fail}}} {n_{\mathrm{opp}}},

but the logical opportunity in the denominator may be:

  • one syndrome-extraction round;
  • one complete QEC cycle;
  • one logical qubit-round;
  • one prepared or measured logical state;
  • one protected gate or lattice-surgery operation;
  • one Clifford in a randomized sequence;
  • one spacetime volume; or
  • one complete logical circuit.

These quantities are not interchangeable. For small independent failure probabilities pip_i, a circuit with NiN_i opportunities of class ii has

Pfail,circ≈1−exp⁡(−∑iNipi)≈∑iNipi.P_{\mathrm{fail,circ}} \approx 1- \exp\left(-\sum_iN_ip_i\right) \approx \sum_iN_ip_i.

The approximation fails when faults are correlated, when recovery couples opportunities, or when the operation classes overlap. A per-round number cannot be multiplied by gate count unless the gate genuinely consists of those rounds under the same noise and decoding conditions.

Suppose a binary logical observable flips independently with probability ϵL\epsilon_L per cycle. After rr cycles, its correlation decays as

⟨PL(0)PL(r)⟩=(1−2ϵL)r.\langle P_L(0)P_L(r)\rangle = (1-2\epsilon_L)^r.

With preparation and measurement nuisance parameters, a common fit is

Pcorrect(r)=A+B(1−2ϵL)r.P_{\mathrm{correct}}(r) = A+B(1-2\epsilon_L)^r.

AA and BB do not magically remove SPAM; they model fixed boundary effects. Multiple exponentials, oscillations, leakage, nonstationarity, and decoder memory can invalidate this form.

If one cycle lasts τcyc\tau_{\mathrm{cyc}}, the equivalent continuous-time parity-flip rate is

ΓL=−ln⁡(1−2ϵL)2τcyc≈ϵLτcyc\Gamma_L = -\frac{\ln(1-2\epsilon_L)} {2\tau_{\mathrm{cyc}}} \approx \frac{\epsilon_L}{\tau_{\mathrm{cyc}}}

for ϵL≪1\epsilon_L\ll1. Comparing codes only per cycle can favor a slower cycle; comparing only per second can hide a useful architectural primitive. Report both.

Initialization and final measurement rounds often have different fault mechanisms from steady-state memory rounds. If the probability of a correct logical outcome after rr bulk rounds is fit as

Pcorrect(r)=A+Bλr,P_{\mathrm{correct}}(r) = A+B\lambda^r,

then λ\lambda describes a bulk decay mode only if the boundaries are stable and the chosen range has reached that mode. Quoting ϵL=(1−λ)/2\epsilon_L=(1-\lambda)/2 from two sequence lengths without residual checks can turn transients into a steady-state error rate.

In a stabilizer memory with a Pauli-frame decoder, a trial fails when the inferred recovery differs from the actual fault by a nontrivial logical operator. This produces a logical Pauli distribution

pL=(pI,pX,pY,pZ),∑PpP=1.\mathbf p_L = \left( p_I,p_X,p_Y,p_Z \right), \qquad \sum_Pp_P=1.

The scalar 1−pI1-p_I is meaningful for that Pauli-channel model. Under coherent logical rotations, leakage, or non-Markovian behavior, one number does not describe how errors compose.

For kk logical qubits with logical Pauli basis {Pi}\{P_i\}, a Pauli transfer matrix can represent the decoded channel:

(RL)ij=12kTr⁡[PiL(Pj)].(R_L)_{ij} = \frac{1}{2^k} \operatorname{Tr} \left[ P_i\mathcal L(P_j) \right].

The diagonal entries describe contraction of Pauli observables; off-diagonal entries expose coherent mixing and non-Pauli structure. Logical process tomography can estimate this matrix for small kk, but its preparation, measurement, gauge, and sampling costs grow quickly.

Average logical fidelity is useful for a declared ensemble. It should not be substituted for a worst-case error bound without assumptions. Coherent errors can have average infidelity quadratic in a small rotation angle while their worst-case effect is linear in that angle. Threshold and resource models must use a metric compatible with their noise assumptions.

The code space is not always preserved. Report at least

pcode,pleak,perase,plogical.p_{\mathrm{code}}, \qquad p_{\mathrm{leak}}, \qquad p_{\mathrm{erase}}, \qquad p_{\mathrm{logical}}.

An erasure flag can help the decoder and is operationally different from unflagged leakage. A protocol that discards every detected loss measures a conditional logical channel plus an acceptance rate, not an unconditional error-corrected channel.

Detection events are evidence that the syndrome circuit is sensitive to faults. They are not themselves logical failures, and a low detection-event rate is not necessarily good: a broken check can report no events while failing to protect the state.

Three benchmark modes should be named separately:

  1. Detection: acquire flags or syndrome changes and characterize their statistics without attempting to deliver every logical output.
  2. Postselected detection: accept only a declared subset of records and report the conditional output quality and acceptance.
  3. Correction: decode every in-scope record, update a frame or recovery, and score the delivered logical output.

If aa is the acceptance probability and pL∣Ap_{L|A} the logical failure probability conditioned on acceptance, one trial has probabilities

Pgood=a(1−pL∣A),Pbad=apL∣A,Preject=1−a.P_{\mathrm{good}} = a(1-p_{L|A}), \qquad P_{\mathrm{bad}} = ap_{L|A}, \qquad P_{\mathrm{reject}} = 1-a.

The pair (a,pL∣A)(a,p_{L|A}) must be reported. If rejected trials are restarted, the mean number of attempts per accepted result is 1/a1/a under independent, stationary trials. A lower conditional error obtained by driving aa toward zero is not free error correction.

Selection can also bias the noise distribution. Accepted runs may exclude precisely the leakage, loss, or high-weight events that a scalable machine must handle. That may be appropriate for a heralded primitive, but the benchmark claim must stay conditional.

For code sizes dd and d+Δdd+\Delta d, define the measured suppression factor

Λd,d+Δd=pL(d)pL(d+Δd).\Lambda_{d,d+\Delta d} = \frac{p_L(d)} {p_L(d+\Delta d)}.

Λ>1\Lambda>1 means the larger tested code performs better for the declared metric. The comparison supports below-threshold behavior only when the code family, logical task, physical operating point, cycle definition, noise exposure, decoder information, and acceptance rule are matched. If the larger code gets a better-calibrated subset of hardware or a slower cycle, the change must be modeled rather than attributed solely to distance.

One favorable pair is finite-size evidence, not an asymptotic proof. Stronger evidence includes several increasing sizes, both logical bases, multiple round counts, stable suppression across time blocks, and a circuit-level model that predicts held-out data. Threshold Theorem explains why a finite crossing, a pseudothreshold, and an asymptotic threshold are distinct concepts.

Because Λ\Lambda is a ratio, uncertainty in the smaller logical error can dominate. For approximately independent estimates with small relative error,

Var⁡(ln⁡Λ^)≈Var⁡(p^d)pd2+Var⁡(p^d+Δd)pd+Δd2.\operatorname{Var}(\ln\widehat\Lambda) \approx \frac{\operatorname{Var}(\widehat p_d)} {p_d^2} + \frac{\operatorname{Var}(\widehat p_{d+\Delta d})} {p_{d+\Delta d}^2}.

Near zero failures this approximation is poor; use a joint likelihood, Bayesian posterior with a declared prior, or profile/bootstrap interval that respects the binomial boundary. Report the confidence interval for Λ\Lambda, not only point estimates for each error rate.

Logical memory should be tested over enough rounds to expose the steady-state regime and rare correlated events. Scaling only at one short duration can be dominated by preparation and readout. Scaling only at a very long duration can saturate near random guessing and lose sensitivity. A two-dimensional grid in distance and round count is more informative than one diagonal slice.

A logical system reaches break-even when it outperforms a declared unencoded or physical reference on a matched task and metric. For an error probability,

GB=prefpLG_{\mathcal B} = \frac{p_{\mathrm{ref}}} {p_L}

is a gain relative to baseline B\mathcal B; GB>1G_{\mathcal B}>1 indicates break-even. The result is incomplete unless the denominator is named.

Possible references include:

  • the best physical carrier among the encoded components;
  • the mean or median constituent carrier;
  • an unencoded carrier exposed for the same wall-clock time;
  • an uncorrected encoded state using the same preparation and measurement;
  • the best available physical implementation of the same gate; or
  • a module-level reference that includes ancillas, controls, and retries.

No one reference answers every engineering question. A fair comparison aligns the input ensemble, elapsed time, measurement rule, acceptance policy, and access to calibration information. Comparing one logical QEC round with one physical gate is not a memory break-even test.

Break-even and error suppression are logically independent. Increasing distance can reduce pLp_L while the logical memory still underperforms the best physical carrier. A fixed code can beat a physical reference without belonging to a scalable family. Neither statement alone establishes a fault-tolerant universal gate set.

For one or a few logical qubits, prepare a spanning set of logical inputs, apply the complete gate gadget, and measure a spanning set of logical observables. This reconstructs or bounds the logical process. Include idle or identity gadgets of equal duration, because a high gate fidelity can coexist with a poor logical memory and vice versa.

Tomography is diagnostic but not scalable. It also depends on trusted logical preparations and measurements unless gate-set methods are used, and its result may be gauge-dependent. The full syndrome record can help diagnose faults but must not be used to tune the gate on the held-out test data.

Logical randomized benchmarking samples sequences from a logical gate group, implements each logical element fault tolerantly, appends an inversion, and fits survival versus sequence length. Under the usual single-mode model,

Psurv(m)=A+BαLm.P_{\mathrm{surv}}(m) = A+B\alpha_L^m.

For a dLd_L-dimensional logical system, the associated average error parameter is

rL=dL−1dL(1−αL).r_L = \frac{d_L-1}{d_L} (1-\alpha_L).

The implementation must specify how a sampled logical Clifford is compiled, where error correction occurs, whether syndrome ancillas are reset, and which records are accepted. The sequence ensemble measures an average over that implemented distribution, not every logical gate and not an arbitrary algorithm.

Logical QEC machinery can preserve hidden syndrome memory between gates, so even gate- and time-independent physical noise can produce nonexponential logical-RB behavior. Residuals, alternative decay models, leakage observables, and sequence-to-sequence variation are therefore essential checks.

An interleaved logical benchmark alternates a target logical gadget with reference randomizers. Its interpretation inherits the assumptions and bounds of physical interleaved RB. A logical cycle benchmark instead targets a parallel protected layer and can expose simultaneous-operation and crosstalk effects. In both cases, compare against an equal-duration logical identity and report the compiled primitive distribution.

Non-Clifford gates require special care. State injection can move error from a resource state into the data; adaptive corrections add reaction latency; factory postselection changes acceptance. Benchmark the delivered injected gate or magic state and separately report factory quality, throughput, and resource cost.

The decoder converts the syndrome record into a logical decision, so changing it changes the measured logical channel. Report:

  • decoder algorithm and version;
  • the exact observations supplied, including soft readout, erasure, leakage, and future-round information;
  • training, calibration, and hyperparameter data;
  • whether decoding is offline, streaming, or deadline constrained;
  • handling of ties, timeouts, and decoder failures; and
  • latency and throughput distributions when they affect execution.

Training and evaluation records must be separated. A decoder fitted to every observed logical outcome can memorize drift or exploit information unavailable in operation. Time-ordered holdout blocks are often more revealing than a random shot split when calibration drifts.

An offline decoder with access to syndrome rounds after the nominal endpoint can estimate a retrospective memory channel. A real-time decoder that must choose a feed-forward action before a deadline measures an operational gate or computation channel. Both are legitimate, but they answer different questions.

Benchmark the decoder jointly with the physical implementation for delivered logical performance, and benchmark it separately on frozen datasets for algorithmic comparison. This prevents a faster but less accurate decoder from being ranked by accuracy alone, or a more accurate offline decoder from being presented as deployable without a latency analysis.

Failures are binomial only under a contract

Section titled “Failures are binomial only under a contract”

If nn independent, identically distributed logical opportunities produce kk failures, the binomial likelihood is

L(pL)∝pLk(1−pL)n−k.\mathcal L(p_L) \propto p_L^k(1-p_L)^{n-k}.

The estimator p^L=k/n\widehat p_L=k/n is simple, but QEC data often violate the independence assumption. Many rounds share one shot, qubits share a burst event, and chronological blocks share calibration drift. Treat the shot, sequence, device region, or time block as a sampling cluster when appropriate. Resampling individual rounds can produce confidence intervals that are much too narrow.

For ordinary binomial data, use Wilson, likelihood-ratio, or exact intervals rather than a symmetric normal interval near zero. Publish failure counts and opportunities so readers can recompute the interval.

With k=0k=0 failures in nn independent opportunities, the one-sided upper confidence bound at confidence 1−α1-\alpha is

pU=1−α1/n≈−ln⁡αn.p_U = 1-\alpha^{1/n} \approx -\frac{\ln\alpha}{n}.

At 95% confidence this is approximately 3/n3/n. To bound a logical failure probability below 10−610^{-6} with zero observed failures requires about three million independent opportunities, and correlation increases the effective requirement. Reporting “zero error” after a few thousand trials is therefore not evidence for a very low rate.

For randomized logical circuits, three sample counts matter:

  • independent logical sequences or circuit instances;
  • repeated shots per fixed sequence; and
  • chronological repetitions across recalibration or drift blocks.

More shots reduce projection noise for the chosen sequences. They do not measure sequence-to-sequence variability or long-time drift. A hierarchical model or nested bootstrap should reflect all sampled levels.

When direct experiment or Monte Carlo cannot reach the target logical error rate, importance sampling, splitting, fault injection, or extrapolation may be used. Report the proposal distribution, likelihood weights, effective sample size, fault-order truncation, and validation region. A simulated logical failure estimate is evidence about the declared model, not direct evidence that hardware has the same tail.

Predeclare sequence lengths, code sizes, primary metrics, stopping rules, and comparisons. Searching many decoder settings, bases, time windows, and postselection thresholds and reporting only the best one creates selection bias. Use a validation set or adjust uncertainty for the selection process.

Reliability Must Travel with Rate and Resources

Section titled “Reliability Must Travel with Rate and Resources”

A logical primitive is useful only if it can deliver outputs. Report the physical and classical footprint, elapsed time, acceptance, and throughput beside reliability. If one trial takes ttrialt_{\mathrm{trial}}, acceptance is aa, and accepted logical failure is pL∣Ap_{L|A}, the rate of accepted correct outputs under stationary serial trials is

Rgood=a(1−pL∣A)ttrial.R_{\mathrm{good}} = \frac{a(1-p_{L|A})} {t_{\mathrm{trial}}}.

Parallel hardware changes the numerator and footprint together. A benchmark should therefore report a performance vector such as

ML=(pL,τL,Rgood,Qphys,VQt,Paccept),\mathbf M_L = \left( p_L, \tau_L, R_{\mathrm{good}}, Q_{\mathrm{phys}}, V_{Qt}, P_{\mathrm{accept}} \right),

not collapse everything into one score. Resource Estimation explains how these measured primitive quantities enter an application-scale cost model.

Classical decoding is part of the resource boundary when it is necessary to interpret or control the logical state. Report acquisition bandwidth, compute hardware, latency tails, backlog, and energy when they constrain the experiment. “Real time” should mean that the full closed loop meets the declared deadline, not merely that one decoder kernel is fast in isolation.

Worked Benchmark: Suppression Without Break-Even

Section titled “Worked Benchmark: Suppression Without Break-Even”

Suppose matched logical-memory experiments produce the following independent window counts:

DistanceOpportunities nnFailures kkCycle time
32.0×1062.0\times10^636000.9 μs0.9\,\mu\mathrm{s}
52.0×1062.0\times10^612001.1 μs1.1\,\mu\mathrm{s}
74.0×1064.0\times10^69601.4 μs1.4\,\mu\mathrm{s}

The per-cycle estimates are

p^L(3)=1.8×10−3,p^L(5)=6.0×10−4,p^L(7)=2.4×10−4.\widehat p_L(3)=1.8\times10^{-3}, \qquad \widehat p_L(5)=6.0\times10^{-4}, \qquad \widehat p_L(7)=2.4\times10^{-4}.

The suppression factors are

Λ^3,5=3.0,Λ^5,7=2.5.\widehat\Lambda_{3,5}=3.0, \qquad \widehat\Lambda_{5,7}=2.5.

Approximate binomial standard errors are 3.0×10−53.0\times10^{-5}, 1.7×10−51.7\times10^{-5}, and 7.7×10−67.7\times10^{-6}, respectively. A real analysis would use clustered intervals if windows share shots or drift blocks, then propagate the joint uncertainty to each ratio.

Per-time rates are approximately

ΓL(3)≈2.0×103 s−1,ΓL(5)≈5.5×102 s−1,\Gamma_L(3)\approx2.0\times10^3\,\mathrm{s}^{-1}, \qquad \Gamma_L(5)\approx5.5\times10^2\,\mathrm{s}^{-1}, ΓL(7)≈1.7×102 s−1.\Gamma_L(7)\approx1.7\times10^2\,\mathrm{s}^{-1}.

Thus both per-cycle and per-time metrics improve with distance. Now compare distance 7 with a physical reference that fails with probability 2.0×10−42.0\times10^{-4} over the same 1.4 μs1.4\,\mu\mathrm{s} interval and matched state ensemble. The break-even gain is

GB=2.0×10−42.4×10−4≈0.83.G_{\mathcal B} = \frac{2.0\times10^{-4}} {2.4\times10^{-4}} \approx0.83.

The logical family shows finite-size error suppression, yet the distance-7 memory has not beaten this physical reference. That is a coherent result, not a contradiction. It also says nothing about protected-gate quality until those gates are benchmarked.

No single protocol is complete. A defensible suite for an encoded processor contains at least:

  1. Syndrome integrity: detection-event rates, correlations, leakage or loss flags, reset behavior, and response to injected faults.
  2. Logical memory: all relevant logical axes over multiple round counts, code sizes, time blocks, and simultaneous-load conditions.
  3. Scaling: matched suppression factors with uncertainty and finite-size caveats.
  4. Break-even: one or more named references aligned in task, time, acceptance, and measurement.
  5. Logical SPAM: preparation and measurement confusion or channel metrics, including rejected outcomes.
  6. Logical gates: identity and target gadgets, random-sequence or cycle-level tests, and controlled fault injection.
  7. Integrated circuits: representative entangling, adaptive, and non-Clifford workloads with end-to-end logical scoring.
  8. Delivery metrics: cycle and reaction time, decoder latency and throughput, physical footprint, acceptance, and correct-output rate.

The suite should include negative controls. Disable correction, use a simpler decoder, omit soft information, or replace a fault-tolerant gadget with a non-fault-tolerant one while keeping the task matched. These ablations identify which part of the stack produced the improvement.

A syndrome response shows sensitivity to faults. Improvement of the decoded logical channel shows correction. Postselected improvement is conditional and must include acceptance.

Per round, per second, per logical Clifford, and per circuit answer different questions. State the opportunity count and operation boundary.

Preserving ∣0L⟩|0_L\rangle can miss dephasing entirely. Use a state ensemble matched to the claimed logical channel.

Treating a fitted exponential as guaranteed

Section titled “Treating a fitted exponential as guaranteed”

Syndrome memory, leakage, coherent errors, drift, and boundary transients can produce several modes or nonexponential behavior. Inspect residuals and vary the fit window.

Decoder weights, neural parameters, postselection thresholds, and calibration choices must be fixed before held-out evaluation. Random splitting can still leak chronological drift.

pL(d+2)<pL(d)p_L(d+2)<p_L(d) compares two encoded systems. Break-even compares an encoded system with a named reference. Neither alone proves a universal logical gate set.

A lower error per cycle can come from a much slower cycle. A lower conditional error can come from rejecting most trials. Report delivery rate and resource cost.

Zero observed failures implies an upper confidence bound controlled by sample size. It is never evidence for exactly zero error.

Future syndrome access and unlimited latency can improve retrospective decoding. An adaptive computation must satisfy the online information and deadline contract.

1. Convert a memory decay to error per cycle

Section titled “1. Convert a memory decay to error per cycle”

A logical parity correlation decays by a fitted factor λ=0.9984\lambda=0.9984 per cycle. Find the independent parity-flip probability per cycle.

Solution

The model uses λ=1−2ϵL\lambda=1-2\epsilon_L, so

ϵL=1−0.99842=8.0×10−4.\epsilon_L = \frac{1-0.9984}{2} =8.0\times10^{-4}.

This conversion is valid only for the stated single-mode binary-flip model. It is not a general conversion from every fitted decay to a logical channel.

2. Compare per-cycle and per-time performance

Section titled “2. Compare per-cycle and per-time performance”

Code A has pL=4×10−4p_L=4\times10^{-4} per 1 μs1\,\mu\mathrm{s} cycle. Code B has pL=2.5×10−4p_L=2.5\times10^{-4} per 2 μs2\,\mu\mathrm{s} cycle. Which is better per cycle and which is better per unit time in the small-error limit?

Solution

Code B has the smaller error per cycle. The approximate rates are

ΓA≈4×10−410−6 s=400 s−1,\Gamma_A \approx \frac{4\times10^{-4}}{10^{-6}\,\mathrm{s}} =400\,\mathrm{s}^{-1}, ΓB≈2.5×10−42×10−6 s=125 s−1.\Gamma_B \approx \frac{2.5\times10^{-4}}{2\times10^{-6}\,\mathrm{s}} =125\,\mathrm{s}^{-1}.

Code B is also better per unit time. The question remains meaningful because the cycle durations were reported.

A benchmark observes no logical failures in 500,000500{,}000 independent trials. Find the approximate one-sided 95% upper bound on pLp_L. Does this establish pL<10−6p_L<10^{-6} at 95% confidence?

Solution

Using pU≈−ln⁡(0.05)/np_U\approx-\ln(0.05)/n gives

pU≈2.996500,000≈6.0×10−6.p_U \approx \frac{2.996}{500{,}000} \approx6.0\times10^{-6}.

No. The data are compatible with probabilities above 10−610^{-6} at that confidence. Correlation would reduce the effective sample size and weaken the bound further.

A detected-code experiment accepts 20% of trials and has conditional logical failure pL∣A=10−3p_{L|A}=10^{-3}. Each trial takes 5 ms. Find the rate of accepted correct outputs and the mean number of attempts per accepted output.

Solution

The correct-output rate is

Rgood=0.2(1−10−3)5×10−3 s≈39.96 s−1.R_{\mathrm{good}} = \frac{0.2(1-10^{-3})}{5\times10^{-3}\,\mathrm{s}} \approx39.96\,\mathrm{s}^{-1}.

The mean number of attempts per accepted output is 1/0.2=51/0.2=5. Reporting only 10−310^{-3} would hide most of the delivery cost.

A distance-3 memory has pL=9×10−4p_L=9\times10^{-4} and distance 5 has pL=3×10−4p_L=3\times10^{-4} under matched conditions. The physical reference has pref=2×10−4p_{\mathrm{ref}}=2\times10^{-4}. State the supported claims.

Solution

The measured suppression factor is

Λ3,5=9×10−43×10−4=3.\Lambda_{3,5} = \frac{9\times10^{-4}}{3\times10^{-4}} =3.

Thus the larger code suppresses the declared logical error under the matched finite-size test. Its break-even gain is

GB=2×10−43×10−4=23<1,G_{\mathcal B} = \frac{2\times10^{-4}}{3\times10^{-4}} =\frac{2}{3}<1,

so it does not beat that physical reference. Neither comparison establishes logical-gate fault tolerance.

6. Diagnose a nonexponential logical-RB curve

Section titled “6. Diagnose a nonexponential logical-RB curve”

A logical randomized benchmark has oscillatory residuals around a one- exponential fit, and the oscillation phase correlates with the syndrome state left by the previous logical Clifford. What assumption is suspect, and what should be reported?

Solution

The single stationary decay-mode assumption is suspect. Syndrome ancillas, decoder state, or frame history can retain memory and act as hidden degrees of freedom coupled to the logical system. Report the residuals, sequence-level data, syndrome-reset policy, and alternative decay model; do not reduce the curve to one logical error parameter without a validated model.

List the controls needed to compare a fault-tolerant logical Hadamard gadget with a non-fault-tolerant implementation.

Solution

Use the same logical code and input ensemble; define both gadget boundaries; match or explicitly report elapsed time, syndrome rounds, preparation and measurement, decoder information, acceptance, and physical operating point; include equal-duration identity controls; use held-out trials; report logical channel metrics and uncertainty; inject representative single faults to test containment; and account for all qubit, controller, and retry resources. A fidelity improvement under nominal noise and a containment improvement under fault injection are complementary results.

Design a minimal experiment that can distinguish boundary SPAM from bulk logical memory decay while testing suppression from distance 3 to distance 5.

Solution

Use matched physical operating conditions and syndrome schedules for both distances; prepare at least the logical XX and ZZ eigenstate ensembles; sample several round counts including zero or a short boundary control and enough long counts to expose bulk decay; interleave distances chronologically to control drift; freeze the decoder before evaluation; record all syndrome, leakage, acceptance, and timing data; fit boundary amplitudes separately from bulk decay; inspect residuals; estimate clustered confidence intervals; and report the suppression-factor interval in both per-cycle and per-time units.

The need to measure complete logical channels and declare denominators is settled methodology. The best scalable protocols are still developing. Logical randomized benchmarking can inherit hidden syndrome memory; non-Clifford resource states can be expensive to test at the multiplicative precision required by fault-tolerant budgets; and extremely low logical error rates make direct sampling costly. Scalable accreditation, rare-event simulation, and benchmark suites for integrated logical processors remain active research.

Experimental capabilities are also changing quickly. Increasing-distance surface-code memories, bosonic break-even results, protected logical gates, real-time decoding, loss-aware operation, and multi-logical-qubit processors probe different parts of the benchmark suite. Their current evidence belongs in Error-Correction Case Studies and the Fault-Tolerant Quantum Computing Frontier.

  • Why Quantum Error Correction Is Possible defines the ideal correctability condition behind the measured logical channel.
  • Decoders develops logical-class inference, decoder families, soft information, model mismatch, and real-time systems metrics.
  • Threshold Theorem distinguishes a finite suppression measurement from an asymptotic accuracy threshold.
  • Fault-Tolerant Gates supplies the error-containment property that nominal logical fidelity alone cannot establish.
  • Surface Code develops repeated checks, spacetime detection events, logical strings, distance, and code-specific suppression.
  • Resource Estimation consumes measured logical failure, timing, throughput, and footprint as inputs to application-scale projections.
  • Randomized Benchmarking derives reference and interleaved RB, nuisance parameters, sampling hierarchy, and failure modes.
  • Cycle Benchmarking develops dressed-cycle estimands and simultaneous-layer protocols that can be lifted to logical cycles.
  • Why Benchmarking Is Hard explains benchmark boundaries, context dependence, drift, compiler freedom, and selection risk across the whole stack.
  • Reporting Standards supplies the provenance, uncertainty, correction, and artifact record for published logical results.
  • Error-Correction Case Studies applies the contract to dated repetition-code, surface-code, and bosonic experiments.
  1. J. Combes, C. Granade, C. Ferrie, and S. T. Flammia, “Logical randomized benchmarking,” arXiv:1702.03688 (2017), arXiv:1702.03688.
  2. A. Ceasura, P. Iyer, J. J. Wallman, and H. Pashayan, “Non-exponential behaviour in logical randomized benchmarking,” arXiv:2212.05488 (2022), arXiv:2212.05488.
  3. E. Knill et al., “Randomized benchmarking of quantum gates,” Physical Review A 77, 012307 (2008), doi:10.1103/PhysRevA.77.012307.
  4. E. Magesan, J. M. Gambetta, and J. Emerson, “Scalable and robust randomized benchmarking of quantum processes,” Physical Review Letters 106, 180504 (2011), doi:10.1103/PhysRevLett.106.180504.
  5. J. Helsen, I. Roth, E. Onorati, A. H. Werner, and J. Eisert, “General framework for randomized benchmarking,” PRX Quantum 3, 020357 (2022), doi:10.1103/PRXQuantum.3.020357.
  6. R. Harper and S. T. Flammia, “Fault-tolerant logical gates in the IBM quantum experience,” Physical Review Letters 122, 080504 (2019), doi:10.1103/PhysRevLett.122.080504.
  7. Google Quantum AI, “Suppressing quantum errors by scaling a surface code logical qubit,” Nature 614, 676–681 (2023), doi:10.1038/s41586-022-05434-1.
  8. Google Quantum AI and Collaborators, “Quantum error correction below the surface code threshold,” Nature 638, 920–926 (2025), doi:10.1038/s41586-024-08449-y.
  9. J. Kelly et al., “State preservation by repetitive error detection in a superconducting quantum circuit,” Nature 519, 66–69 (2015), doi:10.1038/nature14270.
  10. L. Egan et al., “Fault-tolerant control of an error-corrected qubit,” Nature 598, 281–286 (2021), doi:10.1038/s41586-021-03928-y.
  11. C. Ryan-Anderson et al., “Realization of real-time fault-tolerant quantum error correction,” Physical Review X 11, 041058 (2021), doi:10.1103/PhysRevX.11.041058.
  12. J. F. Marques et al., “Logical-qubit operations in an error-detecting surface code,” Nature Physics 18, 80–86 (2022), doi:10.1038/s41567-021-01423-9.
  13. P. Reinhold et al., “Error-corrected gates on an encoded qubit,” Nature Physics 16, 822–826 (2020), doi:10.1038/s41567-020-0931-8.
  14. M. H. Abobeih et al., “Fault-tolerant operation of a logical qubit in a diamond quantum processor,” Nature 606, 884–889 (2022), doi:10.1038/s41586-022-04819-6.
  15. E. H. Chen et al., “Calibrated decoders for experimental quantum error correction,” Physical Review Letters 128, 110504 (2022), doi:10.1103/PhysRevLett.128.110504.
  16. V. V. Sivak et al., “Real-time quantum error correction beyond break-even,” Nature 616, 50–55 (2023), doi:10.1038/s41586-023-05782-6.
  17. Z. Ni et al., “Beating the break-even point with a discrete-variable-encoded logical qubit,” Nature 616, 56–60 (2023), doi:10.1038/s41586-023-05784-4.
  18. S. Lee, M. Yuan, S. Chen, K. Tsubouchi, and L. Jiang, “Efficient benchmarking of logical magic state,” Physical Review Letters 136, 050602 (2026), doi:10.1103/fwjt-mw2c.