Skip to content

Claims, Hype, and Evidence Standards

Quantum information combines mathematics, physics, computer science, engineering, and commercial development. Those communities produce different kinds of evidence. A theorem, a device benchmark, a laboratory demonstration, and a product forecast can all be valuable, but they do not answer the same question.

This page supplies the canonical claim-audit framework for the volume. Its purpose is not automatic skepticism. Its purpose is calibrated belief: state exactly what was shown, under which assumptions, with which comparison, and how far the conclusion travels.

A claim should be small enough to test and specific enough to be wrong. “Quantum computing is faster” is not a testable scientific claim. A usable statement identifies at least:

C=(T, I, O, S, B, R, V, U, Ω, t),\mathcal C = \bigl( T,\, I,\, O,\, S,\, B,\, R,\, V,\, U,\, \Omega,\, t \bigr),

where:

  • TT is the task;
  • II is the input model or promise;
  • OO is the required output;
  • SS is the success criterion;
  • BB is the comparison baseline;
  • RR is the resource ledger;
  • VV is the verification procedure;
  • UU is the uncertainty statement;
  • Ω\Omega is the scope of generalization;
  • tt is the date, software version, calibration epoch, or other time boundary.

This tuple is a checklist, not standard notation. Its value is diagnostic. If one entry is missing, the reader knows exactly what must be supplied before interpreting the headline.

For example, “processor Q solved problem P in ten seconds” remains ambiguous until one knows whether ten seconds includes state preparation, compilation, repeated shots, network latency, decoding, and classical postprocessing; what accuracy was required; which classical algorithm and hardware were used; and how the answer was verified.

Claim audit from task, baseline, resources, verification, uncertainty, and scope to a bounded evidence label

A defensible claim is the intersection of a task contract and an evidence contract. Moving from a component result to a protocol or end-to-end application requires new evidence at each boundary; a stronger adjective cannot perform that promotion.

Use the narrowest label that the evidence supports.

LabelRequired coreWhat it can establishWhat it does not establish by itself
Mathematical theoremdefinitions, assumptions, proofa conclusion follows within a formal modelphysical realizability, efficient constants, or experimental performance
Algorithmic speedup under assumptionsproblem, access model, algorithm, complexity comparisonasymptotic or finite-cost improvement in a declared modeladvantage after data loading, error correction, or hardware overhead
Simulation evidencemodel, numerical method, convergence and validation testsbehavior of the simulated model in the checked regimebehavior of an unmodeled device or the asymptotic limit
Experimental demonstrationapparatus, protocol, calibration, data, uncertaintya component or protocol operated under reported conditionsscalability, fault tolerance, or usefulness
Benchmark resultfixed test, implementation rules, score and uncertaintyperformance on that test at that timeuniversal device quality or application performance
Resource estimatealgorithm, architecture, error model, compiler assumptionsprojected resources under a declared scenariothat the assumed hardware or logical rates will be achieved
Engineering projectionmilestones, dependencies, risks, schedule modela technically motivated patha demonstrated capability or reliable date
Commercial claimnamed product, terms, workload and evidencewhat an organization asserts or offersindependent scientific validation
Speculative applicationmechanism, assumptions, unresolved dependenciesa research direction worth investigatingfeasibility, advantage, demand, or deployment

The labels are not a single ladder. A theorem may be deeper than a benchmark while saying less about present hardware. A benchmark may be highly reproducible while saying little about an application. A resource estimate may be internally rigorous while depending on device parameters that do not yet exist.

A paper can contain several evidence types:

  • a theorem proving hardness under complexity assumptions;
  • a numerical estimate for finite instances;
  • an experimental sampling result;
  • a benchmark against selected classical codes;
  • a projection to a future fault-tolerant machine.

These should remain separate sentences. Combining them into “the experiment proved scalable quantum advantage” silently crosses several unsupported boundaries.

Name the input-output relation. “Optimization,” “simulation,” and “machine learning” are domains, not tasks.

For an optimization problem, specify whether the output must be:

  • an exact optimum;
  • a feasible point;
  • an approximation with a certified ratio;
  • a low objective value relative to a reference;
  • or samples from a distribution over candidate solutions.

For a sampling task, specify the target distribution and distance criterion. For sensing, specify the parameter, prior or local regime, estimator, and loss function. For communication, specify message type, error, rate, secrecy definition, and adversary.

Input access can determine the apparent speedup. Distinguish:

  • an explicit classical list;
  • an oracle;
  • a succinct circuit;
  • a quantum state supplied by another process;
  • a stream of experimental data;
  • a promise distribution;
  • or a family of Hamiltonians with local access.

An algorithm that assumes coherent oracle access to data is not automatically an algorithm for a classical database. The cost of constructing the oracle can dominate the calculation.

State the tolerance before seeing the answer. Common criteria include

Pr⁡(correct)≥1−δ,\Pr(\text{correct})\geq1-\delta,

an additive error

∣f^−f∣≤ϵ,\lvert\widehat f-f\rvert\leq\epsilon,

or distributional accuracy

D(Pout,Ptarget)≤ϵ.D(P_{\mathrm{out}},P_{\mathrm{target}}) \leq \epsilon.

The distance DD, confidence level, treatment of failed runs, and aggregation across instances matter. A mean score can conceal a heavy tail or a subgroup on which the method fails.

A fair comparison holds the task and success criterion fixed. It then reports:

  • the classical algorithm and version;
  • compiler and numerical-library choices;
  • classical hardware and parallelism;
  • quantum compiler and calibration;
  • preprocessing and postprocessing;
  • time limits and stopping rules;
  • whether methods were tuned on the test set;
  • and the date of the comparison.

“Best known classical” is time-indexed, not permanent. Classical algorithms, tensor-network methods, sampling methods, hardware accelerators, and implementation quality can improve after a quantum result is published.

5. Is the device noisy, error-mitigated, error-detected, or fault-tolerant?

Section titled “5. Is the device noisy, error-mitigated, error-detected, or fault-tolerant?”

These regimes are distinct.

RegimeWhat happens to errors?Appropriate claim
noisy physical executionerrors accumulate in physical operationsdevice- or circuit-level experimental result
error mitigationbiased estimators are corrected or extrapolated without encoding a protected logical statemitigated estimate with overhead and assumptions
error detection or postselectionflagged runs are rejectedconditional performance plus acceptance probability
quantum error correctionlogical information is encoded and syndromes are used for recoverylogical performance for a specified code and decoder
fault-tolerant operationerrors are controlled throughout preparation, gates, measurement, and recoveryscalable logical claim within a threshold and architecture model

Postselection can increase conditional fidelity while decreasing yield. Report both. Error mitigation can improve an observable estimate while introducing sampling overhead and model dependence. It is not a substitute term for fault tolerance.

A logical algorithm and a physical implementation are separated by an architecture. A resource estimate should identify:

  • code family and distance;
  • physical error model and correlations;
  • syndrome circuit and decoder;
  • logical gate construction;
  • magic-state or other non-Clifford resources;
  • leakage handling;
  • routing and connectivity;
  • cycle time and classical feedback latency;
  • failure budget assigned to each subroutine.

For a circuit with NlocN_{\mathrm{loc}} relevant failure locations and small logical failure probability pLp_L, a first budgeting estimate is

Pfail≲NlocpL.P_{\mathrm{fail}} \lesssim N_{\mathrm{loc}}p_L.

This is a union-bound style estimate, not a universal equality. Correlated errors, nonuniform locations, retries, verification, and decoder behavior require a more detailed model.

Use a vector, not one headline number:

R=(Nphys,Nlog,D,Nshots,T,E,M,Cclass,Psucc),\mathbf R = \bigl( N_{\mathrm{phys}}, N_{\mathrm{log}}, D, N_{\mathrm{shots}}, T, E, M, C_{\mathrm{class}}, P_{\mathrm{succ}} \bigr),

where entries may denote physical and logical systems, circuit depth, samples, wall-clock time, energy, memory, classical computation, and success probability.

A useful physical-qubit ledger is

Nphys=Ndata+Nsyndrome+Nfactory+Nrouting+Nreserve.N_{\mathrm{phys}} = N_{\mathrm{data}} + N_{\mathrm{syndrome}} + N_{\mathrm{factory}} + N_{\mathrm{routing}} + N_{\mathrm{reserve}}.

The terms depend on architecture. Omitting factories, routing, spare capacity, or interconnects can make a precise-looking estimate incomplete.

Likewise, end-to-end time is not only logical depth:

Tend=Tprepare+Tcompile+Texecute+Tdecode+Tverify+Tclassical.T_{\mathrm{end}} = T_{\mathrm{prepare}} + T_{\mathrm{compile}} + T_{\mathrm{execute}} + T_{\mathrm{decode}} + T_{\mathrm{verify}} + T_{\mathrm{classical}}.

Some terms can overlap; some can be amortized over many instances. The accounting convention must be explicit.

Verification asks whether the claimed output or process is correct relative to a target. The difficulty can grow with system size. If exact classical verification is feasible only for small instances, a large-instance claim needs another strategy:

  • hidden tests or held-out circuits;
  • interactive verification;
  • cross-checks among independent methods;
  • conserved quantities or rigorous bounds;
  • classically tractable subfamilies;
  • extrapolation with a validated error model;
  • local observables with known limits;
  • cryptographic verification assumptions.

Verification and classical simulation are not identical. Some outputs are hard to generate but easy to check; some sampling distributions are hard both to generate and to verify from finite samples.

A benchmark can be scientifically important without being an application. Ask:

  • Is the input representative of a real workload?
  • Is the output useful at the achieved accuracy?
  • Does the workflow include data acquisition and preparation?
  • Is there a competitive non-quantum method?
  • Are reliability, latency, energy, throughput, and cost included?
  • Does a downstream user receive a better decision or product?

“Application-inspired” and “application-ready” are not synonyms.

Asymptotic versus finite-instance advantage

Section titled “Asymptotic versus finite-instance advantage”

Let CQ(n,ϵ)C_Q(n,\epsilon) and CC(n,ϵ)C_C(n,\epsilon) denote quantum and classical costs for problem size nn and accuracy ϵ\epsilon. A speedup ratio is

S(n,ϵ)=CC(n,ϵ)CQ(n,ϵ).\mathcal S(n,\epsilon) = \frac{C_C(n,\epsilon)} {C_Q(n,\epsilon)}.

An asymptotic speedup concerns scaling as nn grows. A finite-instance advantage concerns S>1\mathcal S>1 over a declared set of instances and hardware. Neither implies the other:

  • a better asymptotic scaling can lose at all practical sizes because of constants and overhead;
  • a finite benchmark win can disappear when a classical implementation improves;
  • a sampling separation can be real while lacking a known end-user application.

State whether the comparison concerns query complexity, gate complexity, sample complexity, memory, wall-clock time, energy, cost, or another resource.

The classical comparator is part of the experiment

Section titled “The classical comparator is part of the experiment”

The comparator should receive equivalent information and target the same error. A quantum method that estimates an observable to error ϵ\epsilon should not be compared with a classical method required to reconstruct an entire state. Conversely, a classical heuristic should not be granted privileged preprocessing that is omitted from its runtime.

Classical Information Review supplies the source, channel, code, decoder, access, error, security, and cost ledger for that baseline; this page owns the evidence label and strength of the resulting claim.

Strong reports include several baselines:

  1. a simple reproducible baseline;
  2. a competitive tuned baseline;
  3. a best-known method from the literature;
  4. an ablation showing which quantum component matters.

Negative results are informative. If a classical method wins after careful tuning, that constrains the region in which a quantum advantage claim remains plausible.

Algorithms stated with state-preparation oracles often exclude data loading. End-to-end claims must restore that cost. Similarly, an exponentially large state space does not produce an exponentially large readable output. If the desired output contains mm classical numbers to fixed precision, producing those numbers already incurs output and sampling costs.

Lower Bounds and Limitations owns the formal cross-resource results and their licensed implications; this page retains evidence labels, reporting discipline, uncertainty, matched comparisons, and the language warranted by each level of support.

Uncertainty, Selection, and Reproducibility

Section titled “Uncertainty, Selection, and Reproducibility”

For independent bounded observations XiX_i with sample mean

X‾=1N∑i=1NXi,\overline X = \frac1N \sum_{i=1}^N X_i,

the standard error commonly scales as N−1/2N^{-1/2} under regular conditions. Quantum hardware data can violate simple assumptions through drift, temporal correlations, calibration changes, and adaptive stopping. Report the sampling unit and dependence structure, not only the number of shots.

A complete uncertainty statement distinguishes:

  • shot noise;
  • calibration uncertainty;
  • model mismatch;
  • device drift;
  • finite-instance variation;
  • optimizer randomness;
  • classical numerical error;
  • and uncertainty in extrapolated hardware parameters.

If runs are discarded, report:

Paccept=NacceptedNattemptedP_{\mathrm{accept}} = \frac{N_{\mathrm{accepted}}} {N_{\mathrm{attempted}}}

alongside conditional accuracy. If many circuits, observables, stopping times, or mitigation settings were tried, state how the reported result was selected. Choosing the best-looking configuration after inspection can bias a benchmark even when every individual measurement is correct.

For a durable experimental or benchmark claim, preserve:

  • circuit or pulse descriptions;
  • raw and processed data;
  • calibration records and timestamps;
  • software, firmware, compiler, and dependency versions;
  • random seeds or instance generators;
  • exclusion and postselection rules;
  • uncertainty code;
  • classical baseline code and hardware details;
  • enough metadata to reconstruct figures and tables.

Independent reproduction is stronger than rerunning the same pipeline on the same hidden assumptions. When proprietary constraints prevent full release, list exactly what cannot be audited.

Characterization, Verification, Validation, and Benchmarking

Section titled “Characterization, Verification, Validation, and Benchmarking”

These words are related but not interchangeable.

ActivityQuestionTypical output
characterizationWhat state, process, detector, or noise model describes the component?estimated parameters or reconstructed model
verificationIs the implementation close enough to a specified target?acceptance decision, fidelity bound, or witness
validationDoes the model or system predict the relevant real behavior for its intended use?agreement across held-out observables, regimes, or tasks
benchmarkingHow does performance score on a standardized workload or protocol?metric with implementation rules and uncertainty

No single protocol measures everything. Randomized benchmarking can estimate an averaged gate-performance quantity while reducing sensitivity to state-preparation and measurement errors, but it may not diagnose coherent, correlated, leakage, or context-dependent errors relevant to a specific circuit. Cross-Entropy Benchmarking can detect random-circuit output correlation, but its conversion to fidelity is conditional and the score alone does not certify distributional closeness or computational advantage. Noise in Quantum Information develops these distinctions and the diagnostic hierarchy. Metrics for Quantum Hardware defines the estimands, assumptions, uncertainty, and blind spots of the main characterization protocols. Quantum Volume and Application Benchmarks explains why quantum volume combines several system capabilities in random square circuits but is not an application runtime or a logical-qubit metric.

Application-oriented benchmarks move closer to workloads, yet become more sensitive to compiler choices, instance selection, and classical baselines. A benchmark suite is usually more informative than a universal scalar.

Audit definitions, quantifiers, asymptotic regime, oracle model, error tolerance, and complexity assumptions. A conditional hardness result should name the conjecture or hierarchy assumption. “Provable” does not mean assumption-free.

Audit discretization, truncation, convergence, finite-size effects, numerical conditioning, stochastic error, and comparison with exact limits. Agreement between two codes is stronger when they use independent algorithms and representations. Error Estimates and Convergence Tests provide the numerical foundation.

Audit controls, calibration, blinding where relevant, drift, uncertainty, excluded data, and alternative explanations. Separate direct observables from model-dependent reconstruction. Data Interpretation and Pitfalls develops this logic for quantum-matter probes.

Audit the benchmark specification, implementation freedom, pass threshold, statistical rule, instance generator, compiler, and scope. A metric measured after aggressive circuit simplification may be valid, but the simplification policy is part of the benchmark.

Audit every interface between algorithm, logical circuit, code, architecture, and hardware. Report sensitivity:

∂ln⁡R∂ln⁡θi,\frac{\partial\ln R} {\partial\ln \theta_i},

or at least rerun the estimate across plausible values of each important assumption θi\theta_i. A range or scenario table is usually more honest than a single integer.

Audit dependencies and failure modes. A roadmap milestone should identify what observation would falsify the schedule, which advances can proceed in parallel, and which bottleneck controls the critical path.

Look for a task specification, service-level definition, access conditions, price basis, queueing and latency, data policy, independently reproducible result, and contractually meaningful metric. Press language is not a substitute for methods.

Separate a physical mechanism from a deployable workflow. Label missing algorithms, data interfaces, scaling results, hardware, verification, regulation, or market assumptions. Speculation can be valuable when its dependencies are visible.

Claim: “A quantum processor sampled a target circuit distribution faster than a classical computer.”

Required audit:

  1. define the circuit ensemble and target distribution;
  2. define the distance or score and pass criterion;
  3. report device size, depth, fidelity proxy, shots, and postselection;
  4. document the classical simulator, approximation error, hardware, and date;
  5. explain how large-instance correctness is inferred;
  6. distinguish benchmark advantage from useful application value.

The supported label may be experimental benchmark advantage over named classical implementations. A theorem about asymptotic hardness and an experiment at finite noise are separate supporting claims.

Claim: “Logical error decreases as code distance increases.”

Useful evidence compares matched logical experiments at several distances under a stable protocol. A common phenomenological fit is

pL(d,p)≈A(ppth)(d+1)/2,p_L(d,p) \approx A \left( \frac{p}{p_{\mathrm{th}}} \right)^{(d+1)/2},

but its meaning depends on code family, decoder, circuit, noise, leakage, and fitted range. Decreasing logical error with distance is evidence for below-threshold behavior in that experiment. It is not by itself a demonstration of an arbitrary fault-tolerant algorithm, a universal threshold, or economical scaling.

Here AA is a prefactor, not an additive error floor.

Claim: “The sensor beats the standard quantum limit.”

Define the classical or unentangled comparator and match:

  • total probe number;
  • interrogation time;
  • energy or photon flux;
  • bandwidth;
  • loss and detector efficiency;
  • prior information;
  • dead time and duty cycle;
  • estimator bias and confidence.

A useful gain is a ratio such as

G=RclassicalRquantum,G = \frac{ \mathcal R_{\mathrm{classical}} }{ \mathcal R_{\mathrm{quantum}} },

where R\mathcal R is the same risk or variance under matched resources. G>1G>1 supports an advantage for that task and operating point. It does not establish a universal metrological gain or superior deployed-system cost.

  • A headline names no task or error criterion.
  • “Exponential” refers only to Hilbert-space dimension.
  • The classical comparator is old, untuned, or given a harder output requirement.
  • Hardware time excludes compilation, queueing, shots, decoding, or classical optimization.
  • Only accepted postselected runs are reported.
  • Error bars include shot noise but omit drift or calibration.
  • A component fidelity is multiplied into an application claim without validation.
  • Physical qubits are described as logical qubits.
  • Error mitigation is called fault tolerance.
  • A resource estimate reports one number with no parameter sensitivity.
  • Verification is performed only where the result is classically easy, then assumed at larger sizes without a validated bridge.
  • A benchmark score is treated as universal device rank.
  • A commercial roadmap is cited as evidence that a scientific milestone has occurred.
  • “Potentially useful” becomes “useful” in a summary.
  • A press release broadens the claim beyond the paper’s stated scope.

A concise defensible statement can use this form:

For task TT on input family II, system or method QQ produced output OO with success criterion SS, using resources RR. It outperformed or agreed with baseline BB under matched conditions. Correctness was assessed by verification VV, with uncertainty UU. The result supports claim label LL for scope Ω\Omega as of date/version tt. It does not by itself establish nearest stronger claim.

The last sentence is not self-sabotage. It prevents readers from assigning the result a stronger meaning than the evidence carries.

A paper proves that a circuit family is hard to sample classically unless the polynomial hierarchy collapses, numerically simulates 40-qubit instances, and runs 80-qubit instances on hardware. Assign labels to the three results.

Solution

The hardness statement is an algorithmic or complexity-theoretic result under explicit assumptions. The 40-qubit calculations are simulation evidence for finite instances and can support validation of the implementation in that regime. The 80-qubit runs are an experimental demonstration and possibly a benchmark result if a fixed score and comparison protocol are supplied.

None alone proves the others. The theorem may not model finite noise; the simulation may not reach the experimental regime; and the experiment needs a verification bridge before it inherits the theoretical hardness conclusion.

Rewrite “our quantum optimizer is 100 times faster” as a claim contract. List the missing fields rather than inventing values.

Solution

A repaired skeleton is:

For optimization task [missing], on instance distribution [missing] at sizes [missing], the quantum workflow returned outputs meeting [objective/feasibility/approximation criterion] in [time-accounting convention]. This was 100 times lower than [named classical implementation and hardware] under [matched input, accuracy, tuning, and stopping rules]. The quantum cost included [state preparation, circuit execution, shots, optimization, and postprocessing status]. Results were verified by [method] with [uncertainty] as of [software, hardware, and date].

The original sentence omits the task, instances, success criterion, time boundary, baseline, resource ledger, verification, uncertainty, and date.

Method A has conditional success probability 0.990.99 and acceptance probability 0.020.02. Method B always returns an answer with success probability 0.900.90. If attempts have equal cost and rejected runs are repeated, compare the expected number of correct accepted answers per attempt.

Solution

For A,

P(correct accepted)=0.02×0.99=0.0198.P(\text{correct accepted}) = 0.02\times0.99 = 0.0198.

For B,

P(correct output)=0.90.P(\text{correct output}) =0.90.

A has much higher conditional fidelity but far lower throughput per attempt. Which method is preferable depends on whether rejection is allowed, attempt cost, latency, and the consequence of an incorrect answer. Reporting only 0.990.99 would hide the dominant resource cost.

A logical computation contains 10810^8 locations. Using the simple bound

Pfail≲NlocpL,P_{\mathrm{fail}} \lesssim N_{\mathrm{loc}}p_L,

what logical error per location is required to keep total failure below 1%1\%?

Solution

Require

108pL≤10−2,10^8p_L \leq 10^{-2},

so

pL≤10−10.p_L \leq 10^{-10}.

This is a rough budgeting bound. A real estimate should allocate failure across nonuniform operations, state factories, memory, measurement, routing, retries, and correlated faults.

A processor obtains a high quantum-volume score. State three conclusions that the score can support and three it cannot support by itself.

Solution

It can support:

  1. successful execution of the benchmark’s random square circuits up to a reported width and depth;
  2. a system-level comparison under the benchmark’s compilation and pass rules;
  3. evidence that gate quality, connectivity, and compilation jointly permit that circuit regime.

It cannot by itself establish:

  1. performance on a particular chemistry, optimization, or simulation workload;
  2. a logical error rate or fault-tolerant capability;
  3. universal superiority over another processor or classical computer.

Those conclusions require workload-specific, logical, or comparative evidence.

A projected runtime scales as

T∝NTd2tcPfactory,T \propto \frac{N_Td^2t_c}{P_{\mathrm{factory}}},

where NTN_T is a non-Clifford count, dd is code distance, tct_c is cycle time, and PfactoryP_{\mathrm{factory}} is the number of parallel factories. Which parameters deserve the most careful sensitivity analysis?

Solution

The logarithmic sensitivities are

∂ln⁡T∂ln⁡NT=1,∂ln⁡T∂ln⁡d=2,∂ln⁡T∂ln⁡tc=1,∂ln⁡T∂ln⁡Pfactory=−1.\begin{aligned} \frac{\partial\ln T}{\partial\ln N_T}&=1, & \frac{\partial\ln T}{\partial\ln d}&=2, \\ \frac{\partial\ln T}{\partial\ln t_c}&=1, & \frac{\partial\ln T}{\partial\ln P_{\mathrm{factory}}}&=-1. \end{aligned}

The quadratic dependence makes uncertainty in dd especially influential within this simplified model. However, dd is itself chosen from a physical error model and failure budget, while factory parallelism consumes qubits and routing. A credible estimate varies coupled assumptions rather than changing one parameter while holding an inconsistent architecture fixed.

  • What Is Quantum Information? defines the task-and-resource viewpoint used by this audit.
  • Quantum Information and Computation maps physical, logical, protocol, and evidence layers.
  • Quantum Algorithms and Complexity specializes this page’s evidence discipline to a ten-field algorithm and complexity claim record.
  • Quantum Software Stack traces source, target state, dispatched executable, measurement records, postprocessing, and provenance so an implementation claim can be audited at the correct boundary.
  • Circuit Intermediate Representations shows which specification, profile, target epoch, output schema, and pass record are needed to audit what an executable actually denotes.
  • Resource Estimation Tools supplies the layered contract, uncertainty ledger, validation ladder, and reproducibility bundle required to keep a model-based estimate distinct from a hardware forecast.
  • Why Benchmarking Is Hard shows why a score depends on its task ensemble, implementation policy, device epoch, reference, statistics, resources, and scope.
  • Reporting Standards supplies the common disclosure record and result-type profiles needed to preserve the evidence behind an audited claim.
  • Claims and Evidence Checklist turns this framework into a gate-based review card, comparator test, domain module, and bounded disposition.
  • Negative Results and Limitations distinguishes formal barriers, empirical nulls, comparator reversals, resource bottlenecks, and claims that remain open.
  • Quantum Illumination applies this audit to a particularly delicate sensing claim by separating an ideal 6 dB error-exponent theorem, structured receivers, matched laboratory advantages, and unestablished field-radar capability.
  • Algorithmic Benchmarking states what must be measured before a circuit run becomes evidence about a complete algorithmic task, and why that evidence alone is not an advantage demonstration.
  • Verification of Quantum Advantage gives the dedicated audit for correctness, conditional hardness, a dated classical frontier, matched statistical separation, adversarial challenge, and independent reproduction.
  • Cross-Entropy Benchmarking shows why a measured random-circuit correlation, a model-dependent fidelity parameter, and computational-advantage evidence are three different claims.
  • Evidence Labels supplies site-wide epistemic labels.
  • Page Status Labels distinguishes settled, active, conjectural, speculative, and controversial material.
  • Data Interpretation and Pitfalls develops cross-probe causal alternatives, mixtures, contacts, surface–bulk distinctions, and history dependence.
  • Measurement Tomography covers reconstruction of effective measurement operators from calibration data.
  • Living Review Archive separates durable AMO concepts from time-sensitive platform records.

The audit principles on this page are durable, but benchmark protocols, classical baselines, hardware records, software stacks, and preferred terminology evolve. Named performance comparisons should be dated and revisited. New evidence should strengthen or narrow a claim by changing an explicit field in the claim contract, not by silently replacing its label.

  • National Academies of Sciences, Engineering, and Medicine, Quantum Computing: Progress and Prospects, National Academies Press, 2019, doi:10.17226/25196.
  • J. Preskill, “Quantum computing in the NISQ era and beyond,” Quantum 2, 79, 2018, doi:10.22331/q-2018-08-06-79.
  • J. Eisert, D. Hangleiter, N. Walk, I. Roth, D. Markham, R. Parekh, U. Chabaud, and E. Kashefi, “Quantum certification and benchmarking,” Nature Reviews Physics 2, 382–390, 2020, doi:10.1038/s42254-020-0186-4.
  • M. Kliesch and I. Roth, “Theory of quantum system certification,” PRX Quantum 2, 010201, 2021, doi:10.1103/PRXQuantum.2.010201.
  • A. Hashim, L. B. Nguyen, N. Goss, B. Marinelli, R. K. Naik, T. Chistolini, J. Hines, et al., “Practical introduction to benchmarking and characterization of quantum computers,” PRX Quantum 6, 030202, 2025, doi:10.1103/PRXQuantum.6.030202.
  • A. W. Cross, L. S. Bishop, S. Sheldon, P. D. Nation, and J. M. Gambetta, “Validating quantum computers using randomized model circuits,” Physical Review A 100, 032328, 2019, doi:10.1103/PhysRevA.100.032328.
  • T. Proctor, K. Rudinger, K. Young, E. Nielsen, and R. Blume-Kohout, “Measuring the capabilities of quantum computers,” Nature Physics 18, 75–79, 2022, doi:10.1038/s41567-021-01409-7.
  • G. D. Kahanamoku-Meyer, S. Choi, U. V. Vazirani, and N. Y. Yao, “Classically verifiable quantum advantage from a computational Bell test,” Nature Physics 18, 918–924, 2022, doi:10.1038/s41567-022-01643-7.
  • C. Gidney and M. Ekerå, “How to factor 2048 bit RSA integers in 8 hours using 20 million noisy qubits,” Quantum 5, 433, 2021, doi:10.22331/q-2021-04-15-433.
  • E. Knill and R. Laflamme, “Theory of quantum error-correcting codes,” Physical Review A 55, 900–911, 1997, doi:10.1103/PhysRevA.55.900.
  • A. Montanaro, “Quantum algorithms: an overview,” npj Quantum Information 2, 15023, 2016, doi:10.1038/npjqi.2015.23.
  • K. Bharti, A. Cervera-Lierta, T. H. Kyaw, T. Haug, S. Alperin-Lea, A. Anand, M. Degroote, et al., “Noisy intermediate-scale quantum algorithms,” Reviews of Modern Physics 94, 015004, 2022, doi:10.1103/RevModPhys.94.015004.
  • A. Miessen, D. J. Egger, I. Tavernelli, and G. Mazzola, “Benchmarking digital quantum simulations above hundreds of qubits using quantum critical dynamics,” PRX Quantum 5, 040320, 2024, doi:10.1103/PRXQuantum.5.040320.

Trustworthy quantum-information reporting begins by labeling the evidence correctly. A theorem, simulation, experiment, benchmark, resource estimate, projection, commercial statement, and speculative application carry different warrants. Every consequential claim should expose its task, input, output, success criterion, baseline, resources, verification, uncertainty, scope, and date.

The practical discipline is simple: compare like with like, count the full workflow, report rejected runs and uncertainty, preserve reproducibility metadata, and state the nearest stronger conclusion that has not yet been established.