Skip to content

Claims and Evidence Checklist

The claims and evidence checklist is a practical decision procedure for turning a headline into an assessable statement, matching that statement to the evidence actually supplied, and recording the narrowest conclusion that survives review.

It is designed for reading papers, technical reports, benchmark releases, resource estimates, vendor announcements, and proposals. Its output is not a numerical credibility score. It is a disposition:

DispositionMeaning
supported as writtenthe evidence directly supports the stated claim within its declared scope
supported after narrowinga more limited statement is supported once an overbroad word, population, baseline, or inference is removed
unresolvedthe claim may be true, but required evidence or disclosure is missing
contradictedthe supplied evidence is materially inconsistent with the claim
not assessablethe claim is too undefined, inaccessible, or unfalsifiable to evaluate

Claims, Hype, and Evidence Standards owns the conceptual claim contract, evidence labels, and reasons the fields matter. Reporting Standards owns the report, artifact, provenance, and reproducibility contract. The present page owns the review worksheet: which questions to ask, where to stop, how to test comparator compatibility, and how to write a bounded decision record.

Do not audit a paper, company, platform, or research program as one indivisible object. Extract each consequential sentence separately. One source may contain:

  • a proved asymptotic theorem;
  • a numerical result for selected instances;
  • a laboratory demonstration;
  • a benchmark against one implementation;
  • a future hardware projection;
  • an application or commercial interpretation.

Those claims can receive different dispositions. A rigorous theorem does not automatically validate the experimental implementation, and a strong experiment does not automatically validate a projection built on additional assumptions.

Use one of four field states:

Field stateUse
reportedthe value, definition, and source are supplied
not applicablethe field genuinely does not apply, with a reason
unavailablethe reviewer cannot obtain the information
withheldthe source intentionally restricts it and states why

Blank is not a field state. It can mean zero, unknown, forgotten, proprietary, or not applicable. Preserve the distinction because it determines whether the claim is false, unresolved, or simply outside the available evidence.

A points system is tempting but usually misleading. Ten well-reported calibration fields cannot repair a classical baseline that solves a different task. Open code cannot repair an output that was never verified. A large sample count cannot repair pseudoreplication when all shots came from one device state.

Use gates instead:

  1. Is there an exact, falsifiable claim?
  2. Are candidate and baseline solving the same task to the same quality?
  3. Does the evidence identify the relevant population and uncertainty?
  4. Can the output or inferred process be checked?
  5. Does the public sentence stay inside the tested scope?

If a gate fails, stop the stronger conclusion and record what narrower claim, if any, still passes.

Complete one card for every headline or decision-relevant claim.

FieldEntry to record
claim IDstable label connecting the review, source text, figures, and artifacts
exact source sentenceverbatim claim, location, version, and date
proposed evidence labeltheorem, simulation, experiment, benchmark, resource estimate, projection, commercial claim, or speculation
taskexact input-output relation
input populationinstances, devices, circuits, states, seeds, time window, or promise distribution
output and qualityrequired answer, tolerance, confidence, fidelity, approximation, or utility
comparatoralgorithm, device, protocol, theory value, or no-comparison reference
resourcestime, samples, qubits or modes, energy, memory, control, classical work, retries, and infrastructure
evidencedata, proof, simulation, calibration, benchmark record, or cited result
verificationhow correctness or model adequacy is checked
uncertaintyestimand, estimator, independent unit, interval, and systematic contributions
exclusionspostselection, failed jobs, omitted devices, data cleaning, and stopping rules
scopeconditions, populations, dates, and versions to which the result applies
nearest stronger claimplausible conclusion that this evidence does not yet establish
dispositionsupported, narrowed, unresolved, contradicted, or not assessable
next evidencesmallest additional test or disclosure that could change the disposition

The “nearest stronger claim” field is especially useful. It prevents a reader from silently upgrading “component demonstration” to “system capability,” “finite-instance benchmark” to “asymptotic advantage,” or “below-threshold memory” to “fault-tolerant computation.”

Ask what object actually has the reported property:

  • a physical gate or an inferred gate model;
  • one calibrated qubit pair or the processor;
  • one circuit family or an algorithm;
  • a simulated Hamiltonian or a material;
  • an encoded memory or a logical computer;
  • a conditional accepted data set or all attempted runs.

“The system achieved 99.9%99.9\% fidelity” is not yet testable. Fidelity is a relation between specified objects under a specified convention. The sentence must name the state, channel, gate set, estimator, and treatment of state preparation and measurement.

“Quantum chemistry,” “optimization,” “machine learning,” and “networking” are domains. A task has inputs, outputs, and a success rule. Examples include:

  • estimate a ground-state energy within ϵ=1 mEh\epsilon=1\ {\rm m}E_h for a declared molecular Hamiltonian;
  • return a feasible cut whose objective is at least a declared fraction of a reference value;
  • sample from a target distribution within a named distance;
  • distribute one Bell pair above a fidelity threshold between named endpoint memories;
  • estimate a magnetic field with specified bandwidth, spatial resolution, and mean-squared error.

If the source does not permit this conversion, mark the claim not assessable, not false.

Write two sentences:

  1. Observation: what quantity was measured or computed?
  2. Interpretation: what physical, computational, or practical conclusion is inferred?

For example, a randomized benchmark may observe a fitted sequence-decay parameter. Interpreting it as average gate infidelity requires a protocol model. Interpreting that gate infidelity as application performance requires another bridge. Each bridge carries assumptions and should be audited separately.

Hardware calibration, cloud service terms, software stacks, and best-known classical methods change. Record:

  • acquisition dates and calibration epoch;
  • software and compiler versions;
  • source revision, correction, or retraction status;
  • date on which the comparator search was performed;
  • date after which the claim should be rechecked.

A dated benchmark can remain correct historical evidence even after it stops describing current capability.

For every comparative claim, classify each row as matched, reconciled, or incompatible.

Comparison fieldQuestions
taskare both methods solving the same input-output problem?
instance populationare sizes, distributions, promises, and excluded cases aligned?
output qualityare accuracy, fidelity, approximation, confidence, and failure probability matched?
access modeldo both receive the same data, oracle, state, Hamiltonian, or side information?
tuning budgetare hyperparameter search, compilation, calibration, and manual intervention counted?
stopping ruledo both stop on the same quality or resource condition?
time boundaryare setup, queueing, execution, decoding, verification, and postprocessing treated consistently?
hardware boundaryare processors, accelerators, networking, control, and classical hosts included consistently?
repetitionsare failed attempts, shots, seeds, restarts, and confidence targets aligned?
energy and infrastructureare cooling, control, communication, and amortization conventions disclosed?
date and implementation qualityis the comparator current, competently implemented, and versioned?

Matched means the same convention is used. Reconciled means a documented conversion or sensitivity analysis places the quantities on a common basis. Incompatible means the headline ratio is not interpretable as the stated comparison.

Keep conditional and delivered performance separate

Section titled “Keep conditional and delivered performance separate”

Suppose NattN_{\rm att} attempts produce NaccN_{\rm acc} accepted records, of which NgoodN_{\rm good} satisfy the task criterion. Define

pacc=NaccNatt,pgood∣acc=NgoodNacc.p_{\rm acc} = \frac{N_{\rm acc}}{N_{\rm att}}, \qquad p_{\rm good|acc} = \frac{N_{\rm good}}{N_{\rm acc}}.

The probability of delivering a good result per attempt is

pdel=paccpgood∣acc=NgoodNatt.p_{\rm del} = p_{\rm acc}p_{\rm good|acc} = \frac{N_{\rm good}}{N_{\rm att}}.

Conditional quality can improve while delivered quality declines. This applies to error-detected circuits, heralded links, filtered sensing records, optimizer restarts, and any protocol with abstention.

For a wall-clock attempt rate RattR_{\rm att}, the delivered good-result rate is

Rdel=Rattpdel.R_{\rm del} = R_{\rm att}p_{\rm del}.

A comparison may legitimately optimize conditional fidelity, delivered rate, or a utility combining both. It must say which.

Let CfixedC_{\rm fixed} be one-time setup cost, CattC_{\rm att} the cost per attempt, CaccC_{\rm acc} additional work per accepted record, and NdelN_{\rm del} the number of useful delivered outputs. A transparent average cost is

Cdel=Cfixed+NattCatt+NaccCaccNdel.C_{\rm del} = \frac{ C_{\rm fixed} + N_{\rm att}C_{\rm att} + N_{\rm acc}C_{\rm acc} }{ N_{\rm del} }.

The cost can represent time, energy, money, samples, or a resource vector. A fixed cost may be amortized only over the declared workload. Compilation for one instance cannot be divided by a hypothetical million future instances while the baseline is charged setup per current run.

For a claimed speedup or superiority:

  • identify the comparator search procedure and cutoff date;
  • include specialized as well as generic methods;
  • report implementation language, libraries, precision, and hardware;
  • give both methods a defensible tuning budget;
  • inspect scaling across instances rather than one favorable point;
  • include verification and output-conversion costs;
  • disclose whether the candidate influenced which baseline or instances were selected.

“Best known” is a review conclusion supported by a dated search, not a permanent property of an algorithm.

Shots are often conditional repeats within one higher-level unit. The independent unit may instead be:

  • a randomly drawn circuit;
  • a problem instance;
  • a device or fabrication lot;
  • a calibration epoch;
  • a day of acquisition;
  • an optimizer seed;
  • a biological or material sample;
  • a network path or weather interval.

One million shots from one circuit on one calibration do not establish variation across circuits or days. The uncertainty calculation should match the population named in the claim.

The estimand is the quantity the claim intends to describe. The estimator is the rule applied to observed records. Record:

  • the target population;
  • weighting across instances or devices;
  • whether the statistic is mean, median, quantile, worst case, or selected best;
  • the fit model and parameter convention;
  • interval type and coverage or credibility level;
  • known systematic effects and model discrepancy.

A narrow bootstrap interval around a biased estimator does not include model error merely because it has many resamples.

List the choices explored before the reported result was selected:

  • qubit subsets and couplers;
  • circuit depths and benchmark families;
  • data-cleaning thresholds;
  • optimizer seeds and ansatzes;
  • mitigation parameters;
  • time windows;
  • metrics and fit ranges;
  • problem instances and classical baselines.

Then ask whether the uncertainty or validation set accounts for that search. Reporting the winning configuration as though it had been prespecified can produce severe winner’s-curse bias.

Useful stress tests include:

  • held-out circuits, instances, observables, or acquisition periods;
  • injected faults or synthetic data with known truth;
  • classically tractable limits;
  • conservation laws and sum rules;
  • randomized labels or null inputs;
  • independent analysis pipelines;
  • hardware-disabled or correction-disabled controls;
  • blinded parameter choices.

A method tuned and evaluated on the same small set demonstrates fit to that set. Generalization requires new evidence.

Verification must reach the claim’s output:

ClaimMinimum relevant check
state preparationinformationally relevant observables, witnesses, or validated tomography
gate or channelprotocol-appropriate process evidence with model and SPAM limitations
algorithmic answerobjective, residual, certificate, known solution, or independent cross-check
sampling distributionfinite-sample test tied to a declared distance or score
simulatorconvergence, discretization, limiting cases, and independent method where possible
resource estimatearithmetic reproduction plus sensitivity to architecture assumptions
sensorcalibrated response, uncertainty budget, bandwidth, and matched classical limit
secure protocolproof assumptions plus implementation parameters and finite-size analysis
logical memorylogical channel, matched reference, decoder, acceptance, and elapsed time

A proxy metric may be useful, but it supports a proxy claim until its connection to the target is validated.

Many overclaims occur after the measurement is complete. Ask whether the conclusion moves:

  • from selected instances to an instance family;
  • from one device to a platform;
  • from a component to a system;
  • from short duration to sustained operation;
  • from simulated noise to physical noise;
  • from average behavior to worst-case guarantees;
  • from a benchmark to an application;
  • from a present result to a deployment forecast.

Each move needs an argument, data, or model. If none is supplied, retain the local observation and narrow the interpretation.

SourceWhat it can directly establish
proof or formal papertheorem under its definitions and assumptions
original experimental article and datareported apparatus, protocol, observations, and analysis
benchmark releasescore under that benchmark version and rules
resource-estimate studyprojection under its algorithm and architecture scenario
review articlesynthesis and context up to its search date
press release or company pagethat the organization made the stated claim
news reporta secondary account and route to primary sources
roadmapintended milestones, dependencies, and organizational expectations

Peer review raises neither every claim to fact nor every preprint to unreliable. Record status, correction history, methods access, and independent support. A press release may be the primary source for a product announcement but remains secondary evidence for the underlying scientific performance.

For every number in the final sentence, ask:

  1. Where was it first measured, proved, or computed?
  2. Is the cited source the original source?
  3. Did a later source change the definition or denominator?
  4. Is the quotation faithful to the source’s caveats?
  5. Has a correction, retraction, superseding version, or improved comparator appeared?

Circular citation can make one estimate look independently confirmed when several reports all trace to the same unpublished assumption.

Record interests without using them as a verdict

Section titled “Record interests without using them as a verdict”

Funding, employment, equity, benchmark authorship, and access restrictions can shape incentives and available evidence. Record them. Do not replace methodological review with an inference about motive. A conflicted source can provide excellent evidence; an apparently disinterested source can provide poor evidence. The practical response is stronger transparency and independent checking.

Use the common gates first, then add the relevant module.

  • Is the advantage asymptotic, finite-instance, or measured wall-clock?
  • What input-access model is assumed?
  • Are data loading, state preparation, repeated shots, decoding, and output extraction included?
  • Does the classical baseline solve the same task to the same quality?
  • Is the hardness statement worst-case, average-case, conjectural, or tied to the sampled instances?
  • Can the output be verified at the claimed scale?

Stop an advantage claim if the baseline is stale, quality is unmatched, or the quantum output is not checked. A circuit can remain a valuable experimental demonstration after that stronger label is removed.

  • Which qubits, gates, topology, concurrency pattern, and calibration epoch?
  • Is the metric directly observed or inferred through a model?
  • Are leakage, crosstalk, drift, and SPAM inside the protocol’s scope?
  • Does the benchmark sample the operations needed by the target workload?
  • Are uncertainty and device-selection procedures reported?
  • Does a component metric predict a system-level outcome on held-out circuits?

Stop a platform-wide claim when only a selected subset was measured.

  • Is the target a mathematical model, a physical system, or both?
  • Which discretization, truncation, boundary condition, and solver tolerance?
  • Are convergence and finite-size studies shown?
  • Is sampling or optimization error separated from model discrepancy?
  • Are sign problems, entanglement growth, or conditioning restricting the tested regime?
  • Is the physical interpretation validated against independent observables?

Agreement with another implementation of the same approximation checks code more directly than it checks nature.

  • Which algorithm, compiler, code, physical error model, decoder, cycle time, and logical failure budget?
  • Are factories, routing, storage, interconnects, retries, and classical control included?
  • Which quantities are measured, assumed, extrapolated, or targeted?
  • How does the answer change under plausible parameter ranges?
  • Is the estimate optimized for qubits, runtime, space-time volume, energy, or another objective?
  • Are future hardware dependencies labeled as dependencies?

Use Resource Estimation Tools for the full scenario and sensitivity contract.

  • Is the result detection, postselection, correction, break-even, below-threshold scaling, or a logical operation?
  • Which logical states and Pauli sectors were tested?
  • What is the code, syndrome circuit, decoder, and frame convention?
  • Is the metric per cycle, per elapsed time, or both?
  • What physical or uncorrected reference defines break-even?
  • Are rare correlated events, leakage, decoder latency, and rejected runs included?

Error-Correction Case Studies provides the matched logical-channel comparison.

  • What parameter, operating range, bandwidth, spatial resolution, and loss function?
  • Which resources define the standard quantum limit?
  • Are preparation time, dead time, readout, calibration, and estimator bias included?
  • Is sensitivity distinguished from precision, accuracy, and detection limit?
  • Are technical noise and drift included in the same bandwidth?
  • Does the quantum enhancement survive loss and the complete measurement cycle?

A squeezed input is a resource demonstration; it becomes a sensor advantage only through a matched end-to-end uncertainty comparison.

  • What object is delivered: clicks, raw bits, entanglement, a teleported state, secret key, or a remote gate?
  • Are rates conditional, sifted, secret, or end-to-end delivered?
  • What trust, adversary, authentication, and composability model applies?
  • Are finite-size effects, loss, side channels, memory age, and classical latency included?
  • Is the comparison direct transmission, another network architecture, or a classical service?
  • Are security proof assumptions connected to measured device parameters?

An optical event and a delivered secure or logical service are different evidence levels.

A complete disposition can use this form:

For [task and population], under [versioned conditions], [method] estimated [quantity] as [value and uncertainty] relative to [baseline], using [resource boundary] and [verification]. This supports [evidence label] over [scope]. It does not yet establish [nearest stronger claim] because [specific missing bridge].

The template is not a demand for one very long public sentence. It is a test: every clause should be recoverable from the report.

Suppose a source says “the quantum optimizer is 100 times faster,” but the data show a 100-fold reduction in device execution time for selected instances, excluding compilation, queueing, repeated shots, and classical postprocessing. The evidence does not necessarily contradict the measured device-time ratio. It contradicts or leaves unresolved the end-to-end optimizer claim. A defensible narrowing is:

On the selected instances, the reported quantum device-execution segment was 100 times shorter than the declared baseline segment; no end-to-end time-to-solution comparison was established.

Preserving the valid local result is more informative than replacing every overstatement with “false.”

Do not bury the review in a catalog of every desirable experiment. Name the smallest missing item that blocks the claim:

  • matched output quality;
  • a current competitive baseline;
  • an independent test set;
  • unconditional acceptance and retry accounting;
  • a logical rather than component metric;
  • uncertainty over devices rather than shots;
  • verification beyond the classically tractable regime;
  • sensitivity to one dominant architecture assumption.

This turns the checklist into a research plan.

  1. Copy the exact claim.
  2. Assign a provisional evidence label.
  3. Identify task, output quality, comparator, and date.
  4. Find the primary source.
  5. Check whether the nearest stronger claim is being implied.

Use not assessable or unresolved when these fields are absent. Do not fill them with charitable guesses.

  1. Complete the claim card.
  2. Build the comparator compatibility table.
  3. Identify the independent unit and uncertainty semantics.
  4. Trace exclusions, retries, and selected configurations.
  5. Locate verification and robustness tests.
  6. Write the bounded disposition.
  1. Recompute key values from released records.
  2. Inspect executable versions, calibration context, and provenance.
  3. Reproduce the analysis where possible.
  4. Test alternative fit, stopping, and exclusion choices.
  5. Update the classical or physical baseline.
  6. Seek independent replication or a genuinely different validation method.
  7. Record which unavailable artifacts limit the conclusion.

A full review can remain inconclusive. Honest unresolved status is an outcome, not a failed review.

  • The exact sentence, source location, version, and date are recorded.
  • The grammatical subject is a specific object, protocol, or population.
  • The task, input, output, and success criterion are explicit.
  • Observation, model-based inference, and projection are separate.
  • The provisional evidence label matches the source type.
  • Candidate and baseline solve the same task.
  • Output quality and failure probability are aligned.
  • Access models and side information are aligned.
  • Setup, tuning, execution, retries, and postprocessing use one boundary.
  • Hardware, parallelism, precision, and versions are reported.
  • The comparator search is competent and dated.
  • The number traces to a primary source or its dependency is explicit.
  • The measured or computed output is checked at the level claimed.
  • Proxy metrics are not silently promoted to task success.
  • Holdout tests, negative controls, or independent methods are present where generalization is claimed.
  • Corrections, retractions, and superseding results were checked.
  • The estimand, estimator, and independent unit are named.
  • Interval type, level, and uncertainty sources are reported.
  • Device, instance, seed, and analysis selection are disclosed.
  • Rejected runs, failed jobs, and stopping rules are counted.
  • Model discrepancy and systematic effects are not hidden inside shot error.
  • Resources are a vector rather than one convenient count.
  • Conditional quality and acceptance probability are both visible.
  • Cost and rate are normalized to a useful delivered output.
  • Amortization uses the declared workload.
  • Verification and classical control costs remain inside the boundary.
  • The tested population, operating regime, and time boundary are explicit.
  • Component, system, benchmark, application, and projection levels remain distinct.
  • The nearest stronger unestablished claim is stated.
  • Conflicts and access restrictions are recorded without replacing methodological review.
  • The final disposition names one decisive reason and the next useful evidence.
  • Scoring the source or institution instead of auditing a sentence.
  • Treating peer review as a substitute for checking the inference.
  • Calling missing evidence proof that the claim is false.
  • Reconstructing unstated assumptions in the most favorable way.
  • Comparing candidate and baseline at different output quality.
  • Counting shots as independent evidence about devices, days, or instances.
  • Accepting a proxy because it correlates with the target in a different regime.
  • Reporting only conditional quality after postselection.
  • Treating an error bar as complete without interval semantics.
  • Ignoring the search over qubits, seeds, metrics, and fit windows.
  • Calling a resource projection a demonstrated capability.
  • Treating a commercial roadmap date as a scientific confidence interval.
  • Letting a disclosure checklist obscure one fatal task or comparator mismatch.
  • Writing “more evidence is needed” without naming the smallest decisive test.

A report states, “The processor has 99.8%99.8\% fidelity,” and cites an interleaved randomized benchmark on one selected two-qubit gate pair. Give the provisional disposition and a bounded replacement sentence.

Solution

The processor-wide claim is supported after narrowing at best. The source object is one selected gate pair, the protocol estimates a benchmark-specific average error parameter, and selection, SPAM, leakage, crosstalk, parallel operation, and other processor gates may be outside scope.

A bounded sentence is:

During the reported calibration epoch, interleaved randomized benchmarking on the selected two-qubit pair produced the reported 99.8%99.8\% protocol-derived fidelity estimate. This is evidence about that gate pair under the benchmark model, not a processor-wide fidelity or application success probability.

A quantum method returns feasible solutions with mean objective 120 in 0.5 s0.5\ {\rm s}. A classical method certifies the optimum, mean objective 128, in 4 s4\ {\rm s}. May the source claim an eightfold speedup?

Solution

Not from these data. The outputs have different quality and possibly different tasks: one returns a feasible point of quality 120, while the other certifies the optimum of quality 128. The runtime ratio

40.5=8\frac{4}{0.5}=8

is arithmetically correct but not a matched time-to-solution speedup.

The review should request either the time the classical method needs to reach quality 120 under the same stopping rule, or the time and success probability for the quantum method to reach and certify quality 128. Until then, the speedup claim is unresolved or contradicted as written, while the two local runtime observations can remain supported.

A protocol attempts 50,000 runs, accepts 2,000, and obtains 1,900 good outputs among accepted runs. Compute paccp_{\rm acc}, pgood∣accp_{\rm good|acc}, and pdelp_{\rm del}. Which percentage belongs in an unconditional delivery claim?

Solution

The acceptance probability is

pacc=200050000=0.04.p_{\rm acc} = \frac{2000}{50000} = 0.04.

Conditional quality is

pgood∣acc=19002000=0.95.p_{\rm good|acc} = \frac{1900}{2000} = 0.95.

The good-output probability per attempt is

pdel=190050000=0.038.p_{\rm del} = \frac{1900}{50000} = 0.038.

The 95%95\% figure describes accepted records. The unconditional delivery claim must use 3.8%3.8\%, together with the attempt rate or cost if throughput is at issue.

An experiment collects 10610^6 shots from each of 20 circuits, all during one calibration window. It claims the result generalizes across devices and months. Identify the units supported by the data and the missing evidence.

Solution

Shots estimate within-circuit outcome probabilities during one operating window. The 20 circuits provide some variation across that chosen circuit set, although their sampling procedure must still be known. There is only one device instance and one calibration epoch.

The data do not estimate device-to-device or month-to-month variation. Independent devices, fabrication lots, or acquisition epochs are needed for those populations. More shots in the same window reduce shot uncertainty but do not supply the missing higher-level replication.

A study projects a useful algorithm using 10610^6 physical qubits and ten days of runtime. The estimate assumes a physical error rate ten times lower than the authors’ cited device result and omits classical decoding power. Is the claim false?

Solution

The projection is not false merely because it uses future parameters. It is a resource estimate under assumptions, not a demonstrated capability.

The review should verify the arithmetic, label the lower physical error rate as an unachieved dependency, request sensitivity to plausible error rates, and add decoder throughput, latency, and power to the resource boundary. A claim that “the algorithm can currently run with 10610^6 qubits in ten days” would be contradicted by the stated status; a claim that “the scenario requires these resources if the assumptions hold” may be supported.

A distance-5 memory has a lower logical error per round than distance 3 for four rounds. A logical entangling primitive is demonstrated in a separate circuit. The source says, “We achieved universal fault-tolerant quantum computing.” Write the disposition.

Solution

The universal-computing claim is contradicted as written or supported only after substantial narrowing. The evidence supports a matched finite-distance memory-scaling result for the four-round protocol and a separate logical-operation building block, assuming their detailed fault conditions pass review.

It does not establish composition into a universal gate set at a sustained logical error rate, non-Clifford resource production, decoder and frame throughput, arbitrarily deep execution, or a useful algorithm. The smallest next test is not simply another component benchmark; it is an integrated logical workload that composes the required primitives under one error and resource contract.

  1. National Academies of Sciences, Engineering, and Medicine, Quantum Computing: Progress and Prospects (National Academies Press, 2019), doi:10.17226/25196.
  2. National Academies of Sciences, Engineering, and Medicine, Reproducibility and Replicability in Science (National Academies Press, 2019), doi:10.17226/25303.
  3. Joint Committee for Guides in Metrology, Evaluation of Measurement Data: Guide to the Expression of Uncertainty in Measurement, JCGM 100:2008 (2008), BIPM publication.
  4. M. D. Wilkinson et al., “The FAIR Guiding Principles for scientific data management and stewardship,” Scientific Data 3, 160018 (2016), doi:10.1038/sdata.2016.18.
  5. T. Lebo, S. Sahoo, and D. McGuinness, editors, “PROV-O: The PROV Ontology,” W3C Recommendation (2013), W3C PROV-O.
  6. Association for Computing Machinery, “Artifact Review and Badging, Version 1.1” (2020), ACM policy.
  7. R. D. Peng, “Reproducible research in computational science,” Science 334, 1226–1227 (2011), doi:10.1126/science.1213847.
  8. G. K. Sandve et al., “Ten simple rules for reproducible computational research,” PLoS Computational Biology 9, e1003285 (2013), doi:10.1371/journal.pcbi.1003285.
  9. M. R. Munafò et al., “A manifesto for reproducible science,” Nature Human Behaviour 1, 0021 (2017), doi:10.1038/s41562-016-0021.
  10. R. L. Wasserstein and N. A. Lazar, “The ASA statement on p-values: Context, process, and purpose,” The American Statistician 70, 129–133 (2016), doi:10.1080/00031305.2016.1154108.
  11. J. Eisert et al., “Quantum certification and benchmarking,” Nature Reviews Physics 2, 382–390 (2020), doi:10.1038/s42254-020-0186-4.
  12. M. Kliesch and I. Roth, “Theory of quantum system certification,” PRX Quantum 2, 010201 (2021), doi:10.1103/PRXQuantum.2.010201.
  13. T. Proctor, K. Rudinger, K. Young, E. Nielsen, and R. Blume-Kohout, “Measuring the capabilities of quantum computers,” Nature Physics 18, 75–79 (2022), doi:10.1038/s41567-021-01409-7.
  14. T. Lubinski et al., “Application-oriented performance benchmarks for quantum computing,” IEEE Transactions on Quantum Engineering 4, 1–32 (2023), doi:10.1109/TQE.2023.3253761.
  15. D. Mills et al., “Application-motivated, holistic benchmarking of a full quantum computing stack,” Quantum 5, 415 (2021), doi:10.22331/q-2021-03-22-415.
  16. T. Proctor et al., “Benchmarking quantum computers,” Nature Reviews Physics 7, 105–118 (2025), doi:10.1038/s42254-024-00796-z.