Claims and Evidence Checklist
Short Definition
Section titled “Short Definition”The claims and evidence checklist is a practical decision procedure for turning a headline into an assessable statement, matching that statement to the evidence actually supplied, and recording the narrowest conclusion that survives review.
It is designed for reading papers, technical reports, benchmark releases, resource estimates, vendor announcements, and proposals. Its output is not a numerical credibility score. It is a disposition:
| Disposition | Meaning |
|---|---|
| supported as written | the evidence directly supports the stated claim within its declared scope |
| supported after narrowing | a more limited statement is supported once an overbroad word, population, baseline, or inference is removed |
| unresolved | the claim may be true, but required evidence or disclosure is missing |
| contradicted | the supplied evidence is materially inconsistent with the claim |
| not assessable | the claim is too undefined, inaccessible, or unfalsifiable to evaluate |
Claims, Hype, and Evidence Standards owns the conceptual claim contract, evidence labels, and reasons the fields matter. Reporting Standards owns the report, artifact, provenance, and reproducibility contract. The present page owns the review worksheet: which questions to ask, where to stop, how to test comparator compatibility, and how to write a bounded decision record.
How to Use the Checklist
Section titled “How to Use the Checklist”Work claim by claim
Section titled “Work claim by claim”Do not audit a paper, company, platform, or research program as one indivisible object. Extract each consequential sentence separately. One source may contain:
- a proved asymptotic theorem;
- a numerical result for selected instances;
- a laboratory demonstration;
- a benchmark against one implementation;
- a future hardware projection;
- an application or commercial interpretation.
Those claims can receive different dispositions. A rigorous theorem does not automatically validate the experimental implementation, and a strong experiment does not automatically validate a projection built on additional assumptions.
Record missing information explicitly
Section titled “Record missing information explicitly”Use one of four field states:
| Field state | Use |
|---|---|
| reported | the value, definition, and source are supplied |
| not applicable | the field genuinely does not apply, with a reason |
| unavailable | the reviewer cannot obtain the information |
| withheld | the source intentionally restricts it and states why |
Blank is not a field state. It can mean zero, unknown, forgotten, proprietary, or not applicable. Preserve the distinction because it determines whether the claim is false, unresolved, or simply outside the available evidence.
Do not average away a fatal mismatch
Section titled “Do not average away a fatal mismatch”A points system is tempting but usually misleading. Ten well-reported calibration fields cannot repair a classical baseline that solves a different task. Open code cannot repair an output that was never verified. A large sample count cannot repair pseudoreplication when all shots came from one device state.
Use gates instead:
- Is there an exact, falsifiable claim?
- Are candidate and baseline solving the same task to the same quality?
- Does the evidence identify the relevant population and uncertainty?
- Can the output or inferred process be checked?
- Does the public sentence stay inside the tested scope?
If a gate fails, stop the stronger conclusion and record what narrower claim, if any, still passes.
The Claim Card
Section titled “The Claim Card”Complete one card for every headline or decision-relevant claim.
| Field | Entry to record |
|---|---|
| claim ID | stable label connecting the review, source text, figures, and artifacts |
| exact source sentence | verbatim claim, location, version, and date |
| proposed evidence label | theorem, simulation, experiment, benchmark, resource estimate, projection, commercial claim, or speculation |
| task | exact input-output relation |
| input population | instances, devices, circuits, states, seeds, time window, or promise distribution |
| output and quality | required answer, tolerance, confidence, fidelity, approximation, or utility |
| comparator | algorithm, device, protocol, theory value, or no-comparison reference |
| resources | time, samples, qubits or modes, energy, memory, control, classical work, retries, and infrastructure |
| evidence | data, proof, simulation, calibration, benchmark record, or cited result |
| verification | how correctness or model adequacy is checked |
| uncertainty | estimand, estimator, independent unit, interval, and systematic contributions |
| exclusions | postselection, failed jobs, omitted devices, data cleaning, and stopping rules |
| scope | conditions, populations, dates, and versions to which the result applies |
| nearest stronger claim | plausible conclusion that this evidence does not yet establish |
| disposition | supported, narrowed, unresolved, contradicted, or not assessable |
| next evidence | smallest additional test or disclosure that could change the disposition |
The “nearest stronger claim” field is especially useful. It prevents a reader from silently upgrading “component demonstration” to “system capability,” “finite-instance benchmark” to “asymptotic advantage,” or “below-threshold memory” to “fault-tolerant computation.”
Pass 1: Make the Claim Testable
Section titled “Pass 1: Make the Claim Testable”Identify the grammatical subject
Section titled “Identify the grammatical subject”Ask what object actually has the reported property:
- a physical gate or an inferred gate model;
- one calibrated qubit pair or the processor;
- one circuit family or an algorithm;
- a simulated Hamiltonian or a material;
- an encoded memory or a logical computer;
- a conditional accepted data set or all attempted runs.
“The system achieved fidelity” is not yet testable. Fidelity is a relation between specified objects under a specified convention. The sentence must name the state, channel, gate set, estimator, and treatment of state preparation and measurement.
Replace field names with tasks
Section titled “Replace field names with tasks”“Quantum chemistry,” “optimization,” “machine learning,” and “networking” are domains. A task has inputs, outputs, and a success rule. Examples include:
- estimate a ground-state energy within for a declared molecular Hamiltonian;
- return a feasible cut whose objective is at least a declared fraction of a reference value;
- sample from a target distribution within a named distance;
- distribute one Bell pair above a fidelity threshold between named endpoint memories;
- estimate a magnetic field with specified bandwidth, spatial resolution, and mean-squared error.
If the source does not permit this conversion, mark the claim not assessable, not false.
Separate observations from interpretation
Section titled “Separate observations from interpretation”Write two sentences:
- Observation: what quantity was measured or computed?
- Interpretation: what physical, computational, or practical conclusion is inferred?
For example, a randomized benchmark may observe a fitted sequence-decay parameter. Interpreting it as average gate infidelity requires a protocol model. Interpreting that gate infidelity as application performance requires another bridge. Each bridge carries assumptions and should be audited separately.
Name the time boundary
Section titled “Name the time boundary”Hardware calibration, cloud service terms, software stacks, and best-known classical methods change. Record:
- acquisition dates and calibration epoch;
- software and compiler versions;
- source revision, correction, or retraction status;
- date on which the comparator search was performed;
- date after which the claim should be rechecked.
A dated benchmark can remain correct historical evidence even after it stops describing current capability.
Pass 2: Test Comparator Compatibility
Section titled “Pass 2: Test Comparator Compatibility”Build the compatibility table
Section titled “Build the compatibility table”For every comparative claim, classify each row as matched, reconciled, or incompatible.
| Comparison field | Questions |
|---|---|
| task | are both methods solving the same input-output problem? |
| instance population | are sizes, distributions, promises, and excluded cases aligned? |
| output quality | are accuracy, fidelity, approximation, confidence, and failure probability matched? |
| access model | do both receive the same data, oracle, state, Hamiltonian, or side information? |
| tuning budget | are hyperparameter search, compilation, calibration, and manual intervention counted? |
| stopping rule | do both stop on the same quality or resource condition? |
| time boundary | are setup, queueing, execution, decoding, verification, and postprocessing treated consistently? |
| hardware boundary | are processors, accelerators, networking, control, and classical hosts included consistently? |
| repetitions | are failed attempts, shots, seeds, restarts, and confidence targets aligned? |
| energy and infrastructure | are cooling, control, communication, and amortization conventions disclosed? |
| date and implementation quality | is the comparator current, competently implemented, and versioned? |
Matched means the same convention is used. Reconciled means a documented conversion or sensitivity analysis places the quantities on a common basis. Incompatible means the headline ratio is not interpretable as the stated comparison.
Keep conditional and delivered performance separate
Section titled “Keep conditional and delivered performance separate”Suppose attempts produce accepted records, of which satisfy the task criterion. Define
The probability of delivering a good result per attempt is
Conditional quality can improve while delivered quality declines. This applies to error-detected circuits, heralded links, filtered sensing records, optimizer restarts, and any protocol with abstention.
For a wall-clock attempt rate , the delivered good-result rate is
A comparison may legitimately optimize conditional fidelity, delivered rate, or a utility combining both. It must say which.
Normalize cost to the useful output
Section titled “Normalize cost to the useful output”Let be one-time setup cost, the cost per attempt, additional work per accepted record, and the number of useful delivered outputs. A transparent average cost is
The cost can represent time, energy, money, samples, or a resource vector. A fixed cost may be amortized only over the declared workload. Compilation for one instance cannot be divided by a hypothetical million future instances while the baseline is charged setup per current run.
Audit “best known”
Section titled “Audit “best known””For a claimed speedup or superiority:
- identify the comparator search procedure and cutoff date;
- include specialized as well as generic methods;
- report implementation language, libraries, precision, and hardware;
- give both methods a defensible tuning budget;
- inspect scaling across instances rather than one favorable point;
- include verification and output-conversion costs;
- disclose whether the candidate influenced which baseline or instances were selected.
“Best known” is a review conclusion supported by a dated search, not a permanent property of an algorithm.
Pass 3: Stress-Test the Inference
Section titled “Pass 3: Stress-Test the Inference”Find the independent experimental unit
Section titled “Find the independent experimental unit”Shots are often conditional repeats within one higher-level unit. The independent unit may instead be:
- a randomly drawn circuit;
- a problem instance;
- a device or fabrication lot;
- a calibration epoch;
- a day of acquisition;
- an optimizer seed;
- a biological or material sample;
- a network path or weather interval.
One million shots from one circuit on one calibration do not establish variation across circuits or days. The uncertainty calculation should match the population named in the claim.
Name the estimand and estimator
Section titled “Name the estimand and estimator”The estimand is the quantity the claim intends to describe. The estimator is the rule applied to observed records. Record:
- the target population;
- weighting across instances or devices;
- whether the statistic is mean, median, quantile, worst case, or selected best;
- the fit model and parameter convention;
- interval type and coverage or credibility level;
- known systematic effects and model discrepancy.
A narrow bootstrap interval around a biased estimator does not include model error merely because it has many resamples.
Audit selection
Section titled “Audit selection”List the choices explored before the reported result was selected:
- qubit subsets and couplers;
- circuit depths and benchmark families;
- data-cleaning thresholds;
- optimizer seeds and ansatzes;
- mitigation parameters;
- time windows;
- metrics and fit ranges;
- problem instances and classical baselines.
Then ask whether the uncertainty or validation set accounts for that search. Reporting the winning configuration as though it had been prespecified can produce severe winner’s-curse bias.
Look for a holdout or negative control
Section titled “Look for a holdout or negative control”Useful stress tests include:
- held-out circuits, instances, observables, or acquisition periods;
- injected faults or synthetic data with known truth;
- classically tractable limits;
- conservation laws and sum rules;
- randomized labels or null inputs;
- independent analysis pipelines;
- hardware-disabled or correction-disabled controls;
- blinded parameter choices.
A method tuned and evaluated on the same small set demonstrates fit to that set. Generalization requires new evidence.
Verify at the level claimed
Section titled “Verify at the level claimed”Verification must reach the claim’s output:
| Claim | Minimum relevant check |
|---|---|
| state preparation | informationally relevant observables, witnesses, or validated tomography |
| gate or channel | protocol-appropriate process evidence with model and SPAM limitations |
| algorithmic answer | objective, residual, certificate, known solution, or independent cross-check |
| sampling distribution | finite-sample test tied to a declared distance or score |
| simulator | convergence, discretization, limiting cases, and independent method where possible |
| resource estimate | arithmetic reproduction plus sensitivity to architecture assumptions |
| sensor | calibrated response, uncertainty budget, bandwidth, and matched classical limit |
| secure protocol | proof assumptions plus implementation parameters and finite-size analysis |
| logical memory | logical channel, matched reference, decoder, acceptance, and elapsed time |
A proxy metric may be useful, but it supports a proxy claim until its connection to the target is validated.
Test the transport step
Section titled “Test the transport step”Many overclaims occur after the measurement is complete. Ask whether the conclusion moves:
- from selected instances to an instance family;
- from one device to a platform;
- from a component to a system;
- from short duration to sustained operation;
- from simulated noise to physical noise;
- from average behavior to worst-case guarantees;
- from a benchmark to an application;
- from a present result to a deployment forecast.
Each move needs an argument, data, or model. If none is supplied, retain the local observation and narrow the interpretation.
Source and Provenance Check
Section titled “Source and Provenance Check”Match source type to the proposition
Section titled “Match source type to the proposition”| Source | What it can directly establish |
|---|---|
| proof or formal paper | theorem under its definitions and assumptions |
| original experimental article and data | reported apparatus, protocol, observations, and analysis |
| benchmark release | score under that benchmark version and rules |
| resource-estimate study | projection under its algorithm and architecture scenario |
| review article | synthesis and context up to its search date |
| press release or company page | that the organization made the stated claim |
| news report | a secondary account and route to primary sources |
| roadmap | intended milestones, dependencies, and organizational expectations |
Peer review raises neither every claim to fact nor every preprint to unreliable. Record status, correction history, methods access, and independent support. A press release may be the primary source for a product announcement but remains secondary evidence for the underlying scientific performance.
Trace the dependency chain
Section titled “Trace the dependency chain”For every number in the final sentence, ask:
- Where was it first measured, proved, or computed?
- Is the cited source the original source?
- Did a later source change the definition or denominator?
- Is the quotation faithful to the source’s caveats?
- Has a correction, retraction, superseding version, or improved comparator appeared?
Circular citation can make one estimate look independently confirmed when several reports all trace to the same unpublished assumption.
Record interests without using them as a verdict
Section titled “Record interests without using them as a verdict”Funding, employment, equity, benchmark authorship, and access restrictions can shape incentives and available evidence. Record them. Do not replace methodological review with an inference about motive. A conflicted source can provide excellent evidence; an apparently disinterested source can provide poor evidence. The practical response is stronger transparency and independent checking.
Domain Modules
Section titled “Domain Modules”Use the common gates first, then add the relevant module.
Algorithms and advantage
Section titled “Algorithms and advantage”- Is the advantage asymptotic, finite-instance, or measured wall-clock?
- What input-access model is assumed?
- Are data loading, state preparation, repeated shots, decoding, and output extraction included?
- Does the classical baseline solve the same task to the same quality?
- Is the hardness statement worst-case, average-case, conjectural, or tied to the sampled instances?
- Can the output be verified at the claimed scale?
Stop an advantage claim if the baseline is stale, quality is unmatched, or the quantum output is not checked. A circuit can remain a valuable experimental demonstration after that stronger label is removed.
Hardware and benchmarks
Section titled “Hardware and benchmarks”- Which qubits, gates, topology, concurrency pattern, and calibration epoch?
- Is the metric directly observed or inferred through a model?
- Are leakage, crosstalk, drift, and SPAM inside the protocol’s scope?
- Does the benchmark sample the operations needed by the target workload?
- Are uncertainty and device-selection procedures reported?
- Does a component metric predict a system-level outcome on held-out circuits?
Stop a platform-wide claim when only a selected subset was measured.
Simulation
Section titled “Simulation”- Is the target a mathematical model, a physical system, or both?
- Which discretization, truncation, boundary condition, and solver tolerance?
- Are convergence and finite-size studies shown?
- Is sampling or optimization error separated from model discrepancy?
- Are sign problems, entanglement growth, or conditioning restricting the tested regime?
- Is the physical interpretation validated against independent observables?
Agreement with another implementation of the same approximation checks code more directly than it checks nature.
Resource estimates
Section titled “Resource estimates”- Which algorithm, compiler, code, physical error model, decoder, cycle time, and logical failure budget?
- Are factories, routing, storage, interconnects, retries, and classical control included?
- Which quantities are measured, assumed, extrapolated, or targeted?
- How does the answer change under plausible parameter ranges?
- Is the estimate optimized for qubits, runtime, space-time volume, energy, or another objective?
- Are future hardware dependencies labeled as dependencies?
Use Resource Estimation Tools for the full scenario and sensitivity contract.
Error correction
Section titled “Error correction”- Is the result detection, postselection, correction, break-even, below-threshold scaling, or a logical operation?
- Which logical states and Pauli sectors were tested?
- What is the code, syndrome circuit, decoder, and frame convention?
- Is the metric per cycle, per elapsed time, or both?
- What physical or uncorrected reference defines break-even?
- Are rare correlated events, leakage, decoder latency, and rejected runs included?
Error-Correction Case Studies provides the matched logical-channel comparison.
Sensing and metrology
Section titled “Sensing and metrology”- What parameter, operating range, bandwidth, spatial resolution, and loss function?
- Which resources define the standard quantum limit?
- Are preparation time, dead time, readout, calibration, and estimator bias included?
- Is sensitivity distinguished from precision, accuracy, and detection limit?
- Are technical noise and drift included in the same bandwidth?
- Does the quantum enhancement survive loss and the complete measurement cycle?
A squeezed input is a resource demonstration; it becomes a sensor advantage only through a matched end-to-end uncertainty comparison.
Communication, cryptography, and networks
Section titled “Communication, cryptography, and networks”- What object is delivered: clicks, raw bits, entanglement, a teleported state, secret key, or a remote gate?
- Are rates conditional, sifted, secret, or end-to-end delivered?
- What trust, adversary, authentication, and composability model applies?
- Are finite-size effects, loss, side channels, memory age, and classical latency included?
- Is the comparison direct transmission, another network architecture, or a classical service?
- Are security proof assumptions connected to measured device parameters?
An optical event and a delivered secure or logical service are different evidence levels.
Writing the Disposition
Section titled “Writing the Disposition”Use a bounded result sentence
Section titled “Use a bounded result sentence”A complete disposition can use this form:
For [task and population], under [versioned conditions], [method] estimated [quantity] as [value and uncertainty] relative to [baseline], using [resource boundary] and [verification]. This supports [evidence label] over [scope]. It does not yet establish [nearest stronger claim] because [specific missing bridge].
The template is not a demand for one very long public sentence. It is a test: every clause should be recoverable from the report.
Distinguish narrowing from contradiction
Section titled “Distinguish narrowing from contradiction”Suppose a source says “the quantum optimizer is 100 times faster,” but the data show a 100-fold reduction in device execution time for selected instances, excluding compilation, queueing, repeated shots, and classical postprocessing. The evidence does not necessarily contradict the measured device-time ratio. It contradicts or leaves unresolved the end-to-end optimizer claim. A defensible narrowing is:
On the selected instances, the reported quantum device-execution segment was 100 times shorter than the declared baseline segment; no end-to-end time-to-solution comparison was established.
Preserving the valid local result is more informative than replacing every overstatement with “false.”
State the smallest decisive gap
Section titled “State the smallest decisive gap”Do not bury the review in a catalog of every desirable experiment. Name the smallest missing item that blocks the claim:
- matched output quality;
- a current competitive baseline;
- an independent test set;
- unconditional acceptance and retry accounting;
- a logical rather than component metric;
- uncertainty over devices rather than shots;
- verification beyond the classically tractable regime;
- sensitivity to one dominant architecture assumption.
This turns the checklist into a research plan.
Three Review Depths
Section titled “Three Review Depths”Five-minute screen
Section titled “Five-minute screen”- Copy the exact claim.
- Assign a provisional evidence label.
- Identify task, output quality, comparator, and date.
- Find the primary source.
- Check whether the nearest stronger claim is being implied.
Use not assessable or unresolved when these fields are absent. Do not fill them with charitable guesses.
Thirty-minute audit
Section titled “Thirty-minute audit”- Complete the claim card.
- Build the comparator compatibility table.
- Identify the independent unit and uncertainty semantics.
- Trace exclusions, retries, and selected configurations.
- Locate verification and robustness tests.
- Write the bounded disposition.
Full technical review
Section titled “Full technical review”- Recompute key values from released records.
- Inspect executable versions, calibration context, and provenance.
- Reproduce the analysis where possible.
- Test alternative fit, stopping, and exclusion choices.
- Update the classical or physical baseline.
- Seek independent replication or a genuinely different validation method.
- Record which unavailable artifacts limit the conclusion.
A full review can remain inconclusive. Honest unresolved status is an outcome, not a failed review.
Compact Checklist
Section titled “Compact Checklist”Claim identity
Section titled “Claim identity”- The exact sentence, source location, version, and date are recorded.
- The grammatical subject is a specific object, protocol, or population.
- The task, input, output, and success criterion are explicit.
- Observation, model-based inference, and projection are separate.
- The provisional evidence label matches the source type.
Comparison
Section titled “Comparison”- Candidate and baseline solve the same task.
- Output quality and failure probability are aligned.
- Access models and side information are aligned.
- Setup, tuning, execution, retries, and postprocessing use one boundary.
- Hardware, parallelism, precision, and versions are reported.
- The comparator search is competent and dated.
Evidence and verification
Section titled “Evidence and verification”- The number traces to a primary source or its dependency is explicit.
- The measured or computed output is checked at the level claimed.
- Proxy metrics are not silently promoted to task success.
- Holdout tests, negative controls, or independent methods are present where generalization is claimed.
- Corrections, retractions, and superseding results were checked.
Statistics and selection
Section titled “Statistics and selection”- The estimand, estimator, and independent unit are named.
- Interval type, level, and uncertainty sources are reported.
- Device, instance, seed, and analysis selection are disclosed.
- Rejected runs, failed jobs, and stopping rules are counted.
- Model discrepancy and systematic effects are not hidden inside shot error.
Resources and delivery
Section titled “Resources and delivery”- Resources are a vector rather than one convenient count.
- Conditional quality and acceptance probability are both visible.
- Cost and rate are normalized to a useful delivered output.
- Amortization uses the declared workload.
- Verification and classical control costs remain inside the boundary.
Scope and communication
Section titled “Scope and communication”- The tested population, operating regime, and time boundary are explicit.
- Component, system, benchmark, application, and projection levels remain distinct.
- The nearest stronger unestablished claim is stated.
- Conflicts and access restrictions are recorded without replacing methodological review.
- The final disposition names one decisive reason and the next useful evidence.
Common Mistakes
Section titled “Common Mistakes”- Scoring the source or institution instead of auditing a sentence.
- Treating peer review as a substitute for checking the inference.
- Calling missing evidence proof that the claim is false.
- Reconstructing unstated assumptions in the most favorable way.
- Comparing candidate and baseline at different output quality.
- Counting shots as independent evidence about devices, days, or instances.
- Accepting a proxy because it correlates with the target in a different regime.
- Reporting only conditional quality after postselection.
- Treating an error bar as complete without interval semantics.
- Ignoring the search over qubits, seeds, metrics, and fit windows.
- Calling a resource projection a demonstrated capability.
- Treating a commercial roadmap date as a scientific confidence interval.
- Letting a disclosure checklist obscure one fatal task or comparator mismatch.
- Writing “more evidence is needed” without naming the smallest decisive test.
Exercises
Section titled “Exercises”1. Repair a gate-fidelity claim
Section titled “1. Repair a gate-fidelity claim”A report states, “The processor has fidelity,” and cites an interleaved randomized benchmark on one selected two-qubit gate pair. Give the provisional disposition and a bounded replacement sentence.
Solution
The processor-wide claim is supported after narrowing at best. The source object is one selected gate pair, the protocol estimates a benchmark-specific average error parameter, and selection, SPAM, leakage, crosstalk, parallel operation, and other processor gates may be outside scope.
A bounded sentence is:
During the reported calibration epoch, interleaved randomized benchmarking on the selected two-qubit pair produced the reported protocol-derived fidelity estimate. This is evidence about that gate pair under the benchmark model, not a processor-wide fidelity or application success probability.
2. Audit a mismatched optimizer race
Section titled “2. Audit a mismatched optimizer race”A quantum method returns feasible solutions with mean objective 120 in . A classical method certifies the optimum, mean objective 128, in . May the source claim an eightfold speedup?
Solution
Not from these data. The outputs have different quality and possibly different tasks: one returns a feasible point of quality 120, while the other certifies the optimum of quality 128. The runtime ratio
is arithmetically correct but not a matched time-to-solution speedup.
The review should request either the time the classical method needs to reach quality 120 under the same stopping rule, or the time and success probability for the quantum method to reach and certify quality 128. Until then, the speedup claim is unresolved or contradicted as written, while the two local runtime observations can remain supported.
3. Include postselection
Section titled “3. Include postselection”A protocol attempts 50,000 runs, accepts 2,000, and obtains 1,900 good outputs among accepted runs. Compute , , and . Which percentage belongs in an unconditional delivery claim?
Solution
The acceptance probability is
Conditional quality is
The good-output probability per attempt is
The figure describes accepted records. The unconditional delivery claim must use , together with the attempt rate or cost if throughput is at issue.
4. Find the independent unit
Section titled “4. Find the independent unit”An experiment collects shots from each of 20 circuits, all during one calibration window. It claims the result generalizes across devices and months. Identify the units supported by the data and the missing evidence.
Solution
Shots estimate within-circuit outcome probabilities during one operating window. The 20 circuits provide some variation across that chosen circuit set, although their sampling procedure must still be known. There is only one device instance and one calibration epoch.
The data do not estimate device-to-device or month-to-month variation. Independent devices, fabrication lots, or acquisition epochs are needed for those populations. More shots in the same window reduce shot uncertainty but do not supply the missing higher-level replication.
5. Classify a resource estimate
Section titled “5. Classify a resource estimate”A study projects a useful algorithm using physical qubits and ten days of runtime. The estimate assumes a physical error rate ten times lower than the authors’ cited device result and omits classical decoding power. Is the claim false?
Solution
The projection is not false merely because it uses future parameters. It is a resource estimate under assumptions, not a demonstrated capability.
The review should verify the arithmetic, label the lower physical error rate as an unachieved dependency, request sensitivity to plausible error rates, and add decoder throughput, latency, and power to the resource boundary. A claim that “the algorithm can currently run with qubits in ten days” would be contradicted by the stated status; a claim that “the scenario requires these resources if the assumptions hold” may be supported.
6. Narrow a QEC headline
Section titled “6. Narrow a QEC headline”A distance-5 memory has a lower logical error per round than distance 3 for four rounds. A logical entangling primitive is demonstrated in a separate circuit. The source says, “We achieved universal fault-tolerant quantum computing.” Write the disposition.
Solution
The universal-computing claim is contradicted as written or supported only after substantial narrowing. The evidence supports a matched finite-distance memory-scaling result for the four-round protocol and a separate logical-operation building block, assuming their detailed fault conditions pass review.
It does not establish composition into a universal gate set at a sustained logical error rate, non-Clifford resource production, decoder and frame throughput, arbitrarily deep execution, or a useful algorithm. The smallest next test is not simply another component benchmark; it is an integrated logical workload that composes the required primitives under one error and resource contract.
References
Section titled “References”- National Academies of Sciences, Engineering, and Medicine, Quantum Computing: Progress and Prospects (National Academies Press, 2019), doi:10.17226/25196.
- National Academies of Sciences, Engineering, and Medicine, Reproducibility and Replicability in Science (National Academies Press, 2019), doi:10.17226/25303.
- Joint Committee for Guides in Metrology, Evaluation of Measurement Data: Guide to the Expression of Uncertainty in Measurement, JCGM 100:2008 (2008), BIPM publication.
- M. D. Wilkinson et al., “The FAIR Guiding Principles for scientific data management and stewardship,” Scientific Data 3, 160018 (2016), doi:10.1038/sdata.2016.18.
- T. Lebo, S. Sahoo, and D. McGuinness, editors, “PROV-O: The PROV Ontology,” W3C Recommendation (2013), W3C PROV-O.
- Association for Computing Machinery, “Artifact Review and Badging, Version 1.1” (2020), ACM policy.
- R. D. Peng, “Reproducible research in computational science,” Science 334, 1226–1227 (2011), doi:10.1126/science.1213847.
- G. K. Sandve et al., “Ten simple rules for reproducible computational research,” PLoS Computational Biology 9, e1003285 (2013), doi:10.1371/journal.pcbi.1003285.
- M. R. Munafò et al., “A manifesto for reproducible science,” Nature Human Behaviour 1, 0021 (2017), doi:10.1038/s41562-016-0021.
- R. L. Wasserstein and N. A. Lazar, “The ASA statement on p-values: Context, process, and purpose,” The American Statistician 70, 129–133 (2016), doi:10.1080/00031305.2016.1154108.
- J. Eisert et al., “Quantum certification and benchmarking,” Nature Reviews Physics 2, 382–390 (2020), doi:10.1038/s42254-020-0186-4.
- M. Kliesch and I. Roth, “Theory of quantum system certification,” PRX Quantum 2, 010201 (2021), doi:10.1103/PRXQuantum.2.010201.
- T. Proctor, K. Rudinger, K. Young, E. Nielsen, and R. Blume-Kohout, “Measuring the capabilities of quantum computers,” Nature Physics 18, 75–79 (2022), doi:10.1038/s41567-021-01409-7.
- T. Lubinski et al., “Application-oriented performance benchmarks for quantum computing,” IEEE Transactions on Quantum Engineering 4, 1–32 (2023), doi:10.1109/TQE.2023.3253761.
- D. Mills et al., “Application-motivated, holistic benchmarking of a full quantum computing stack,” Quantum 5, 415 (2021), doi:10.22331/q-2021-03-22-415.
- T. Proctor et al., “Benchmarking quantum computers,” Nature Reviews Physics 7, 105–118 (2025), doi:10.1038/s42254-024-00796-z.
Further Connections
Section titled “Further Connections”- Why Benchmarking Is Hard explains why a score is a conditional projection of a device, workload, implementation policy, and analysis.
- Algorithmic Benchmarking owns end-to-end task and comparator design.
- Verification of Quantum Advantage develops the stronger evidence chain needed for computational-advantage claims.
- Negative Results and Limitations turns failed, narrowed, or reversed claims into bounded records with explicit assumptions and update triggers.
- Metrics for Quantum Hardware supplies operational definitions for component, cycle, logical, and workload metrics.
- Evidence Labels supplies the site-wide epistemic vocabulary used in the final disposition.