Skip to content

Reporting Standards

Reporting standards specify the information and artifacts that must accompany a quantum-computing result so that another reader can interpret, audit, reanalyze, compare, and, where access permits, reproduce it. A complete report connects the headline claim to the protocol, physical system, executable, records, analysis, uncertainty, and resource accounting:

R=(C,P,H,X,D,A,U,Q).\mathcal R = \left( C,P,H,X,D,A,U,Q \right).

Here:

  • CC is the precise claim;
  • PP is the experimental or computational protocol;
  • HH is the hardware and operating context;
  • XX is the exact executable and its compilation history;
  • DD is the acquired or simulated record;
  • AA is the analysis procedure;
  • UU is the uncertainty and model-limit statement;
  • QQ is the reported quantity, decision, or bound.

The tuple is not a suggestion that one universal file format suits every experiment. It identifies the evidence links that must not be broken. A paper containing a fidelity number but no target operation, acquisition date, shot hierarchy, compiler settings, or uncertainty does not define a result that can be compared.

Transparent reporting is not the same as publishing everything without restriction. Security, privacy, proprietary access, export control, and hardware protection can limit disclosure. A transparent report states what is withheld, why, how qualified reviewers may request access, and which claims cannot be independently checked as a consequence.

This page owns the report contract:

  • a common minimum record for hardware, circuits, execution, analysis, uncertainty, resources, code, and data;
  • additional profiles for characterization, tomography, certification, benchmarking, hybrid algorithms, logical experiments, simulations, resource estimates, and advantage claims;
  • machine-readable manifests, artifact identity, provenance, missing-data semantics, corrections, and versioning;
  • rules for postselection, error mitigation, optimizer settings, comparison, and restricted artifacts;
  • checklists and worked audits for authors, reviewers, and benchmark maintainers.

Why Benchmarking Is Hard owns the conceptual benchmark contract and explains why capability is a conditional surface. Device Characterization owns predictive model identification and validation. Algorithmic Benchmarking owns end-to-end task design, and Verification of Quantum Advantage owns the evidence chain for an advantage claim.

Claims, Hype, and Evidence Standards owns public claim classification and claim audits. Reproducible Notebooks owns executable notebook admission and validation. This page answers the cross-cutting question: what must travel with a result, regardless of which valid protocol produced it? Network Verification owns network-specific trial eligibility, destructive sample assignment, source-independence assumptions, resource identity, causal transcripts, and service acceptance; the present page owns how that evidence is packaged.

A mature report should be:

  1. identifiable: every device, data set, executable, and software release is versioned or otherwise uniquely located;
  2. interpretable: quantities, populations, conditions, units, and estimands are defined;
  3. auditable: the path from raw records to each displayed result is traceable;
  4. reanalyzable: sufficient data and code, or a justified access route, permit independent analysis;
  5. comparable: the report exposes the conditions that must match before two results can be contrasted.

Availability alone does not establish these properties. A public repository without a tagged version, environment, raw-data map, or run instructions is available but not reproducible. A polished PDF can be interpretable while remaining impossible to audit. A complete reporting package needs both human-readable explanation and machine-actionable metadata.

Each recommended field should carry one of four statuses:

statusmeaning
reporteda value and its definition are supplied
not applicablethe field does not apply, with a reason
unavailablethe value existed but cannot be obtained or released
withheldthe value is intentionally restricted, with policy and access information

An empty cell or missing key has no defined meaning. It can represent oversight, zero, not applicable, unknown, or confidential. Schema validation should reject unexplained absence for required fields.

“Proprietary” is not a complete value. For a withheld compiler or pulse schedule, the report can still provide an interface version, input and output hashes, optimization settings, declared transformations, resource summaries, and an independent escrow or review procedure.

Communities use repeatability, reproducibility, and replicability in conflicting ways. A report should name the convention it adopts. This page uses the National Academies convention:

  • reproducibility: consistent results from the same input data, computational steps, methods, code, and analysis conditions;
  • replicability: consistent results in a new study that acquires its own data to answer the same scientific question.

For an online quantum processor, exact physical repetition may be impossible: the device drifts, calibration changes, and access to an old backend may end. That does not excuse an incomplete record. It makes three targets important:

  1. reproduce the analysis from archived records;
  2. rerun the exact executable on an identified compatible system;
  3. replicate the scientific conclusion with a new device record.

The report should state which target is supported. “Reproducible” without an object and tolerance is too vague.

Every headline result should have a ledger entry. A useful one-sentence form is:

Under conditions SS, protocol PP on object OO estimated quantity qq for population GG as q^\widehat q with uncertainty statement UU, using decision rule RR.

The ledger separates:

  • primary confirmatory claims;
  • secondary prespecified analyses;
  • exploratory analyses chosen after seeing the data;
  • derived projections that depend on additional models.

For each claim, record:

fieldquestion
identifierwhich stable label connects text, table, data, and code?
result classmeasurement, estimate, bound, benchmark, simulation, projection, or decision?
objectwhich qubits, gate, circuit family, task, code, or model?
populationshots, random circuits, instances, devices, time window, or seeds?
estimandwhat mathematical quantity is intended?
estimatorhow is it computed from records?
directionis larger, smaller, or proximity to a target better?
uncertaintywhich interval, level, method, and uncertainty sources?
decisionwhat threshold or comparison supports the claim?
scopewhere should the result not be generalized?

A claim selected because it looked strongest needs selection-aware inference. The ledger should expose searched devices, qubit subsets, circuit families, metrics, time windows, optimizer seeds, and analysis variants. Reporting only the winner hides the population over which selection occurred.

Identify the Hardware and Operating Context

Section titled “Identify the Hardware and Operating Context”

A hardware description should support both interpretation and later matching. At minimum, report:

  • provider and stable backend or instrument identifier;
  • platform and architecture;
  • acquisition dates and time zone;
  • physical qubits, modes, couplers, or links used;
  • topology and active connectivity;
  • native operations, measurement capabilities, reset, feedforward, and dynamic-circuit support;
  • control or service interface version;
  • relevant environmental and operating conditions;
  • access mode, reservation policy, and whether jobs shared a calibration epoch.

A product family name is not a device identity. Cloud backends can change qubit count, couplers, firmware, and calibration while keeping a familiar label.

The report should bind each job or batch to a calibration snapshot or calibration epoch. Include:

  • snapshot identifier and timestamp;
  • age of the snapshot when execution began;
  • relevant frequencies, coherence estimates, assignment metrics, leakage, and gate or cycle metrics;
  • protocols and uncertainties behind those metrics;
  • whether recalibration occurred within the acquisition;
  • invalidation, rollback, or quarantine events.

Do not paste a large provider dashboard and call it context. Select the calibrations relevant to the executed operations, preserve the complete machine-readable snapshot when possible, and state which dashboard quantities were not independently verified.

Declare whether the reported system includes:

  • state preparation and reset;
  • compiler and placement;
  • pulse generation and controller latency;
  • measurement and discrimination;
  • queue and network delay;
  • classical feedback and optimization;
  • error mitigation and postprocessing.

Two “runtime” or “fidelity” numbers with different boundaries are not the same metric.

Preserve the Executable, Not Only the Source

Section titled “Preserve the Executable, Not Only the Source”

A high-level circuit does not uniquely determine what the device executed. The report should preserve a chain such as

Xsource⟶Xmapped⟶Xscheduled⟶Xcontroller.X_{\mathrm{source}} \longrightarrow X_{\mathrm{mapped}} \longrightarrow X_{\mathrm{scheduled}} \longrightarrow X_{\mathrm{controller}}.

For every material stage, record:

  • file or object identifier and cryptographic digest;
  • language or intermediate-representation version;
  • compiler, transpiler, and plugin versions;
  • target profile and calibration snapshot;
  • optimization level and explicit passes;
  • placement, routing, scheduling, and gate-synthesis options;
  • random seeds and tie-breaking rules;
  • error mitigation, dynamical decoupling, twirling, and pulse substitutions;
  • any manual edit after automated compilation.

The reported gate count and depth must name the stage at which they were measured. Define depth with respect to a gate alphabet, scheduling model, and parallelism rule. A logical two-qubit-gate depth is not a pulse-schedule duration.

Report the complete input contract:

  • algorithm and variant;
  • problem instance, encoding, and instance generator;
  • register sizes and ancilla conventions;
  • initial state and state-preparation method;
  • observables, measurement bases, and classical output decoding;
  • parameter values or a parameter file;
  • dynamic branches, mid-circuit measurements, resets, and feedforward;
  • expected ideal output or independently generated reference when available.

For randomized suites, preserve the generator version and realized random instances. A seed is useful only when the generator and its dependencies are stable enough to regenerate the same objects.

“We used 10,000 shots” is not a sampling design. Quantum experiments often have nested random units:

epoch⊃job⊃batch⊃circuit⊃shot.\text{epoch} \supset \text{job} \supset \text{batch} \supset \text{circuit} \supset \text{shot}.

Report:

  • number of calibration epochs;
  • number of dates, devices, or independent runs;
  • jobs and batches per condition;
  • independently sampled circuits, sequences, instances, and optimizer seeds;
  • shots per executable;
  • acquisition order and randomization;
  • stopping rules and adaptive choices;
  • retries, timeouts, cancellations, and failed jobs;
  • duplicate executions and whether they were averaged or selected;
  • timestamps at the finest available useful level.

The independent unit for uncertainty may be the random circuit or time block, not the shot. Increasing shots reduces conditional sampling noise but does not estimate variation across circuits, calibrations, or days.

Job failures, invalid responses, and dropped records are data about the system. Report their counts, causes, and treatment. Retrying until success changes time to solution and can select favorable operating periods. Both attempted and accepted denominators belong in the record.

Every exclusion rule should be stated before the outcome it can affect is examined, or labeled exploratory. Report:

  • eligibility rule and implementation;
  • whether the rule uses settings, outcomes, timestamps, or diagnostic data;
  • total and retained counts in every comparison group;
  • retention uncertainty and dependence on input or context;
  • the result before and after exclusion when meaningful;
  • sensitivity to plausible alternative rules.

If Ki∈{0,1}K_i\in\{0,1\} marks retention and ZiZ_i is an outcome, the conditional mean is

μ^keep=∑iKiZi∑iKi,\widehat\mu_{\mathrm{keep}} = \frac{ \sum_i K_i Z_i }{ \sum_i K_i },

while the retention fraction is

r^=1N∑iKi.\widehat r = \frac1N \sum_i K_i.

Neither number substitutes for the other. A high conditional fidelity with small or state-dependent r^\widehat r may be operationally poor. If postselection is part of a heralded protocol, the herald timing and causal order must be explicit.

Exclusions based on a device-health threshold can also bias comparisons if one platform spends more time below threshold. Report availability and discarded wall-clock time.

Error mitigation changes both the estimator and its cost. Report:

  • method and software implementation;
  • calibration or training data;
  • noise-scaling factors, twirling ensemble, symmetry checks, or quasiprobability coefficients;
  • shot allocation across transformed circuits;
  • fit model and extrapolation order;
  • clipping, projection, regularization, and physicality constraints;
  • all attempted settings, not only the best;
  • variance or sampling-overhead amplification;
  • raw and mitigated results side by side.

Error Mitigation Overview owns why these fields define the estimator license, uncertainty, and quality–cost claim; this page retains the durable reporting, provenance, raw-data, executable, and versioning contract.

Probabilistic Error Cancellation owns the implemented coefficients, signed acquisition, one-norm, squared overhead, model-mismatch audit, and no-cancel license; this page retains durable artifacts, provenance, raw data, executable records, and versioning.

For a linear quasiprobability estimator

q^mit=∑jηjq^j,\widehat q_{\mathrm{mit}} = \sum_j \eta_j \widehat q_j,

the sampling overhead depends on

γ=∑j∣ηj∣.\gamma = \sum_j |\eta_j|.

In common sampling constructions, variance cost can scale as γ2\gamma^2. Reporting only the mitigated point estimate suppresses the resource tradeoff that defines the method.

For zero-noise extrapolation, publish the unextrapolated observations at every scale, requested and effective gains, full covariance, model and range sensitivity, held-out checks, no-fit decisions, and total cost. The specialist owns why those fields license a ZNE claim; this page retains durable provenance, raw-data, versioning, and executable reporting requirements. An extrapolated value outside the physical range must not be silently clipped.

Classical Optimizer and Hybrid-Loop Settings

Section titled “Classical Optimizer and Hybrid-Loop Settings”

Hybrid algorithms are joint quantum–classical procedures. Report:

  • optimizer name, version, and hyperparameters;
  • initialization strategy and every initial point or seed;
  • objective estimator and shot allocation;
  • gradient method and finite-difference step if applicable;
  • parameter constraints and transforms;
  • stopping rule, evaluation budget, and failure rule;
  • number of restarts and selection criterion;
  • batching, caching, asynchronous execution, and reused measurements;
  • complete objective history, not only the best iteration;
  • classical hardware, precision, threads, accelerators, and software stack.

If the best of MM seeds is reported,

q^best=min⁡1≤m≤Mq^m,\widehat q_{\mathrm{best}} = \min_{1\leq m\leq M} \widehat q_m,

then MM, all seed outcomes, and the selection cost are part of the result. The best run is not an estimate of typical performance.

Classical baselines require an equally explicit tuning budget. Comparing a heavily tuned quantum workflow with a default classical implementation is not a matched experiment.

For every reported quantity, give:

  1. the exact mathematical definition;
  2. the object and population over which it is defined;
  3. the estimator computed from finite records;
  4. units, normalization, and direction of improvement;
  5. uncertainty method and confidence or credibility level;
  6. treatment of missing records, zero counts, and boundary values.

For KK independently sampled circuit scores ZkZ_k, the suite mean is

μ^=1K∑k=1KZk.\widehat\mu = \frac1K \sum_{k=1}^{K} Z_k.

An estimator of its between-circuit variance contribution is

Var⁡^ ⁣(μ^)=1K(K−1)∑k=1K(Zk−μ^)2.\widehat{\operatorname{Var}} \!\left( \widehat\mu \right) = \frac1{K(K-1)} \sum_{k=1}^{K} \left( Z_k-\widehat\mu \right)^2.

Shot uncertainty inside each ZkZ_k may add another component. Treating all shots from all circuits as exchangeable generally understates uncertainty.

If a metric is estimated after fitting a model, report fit range, weights, constraints, nuisance parameters, residuals, and goodness-of-fit diagnostics. A standard protocol name does not uniquely specify these choices.

A point estimate should be accompanied by a decomposition such as

utot2=ushot2+uinstance2+udrift2+ucal2+umodel2,u_{\mathrm{tot}}^2 = u_{\mathrm{shot}}^2 + u_{\mathrm{instance}}^2 + u_{\mathrm{drift}}^2 + u_{\mathrm{cal}}^2 + u_{\mathrm{model}}^2,

when those components can reasonably be separated. Correlated contributions require covariance terms rather than addition in quadrature.

State:

  • interval type and coverage or credibility level;
  • analytic, bootstrap, Bayesian, concentration, or other construction;
  • resampling unit and number of resamples;
  • priors and sensitivity for Bayesian intervals;
  • covariance among jointly reported quantities;
  • systematic and calibration bounds;
  • model-selection and extrapolation uncertainty;
  • multiplicity correction or simultaneous coverage when many results are searched.

An error bar whose construction is not named cannot be audited. Standard deviation, standard error, confidence interval, posterior credible interval, and prediction interval answer different questions.

For nonlinear propagation q=f(θ)q=f(\boldsymbol\theta), first-order covariance is

Cov⁡(q^)≈JfCov⁡ ⁣(θ^)JfT,\operatorname{Cov}(\widehat q) \approx J_f \operatorname{Cov} \!\left( \widehat{\boldsymbol\theta} \right) J_f^{\mathsf T},

but bootstrap or posterior propagation is preferable near boundaries, singularities, or multimodal fits.

Report resources as a vector rather than one flattering scalar:

R=(nphys,nlogical,G,D,S,TQPU,Twall,Cclass,M,E).\mathbf R = \left( n_{\mathrm{phys}}, n_{\mathrm{logical}}, G, D, S, T_{\mathrm{QPU}}, T_{\mathrm{wall}}, C_{\mathrm{class}}, M, E \right).

Possible components are physical and logical qubits, gate counts GG, depth DD, shots SS, quantum-processing time, end-to-end wall time, classical compute, memory, and energy. Not every study needs every component, but the boundary and omissions must be explicit.

Include:

  • compilation and state-preparation cost;
  • calibration and characterization amortization;
  • queueing, network, reset, measurement, and feedback;
  • discarded and failed attempts;
  • mitigation and verification overhead;
  • classical preprocessing, optimization, decoding, and postprocessing;
  • reference-solution and baseline cost;
  • whether a projected cost is measured, simulated, fitted, or assumed.

For a success probability paccp_{\mathrm{acc}} and cost CtryC_{\mathrm{try}} per independent attempt, a simple expected cost to accepted result is

Cacc=Ctrypacc,C_{\mathrm{acc}} = \frac{ C_{\mathrm{try}} }{ p_{\mathrm{acc}} },

under identical independent attempts. Reporting only successful-run cost hides retries and postselection.

Evidence lineage from a scoped claim through protocol, executable, records, analysis, and released artifacts

A report is an evidence graph. Stable identifiers and hashes bind the scoped claim to a protocol, operating context, exact executable, raw records, analysis environment, and released result. Corrections create new versions without erasing the earlier lineage.

Every report needs the common record above. Additional fields depend on the result class.

profileadditional reporting focus
component characterizationmodel family, experiment design, identifiability, residuals, held-out prediction
tomographytrusted SPAM, informational completeness, physical estimator, reconstruction uncertainty
certificationnull model, trust assumptions, finite-data test, losses, loopholes, decision threshold
randomized benchmarkensemble, realized sequences, fit range, sequence and shot hierarchy, compilation
application or hybrid algorithmtask contract, inputs, optimizer history, acceptance, end-to-end resources
logical experimentcode and decoder, rounds, events, leakage, logical denominator, failure definition
simulationequations, discretization, approximations, convergence, reference validation, numerical error
resource estimatelogical algorithm, error budget, code model, factories, schedule, scenario uncertainty
advantage claimcorrectness, scoped hardness, dated classical frontier, matched resources, independent checks

Calling a report “compliant” without naming its profile is incomplete. A profile can also be extended for a platform or protocol, but extensions should not weaken the common fields.

In addition to the common record, report:

  • estimand and physical system boundary;
  • Hamiltonian, channel, generator, drift, or hidden-state model;
  • preparation and measurement assumptions;
  • experiment settings and sensitivity design;
  • identifiable parameter combinations and gauge treatment;
  • likelihood, constraints, priors, and optimization;
  • parameter covariance or joint uncertainty;
  • residuals versus time, setting, sequence length, and context;
  • held-out prediction and alternative-model comparison;
  • proposed calibration action and independent validation.

A table of T1T_1, T2T_2, frequency, and “gate error” does not establish how those quantities were estimated or whether they remained valid during the reported workload. The canonical Device Characterization page develops the inference workflow.

For tomography, add:

  • state, process, measurement, or instrument object;
  • Hilbert-space dimension and leakage convention;
  • preparations, settings, effects, and completeness test;
  • trusted calibration and SPAM boundary;
  • raw counts for every setting;
  • linear, constrained, likelihood, Bayesian, or other estimator;
  • positivity, complete-positivity, trace, or normalization constraints;
  • bias and uncertainty procedure;
  • residual and goodness-of-fit tests;
  • derived quantities with propagated uncertainty.

Do not publish only a reconstructed matrix rendered as a color plot. Supply the numeric object, basis and vectorization conventions, count table, and estimator.

For certification, add:

  • exact null class;
  • trusted, one-sided-device-independent, or device-independent assumptions;
  • test statistic and null bound;
  • setting generation and causal order;
  • treatment of no-click, inconclusive, and leaked outcomes;
  • stopping rule, memory assumptions, and finite-data method;
  • multiplicity or witness-selection control;
  • confidence statement or valid pp-value;
  • the precise population conclusion.

A reconstructed entangled state and a loophole-aware Bell certificate are different evidence products. Certification of Entanglement owns that trust hierarchy.

Report:

  • protocol and exact variant;
  • target gates, cycles, layers, or circuit family;
  • random ensemble and sampling algorithm;
  • realized sequences or sufficient regeneration artifacts;
  • sequence lengths and number of independent sequences per length;
  • shots per sequence and batching;
  • inverse or mirror construction;
  • compilation freedom and target snapshot;
  • survival, polarization, heavy-output, or score definition;
  • fit family, fit range, weights, nuisance parameters, and residuals;
  • sequence-aware uncertainty;
  • leakage, drift, and nonexponential checks;
  • interleaved, simultaneous, twirled, or mitigation conditions;
  • every conversion from decay parameter to fidelity-like quantity.

If random sequences are resampled for each device, comparisons include between-suite variation. If the same sequences are reused, the pairing should be preserved in analysis. Publishing only fitted decay constants prevents either audit.

Report the full task:

  • input distribution and every realized instance;
  • encoding, Hamiltonian or oracle construction, and data access;
  • algorithmic variant, ansatz, and approximation choices;
  • output and acceptance criterion;
  • quantum and classical workflow graph;
  • optimizer settings, seeds, histories, and restarts;
  • error mitigation, postselection, retries, and verification;
  • solution quality and independent reference;
  • success rate and time or cost to accepted answer;
  • scaling design and which resources were measured or projected;
  • classical baselines with matched input, accuracy, tuning, and hardware budgets.

Kernel fidelity is not an application result. Conversely, an application score without the component and resource record cannot diagnose why systems differ.

For variational energy estimation, distinguish:

Eraw,Emit,Ebest,Emean,E_{\mathrm{raw}}, \qquad E_{\mathrm{mit}}, \qquad E_{\mathrm{best}}, \qquad E_{\mathrm{mean}},

where “best” and “mean” must name the seed, iteration, and instance populations. State whether comparison with a reference includes basis, active-space, relativistic, and other modeling errors.

Report:

  • physical layout, code family, distance, boundaries, and logical operators;
  • stabilizer or gauge-check schedule;
  • number of rounds and initialization/finalization conventions;
  • circuit-level operations and detector definition;
  • decoder, version, weights, training data, and latency;
  • leakage, erasure, reset, feedforward, and postselection;
  • number of shots, detected events, accepted shots, and logical opportunities;
  • logical failure definition and confidence interval;
  • physical comparison point and matching circuit context;
  • whether thresholds or scaling exponents are fitted, assumed, or simulated.

The denominator is critical. Per-round, per-cycle, per-shot, and per-logical operation error rates are not interchangeable. If FF logical failures occur over NN shots with RR protected rounds, both

p^shot=FN\widehat p_{\mathrm{shot}} = \frac{F}{N}

and an inferred per-round model require their own assumptions. The shortcut F/(NR)F/(NR) is valid only in a stated rare, independent, constant-hazard approximation.

Report:

  • equations, Hamiltonian, channel, or circuit semantics;
  • initial and boundary conditions;
  • units and parameter provenance;
  • discretization, basis truncation, time step, bond dimension, cutoff, or sampling method;
  • solver and library versions;
  • numerical precision and hardware;
  • convergence sweeps and stopping tolerances;
  • normalization, conservation, symmetry, and limiting-case checks;
  • comparison with an independent method or exact small case;
  • stochastic seeds and sample count;
  • wall time, memory, and accelerator use;
  • machine-readable source data behind plots.

Separate physical-model uncertainty from numerical error:

Δtotal≢Δnumerical.\Delta_{\mathrm{total}} \not\equiv \Delta_{\mathrm{numerical}}.

A numerically converged answer can be wrong for the intended device because the model omits drift, leakage, correlations, or control distortion.

Report a layered estimate:

  • source problem and target success or precision;
  • logical algorithm and oracle/input assumptions;
  • logical qubits, non-Clifford counts, measurements, and dependency depth;
  • approximation, synthesis, and algorithmic failure budgets;
  • code family, distance-selection rule, cycle time, and physical error model;
  • factory protocol, throughput, footprint, routing, and storage;
  • decoder and classical-control assumptions;
  • schedule, concurrency, and bottleneck calculation;
  • physical qubits, runtime, energy boundary, and availability assumptions;
  • scenario ranges, sensitivity analysis, and correlated uncertainties;
  • measured inputs versus road-map assumptions.

The canonical Resource Estimation Tools page owns the estimation workflow. A report must preserve enough of that model to recompute the headline scenario after any hardware assumption changes.

An advantage claim adds unusually strong obligations:

  • correctness evidence for the quantum output;
  • exact task and scoped computational claim;
  • hardness assumptions and known loopholes;
  • dated inventory of classical algorithms, implementations, and hardware;
  • matched accuracy, success probability, and resource boundaries;
  • tuning and preprocessing budgets on both sides;
  • uncertainty in the crossover or separation;
  • adversarial alternative explanations and spoofing tests;
  • independent reproduction or challenge access;
  • explicit separation of measured facts from extrapolation.

A moving classical frontier is part of the experiment. Archive baseline code, compiler flags, data layout, hardware counters, and dated literature search. The comparison should be updateable without rerunning unrelated parts of the quantum experiment.

Human prose supplies interpretation; a manifest supplies structure and validation. A minimal shape might be:

{
"schema_version": "qi-report/1.0",
"record_id": "urn:example:report:2026-08-10:001",
"claim_id": "C1",
"profile": "randomized-benchmark",
"system": {
"backend_id": "provider:backend:revision",
"calibration_snapshot": "sha256:..."
},
"executable": {
"source": "artifact:circuit-source",
"source_sha256": "...",
"scheduled": "artifact:scheduled-circuit",
"scheduled_sha256": "..."
},
"acquisition": {
"started_at": "2026-08-10T18:00:00Z",
"independent_sequences": 100,
"shots_per_sequence": 1000
},
"analysis": {
"release": "doi:...",
"environment_digest": "sha256:..."
},
"result": {
"metric": "declared_metric_id",
"estimate": 0.0,
"interval": {
"type": "declared_interval_type",
"level": 0.95,
"lower": 0.0,
"upper": 0.0
}
},
"artifacts": [
{
"id": "artifact:raw-counts",
"status": "reported",
"sha256": "...",
"media_type": "application/parquet"
}
]
}

The values are schematic, not a universal schema. A community schema should define types, units, required fields, enumerations, and migrations. It should also distinguish zero from missing and a point value from a bound.

For an artifact with bytes BB, record a digest such as

h=SHA256⁡(B).h = \operatorname{SHA256}(B).

A digest establishes identity, not correctness. It answers whether two artifacts are byte-for-byte the same. Correctness still requires validation, authorship, provenance, and review.

Use:

  • a persistent identifier for the released package;
  • a version-specific identifier for the exact release;
  • content digests for files and immutable objects;
  • semantic version or schema version for software and metadata;
  • internal claim and artifact identifiers for cross-reference.

“Latest” is not a reproducible version. A repository branch can move after publication.

Provenance records which entities were used or generated by which activities, and which agents were responsible. A minimal lineage is:

source→compileexecutable→runrecords→analyzeresult.\text{source} \xrightarrow{\text{compile}} \text{executable} \xrightarrow{\text{run}} \text{records} \xrightarrow{\text{analyze}} \text{result}.

Each arrow needs an activity identifier, software environment, inputs, outputs, timestamp, and parameters. The W3C PROV model and RO-Crate provide general-purpose structures that can be specialized rather than inventing an unrelated provenance vocabulary for every benchmark.

Lineage should include negative and intermediate results when they affect selection. If 50 layouts were compiled and one was chosen, the chosen executable has selection provenance from all 50 candidates and the ranking rule.

Publish the least processed record that can be released, together with:

  • schema, field definitions, units, and controlled vocabularies;
  • row or event identifiers and timestamps;
  • mapping from settings to records;
  • quality flags without deleting failed rows;
  • transformation history;
  • checksums and file sizes;
  • license, citation, and retention policy.

For ordinary circuit experiments, per-circuit counts may be sufficient for many analyses. Drift, stopping, memory, or adaptive claims may require shot or block order. State explicitly which information was aggregated away.

Archive:

  • analysis and figure-generation code;
  • compiler and workload-generation code when material;
  • tests and expected outputs;
  • environment lockfile or container recipe;
  • platform and architecture constraints;
  • run instructions and approximate resource needs;
  • exact tagged release and software citation.

Code availability is not code validation. Report which tests ran, on what environment, and whether an independent party reproduced the figures or tables.

Record operating system, architecture, language runtime, dependencies, accelerators, drivers, and environment variables that affect behavior. Containers improve portability but are not timeless: images need immutable digests, base-image provenance, and an archive independent of one registry.

FAIR data are findable, accessible, interoperable, and reusable. “Accessible” permits authenticated or controlled access described by metadata. FAIR does not require that sensitive or proprietary records be public.

When artifacts are restricted, report:

  • artifact inventory and metadata;
  • responsible custodian;
  • reason and legal or policy basis;
  • eligibility and request procedure;
  • expected response time;
  • whether reviewers received access;
  • synthetic or redacted alternatives;
  • expiration or future-release condition;
  • which analyses cannot be checked without access.

An unavailable artifact can still have a persistent metadata record and citation. A statement that data are “available on reasonable request” should define reasonable, identify the decision maker, and provide a durable request route.

Tables and Figures Must Point Back to Data

Section titled “Tables and Figures Must Point Back to Data”

Every table and figure should identify:

  • source data object and version;
  • analysis script and release;
  • plotted population and exclusions;
  • axis units and transformations;
  • uncertainty-bar meaning;
  • smoothing, aggregation, or normalization;
  • number of independent units;
  • whether displayed examples were selected.

Publish machine-readable table values and source points. Rasterized curves and rounded PDF tables should not be the only data release. Insets, cropped axes, and logarithmic scales must not obscure zero, failure counts, or outliers relevant to interpretation.

Representative examples need a selection rule. A “typical circuit” chosen by eye after inspecting all outcomes is an exploratory illustration.

Comparability Requires a Compatibility Audit

Section titled “Comparability Requires a Compatibility Audit”

Before ranking two results, compare:

dimensioncompatibility question
metricsame definition, normalization, estimator, and direction?
objectsame operation, layer, workload, code, or task?
populationsame input, circuit, instance, seed, and time distributions?
system boundarysame inclusion of SPAM, compiler, queue, classical work, and verification?
controlssame optimization, mitigation, postselection, and stopping freedom?
statisticscomparable independent units, uncertainty, and selection correction?
resourcessame cost boundary and amortization?
validityacquired under relevant and overlapping operating conditions?

If dimensions differ, report a conditional comparison rather than a universal ranking. A Pareto table is often more honest than collapsing quality, time, qubits, and availability into one score.

Normalized ratios can inherit uncertainty and denominator instability. For

R=qAqB,R = \frac{q_A}{q_B},

the analysis should preserve covariance when qAq_A and qBq_B share instances or references. Independent error propagation can be either too conservative or anti-conservative depending on that covariance.

A released evidence package should be immutable. Corrections create a new version that:

  • identifies the superseded release;
  • lists changed artifacts and digests;
  • explains the reason and impact on every claim;
  • preserves access to the earlier version when policy permits;
  • reruns validation and updates the manifest;
  • distinguishes erratum, reanalysis, new data, and retraction.

Do not silently replace a data file behind an unchanged DOI or URL. Mutable dashboards should offer dated snapshots for cited results.

A correction that changes an estimate but not the conclusion is still material to reproducibility. Conversely, a calibration update after the study does not retroactively change the archived experimental record.

Consider the statement:

A variational quantum algorithm reached chemical accuracy on a 20-qubit processor in 90 seconds.

It leaves the claim undefined. A complete report would identify:

  1. molecule, geometry, basis, active space, mapping, Hamiltonian coefficients, and reference energy;
  2. ansatz, parameter count, circuit source, mapped executable, qubit layout, compiler, and calibration snapshot;
  3. processor revision, acquisition dates, shots, jobs, failed attempts, and drift controls;
  4. optimizer, initialization, seed ensemble, objective history, stopping rule, and best-run selection;
  5. mitigation and postselection methods, raw and transformed energies, and retention;
  6. statistical, seed, calibration, model, and reference uncertainties;
  7. whether “chemical accuracy” covers electronic-model error or only error relative to a finite-basis exact diagonalization;
  8. whether 90 seconds is QPU execution, successful-job wall time, or complete time to accepted answer;
  9. classical baseline algorithm, implementation, hardware, accuracy, and tuning budget;
  10. code, data, environment, manifests, and stable artifact identifiers.

The repaired headline might be longer, but it would say which quantity crossed which threshold under which boundary. The detailed ledger and artifact package can then carry the full record without overloading the abstract.

Before release, verify:

  • every headline sentence maps to a claim identifier;
  • every reported number has a definition, population, and unit;
  • hardware, calibration, compiler, executable, and analysis versions are fixed;
  • acquisition hierarchy, failures, exclusions, and retention are complete;
  • raw and transformed results are distinguishable;
  • uncertainty construction and independent units are stated;
  • optimizer, mitigation, and selection freedoms are disclosed;
  • quantum and classical resource boundaries match the claim;
  • data, code, and environment releases have persistent versioned identifiers;
  • restricted artifacts have metadata and access procedures;
  • figures and tables map to source data and scripts;
  • the manifest validates and every digest resolves;
  • limitations and non-generalization scope are explicit;
  • correction and contact routes are durable.

A reviewer should be able to answer:

  1. What exact result is being claimed?
  2. What was executed, on which system, and when?
  3. What are the independent random units?
  4. Which records were excluded, retried, postselected, or transformed?
  5. How was uncertainty obtained, and does it cover selection?
  6. Which assumptions convert the estimator into the stated physical claim?
  7. Are quality, time, and resource boundaries matched across comparisons?
  8. Can each figure and table be regenerated?
  9. Which artifacts are inaccessible, and how does that limit review?
  10. What observation would falsify or materially revise the conclusion?

Failure to answer one question does not automatically invalidate the science, but it should narrow the accepted claim and be recorded as a review limitation.

Reporting a platform label instead of a system

Section titled “Reporting a platform label instead of a system”

Backends evolve. Bind results to a revision, date, qubit subset, and calibration epoch.

Publishing source circuits but not executed circuits

Section titled “Publishing source circuits but not executed circuits”

Placement, routing, scheduling, synthesis, and mitigation can dominate the result. Preserve the lowered artifacts and pass configuration.

Random circuits, instances, seeds, days, and devices may be the independent units. Report the hierarchy.

Success-conditioned reports overstate availability and understate time to solution.

Reporting only mitigated or postselected values

Section titled “Reporting only mitigated or postselected values”

Readers need raw values, retention, overhead, fit sensitivity, and unconditional performance.

State interval type, level, resampling unit, assumptions, and systematic budget.

Archive a release and cite the exact version. A repository homepage is not an artifact identity.

Data need schema, provenance, code, environment, and validation. Public bytes alone are not an analysis.

Layout search, compiler tuning, mitigation choice, optimizer restarts, and metric selection are experimental resources and selection mechanisms.

Publish a new version with a change log and claim-impact statement.

General principles for measurement uncertainty, data and software citation, provenance, FAIR stewardship, and computational artifact review are mature. Quantum-specific benchmark reporting is less standardized. Different protocols, providers, and suites still use incompatible schemas, runtime boundaries, circuit abstractions, and calibration records.

Active work includes:

  • interoperable benchmark and calibration schemas;
  • portable representations of dynamic and pulse-level executables;
  • privacy-preserving release of cloud and hardware telemetry;
  • provenance across hybrid HPC–quantum workflows;
  • long-term replay of provider-dependent jobs;
  • uncertainty standards for drift, adaptive experiments, and selected benchmark winners;
  • machine-checkable resource and advantage claims.

Recent empirical work on quantum-computing artifacts indicates that code, environment, and clean execution remain unevenly available. Such surveys are useful diagnostics, but their sampling frames and automated detection methods must themselves be reported under the same principles.

  • Variational Quantum Algorithms defines the adaptive training, estimator, stopping, validation, and all-seed resource record required by the hybrid-algorithm profile.
  • VQE specializes that record to encoded Hamiltonians, energy estimators, variational-bound claims, model error, and independent final measurements.
  • Quantum Chemistry Case Studies shows how that record changes the interpretation of small-molecule experiments, logical demonstrations, and fault-tolerant chemistry estimates.
  • Quantum Software Stack identifies the artifacts and responsibility boundaries that provenance must connect.
  • Circuit Intermediate Representations develops typed, versioned executable representations across lowering.
  • Metrics for Quantum Hardware supplies the metric definitions whose objects and assumptions must appear in reports.
  • Calibration Loops owns calibration epochs, validity predicates, publication, rollback, and quarantine.
  • Reproducible Notebooks gives the repository-specific contract for executable notebook artifacts.
  • Claims, Hype, and Evidence Standards turns the evidence record into a correctly scoped public statement.
  • Claims and Evidence Checklist supplies the reader-facing gate procedure for deciding whether that statement is supported, must be narrowed, or remains unresolved.
  1. National Academies of Sciences, Engineering, and Medicine, Reproducibility and Replicability in Science (National Academies Press, 2019), doi:10.17226/25303.
  2. M. D. Wilkinson et al., “The FAIR Guiding Principles for scientific data management and stewardship,” Scientific Data 3, 160018 (2016), doi:10.1038/sdata.2016.18.
  3. Data Citation Synthesis Group, “Joint Declaration of Data Citation Principles,” FORCE11 (2014), doi:10.25490/a97f-egyk.
  4. A. M. Smith, D. S. Katz, and K. E. Niemeyer, “Software citation principles,” PeerJ Computer Science 2, e86 (2016), doi:10.7717/peerj-cs.86.
  5. T. Lebo, S. Sahoo, and D. McGuinness, editors, “PROV-O: The PROV Ontology,” W3C Recommendation (2013), W3C PROV-O.
  6. S. Soiland-Reyes et al., “Packaging research artefacts with RO-Crate,” Data Science 5, 97–138 (2022), doi:10.3233/DS-210053.
  7. Association for Computing Machinery, “Artifact Review and Badging, Version 1.1,” (2020), ACM policy.
  8. Joint Committee for Guides in Metrology, Evaluation of Measurement Data: Guide to the Expression of Uncertainty in Measurement, JCGM 100:2008 (2008), BIPM publication.
  9. E. Marceaux et al., “A practical introduction to benchmarking and characterization of quantum computers,” PRX Quantum 6, 030202 (2025), doi:10.1103/PRXQuantum.6.030202.
  10. T. J. Proctor et al., “Benchmarking quantum computers,” Nature Reviews Physics (2025), doi:10.1038/s42254-024-00796-z.
  11. J. Eisert et al., “Quantum certification and benchmarking,” Nature Reviews Physics 2, 382–390 (2020), doi:10.1038/s42254-020-0186-4.
  12. J. M. Lorenz et al., “Systematic benchmarking of quantum computers: status and recommendations,” arXiv:2503.04905 (2025), arXiv:2503.04905.
  13. T. Lubinski et al., “Application-oriented performance benchmarks for quantum computing,” IEEE Transactions on Quantum Engineering 4, 1–32 (2023), doi:10.1109/TQE.2023.3253761.
  14. D. Mills et al., “Application-motivated, holistic benchmarking of a full quantum computing stack,” Quantum 5, 415 (2021), doi:10.22331/q-2021-03-22-415.
  15. M. Amico et al., “Defining standard strategies for quantum benchmarks,” arXiv:2303.02108 (2023), arXiv:2303.02108.
  16. A. W. Cross et al., “Validating quantum computers using randomized model circuits,” Physical Review A 100, 032328 (2019), doi:10.1103/PhysRevA.100.032328.
  17. E. Nielsen, K. Rudinger, T. Proctor, R. Blume-Kohout, and K. Young, “Gate set tomography,” Quantum 5, 557 (2021), doi:10.22331/q-2021-10-05-557.
  18. S. J. van Enk and R. Blume-Kohout, “When quantum tomography goes wrong: drift of quantum sources and other errors,” New Journal of Physics 15, 025024 (2013), doi:10.1088/1367-2630/15/2/025024.
  19. J. J. Wallman and S. T. Flammia, “Randomized benchmarking with confidence,” New Journal of Physics 16, 103032 (2014), doi:10.1088/1367-2630/16/10/103032.
  20. N. Quetschlich et al., “MQT Bench: Benchmarking software and design automation tools for quantum computing,” Quantum 7, 1062 (2023), doi:10.22331/q-2023-07-20-1062.
  21. D. Köster et al., “Works on my QPU: Reproducibility in quantum computing research,” arXiv:2607.08348 (2026), arXiv:2607.08348.
  22. R. D. Peng, “Reproducible research in computational science,” Science 334, 1226–1227 (2011), doi:10.1126/science.1213847.
  23. G. K. Sandve et al., “Ten simple rules for reproducible computational research,” PLoS Computational Biology 9, e1003285 (2013), doi:10.1371/journal.pcbi.1003285.
  24. M. R. Munafò et al., “A manifesto for reproducible science,” Nature Human Behaviour 1, 0021 (2017), doi:10.1038/s41562-016-0021.

A paper states, “Our two-qubit gates have 99.7%99.7\% fidelity.” List eight fields needed before this becomes an interpretable result.

Solution

At minimum, ask for:

  1. the physical or logical gate object and target;
  2. device, qubit pair, date, and calibration snapshot;
  3. metric definition and conversion convention;
  4. characterization protocol and variant;
  5. isolated or simultaneous circuit context;
  6. treatment of SPAM, leakage, loss, and mitigation;
  7. sample hierarchy, fit model, and uncertainty;
  8. compiler, pulse, and executable versions.

The report should also state selection across qubit pairs or dates and provide the underlying counts or protocol records. “Fidelity” alone does not identify average gate fidelity, process fidelity, entanglement fidelity, or a fitted RB quantity.

An RB experiment samples K=30K=30 random sequences at one length and takes 2,0002{,}000 shots per sequence. Which number primarily controls uncertainty in the ensemble mean after shot noise is already small?

Solution

The 30 independently sampled sequences control the uncertainty from variation over the randomized ensemble. The 60,000 shots are not exchangeable draws from that ensemble; they are nested inside 30 sequences. Once each sequence score is precise, taking more shots does little to estimate between-sequence variation. The report should provide sequence-level scores or a hierarchical analysis.

A protocol retains 7,2007{,}200 of 10,00010{,}000 trials and has conditional success probability 0.950.95. Compute the retention fraction and the fraction of all attempts that are both retained and successful.

Solution

The retention fraction is

r^=720010000=0.72.\widehat r = \frac{7200}{10000} = 0.72.

The unconditional accepted-success fraction is

0.72×0.95=0.684.0.72\times0.95 = 0.684.

Reporting only 0.950.95 hides nearly one third of attempted trials. A complete report also gives retention by input and context, uncertainty, the eligibility rule, and the cost of detecting or restarting rejected trials.

A hybrid optimizer runs 40 seeds and reports the lowest energy. What should accompany that number?

Solution

Report all 40 initializations or their generator, complete objective histories, stopping and failure rules, per-seed quantum and classical cost, and the selection rule. Summaries should include the distribution across seeds, not only the minimum. Any interval for expected workflow performance must account for seed variation and best-of-40 selection. The total cost includes all 40 runs.

5. Convert trial cost to accepted-result cost

Section titled “5. Convert trial cost to accepted-result cost”

One attempt costs 1212 QPU-seconds and succeeds independently with probability 0.300.30. What is the expected QPU cost per accepted result under the simple repeat-until-success model?

Solution

The expected number of attempts is 1/0.301/0.30, so

Cacc=120.30=40 QPU-seconds.C_{\mathrm{acc}} = \frac{12}{0.30} = 40\ \text{QPU-seconds}.

This excludes queueing, verification, reset, and classical work. If failures have different costs or attempts are correlated, the simple geometric model must be replaced.

6. Decide whether two runtimes are comparable

Section titled “6. Decide whether two runtimes are comparable”

System A reports circuit execution time only. System B reports compilation, queue, execution, measurement, and classical postprocessing. May the two numbers be ranked as “time to solution”?

Solution

No. Their system boundaries differ. The authors can compare overlapping execution components if those definitions and workload conditions match, or recompute both end-to-end times under one boundary. Until then, the comparison must be labeled conditional and cannot support a universal time-to-solution ranking.

An analysis bug changes a benchmark estimate from 0.740.74 to 0.710.71 but leaves the pass/fail decision unchanged. Should the original data package be overwritten?

Solution

No. Release a new immutable version that cites and supersedes the old one, identifies the bug, lists changed files and digests, and states that the decision is unchanged. Preserve the original release when policy permits. Readers must be able to determine which estimate and analysis version a citation used.

A provider cannot release pulse telemetry for security reasons. What can a transparent report still provide?

Solution

It can publish an artifact inventory, schema and time coverage, immutable identifier or digest, reason and policy basis for restriction, custodian, reviewer-access statement, request procedure, redacted or aggregated alternative, and a claim-impact statement. It should release all unrestricted circuits, counts, analysis code, calibration summaries, and validation records. The report must state which hardware or drift claims cannot be independently audited without the telemetry.