Skip to content

Reproducibility Checklist

A computational result is ready to support a scientific statement only when another reader can identify the statement, reconstruct the calculation, run its checks, and trace every published number or figure back to retained evidence. This checklist helps assess computational claims across the collection, including approximation theory, scattering, semiclassics, tunneling, and driven dynamics.

Reproducibility is not the same as correctness. A fully repeatable program can implement the wrong Hamiltonian, and a correct number can be unreproducible if its environment and parameters are missing. Trust requires both an executable chain and independent validation.

Use this page before publishing a notebook, script, numerical table, or generated figure. The site-wide Notebook Index owns the common metadata contract, Reproducibility Status owns status labels, and Validation Tests owns the test taxonomy. This page turns those principles into an end-to-end computational audit.

QuestionCanonical homeRole here
What metadata must every notebook record?Notebook Indexapply the contract to one release
Which environment details and dependencies matter?Environmentsverify that the environment is reconstructible
Which status label may be assigned?Reproducibility Statuscollect the evidence needed for that decision
What kinds of executable checks exist?Validation Testschoose checks with independent failure modes
How should refinement and uncertainty be quantified?Convergence Tests and Error Estimatesconfirm that the reported observable is resolved
Is the physical approximation itself controlled?Error-Estimate Checklistkeep model error separate from numerical error
Is one computational result ready to release?this pageperform the end-to-end audit

This checklist does not create a second status system or duplicate numerical-method derivations. It asks whether the evidence required by those canonical pages is present, connected, and current.

Terminology varies across disciplines. This page follows the operational distinction used by the National Academies report on reproducibility and replicability:

  • computational reproducibility: consistent results are obtained from the same input data, code, computational steps, and stated conditions;
  • replicability: a new study addressing the same scientific question obtains consistent results using independently collected evidence or an independently constructed calculation;
  • repeatability: a narrower engineering term often used for rerunning under nearly identical conditions.

For a deterministic quantum calculation, the first release target is computational reproducibility. An independent finite-difference check of a basis calculation moves toward replication because it changes the numerical representation and its failure modes. Merely rerunning the same executable twice tests less.

State what agreement is required.

LevelRequirementAppropriate use
bitwiseoutput bytes or hashes are identicaldeterministic archives on a fixed stack
numericaldeclared quantities agree within justified tolerancesfloating-point work across compatible platforms
scientificconclusions and uncertainty statements agreeindependent methods or environments

Bitwise identity is strong provenance evidence, but it is neither necessary nor sufficient for scientific agreement. Parallel reductions, linear-algebra libraries, and hardware instructions can alter low-order bits without changing a resolved observable. Conversely, identical wrong code produces identical wrong bytes.

Start from the scientific claim and work backward:

claim
<- quoted value, table, or figure
<- retained machine-readable output
<- validated computation
<- code plus parameters and inputs
<- declared environment
<- source revision, run record, and license

Every arrow must be inspectable. A figure that cannot be regenerated from retained data is detached from its evidence. A CSV with no parameter record is detached from the physical problem. A script with no validation target is executable, but not yet trustworthy.

Before running anything, write one sentence of the form:

Running command in environment with inputs and parameters must regenerate outputs, pass named checks at stated tolerances, and support this claim.

This sentence prevents the audit from collapsing into the vague question “does the notebook run?” A notebook may execute successfully while producing the wrong branch, normalization, phase convention, or asymptotic regime.

A reference-grade computational result should retain the following small, inspectable packet.

ArtifactRequired contentWhy it matters
explanatory pagequestion, conventions, result, limitations, canonical linksstates what the computation means
executable sourcenotebook, script, or small programreconstructs the transformation
parameter recordphysical and numerical controlsprevents hidden defaults
environment recordruntime, packages, versions, optional dependenciesreconstructs execution conditions
input recordlocal data, provenance, checksum, preprocessingidentifies what entered the calculation
retained outputscompact CSV, JSON, or text summariespreserves evidence independently of plotting software
validation reportchecks, targets, tolerances, pass/fail resultsdistinguishes execution from validation
figure sourceplotting or TikZ source tied to retained datamakes visual evidence regenerable
provenance recorddate, source revision, command, runtime classidentifies the validated run
license informationcode, data, and third-party termsdetermines whether others may reproduce and reuse it

Large opaque binary outputs are rarely the best release artifact. Retain the smallest lossless or scientifically sufficient data product from which the reported table or figure can be regenerated.

The checkboxes below are ordered. A missing physics convention cannot be repaired by a more detailed package list, and a passing regression test cannot compensate for an unconverged observable.

  • The page states the exact quantity or qualitative conclusion supported by computation.
  • The claim names its parameter regime and excludes regimes not tested.
  • The canonical analytic page is linked instead of rederived in the notebook.
  • Exploratory observations are labeled separately from validated results.
  • The required matching level is declared: bitwise, numerical, or scientific.
  • The number of printed digits is justified by the error estimate.

A claim such as “the method works” is too broad. Prefer “for 0.05≤g≤0.200.05\le g\le0.20, the lowest gap is stable under the retained basis refinements and differs from the one-loop estimate by the reported amount.”

  • The Hamiltonian or evolution equation is stated.
  • Units or dimensionless scales are explicit.
  • Initial, boundary, and asymptotic conditions are explicit.
  • Basis ordering, Fourier sign, phase, branch, and normalization conventions are recorded where relevant.
  • Symmetry sector, particle statistics, reduced mass, and gauge choice are named when they affect the answer.
  • The reported observable is defined mathematically.
  • Model assumptions and omitted physics are separated from numerical approximations.

For scattering, include incident-state normalization, current factors, partial-wave convention, matching radius, and phase unwrapping. For tunneling, specify whether the output is an amplitude, rate, transmission probability, or spectral splitting. For driven systems, record the time origin, pulse order, and effective-Hamiltonian branch.

  • The representation is named: grid, basis, finite elements, sparse matrix, trajectories, or another explicit form.
  • Every independent cutoff is recorded.
  • Domain size and boundary treatment are recorded separately from resolution.
  • Solver, algorithm, and matrix structure are stated.
  • Arithmetic precision and relevant summation or diagonalization choices are stated.
  • Stopping criteria and solver tolerances are explicit.
  • The refinement path is recorded, including jointly changed controls.
  • Runtime and memory class are stated when they affect practical reproduction.

Do not write “high resolution” or “tight tolerance” without numbers. A basis size N=200N=200 is not reproducible if the basis scale, ordering, symmetry projection, and matrix-element construction remain implicit.

  • Runtime and language versions are recorded.
  • Direct package versions are recorded.
  • A lockfile or constrained environment file is retained when the result supports a published figure or benchmark.
  • Optional plotting, acceleration, and file-format packages are distinguished from required dependencies.
  • Operating system, architecture, accelerator, or linear-algebra backend is recorded when results are sensitive to it.
  • Paths are relative to the repository or supplied input directory.
  • Ordinary execution does not depend on a contributor’s private files, credentials, or shell history.
  • The run command works from a documented working directory.

Prefer the smallest dependency set that honestly implements the calculation. Fewer dependencies reduce installation ambiguity and long-term API drift, but replacing a tested domain library with opaque custom code does not improve reproducibility.

  • Every input file has a stable name and documented role.
  • External data have a source, version or access date, license, and checksum when practical.
  • Download and preprocessing steps are scripted or described exactly.
  • Raw inputs are distinguishable from processed and generated files.
  • Physical constants identify their source and unit convention.
  • Network access is not silently required during the validated run.
  • Generated output cannot overwrite irreplaceable source data without an explicit option.

A checksum establishes file identity, not scientific quality. It answers “is this the same input?” rather than “is this the right input?”

  • The random-number generator and library are named.
  • Seeds or seed-generation rules are recorded.
  • Independent streams are constructed deliberately for parallel work.
  • Sample count, burn-in, thinning, and autocorrelation treatment are recorded when relevant.
  • Statistical uncertainty is estimated from independent information, not inferred from a fixed seed.
  • Multiple seeds or batches are used when the claim concerns an ensemble rather than one trajectory.
  • Deterministic calculations state that no randomness is used.

A fixed seed reproduces one pseudorandom stream. It does not estimate variance, expose seed sensitivity, or remove Monte Carlo bias.

  • At least one structural check is executable.
  • At least one exact limit, analytic benchmark, or manufactured solution is checked when available.
  • Every relevant numerical control is refined.
  • The observable, not only an internal residual, is tested.
  • At least two checks have meaningfully different failure modes when practical.
  • Tolerances are declared before inspecting the final pass/fail result.
  • Tolerances are scale-aware and justified by truncation, roundoff, or statistical uncertainty.
  • Validation fails visibly and returns a nonzero status for scripts used in automation.
  • Known failure regions are retained rather than removed from the report.

For a scalar reference value, a common numerical criterion is

∣q−qref∣≤ϵabs+ϵrel∣qref∣.|q-q_{\mathrm{ref}}| \le \epsilon_{\mathrm{abs}} + \epsilon_{\mathrm{rel}} |q_{\mathrm{ref}}|.

The absolute term controls comparisons near zero. The relative term scales comparisons away from zero. Neither should be copied from a package default without relating it to the physical question.

For a quantity expected to be zero by symmetry, use an explicit physical or numerical scale q∗q_*:

∣q∣q∗≤ϵmathrmsym.\frac{|q|}{q_*} \le \epsilon_{mathrm{sym}}.

For stochastic estimates qˉ1\bar q_1 and qˉ2\bar q_2 with standard errors s1s_1 and s2s_2, a compatibility diagnostic may use

∣qˉ1−qˉ2∣≤ks12+s22,|\bar q_1-\bar q_2| \le k \sqrt{s_1^2+s_2^2},

with the confidence interpretation of kk stated. This test assumes the uncertainty estimates and independence assumptions are credible; it is not a universal substitute for distributional checks.

  • Quoted values appear in a retained machine-readable table.
  • Columns have units or explicit dimensionless definitions.
  • Output precision does not imply unsupported accuracy.
  • Figures are generated from tracked source and retained data.
  • Axes, scales, parameter values, and normalization are labeled.
  • Logarithmic scales and transformed quantities are stated.
  • Curves remain distinguishable without relying on color alone.
  • Captions state what was computed and what validation supports it.
  • Hand editing after export is absent or fully documented.
  • Intermediate caches are either reproducibly generated or excluded from the evidence chain.

The final SVG or PNG is a presentation artifact. The scientific evidence is the linked combination of data, plotting source, parameters, and validation record.

  • The source runs from a fresh checkout or clean kernel.
  • Notebook cells execute in documented order without hidden state.
  • The retained command regenerates outputs in a temporary or declared directory.
  • The validation summary is visible at the end of the run.
  • The run does not rely on stale outputs already present in the repository.
  • A second run is idempotent or documents why output changes.
  • Expected warnings are recorded and unexpected warnings are investigated.
  • The validated run date and source revision are recorded.

When practical, move existing generated outputs aside before the clean run. Otherwise a script that silently fails to regenerate a file can appear successful because an old artifact remains available.

  • Code has an explicit license or SPDX identifier.
  • Data have reuse terms compatible with redistribution.
  • Third-party code, data, and formulas are cited.
  • Software versions are cited when their implementation materially affects the result.
  • The page records a review date.
  • Dependency changes, convention changes, or failed validations trigger status review.
  • Rebaselining a regression value requires a documented scientific or technical reason.
  • Superseded artifacts remain traceable or receive a migration note.

“Available online” is not a license. Without explicit reuse terms, another researcher may be able to inspect an artifact but not legally redistribute or adapt it.

The explanatory page remains primary, but a small YAML manifest reduces ambiguity and makes audits easier to automate. The following example follows the site-wide metadata contract without defining a new status vocabulary.

schema_version: 1
artifact:
id: 'magnus-expansion-error'
kind: 'python-script'
source: 'public/notebooks/approximation-scattering-semiclassics/magnus-expansion-error.py'
canonical_page: 'labs/magnus-expansion-error'
purpose:
claim: 'Cumulative Magnus errors recover the first four expected weak-drive powers.'
matching_level: 'numerical'
execution:
working_directory: 'repository-root'
command: 'python public/notebooks/approximation-scattering-semiclassics/magnus-expansion-error.py'
expected_runtime_seconds: 1
environment:
python: '3.12.13'
packages:
numpy: '2.3.5'
randomness: 'none'
parameters:
eta_min: 0.010
eta_max: 2.000
stroboscopic_eta: 0.16
periods: 200
validation:
- name: 'Magnus-4 weak-drive power'
target: 5.0
tolerance: 0.03
result: 5.000917
status: 'pass'
- name: 'maximum one-period unitarity defect'
upper_bound: 2.0e-14
result: 1.12e-15
status: 'pass'
outputs:
- 'public/data/approximation-scattering-semiclassics/magnus-expansion-error-sweep.csv'
- 'public/data/approximation-scattering-semiclassics/magnus-expansion-stroboscopic.csv'
- 'public/data/approximation-scattering-semiclassics/magnus-effective-hamiltonian-branches.csv'
provenance:
last_run: '2026-07-16'
source_revision: 'record-at-release'
license:
code: 'MIT'
reproducibility_status: 'assign-using-canonical-status-page'

The manifest is useful only if it is kept with the source revision it describes. A precise manifest attached to later, changed code can be more misleading than no manifest.

The author and reviewer should not perform exactly the same check in exactly the same way.

Read the page without running code. Confirm that the claim, model, conventions, parameter regime, outputs, validation targets, and limitations can be understood from retained text and metadata.

Create or select the declared environment, use a clean checkout, run the stated command, and verify that expected outputs are regenerated. Record platform differences and warnings.

Inspect the validation logic. Ask whether each check could pass while the reported claim remained wrong. Add a check with a different failure mode when the answer is yes.

Examples:

  • pair norm conservation with a phase-sensitive observable;
  • pair an eigensolver residual with basis and domain refinement;
  • pair a seeded regression with independent-seed uncertainty;
  • pair a WKB slope with a prefactor-sensitive ratio;
  • pair a matrix identity with an independent discretization.

Choose at least one quoted number and one figure. Trace each backward through retained data and code to the parameter and environment record. Regenerate the figure when practical.

Assign or update the label using Reproducibility Status. A warning must state whether it affects the scientific conclusion. Do not loosen a tolerance silently to recover a preferred label.

The Magnus Expansion Error laboratory provides a compact deterministic example.

Audit itemRetained evidenceResidual limitation
analytic targetexact product of two Pauli exponentialsstep drive is a deliberately small model
sourceNumPy-only Python program with MIT SPDX headeroptional plotting adds Matplotlib
parametersfixed 18-point pulse-scale sweep and two named diagnostic couplingsno adaptive exploration is claimed
environmentPython and NumPy versions printed by the runlow-order bits may vary with numerical libraries
randomnessnonedeterministic execution does not prove correctness
validationweak-drive slopes, unitarity, pulse reversal, commuting control, branch reconstructiontests a two-dimensional Hilbert space only
retained datathree CSV filestables summarize rather than archive every intermediate matrix
figurestwo SVGs generated from CSV tables with retained TikZ sourcesvisual agreement is not itself a pass criterion
long-time checkpowers through 200 periodsno claim is made beyond that horizon

The strongest part of this packet is not that every check passes. It is that the checks answer different questions and the limitations remain visible.

Do not release a computational claim as verified when any critical item below fails.

Critical gateFail conditionRequired response
identifiable claimoutput has no precise scientific interpretationnarrow and rewrite the claim
reconstructible modelHamiltonian, units, conditions, or conventions are missingdocument before rerunning
executable sourceclean run cannot be startedrepair environment or mark broken
resolved observablerelevant refinements are absent or numerical error is too largeextend convergence study or weaken claim
independent validationall checks share the same likely failureadd an exact limit, symmetry, or alternative method
traceable outputtable or figure cannot be regeneratedrestore data and generation source
legal reusecode or data terms are absent or incompatibleclarify licensing before redistribution
current provenancerun predates material code, dependency, or convention changesrerun and review status

Noncritical differences, such as platform-dependent low-order bits inside a much looser scientific tolerance, may support reproduced_with_warnings when explained. The canonical status page owns that label and its evidence requirements.

Reopen the audit when any of the following changes:

  • the Hamiltonian, convention, or canonical formula;
  • a dependency major version or numerical backend;
  • input data, preprocessing, or constants;
  • a solver, tolerance, grid, basis, or random-number generator;
  • the plotting path or retained output schema;
  • a validation threshold or regression reference;
  • the page’s scientific claim;
  • the license or availability of a dependency or dataset.

Do not update a regression value merely because the new run differs. First classify the change as a bug fix, convention change, algorithmic improvement, dependency drift, expected platform variation, or actual regression. Retain the reason with the new baseline.

  • Treating a successfully executed notebook as a validated result.
  • Recording a seed but no sampling uncertainty.
  • Pinning every transitive dependency while omitting the Hamiltonian convention.
  • Reporting an eigensolver residual as if it included basis and domain error.
  • Validating a time integrator only through norm conservation.
  • Comparing relative error to a reference value that crosses zero.
  • Regenerating a figure from edited plotting data rather than retained raw output.
  • Depending on cells executed out of order or variables from a previous session.
  • Using an absolute local path or undocumented network download.
  • Keeping only a screenshot of a table or plot.
  • Silently changing tolerances until a test passes.
  • Assuming a public repository grants a license to reuse code or data.
  • Declaring exact cross-platform byte identity when the scientific claim needs only tolerance-level agreement.
  • Deleting failed parameter points and thereby hiding the method’s validity boundary.

A symmetry-forbidden matrix element has reference value qref=2×10−12q_{\mathrm{ref}}=2\times10^{-12}. A reproduced run gives q=3.2×10−11q=3.2\times10^{-11}. Why is a relative tolerance alone misleading, and how would an absolute scale improve the test?

Solution

The relative discrepancy is

∣q−qref∣∣qref∣=15,\frac{|q-q_{\mathrm{ref}}|}{|q_{\mathrm{ref}}|} = 15,

which looks catastrophic even though both values may be negligible on the physical scale. If the natural matrix-element scale is q∗=1q_*=1, an absolute criterion such as

∣q∣≤10−10q∗|q| \le 10^{-10}q_*

classifies the reproduced value correctly as a small symmetry residual. The scale and threshold must be justified by the computation; choosing them after seeing the result would invalidate the audit.

A Monte Carlo notebook reproduces the same energy to every printed digit when rerun with seed 17. Does this establish a reproducible uncertainty estimate?

Solution

No. It establishes repeatability of one pseudorandom stream. It does not measure estimator variance, autocorrelation, burn-in bias, time-step bias, or sensitivity to the seed. The release should retain the seed for regression and also use independent seeds or batches to estimate statistical uncertainty. Systematic biases require separate checks.

An eigenvalue calculation reports residual ∥HNv−Ev∥=10−13\|H_Nv-Ev\|=10^{-13} in a basis of size N=40N=40, but the energy changes by 10−310^{-3} at N=80N=80. Which claim is supported?

Solution

The small residual shows that the finite N=40N=40 matrix eigenproblem was solved accurately. It does not show that the finite basis represents the continuum Hamiltonian accurately. The 10−310^{-3} basis change is the relevant evidence for the physical energy and dominates the algebraic residual. The energy is not resolved below that scale until basis refinement stabilizes.

Two platforms produce different CSV hashes, but every reported observable agrees within 3×10−133\times10^{-13}, while the justified release tolerance is 10−910^{-9}. How should this be reported?

Solution

Bitwise reproduction failed, but numerical reproduction passed comfortably. Record the platforms, library versions, hash difference, maximum observable discrepancy, and tolerance. If no scientific conclusion changes, the difference is a provenance warning rather than a validation failure. The assigned status should follow the canonical status policy.

A figure has a plotting script and labeled axes, but the plotted CSV is absent and the script downloads a changing URL without a version. Identify the release blockers.

Solution

The figure is detached from stable input evidence. A reviewer cannot know which data version produced it, verify the download, or regenerate the same output later. Retain or archive the exact input when licensing permits, record its source version and checksum, script preprocessing, and connect the plotting source to that retained file. Labeled axes do not repair missing provenance.

  1. National Academies of Sciences, Engineering, and Medicine, Reproducibility and Replicability in Science, National Academies Press (2019), doi:10.17226/25303.
  2. G. K. Sandve, A. Nekrutenko, J. Taylor, and E. Hovig, “Ten Simple Rules for Reproducible Computational Research,” PLOS Computational Biology 9, e1003285 (2013), doi:10.1371/journal.pcbi.1003285.
  3. G. Wilson et al., “Good Enough Practices in Scientific Computing,” PLOS Computational Biology 13, e1005510 (2017), doi:10.1371/journal.pcbi.1005510.
  4. M. D. Wilkinson et al., “The FAIR Guiding Principles for scientific data management and stewardship,” Scientific Data 3, 160018 (2016), doi:10.1038/sdata.2016.18.
  5. R. D. Peng, “Reproducible research in computational science,” Science 334, 1226–1227 (2011), doi:10.1126/science.1213847.
  6. V. Stodden et al., “Enhancing reproducibility for computational methods,” Science 354, 1240–1241 (2016), doi:10.1126/science.aah6168.
  7. The Turing Way Community, The Turing Way: A Handbook for Reproducible, Ethical and Collaborative Data Science, doi:10.5281/zenodo.3233853.