Reproducibility Checklist
A computational result is ready to support a scientific statement only when another reader can identify the statement, reconstruct the calculation, run its checks, and trace every published number or figure back to retained evidence. This checklist helps assess computational claims across the collection, including approximation theory, scattering, semiclassics, tunneling, and driven dynamics.
Reproducibility is not the same as correctness. A fully repeatable program can implement the wrong Hamiltonian, and a correct number can be unreproducible if its environment and parameters are missing. Trust requires both an executable chain and independent validation.
Use this page before publishing a notebook, script, numerical table, or generated figure. The site-wide Notebook Index owns the common metadata contract, Reproducibility Status owns status labels, and Validation Tests owns the test taxonomy. This page turns those principles into an end-to-end computational audit.
Canonical Boundaries
Section titled “Canonical Boundaries”| Question | Canonical home | Role here |
|---|---|---|
| What metadata must every notebook record? | Notebook Index | apply the contract to one release |
| Which environment details and dependencies matter? | Environments | verify that the environment is reconstructible |
| Which status label may be assigned? | Reproducibility Status | collect the evidence needed for that decision |
| What kinds of executable checks exist? | Validation Tests | choose checks with independent failure modes |
| How should refinement and uncertainty be quantified? | Convergence Tests and Error Estimates | confirm that the reported observable is resolved |
| Is the physical approximation itself controlled? | Error-Estimate Checklist | keep model error separate from numerical error |
| Is one computational result ready to release? | this page | perform the end-to-end audit |
This checklist does not create a second status system or duplicate numerical-method derivations. It asks whether the evidence required by those canonical pages is present, connected, and current.
Operational Vocabulary
Section titled “Operational Vocabulary”Terminology varies across disciplines. This page follows the operational distinction used by the National Academies report on reproducibility and replicability:
- computational reproducibility: consistent results are obtained from the same input data, code, computational steps, and stated conditions;
- replicability: a new study addressing the same scientific question obtains consistent results using independently collected evidence or an independently constructed calculation;
- repeatability: a narrower engineering term often used for rerunning under nearly identical conditions.
For a deterministic quantum calculation, the first release target is computational reproducibility. An independent finite-difference check of a basis calculation moves toward replication because it changes the numerical representation and its failure modes. Merely rerunning the same executable twice tests less.
Three matching levels
Section titled “Three matching levels”State what agreement is required.
| Level | Requirement | Appropriate use |
|---|---|---|
| bitwise | output bytes or hashes are identical | deterministic archives on a fixed stack |
| numerical | declared quantities agree within justified tolerances | floating-point work across compatible platforms |
| scientific | conclusions and uncertainty statements agree | independent methods or environments |
Bitwise identity is strong provenance evidence, but it is neither necessary nor sufficient for scientific agreement. Parallel reductions, linear-algebra libraries, and hardware instructions can alter low-order bits without changing a resolved observable. Conversely, identical wrong code produces identical wrong bytes.
The Claim-to-Artifact Chain
Section titled “The Claim-to-Artifact Chain”Start from the scientific claim and work backward:
claim <- quoted value, table, or figure <- retained machine-readable output <- validated computation <- code plus parameters and inputs <- declared environment <- source revision, run record, and licenseEvery arrow must be inspectable. A figure that cannot be regenerated from retained data is detached from its evidence. A CSV with no parameter record is detached from the physical problem. A script with no validation target is executable, but not yet trustworthy.
Define the reproduction target
Section titled “Define the reproduction target”Before running anything, write one sentence of the form:
Running command in environment with inputs and parameters must regenerate outputs, pass named checks at stated tolerances, and support this claim.
This sentence prevents the audit from collapsing into the vague question “does the notebook run?” A notebook may execute successfully while producing the wrong branch, normalization, phase convention, or asymptotic regime.
Minimum Release Packet
Section titled “Minimum Release Packet”A reference-grade computational result should retain the following small, inspectable packet.
| Artifact | Required content | Why it matters |
|---|---|---|
| explanatory page | question, conventions, result, limitations, canonical links | states what the computation means |
| executable source | notebook, script, or small program | reconstructs the transformation |
| parameter record | physical and numerical controls | prevents hidden defaults |
| environment record | runtime, packages, versions, optional dependencies | reconstructs execution conditions |
| input record | local data, provenance, checksum, preprocessing | identifies what entered the calculation |
| retained outputs | compact CSV, JSON, or text summaries | preserves evidence independently of plotting software |
| validation report | checks, targets, tolerances, pass/fail results | distinguishes execution from validation |
| figure source | plotting or TikZ source tied to retained data | makes visual evidence regenerable |
| provenance record | date, source revision, command, runtime class | identifies the validated run |
| license information | code, data, and third-party terms | determines whether others may reproduce and reuse it |
Large opaque binary outputs are rarely the best release artifact. Retain the smallest lossless or scientifically sufficient data product from which the reported table or figure can be regenerated.
Release Checklist
Section titled “Release Checklist”The checkboxes below are ordered. A missing physics convention cannot be repaired by a more detailed package list, and a passing regression test cannot compensate for an unconverged observable.
1. Claim and scope
Section titled “1. Claim and scope”- The page states the exact quantity or qualitative conclusion supported by computation.
- The claim names its parameter regime and excludes regimes not tested.
- The canonical analytic page is linked instead of rederived in the notebook.
- Exploratory observations are labeled separately from validated results.
- The required matching level is declared: bitwise, numerical, or scientific.
- The number of printed digits is justified by the error estimate.
A claim such as “the method works” is too broad. Prefer “for , the lowest gap is stable under the retained basis refinements and differs from the one-loop estimate by the reported amount.”
2. Physics specification
Section titled “2. Physics specification”- The Hamiltonian or evolution equation is stated.
- Units or dimensionless scales are explicit.
- Initial, boundary, and asymptotic conditions are explicit.
- Basis ordering, Fourier sign, phase, branch, and normalization conventions are recorded where relevant.
- Symmetry sector, particle statistics, reduced mass, and gauge choice are named when they affect the answer.
- The reported observable is defined mathematically.
- Model assumptions and omitted physics are separated from numerical approximations.
For scattering, include incident-state normalization, current factors, partial-wave convention, matching radius, and phase unwrapping. For tunneling, specify whether the output is an amplitude, rate, transmission probability, or spectral splitting. For driven systems, record the time origin, pulse order, and effective-Hamiltonian branch.
3. Numerical specification
Section titled “3. Numerical specification”- The representation is named: grid, basis, finite elements, sparse matrix, trajectories, or another explicit form.
- Every independent cutoff is recorded.
- Domain size and boundary treatment are recorded separately from resolution.
- Solver, algorithm, and matrix structure are stated.
- Arithmetic precision and relevant summation or diagonalization choices are stated.
- Stopping criteria and solver tolerances are explicit.
- The refinement path is recorded, including jointly changed controls.
- Runtime and memory class are stated when they affect practical reproduction.
Do not write “high resolution” or “tight tolerance” without numbers. A basis size is not reproducible if the basis scale, ordering, symmetry projection, and matrix-element construction remain implicit.
4. Environment and dependencies
Section titled “4. Environment and dependencies”- Runtime and language versions are recorded.
- Direct package versions are recorded.
- A lockfile or constrained environment file is retained when the result supports a published figure or benchmark.
- Optional plotting, acceleration, and file-format packages are distinguished from required dependencies.
- Operating system, architecture, accelerator, or linear-algebra backend is recorded when results are sensitive to it.
- Paths are relative to the repository or supplied input directory.
- Ordinary execution does not depend on a contributor’s private files, credentials, or shell history.
- The run command works from a documented working directory.
Prefer the smallest dependency set that honestly implements the calculation. Fewer dependencies reduce installation ambiguity and long-term API drift, but replacing a tested domain library with opaque custom code does not improve reproducibility.
5. Inputs and provenance
Section titled “5. Inputs and provenance”- Every input file has a stable name and documented role.
- External data have a source, version or access date, license, and checksum when practical.
- Download and preprocessing steps are scripted or described exactly.
- Raw inputs are distinguishable from processed and generated files.
- Physical constants identify their source and unit convention.
- Network access is not silently required during the validated run.
- Generated output cannot overwrite irreplaceable source data without an explicit option.
A checksum establishes file identity, not scientific quality. It answers “is this the same input?” rather than “is this the right input?”
6. Randomness and sampling
Section titled “6. Randomness and sampling”- The random-number generator and library are named.
- Seeds or seed-generation rules are recorded.
- Independent streams are constructed deliberately for parallel work.
- Sample count, burn-in, thinning, and autocorrelation treatment are recorded when relevant.
- Statistical uncertainty is estimated from independent information, not inferred from a fixed seed.
- Multiple seeds or batches are used when the claim concerns an ensemble rather than one trajectory.
- Deterministic calculations state that no randomness is used.
A fixed seed reproduces one pseudorandom stream. It does not estimate variance, expose seed sensitivity, or remove Monte Carlo bias.
7. Validation and tolerances
Section titled “7. Validation and tolerances”- At least one structural check is executable.
- At least one exact limit, analytic benchmark, or manufactured solution is checked when available.
- Every relevant numerical control is refined.
- The observable, not only an internal residual, is tested.
- At least two checks have meaningfully different failure modes when practical.
- Tolerances are declared before inspecting the final pass/fail result.
- Tolerances are scale-aware and justified by truncation, roundoff, or statistical uncertainty.
- Validation fails visibly and returns a nonzero status for scripts used in automation.
- Known failure regions are retained rather than removed from the report.
For a scalar reference value, a common numerical criterion is
The absolute term controls comparisons near zero. The relative term scales comparisons away from zero. Neither should be copied from a package default without relating it to the physical question.
For a quantity expected to be zero by symmetry, use an explicit physical or numerical scale :
For stochastic estimates and with standard errors and , a compatibility diagnostic may use
with the confidence interpretation of stated. This test assumes the uncertainty estimates and independence assumptions are credible; it is not a universal substitute for distributional checks.
8. Outputs, tables, and figures
Section titled “8. Outputs, tables, and figures”- Quoted values appear in a retained machine-readable table.
- Columns have units or explicit dimensionless definitions.
- Output precision does not imply unsupported accuracy.
- Figures are generated from tracked source and retained data.
- Axes, scales, parameter values, and normalization are labeled.
- Logarithmic scales and transformed quantities are stated.
- Curves remain distinguishable without relying on color alone.
- Captions state what was computed and what validation supports it.
- Hand editing after export is absent or fully documented.
- Intermediate caches are either reproducibly generated or excluded from the evidence chain.
The final SVG or PNG is a presentation artifact. The scientific evidence is the linked combination of data, plotting source, parameters, and validation record.
9. Clean execution
Section titled “9. Clean execution”- The source runs from a fresh checkout or clean kernel.
- Notebook cells execute in documented order without hidden state.
- The retained command regenerates outputs in a temporary or declared directory.
- The validation summary is visible at the end of the run.
- The run does not rely on stale outputs already present in the repository.
- A second run is idempotent or documents why output changes.
- Expected warnings are recorded and unexpected warnings are investigated.
- The validated run date and source revision are recorded.
When practical, move existing generated outputs aside before the clean run. Otherwise a script that silently fails to regenerate a file can appear successful because an old artifact remains available.
10. License, citation, and maintenance
Section titled “10. License, citation, and maintenance”- Code has an explicit license or SPDX identifier.
- Data have reuse terms compatible with redistribution.
- Third-party code, data, and formulas are cited.
- Software versions are cited when their implementation materially affects the result.
- The page records a review date.
- Dependency changes, convention changes, or failed validations trigger status review.
- Rebaselining a regression value requires a documented scientific or technical reason.
- Superseded artifacts remain traceable or receive a migration note.
“Available online” is not a license. Without explicit reuse terms, another researcher may be able to inspect an artifact but not legally redistribute or adapt it.
A Machine-Readable Manifest
Section titled “A Machine-Readable Manifest”The explanatory page remains primary, but a small YAML manifest reduces ambiguity and makes audits easier to automate. The following example follows the site-wide metadata contract without defining a new status vocabulary.
schema_version: 1artifact: id: 'magnus-expansion-error' kind: 'python-script' source: 'public/notebooks/approximation-scattering-semiclassics/magnus-expansion-error.py' canonical_page: 'labs/magnus-expansion-error'purpose: claim: 'Cumulative Magnus errors recover the first four expected weak-drive powers.' matching_level: 'numerical'execution: working_directory: 'repository-root' command: 'python public/notebooks/approximation-scattering-semiclassics/magnus-expansion-error.py' expected_runtime_seconds: 1environment: python: '3.12.13' packages: numpy: '2.3.5' randomness: 'none'parameters: eta_min: 0.010 eta_max: 2.000 stroboscopic_eta: 0.16 periods: 200validation: - name: 'Magnus-4 weak-drive power' target: 5.0 tolerance: 0.03 result: 5.000917 status: 'pass' - name: 'maximum one-period unitarity defect' upper_bound: 2.0e-14 result: 1.12e-15 status: 'pass'outputs: - 'public/data/approximation-scattering-semiclassics/magnus-expansion-error-sweep.csv' - 'public/data/approximation-scattering-semiclassics/magnus-expansion-stroboscopic.csv' - 'public/data/approximation-scattering-semiclassics/magnus-effective-hamiltonian-branches.csv'provenance: last_run: '2026-07-16' source_revision: 'record-at-release'license: code: 'MIT'reproducibility_status: 'assign-using-canonical-status-page'The manifest is useful only if it is kept with the source revision it describes. A precise manifest attached to later, changed code can be more misleading than no manifest.
Reviewer Protocol
Section titled “Reviewer Protocol”The author and reviewer should not perform exactly the same check in exactly the same way.
Pass 1: documentary audit
Section titled “Pass 1: documentary audit”Read the page without running code. Confirm that the claim, model, conventions, parameter regime, outputs, validation targets, and limitations can be understood from retained text and metadata.
Pass 2: clean reconstruction
Section titled “Pass 2: clean reconstruction”Create or select the declared environment, use a clean checkout, run the stated command, and verify that expected outputs are regenerated. Record platform differences and warnings.
Pass 3: scientific validation
Section titled “Pass 3: scientific validation”Inspect the validation logic. Ask whether each check could pass while the reported claim remained wrong. Add a check with a different failure mode when the answer is yes.
Examples:
- pair norm conservation with a phase-sensitive observable;
- pair an eigensolver residual with basis and domain refinement;
- pair a seeded regression with independent-seed uncertainty;
- pair a WKB slope with a prefactor-sensitive ratio;
- pair a matrix identity with an independent discretization.
Pass 4: artifact trace
Section titled “Pass 4: artifact trace”Choose at least one quoted number and one figure. Trace each backward through retained data and code to the parameter and environment record. Regenerate the figure when practical.
Pass 5: status decision
Section titled “Pass 5: status decision”Assign or update the label using Reproducibility Status. A warning must state whether it affects the scientific conclusion. Do not loosen a tolerance silently to recover a preferred label.
Worked Audit: Magnus Benchmark
Section titled “Worked Audit: Magnus Benchmark”The Magnus Expansion Error laboratory provides a compact deterministic example.
| Audit item | Retained evidence | Residual limitation |
|---|---|---|
| analytic target | exact product of two Pauli exponentials | step drive is a deliberately small model |
| source | NumPy-only Python program with MIT SPDX header | optional plotting adds Matplotlib |
| parameters | fixed 18-point pulse-scale sweep and two named diagnostic couplings | no adaptive exploration is claimed |
| environment | Python and NumPy versions printed by the run | low-order bits may vary with numerical libraries |
| randomness | none | deterministic execution does not prove correctness |
| validation | weak-drive slopes, unitarity, pulse reversal, commuting control, branch reconstruction | tests a two-dimensional Hilbert space only |
| retained data | three CSV files | tables summarize rather than archive every intermediate matrix |
| figures | two SVGs generated from CSV tables with retained TikZ sources | visual agreement is not itself a pass criterion |
| long-time check | powers through 200 periods | no claim is made beyond that horizon |
The strongest part of this packet is not that every check passes. It is that the checks answer different questions and the limitations remain visible.
Release Decision
Section titled “Release Decision”Do not release a computational claim as verified when any critical item below fails.
| Critical gate | Fail condition | Required response |
|---|---|---|
| identifiable claim | output has no precise scientific interpretation | narrow and rewrite the claim |
| reconstructible model | Hamiltonian, units, conditions, or conventions are missing | document before rerunning |
| executable source | clean run cannot be started | repair environment or mark broken |
| resolved observable | relevant refinements are absent or numerical error is too large | extend convergence study or weaken claim |
| independent validation | all checks share the same likely failure | add an exact limit, symmetry, or alternative method |
| traceable output | table or figure cannot be regenerated | restore data and generation source |
| legal reuse | code or data terms are absent or incompatible | clarify licensing before redistribution |
| current provenance | run predates material code, dependency, or convention changes | rerun and review status |
Noncritical differences, such as platform-dependent low-order bits inside a much looser scientific tolerance, may support reproduced_with_warnings when explained. The canonical status page owns that label and its evidence requirements.
Maintenance Triggers
Section titled “Maintenance Triggers”Reopen the audit when any of the following changes:
- the Hamiltonian, convention, or canonical formula;
- a dependency major version or numerical backend;
- input data, preprocessing, or constants;
- a solver, tolerance, grid, basis, or random-number generator;
- the plotting path or retained output schema;
- a validation threshold or regression reference;
- the page’s scientific claim;
- the license or availability of a dependency or dataset.
Do not update a regression value merely because the new run differs. First classify the change as a bug fix, convention change, algorithmic improvement, dependency drift, expected platform variation, or actual regression. Retain the reason with the new baseline.
Common Failure Modes
Section titled “Common Failure Modes”- Treating a successfully executed notebook as a validated result.
- Recording a seed but no sampling uncertainty.
- Pinning every transitive dependency while omitting the Hamiltonian convention.
- Reporting an eigensolver residual as if it included basis and domain error.
- Validating a time integrator only through norm conservation.
- Comparing relative error to a reference value that crosses zero.
- Regenerating a figure from edited plotting data rather than retained raw output.
- Depending on cells executed out of order or variables from a previous session.
- Using an absolute local path or undocumented network download.
- Keeping only a screenshot of a table or plot.
- Silently changing tolerances until a test passes.
- Assuming a public repository grants a license to reuse code or data.
- Declaring exact cross-platform byte identity when the scientific claim needs only tolerance-level agreement.
- Deleting failed parameter points and thereby hiding the method’s validity boundary.
Exercises
Section titled “Exercises”1. Relative tolerance near zero
Section titled “1. Relative tolerance near zero”A symmetry-forbidden matrix element has reference value . A reproduced run gives . Why is a relative tolerance alone misleading, and how would an absolute scale improve the test?
Solution
The relative discrepancy is
which looks catastrophic even though both values may be negligible on the physical scale. If the natural matrix-element scale is , an absolute criterion such as
classifies the reproduced value correctly as a small symmetry residual. The scale and threshold must be justified by the computation; choosing them after seeing the result would invalidate the audit.
2. Fixed seed versus uncertainty
Section titled “2. Fixed seed versus uncertainty”A Monte Carlo notebook reproduces the same energy to every printed digit when rerun with seed 17. Does this establish a reproducible uncertainty estimate?
Solution
No. It establishes repeatability of one pseudorandom stream. It does not measure estimator variance, autocorrelation, burn-in bias, time-step bias, or sensitivity to the seed. The release should retain the seed for regression and also use independent seeds or batches to estimate statistical uncertainty. Systematic biases require separate checks.
3. Residual without refinement
Section titled “3. Residual without refinement”An eigenvalue calculation reports residual in a basis of size , but the energy changes by at . Which claim is supported?
Solution
The small residual shows that the finite matrix eigenproblem was solved accurately. It does not show that the finite basis represents the continuum Hamiltonian accurately. The basis change is the relevant evidence for the physical energy and dominates the algebraic residual. The energy is not resolved below that scale until basis refinement stabilizes.
4. Cross-platform output
Section titled “4. Cross-platform output”Two platforms produce different CSV hashes, but every reported observable agrees within , while the justified release tolerance is . How should this be reported?
Solution
Bitwise reproduction failed, but numerical reproduction passed comfortably. Record the platforms, library versions, hash difference, maximum observable discrepancy, and tolerance. If no scientific conclusion changes, the difference is a provenance warning rather than a validation failure. The assigned status should follow the canonical status policy.
5. Trace a figure
Section titled “5. Trace a figure”A figure has a plotting script and labeled axes, but the plotted CSV is absent and the script downloads a changing URL without a version. Identify the release blockers.
Solution
The figure is detached from stable input evidence. A reviewer cannot know which data version produced it, verify the download, or regenerate the same output later. Retain or archive the exact input when licensing permits, record its source version and checksum, script preprocessing, and connect the plotting source to that retained file. Labeled axes do not repair missing provenance.
Cross-Links
Section titled “Cross-Links”- Computational Notebooks defines the notebook family and its validation contract.
- Error-Estimate Checklist audits physical approximation error separately from computational reproducibility.
- Notebook Index defines the site-wide metadata contract.
- Environments owns dependency, path, and platform policy.
- Reproducibility Status defines labels and review triggers.
- Validation Tests gives the executable test taxonomy.
- Analytic Benchmarks and Numerical Benchmarks provide reusable targets.
- Convergence Tests and Error Estimates develop the quantitative refinement workflow.
- Citation Style covers software, data, and version citation.
References
Section titled “References”- National Academies of Sciences, Engineering, and Medicine, Reproducibility and Replicability in Science, National Academies Press (2019), doi:10.17226/25303.
- G. K. Sandve, A. Nekrutenko, J. Taylor, and E. Hovig, “Ten Simple Rules for Reproducible Computational Research,” PLOS Computational Biology 9, e1003285 (2013), doi:10.1371/journal.pcbi.1003285.
- G. Wilson et al., “Good Enough Practices in Scientific Computing,” PLOS Computational Biology 13, e1005510 (2017), doi:10.1371/journal.pcbi.1005510.
- M. D. Wilkinson et al., “The FAIR Guiding Principles for scientific data management and stewardship,” Scientific Data 3, 160018 (2016), doi:10.1038/sdata.2016.18.
- R. D. Peng, “Reproducible research in computational science,” Science 334, 1226–1227 (2011), doi:10.1126/science.1213847.
- V. Stodden et al., “Enhancing reproducibility for computational methods,” Science 354, 1240–1241 (2016), doi:10.1126/science.aah6168.
- The Turing Way Community, The Turing Way: A Handbook for Reproducible, Ethical and Collaborative Data Science, doi:10.5281/zenodo.3233853.