Skip to content

Validation Tests

Validation tests are executable checks that decide whether a computational artifact can be trusted for its stated purpose. They are narrower than general benchmarks: a benchmark defines the reference problem, while a validation test says exactly what the notebook or script must verify before its output is cited.

A generated plot, table, or notebook result should not be treated as evidence unless the relevant validation tests pass or the page explicitly records the failure.

Every reviewed computational artifact should state:

  • the physical model and canonical page being tested;
  • the formula, theorem, or benchmark target;
  • the numerical method and approximation;
  • the environment or dependency recipe;
  • the input parameters, units, and random seeds when relevant;
  • the test quantity;
  • the tolerance and why it is appropriate;
  • the observed result and status.

The tolerance belongs to the test. It should not be chosen after seeing the result.

Test familyWhat it checksExample pass condition
Smoke testThe notebook or script runs from a clean environmentno missing imports, files, or out-of-order state
Shape and domain testArrays, matrices, bases, and grids have the intended dimensionsHamiltonian is square on the declared basis
Hermiticity testClosed-system Hamiltonians are Hermitian in the chosen inner product∥H−H†∥\lVert H-H^\dagger\rVert below tolerance
Normalization testStates use the correct discrete or continuous measure∣⟨ψ∣ψ⟩−1∣\lvert\langle\psi\vert\psi\rangle-1\rvert below tolerance
Conservation testDynamics preserve the invariant expected from the modelnorm, trace, energy, or current drift below tolerance
Analytic-target testThe result matches an exact formula in a controlled caselow oscillator energies match En=ℏω(n+1/2)E_n=\hbar\omega(n+1/2)
Convergence testThe result stabilizes under refinementerror decreases when grid spacing, time step, or basis cutoff is refined
Regression testA validated output has not drifted unexpectedlyselected eigenvalues or summary diagnostics match stored references
Stochastic testSampling errors are quantifieduncertainty decreases with sample size and multiple seeds agree

Most artifacts need more than one family. For example, a split-operator wave-packet notebook should pass smoke, normalization, unitarity, Fourier-grid convention, convergence, and regression tests.

For a finite-dimensional closed Hamiltonian calculation, check that the Hamiltonian is Hermitian and that the evolution is unitary:

H†=H,U†U≈I.H^\dagger=H, \qquad U^\dagger U\approx I.

For wavefunction propagation, monitor both norm and physically meaningful observables. Norm conservation alone does not validate phases, dispersion, tunneling probabilities, or boundary effects.

For bound-state calculations, check:

  • boundary conditions or endpoint regularity;
  • orthonormality under the correct measure;
  • residuals ∥Hψn−Enψn∥\lVert H\psi_n-E_n\psi_n\rVert;
  • convergence of low-lying eigenvalues under refinement;
  • symmetry labels such as parity or angular momentum when applicable.

Do not validate a spectrum only by visual agreement with expected node shapes. A visually plausible eigenfunction can still use the wrong normalization, boundary condition, or grid operator.

For conservative one-dimensional scattering, check current conservation:

R+T≈1.R+T\approx1.

The test must use current ratios, not merely squared amplitudes, when the asymptotic wavenumbers differ. For wave-packet scattering, also check packet bandwidth, reflection from numerical boundaries, and agreement with stationary formulas only in the appropriate narrow-packet limit.

For density-operator calculations, check:

  • Hermiticity of ρ\rho;
  • trace preservation;
  • positivity within numerical tolerance;
  • complete positivity when a Kraus map or Lindblad generator is claimed;
  • entropy or purity behavior only when the model assumptions justify it.

A trace-preserving evolution can still be unphysical if it creates negative eigenvalues beyond tolerance.

Quantum circuits and finite-dimensional gates

Section titled “Quantum circuits and finite-dimensional gates”

For circuit or gate simulations, check:

  • state normalization;
  • unitary matrices satisfy U†U≈IU^\dagger U\approx I;
  • tensor-factor ordering matches the convention page;
  • known identities hold for small circuits;
  • global phase is not mistaken for an observable error.

Use Quantum Gates and Tensor Product Ordering when translating between packages.

A compact validation record should be attached to any notebook or generated table that is cited by a page.

FieldMeaning
artifactnotebook, script, generated data file, or figure
canonical_targetconcept, formula, benchmark, or model page
test_idstable identifier for the check
methodnumerical or symbolic method being exercised
parametersphysical and numerical parameters
tolerancepredeclared pass threshold
observedmeasured diagnostic
statuspassed, warning, failed, skipped, or conceptual
last_rundate and environment record
notesfailure diagnosis, limitations, or follow-up

The record may live in notebook metadata, a small report file, or a page table. It should be easy to find from the artifact and from the page that cites the artifact.

Use status labels conservatively:

StatusMeaning
passedThe stated test passed with the recorded environment and tolerance
warningThe test passed but has a limitation that affects interpretation
failedThe result should not be used as evidence until repaired
skippedThe test is applicable but was not run; state why
conceptualThe artifact is explanatory and not intended as numerical evidence

Do not convert a failure into a warning just because the plotted output looks reasonable.

When a notebook family is large enough to be run automatically, the minimal continuous-validation layer should include:

  1. import and environment smoke tests;
  2. execution of lightweight notebooks or stripped validation scripts;
  3. unit tests for reusable numerical helpers;
  4. comparison with analytic or stored benchmark diagnostics;
  5. report generation that records failures without hiding skipped checks.

Heavy notebooks may run on a slower schedule. If a result is too expensive for ordinary validation, it should still have a smaller representative test.

  • Passing norm conservation while the time step gives the wrong phase.
  • Matching one eigenvalue while the basis ordering is wrong.
  • Passing a regression test after updating the stored answer without diagnosis.
  • Using absolute error where a relative or scale-aware error is required.
  • Testing a dimensionless version while citing a dimensional formula without checking scales.
  • Running from a notebook kernel that already contains hidden state.
  • Treating one random seed as a statistical uncertainty estimate.
  • J. M. Thijssen, Computational Physics, 2nd ed., Cambridge University Press, 2007.
  • R. J. LeVeque, Finite Difference Methods for Ordinary and Partial Differential Equations, SIAM, 2007.
  • L. N. Trefethen and D. Bau III, Numerical Linear Algebra, SIAM, 1997.
  • Y. Saad, Numerical Methods for Large Eigenvalue Problems, 2nd ed., SIAM, 2011.
  • The Turing Way Community, The Turing Way: A Handbook for Reproducible, Ethical and Collaborative Data Science.
  1. A wave-packet propagation notebook conserves norm to machine precision but disagrees with the analytic spreading width. Should it be marked passed?
Solution

Not for the propagation claim. Norm conservation is necessary, but the spreading width tests the kinetic phase and Fourier-grid convention. The notebook may pass a norm test while failing a dispersion test. The validation record should mark the norm test as passed and the spreading-width test as failed or under investigation.

  1. A Monte Carlo notebook reproduces a reference value for one fixed seed. What validation is still missing?
Solution

A fixed seed is useful as a regression test, but it does not estimate statistical uncertainty. The notebook should also run multiple seeds or batches, state the estimator, report an uncertainty measure, and check that the uncertainty decreases with sample size in the expected way.