Skip to content

Benchmark Problems

Benchmark problems are controlled calculations that decide whether a numerical method, notebook, or package workflow is trustworthy for a specific task. A benchmark is not merely an example with a known answer. It has a target, a tolerance, a refinement path, and a failure diagnosis.

The general numerical-method discussion lives in Benchmark Problems. This page defines the Reference taxonomy used for notebooks and package-dependent artifacts.

Use benchmark categories to say what a computation actually tests:

  • analytic eigenvalue benchmarks,
  • normalization benchmarks,
  • Hermiticity and commutator benchmarks,
  • unitarity and time-evolution benchmarks,
  • scattering current-conservation benchmarks,
  • open-system trace-preservation and positivity benchmarks,
  • quantum circuit identity benchmarks,
  • many-body exact-diagonalization benchmarks,
  • convergence and finite-size benchmarks,
  • regression benchmarks for previously validated artifacts.

A single notebook may need several categories. For example, a wave-packet scattering notebook should check norm conservation, current accounting, boundary effects, and agreement with a stationary transmission formula in the narrow-packet limit.

Every benchmark report should include:

  • benchmark ID,
  • physical model and canonical page,
  • numerical method,
  • environment,
  • parameters and units,
  • reference result,
  • pass criterion,
  • observed result,
  • interpretation,
  • status label.

The interpretation line should say what has been tested and what has not. Passing a low-energy oscillator spectrum does not validate high-energy cutoff behavior or arbitrary time evolution.

Use short IDs with a domain prefix:

PrefixDomain
WMCSwave mechanics and canonical systems
DYNdynamics and formulations
SCATscattering and semiclassics
OPENdensity matrices and open systems
QIquantum information
MBmany-body and statistical quantum mechanics
AMOatoms, molecules, and light
QMTRquantum matter

IDs should stay stable once cited by a notebook or page.

Declare pass criteria before running the benchmark. Examples:

  • first five eigenvalues converge at the expected order,
  • discrete states are orthonormal under the correct quadrature weight,
  • norm drift stays below a stated threshold over a stated time window,
  • a Lindblad evolution preserves trace and positivity within tolerance,
  • R+T=1R+T=1 for conservative one-dimensional scattering,
  • a circuit simulation preserves state norm and matches a known unitary identity,
  • finite-size scaling is monotone or follows the stated asymptotic law.

Tolerances should be tied to the method, scale, and intended use. A visually plausible plot is not an acceptance criterion.

A benchmark can be used as reference evidence only after:

  1. the analytic or trusted reference result is cited,
  2. the notebook or script records the environment,
  3. at least one refinement or independent check is present,
  4. the status is reproduced or reproduced_with_warnings,
  5. failure modes are described.
  • J. M. Thijssen, Computational Physics, 2nd ed., Cambridge University Press, 2007.
  • R. J. LeVeque, Finite Difference Methods for Ordinary and Partial Differential Equations, SIAM, 2007.
  • D. J. Tannor, Introduction to Quantum Mechanics: A Time-Dependent Perspective, University Science Books, 2007.
  • L. N. Trefethen and D. Bau III, Numerical Linear Algebra, SIAM, 1997.