Floating-Point Arithmetic
Floating-point arithmetic represents real and complex numbers with finitely many bits. It is the arithmetic behind almost every numerical wavefunction, Hamiltonian matrix, expectation value, time-evolution calculation, and plot.
The essential point is simple: floating-point numbers approximate real numbers, and each elementary operation can introduce a small rounding error. Most calculations tolerate this well. Some calculations amplify the error so strongly that a visually reasonable answer is mathematically unreliable.
This page gives the numerical hygiene needed before trusting computational quantum-mechanics results. For the broader question of how to combine roundoff with truncation, solver, and statistical uncertainty, see Error Estimates.
Why Quantum Mechanics Needs It
Section titled “Why Quantum Mechanics Needs It”Quantum calculations often combine small differences with large cancellations:
- normalizing a wavefunction from a quadrature or grid sum;
- subtracting nearby energy eigenvalues to estimate a splitting;
- computing a variance as ;
- forming finite-difference derivatives from nearby function values;
- checking orthogonality of nearly degenerate eigenvectors;
- evolving phases such as over long times;
- summing many oscillatory amplitudes.
In exact mathematics these are ordinary algebraic operations. On a computer, they are operations in a finite set of representable numbers. A mature numerical result should therefore report not only a formula, but also the scale of the error being controlled.
Floating-Point Numbers
Section titled “Floating-Point Numbers”A normalized binary floating-point number has the schematic form
where is a finite-precision significand and is an integer exponent in a finite range. The exponent allows very large and very small magnitudes; the finite significand means that only finitely many numbers are representable between any two powers of two.
The spacing is not uniform. Near , adjacent double-precision numbers are separated by roughly . Near a number of size , the spacing is roughly scaled by .
This relative-spacing behavior is usually good for physics: a value near and a value near can each be represented with a comparable number of significant binary digits. But it also means that adding a tiny number to a huge number may do nothing if the tiny number is below the local spacing.
Machine Epsilon and Unit Roundoff
Section titled “Machine Epsilon and Unit Roundoff”Two closely related constants are often mixed together.
For IEEE binary64 arithmetic, commonly called double precision, the spacing between and the next larger representable number is
For rounding to nearest, the maximum relative error from rounding one normal real number to the nearest representable value is approximately
Many texts call one or the other quantity “machine epsilon.” When reading code, documentation, or papers, check the convention. For error estimates, the important idea is that double precision gives roughly sixteen decimal digits before any amplification by the algorithm or problem.
The Basic Rounding Model
Section titled “The Basic Rounding Model”A standard local model for one arithmetic operation is
where is one of , , , or division, and denotes the computed floating-point result.
This model is not a license to ignore details. It assumes normal results and excludes overflow, severe underflow, and exceptional cases. Still, it is the right first diagnostic: each operation is nearly correct, but a long algorithm can amplify or accumulate those small errors.
Absolute error measures the size of the difference:
Relative error measures the difference compared with the scale of the target:
Relative error is usually more meaningful, but it becomes delicate when the true quantity is zero or extremely small. Quantum calculations often care about small residuals, small gaps, and small probabilities, so one must choose tolerances with scale in mind.
Catastrophic Cancellation
Section titled “Catastrophic Cancellation”Subtraction is not automatically unstable. The danger appears when two nearly equal rounded quantities are subtracted.
Suppose the target is
with . If the inputs have small relative errors, the relative error in can be amplified roughly by
When is tiny compared with and , the condition factor is large. The leading digits cancel, and the remaining digits may mostly be inherited rounding error.
This is why the formula
can be numerically poor for a state that is almost an energy eigenstate. The variance is small because two large quantities nearly cancel. For a normalized state, the equivalent expression
is often a better diagnostic because it computes the small quantity as a norm of a residual.
Summation and Inner Products
Section titled “Summation and Inner Products”Sums are not associative in floating-point arithmetic:
in general. The order of summation can matter, especially when terms have mixed signs, very different magnitudes, or rapidly oscillating phases.
This affects:
- wavefunction normalization sums;
- numerical inner products;
- expectation values ;
- path-like sums of oscillatory amplitudes;
- Monte Carlo estimates with cancellations.
Useful habits include summing from small to large magnitude when possible, using pairwise summation or compensated summation, and checking whether the answer changes under a harmless reordering. For complex vectors, separately tracking real and imaginary parts may expose cancellation that a final magnitude hides.
Scale and Nondimensionalization
Section titled “Scale and Nondimensionalization”Floating-point arithmetic rewards sensible units. A Hamiltonian matrix whose entries range from to is usually harder to treat accurately than an equivalent nondimensional form with entries of order .
Before a numerical calculation, ask:
- What is the natural length, energy, and time scale?
- Are the variables dimensionless?
- Are two large terms being subtracted to reveal a small physical effect?
- Is the desired answer many orders of magnitude smaller than intermediate quantities?
- Does the tolerance use the same units and scale as the reported residual?
For example, harmonic-oscillator calculations are often cleaner in units where . The physics is unchanged, but the numerical matrix entries and expected eigenvalues have a natural scale.
Overflow, Underflow, and Tiny Probabilities
Section titled “Overflow, Underflow, and Tiny Probabilities”A floating-point type has a finite exponent range. If a result is too large, it overflows. If it is too small, it underflows to a subnormal number or to zero.
This matters in quantum mechanics whenever exponentials appear:
Real decays can underflow long before the exact mathematical value is zero. Oscillatory phases do not underflow, but very large phase arguments can lose meaningful low-order bits before argument reduction. Numerically stable codes often rescale wavefunctions, subtract a reference energy, work with logarithms, or factor out a known phase.
If the physical result is an exponentially small tunneling probability, the question is not merely whether the final number prints as nonzero. The calculation must preserve the exponent and prefactor at the intended accuracy.
Discretization Versus Roundoff
Section titled “Discretization Versus Roundoff”Floating-point error is only one source of numerical error. In a finite-difference approximation, shrinking the grid spacing often reduces truncation error but increases roundoff amplification. The broader continuum-to-finite-model step is Discretization.
For a centered second derivative,
the numerator subtracts nearby values when is small. A schematic error balance is
The first term is discretization error; the second term is roundoff amplification. Making smaller forever is not a convergence strategy. A good calculation checks a range of grid spacings and looks for a stable window.
Hermiticity, Symmetry, and Physical Constraints
Section titled “Hermiticity, Symmetry, and Physical Constraints”Floating-point operations can break exact algebraic identities at the last few digits. A matrix intended to be Hermitian may satisfy
by a tiny amount after assembly.
If Hermiticity is exact in the mathematical model, it is reasonable to enforce it numerically by replacing with
But this should be done as a documented projection onto an exact constraint, not as a way to hide a modeling error. If a magnetic field, absorbing boundary, effective non-Hermitian Hamiltonian, or open-system approximation genuinely makes the operator non-Hermitian, symmetrizing would change the physics.
The same caution applies to normalization, unitarity, positivity, and trace preservation. Small repairs can be useful, but they should be paired with residual checks that reveal how large the repair was.
Eigenvalue Splittings and Near Degeneracy
Section titled “Eigenvalue Splittings and Near Degeneracy”Eigenvalues may be computed accurately while individual eigenvectors inside a nearly degenerate subspace are unstable. A tiny perturbation can rotate the basis inside that subspace without changing the physically meaningful subspace much.
When a calculation reports a small splitting
compare with:
- the residual norms of the computed eigenpairs;
- the sensitivity of the result under grid, basis, or tolerance changes;
- the scale of the Hamiltonian entries;
- any exact or approximate symmetry that predicts degeneracy.
A small printed difference between two large computed energies is not automatically a physical splitting. It may be roundoff, truncation, symmetry breaking, or a real effect. The calculation must distinguish these possibilities.
Reproducibility and Order of Operations
Section titled “Reproducibility and Order of Operations”Floating-point results can depend on evaluation order. Parallel reductions, vectorized code, different processor instructions, and different library versions can change the last bits. Usually this is harmless. It becomes important when a conclusion rests on a small difference, a threshold comparison, or an apparent symmetry violation.
For reproducible research, record:
- the discretization or basis size;
- the precision used;
- the algorithm and library where relevant;
- tolerances and stopping criteria;
- residual checks;
- benchmark comparisons.
The goal is not to make every last bit identical across machines. The goal is to make the physical conclusion stable under numerically reasonable changes.
Practical Checklist
Section titled “Practical Checklist”Before trusting a numerical quantum result, ask:
- Are the variables scaled so typical numbers are not extreme?
- Is the result a small difference of large quantities?
- Are residuals reported in a norm with a meaningful scale?
- Is a convergence check separating discretization error from roundoff?
- Does changing precision or summation order change the conclusion?
- Are exact constraints such as Hermiticity or normalization preserved or checked?
- Is the requested tolerance realistic compared with double precision and problem conditioning?
For eigenvalue calculations, continue with Matrix Diagonalization. For the language of norms and residuals, see Norms and Metrics.
Common Mistakes
Section titled “Common Mistakes”- Treating double precision as exactly sixteen correct digits in every final answer.
- Comparing floating-point numbers for exact equality after a calculation.
- Using an absolute tolerance without considering the scale and units of the quantity.
- Making a grid spacing smaller until roundoff dominates.
- Computing a small variance or energy splitting by subtracting two large noisy numbers.
- Assuming a residual of is meaningful without knowing the matrix norm or problem scale.
- Mistaking loss of orthogonality in a numerical basis for a physical effect.
- Reporting a numerical result without enough information to reproduce the calculation.
Cross-Links
Section titled “Cross-Links”- Norms and Metrics
- Inner Products
- Matrices as Linear Maps
- Singular Value Decomposition
- Conditioning and Stability
- Error Estimates
- Discretization
- Finite Difference Methods
- Matrix Diagonalization
- Asymptotic Analysis
- Small Parameters and Error Estimates
References
Section titled “References”- N. J. Higham, Accuracy and Stability of Numerical Algorithms, 2nd ed., SIAM, 2002.
- D. Goldberg, “What every computer scientist should know about floating-point arithmetic,” ACM Computing Surveys 23, 5-48, 1991.
- IEEE Computer Society, IEEE Standard for Floating-Point Arithmetic, IEEE Std 754-2019, 2019.
- L. N. Trefethen and D. Bau, Numerical Linear Algebra, SIAM, 1997.
- W. H. Press, S. A. Teukolsky, W. T. Vetterling, and B. P. Flannery, Numerical Recipes: The Art of Scientific Computing, 3rd ed., Cambridge University Press, 2007.
Exercises
Section titled “Exercises”- In double precision, distinguish the spacing between and the next larger floating-point number from the unit roundoff for rounding to nearest.
Solution
For IEEE binary64 arithmetic, the spacing between and the next larger representable number is
With rounding to nearest, a rounded number is at most half a spacing away from the exact value near , so the unit roundoff is
Different authors call one or the other quantity “machine epsilon,” so the convention should be checked.
- Explain why computing can be unreliable for a state close to an energy eigenstate.
Solution
For a state close to an eigenstate, and are nearly equal. Their difference is small compared with either term, so subtraction can amplify the relative error in the inputs. A more stable diagnostic is to compute the residual norm
for a normalized state, because it forms the small quantity directly as a norm.
- Why is decreasing a finite-difference grid spacing not always an improvement?
Solution
For many centered finite differences, the truncation error decreases as a power of , but roundoff can be amplified by division by powers of . A schematic second-derivative balance is
At first, decreasing reduces the term. Eventually, the term grows and the result gets worse. A convergence study should look for a stable window rather than always taking the smallest possible spacing.
- Give one reason summing complex amplitudes can be more delicate than summing probabilities.
Solution
Complex amplitudes can cancel through phase. If many terms have similar magnitudes but different phases, the final result can be much smaller than the sum of magnitudes. Then small rounding errors in individual terms or in the summation order can become visible in the relative error of the final amplitude. Probabilities are nonnegative in ordinary sums, so this particular phase-cancellation mechanism is absent.
- A computed Hamiltonian should be Hermitian, but the assembled matrix has tiny violations of . When is symmetrizing reasonable, and what should still be checked?
Solution
Symmetrizing by replacing with is reasonable when Hermiticity is an exact property of the mathematical model and the violation comes only from rounding or assembly order. The size of the correction should still be reported or checked. If the non-Hermitian part comes from the physical model, such as an absorbing boundary or effective decay term, symmetrizing would change the problem.