Variational and Bound Methods
Variational methods replace an intractable spectral problem by an optimization problem over a controlled set of states. Their central promise is unusually concrete: for a Hamiltonian bounded from below, the energy expectation of every admissible trial state lies at or above the bottom of the spectrum. A well-chosen trial state therefore gives both an approximation and a rigorous one-sided bound.
The promise has limits. The inequality may be exact while the chosen ansatz is poor, the optimizer stops at the wrong point, matrix elements are evaluated inaccurately, or an apparently excellent energy hides a deficient wavefunction. A trustworthy variational calculation keeps those logically distinct issues visible.
This chapter develops that ledger. The proof of the ground-state bound has its canonical home in the Variational Principle, and finite-dimensional subspace calculations belong to the Rayleigh–Ritz Method. This overview explains how the pieces fit together, what can be certified, and what must still be checked.
The Central Statement
Section titled “The Central Statement”Let be self-adjoint and bounded below. For a nonzero state on which the energy quadratic form is finite, define the Rayleigh quotient
If
is the bottom of the spectrum, then
When is an isolated eigenvalue, it is the ground-state energy. If the bottom of the spectrum is not an eigenvalue, the inequality still holds, but no normalizable state need attain equality. Most of this chapter concerns bound-state problems for which a normalizable ground state exists.
The quotient is unchanged by nonzero rescaling:
Normalization is therefore convenient rather than essential. Admissibility is essential: the trial state must obey the problem’s boundary and regularity conditions and must have finite energy.
Operator domains and form domains
Section titled “Operator domains and form domains”For an unbounded Hamiltonian, writing can conceal a domain issue. The clean mathematical statement uses the closed quadratic form associated with . Its form domain can be larger than the operator domain of .
For familiar Schrödinger Hamiltonians, this distinction means that a trial wavefunction need not possess every derivative required for to exist as a Hilbert-space vector, but its kinetic and potential quadratic forms must still be finite. Piecewise-smooth basis functions can therefore be admissible even when applying the differential operator pointwise would be awkward. The precise form domain depends on the potential and boundary conditions.
Why a Minimum Knows about the Ground State
Section titled “Why a Minimum Knows about the Ground State”In the simplest discrete, nondegenerate case, expand a normalized trial state in exact energy eigenstates:
Its energy is
Every term after is nonnegative. Excited-state contamination raises the expectation value, and equality occurs only when the state lies entirely in the ground eigenspace. The Variational Principle gives the full proof and its assumptions.
This argument is more than an algebraic trick. It identifies what energy minimization is doing physically: the objective penalizes components with large excitation energy. It does not penalize all state errors equally. Components just above the ground state are energetically cheap, which becomes important when the gap is small.
A parameterized trial family produces an energy landscape . Optimization removes parameter error within that family, but the best value can remain above because the exact ground state is not contained in the ansatz.
Where the Approximation Enters
Section titled “Where the Approximation Enters”The variational inequality itself is not approximate. Approximation enters when the full form domain is replaced by a tractable trial set . Define
Then
A numerical calculation usually adds two more layers. If is the best value actually found and is the reported, numerically evaluated value, then the error ledger is
The first two terms are nonnegative when the definitions are exact. The evaluation term need not be. Quadrature bias, stochastic error, basis-conditioning error, Hamiltonian truncation, or measurement noise can make a reported number fall below the true energy expectation. The upper-bound guarantee applies to the exact Rayleigh quotient of an admissible state for the stated Hamiltonian, not automatically to every number emitted by an optimizer.
Four objects that should not be conflated
Section titled “Four objects that should not be conflated”| Object | Meaning | Typical source of error |
|---|---|---|
| Hamiltonian | The model whose spectrum is being bounded | Model or effective-theory error |
| Trial set | States the calculation is allowed to explore | Ansatz or basis incompleteness |
| Objective evaluator | Procedure for computing | Quadrature, sampling, truncation, or noise |
| Optimizer | Procedure for searching | Local minima, flat directions, poor conditioning |
A low variational energy for one Hamiltonian does not bound the ground-state energy of a different Hamiltonian. This matters when pseudopotentials, finite boxes, basis cutoffs, or effective interactions are introduced.
Trial Sets and Ansätze
Section titled “Trial Sets and Ansätze”An ansatz is a deliberately restricted representation of a state. Its design should encode as much exact structure as possible before optimization begins. The detailed craft belongs to Trial Wavefunctions; the main families are summarized here.
Nonlinear parameterized states
Section titled “Nonlinear parameterized states”A finite set of parameters defines
Widths, screening constants, correlation lengths, orbital exponents, and circuit angles are common examples. The objective
is generally nonlinear and can have several stationary points. A stationary point is not necessarily the global minimum over the trial family.
Energy gradients, projected tangents, parameter metrics, Hessians, constraints, and optimizer diagnostics are developed in Variational Parameters.
Linear trial subspaces
Section titled “Linear trial subspaces”Choose basis states and vary their coefficients:
With
stationarity gives the generalized eigenvalue problem
Solving this matrix problem exactly performs the energy minimization exactly inside the chosen span. Basis truncation remains the approximation. This is the Rayleigh–Ritz Method.
Mixed linear and nonlinear ansätze
Section titled “Mixed linear and nonlinear ansätze”Many high-accuracy calculations combine both structures:
For fixed , the coefficients can be found by generalized diagonalization. An outer nonlinear loop then changes orbital shapes, Gaussian exponents, or other basis parameters. Separating the exact inner linear solve from the nonlinear outer search often improves stability and interpretability.
Stochastic evaluation
Section titled “Stochastic evaluation”In many-body configuration space, deterministic integration can be impossible. Variational Monte Carlo samples and estimates the energy from the local energy. It preserves the conceptual variational structure, but finite-sample uncertainty and sampling bias must be reported separately. See Variational Monte Carlo Preview.
Time-dependent trial manifolds
Section titled “Time-dependent trial manifolds”The Time-Dependent Variational Principle projects Schrödinger evolution onto a manifold of parameterized states. Its purpose is not generally to minimize the instantaneous energy. It chooses the tangent-space motion that best satisfies an action or residual condition. The word “variational” therefore refers to a broader stationarity principle than the static ground-state bound.
Admissibility Comes Before Optimization
Section titled “Admissibility Comes Before Optimization”A flexible ansatz can still be invalid. Before evaluating an energy, check that each trial state has the properties required by the physical problem.
Normalizability and finite energy
Section titled “Normalizability and finite energy”A wavefunction may be square-integrable while its kinetic energy diverges. A cusp, discontinuity, or slowly decaying tail must be assessed using the quadratic form, not normalization alone.
For a one-dimensional kinetic term,
the form expectation is often written
after boundary terms have been shown to vanish. This expression makes the finite-gradient requirement visible.
Boundary conditions
Section titled “Boundary conditions”Trial functions must respect hard walls, periodicity, self-adjoint boundary conditions, and regularity at singular points. Violating a boundary condition can enlarge the trial set beyond the Hamiltonian’s domain and destroy the claimed bound.
Exchange symmetry
Section titled “Exchange symmetry”For identical particles, the trial state must lie in the correct bosonic or fermionic sector. Antisymmetry is structural, not a small correction. Slater determinants enforce it for fermions, while correlation factors can be multiplied into an antisymmetric reference without changing particle-exchange parity.
Conserved symmetry sectors
Section titled “Conserved symmetry sectors”Suppose preserves a sector , such as fixed parity or angular momentum. Minimization within that sector gives
An admissible trial state in satisfies
The sector ground state may be an excited state of the full Hamiltonian. Symmetry adaptation can therefore target selected excitations while retaining a variational bound, provided the sector is truly invariant.
Choosing Parameters by Scale Analysis
Section titled “Choosing Parameters by Scale Analysis”Optimization works better when parameters reflect physical scales. Consider a schematic one-dimensional Hamiltonian
For a normalized trial family of width , dimensional reasoning gives
Balancing the two terms predicts
and the energy scale
The omitted numerical factors depend on the trial shape. Even so, scale analysis supplies sensible initial parameters, exposes unit mistakes, and suggests dimensionless optimization variables. A parameter search across many orders of magnitude is usually better performed in than in itself.
Upper Bounds and Systematic Improvement
Section titled “Upper Bounds and Systematic Improvement”If two trial sets are nested,
then
Enlarging the allowed set cannot worsen the exact optimized energy because every old trial state remains available. This monotonicity is one of the strongest diagnostics in basis-set calculations.
The statement needs all of its qualifiers:
- the Hamiltonian and its matrix elements must be unchanged;
- the trial sets must actually be nested;
- each minimization must be sufficiently converged;
- numerical evaluation must be accurate enough to resolve the difference.
Changing a nonlinear ansatz from one family to another does not generally produce a monotone sequence. Neither family may contain the other.
A bound is not an error bar
Section titled “A bound is not an error bar”From
alone, one cannot infer the size of . A certified lower bound, an exact comparison value, a converged numerical benchmark, or an additional theorem is needed to bracket the error. Agreement between two upper bounds is evidence, not a proof, because two ansätze can share the same bias.
Excited states and the min–max principle
Section titled “Excited states and the min–max principle”The ground-state inequality does not by itself make an arbitrary excited-state ansatz an upper bound to a chosen excited level. For discrete eigenvalues ordered as , the relevant statement is the min–max characterization
Ritz eigenvalues inherit upper-bound properties under the theorem’s hypotheses. By contrast, minimizing over states orthogonal to inaccurate approximations of lower eigenstates does not automatically produce a rigorous bound to the next exact level. Upper Bounds and the Min–Max Principle gives the theorem, proof, essential-spectrum caveats, and interlacing results; the Rayleigh–Ritz Method develops the matrix construction.
Energy Accuracy and State Accuracy
Section titled “Energy Accuracy and State Accuracy”The variational energy is often more accurate than the trial state. This useful property can become a trap if energy agreement is treated as a universal validation.
Assume a nondegenerate ground state and a positive spectral gap
For a normalized trial state,
Therefore
This converts an energy error into a fidelity guarantee only when and a usable gap bound are known. If is small, a tiny energy error can coexist with substantial mixing. For a degenerate ground space, replace the overlap with , where projects onto that space.
For a bounded observable , the pure-state trace-distance inequality gives
Combining the two bounds yields
No corresponding conclusion follows for an arbitrary unbounded observable without extra assumptions. Short-distance densities, high moments, cusp-sensitive quantities, and tails can converge much more slowly than the energy.
A Two-Level Diagnostic Model
Section titled “A Two-Level Diagnostic Model”Let
and consider
The energy landscape is
Several general lessons are visible in this small model:
- The minimum occurs at the exact ground state, modulo phase and periodic parameter redundancy.
- The energy error is quadratic in the small state-amplitude error .
- The energy is independent of , although observables with off-diagonal matrix elements can depend on .
- A small gap makes the landscape shallow and parameter estimation ill-conditioned.
The energy variance is
It vanishes at both exact eigenstates. Variance can diagnose whether a state is close to some eigenstate, but energy or symmetry information is needed to identify which one.
Diagnostics Beyond the Energy
Section titled “Diagnostics Beyond the Energy”A mature variational calculation reports how the result was challenged. No single diagnostic is sufficient in every problem.
Nested-space convergence
Section titled “Nested-space convergence”Increase a basis cutoff or enlarge a genuinely nested trial space. The optimized energy should decrease toward a stable value. Nonmonotonic behavior signals incomplete optimization, numerical error, or non-nested approximations.
Residual and energy variance
Section titled “Residual and energy variance”For a normalized trial state with energy , define
When lies in the operator domain of ,
An exact eigenstate has zero residual. A small residual is stronger evidence of local eigen-equation accuracy than a low energy alone, although interpreting it quantitatively still requires spectral information.
In a projected Rayleigh–Ritz calculation, the residual is orthogonal to the trial subspace at a stationary Ritz vector:
This Galerkin orthogonality does not imply outside the subspace.
Symmetry and exact identities
Section titled “Symmetry and exact identities”Check conserved quantum numbers, exchange symmetry, parity, and known commutator identities. The Hellmann–Feynman Theorem and virial relations can supply useful diagnostics, but only when their assumptions and parameter dependence are handled correctly.
Independent trial families
Section titled “Independent trial families”Compare ansätze with different structural biases: Gaussian versus exponential tails, coordinate-space versus basis-space representations, or deterministic versus stochastic evaluation. Agreement is more informative when the approximations fail differently.
Observable convergence
Section titled “Observable convergence”Track the observables relevant to the scientific question, not just the optimized objective. Energy convergence does not guarantee convergence of radii, contact densities, transition matrix elements, entanglement, or response functions.
Numerical conditioning
Section titled “Numerical conditioning”In a nonorthogonal basis, inspect the spectrum or condition number of the overlap matrix . Near-linear dependence can produce apparently dramatic energy improvements that are numerical artifacts. Removing redundant directions or using a rank-revealing orthogonalization is part of the method, not housekeeping.
Method Roadmap
Section titled “Method Roadmap”| Question | Natural formulation | Continue with |
|---|---|---|
| Why is the ground-state estimate an upper bound? | Rayleigh quotient over the energy form domain | Variational Principle |
| How should physical information enter an ansatz? | Boundary, symmetry, scale, cusp, and tail constraints | Trial Wavefunctions |
| How are nonlinear ansatz coordinates optimized and diagnosed? | Gradients, tangent metrics, Hessians, and constrained searches | Variational Parameters |
| What does a complete one-parameter calculation look like? | Gaussian width for a problem with a known exact answer | Variational Estimate for the Harmonic Oscillator |
| How do Gaussian widths generalize to coupled coordinates? | Positive-definite width matrices and correlated Gaussian bases | Gaussian Variational Methods |
| How can a divergent weak-coupling series be reorganized around an adjustable reference problem? | Formal δ interpolation, order-dependent re-expansion, and optimization | Variational Perturbation Theory |
| How is a linear trial space optimized? | Hamiltonian and overlap matrices | Rayleigh–Ritz Method |
| Which excited-state Ritz values remain rigorous upper bounds? | Ordered subspace optimization below essential spectrum | Upper Bounds and the Min–Max Principle |
| How does screening improve a two-electron estimate? | Effective-charge product state | Variational Estimate for the Helium Atom |
| How can a parameterized state follow real-time motion? | Tangent-space projection or action stationarity | Time-Dependent Variational Principle |
| How are high-dimensional expectations evaluated stochastically? | Sampling and local energy | Variational Monte Carlo Preview |
| How should a suspicious result be diagnosed? | Symptom, failed assumption, quantitative test, and repair | Common Variational Pitfalls |
Connections to Other Approximation Strategies
Section titled “Connections to Other Approximation Strategies”Perturbation theory
Section titled “Perturbation theory”Perturbation theory expands around a solved Hamiltonian and can deliver analytic corrections order by order. A truncated perturbation series is not generally an upper bound. Variational methods instead optimize over states and can remain useful when no small coupling is available, but their accuracy depends strongly on ansatz design. The two approaches can be combined by using perturbative insight to construct or improve a trial family. Variational Perturbation Theory develops a more systematic hybrid in which an adjustable reference Hamiltonian is introduced, the series is re-expanded, and the residual parameter dependence is optimized at each order.
Semiclassical methods
Section titled “Semiclassical methods”WKB and related methods organize an asymptotic expansion in a scale ratio such as the de Broglie wavelength divided by a potential-variation length. Their errors have a different character from variational ansatz bias. Comparing a variational upper bound with a semiclassical estimate can be informative, but the semiclassical value does not automatically provide a lower bound.
Many-body and electronic-structure methods
Section titled “Many-body and electronic-structure methods”Hartree–Fock, configuration interaction, matrix-product states, tensor-network methods, and many neural-network quantum states are variational when they minimize the expectation value of a specified Hamiltonian over a restricted state class. Each changes the expressive family and the optimization problem; none removes the need to separate model, ansatz, evaluation, and optimization errors.
Variational quantum eigensolvers
Section titled “Variational quantum eigensolvers”A variational quantum eigensolver prepares a parameterized state and estimates Hamiltonian expectation values. In the ideal mathematical description, the exact expectation of the prepared state obeys the same upper bound. Finite measurement statistics, device noise, imperfect Hamiltonian decomposition, and error mitigation can make the reported estimator nonvariational. The word “variational” should not be used as a substitute for an uncertainty budget.
Common Failure Modes
Section titled “Common Failure Modes”- Using an inadmissible state. Normalization alone does not establish finite energy or correct boundary conditions.
- Treating an upper bound as a symmetric error bar. The variational principle controls the sign, not the magnitude, of the energy error.
- Optimizing a modified objective. Penalty terms, truncated Hamiltonians, or noisy estimators may not bound the original Hamiltonian.
- Stopping at a stationary point. A zero gradient can identify a maximum, saddle, or local minimum.
- Comparing non-nested ansätze as though convergence must be monotone. Monotonicity follows from set inclusion, not from increasing a parameter count.
- Ignoring overlap-matrix conditioning. Nearly dependent basis vectors can corrupt generalized eigenvalues.
- Forgetting symmetry sectors. A low energy in the wrong sector does not approximate the targeted state.
- Using approximate orthogonality to claim an excited-state bound. Excited-state guarantees require the min–max structure or exact constraints.
- Equating energy accuracy with wavefunction accuracy. Small gaps and insensitive observables can hide substantial state error.
- Optimizing before checking scales. Poorly scaled parameters create flat directions and avoidable numerical instability.
The Common Variational Pitfalls guide turns these warnings into symptom-to-test-to-repair workflows, including boundary-domain failures, state-versus-energy accuracy, noisy estimators, and comparisons with perturbation theory.
Exercises
Section titled “Exercises”1. Scale invariance of the Rayleigh quotient
Section titled “1. Scale invariance of the Rayleigh quotient”Show that for every nonzero complex . Explain why minimizing the unnormalized numerator alone is not an equivalent procedure.
Solution
The numerator and denominator transform as
The factors cancel in the quotient. By contrast, the unnormalized numerator can be driven toward zero by taking , regardless of the state’s shape. One must either constrain the norm or minimize the quotient.
2. Nested trial spaces
Section titled “2. Nested trial spaces”Let . Prove that their exact optimized energies obey . Does this prove that either value is close to ?
Solution
The infimum over is taken over every state in plus possibly more states. Therefore
Together with the variational principle,
This proves monotone improvement of the upper bound, not closeness to . Both trial spaces can omit an important feature and remain far above the exact ground energy.
3. Two-level energy and variance
Section titled “3. Two-level energy and variance”For the two-level trial state in the text, derive and . Identify every value of for which the variance vanishes.
Solution
The occupation probabilities are and , so
Similarly,
Subtracting gives
The variance vanishes when or , corresponding to either exact energy eigenstate. Zero variance alone does not select the ground state.
4. Fidelity from an energy estimate
Section titled “4. Fidelity from an energy estimate”A normalized trial state has energy for a Hamiltonian with a nondegenerate ground state and gap . Derive a lower bound on the ground-state fidelity. What does the bound say if ?
Solution
Writing , the gap inequality gives
Hence
If , the right side is nonpositive and the inequality supplies no nontrivial fidelity information. The variational energy remains an upper bound, but the gap estimate is then too weak to certify overlap.
5. A symmetry-sector bound
Section titled “5. A symmetry-sector bound”Suppose and . Let be the range of . Show that minimizing the Rayleigh quotient over bounds the lowest energy in that invariant sector, even if that state is not the absolute ground state.
Solution
Because , the range of is invariant under . Restricting to gives a self-adjoint sector Hamiltonian under the appropriate domain assumptions. Applying the variational principle to this restriction yields
The value is the bottom of the spectrum in that sector. Another sector can contain a lower state, so need not equal the absolute ground energy.
6. Width scaling in a quartic trap
Section titled “6. Width scaling in a quartic trap”For with , use scale balancing to determine the dependence of the optimal width and energy on , , and . Numerical constants are not required.
Solution
For a trial state of width ,
Balancing the two contributions gives
so
Substitution gives
An explicit ansatz fixes the dimensionless prefactors but cannot change these scalings.
References
Section titled “References”- W. Ritz, “Über eine neue Methode zur Lösung gewisser Variationsprobleme der mathematischen Physik”, Journal für die reine und angewandte Mathematik 135, 1–61 (1909). The foundational finite-basis variational construction.
- J. K. L. MacDonald, “Successive Approximations by the Rayleigh–Ritz Variation Method”, Physical Review 43, 830–833 (1933). A classic account of variational eigenvalue bounds and successive subspace approximations.
- M. Reed and B. Simon, Methods of Modern Mathematical Physics, Vol. IV: Analysis of Operators (Academic Press, 1978), especially the spectral and min–max treatment. Bibliographic record.
- R. K. Nesbet, Variational Principles and Methods in Theoretical Physics and Chemistry (Cambridge University Press, 2003). A graduate-level survey spanning bound states, time-dependent theory, and scattering applications.
- E. A. Hylleraas, “Neue Berechnung der Energie des Heliums im Grundzustande, sowie des tiefsten Terms von Ortho-Helium”, Zeitschrift für Physik 54, 347–366 (1929). A landmark correlated variational calculation for helium.
- R. McLachlan, “A variational solution of the time-dependent Schrödinger equation”, Molecular Physics 8, 39–44 (1964). A foundational residual-minimization formulation for variational dynamics.
- W. M. C. Foulkes, L. Mitas, R. J. Needs, and G. Rajagopal, “Quantum Monte Carlo simulations of solids”, Reviews of Modern Physics 73, 33–83 (2001). A detailed review of variational and diffusion Monte Carlo, including optimization and statistical issues.