Variational Monte Carlo Preview
Variational Monte Carlo, or VMC, evaluates expectation values of a parameterized trial state by stochastic sampling. Its central move is to rewrite a high-dimensional quantum expectation as an ordinary average over configurations distributed according to .
The method combines three logically distinct ingredients:
- a variational family that encodes the physics one is willing to represent;
- a Monte Carlo estimator of the energy and other observables for fixed ;
- a stochastic optimization that changes using noisy estimates.
Only the first ingredient is covered by the exact variational upper-bound theorem. Sampling and optimization introduce additional uncertainty and bias that must be analyzed separately.
This page owns the physics of the VMC estimator: the sampling density, local energy, zero-variance identity, covariance gradient, and geometric optimization preview. Variational Many-Body States owns the state-family comparison, including determinants, Jastrow factors, projections, MPS, and neural amplitudes. General probability results belong to Monte Carlo Basics. Production Markov-chain design, blocking analysis, parallel implementation, and benchmark code belong to the future computational treatment.
Configuration-space formulation
Section titled “Configuration-space formulation”Let denote a complete configuration. Depending on the problem, it may collect particle positions,
occupation numbers, spin labels, lattice configurations, or a mixture of continuous and discrete variables. The symbol
will stand for the corresponding integrals and sums.
Choose a nonzero trial wavefunction with real parameters . Its squared norm is
and its normalized configuration density is
The wavefunction itself may be real, complex, positive, sign-changing, or antisymmetric. The sampling density is nevertheless nonnegative.
Local estimators
Section titled “Local estimators”For an operator , define its local estimator by
wherever . Then
The identity follows by multiplying by . It is the bridge from a quantum expectation value to a statistical average.
If is diagonal in the sampled configuration basis, then is simply its diagonal value. For a differential or off-diagonal operator, the local estimator contains derivatives or ratios of wavefunction amplitudes at connected configurations.
The formula is formal at nodes of . A node has zero probability under , but nearby singularities can still produce large fluctuations or even an infinite estimator variance. A set of probability zero is not automatically harmless.
The local energy
Section titled “The local energy”For the Hamiltonian,
and the exact variational energy at fixed parameters is
For a continuum Hamiltonian
the local energy is
Where a smooth branch of exists, one may use
For a complex wavefunction, the last dot product does not complex-conjugate the first factor. This logarithmic form often exposes short-distance cancellations and avoids separately evaluating very large amplitudes, although numerical implementation belongs to the computational treatment.
For self-adjoint , the expectation is real. The pointwise local energy can nevertheless be complex when is complex. In an exact calculation its imaginary part averages to zero; a persistent sampled imaginary mean is therefore a useful warning about boundary terms, coding errors, or unconverged statistics.
The VMC loop. For fixed , configurations are sampled from and converted into local estimates. Optimization then updates the trial state. Sampling error, optimization error, and ansatz bias enter at different stages.
Zero variance and the Schrödinger residual
Section titled “Zero variance and the Schrödinger residual”The local-energy variance is
Substituting the definitions gives
Thus the local-energy variance is the squared norm of the stationary Schrödinger residual, divided by the state norm. When lies in the operator domain of , this is also the usual energy variance.
If the trial state is an exact eigenstate,
then
wherever the ratio is defined, and . Conversely, a normalized pure state with zero residual variance is an eigenstate. This is the zero-variance principle.
The result is stronger than the statement that an eigenstate has the correct mean energy: every sampled configuration returns the same local energy. Near a good eigenstate, reduced local-energy fluctuations often make VMC statistically efficient.
Two cautions matter:
- A small energy error does not universally imply a proportionally small variance, or vice versa, without information about spectral gaps and state overlap.
- The Rayleigh quotient may be finite for a state in the quadratic-form domain even when is not square-integrable. In that case the energy exists but the local-energy variance can diverge.
Correct cusp, boundary, and nodal behavior is therefore both physical and statistical. It can remove singular cancellations that would otherwise create heavy local-energy tails.
Sampling the Born density
Section titled “Sampling the Born density”Independent samples from are rarely available in a complicated many-body problem. VMC commonly uses a Markov chain with stationary density .
Given a proposal density , the Metropolis–Hastings acceptance probability is
For a symmetric proposal,
Normalization cancels from the ratio, so VMC does not require prior knowledge of . Detailed balance is a sufficient route to the desired stationary distribution, but stationarity alone is not enough: the chain must also explore every relevant region on the time scale of the calculation.
Nodes or separated modes can obstruct mixing. A local proposal may remain trapped in one nodal pocket or one metastable region even though its acceptance rate looks healthy. Acceptance rate is therefore not an ergodicity certificate.
Fermions and signs
Section titled “Fermions and signs”For a fixed antisymmetric trial wavefunction, ordinary VMC samples the positive density . The fermionic sign or phase enters the local energy through amplitude ratios and derivatives, not through a fluctuating signed sampling weight. In this limited sense, evaluating a specified fermionic trial state does not have the same average-sign problem as many projection or path-integral methods.
The hard fermionic physics has not disappeared. The antisymmetry and nodal surface must be represented by the ansatz, and poor nodes can dominate the variational bias. Fixed-node projection is a separate approximation associated with diffusion Monte Carlo, not a property of VMC itself.
Energy estimator and uncertainty
Section titled “Energy estimator and uncertainty”After equilibration, suppose a chain provides configurations
The sample-mean energy estimator is
For a complex local energy, one normally reports the real part and separately checks that the sampled imaginary mean is statistically consistent with zero.
If the samples were independent and the variance finite, the standard error would scale as . Markov-chain samples are correlated, so the relevant quantity is the autocovariance of the real local-energy fluctuations. Let
and define
The integrated autocorrelation time is
when the sum converges. For a long stationary chain,
with
For a real local energy, . For a complex local energy, the residual norm also contains fluctuations of the imaginary part, whereas the quoted real-energy standard error does not.
The central-limit approximation requires adequate mixing and sufficiently light tails. If has rare, extreme excursions or infinite variance, a familiar-looking standard error can be meaningless. Independent chains, trace diagnostics, blocking or batching, and tail inspection are not optional decorations.
The reusable probability theory and error-bar machinery are developed in Monte Carlo Basics. This page keeps the emphasis on what those quantities mean for a variational quantum state.
What remains of the upper bound?
Section titled “What remains of the upper bound?”For every fixed admissible trial state, the exact Rayleigh quotient obeys
The finite-sample estimate does not obey the inequality sample by sample. It has the form
where statistical fluctuation, incomplete equilibration, and implementation error have no common one-sided sign. A reported VMC number can fall below even when the underlying variational energy is above it.
Calling a Monte Carlo estimate an “upper bound” is justified only after attaching a statistically and systematically defensible one-sided uncertainty statement. A symmetric one-standard-error interval is not a rigorous bound.
Optimization creates another subtlety. If many noisy estimates are compared and the smallest is selected, the winner is preferentially associated with a downward fluctuation. Reusing the same sample to optimize and to report the final energy can therefore produce selection bias. A fresh validation sample at the final parameters helps separate optimization noise from final estimation.
Worked example: Gaussian oscillator recast as VMC
Section titled “Worked example: Gaussian oscillator recast as VMC”The Variational Estimate for the Harmonic Oscillator owns the deterministic derivation. Here the same family is recast as a sampling problem.
Take
so that
For
the local energy is
Sampling from and using gives
Let
The exact ground-state width is . At that point the coefficient of in vanishes, so every sample returns
For a nonoptimal width, the local energy fluctuates with . Its exact variance is
This example displays all three layers cleanly. The family contains the exact ground state, the local-energy variance diagnoses the residual, and finite sampling estimates the same analytic Rayleigh quotient with a parameter-dependent error bar.
Energy gradients from samples
Section titled “Energy gradients from samples”The deterministic geometry of parameter optimization is developed in Variational Parameters. VMC makes its gradients into covariance estimators.
For real parameters, define logarithmic derivatives
Thus
Assume is independent of , the required derivatives can be moved through the integrals, and boundary terms are controlled. Differentiating the Rayleigh quotient and using self-adjointness gives
Equivalently,
For a real positive trial state, this reduces to twice the covariance of and . The formula is valuable because it does not require a separate pointwise derivative of : the Hamiltonian’s self-adjointness has converted the derivative into a covariance.
Replacing the expectations by sample means produces a stochastic gradient. Its noise is correlated across parameters because every component uses the same configurations. Near an eigenstate, the factor becomes small, reflecting the zero-variance principle.
Finite-sample covariance estimators can still be biased at order , and adaptive reuse of samples complicates the error analysis. Gradient convergence should be checked separately from energy convergence.
Stochastic reconfiguration and natural gradient
Section titled “Stochastic reconfiguration and natural gradient”The logarithmic derivatives also give the pullback of the projective state metric:
For any real vector ,
Null directions correspond to parameter combinations that do not change the physical ray to first order, such as redundant coordinates or pure normalization and phase changes.
A natural-gradient or stochastic-reconfiguration step solves schematically
where controls the step size. This chooses a small change in the quantum state rather than a small change measured only by coordinate distance.
The same metric appears in imaginary-time projection and the Time-Dependent Variational Principle. Stochastic reconfiguration is therefore not merely a preconditioner chosen by convenience; in its ideal form it is a sampled state-space projection.
In practice, is itself noisy and often ill-conditioned. Gauge fixing, rank truncation, diagonal shifts, trust regions, or pseudoinverses may be needed. Every such regularization changes the update and should be included in sensitivity tests.
Other optimization strategies include direct stochastic gradients, variance minimization, and the linear method. No optimizer repairs a trial family that omits the relevant symmetry, nodes, correlations, or asymptotic behavior.
Trial-state design for VMC
Section titled “Trial-state design for VMC”A VMC ansatz must be physically admissible, but it must also make the sampled quantities tractable. Useful requirements include:
- inexpensive evaluation of amplitude ratios or logarithmic amplitudes;
- exact enforcement of required exchange symmetry and quantum numbers;
- correct boundary, cusp, and asymptotic behavior where known;
- stable derivatives for and ;
- support broad enough to cover every relevant configuration sector;
- a parameterization whose metric is not needlessly singular.
Common continuum fermion states combine a Slater determinant with a symmetric Jastrow factor,
The determinant enforces antisymmetry, while the Jastrow factor represents symmetric correlations and can encode cusp conditions. Backflow coordinates, pairing determinants, and Pfaffian forms enrich the nodal and pairing structure.
On lattices, correlator products, projected mean-field states, tensor-network amplitudes, and neural-network wavefunctions can all be sampled when their amplitudes and local connections are evaluable. Expressivity is not a guarantee of accuracy: a larger family may be harder to mix, optimize, condition, and validate.
Neural quantum states are therefore best viewed as flexible VMC ansätze, not as a separate variational theorem. Their claims require the same energy, variance, symmetry, convergence, and independent-benchmark checks as traditional trial states.
Observables beyond the energy
Section titled “Observables beyond the energy”The local-estimator identity applies to any operator for which the expectation exists. For a Hermitian observable ,
Unlike the energy, a generic observable has no variational upper-bound property. It also need not be stationary to first order in the wavefunction error. A trial state can therefore have an excellent energy while giving a noticeably biased density, correlation function, response, or transition matrix element.
Variance and mixing are observable-dependent. A chain that estimates the energy efficiently may have a much longer autocorrelation time for a collective order parameter. Error bars must be computed for the observable actually reported.
Off-diagonal estimators may involve large amplitude ratios and heavier tails than the energy. Their finite variance should be checked rather than assumed.
Error budget
Section titled “Error budget”Let minimize the exact variational energy within the chosen family, and let be the parameters returned by a noisy optimization. Then a useful conceptual decomposition is
For an exact global minimum within an admissible family, the first two terms are nonnegative. The last term has either sign and includes more than finite-sample variance if equilibration or implementation is imperfect.
A complete VMC error ledger separates:
- ansatz bias: the family cannot represent the target state;
- optimization error: the algorithm does not reach the best state in the family;
- sampling variance: a finite correlated sample fluctuates;
- equilibration and mixing bias: the chain does not represent ;
- estimator pathology: local values have heavy or nonintegrable tails;
- implementation error: derivatives, Hamiltonian terms, boundary conditions, or precision are wrong;
- model or discretization error: the simulated Hamiltonian differs from the intended physical problem.
Only sampling variance is expected to shrink automatically as under ordinary central-limit conditions.
Minimum reporting standard
Section titled “Minimum reporting standard”A trustworthy VMC result should report enough information to reconstruct the error logic:
- the Hamiltonian, boundary conditions, and sampled configuration space;
- the trial-state form, symmetry sector, parameter count, and optimization objective;
- the sampling distribution and proposal class;
- equilibration checks, chain lengths, independent-chain count, and autocorrelation or blocking analysis;
- the final energy with uncertainty and the local-energy variance;
- optimization convergence and regularization sensitivity;
- convergence under ansatz enlargement and comparison with independent benchmarks;
- separate validation samples when optimization and final estimation share data;
- random seeds and other reproducibility metadata in computational work.
An energy with many printed digits but no autocorrelation or ansatz-convergence evidence is not a high-precision result.
Boundary with computational methods
Section titled “Boundary with computational methods”This preview establishes why the VMC estimators work and what their errors mean. A full computational treatment should own:
- stable determinant, Pfaffian, and neural-amplitude updates;
- proposal design and drift-diffusion sampling;
- burn-in, blocking, effective sample size, and convergence diagnostics;
- parallel and population sampling;
- robust stochastic-reconfiguration and linear-method solvers;
- automatic differentiation and custom local-energy kernels;
- correlated sampling and reweighting diagnostics;
- benchmark implementations with exact or independently converged answers.
Diffusion Monte Carlo, path-integral Monte Carlo, worldline methods, and auxiliary-field methods use different stochastic representations. They should not be inferred from the sampling identity alone.
Common mistakes
Section titled “Common mistakes”- Treating as an exact upper bound. The underlying Rayleigh quotient is variational; its noisy estimate can fluctuate below .
- Dividing the raw variance by . Correlated samples require an autocorrelation correction or equivalent blocking analysis.
- Using acceptance rate as the only chain diagnostic. A chain can accept often while remaining trapped in one mode or nodal pocket.
- Ignoring local-energy tails. A finite-looking mean does not guarantee a finite or well-estimated variance.
- Dropping complex conjugation in the gradient. For complex trial states, the covariance uses and a final real part.
- Reusing optimization samples without acknowledging selection bias. Fresh final samples help separate fitting from evaluation.
- Calling every fermionic difficulty a sign problem. Fixed-state VMC samples a positive density, while nodal bias and projection-method sign problems are distinct issues.
- Optimizing only the energy mean. Variance, symmetries, observables, and ansatz convergence reveal failures hidden by a favorable energy.
- Regularizing the metric silently. A diagonal shift or rank cutoff changes the parameter update.
- Assuming a more expressive ansatz is automatically better. Expressivity can worsen conditioning, mixing, and optimization.
Exercises
Section titled “Exercises”1. Derive the local-estimator identity
Section titled “1. Derive the local-estimator identity”Starting from the normalized expectation of an operator , derive
State the support caveat.
Solution
By definition,
Where , multiply and divide the integrand by :
The ratio is undefined at nodes. Nodes themselves have zero measure, but singular behavior near them must still be integrable for the mean and variance under discussion to exist.
2. Prove the zero-variance identity
Section titled “2. Prove the zero-variance identity”Show that
Why does zero variance imply an eigenstate?
Solution
Use
Then
Since , the identity follows. A squared Hilbert-space norm vanishes only when its vector vanishes, so zero variance gives
Thus a nonzero state in the operator domain is an eigenstate with eigenvalue .
3. Oscillator mean and variance
Section titled “3. Oscillator mean and variance”For the Gaussian oscillator example, use
to derive both and .
Solution
Write
where
The mean is
The constant does not contribute to the variance, and
Therefore
4. Effective sample size for exponential correlation
Section titled “4. Effective sample size for exponential correlation”Suppose with . Compute and . What fraction of remains effective when ?
Solution
The geometric series gives
Hence
For ,
Nine stored configurations then carry roughly the mean-estimation information of one independent draw, under this idealized correlation model.
5. Derive the covariance gradient
Section titled “5. Derive the covariance gradient”For real , use to derive
Solution
Let
Self-adjointness gives
Similarly,
Differentiating gives
Since the exact local-energy mean equals the real number , this is the stated covariance form.
6. A sampled value below the ground-state energy
Section titled “6. A sampled value below the ground-state energy”A VMC run returns with estimated standard error , while the exact ground-state energy is known to be . Does this violate the variational principle? What should be checked?
Solution
No. The variational principle constrains the exact Rayleigh quotient of the trial state, not every finite-sample realization . The observed difference is only one third of the quoted standard error.
One should check that the error estimate accounts for autocorrelation, that the chain is equilibrated and mixing, that the local-energy variance is finite, and that no implementation or discretization bias is present. If the parameters were selected using the same noisy sample, a fresh validation run is also appropriate. Only a controlled one-sided confidence statement could turn the stochastic result into a probabilistic upper bound.
References
Section titled “References”- W. L. McMillan, “Ground State of Liquid He4,” Physical Review 138, A442–A451 (1965), doi:10.1103/PhysRev.138.A442.
- N. Metropolis, A. W. Rosenbluth, M. N. Rosenbluth, A. H. Teller, and E. Teller, “Equation of State Calculations by Fast Computing Machines,” Journal of Chemical Physics 21, 1087–1092 (1953), doi:10.1063/1.1699114.
- W. K. Hastings, “Monte Carlo sampling methods using Markov chains and their applications,” Biometrika 57, 97–109 (1970), doi:10.1093/biomet/57.1.97.
- D. M. Ceperley, G. V. Chester, and M. H. Kalos, “Monte Carlo simulation of a many-fermion system,” Physical Review B 16, 3081–3099 (1977), doi:10.1103/PhysRevB.16.3081.
- C. J. Umrigar, K. G. Wilson, and J. W. Wilkins, “Optimized trial wave functions for quantum Monte Carlo calculations,” Physical Review Letters 60, 1719–1722 (1988), doi:10.1103/PhysRevLett.60.1719.
- S. Sorella, “Wave function optimization in the variational Monte Carlo method,” Physical Review B 71, 241103(R) (2005), doi:10.1103/PhysRevB.71.241103.
- W. M. C. Foulkes, L. Mitas, R. J. Needs, and G. Rajagopal, “Quantum Monte Carlo simulations of solids,” Reviews of Modern Physics 73, 33–83 (2001), doi:10.1103/RevModPhys.73.33.
- G. Carleo and M. Troyer, “Solving the quantum many-body problem with artificial neural networks,” Science 355, 602–606 (2017), doi:10.1126/science.aag2302.