Skip to content

Variance and Standard Deviation

For a state ρ\rho and self-adjoint observable AA, define the centered observable

δρA≡A−⟨A⟩ρI.\delta_\rho A \equiv A-\langle A\rangle_\rho I.

The variance and standard deviation are

Var⁡ρ(A)≡(ΔρA)2=⟨(δρA)2⟩ρ,\operatorname{Var}_\rho(A) \equiv (\Delta_\rho A)^2 = \left\langle (\delta_\rho A)^2 \right\rangle_\rho, ΔρA=Var⁡ρ(A).\Delta_\rho A = \sqrt{\operatorname{Var}_\rho(A)}.

When the second moment exists,

Var⁡ρ(A)=⟨A2⟩ρ−⟨A⟩ρ2.\operatorname{Var}_\rho(A) = \langle A^2\rangle_\rho - \langle A\rangle_\rho^2.

Variance has the squared units of AA; standard deviation has the same units as AA.

SettingVariance
Centered observableVar⁡ρ(A)=⟨(A−⟨A⟩ρI)2⟩ρ\operatorname{Var}_\rho(A)=\langle(A-\langle A\rangle_\rho I)^2\rangle_\rho
Raw momentsVar⁡ρ(A)=⟨A2⟩ρ−⟨A⟩ρ2\operatorname{Var}_\rho(A)=\langle A^2\rangle_\rho-\langle A\rangle_\rho^2
Pure-state norm(ΔψA)2=∥(A−⟨A⟩ψI)∣ψ⟩∥2(\Delta_\psi A)^2=\lVert(A-\langle A\rangle_\psi I)\lvert\psi\rangle\rVert^2
Density operator(ΔρA)2=Tr⁡(ρA2)−[Tr⁡(ρA)]2(\Delta_\rho A)^2=\operatorname{Tr}(\rho A^2)-[\operatorname{Tr}(\rho A)]^2
Discrete outcomes(ΔA)2=∑a(a−⟨A⟩)2p(a)(\Delta A)^2=\sum_a(a-\langle A\rangle)^2p(a)
Continuous outcomes(ΔA)2=∫(λ−⟨A⟩)2pA(λ) dλ(\Delta A)^2=\int(\lambda-\langle A\rangle)^2p_A(\lambda)\,d\lambda
Affine transformationVar⁡(αA+βI)=∣α∣2Var⁡(A)\operatorname{Var}(\alpha A+\beta I)=\lvert\alpha\rvert^2\operatorname{Var}(A) for real-valued observables
Two outcomes a,ba,bVar⁡(A)=p(1−p)(a−b)2\operatorname{Var}(A)=p(1-p)(a-b)^2
Bounded spectrum [amin⁡,amax⁡][a_{\min},a_{\max}]0≤Var⁡(A)≤(amax⁡−amin⁡)2/40\leq\operatorname{Var}(A)\leq(a_{\max}-a_{\min})^2/4

The centered and raw-moment formulas are mathematically equivalent, but the centered formula is often more stable numerically.

Variance is the second central moment of the Born distribution for measuring AA in the specified state. It measures the distribution’s spread around its mean. It is not, by itself:

  • an apparatus calibration error;
  • a measurement disturbance;
  • the error of an estimate of the mean;
  • the width of the wavefunction in every representation;
  • a complete description of the outcome distribution.

Two states can have the same expectation and variance while differing in skewness, tails, multimodality, or higher moments.

The standard deviation is the root-mean-square distance from the expectation:

ΔρA=[∫R(λ−⟨A⟩ρ)2 dμρA(λ)]1/2,\Delta_\rho A = \left[ \int_{\mathbb R} (\lambda-\langle A\rangle_\rho)^2 \,d\mu_\rho^A(\lambda) \right]^{1/2},

where μρA\mu_\rho^A is the Born probability measure of AA.

For a normalized pure state,

⟨A⟩ψ=⟨ψ∣A∣ψ⟩.\langle A\rangle_\psi = \langle\psi\rvert A\lvert\psi\rangle.

The norm form

(ΔψA)2=∥(A−⟨A⟩ψI)∣ψ⟩∥2(\Delta_\psi A)^2 = \left\lVert (A-\langle A\rangle_\psi I) \lvert\psi\rangle \right\rVert^2

makes nonnegativity immediate:

(ΔψA)2≥0.(\Delta_\psi A)^2\geq0.

For a density operator,

⟨A⟩ρ=Tr⁡(ρA),\langle A\rangle_\rho = \operatorname{Tr}(\rho A),

and

(ΔρA)2=Tr⁡(ρA2)−[Tr⁡(ρA)]2.(\Delta_\rho A)^2 = \operatorname{Tr}(\rho A^2) - \left[ \operatorname{Tr}(\rho A) \right]^2.

The trace formula includes pure states by setting ρ=∣ψ⟩⟨ψ∣\rho=\lvert\psi\rangle\langle\psi\rvert.

If

A=∑aaPa,p(a)=Tr⁡(ρPa),A = \sum_a aP_a, \qquad p(a) = \operatorname{Tr}(\rho P_a),

then

⟨A⟩=∑aap(a)\langle A\rangle = \sum_aap(a)

and

(ΔA)2=∑a(a−⟨A⟩)2p(a).(\Delta A)^2 = \sum_a (a-\langle A\rangle)^2 p(a).

Degeneracy is already included when PaP_a projects onto the full eigenspace. Do not sum separate probabilities for arbitrarily chosen vectors inside that eigenspace unless the measurement actually resolves them.

If the Born law has a density pA(λ)p_A(\lambda),

(ΔA)2=∫−∞∞(λ−⟨A⟩)2pA(λ) dλ.(\Delta A)^2 = \int_{-\infty}^{\infty} (\lambda-\langle A\rangle)^2 p_A(\lambda)\,d\lambda.

For one-dimensional position,

(ΔX)2=∫−∞∞(x−⟨X⟩)2∣ψ(x)∣2 dx.(\Delta X)^2 = \int_{-\infty}^{\infty} (x-\langle X\rangle)^2 \lvert\psi(x)\rvert^2\,dx.

For a spectrum with discrete and continuous parts, include both:

(ΔA)2=∑n(an−⟨A⟩)2wn+∫R(λ−⟨A⟩)2pac(λ) dλ.\begin{aligned} (\Delta A)^2 ={}& \sum_n (a_n-\langle A\rangle)^2w_n\\ &+ \int_{\mathbb R} (\lambda-\langle A\rangle)^2 p_{\mathrm{ac}}(\lambda)\,d\lambda. \end{aligned}

Normalization does not guarantee a finite variance. For a pure state and self-adjoint AA, a finite second spectral moment requires

∫Rλ2 dμψA(λ)<∞.\int_{\mathbb R} \lambda^2 \,d\mu_\psi^A(\lambda) <\infty.

Equivalently,

∣ψ⟩∈D(A).\lvert\psi\rangle \in \mathcal D(A).

Then A∣ψ⟩A\lvert\psi\rangle is a Hilbert-space vector and

⟨A2⟩ψ=∥A∣ψ⟩∥2\langle A^2\rangle_\psi = \lVert A\lvert\psi\rangle\rVert^2

in the spectral or quadratic-form sense. The literal vector expression ⟨ψ∣A2∣ψ⟩\langle\psi\rvert A^2\lvert\psi\rangle would demand ∣ψ⟩∈D(A2)\lvert\psi\rangle\in\mathcal D(A^2), which is stronger than necessary.

For a density operator, a standard sufficient condition is

Tr⁡(ρA2)<∞\operatorname{Tr}(\rho A^2)<\infty

with the product understood through the positive spectral calculus. Formal matrix multiplication cannot override domain or convergence failures in infinite dimensions.

If the second moment diverges, write

Var⁡ρ(A)=∞\operatorname{Var}_\rho(A)=\infty

when that extended value is meaningful. Do not subtract two divergent quantities in the raw-moment formula.

For a pure state with finite variance,

ΔψA=0\Delta_\psi A=0

if and only if

(A−⟨A⟩ψI)∣ψ⟩=0.(A-\langle A\rangle_\psi I) \lvert\psi\rangle=0.

Thus the state lies in the eigenspace with eigenvalue ⟨A⟩ψ\langle A\rangle_\psi.

For a mixed state, zero variance means the support of ρ\rho lies entirely in one eigenspace of AA. The state can still be mixed within a degenerate eigenspace. Therefore zero variance for one observable does not imply that the state is pure.

For real α\alpha and β\beta,

⟨αA+βI⟩=α⟨A⟩+β,\langle\alpha A+\beta I\rangle = \alpha\langle A\rangle+\beta,

while

Var⁡(αA+βI)=α2Var⁡(A).\operatorname{Var}(\alpha A+\beta I) = \alpha^2 \operatorname{Var}(A).

Consequently,

Δ(αA+βI)=∣α∣ΔA.\Delta(\alpha A+\beta I) = \lvert\alpha\rvert\Delta A.

Adding a constant shifts every outcome without changing the spread. Rescaling the measurement scale rescales standard deviation by the magnitude of the same factor.

The dimensions are

[Var⁡(A)]=[A]2,[ΔA]=[A].[\operatorname{Var}(A)] = [A]^2, \qquad [\Delta A]=[A].

This is a fast way to detect the common error of identifying variance itself with ΔA\Delta A.

Suppose AA has only outcomes aa and bb, with probabilities pp and 1−p1-p. Then

⟨A⟩=pa+(1−p)b\langle A\rangle = pa+(1-p)b

and

Var⁡(A)=p(1−p)(a−b)2.\operatorname{Var}(A) = p(1-p)(a-b)^2.

The variance is maximal at p=1/2p=1/2 and vanishes at p=0p=0 or p=1p=1.

For a Pauli observable

σn=n⋅σ,∥n∥=1,\sigma_{\mathbf n} = \mathbf n\cdot\boldsymbol{\sigma}, \qquad \lVert\mathbf n\rVert=1,

one has σn2=I\sigma_{\mathbf n}^2=I, so

(Δσn)2=1−⟨σn⟩2.(\Delta\sigma_{\mathbf n})^2 = 1- \langle\sigma_{\mathbf n}\rangle^2.

For a spin component

Sn=ℏ2σn,S_{\mathbf n} = \frac{\hbar}{2}\sigma_{\mathbf n}, (ΔSn)2=ℏ24[1−⟨σn⟩2].(\Delta S_{\mathbf n})^2 = \frac{\hbar^2}{4} \left[ 1- \langle\sigma_{\mathbf n}\rangle^2 \right].

Let a preparation be described by

ρ=∑kqkρk,qk≥0,∑kqk=1.\rho = \sum_kq_k\rho_k, \qquad q_k\geq0, \qquad \sum_kq_k=1.

Then

Var⁡ρ(A)=∑kqkVar⁡ρk(A)+∑kqk(⟨A⟩ρk−⟨A⟩ρ)2.\begin{aligned} \operatorname{Var}_\rho(A) ={}& \sum_kq_k \operatorname{Var}_{\rho_k}(A)\\ &+ \sum_kq_k \left( \langle A\rangle_{\rho_k} - \langle A\rangle_\rho \right)^2. \end{aligned}

The first term is the average spread within the component states. The second is the classical spread of their means. Mixing can therefore increase variance even when each component has zero variance.

This decomposition depends on the chosen ensemble representation of ρ\rho, although the total on the left does not. Different ensembles for the same density operator can divide the total into different within- and between-component contributions.

If the spectrum of AA lies in the finite interval [amin⁡,amax⁡][a_{\min},a_{\max}], then

0≤Var⁡ρ(A)≤(amax⁡−amin⁡)24.0 \leq \operatorname{Var}_\rho(A) \leq \frac{ (a_{\max}-a_{\min})^2 }{4}.

The upper bound is reached by a distribution with equal probability at the two endpoints when the endpoint eigenspaces are available. A result outside this range signals an incorrect state, operator, probability distribution, or calculation.

For two observables AA and BB, the standard deviations enter the Robertson bound

ΔA ΔB≥12∣⟨[A,B]⟩∣.\Delta A\,\Delta B \geq \frac{1}{2} \left\lvert \langle[A,B]\rangle \right\rvert.

The stronger Schrödinger form also contains the covariance of the centered observables. These inequalities constrain a pair of outcome distributions. They do not redefine variance, and they do not say that measuring AA necessarily disturbs BB by an amount ΔB\Delta B.

The derivations and equality conditions belong to General Uncertainty Relations.

For NN independent measurements of the same observable in the same state, the sample mean has standard deviation

ΔA‾=ΔAN.\Delta\overline A = \frac{\Delta A}{\sqrt N}.

This is the standard error of the sample mean under the independent, identically distributed model. It is distinct from the single-shot quantum spread ΔA\Delta A.

If independent zero-mean detector noise with variance σdet2\sigma_{\mathrm{det}}^2 is added to the ideal outcome, then

Var⁡(Aobs)=Var⁡(A)+σdet2.\operatorname{Var}(A_{\mathrm{obs}}) = \operatorname{Var}(A) + \sigma_{\mathrm{det}}^2.

Correlated, biased, state-dependent, or deconvolved noise requires a more specific measurement model. One cannot identify an observed histogram width with intrinsic quantum variance without calibration.

The raw-moment formula

⟨A2⟩−⟨A⟩2\langle A^2\rangle-\langle A\rangle^2

can subtract two large, nearly equal floating-point numbers. Roundoff may then produce a small negative result even though exact variance is nonnegative.

Prefer one of the following:

  • evaluate the centered norm ∥(A−⟨A⟩I)∣ψ⟩∥2\lVert(A-\langle A\rangle I)\lvert\psi\rangle\rVert^2;
  • sum centered outcome deviations;
  • use a stable online variance algorithm for sampled data;
  • symmetrize numerically Hermitian inputs;
  • treat only a negative value consistent with a documented tolerance as roundoff.

A materially negative variance is not repaired by an absolute value. Diagnose normalization, Hermiticity, basis, and arithmetic first.

  • The state is normalized and valid.
  • AA is self-adjoint when interpreted as an observable.
  • State and operator act on the same Hilbert space.
  • The second spectral moment is finite.
  • Discrete probabilities include every eigenspace and continuous formulas use the correct measure.
  • Trace expressions with unbounded operators are defined.
  • Statistical sample formulas assume stable, independent preparations unless a different model is stated.
  • Detector-noise decompositions assume the specified independence and bias conditions.

Variance and standard deviation are exact properties of a Born distribution. They summarize spread but do not determine the complete distribution. They can fail to exist for normalizable states with heavy spectral tails.

The variance of a non-self-adjoint operator can be defined in several inequivalent ways. This card concerns self-adjoint observables. For general operators, expressions such as

⟨(B−⟨B⟩)†(B−⟨B⟩)⟩\left\langle (B-\langle B\rangle)^\dagger (B-\langle B\rangle) \right\rangle

must be labeled explicitly rather than silently called the same variance.

  • Verify Var⁡(A)≥0\operatorname{Var}(A)\geq0 within numerical tolerance.
  • Check [ΔA]=[A][\Delta A]=[A] and [Var⁡(A)]=[A]2[\operatorname{Var}(A)]=[A]^2.
  • In an eigenstate or one-eigenspace mixture, the variance must vanish.
  • For A=αIA=\alpha I, the variance must vanish in every state.
  • Shifting AA by βI\beta I must leave the variance unchanged.
  • Scaling AA by α\alpha must scale variance by α2\alpha^2.
  • A bounded-spectrum result must obey the range bound.
  • A unitary basis change applied to both state and observable must preserve the result.
  • For a two-outcome law, compare with p(1−p)(a−b)2p(1-p)(a-b)^2.
  • When using the raw-moment form, compare with a centered calculation if cancellation is plausible.

Let

∣ψ⟩=α∣+⟩z+β∣−⟩z,∣α∣2+∣β∣2=1.\lvert\psi\rangle = \alpha\lvert+\rangle_z + \beta\lvert-\rangle_z, \qquad \lvert\alpha\rvert^2+\lvert\beta\rvert^2=1.

For

Sz=ℏ2σz,S_z = \frac{\hbar}{2}\sigma_z, ⟨Sz⟩=ℏ2(∣α∣2−∣β∣2),\langle S_z\rangle = \frac{\hbar}{2} \left( \lvert\alpha\rvert^2 - \lvert\beta\rvert^2 \right),

while Sz2=ℏ2I/4S_z^2=\hbar^2I/4. Hence

(ΔSz)2=ℏ2∣α∣2∣β∣2.(\Delta S_z)^2 = \hbar^2 \lvert\alpha\rvert^2 \lvert\beta\rvert^2.

The spread vanishes for either eigenstate and is maximal for equal outcome probabilities.

For the normalized density

p(x)=12πσexp⁡[−(x−x0)22σ2],p(x) = \frac{1}{\sqrt{2\pi}\sigma} \exp \left[ - \frac{(x-x_0)^2}{2\sigma^2} \right], ⟨X⟩=x0,(ΔX)2=σ2,ΔX=σ.\langle X\rangle=x_0, \qquad (\Delta X)^2=\sigma^2, \qquad \Delta X=\sigma.

The symbol σ\sigma in this parameterization is the standard deviation, not the variance.

Variance and Standard Deviation owns the derivation from the Born measure, the pure-state norm form, finite-moment conditions, zero-variance characterization, mixture decomposition, and statistical interpretation.

Variance and Covariance owns the underlying probability identities. Correlations and Covariance develops joint quantum moments, and General Uncertainty Relations owns the operator inequalities.

  • Writing ΔA=⟨A2⟩−⟨A⟩2\Delta A=\langle A^2\rangle-\langle A\rangle^2.
  • Confusing ⟨A⟩2\langle A\rangle^2 with ⟨A2⟩\langle A^2\rangle.
  • Treating zero expectation as zero variance.
  • Squaring each matrix element of AA instead of computing A2A^2.
  • Forgetting degeneracy or a continuous part of the spectrum.
  • Assuming a normalized state has a finite second moment.
  • Requiring ∣ψ⟩∈D(A2)\lvert\psi\rangle\in\mathcal D(A^2) when the norm form only requires ∣ψ⟩∈D(A)\lvert\psi\rangle\in\mathcal D(A).
  • Calling intrinsic spread an apparatus error or measurement disturbance.
  • Omitting the square when converting physical units.
  • Inferring a full probability law from only its mean and variance.
  • Hiding a materially negative numerical result with an absolute value.
  • Assuming every ensemble decomposition of a mixed state assigns the same within-state and between-state variances.
  • Applying ΔA‾=ΔA/N\Delta\overline A=\Delta A/\sqrt N to correlated or drifting trials without qualification.
  • J. J. Sakurai and J. Napolitano, Modern Quantum Mechanics, 3rd ed., Cambridge University Press, 2020, chs. 1 and 3.
  • R. Shankar, Principles of Quantum Mechanics, 2nd ed., Springer, 1994, chs. 4 and 9.
  • L. E. Ballentine, Quantum Mechanics: A Modern Development, 2nd ed., World Scientific, 2014, chs. 2 and 3.
  • M. Reed and B. Simon, Methods of Modern Mathematical Physics I: Functional Analysis, rev. ed., Academic Press, 1980, ch. VIII.

An observable has outcomes −1-1, 00, and 22 with probabilities 1/41/4, 1/21/2, and 1/41/4. Find its expectation, variance, and standard deviation.

Solution

The expectation is

⟨A⟩=−14+0+12=14.\langle A\rangle = - \frac14 + 0 + \frac12 = \frac14.

The second moment is

⟨A2⟩=14+0+1=54.\langle A^2\rangle = \frac14 + 0 + 1 = \frac54.

Therefore

Var⁡(A)=54−(14)2=1916,\operatorname{Var}(A) = \frac54 - \left(\frac14\right)^2 = \frac{19}{16},

and

ΔA=194.\Delta A = \frac{\sqrt{19}}{4}.

Let ρ1\rho_1 and ρ2\rho_2 be eigenstates of AA with eigenvalues a1a_1 and a2a_2. For

ρ=qρ1+(1−q)ρ2,\rho = q\rho_1 + (1-q)\rho_2,

show that the total variance is entirely between components.

Solution

Each component is supported in one eigenspace, so

Var⁡ρ1(A)=Var⁡ρ2(A)=0.\operatorname{Var}_{\rho_1}(A) = \operatorname{Var}_{\rho_2}(A) = 0.

The two-outcome formula gives

Var⁡ρ(A)=q(1−q)(a1−a2)2.\operatorname{Var}_\rho(A) = q(1-q)(a_1-a_2)^2.

This equals the variance of the component means, so the within-component term in the law of total variance vanishes.

An observable has single-shot standard deviation ΔA=3.0\Delta A=3.0 in the prepared state. Assuming independent trials, how many measurements are needed to make the standard deviation of the sample mean at most 0.030.03?

Solution

Use

ΔA‾=ΔAN.\Delta\overline A = \frac{\Delta A}{\sqrt N}.

The requirement is

3.0N≤0.03,\frac{3.0}{\sqrt N} \leq 0.03,

so N≥100\sqrt N\geq100 and therefore

N≥104.N\geq10^4.

This calculation assumes independent, identically distributed trials and does not include systematic detector error.