Skip to content

Variance and Covariance

Variance measures the spread of one random variable. Covariance measures how two random variables fluctuate together.

For a real random variable XX with mean

μX=E[X],\mu_X=\mathbb E[X],

the variance is

Var⁡(X)=E[(X−μX)2].\operatorname{Var}(X) = \mathbb E[(X-\mu_X)^2].

The standard deviation is

σX=Var⁡(X).\sigma_X = \sqrt{\operatorname{Var}(X)}.

In quantum mechanics, standard deviations become the ΔA\Delta A appearing in uncertainty relations. The operator-specific treatment is Variance and Standard Deviation; this page supplies the probability language underneath it.

Variance is defined only when the required expectations exist. A random variable can have a well-defined distribution but no finite variance.

For a discrete random variable with values xix_i and probabilities pip_i,

E[X]=∑ixipi,E[X2]=∑ixi2pi.\mathbb E[X] = \sum_i x_i p_i, \qquad \mathbb E[X^2] = \sum_i x_i^2p_i.

For a continuous random variable with density fX(x)f_X(x),

E[X]=∫−∞∞xfX(x) dx,E[X2]=∫−∞∞x2fX(x) dx,\mathbb E[X] = \int_{-\infty}^{\infty}x f_X(x)\,dx, \qquad \mathbb E[X^2] = \int_{-\infty}^{\infty}x^2 f_X(x)\,dx,

when the integrals converge in the required sense.

Finite normalization does not imply finite variance. A heavy-tailed distribution can have total probability one but divergent second moment.

Expanding the square gives the computational form

Var⁡(X)=E[X2]−E[X]2.\operatorname{Var}(X) = \mathbb E[X^2] - \mathbb E[X]^2.

Derivation:

E[(X−μX)2]=E[X2−2μXX+μX2]=E[X2]−2μXE[X]+μX2=E[X2]−μX2.\begin{aligned} \mathbb E[(X-\mu_X)^2] &= \mathbb E[X^2-2\mu_X X+\mu_X^2]\\ &= \mathbb E[X^2] -2\mu_X\mathbb E[X] +\mu_X^2\\ &= \mathbb E[X^2]-\mu_X^2. \end{aligned}

Variance is nonnegative:

Var⁡(X)≥0,\operatorname{Var}(X)\ge0,

because it is the expectation of a square. It has the squared units of XX, while σX\sigma_X has the same units as XX.

The standard deviation

σX=Var⁡(X)\sigma_X=\sqrt{\operatorname{Var}(X)}

is the most common quoted measure of spread because it has the same units as the measured quantity.

It is not the same thing as an experimental error. In probability theory it describes the spread of the distribution. In quantum mechanics, ΔA\Delta A describes the spread of ideal measurement outcomes predicted by a state, even before apparatus imperfections are considered.

Let XX take values +1+1 and −1-1 with

P(X=+1)=q,P(X=−1)=1−q.\mathbb P(X=+1)=q, \qquad \mathbb P(X=-1)=1-q.

Then

E[X]=2q−1.\mathbb E[X] = 2q-1.

Since X2=1X^2=1 always,

E[X2]=1.\mathbb E[X^2]=1.

Therefore

Var⁡(X)=1−(2q−1)2=4q(1−q).\operatorname{Var}(X) = 1-(2q-1)^2 = 4q(1-q).

The variance is zero when one outcome is certain and maximal when q=1/2q=1/2.

For a normal density with mean μ\mu and standard deviation ss,

f(x)=12π sexp⁡ ⁣[−(x−μ)22s2],f(x) = \frac{1}{\sqrt{2\pi}\,s} \exp\!\left[ - \frac{(x-\mu)^2}{2s^2} \right],

one has

E[X]=μ,Var⁡(X)=s2.\mathbb E[X]=\mu, \qquad \operatorname{Var}(X)=s^2.

Different wavefunction conventions sometimes use different width parameters. The variance is the reliable way to identify the physical width of the probability density.

For the broader Gaussian toolkit, including multivariate covariance matrices and Gaussian integrals, see Gaussian Distributions.

For two real random variables XX and YY on the same probability space, the covariance is

Cov⁡(X,Y)=E[(X−μX)(Y−μY)].\operatorname{Cov}(X,Y) = \mathbb E[(X-\mu_X)(Y-\mu_Y)].

Equivalently,

Cov⁡(X,Y)=E[XY]−E[X]E[Y].\operatorname{Cov}(X,Y) = \mathbb E[XY] - \mathbb E[X]\mathbb E[Y].

Covariance is positive when XX and YY tend to fluctuate in the same direction, negative when they tend to fluctuate in opposite directions, and zero when their linear fluctuation is absent.

The units of Cov⁡(X,Y)\operatorname{Cov}(X,Y) are the units of XX times the units of YY.

Covariance depends on the joint distribution of XX and YY, not only on their separate distributions.

For a joint density fX,Y(x,y)f_{X,Y}(x,y),

E[XY]=∫−∞∞∫−∞∞xy fX,Y(x,y) dx dy.\mathbb E[XY] = \int_{-\infty}^{\infty} \int_{-\infty}^{\infty} xy\,f_{X,Y}(x,y)\,dx\,dy.

The marginal densities

fX(x)=∫−∞∞fX,Y(x,y) dy,fY(y)=∫−∞∞fX,Y(x,y) dxf_X(x) = \int_{-\infty}^{\infty} f_{X,Y}(x,y)\,dy, \qquad f_Y(y) = \int_{-\infty}^{\infty} f_{X,Y}(x,y)\,dx

do not determine the covariance by themselves.

This point matters in quantum mechanics because incompatible observables generally do not come with a single classical joint distribution of pre-existing values; see Classical Probability versus Quantum Probability.

When σX>0\sigma_X>0 and σY>0\sigma_Y>0, the correlation coefficient is

ρXY=Cov⁡(X,Y)σXσY.\rho_{XY} = \frac{\operatorname{Cov}(X,Y)}{\sigma_X\sigma_Y}.

It is dimensionless and satisfies

−1≤ρXY≤1.-1\le \rho_{XY}\le 1.

The bound follows from the Cauchy–Schwarz inequality applied to the centered variables X−μXX-\mu_X and Y−μYY-\mu_Y.

Correlation rescales covariance so that variables with different units can be compared. It measures linear association, not all possible dependence.

If XX and YY are independent and have finite second moments, then

Cov⁡(X,Y)=0.\operatorname{Cov}(X,Y)=0.

The converse is false. For example, let XX be uniformly distributed on [−1,1][-1,1] and let

Y=X2.Y=X^2.

Then

E[X]=0,E[XY]=E[X3]=0,\mathbb E[X]=0, \qquad \mathbb E[XY] = \mathbb E[X^3]=0,

so Cov⁡(X,Y)=0\operatorname{Cov}(X,Y)=0. But YY is fully determined by XX, so they are not independent.

Zero covariance means no linear covariance, not no relationship.

For random variables X1,…,XnX_1,\ldots,X_n, the covariance matrix is

Σij=Cov⁡(Xi,Xj).\Sigma_{ij} = \operatorname{Cov}(X_i,X_j).

It is symmetric:

Σij=Σji.\Sigma_{ij}=\Sigma_{ji}.

It is also positive semidefinite. For any real vector a\mathbf a,

aTΣa=Var⁡(∑iaiXi)≥0.\mathbf a^T\Sigma\mathbf a = \operatorname{Var} \left( \sum_i a_iX_i \right) \ge0.

This property is the probability-theory ancestor of many positivity constraints in quantum mechanics, including uncertainty inequalities and positivity of density operators.

For a quantum observable AA in a state, the outcome variance is the classical variance of the Born-rule measurement distribution:

(ΔA)2=⟨A2⟩−⟨A⟩2.(\Delta A)^2 = \langle A^2\rangle - \langle A\rangle^2.

For two commuting observables with a joint measurement, covariance has the ordinary interpretation as covariance of two jointly measured outcomes.

For noncommuting observables, one must be more careful. A common symmetrized quantum covariance is

Cov⁡ψ(A,B)=12⟨{A~,B~}⟩,\operatorname{Cov}_\psi(A,B) = \frac12 \left\langle \{\widetilde A,\widetilde B\} \right\rangle,

where

A~=A−⟨A⟩I,B~=B−⟨B⟩I,\widetilde A=A-\langle A\rangle I, \qquad \widetilde B=B-\langle B\rangle I,

and

{A~,B~}=A~B~+B~A~.\{\widetilde A,\widetilde B\} = \widetilde A\widetilde B+\widetilde B\widetilde A.

This is not the same as assuming that AA and BB have simultaneous hidden classical values. It is an operator expression defined by the state and observables.

The general uncertainty relation includes both spread and noncommutativity. A basic form is

ΔA ΔB≥12∣⟨[A,B]⟩∣.\Delta A\,\Delta B \ge \frac12 \left\lvert \langle[A,B]\rangle \right\rvert.

The full discussion belongs in General Uncertainty Relations.

  • Treating variance as the mean absolute deviation.
  • Forgetting to subtract the square of the mean in E[X2]−E[X]2\mathbb E[X^2]-\mathbb E[X]^2.
  • Assuming zero mean implies zero variance.
  • Forgetting that variance has squared units.
  • Computing covariance from marginal distributions without a joint distribution.
  • Treating zero covariance as independence.
  • Reading quantum uncertainty as ordinary apparatus error.
  • Assigning covariance to noncommuting observables without specifying the measurement or operator convention.
  • W. Feller, An Introduction to Probability Theory and Its Applications, Volume I, 3rd ed., Wiley, 1968.
  • P. Billingsley, Probability and Measure, 3rd ed., Wiley, 1995.
  • R. Durrett, Probability: Theory and Examples, 5th ed., Cambridge University Press, 2019.
  • R. Shankar, Principles of Quantum Mechanics, 2nd ed., Springer, 1994.
  • J. J. Sakurai and J. Napolitano, Modern Quantum Mechanics, 3rd ed., Cambridge University Press, 2020.
  1. A random variable takes values 00 and 11 with P(X=1)=q\mathbb P(X=1)=q. Compute its variance.
Solution

Here

E[X]=q,E[X2]=q,\mathbb E[X]=q, \qquad \mathbb E[X^2]=q,

because X2=XX^2=X for values 00 and 11. Therefore

Var⁡(X)=q−q2=q(1−q).\operatorname{Var}(X) = q-q^2 = q(1-q).
  1. Let XX take values +1+1 and −1-1 with equal probability. Compute E[X]\mathbb E[X], E[X2]\mathbb E[X^2], and Var⁡(X)\operatorname{Var}(X).
Solution

The mean is

E[X]=12(1)+12(−1)=0.\mathbb E[X] = \frac12(1)+\frac12(-1) = 0.

Since X2=1X^2=1 always,

E[X2]=1.\mathbb E[X^2]=1.

Thus

Var⁡(X)=1−02=1.\operatorname{Var}(X) = 1-0^2 = 1.
  1. Suppose Var⁡(X)=4\operatorname{Var}(X)=4, Var⁡(Y)=9\operatorname{Var}(Y)=9, and Cov⁡(X,Y)=3\operatorname{Cov}(X,Y)=3. Find the correlation coefficient.
Solution

The standard deviations are σX=2\sigma_X=2 and σY=3\sigma_Y=3. Therefore

ρXY=32⋅3=12.\rho_{XY} = \frac{3}{2\cdot3} = \frac12.
  1. Show that adding a constant does not change variance: Var⁡(X+c)=Var⁡(X)\operatorname{Var}(X+c)=\operatorname{Var}(X).
Solution

The mean of X+cX+c is μX+c\mu_X+c. Therefore

X+c−(μX+c)=X−μX.X+c-(\mu_X+c)=X-\mu_X.

So

Var⁡(X+c)=E[(X−μX)2]=Var⁡(X).\operatorname{Var}(X+c) = \mathbb E[(X-\mu_X)^2] = \operatorname{Var}(X).
  1. For centered variables with E[X]=E[Y]=0\mathbb E[X]=\mathbb E[Y]=0, what does Cov⁡(X,Y)\operatorname{Cov}(X,Y) reduce to?
Solution

Using

Cov⁡(X,Y)=E[XY]−E[X]E[Y],\operatorname{Cov}(X,Y) = \mathbb E[XY]-\mathbb E[X]\mathbb E[Y],

and the centered assumptions, one gets

Cov⁡(X,Y)=E[XY].\operatorname{Cov}(X,Y) = \mathbb E[XY].