Skip to content

Expectation Values

An expectation value is a probability-weighted average. It summarizes the center of a distribution, but it need not be a possible outcome, a typical outcome, or the result of any one trial.

In quantum mechanics, the same mathematical operation averages the outcomes of a specified observable in a specified state. The operator formula is not a new kind of average; the Born rule supplies a probability measure, and the expectation is its first moment.

Let (Ω,F,P)(\Omega,\mathcal F,\mathbb P) be a probability space and let X:Ω→RX:\Omega\to\mathbb R be a random variable. Its expectation is

E[X]=∫ΩX(ω) dP(ω),\mathbb E[X] = \int_\Omega X(\omega)\,d\mathbb P(\omega),

provided the integral is well defined.

A finite expectation exists when XX is integrable:

E[∣X∣]<∞.\mathbb E[\lvert X\rvert] <\infty.

This absolute-integrability condition prevents an undefined cancellation between an infinite positive part and an infinite negative part. A nonnegative random variable can have the extended expectation +∞+\infty, but then it does not have a finite mean.

The underlying vocabulary of outcomes, events, and measures is reviewed in Probability Spaces, Light Version.

If XX takes countably many values xix_i with probabilities pip_i, then

E[X]=∑ixipi,∑ipi=1.\mathbb E[X] = \sum_i x_i p_i, \qquad \sum_i p_i=1.

Absolute integrability means

∑i∣xi∣pi<∞.\sum_i \lvert x_i\rvert p_i <\infty.

If XX has probability density fX(x)f_X(x), then

E[X]=∫−∞∞xfX(x) dx,\mathbb E[X] = \int_{-\infty}^{\infty} x f_X(x)\,dx,

with

∫−∞∞fX(x) dx=1.\int_{-\infty}^{\infty} f_X(x)\,dx=1.

The sum and density formulas are special representations of the same integral over the probability space. Some probability laws are neither purely discrete nor described by an ordinary density; the measure definition covers mixtures and singular distributions as well.

See Probability Densities for normalization, units, and changes of variables.

Whenever the relevant expectations exist, expectation is linear:

E[aX+bY]=a E[X]+b E[Y].\mathbb E[aX+bY] = a\,\mathbb E[X] +b\,\mathbb E[Y].

Independence is not required for linearity.

Constants pass through the average:

E[c]=c.\mathbb E[c]=c.

Expectation is positive and monotone:

X≥0⟹E[X]≥0,X≤Y⟹E[X]≤E[Y].\begin{aligned} X\ge0 &\Longrightarrow \mathbb E[X]\ge0,\\ X\le Y &\Longrightarrow \mathbb E[X]\le\mathbb E[Y]. \end{aligned}

For an event AA, let 1A\mathbf 1_A be its indicator. Then

E[1A]=P(A).\mathbb E[\mathbf 1_A] = \mathbb P(A).

This identity lets probabilities be manipulated as expectations and is often useful in counting arguments and error bounds.

The law of the unconscious statistician, often abbreviated LOTUS, states that

E[g(X)]=∫Ωg(X(ω)) dP(ω).\mathbb E[g(X)] = \int_\Omega g(X(\omega))\,d\mathbb P(\omega).

If XX has density fXf_X, this becomes

E[g(X)]=∫−∞∞g(x)fX(x) dx.\mathbb E[g(X)] = \int_{-\infty}^{\infty} g(x)f_X(x)\,dx.

For a discrete variable,

E[g(X)]=∑ig(xi)pi.\mathbb E[g(X)] = \sum_i g(x_i)p_i.

One does not need to derive the probability distribution of Y=g(X)Y=g(X) merely to compute its expectation. The integrability condition becomes E[∣g(X)∣]<∞\mathbb E[\lvert g(X)\rvert]<\infty.

The nnth raw moment is

μn′=E[Xn],\mu_n' = \mathbb E[X^n],

when it exists. The first raw moment is the mean

μ=E[X].\mu=\mathbb E[X].

The nnth central moment is

μn=E[(X−μ)n].\mu_n = \mathbb E[(X-\mu)^n].

In particular, the variance is

Var⁡(X)=E[(X−μ)2]=E[X2]−μ2.\operatorname{Var}(X) = \mathbb E[(X-\mu)^2] = \mathbb E[X^2]-\mu^2.

The second expression requires the second moment to be finite. A distribution can have a finite mean but infinite variance, or be normalized while having no finite mean at all.

Variance, covariance, and correlation have a canonical treatment in Variance and Covariance. Moment-generating Fourier methods are discussed in Characteristic Functions.

If gg is convex and the expectations exist, then

g(E[X])≤E[g(X)].g\left(\mathbb E[X]\right) \le \mathbb E[g(X)].

For the convex function g(x)=x2g(x)=x^2, this gives

E[X]2≤E[X2],\mathbb E[X]^2 \le \mathbb E[X^2],

which is equivalent to nonnegative variance.

Equality in a strictly convex Jensen inequality requires XX to be constant almost surely. The inequality is a powerful way to compare nonlinear functions of averages with averages of nonlinear functions.

For jointly distributed XX and YY,

E[g(X,Y)]=∫g(x,y) dPX,Y(x,y).\mathbb E[g(X,Y)] = \int g(x,y)\,d\mathbb P_{X,Y}(x,y).

If a joint density exists,

E[g(X,Y)]=∫ ⁣ ⁣∫g(x,y)fX,Y(x,y) dx dy.\begin{aligned} \mathbb E[g(X,Y)] &= \int\!\!\int \\ &\quad g(x,y)f_{X,Y}(x,y)\,dx\,dy. \end{aligned}

If XX and YY are independent and XYXY is integrable, then

E[XY]=E[X]E[Y].\mathbb E[XY] = \mathbb E[X]\mathbb E[Y].

The converse is false: zero covariance or factorization of one product moment does not generally imply independence.

Conditional expectation and the law of total expectation are developed in Conditional Probability.

Three notions of center answer different questions:

  • the mean minimizes expected squared error;
  • a median minimizes expected absolute error;
  • a mode is a most probable value or density maximum.

They coincide for some symmetric unimodal distributions but not in general. For an equal-probability variable taking values −1-1 and +1+1,

E[X]=0,\mathbb E[X]=0,

even though zero never occurs. An expectation can lie between discrete outcomes or in a low-density region of a skewed distribution.

The expectation is a property of a probability law. An experimental or numerical estimate uses a sample. For independent identically distributed values X1,…,XNX_1,\ldots,X_N, define

X‾N=1N∑j=1NXj.\overline X_N = \frac{1}{N} \sum_{j=1}^{N}X_j.

If E[Xj]=μ\mathbb E[X_j]=\mu, then

E[X‾N]=μ.\mathbb E[\overline X_N]=\mu.

If the variance is finite and equal to σ2\sigma^2, independence gives

Var⁡(X‾N)=σ2N.\operatorname{Var}(\overline X_N) = \frac{\sigma^2}{N}.

The standard error therefore scales as σ/N\sigma/\sqrt N. This is a statement about fluctuations of the estimator, not about the spread of individual outcomes. Correlated samples require an effective sample size rather than the naive NN.

For computational estimation, see Monte Carlo Basics.

Let AA be a self-adjoint observable with spectral measure PA(B)P^A(\mathcal B) for Borel sets B⊆R\mathcal B\subseteq\mathbb R. A normalized state ∣ψ⟩\lvert\psi\rangle defines the outcome probability measure

μψA(B)=⟨ψ∣PA(B)∣ψ⟩.\mu_\psi^A(\mathcal B) = \langle\psi\vert P^A(\mathcal B) \lvert\psi\rangle.

The expectation of the measurement outcomes is

⟨A⟩ψ=∫Ra dμψA(a),\langle A\rangle_\psi = \int_{\mathbb R} a\,d\mu_\psi^A(a),

provided

∫R∣a∣ dμψA(a)<∞.\int_{\mathbb R} \lvert a\rvert\,d\mu_\psi^A(a) <\infty.

For a discrete spectral decomposition

A=∑aaPa,A=\sum_a aP_a,

the Born probabilities are

p(a)=⟨ψ∣Pa∣ψ⟩,p(a) = \langle\psi\vert P_a\lvert\psi\rangle,

and

⟨A⟩ψ=∑aa p(a).\langle A\rangle_\psi = \sum_a a\,p(a).

Degeneracy is handled by the projector PaP_a onto the entire eigenspace. Nothing requires choosing a preferred basis inside that eigenspace.

The measurement interpretation belongs in Quantum Expectation Values.

When ∣ψ⟩\lvert\psi\rangle lies in the domain of AA,

⟨A⟩ψ=⟨ψ∣A∣ψ⟩.\langle A\rangle_\psi = \langle\psi\vert A\lvert\psi\rangle.

For a density operator ρ\rho in finite dimensions, or whenever the trace is well defined,

⟨A⟩ρ=Tr⁡(ρA).\langle A\rangle_\rho = \operatorname{Tr}(\rho A).

This includes classical mixtures and reduced states. If

ρ=∑jwj∣ψj⟩⟨ψj∣,\rho=\sum_j w_j \lvert\psi_j\rangle\langle\psi_j\rvert,

then linearity gives

⟨A⟩ρ=∑jwj⟨ψj∣A∣ψj⟩.\langle A\rangle_\rho = \sum_j w_j \langle\psi_j\vert A\lvert\psi_j\rangle.

Different ensemble decompositions of the same ρ\rho give the same expectation because the trace depends only on ρ\rho.

For a normalized position wavefunction,

⟨x⟩=∫−∞∞x∣ψ(x)∣2 dx.\langle x\rangle = \int_{-\infty}^{\infty} x\lvert\psi(x)\rvert^2\,dx.

In momentum representation,

⟨p⟩=∫−∞∞p∣ϕ(p)∣2 dp.\langle p\rangle = \int_{-\infty}^{\infty} p\lvert\phi(p)\rvert^2\,dp.

When ψ\psi lies in the momentum-operator domain, the same quantity is

⟨p⟩=∫−∞∞ψ(x)∗(−iℏddx)ψ(x) dx.\langle p\rangle = \int_{-\infty}^{\infty} \psi(x)^* \left( -i\hbar\frac{d}{dx} \right) \psi(x)\,dx.

The equality follows from the unitary Fourier transform. It also depends on the derivative and boundary conditions being meaningful; the formal differential expression alone is not enough.

The expectation of a self-adjoint observable is real whenever it exists. In the operator-domain form,

⟨ψ∣A∣ψ⟩∗=⟨ψ∣A∣ψ⟩.\langle\psi\vert A\lvert\psi\rangle^* = \langle\psi\vert A\lvert\psi\rangle.

Expectation is linear even when two bounded observables do not commute:

⟨aA+bB⟩=a⟨A⟩+b⟨B⟩.\langle aA+bB\rangle = a\langle A\rangle +b\langle B\rangle.

This does not mean that outcomes of separate incompatible measurements can be added trial by trial to obtain the outcome distribution of A+BA+B. The observable A+BA+B has its own spectral measure. Linearity concerns the first moment, not equality of measurement protocols or full probability laws.

For unbounded operators, common domains and integrability must be checked before applying the formula.

For a self-adjoint AA, the spectral expectation can exist when

∫∣a∣ dμψA(a)<∞.\int \lvert a\rvert\,d\mu_\psi^A(a) <\infty.

The vector expression A∣ψ⟩A\lvert\psi\rangle requires the stronger condition

∫a2 dμψA(a)<∞,\int a^2\,d\mu_\psi^A(a) <\infty,

which characterizes ∣ψ⟩∈D(A)\lvert\psi\rangle\in\mathcal D(A). A finite variance also requires this second moment.

Thus a normalized state need not have finite expectation for every observable, and a finite first moment does not automatically justify every operator manipulation. Domain statements are part of the mathematics, not a technical afterthought.

In the Schrödinger picture,

⟨A⟩t=⟨ψ(t)∣A(t)∣ψ(t)⟩.\langle A\rangle_t = \langle\psi(t)\vert A(t) \lvert\psi(t)\rangle.

Under suitable domain assumptions,

ddt⟨A⟩=iℏ⟨[H,A]⟩+⟨∂A∂t⟩.\frac{d}{dt}\langle A\rangle = \frac{i}{\hbar} \langle[H,A]\rangle +\left\langle \frac{\partial A}{\partial t} \right\rangle.

If AA has no explicit time dependence and commutes with HH, its expectation is conserved. Conservation of the expectation is weaker than certainty of a fixed outcome; the full distribution can matter.

For a continuous probability density:

  1. verify or enforce normalization;
  2. check tail behavior before truncating the domain;
  3. evaluate positive and negative contributions with adequate precision;
  4. refine the quadrature grid and integration range independently;
  5. compare direct and transformed representations when both are available.

For a discretized wavefunction with quadrature weights wjw_j,

⟨A⟩≈∑jwjψj∗(Aψ)j.\langle A\rangle \approx \sum_j w_j\psi_j^*(A\psi)_j.

Ignoring the weights changes the inner product. In nonorthogonal bases, the overlap or mass matrix must be included.

  • Treating the expectation as the outcome of one trial.
  • Assuming the mean must be an allowed value.
  • Confusing mean, median, and mode.
  • Omitting absolute-integrability checks.
  • Inferring independence from E[XY]=E[X]E[Y]\mathbb E[XY]=\mathbb E[X]\mathbb E[Y] alone.
  • Applying LOTUS without checking the integrability of g(X)g(X).
  • Confusing the standard deviation of outcomes with the standard error of a sample mean.
  • Writing ⟨ψ∣A∣ψ⟩\langle\psi\vert A\lvert\psi\rangle for an unbounded observable without checking the domain.
  • Assuming normalization of a state guarantees finite energy or variance.
  • Treating linearity of quantum expectation as an outcome-by-outcome rule for incompatible measurements.
  • Forgetting quadrature or overlap weights in numerical averages.
  1. Let XX be uniform on [−1,1][-1,1]. Compute E[X]\mathbb E[X] and E[X2]\mathbb E[X^2] using LOTUS.
Solution

The density is fX(x)=1/2f_X(x)=1/2 on [−1,1][-1,1]. Therefore,

E[X]=12∫−11x dx=0\mathbb E[X] = \frac{1}{2} \int_{-1}^{1}x\,dx =0

by odd symmetry. Also,

E[X2]=12∫−11x2 dx=13.\begin{aligned} \mathbb E[X^2] &= \frac{1}{2} \int_{-1}^{1}x^2\,dx\\ &= \frac{1}{3}. \end{aligned}

Hence Var⁡(X)=1/3\operatorname{Var}(X)=1/3.

  1. Let AA and BB be events. Use indicator variables to show

    P(A∪B)=P(A)+P(B)−P(A∩B).\begin{aligned} \mathbb P(A\cup B) &= \mathbb P(A)+\mathbb P(B)\\ &\quad-\mathbb P(A\cap B). \end{aligned}
Solution

Pointwise,

1A∪B=1A+1B−1A∩B.\mathbf 1_{A\cup B} = \mathbf 1_A+\mathbf 1_B-\mathbf 1_{A\cap B}.

Take expectations and use linearity:

P(A∪B)=E[1A∪B]=E[1A]+E[1B]−E[1A∩B]=P(A)+P(B)−P(A∩B).\begin{aligned} \mathbb P(A\cup B) &= \mathbb E[\mathbf 1_{A\cup B}]\\ &= \mathbb E[\mathbf 1_A] +\mathbb E[\mathbf 1_B] -\mathbb E[\mathbf 1_{A\cap B}]\\ &= \mathbb P(A)+\mathbb P(B)-\mathbb P(A\cap B). \end{aligned}

No independence assumption is needed.

  1. For

    ∣ψ⟩=cos⁡(θ2)∣0⟩+eiφsin⁡(θ2)∣1⟩,\lvert\psi\rangle = \cos\left(\frac{\theta}{2}\right)\lvert0\rangle + e^{i\varphi} \sin\left(\frac{\theta}{2}\right)\lvert1\rangle,

    compute the expectation and variance of σz\sigma_z, where the eigenvalues of ∣0⟩\lvert0\rangle and ∣1⟩\lvert1\rangle are +1+1 and −1-1.

Solution

The two probabilities are

p+=cos⁡2(θ2),p−=sin⁡2(θ2).p_+ = \cos^2\left(\frac{\theta}{2}\right), \qquad p_- = \sin^2\left(\frac{\theta}{2}\right).

Therefore,

⟨σz⟩=p+−p−=cos⁡θ.\begin{aligned} \langle\sigma_z\rangle &= p_+-p_-\\ &= \cos\theta. \end{aligned}

Since σz2=I\sigma_z^2=I,

Var⁡(σz)=⟨σz2⟩−⟨σz⟩2=1−cos⁡2θ=sin⁡2θ.\begin{aligned} \operatorname{Var}(\sigma_z) &= \langle\sigma_z^2\rangle -\langle\sigma_z\rangle^2\\ &= 1-\cos^2\theta\\ &= \sin^2\theta. \end{aligned}

The phase φ\varphi does not affect a σz\sigma_z measurement.

  1. The density

    fX(x)=1x2,x≥1,f_X(x) = \frac{1}{x^2}, \qquad x\ge1,

    is zero otherwise. Show that it is normalized but has no finite expectation.

Solution

Normalization holds because

∫1∞dxx2=[−1x]1∞=1.\int_1^\infty \frac{dx}{x^2} = \left[ -\frac{1}{x} \right]_1^\infty =1.

The expectation is

E[X]=∫1∞x1x2 dx=∫1∞dxx.\mathbb E[X] = \int_1^\infty x\frac{1}{x^2}\,dx = \int_1^\infty \frac{dx}{x}.

The logarithmic integral diverges, so E[X]=+∞\mathbb E[X]=+\infty. Normalization alone does not guarantee a finite first moment.

  • P. Billingsley, Probability and Measure, 3rd ed., Wiley, 1995.
  • G. Grimmett and D. Stirzaker, Probability and Random Processes, 3rd ed., Oxford University Press, 2001.
  • W. Feller, An Introduction to Probability Theory and Its Applications, Vol. I, 3rd ed., Wiley, 1968.
  • L. E. Ballentine, Quantum Mechanics: A Modern Development, World Scientific, 1998.
  • A. S. Holevo, Probabilistic and Statistical Aspects of Quantum Theory, 2nd ed., Edizioni della Normale, 2011.
  • P. Busch, P. Lahti, J.-P. Pellonpää, and K. Ylinen, Quantum Measurement, Springer, 2016.