Skip to content

Probability Densities

A probability density is not a probability at a point. It is a function that must be integrated over a region to give a probability.

For a real-valued random variable XX, a density fX(x)f_X(x) satisfies

P(X∈Δ)=∫ΔfX(x) dx\mathbb P(X\in\Delta) = \int_\Delta f_X(x)\,dx

for allowed sets of values Δ\Delta. The density tells how probability is distributed with respect to the reference measure dxdx.

This page explains the mathematical rules for densities. The physical Born-rule page for continuous observables is Born Rule for Continuous Spectra, and the wave-mechanics page is Wavefunctions and Probability Density.

Let XX be a real-valued random variable. A probability density for XX with respect to dxdx is a nonnegative function fXf_X such that

PX(Δ)=∫ΔfX(x) dx.\mathbb P_X(\Delta) = \int_\Delta f_X(x)\,dx.

Here PX\mathbb P_X is the distribution, or law, of XX:

PX(Δ)=P(X∈Δ).\mathbb P_X(\Delta) = \mathbb P(X\in\Delta).

The phrase “with respect to dxdx” matters. A density is always density relative to a chosen measure. In elementary real-variable examples the reference measure is usually length dxdx. In three-dimensional position space it is volume d3rd^3r. In spherical coordinates the same volume measure becomes

d3r=r2sin⁡θ dr dθ dϕ.d^3r = r^2\sin\theta\,dr\,d\theta\,d\phi.

Changing the reference measure changes the function that deserves to be called the density.

A density must satisfy

fX(x)≥0f_X(x)\ge0

and

∫−∞∞fX(x) dx=1\int_{-\infty}^{\infty} f_X(x)\,dx=1

for a real-valued variable on the full line. More generally, over a domain DD,

∫Df(x) dx=1.\int_D f(x)\,dx=1.

For a three-dimensional density ρ(r)\rho(\mathbf r),

∫R3ρ(r) d3r=1.\int_{\mathbb R^3} \rho(\mathbf r)\,d^3r = 1.

Normalization is a probability statement, not an arbitrary scale convention. In quantum mechanics, normalizing a wavefunction ensures that the total probability of finding the particle somewhere in the configuration space is one.

For an ordinary continuous density,

P(X=x0)=0.\mathbb P(X=x_0)=0.

Instead, probabilities come from intervals:

P(a≤X≤b)=∫abfX(x) dx.\mathbb P(a\le X\le b) = \int_a^b f_X(x)\,dx.

For a small interval centered at x0x_0,

P(x0−Δx2≤X≤x0+Δx2)≈fX(x0)Δx\mathbb P \left( x_0-\frac{\Delta x}{2} \le X\le x_0+\frac{\Delta x}{2} \right) \approx f_X(x_0)\Delta x

when fXf_X varies slowly over the interval.

This is the operational meaning of density: it gives the leading probability per unit width for sufficiently small bins.

A density carries inverse units of its variable. If XX has units of length, then fXf_X has units of inverse length. This makes

fX(x) dxf_X(x)\,dx

dimensionless.

For a one-dimensional wavefunction,

ρ(x)=∣ψ(x)∣2\rho(x) = \lvert\psi(x)\rvert^2

has units of inverse length, so ψ(x)\psi(x) has units of length−1/2^{-1/2}. In three dimensions, ∣ψ(r)∣2\lvert\psi(\mathbf r)\rvert^2 has units of inverse volume, and ψ(r)\psi(\mathbf r) has units of length−3/2^{-3/2}.

Units are a useful guardrail when changing variables. If p=ℏkp=\hbar k, then a momentum density and a wave-number density cannot be the same numerical function without a Jacobian.

Every real-valued random variable has a cumulative distribution function

FX(x)=P(X≤x).F_X(x) = \mathbb P(X\le x).

If XX has a density fXf_X, then

FX(x)=∫−∞xfX(t) dt.F_X(x) = \int_{-\infty}^{x} f_X(t)\,dt.

At points where FXF_X is differentiable,

fX(x)=dFXdx.f_X(x) = \frac{dF_X}{dx}.

This relation is often the easiest way to derive a density from a known distribution function. It also shows why a density can be changed at isolated points without changing any probabilities: intervals see integrals, not point values.

Suppose Y=g(X)Y=g(X) and gg is one-to-one and differentiable. Conservation of probability gives

fY(y) dy=fX(x) dx.f_Y(y)\,dy = f_X(x)\,dx.

Since x=g−1(y)x=g^{-1}(y),

fY(y)=fX(g−1(y))∣ddyg−1(y)∣.f_Y(y) = f_X(g^{-1}(y)) \left\lvert \frac{d}{dy}g^{-1}(y) \right\rvert.

Equivalently,

fY(y)=fX(x)∣g′(x)∣with y=g(x).f_Y(y) = \frac{f_X(x)}{\lvert g'(x)\rvert} \qquad \text{with } y=g(x).

The absolute value is essential. Densities must stay nonnegative even when a coordinate map reverses orientation.

If the map is many-to-one, sum over all roots xix_i satisfying g(xi)=yg(x_i)=y:

fY(y)=∑xi:g(xi)=yfX(xi)∣g′(xi)∣.f_Y(y) = \sum_{x_i:g(x_i)=y} \frac{f_X(x_i)}{\lvert g'(x_i)\rvert}.

This formula is the source of many square-root singularities in energy, radial, and density-of-states calculations.

For a vector random variable X=(X1,…,Xn)\mathbf X=(X_1,\ldots,X_n), a joint density fX(x)f_{\mathbf X}(\mathbf x) satisfies

P(X∈R)=∫RfX(x) dnx.\mathbb P(\mathbf X\in R) = \int_R f_{\mathbf X}(\mathbf x)\,d^n x.

If y=g(x)\mathbf y=g(\mathbf x) is an invertible differentiable coordinate change, then

fY(y)=fX(g−1(y))∣det⁡∂x∂y∣.f_{\mathbf Y}(\mathbf y) = f_{\mathbf X}(g^{-1}(\mathbf y)) \left\lvert \det \frac{\partial \mathbf x}{\partial \mathbf y} \right\rvert.

The determinant is the multivariable Jacobian. Omitting it changes probabilities.

For example, in three-dimensional position space,

d3r=r2sin⁡θ dr dθ dϕ.d^3r = r^2\sin\theta\,dr\,d\theta\,d\phi.

If ρ(r)\rho(\mathbf r) is the density per unit volume, then the probability in a region is

∫ρ(r,θ,ϕ)r2sin⁡θ dr dθ dϕ,\int \rho(r,\theta,\phi) r^2\sin\theta\,dr\,d\theta\,d\phi,

not just the integral of ρ\rho over dr dθ dϕdr\,d\theta\,d\phi.

For two continuous random variables with joint density fX,Y(x,y)f_{X,Y}(x,y), the marginal density of XX is obtained by integrating out YY:

fX(x)=∫−∞∞fX,Y(x,y) dy.f_X(x) = \int_{-\infty}^{\infty} f_{X,Y}(x,y)\,dy.

Similarly,

fY(y)=∫−∞∞fX,Y(x,y) dx.f_Y(y) = \int_{-\infty}^{\infty} f_{X,Y}(x,y)\,dx.

The joint density contains correlation information that the marginals do not. Two different joint densities can have the same fXf_X and fYf_Y but different covariance; see Variance and Covariance.

Conditional densities are developed in Conditional Probability, but the basic formula, when fY(y)>0f_Y(y)>0, is

fX∣Y(x∣y)=fX,Y(x,y)fY(y).f_{X\mid Y}(x\mid y) = \frac{f_{X,Y}(x,y)}{f_Y(y)}.

This formula belongs to classical probability. Quantum measurement update has analogies to conditioning, but it is not just this formula applied to pre-existing values of all observables.

In one-dimensional wave mechanics, a normalized position-space wavefunction has

∫−∞∞∣ψ(x)∣2 dx=1.\int_{-\infty}^{\infty} \lvert\psi(x)\rvert^2\,dx = 1.

The position density is

ρ(x)=∣ψ(x)∣2.\rho(x) = \lvert\psi(x)\rvert^2.

Thus

P(X∈[a,b])=∫ab∣ψ(x)∣2 dx.\mathbb P(X\in[a,b]) = \int_a^b \lvert\psi(x)\rvert^2\,dx.

The wavefunction ψ\psi is a probability amplitude. The density is its squared magnitude. The phase of ψ\psi is invisible in ρ(x)\rho(x) at one instant, but it still matters for interference, current, and momentum-space structure.

In momentum representation, a normalized momentum-space wavefunction ϕ(p)\phi(p) gives

P(P∈[p1,p2])=∫p1p2∣ϕ(p)∣2 dp.\mathbb P(P\in[p_1,p_2]) = \int_{p_1}^{p_2} \lvert\phi(p)\rvert^2\,dp.

The same state can have different density functions in different representations because the measured variable and reference measure have changed.

A common source of mistakes is the difference between a density per unit volume and a density per unit radius.

For a three-dimensional wavefunction ψ(r)\psi(\mathbf r),

P(r∈R)=∫R∣ψ(r)∣2 d3r.\mathbb P(\mathbf r\in R) = \int_R \lvert\psi(\mathbf r)\rvert^2\,d^3r.

If the state is spherically symmetric, then the radial probability density pr(r)p_r(r) is defined so that

P(r∈[a,b])=∫abpr(r) dr.\mathbb P(r\in[a,b]) = \int_a^b p_r(r)\,dr.

The angular integration gives

pr(r)=4πr2∣ψ(r)∣2p_r(r) = 4\pi r^2\lvert\psi(r)\rvert^2

for a spherically symmetric wavefunction. The extra r2r^2 is not optional; it is the volume element.

For angular-momentum eigenstates written as Rnℓ(r)Yℓm(θ,ϕ)R_{n\ell}(r)Y_{\ell m}(\theta,\phi) with normalized spherical harmonics, the radial density is

pr(r)=r2∣Rnℓ(r)∣2.p_r(r) = r^2\lvert R_{n\ell}(r)\rvert^2.

The convention has changed because RnℓR_{n\ell} is separated from a normalized angular factor.

Some distributions are not represented by ordinary functions. A point mass at x=ax=a can be written formally as

δ(x−a)\delta(x-a)

inside integrals, meaning

∫−∞∞φ(x)δ(x−a) dx=φ(a).\int_{-\infty}^{\infty} \varphi(x)\delta(x-a)\,dx = \varphi(a).

This is useful notation, especially in quantum mechanics, but the delta function is not an ordinary probability density function. The distribution-theory treatment is Delta Function and Distributions.

Mixed distributions can have both a discrete part and a continuous-density part. For example, a detector model may include an atom at “no click” plus a continuous density over measured positions. In such cases, a single ordinary density with respect to dxdx is not the whole probability law.

  • Treating f(x0)f(x_0) as P(X=x0)\mathbb P(X=x_0).
  • Forgetting the Jacobian when changing variables.
  • Omitting volume elements such as r2sin⁡θr^2\sin\theta in spherical coordinates.
  • Comparing densities in different variables as if they had the same units.
  • Assuming every distribution has an ordinary density.
  • Confusing ∣ψ∣2\lvert\psi\rvert^2 with the wavefunction itself.
  • Inferring the full quantum state from a single position probability density.
  • P. Billingsley, Probability and Measure, 3rd ed., Wiley, 1995.
  • R. Durrett, Probability: Theory and Examples, 5th ed., Cambridge University Press, 2019.
  • W. Feller, An Introduction to Probability Theory and Its Applications, Volume II, 2nd ed., Wiley, 1971.
  • G. B. Folland, Real Analysis: Modern Techniques and Their Applications, 2nd ed., Wiley, 1999.
  • R. Shankar, Principles of Quantum Mechanics, 2nd ed., Springer, 1994.
  1. Let f(x)=Ce−axf(x)=Ce^{-ax} for x≥0x\geq0 and f(x)=0f(x)=0 for x<0x\lt0, with a>0a\gt0. Find CC.
Solution

Normalize:

1=∫0∞Ce−ax dx=Ca.1 = \int_0^\infty Ce^{-ax}\,dx = \frac{C}{a}.

Thus

C=a.C=a.
  1. If XX has density fX(x)f_X(x) and Y=2XY=2X, find fY(y)f_Y(y).
Solution

Here x=y/2x=y/2, so

fY(y)=fX(y/2)∣ddyy2∣=12fX(y/2).f_Y(y) = f_X(y/2) \left\lvert \frac{d}{dy}\frac{y}{2} \right\rvert = \frac12 f_X(y/2).

The factor 1/21/2 keeps the total probability normalized.

  1. Let XX be uniformly distributed on [−1,1][-1,1], and let Y=X2Y=X^2. Find the density of YY.
Solution

The density of XX is fX(x)=1/2f_X(x)=1/2 on [−1,1][-1,1]. For 0<y<10\lt y\lt1, the roots of x2=yx^2=y are x=yx=\sqrt y and x=−yx=-\sqrt y. Since g′(x)=2xg'(x)=2x,

fY(y)=1/22y+1/22y=12y.f_Y(y) = \frac{1/2}{2\sqrt y} + \frac{1/2}{2\sqrt y} = \frac{1}{2\sqrt y}.

The density is supported on 0≤y≤10\le y\le1. The singularity at y=0y=0 is integrable.

  1. A three-dimensional spherically symmetric wavefunction is normalized as ψ(r)=Ae−r/a\psi(r)=Ae^{-r/a} with a>0a>0. Find AA for a real positive normalization.
Solution

Use the volume element:

1=4πA2∫0∞r2e−2r/a dr.1 = 4\pi A^2 \int_0^\infty r^2e^{-2r/a}\,dr.

The integral is

∫0∞r2e−λr dr=2λ3,λ=2a.\int_0^\infty r^2e^{-\lambda r}\,dr = \frac{2}{\lambda^3}, \qquad \lambda=\frac{2}{a}.

Thus

∫0∞r2e−2r/a dr=a34.\int_0^\infty r^2e^{-2r/a}\,dr = \frac{a^3}{4}.

So

1=πA2a3,A=1πa3.1 = \pi A^2a^3, \qquad A = \frac{1}{\sqrt{\pi a^3}}.
  1. Explain why changing the value of a continuous density at one isolated point does not change any interval probabilities.
Solution

Interval probabilities are integrals of the density. Changing a function at one point changes its integral by zero with respect to ordinary length measure. Therefore the probability of every interval remains the same.