Skip to content

Relative Entropy

Relative entropy compares one probability distribution with another. In probability and statistics it is usually called the Kullback–Leibler divergence, or KL divergence.

For two discrete probability distributions PP and QQ on the same outcome set, with probabilities pxp_x and qxq_x, the relative entropy of PP with respect to QQ is

D(P∥Q)=∑xpxlog⁡pxqx.D(P\Vert Q) = \sum_x p_x\log\frac{p_x}{q_x}.

It is always nonnegative when it is well defined, and it is zero exactly when the two distributions agree. It is not symmetric, however, and it is not a metric.

Relative entropy is the canonical way to quantify how costly it is to model data drawn from PP as though they were drawn from QQ. It also packages mutual information, likelihood-ratio evidence, statistical distinguishability, and the finite-dimensional quantum relative entropy preview used in quantum information.

Let PP and QQ be probability distributions on a finite or countable set Ω\Omega. Write

px=P({x}),qx=Q({x}).p_x=P(\{x\}), \qquad q_x=Q(\{x\}).

The KL divergence is

D(P∥Q)=∑x∈Ωpxlog⁡pxqx.D(P\Vert Q) = \sum_{x\in\Omega} p_x\log\frac{p_x}{q_x}.

The conventions are:

  • if px=0p_x=0, then the contribution is 00;
  • if px>0p_x>0 and qx=0q_x=0, then D(P∥Q)=+∞D(P\Vert Q)=+\infty.

The support condition can be stated as

P≪Q,P\ll Q,

meaning that every event assigned probability zero by QQ is also assigned probability zero by PP. If this absolute-continuity condition fails, relative entropy is infinite.

The logarithm base fixes the units. Natural logarithms give nats. Base-22 logarithms give bits.

Relative entropy is an expectation value under PP:

D(P∥Q)=EP[log⁡pX(X)qX(X)].D(P\Vert Q) = \mathbb E_P \left[ \log \frac{p_X(X)}{q_X(X)} \right].

The random variable inside the expectation is the log-likelihood ratio favoring PP over QQ after seeing the outcome. Thus D(P∥Q)D(P\Vert Q) is the average log evidence per sample, assuming the data really come from PP.

For independent samples,

X1,…,Xn∼P,X_1,\ldots,X_n\sim P,

the product distributions satisfy

D(P⊗n∥Q⊗n)=nD(P∥Q).D(P^{\otimes n}\Vert Q^{\otimes n}) = nD(P\Vert Q).

This additivity is why relative entropy controls exponential rates in hypothesis testing and large-deviation theory.

The fundamental inequality is

D(P∥Q)≥0.D(P\Vert Q)\ge0.

This is often called Gibbs’ inequality. Equality holds exactly when P=QP=Q.

A direct proof uses

−log⁡u≥1−u,u>0.-\log u\ge1-u, \qquad u>0.

Assume first that P≪QP\ll Q. Then

D(P∥Q)=∑px>0px[−log⁡qxpx]≥∑px>0px[1−qxpx]=1−∑px>0qx≥0.\begin{aligned} D(P\Vert Q) &= \sum_{p_x>0} p_x \left[ - \log \frac{q_x}{p_x} \right]\\ &\ge \sum_{p_x>0} p_x \left[ 1-\frac{q_x}{p_x} \right]\\ &= 1- \sum_{p_x>0}q_x\\ &\ge 0. \end{aligned}

If PP and QQ have the same support, the last line is exactly zero only when qx/px=1q_x/p_x=1 for every xx with px>0p_x>0. If QQ has extra probability outside the support of PP, the inequality is strict. Thus equality holds exactly for identical distributions.

Despite the word “divergence,” relative entropy is not a distance metric. In general,

D(P∥Q)≠D(Q∥P).D(P\Vert Q)\ne D(Q\Vert P).

It also does not satisfy the triangle inequality. The order of the arguments matters:

  • D(P∥Q)D(P\Vert Q) averages log evidence using data drawn from PP;
  • D(Q∥P)D(Q\Vert P) averages log evidence using data drawn from QQ.

These can be very different, especially when one distribution assigns small or zero probability to events that are plausible under the other.

For discrete distributions, define the cross-entropy

H(P,Q)=−∑xpxlog⁡qx.H(P,Q) = - \sum_x p_x\log q_x.

The Shannon entropy of PP is

H(P)=−∑xpxlog⁡px.H(P) = - \sum_x p_x\log p_x.

Then

D(P∥Q)=H(P,Q)−H(P).D(P\Vert Q) = H(P,Q)-H(P).

Thus relative entropy is the extra average code length incurred when using a code optimized for QQ while the true distribution is PP, in the idealized Shannon coding interpretation.

Let PP and QQ be two Bernoulli distributions:

P(1)=q,Q(1)=r.P(1)=q, \qquad Q(1)=r.

Then

D(P∥Q)=qlog⁡qr+(1−q)log⁡1−q1−r.D(P\Vert Q) = q\log\frac{q}{r} + (1-q)\log\frac{1-q}{1-r}.

This expression is not symmetric under swapping qq and rr.

If q=1q=1 and r<1r\lt1, then

D(P∥Q)=log⁡1r.D(P\Vert Q)=\log\frac{1}{r}.

If q>0q>0 and r=0r=0, then D(P∥Q)=+∞D(P\Vert Q)=+\infty, because QQ declares an event impossible that PP can produce.

For probability densities p(x)p(x) and q(x)q(x) with respect to the same reference measure, the relative entropy is

D(P∥Q)=∫p(x)log⁡p(x)q(x) dx,D(P\Vert Q) = \int p(x)\log\frac{p(x)}{q(x)}\,dx,

again with the condition that PP be absolutely continuous with respect to QQ.

Unlike differential entropy, relative entropy is invariant under smooth one-to-one changes of variables. If both densities transform with the same Jacobian, the Jacobian factors cancel inside the ratio

p(x)q(x).\frac{p(x)}{q(x)}.

This is one reason relative entropy is often more robust than differential entropy in continuum calculations.

For two one-dimensional Gaussian distributions with the same variance,

P=N(μ0,σ2),Q=N(μ1,σ2),P=\mathcal N(\mu_0,\sigma^2), \qquad Q=\mathcal N(\mu_1,\sigma^2),

one finds, using natural logarithms,

D(P∥Q)=(μ0−μ1)22σ2.D(P\Vert Q) = \frac{(\mu_0-\mu_1)^2}{2\sigma^2}.

For unequal variances the expression has additional scale terms; see the Gaussian background in Gaussian Distributions. The local curvature of this comparison is Fisher Information.

Let XX and YY have joint distribution PXYP_{XY} and marginals PXP_X, PYP_Y. The classical mutual information is

I(X:Y)=D(PXY∥PXPY).I(X:Y) = D \bigl( P_{XY} \Vert P_XP_Y \bigr).

Written out,

I(X:Y)=∑x,yp(x,y)log⁡p(x,y)pX(x)pY(y).I(X:Y) = \sum_{x,y} p(x,y) \log \frac{p(x,y)}{p_X(x)p_Y(y)}.

This says that mutual information measures how distinguishable the actual joint distribution is from the uncorrelated product distribution with the same marginals.

The nonnegativity of relative entropy immediately gives

I(X:Y)≥0,I(X:Y)\ge0,

with equality exactly when XX and YY are independent.

Relative entropy is not the same as Bayesian updating, but it appears naturally in likelihood comparisons.

Suppose two models assign probabilities p(D)p(D) and q(D)q(D) to data DD. The log-likelihood ratio is

log⁡p(D)q(D).\log\frac{p(D)}{q(D)}.

If repeated data are generated from PP, the average log-likelihood ratio per sample approaches

D(P∥Q)D(P\Vert Q)

under standard law-of-large-numbers assumptions. Thus KL divergence measures the asymptotic expected evidence rate against model QQ when PP is the data-generating distribution.

In parameter estimation, minimizing D(Ptrue∥Qθ)D(P_{\mathrm{true}}\Vert Q_\theta) is the idealized population version of maximum-likelihood fitting when the model family QθQ_\theta may be misspecified.

For finite-dimensional density operators ρ\rho and σ\sigma, the quantum relative entropy is

D(ρ∥σ)=Tr⁡[ρ(log⁡ρ−log⁡σ)],D(\rho\Vert\sigma) = \operatorname{Tr} \bigl[ \rho(\log\rho-\log\sigma) \bigr],

provided the support of ρ\rho is contained in the support of σ\sigma. Otherwise it is infinite.

If ρ\rho and σ\sigma commute, they can be diagonalized in the same basis, and the formula reduces to the classical KL divergence between their eigenvalue distributions.

Quantum mutual information can be written as

I(A:B)ρ=D(ρAB∥ρA⊗ρB).I(A:B)_\rho = D \bigl( \rho_{AB} \Vert \rho_A\otimes\rho_B \bigr).

That identity is developed in Mutual Information. The basic von Neumann entropy entering the formula is introduced in Entropy Overview.

In continuum and field-theoretic settings, relative entropy is often better behaved than individual entanglement entropies, but the correct formulation requires algebras, regulators, and domain care. The finite-QM bridge is Entanglement in QFT Preview.

  • Treating D(P∥Q)D(P\Vert Q) as a symmetric distance.
  • Forgetting the support condition; assigning zero probability to a possible event can make the divergence infinite.
  • Swapping the order of PP and QQ without changing the interpretation.
  • Calling KL divergence an entropy of a single distribution rather than a comparison between two distributions.
  • Confusing differential entropy, which is coordinate dependent, with relative entropy, whose density ratio cancels Jacobians.
  • Using the quantum formula without checking noncommutativity and support.
  • Saying relative entropy is “small” without specifying the logarithm base and scale of the comparison.
  • S. Kullback and R. A. Leibler, “On Information and Sufficiency,” Annals of Mathematical Statistics 22, 79–86, 1951.
  • T. M. Cover and J. A. Thomas, Elements of Information Theory, 2nd ed., Wiley, 2006.
  • I. Csiszar and P. C. Shields, “Information Theory and Statistics: A Tutorial,” Foundations and Trends in Communications and Information Theory 1, 417–528, 2004.
  • D. J. C. MacKay, Information Theory, Inference, and Learning Algorithms, Cambridge University Press, 2003.
  • M. A. Nielsen and I. L. Chuang, Quantum Computation and Quantum Information, Cambridge University Press, 2010.
  • J. Watrous, The Theory of Quantum Information, Cambridge University Press, 2018.
  1. Compute D(P∥Q)D(P\Vert Q) for Bernoulli distributions with P(1)=qP(1)=q and Q(1)=rQ(1)=r.
Solution

The two outcomes have probabilities

P(1)=q,P(0)=1−q,P(1)=q, \qquad P(0)=1-q,

and

Q(1)=r,Q(0)=1−r.Q(1)=r, \qquad Q(0)=1-r.

Therefore

D(P∥Q)=qlog⁡qr+(1−q)log⁡1−q1−r,D(P\Vert Q) = q\log\frac{q}{r} + (1-q)\log\frac{1-q}{1-r},

with the usual zero and support conventions.

  1. Show that D(P∥Q)=0D(P\Vert Q)=0 if P=QP=Q.
Solution

If P=QP=Q, then px=qxp_x=q_x for every xx. Hence

log⁡pxqx=log⁡1=0\log\frac{p_x}{q_x} = \log1 = 0

for every term with px>0p_x>0. Terms with px=0p_x=0 also contribute 00. Therefore

D(P∥Q)=0.D(P\Vert Q)=0.

The converse follows from Gibbs’ inequality: equality in the positivity proof requires the distributions to agree.

  1. Prove additivity for independent product distributions.
Solution

Let P1,Q1P_1,Q_1 be distributions on XX and P2,Q2P_2,Q_2 distributions on YY. For the product distributions,

p(x,y)=p1(x)p2(y),q(x,y)=q1(x)q2(y).p(x,y)=p_1(x)p_2(y), \qquad q(x,y)=q_1(x)q_2(y).

Then

D(P1P2∥Q1Q2)=∑x,yp1(x)p2(y)log⁡p1(x)p2(y)q1(x)q2(y)=∑x,yp1(x)p2(y)[log⁡p1(x)q1(x)+log⁡p2(y)q2(y)]=D(P1∥Q1)+D(P2∥Q2).\begin{aligned} D(P_1P_2\Vert Q_1Q_2) &= \sum_{x,y} p_1(x)p_2(y) \log \frac{p_1(x)p_2(y)}{q_1(x)q_2(y)}\\ &= \sum_{x,y} p_1(x)p_2(y) \left[ \log\frac{p_1(x)}{q_1(x)} + \log\frac{p_2(y)}{q_2(y)} \right]\\ &= D(P_1\Vert Q_1)+D(P_2\Vert Q_2). \end{aligned}
  1. A fair bit XX is copied to Y=XY=X. Show that I(X:Y)=1I(X:Y)=1 bit using relative entropy.
Solution

The joint distribution has

p(0,0)=p(1,1)=12,p(0,0)=p(1,1)=\frac12,

and the other two joint outcomes have probability 00. The marginals are uniform, so the product distribution assigns probability 1/41/4 to each pair.

Therefore, in base 22,

I(X:Y)=∑x,yp(x,y)log⁡2p(x,y)pX(x)pY(y)=12log⁡21/21/4+12log⁡21/21/4=1.\begin{aligned} I(X:Y) &= \sum_{x,y} p(x,y) \log_2 \frac{p(x,y)}{p_X(x)p_Y(y)}\\ &= \frac12\log_2\frac{1/2}{1/4} + \frac12\log_2\frac{1/2}{1/4}\\ &= 1. \end{aligned}

The copied bit contains one bit of shared information.

  1. Derive D(P∥Q)D(P\Vert Q) for two Gaussians with the same variance.
Solution

Let

P=N(μ0,σ2),Q=N(μ1,σ2).P=\mathcal N(\mu_0,\sigma^2), \qquad Q=\mathcal N(\mu_1,\sigma^2).

The log density ratio is

log⁡p(x)q(x)=−(x−μ0)22σ2+(x−μ1)22σ2.\log\frac{p(x)}{q(x)} = - \frac{(x-\mu_0)^2}{2\sigma^2} + \frac{(x-\mu_1)^2}{2\sigma^2}.

Taking expectation under PP gives

D(P∥Q)=12σ2EP[(X−μ1)2−(X−μ0)2].D(P\Vert Q) = \frac{1}{2\sigma^2} \mathbb E_P \left[ (X-\mu_1)^2-(X-\mu_0)^2 \right].

Since

EP[(X−μ0)2]=σ2,\mathbb E_P[(X-\mu_0)^2]=\sigma^2,

and

EP[(X−μ1)2]=σ2+(μ0−μ1)2,\mathbb E_P[(X-\mu_1)^2] = \sigma^2+(\mu_0-\mu_1)^2,

one obtains

D(P∥Q)=(μ0−μ1)22σ2.D(P\Vert Q) = \frac{(\mu_0-\mu_1)^2}{2\sigma^2}.