Relative Entropy
Relative entropy compares one probability distribution with another. In probability and statistics it is usually called the Kullback–Leibler divergence, or KL divergence.
For two discrete probability distributions and on the same outcome set, with probabilities and , the relative entropy of with respect to is
It is always nonnegative when it is well defined, and it is zero exactly when the two distributions agree. It is not symmetric, however, and it is not a metric.
Relative entropy is the canonical way to quantify how costly it is to model data drawn from as though they were drawn from . It also packages mutual information, likelihood-ratio evidence, statistical distinguishability, and the finite-dimensional quantum relative entropy preview used in quantum information.
Definition
Section titled “Definition”Let and be probability distributions on a finite or countable set . Write
The KL divergence is
The conventions are:
- if , then the contribution is ;
- if and , then .
The support condition can be stated as
meaning that every event assigned probability zero by is also assigned probability zero by . If this absolute-continuity condition fails, relative entropy is infinite.
The logarithm base fixes the units. Natural logarithms give nats. Base- logarithms give bits.
Expected Log-Likelihood Ratio
Section titled “Expected Log-Likelihood Ratio”Relative entropy is an expectation value under :
The random variable inside the expectation is the log-likelihood ratio favoring over after seeing the outcome. Thus is the average log evidence per sample, assuming the data really come from .
For independent samples,
the product distributions satisfy
This additivity is why relative entropy controls exponential rates in hypothesis testing and large-deviation theory.
Positivity
Section titled “Positivity”The fundamental inequality is
This is often called Gibbs’ inequality. Equality holds exactly when .
A direct proof uses
Assume first that . Then
If and have the same support, the last line is exactly zero only when for every with . If has extra probability outside the support of , the inequality is strict. Thus equality holds exactly for identical distributions.
Not a Distance
Section titled “Not a Distance”Despite the word “divergence,” relative entropy is not a distance metric. In general,
It also does not satisfy the triangle inequality. The order of the arguments matters:
- averages log evidence using data drawn from ;
- averages log evidence using data drawn from .
These can be very different, especially when one distribution assigns small or zero probability to events that are plausible under the other.
Relation to Entropy and Cross-Entropy
Section titled “Relation to Entropy and Cross-Entropy”For discrete distributions, define the cross-entropy
The Shannon entropy of is
Then
Thus relative entropy is the extra average code length incurred when using a code optimized for while the true distribution is , in the idealized Shannon coding interpretation.
Bernoulli Example
Section titled “Bernoulli Example”Let and be two Bernoulli distributions:
Then
This expression is not symmetric under swapping and .
If and , then
If and , then , because declares an event impossible that can produce.
Continuous Distributions
Section titled “Continuous Distributions”For probability densities and with respect to the same reference measure, the relative entropy is
again with the condition that be absolutely continuous with respect to .
Unlike differential entropy, relative entropy is invariant under smooth one-to-one changes of variables. If both densities transform with the same Jacobian, the Jacobian factors cancel inside the ratio
This is one reason relative entropy is often more robust than differential entropy in continuum calculations.
For two one-dimensional Gaussian distributions with the same variance,
one finds, using natural logarithms,
For unequal variances the expression has additional scale terms; see the Gaussian background in Gaussian Distributions. The local curvature of this comparison is Fisher Information.
Mutual Information as Relative Entropy
Section titled “Mutual Information as Relative Entropy”Let and have joint distribution and marginals , . The classical mutual information is
Written out,
This says that mutual information measures how distinguishable the actual joint distribution is from the uncorrelated product distribution with the same marginals.
The nonnegativity of relative entropy immediately gives
with equality exactly when and are independent.
Bayesian and Statistical Meaning
Section titled “Bayesian and Statistical Meaning”Relative entropy is not the same as Bayesian updating, but it appears naturally in likelihood comparisons.
Suppose two models assign probabilities and to data . The log-likelihood ratio is
If repeated data are generated from , the average log-likelihood ratio per sample approaches
under standard law-of-large-numbers assumptions. Thus KL divergence measures the asymptotic expected evidence rate against model when is the data-generating distribution.
In parameter estimation, minimizing is the idealized population version of maximum-likelihood fitting when the model family may be misspecified.
Quantum Relative Entropy Preview
Section titled “Quantum Relative Entropy Preview”For finite-dimensional density operators and , the quantum relative entropy is
provided the support of is contained in the support of . Otherwise it is infinite.
If and commute, they can be diagonalized in the same basis, and the formula reduces to the classical KL divergence between their eigenvalue distributions.
Quantum mutual information can be written as
That identity is developed in Mutual Information. The basic von Neumann entropy entering the formula is introduced in Entropy Overview.
In continuum and field-theoretic settings, relative entropy is often better behaved than individual entanglement entropies, but the correct formulation requires algebras, regulators, and domain care. The finite-QM bridge is Entanglement in QFT Preview.
Common Mistakes
Section titled “Common Mistakes”- Treating as a symmetric distance.
- Forgetting the support condition; assigning zero probability to a possible event can make the divergence infinite.
- Swapping the order of and without changing the interpretation.
- Calling KL divergence an entropy of a single distribution rather than a comparison between two distributions.
- Confusing differential entropy, which is coordinate dependent, with relative entropy, whose density ratio cancels Jacobians.
- Using the quantum formula without checking noncommutativity and support.
- Saying relative entropy is “small” without specifying the logarithm base and scale of the comparison.
Cross-Links
Section titled “Cross-Links”- Entropy
- Conditional Probability
- Bayes’ Rule
- Gaussian Distributions
- Fisher Information
- Mutual Information
- Entropy Overview
- Entanglement in QFT Preview
- Classical Information Review applies Kullback–Leibler divergence and its support convention to declared classical comparator audits; this page retains the mathematical and statistical interpretation.
References
Section titled “References”- S. Kullback and R. A. Leibler, “On Information and Sufficiency,” Annals of Mathematical Statistics 22, 79–86, 1951.
- T. M. Cover and J. A. Thomas, Elements of Information Theory, 2nd ed., Wiley, 2006.
- I. Csiszar and P. C. Shields, “Information Theory and Statistics: A Tutorial,” Foundations and Trends in Communications and Information Theory 1, 417–528, 2004.
- D. J. C. MacKay, Information Theory, Inference, and Learning Algorithms, Cambridge University Press, 2003.
- M. A. Nielsen and I. L. Chuang, Quantum Computation and Quantum Information, Cambridge University Press, 2010.
- J. Watrous, The Theory of Quantum Information, Cambridge University Press, 2018.
Exercises
Section titled “Exercises”- Compute for Bernoulli distributions with and .
Solution
The two outcomes have probabilities
and
Therefore
with the usual zero and support conventions.
- Show that if .
Solution
If , then for every . Hence
for every term with . Terms with also contribute . Therefore
The converse follows from Gibbs’ inequality: equality in the positivity proof requires the distributions to agree.
- Prove additivity for independent product distributions.
Solution
Let be distributions on and distributions on . For the product distributions,
Then
- A fair bit is copied to . Show that bit using relative entropy.
Solution
The joint distribution has
and the other two joint outcomes have probability . The marginals are uniform, so the product distribution assigns probability to each pair.
Therefore, in base ,
The copied bit contains one bit of shared information.
- Derive for two Gaussians with the same variance.
Solution
Let
The log density ratio is
Taking expectation under gives
Since
and
one obtains