Entropy
Entropy is a function of a probability distribution that measures average uncertainty, average information gained on learning an outcome, or average surprise, depending on context.
For a discrete random variable with probabilities , the Shannon entropy is
The entropy depends on the probabilities, not on the names of the outcomes. A deterministic variable has zero entropy. A uniform distribution over possible outcomes has entropy .
This page gives the classical probability version. The density-operator version is Entropy Overview, and reduced-state entropy in composite quantum systems is Subsystem Entropy.
Discrete Shannon Entropy
Section titled “Discrete Shannon Entropy”Let take values in a finite or countable set with probabilities
The Shannon entropy is
with the convention
This convention is the limit
The logarithm base fixes the units:
- gives entropy in bits;
- gives entropy in nats;
- gives entropy in units of .
Changing the base rescales entropy by a constant. For example,
Entropy as Expected Surprise
Section titled “Entropy as Expected Surprise”The information content, or surprise, of an outcome with probability is
Rare outcomes carry larger surprise than common outcomes. The entropy is the average surprise:
Written out,
This interpretation is useful but should not be overextended. Entropy is not itself a random measurement outcome. It is a functional of the whole distribution.
Deterministic and Uniform Cases
Section titled “Deterministic and Uniform Cases”If is deterministic, then one outcome has probability and all others have probability . Therefore
If is uniform on outcomes, then and
For a fixed finite set of outcomes, this is the maximum possible entropy. Intuitively, uncertainty is largest when no outcome is favored.
Binary Entropy
Section titled “Binary Entropy”For a two-outcome random variable with probabilities and , the entropy in bits is
It satisfies
The binary entropy is symmetric:
It is the classical template behind the entropy of a diagonal qubit density matrix with eigenvalues and .
Joint Entropy
Section titled “Joint Entropy”For a pair of discrete random variables with joint probabilities
the joint entropy is
If and are independent, then
and
If exactly, then the pair contains no more uncertainty than alone:
Thus joint entropy counts uncertainty in the joint outcome, not the number of variables written down.
Conditional Entropy
Section titled “Conditional Entropy”The conditional entropy of given is
Equivalently,
where
when .
The chain rule is
For ordinary classical variables,
If is determined by , then . If and are independent, then .
Quantum conditional entropy can be negative; that is a property of von Neumann entropy for composite quantum states, not of classical conditional Shannon entropy.
Mutual Information Preview
Section titled “Mutual Information Preview”Classical mutual information is
Using the chain rule,
It measures the reduction in uncertainty about one variable obtained by learning the other. It vanishes when and are independent.
Equivalently, mutual information is a relative entropy comparing the joint distribution with the product of its marginals.
The quantum analogue uses von Neumann entropies of density operators:
For the quantum version and its interpretation as total correlation, see Mutual Information.
Entropy and Coarse Graining
Section titled “Entropy and Coarse Graining”Entropy depends on the variables and distinctions being used. If a fine-grained random variable is mapped to a coarser variable
then cannot contain more Shannon entropy than :
Coarse graining can merge outcomes and discard distinctions. The lost information is zero only when is one-to-one on the support of the distribution.
In physics, this warning matters because the entropy assigned to a description depends on what macroscopic variables, measurement outcomes, subsystem splits, or preparation records are being retained.
Differential Entropy Caution
Section titled “Differential Entropy Caution”For a continuous random variable with density , one can define the differential entropy
when the integral is meaningful.
This resembles Shannon entropy, but it is not the same kind of invariant uncertainty measure. There are three major cautions.
First, a density has units. Taking the logarithm of a dimensionful density is shorthand for a convention involving a reference measure or unit scale.
Second, differential entropy changes under coordinate scaling. If
then
Changing from meters to centimeters changes the numerical differential entropy.
Third, differential entropy can be negative. A sharply localized density can have in chosen units. This is not a paradox; it means differential entropy is not a direct continuous analogue of the nonnegative discrete entropy.
For a Gaussian density with variance , the differential entropy is
with the same unit-scale caution. The probability-density facts behind this formula are collected in Gaussian Distributions.
Quantum-Mechanics Interpretation
Section titled “Quantum-Mechanics Interpretation”There are several different entropy-like quantities in quantum mechanics. Keeping them separate prevents many mistakes.
The Shannon entropy of a measurement is the entropy of the Born probabilities for that chosen measurement:
for a projective measurement on a pure state.
The von Neumann entropy of a density operator is
If has eigenvalues , then
Thus von Neumann entropy is the Shannon entropy of the eigenvalue distribution of , not the Shannon entropy of an arbitrary measurement outcome distribution.
A pure state can have zero von Neumann entropy but nonzero measurement entropy in a basis where the outcome is not certain. For example,
has zero von Neumann entropy as a pure state, but a computational-basis measurement has one bit of Shannon entropy.
For reduced states of composite systems, von Neumann entropy can quantify local mixedness, pure-state bipartite entanglement, thermal uncertainty, or correlation-derived quantities depending on the context. The context must be stated.
Common Mistakes
Section titled “Common Mistakes”- Forgetting to state the logarithm base.
- Treating as undefined instead of using its limiting value .
- Calling entropy a property of a single outcome rather than of a distribution.
- Confusing Shannon entropy of a chosen measurement with von Neumann entropy of a quantum state.
- Assuming a positive subsystem entropy always means entanglement, even for mixed joint states.
- Treating differential entropy as coordinate invariant.
- Forgetting that differential entropy can be negative.
- Interpreting entropy without specifying the retained variables, measurement, subsystem split, or coarse graining.
Cross-Links
Section titled “Cross-Links”- Probability Spaces, Light Version
- Random Variables
- Probability Densities
- Expectation Values
- Conditional Probability
- Bayes’ Rule
- Relative Entropy
- Gaussian Distributions
- Entropy Overview
- Subsystem Entropy
- Mutual Information
- Classical Information Review applies Shannon and binary entropy to source-coding and fair-comparator checks; this page retains their definitions and derivations.
References
Section titled “References”- C. E. Shannon, “A Mathematical Theory of Communication,” Bell System Technical Journal 27, 379–423 and 623–656, 1948.
- A. I. Khinchin, Mathematical Foundations of Information Theory, Dover, 1957.
- T. M. Cover and J. A. Thomas, Elements of Information Theory, 2nd ed., Wiley, 2006.
- D. J. C. MacKay, Information Theory, Inference, and Learning Algorithms, Cambridge University Press, 2003.
- J. von Neumann, Mathematical Foundations of Quantum Mechanics, Princeton University Press, 1955.
- M. A. Nielsen and I. L. Chuang, Quantum Computation and Quantum Information, Cambridge University Press, 2010.
Exercises
Section titled “Exercises”- Compute the entropy in bits of a biased coin with .
Solution
The distribution has probabilities and , so
This is the binary entropy .
- Show that the uniform distribution on outcomes has entropy .
Solution
For a uniform distribution, for each of the outcomes. Therefore
- Let be a fair bit and let . Compute , , and in bits.
Solution
Since is a fair bit,
The same is true for :
But the joint pair has only two possible outcomes, and , each with probability . Hence
The two variables are perfectly correlated, so the joint uncertainty is not two bits.
- If with , show that .
Solution
The density of is
Then
- A qubit is in the pure state . Compare its von Neumann entropy with the Shannon entropy of a computational-basis measurement.
Solution
The density operator is pure, so its eigenvalues are and . Therefore
A computational-basis measurement has probabilities
The Shannon entropy of that measurement is
bit. The state entropy and the measurement-outcome entropy answer different questions.