Skip to content

Entropy

Entropy is a function of a probability distribution that measures average uncertainty, average information gained on learning an outcome, or average surprise, depending on context.

For a discrete random variable XX with probabilities px=P(X=x)p_x=\mathbb P(X=x), the Shannon entropy is

H(X)=−∑xpxlog⁡px.H(X) = - \sum_x p_x\log p_x.

The entropy depends on the probabilities, not on the names of the outcomes. A deterministic variable has zero entropy. A uniform distribution over nn possible outcomes has entropy log⁡n\log n.

This page gives the classical probability version. The density-operator version is Entropy Overview, and reduced-state entropy in composite quantum systems is Subsystem Entropy.

Let XX take values in a finite or countable set with probabilities

px≥0,∑xpx=1.p_x\ge0, \qquad \sum_x p_x=1.

The Shannon entropy is

H(X)=−∑xpxlog⁡px,H(X) = - \sum_x p_x\log p_x,

with the convention

0log⁡0=0.0\log0=0.

This convention is the limit

lim⁡p→0+plog⁡p=0.\lim_{p\to0^+}p\log p=0.

The logarithm base fixes the units:

  • log⁡2\log_2 gives entropy in bits;
  • ln⁡\ln gives entropy in nats;
  • log⁡b\log_b gives entropy in units of log⁡b\log b.

Changing the base rescales entropy by a constant. For example,

Hbits(X)=Hnats(X)ln⁡2.H_{\mathrm{bits}}(X) = \frac{H_{\mathrm{nats}}(X)}{\ln2}.

The information content, or surprise, of an outcome with probability pp is

−log⁡p.-\log p.

Rare outcomes carry larger surprise than common outcomes. The entropy is the average surprise:

H(X)=E[−log⁡pX(X)].H(X) = \mathbb E[-\log p_X(X)].

Written out,

E[−log⁡pX(X)]=∑xpx[−log⁡px]=−∑xpxlog⁡px.\mathbb E[-\log p_X(X)] = \sum_x p_x[-\log p_x] = - \sum_x p_x\log p_x.

This interpretation is useful but should not be overextended. Entropy is not itself a random measurement outcome. It is a functional of the whole distribution.

If XX is deterministic, then one outcome has probability 11 and all others have probability 00. Therefore

H(X)=0.H(X)=0.

If XX is uniform on nn outcomes, then px=1/np_x=1/n and

H(X)=−∑x=1n1nlog⁡1n=log⁡n.\begin{aligned} H(X) &= - \sum_{x=1}^n \frac1n\log\frac1n\\ &= \log n. \end{aligned}

For a fixed finite set of nn outcomes, this is the maximum possible entropy. Intuitively, uncertainty is largest when no outcome is favored.

For a two-outcome random variable with probabilities qq and 1−q1-q, the entropy in bits is

H2(q)=−qlog⁡2q−(1−q)log⁡2(1−q).H_2(q) = - q\log_2 q - (1-q)\log_2(1-q).

It satisfies

H2(0)=H2(1)=0,H2(1/2)=1.H_2(0)=H_2(1)=0, \qquad H_2(1/2)=1.

The binary entropy is symmetric:

H2(q)=H2(1−q).H_2(q)=H_2(1-q).

It is the classical template behind the entropy of a diagonal qubit density matrix with eigenvalues qq and 1−q1-q.

For a pair of discrete random variables (X,Y)(X,Y) with joint probabilities

pxy=P(X=x,Y=y),p_{xy} = \mathbb P(X=x,Y=y),

the joint entropy is

H(X,Y)=−∑x,ypxylog⁡pxy.H(X,Y) = - \sum_{x,y}p_{xy}\log p_{xy}.

If XX and YY are independent, then

pxy=pxpy,p_{xy}=p_xp_y,

and

H(X,Y)=H(X)+H(Y).H(X,Y)=H(X)+H(Y).

If Y=XY=X exactly, then the pair (X,Y)(X,Y) contains no more uncertainty than XX alone:

H(X,Y)=H(X).H(X,Y)=H(X).

Thus joint entropy counts uncertainty in the joint outcome, not the number of variables written down.

The conditional entropy of YY given XX is

H(Y∣X)=∑xpxH(Y∣X=x).H(Y\mid X) = \sum_x p_x H(Y\mid X=x).

Equivalently,

H(Y∣X)=−∑x,ypxylog⁡py∣x,H(Y\mid X) = - \sum_{x,y} p_{xy}\log p_{y\mid x},

where

py∣x=P(Y=y∣X=x)p_{y\mid x} = \mathbb P(Y=y\mid X=x)

when px>0p_x>0.

The chain rule is

H(X,Y)=H(X)+H(Y∣X).H(X,Y) = H(X)+H(Y\mid X).

For ordinary classical variables,

H(Y∣X)≥0.H(Y\mid X)\ge0.

If YY is determined by XX, then H(Y∣X)=0H(Y\mid X)=0. If XX and YY are independent, then H(Y∣X)=H(Y)H(Y\mid X)=H(Y).

Quantum conditional entropy can be negative; that is a property of von Neumann entropy for composite quantum states, not of classical conditional Shannon entropy.

Classical mutual information is

I(X:Y)=H(X)+H(Y)−H(X,Y).I(X:Y) = H(X)+H(Y)-H(X,Y).

Using the chain rule,

I(X:Y)=H(Y)−H(Y∣X).I(X:Y) = H(Y)-H(Y\mid X).

It measures the reduction in uncertainty about one variable obtained by learning the other. It vanishes when XX and YY are independent.

Equivalently, mutual information is a relative entropy comparing the joint distribution with the product of its marginals.

The quantum analogue uses von Neumann entropies of density operators:

I(A:B)ρ=S(ρA)+S(ρB)−S(ρAB).I(A:B)_\rho = S(\rho_A)+S(\rho_B)-S(\rho_{AB}).

For the quantum version and its interpretation as total correlation, see Mutual Information.

Entropy depends on the variables and distinctions being used. If a fine-grained random variable XX is mapped to a coarser variable

Y=g(X),Y=g(X),

then YY cannot contain more Shannon entropy than XX:

H(Y)≤H(X).H(Y)\le H(X).

Coarse graining can merge outcomes and discard distinctions. The lost information is zero only when gg is one-to-one on the support of the distribution.

In physics, this warning matters because the entropy assigned to a description depends on what macroscopic variables, measurement outcomes, subsystem splits, or preparation records are being retained.

For a continuous random variable with density fX(x)f_X(x), one can define the differential entropy

h(X)=−∫fX(x)log⁡fX(x) dx,h(X) = - \int f_X(x)\log f_X(x)\,dx,

when the integral is meaningful.

This resembles Shannon entropy, but it is not the same kind of invariant uncertainty measure. There are three major cautions.

First, a density has units. Taking the logarithm of a dimensionful density is shorthand for a convention involving a reference measure or unit scale.

Second, differential entropy changes under coordinate scaling. If

Y=aX,a≠0,Y=aX, \qquad a\ne0,

then

h(Y)=h(X)+log⁡∣a∣.h(Y)=h(X)+\log\lvert a\rvert.

Changing from meters to centimeters changes the numerical differential entropy.

Third, differential entropy can be negative. A sharply localized density can have h(X)<0h(X)\lt0 in chosen units. This is not a paradox; it means differential entropy is not a direct continuous analogue of the nonnegative discrete entropy.

For a Gaussian density with variance σ2\sigma^2, the differential entropy is

h(X)=12log⁡(2πe σ2),h(X) = \frac12 \log \left( 2\pi e\,\sigma^2 \right),

with the same unit-scale caution. The probability-density facts behind this formula are collected in Gaussian Distributions.

There are several different entropy-like quantities in quantum mechanics. Keeping them separate prevents many mistakes.

The Shannon entropy of a measurement is the entropy of the Born probabilities for that chosen measurement:

H(A)=−∑ap(a)log⁡p(a),p(a)=⟨ψ∣Pa∣ψ⟩H(A) = - \sum_a p(a)\log p(a), \qquad p(a)=\langle\psi\vert P_a\lvert\psi\rangle

for a projective measurement on a pure state.

The von Neumann entropy of a density operator is

S(ρ)=−Tr⁡(ρlog⁡ρ).S(\rho) = - \operatorname{Tr}(\rho\log\rho).

If ρ\rho has eigenvalues λk\lambda_k, then

S(ρ)=−∑kλklog⁡λk.S(\rho) = - \sum_k\lambda_k\log\lambda_k.

Thus von Neumann entropy is the Shannon entropy of the eigenvalue distribution of ρ\rho, not the Shannon entropy of an arbitrary measurement outcome distribution.

A pure state can have zero von Neumann entropy but nonzero measurement entropy in a basis where the outcome is not certain. For example,

∣+⟩=∣0⟩+∣1⟩2\lvert+\rangle = \frac{\lvert0\rangle+\lvert1\rangle}{\sqrt2}

has zero von Neumann entropy as a pure state, but a computational-basis measurement has one bit of Shannon entropy.

For reduced states of composite systems, von Neumann entropy can quantify local mixedness, pure-state bipartite entanglement, thermal uncertainty, or correlation-derived quantities depending on the context. The context must be stated.

  • Forgetting to state the logarithm base.
  • Treating 0log⁡00\log0 as undefined instead of using its limiting value 00.
  • Calling entropy a property of a single outcome rather than of a distribution.
  • Confusing Shannon entropy of a chosen measurement with von Neumann entropy of a quantum state.
  • Assuming a positive subsystem entropy always means entanglement, even for mixed joint states.
  • Treating differential entropy as coordinate invariant.
  • Forgetting that differential entropy can be negative.
  • Interpreting entropy without specifying the retained variables, measurement, subsystem split, or coarse graining.
  • C. E. Shannon, “A Mathematical Theory of Communication,” Bell System Technical Journal 27, 379–423 and 623–656, 1948.
  • A. I. Khinchin, Mathematical Foundations of Information Theory, Dover, 1957.
  • T. M. Cover and J. A. Thomas, Elements of Information Theory, 2nd ed., Wiley, 2006.
  • D. J. C. MacKay, Information Theory, Inference, and Learning Algorithms, Cambridge University Press, 2003.
  • J. von Neumann, Mathematical Foundations of Quantum Mechanics, Princeton University Press, 1955.
  • M. A. Nielsen and I. L. Chuang, Quantum Computation and Quantum Information, Cambridge University Press, 2010.
  1. Compute the entropy in bits of a biased coin with P(H)=q\mathbb P(H)=q.
Solution

The distribution has probabilities qq and 1−q1-q, so

H=−qlog⁡2q−(1−q)log⁡2(1−q).H = - q\log_2 q - (1-q)\log_2(1-q).

This is the binary entropy H2(q)H_2(q).

  1. Show that the uniform distribution on nn outcomes has entropy log⁡n\log n.
Solution

For a uniform distribution, px=1/np_x=1/n for each of the nn outcomes. Therefore

H=−∑x=1n1nlog⁡1n=−log⁡1n=log⁡n.\begin{aligned} H &= - \sum_{x=1}^n \frac1n\log\frac1n\\ &= - \log\frac1n\\ &= \log n. \end{aligned}
  1. Let XX be a fair bit and let Y=XY=X. Compute H(X)H(X), H(Y)H(Y), and H(X,Y)H(X,Y) in bits.
Solution

Since XX is a fair bit,

H(X)=1.H(X)=1.

The same is true for YY:

H(Y)=1.H(Y)=1.

But the joint pair has only two possible outcomes, (0,0)(0,0) and (1,1)(1,1), each with probability 1/21/2. Hence

H(X,Y)=1.H(X,Y)=1.

The two variables are perfectly correlated, so the joint uncertainty is not two bits.

  1. If Y=aXY=aX with a≠0a\ne0, show that h(Y)=h(X)+log⁡∣a∣h(Y)=h(X)+\log\lvert a\rvert.
Solution

The density of YY is

fY(y)=1∣a∣fX(ya).f_Y(y) = \frac{1}{\lvert a\rvert} f_X \left( \frac{y}{a} \right).

Then

h(Y)=−∫fY(y)log⁡fY(y) dy=−∫fX(x)log⁡[1∣a∣fX(x)]dx=−∫fX(x)log⁡fX(x) dx+log⁡∣a∣∫fX(x) dx=h(X)+log⁡∣a∣.\begin{aligned} h(Y) &= - \int f_Y(y)\log f_Y(y)\,dy\\ &= - \int f_X(x) \log \left[ \frac{1}{\lvert a\rvert}f_X(x) \right]dx\\ &= - \int f_X(x)\log f_X(x)\,dx + \log\lvert a\rvert \int f_X(x)\,dx\\ &= h(X)+\log\lvert a\rvert. \end{aligned}
  1. A qubit is in the pure state ∣+⟩=(∣0⟩+∣1⟩)/2\lvert+\rangle=(\lvert0\rangle+\lvert1\rangle)/\sqrt2. Compare its von Neumann entropy with the Shannon entropy of a computational-basis measurement.
Solution

The density operator ρ=∣+⟩⟨+∣\rho=\lvert+\rangle\langle+\rvert is pure, so its eigenvalues are 11 and 00. Therefore

S(ρ)=0.S(\rho)=0.

A computational-basis measurement has probabilities

p(0)=p(1)=12.p(0)=p(1)=\frac12.

The Shannon entropy of that measurement is

H=−12log⁡212−12log⁡212=1H = - \frac12\log_2\frac12 - \frac12\log_2\frac12 = 1

bit. The state entropy and the measurement-outcome entropy answer different questions.