Skip to content

Fisher Information

Fisher information measures how sensitively a probability distribution depends on a parameter.

If a model assigns probabilities p(x∣θ)p(x\mid\theta) to outcomes xx, then the Fisher information I(θ)\mathcal I(\theta) quantifies how much one sample can, in principle, tell us about the parameter θ\theta. It is the local curvature of statistical distinguishability and the quantity that appears in the Cramér–Rao lower bound for unbiased estimators.

In quantum mechanics, Fisher information first appears after a measurement has been chosen: a parameter-dependent state and POVM produce an ordinary probability model p(a∣θ)p(a\mid\theta). Classical and Quantum Fisher Information develops the best possible information over allowed measurements; this page owns the classical probability tool needed before that quantum refinement.

A parameterized statistical model is a family of distributions

{Pθ:θ∈Θ}.\{P_\theta:\theta\in\Theta\}.

For discrete outcomes,

px(θ)=Pθ(X=x).p_x(\theta) = \mathbb P_\theta(X=x).

For continuous outcomes,

p(x∣θ)p(x\mid\theta)

is a probability density with respect to a chosen reference measure.

The parameter θ\theta may represent an unknown field strength, phase shift, decay rate, calibration constant, coupling, or model parameter. Fisher information is local in parameter space: it describes sensitivity near a specified value of θ\theta.

For one real parameter, the score is the derivative of the log likelihood for one observation:

Sθ(X)=∂θlog⁡p(X∣θ).S_\theta(X) = \partial_\theta\log p(X\mid\theta).

For a discrete model this means

Sθ(x)=∂θlog⁡px(θ)=∂θpx(θ)px(θ)S_\theta(x) = \partial_\theta\log p_x(\theta) = \frac{\partial_\theta p_x(\theta)}{p_x(\theta)}

when px(θ)>0p_x(\theta)>0.

Under standard regularity conditions, the score has mean zero:

Eθ[Sθ(X)]=0.\mathbb E_\theta[S_\theta(X)]=0.

For a discrete model, the reason is

Eθ[Sθ]=∑xpx(θ)∂θlog⁡px(θ)=∑x∂θpx(θ)=∂θ∑xpx(θ)=0.\begin{aligned} \mathbb E_\theta[S_\theta] &= \sum_x p_x(\theta) \partial_\theta\log p_x(\theta)\\ &= \sum_x \partial_\theta p_x(\theta)\\ &= \partial_\theta \sum_x p_x(\theta)\\ &= 0. \end{aligned}

The regularity assumptions matter: differentiating under the sum or integral must be justified, and the support of the distribution should not jump in a way that invalidates the calculation.

The Fisher information for one parameter is the variance of the score:

I(θ)=Eθ[(∂θlog⁡p(X∣θ))2].\mathcal I(\theta) = \mathbb E_\theta \left[ \left( \partial_\theta\log p(X\mid\theta) \right)^2 \right].

For a discrete model,

I(θ)=∑xpx(θ)[∂θlog⁡px(θ)]2.\mathcal I(\theta) = \sum_x p_x(\theta) \left[ \partial_\theta\log p_x(\theta) \right]^2.

Equivalently,

I(θ)=∑x[∂θpx(θ)]2px(θ),\mathcal I(\theta) = \sum_x \frac{ \left[ \partial_\theta p_x(\theta) \right]^2 }{ p_x(\theta) },

with the usual caution that zero-probability outcomes require support care.

For a density,

I(θ)=∫p(x∣θ)[∂θlog⁡p(x∣θ)]2dx.\mathcal I(\theta) = \int p(x\mid\theta) \left[ \partial_\theta\log p(x\mid\theta) \right]^2dx.

Fisher information is nonnegative. It is zero when the distribution does not change with θ\theta in the relevant local direction.

Fisher information is the local second-order form of relative entropy. For a regular one-parameter model,

D(Pθ∥Pθ+δ)=12I(θ)δ2+O(δ3).D(P_\theta\Vert P_{\theta+\delta}) = \frac12 \mathcal I(\theta)\delta^2 + O(\delta^3).

Thus Fisher information measures how rapidly nearby parameter values become statistically distinguishable.

This statement also explains why Fisher information transforms like a quadratic form under reparameterization. If ϕ=g(θ)\phi=g(\theta) is a smooth one-to-one parameter, then

Iϕ(ϕ)=Iθ(θ)(dθdϕ)2.\mathcal I_\phi(\phi) = \mathcal I_\theta(\theta) \left( \frac{d\theta}{d\phi} \right)^2.

The numerical value of Fisher information depends on the chosen parameter coordinate; the distinguishability line element is the invariant object.

Under stronger regularity conditions,

I(θ)=−Eθ[∂θ2log⁡p(X∣θ)].\mathcal I(\theta) = - \mathbb E_\theta \left[ \partial_\theta^2 \log p(X\mid\theta) \right].

This follows by differentiating the zero-mean score identity:

Eθ[∂θlog⁡p(X∣θ)]=0.\mathbb E_\theta[\partial_\theta\log p(X\mid\theta)]=0.

The second-derivative form is convenient for many likelihood calculations, but it is not the definition. When support depends on θ\theta or boundary terms appear, the score-squared definition is safer.

Let θ^(X1,…,Xn)\hat\theta(X_1,\ldots,X_n) be an unbiased estimator of θ\theta from nn independent samples:

Eθ[θ^]=θ.\mathbb E_\theta[\hat\theta]=\theta.

For a regular one-parameter model, the Cramér–Rao bound says

Var⁡θ(θ^)≥1nI(θ).\operatorname{Var}_\theta(\hat\theta) \ge \frac{1}{n\mathcal I(\theta)}.

The bound says that large Fisher information permits smaller estimator variance, while small Fisher information limits precision.

It is a lower bound under assumptions, not a promise that every estimator reaches it. Attainability can depend on the model, estimator, sample size, and asymptotic regime.

If X1,…,XnX_1,\ldots,X_n are independent and identically distributed, the log likelihood is a sum:

log⁡p(x1,…,xn∣θ)=∑k=1nlog⁡p(xk∣θ).\log p(x_1,\ldots,x_n\mid\theta) = \sum_{k=1}^n \log p(x_k\mid\theta).

Therefore the total score is a sum of independent one-sample scores:

Sθ(n)=∑k=1nSθ(Xk).S_\theta^{(n)} = \sum_{k=1}^n S_\theta(X_k).

Since each score has mean zero, variances add:

In(θ)=nI(θ).\mathcal I_n(\theta) = n\mathcal I(\theta).

This additivity is the source of the familiar 1/n1/n scaling in the Cramér–Rao bound for independent data.

Let XX be Bernoulli with

Pθ(X=1)=θ,Pθ(X=0)=1−θ,0<θ<1.\mathbb P_\theta(X=1)=\theta, \qquad \mathbb P_\theta(X=0)=1-\theta, \qquad 0<\theta<1.

The Fisher information is

I(θ)=(∂θθ)2θ+[∂θ(1−θ)]21−θ=1θ+11−θ=1θ(1−θ).\begin{aligned} \mathcal I(\theta) &= \frac{(\partial_\theta\theta)^2}{\theta} + \frac{[\partial_\theta(1-\theta)]^2}{1-\theta}\\ &= \frac1\theta+\frac{1}{1-\theta}\\ &= \frac{1}{\theta(1-\theta)}. \end{aligned}

The information diverges near θ=0\theta=0 and θ=1\theta=1 because the parameter is near a boundary where a rare contrary outcome is highly informative. Boundary regimes require care when applying asymptotic estimator formulas.

Let

X∼N(μ,σ2),X\sim\mathcal N(\mu,\sigma^2),

where σ\sigma is known and μ\mu is the parameter. The log density is

log⁡p(x∣μ)=−12log⁡(2πσ2)−(x−μ)22σ2.\log p(x\mid\mu) = - \frac12\log(2\pi\sigma^2) - \frac{(x-\mu)^2}{2\sigma^2}.

The score is

∂μlog⁡p(x∣μ)=x−μσ2.\partial_\mu\log p(x\mid\mu) = \frac{x-\mu}{\sigma^2}.

Therefore

I(μ)=Eμ[(X−μ)2σ4]=1σ2.\mathcal I(\mu) = \mathbb E_\mu \left[ \frac{(X-\mu)^2}{\sigma^4} \right] = \frac{1}{\sigma^2}.

For nn independent samples,

Var⁡(μ^)≥σ2n.\operatorname{Var}(\hat\mu) \ge \frac{\sigma^2}{n}.

The sample mean attains this variance in the Gaussian mean model.

For parameters θ=(θ1,…,θm)\boldsymbol\theta=(\theta^1,\ldots,\theta^m), the Fisher information becomes a matrix:

Iij(θ)=Eθ[∂ilog⁡p(X∣θ) ∂jlog⁡p(X∣θ)].\mathcal I_{ij}(\boldsymbol\theta) = \mathbb E_{\boldsymbol\theta} \left[ \partial_i\log p(X\mid\boldsymbol\theta) \, \partial_j\log p(X\mid\boldsymbol\theta) \right].

For an unbiased vector estimator θ^\hat{\boldsymbol\theta} and nn independent samples, the multiparameter Cramér–Rao bound is

Cov⁡(θ^)⪰1nI(θ)−1,\operatorname{Cov}(\hat{\boldsymbol\theta}) \succeq \frac1n \mathcal I(\boldsymbol\theta)^{-1},

when the Fisher matrix is invertible. The symbol ⪰\succeq means that the difference of the left-hand side and right-hand side is positive semidefinite.

Multiparameter estimation is subtler than the one-parameter case. Parameter correlations, nuisance parameters, singular Fisher matrices, and incompatible optimal measurements in the quantum setting can all matter.

In quantum mechanics, a parameter can enter through a state, a Hamiltonian, a channel, a phase shift, or a measurement device. Once a measurement is specified, the outcome probabilities are ordinary classical probabilities.

For a POVM with effects EaE_a and a parameterized density operator ρθ\rho_\theta,

p(a∣θ)=Tr⁡(ρθEa).p(a\mid\theta) = \operatorname{Tr}(\rho_\theta E_a).

The Fisher information of this chosen measurement is

IE(θ)=∑a[∂θp(a∣θ)]2p(a∣θ).\mathcal I_E(\theta) = \sum_a \frac{ \left[ \partial_\theta p(a\mid\theta) \right]^2 }{ p(a\mid\theta) }.

Quantum Fisher information is a different object: it optimizes over measurements, or equivalently is written in terms of the symmetric logarithmic derivative under suitable finite-dimensional assumptions. Classical and Quantum Fisher Information owns that optimization, its state-space geometry, and its attainability caveats.

This page does not develop that theory. It supplies the classical Fisher information needed to understand why quantum metrology is an estimation problem. For the Born probabilities underlying the classical model, see Born Rule and Trace Rule for Expectation Values. For an entangled-state family often used in idealized phase-sensitivity discussions, see GHZ States.

  • Treating Fisher information as a property of data alone rather than of a parameterized model at a parameter value.
  • Forgetting that the value changes under reparameterization.
  • Applying the second-derivative formula when support or boundary terms depend on the parameter.
  • Using the Cramér–Rao bound for a biased estimator without the needed bias correction.
  • Assuming the Cramér–Rao bound is always attainable at finite sample size.
  • Confusing classical Fisher information for a fixed measurement with quantum Fisher information optimized over measurements.
  • Ignoring nuisance parameters and correlations in multiparameter estimation.
  • Treating large Fisher information as automatically useful without checking the experimental model, noise, and allowed measurements.
  • R. A. Fisher, “On the Mathematical Foundations of Theoretical Statistics,” Philosophical Transactions of the Royal Society A 222, 309–368, 1922.
  • C. R. Rao, “Information and the Accuracy Attainable in the Estimation of Statistical Parameters,” Bulletin of the Calcutta Mathematical Society 37, 81–91, 1945.
  • H. Cramér, Mathematical Methods of Statistics, Princeton University Press, 1946.
  • E. L. Lehmann and G. Casella, Theory of Point Estimation, 2nd ed., Springer, 1998.
  • S. M. Kay, Fundamentals of Statistical Signal Processing, Volume I: Estimation Theory, Prentice Hall, 1993.
  • C. W. Helstrom, Quantum Detection and Estimation Theory, Academic Press, 1976.
  • A. S. Holevo, Probabilistic and Statistical Aspects of Quantum Theory, North-Holland, 1982.
  1. Compute the Fisher information for a Bernoulli parameter θ\theta.
Solution

The probabilities are

p1(θ)=θ,p0(θ)=1−θ.p_1(\theta)=\theta, \qquad p_0(\theta)=1-\theta.

Therefore

I(θ)=(∂θp1)2p1+(∂θp0)2p0=1θ+11−θ=1θ(1−θ).\mathcal I(\theta) = \frac{(\partial_\theta p_1)^2}{p_1} + \frac{(\partial_\theta p_0)^2}{p_0} = \frac1\theta+\frac{1}{1-\theta} = \frac{1}{\theta(1-\theta)}.
  1. Compute the Fisher information for the mean μ\mu of a Gaussian with known variance σ2\sigma^2.
Solution

For

p(x∣μ)=12πσexp⁡[−(x−μ)22σ2],p(x\mid\mu) = \frac{1}{\sqrt{2\pi}\sigma} \exp \left[ - \frac{(x-\mu)^2}{2\sigma^2} \right],

the score is

∂μlog⁡p(x∣μ)=x−μσ2.\partial_\mu\log p(x\mid\mu) = \frac{x-\mu}{\sigma^2}.

Thus

I(μ)=Eμ[(X−μ)2σ4]=σ2σ4=1σ2.\mathcal I(\mu) = \mathbb E_\mu \left[ \frac{(X-\mu)^2}{\sigma^4} \right] = \frac{\sigma^2}{\sigma^4} = \frac{1}{\sigma^2}.
  1. Show that Fisher information adds for two independent samples.
Solution

For independent samples, the log likelihood is

log⁡p(x1,x2∣θ)=log⁡p(x1∣θ)+log⁡p(x2∣θ).\log p(x_1,x_2\mid\theta) = \log p(x_1\mid\theta) + \log p(x_2\mid\theta).

The score is the sum

Sθ(2)=Sθ(X1)+Sθ(X2).S^{(2)}_\theta = S_\theta(X_1)+S_\theta(X_2).

Each score has mean zero under the regularity assumptions, and the two scores are independent. Therefore

Var⁡(Sθ(2))=Var⁡(Sθ(X1))+Var⁡(Sθ(X2))=2I(θ).\operatorname{Var}(S^{(2)}_\theta) = \operatorname{Var}(S_\theta(X_1)) + \operatorname{Var}(S_\theta(X_2)) = 2\mathcal I(\theta).
  1. Use the Cramér–Rao bound for estimating a Gaussian mean from nn independent samples with known variance σ2\sigma^2.
Solution

For one sample,

I(μ)=1σ2.\mathcal I(\mu)=\frac{1}{\sigma^2}.

For nn independent samples,

In(μ)=nσ2.\mathcal I_n(\mu)=\frac{n}{\sigma^2}.

Thus any unbiased estimator satisfies

Var⁡(μ^)≥σ2n.\operatorname{Var}(\hat\mu) \ge \frac{\sigma^2}{n}.

The sample mean has variance σ2/n\sigma^2/n, so it attains the bound in this model.

  1. Show that Fisher information is the second-order coefficient of relative entropy for a regular one-parameter model.
Solution

Expand

log⁡p(x∣θ+δ)=log⁡p(x∣θ)+δ ∂θlog⁡p(x∣θ)+δ22∂θ2log⁡p(x∣θ)+O(δ3).\log p(x\mid\theta+\delta) = \log p(x\mid\theta) + \delta\,\partial_\theta\log p(x\mid\theta) + \frac{\delta^2}{2} \partial_\theta^2\log p(x\mid\theta) + O(\delta^3).

Then

D(Pθ∥Pθ+δ)=−δ Eθ[∂θlog⁡p]−δ22Eθ[∂θ2log⁡p]+O(δ3).D(P_\theta\Vert P_{\theta+\delta}) = - \delta\, \mathbb E_\theta[\partial_\theta\log p] - \frac{\delta^2}{2} \mathbb E_\theta[\partial_\theta^2\log p] + O(\delta^3).

The first expectation is zero. Under the regularity condition

I(θ)=−Eθ[∂θ2log⁡p],\mathcal I(\theta) = - \mathbb E_\theta[\partial_\theta^2\log p],

so

D(Pθ∥Pθ+δ)=12I(θ)δ2+O(δ3).D(P_\theta\Vert P_{\theta+\delta}) = \frac12 \mathcal I(\theta)\delta^2 + O(\delta^3).