Skip to content

Probability, Statistics, and Information

Probability theory supplies the language of outcomes, distributions, averages, fluctuations, inference, and sampling error. Quantum mechanics uses that language every time a measurement is fixed, but it does not reduce all observables to random variables on one universal classical sample space. This chapter develops the classical machinery first and marks the quantum boundary only after the machinery is secure.

The chapter has four strands: probability models and random variables; moments and conditioning; information and estimation; and Monte Carlo computation. The final comparison page explains which parts carry directly into a quantum experiment and which parts are changed by amplitudes, noncommutativity, and measurement disturbance.

A probability space is a triple

(Ω,F,P),(\Omega,\mathcal F,\mathbb P),

where Ω\Omega is the outcome space, F\mathcal F is a collection of measurable events, and P\mathbb P is a normalized countably additive measure. A random variable is a measurable map

X:Ω⟶X.X:\Omega\longrightarrow\mathcal X.

Its probability distribution is the pushforward measure

PX(B)=P ⁣(X−1(B)).\mathbb P_X(B) =\mathbb P\!\left(X^{-1}(B)\right).

The distribution can often be used without retaining the microscopic outcome ω∈Ω\omega\in\Omega. This separation among outcome, event, random variable, and value prevents several common category errors.

For a discrete variable, probabilities are masses pi=P(X=xi)p_i=\mathbb P(X=x_i). For a continuous variable with density fXf_X relative to dxdx,

P(X∈B)=∫BfX(x) dx,∫fX(x) dx=1.\mathbb P(X\in B)=\int_B f_X(x)\,dx, \qquad \int f_X(x)\,dx=1.

A density is probability per unit measure and generally has physical units. The value fX(x)f_X(x) is not the probability of the exact point xx. Probability Spaces, Light Version, Random Variables, and Probability Densities form the foundational route.

If Y=g(X)Y=g(X) and gg is differentiable with simple inverse branches, then

fY(y)=∑xi:g(xi)=yfX(xi)∣g′(xi)∣.f_Y(y) =\sum_{x_i:g(x_i)=y} \frac{f_X(x_i)}{\lvert g'(x_i)\rvert}.

The sum over branches is essential when gg is not one-to-one. In several dimensions, the absolute determinant of the inverse Jacobian replaces 1/∣g′∣1/\lvert g'\rvert. A density’s numerical value and units change under reparameterization even though probabilities of corresponding events do not.

This same principle explains why ∣ψ(x)∣2\lvert\psi(x)\rvert^2 and ∣ϕ(p)∣2\lvert\phi(p)\rvert^2 use different measures and why radial densities acquire volume factors. The Born rule itself remains canonical in Probability and the Born Rule.

For an integrable function g(X)g(X),

E[g(X)]={∑ig(xi)pi,discrete,∫g(x)fX(x) dx,continuous.\mathbb E[g(X)] =\begin{cases} \displaystyle\sum_i g(x_i)p_i, &\text{discrete},\\ \displaystyle\int g(x)f_X(x)\,dx, &\text{continuous}. \end{cases}

For real-valued variables, the mean is μX=E[X]\mu_X=\mathbb E[X]. Variance and covariance are

Var⁡(X)=E[(X−μX)2],\operatorname{Var}(X) =\mathbb E[(X-\mu_X)^2],

and

Cov⁡(X,Y)=E[(X−μX)(Y−μY)].\operatorname{Cov}(X,Y) =\mathbb E[(X-\mu_X)(Y-\mu_Y)].

Variance measures mean-square spread, not general unpredictability. Zero covariance does not imply independence except under additional assumptions, notably for jointly Gaussian variables. Covariance matrices must be positive semidefinite because

aTΣa=Var⁡(aTX)≥0.\mathbf a^{\mathsf T}\Sigma\mathbf a =\operatorname{Var}(\mathbf a^{\mathsf T}\mathbf X) \geq0.

Use Expectation Values for integration and linearity, then Variance and Covariance for fluctuations and correlations. Quantum expectation values and uncertainty relations have separate canonical homes in Core Formalism.

For P(B)>0\mathbb P(B)\gt0,

P(A∣B)=P(A∩B)P(B).\mathbb P(A\mid B) =\frac{\mathbb P(A\cap B)}{\mathbb P(B)}.

Bayes’ rule reverses a conditional probability:

P(A∣B)=P(B∣A)P(A)P(B).\mathbb P(A\mid B) =\frac{\mathbb P(B\mid A)\mathbb P(A)} {\mathbb P(B)}.

For a parameter θ\theta and observed data dd,

p(θ∣d)=p(d∣θ)p(θ)∫p(d∣ϑ)p(ϑ) dϑ.p(\theta\mid d) =\frac{p(d\mid\theta)p(\theta)} {\int p(d\mid\vartheta)p(\vartheta)\,d\vartheta}.

The likelihood p(d∣θ)p(d\mid\theta) is not a normalized density over θ\theta until combined with a prior and normalized. Base rates can dominate a seemingly accurate diagnostic when the target event is rare.

Conditional Probability owns conditioning of events and densities. Bayes’ Rule owns inference and parameter updates. Quantum state update may have a conditioning aspect, but a quantum instrument also specifies outcome-dependent physical transformation; it is not determined by Bayes’ rule alone.

The characteristic function

χX(t)=E[eitX]\chi_X(t)=\mathbb E[e^{itX}]

is the Fourier transform of a probability distribution under the probability convention. When the required moments exist,

E[Xn]=1indnχXdtn∣t=0.\mathbb E[X^n] =\frac{1}{i^n} \left.\frac{d^n\chi_X}{dt^n}\right|_{t=0}.

For independent variables, χX+Y(t)=χX(t)χY(t)\chi_{X+Y}(t)=\chi_X(t)\chi_Y(t), turning convolution of distributions into multiplication. Characteristic functions always exist for real random variables because ∣eitX∣=1\lvert e^{itX}\rvert=1, even when ordinary moments do not.

A one-dimensional Gaussian with mean μ\mu and variance σ2\sigma^2 has density

f(x)=12πσ2exp⁡ ⁣[−(x−μ)22σ2]f(x) =\frac{1}{\sqrt{2\pi\sigma^2}} \exp\!\left[-\frac{(x-\mu)^2}{2\sigma^2}\right]

and characteristic function

χX(t)=exp⁡ ⁣(iμt−12σ2t2).\chi_X(t) =\exp\!\left(i\mu t-\frac12\sigma^2t^2\right).

Multivariate Gaussians are controlled by a mean vector and positive-semidefinite covariance matrix, with special care needed for singular covariance. Characteristic Functions and Gaussian Distributions supply the detailed route.

For a discrete distribution pp, Shannon entropy is

H(p)=−∑ipilog⁡pi.H(p)=-\sum_i p_i\log p_i.

The logarithm base sets the unit: base two gives bits and the natural logarithm gives nats. Entropy quantifies average uncertainty under a specified distribution; it is not a universal measure of disorder detached from a model.

For a continuous density,

h(f)=−∫f(x)log⁡f(x) dxh(f)=-\int f(x)\log f(x)\,dx

is differential entropy. It depends on coordinates and reference measure, can be negative, and is not the direct continuous analogue of nonnegative discrete entropy.

Relative entropy compares two distributions:

D(p∥q)=∑ipilog⁡piqi≥0,D(p\Vert q) =\sum_i p_i\log\frac{p_i}{q_i} \geq0,

with D(p∥q)=+∞D(p\Vert q)=+\infty when pp assigns positive mass where qq assigns zero mass. It is generally asymmetric and is not a metric. Unlike differential entropy alone, relative entropy is invariant under smooth one-to-one coordinate changes when both densities transform against the same reference measure.

Entropy and Relative Entropy own the classical definitions. Von Neumann entropy, quantum relative entropy, and entanglement entropy remain in the density-operator and composite-system volumes.

For a regular parameterized density p(x∣θ)p(x\mid\theta), the score is

sθ(x)=∂∂θlog⁡p(x∣θ),s_\theta(x) =\frac{\partial}{\partial\theta} \log p(x\mid\theta),

and the Fisher information is

I(θ)=Eθ[sθ(X)2].I(\theta) =\mathbb E_\theta[s_\theta(X)^2].

Under standard regularity assumptions, the score has zero mean. For an unbiased estimator θ^\widehat\theta based on NN independent observations, the Cramér–Rao bound gives

Var⁡(θ^)≥1NI(θ).\operatorname{Var}(\widehat\theta) \geq\frac{1}{N I(\theta)}.

The bound requires its stated assumptions and need not be attained at finite sample size. Fisher information is local in parameter space and transforms as a metric tensor under smooth reparameterization. Fisher Information develops the classical result and links onward to quantum estimation without duplicating quantum Fisher information.

For independent samples X1,…,XNX_1,\ldots,X_N with finite variance, the sample mean

X‾N=1N∑n=1NXn\overline X_N=\frac1N\sum_{n=1}^N X_n

is unbiased for E[X]\mathbb E[X] and has

Var⁡(X‾N)=Var⁡(X)N.\operatorname{Var}(\overline X_N) =\frac{\operatorname{Var}(X)}{N}.

The standard error therefore scales as N−1/2N^{-1/2}, not N−1N^{-1}. Correlated Markov-chain samples have a smaller effective sample size determined by autocorrelation. Importance sampling can reduce variance by sampling from a proposal that covers the important regions, but support mismatch or highly variable weights can make an estimator unstable.

Monte Carlo Basics owns estimators, uncertainty, importance sampling, and autocorrelation cautions. Numerical implementation, convergence tests, and benchmark design remain in Numerical Mathematics.

For a fixed quantum state ρ\rho and POVM {Ea}\{E_a\},

p(a)=tr⁡(ρEa)p(a)=\operatorname{tr}(\rho E_a)

is an ordinary classical probability distribution over the outcome labels aa. Once the experiment is specified, classical expectation, likelihood, entropy, and sampling tools apply to its recorded outcomes.

The difference appears when one asks for one joint classical model covering all possible quantum measurements.

Classical probabilityQuantum experiment
one sample space and event algebraeach measurement supplies an outcome algebra; incompatible measurements need not share a joint distribution
random variable on the sample spaceself-adjoint operator or POVM specifying an outcome law
law of total probability for exclusive alternativesamplitudes can interfere before probabilities are formed
conditioning changes information about a fixed modelan instrument can also disturb the physical state
joint variables exist by constructionjoint measurability imposes nontrivial compatibility conditions

Commuting projective observables admit a common spectral measure and behave like jointly distributed classical variables for that context. Noncommuting observables generally do not. Classical Probability versus Quantum Probability gives the careful comparison; the Born rule, sequential measurement, and state update remain canonical in Core Formalism.

PageCentral question
Probability Spaces, Light VersionWhat are outcomes, events, measures, and distributions?
Random VariablesHow does a measurable map turn outcomes into values?
Probability DensitiesHow is probability represented relative to a continuous measure?
Expectation ValuesHow are probability-weighted averages defined and manipulated?
Variance and CovarianceHow are spread and linear dependence quantified?
Conditional ProbabilityHow does restricting to known information change a distribution?
Bayes’ RuleHow are likelihood and prior combined into a posterior?
Characteristic FunctionsHow does Fourier analysis encode distributions and independent sums?
Gaussian DistributionsWhy do means and covariances completely determine Gaussian laws?
EntropyHow is average information or uncertainty quantified?
Relative EntropyHow is one distribution compared with another?
Fisher InformationHow much local parameter sensitivity does a statistical model contain?
Monte Carlo BasicsHow are expectations estimated from samples with controlled uncertainty?
Classical Probability versus Quantum ProbabilityWhich classical rules apply to fixed measurements, and where does quantum structure exceed them?
  • Born-rule calculations: probability spaces →\to random variables →\to densities →\to expectation →\to variance, then Probability and the Born Rule.
  • Inference and tomography: conditioning →\to Bayes’ rule →\to Fisher information, followed by the measurement and estimation pages for the experiment at hand.
  • Information theory: entropy →\to relative entropy →\to classical-versus-quantum probability, then Entropy Overview.
  • Stochastic computation: Gaussian distributions →\to expectation and variance →\to Monte Carlo, then Error Estimates and Convergence Tests.
MistakeCorrection
Treating a density value as a point probabilityintegrate the density over an event and track its units
Changing variables without a Jacobian or inverse branchestransform both density and measure and sum over all preimages
Assuming zero covariance implies independencethis requires additional structure, such as joint Gaussianity
Reversing p(d∣θ)p(d\mid\theta) into p(θ∣d)p(\theta\mid d) without a priorapply Bayes’ rule and normalize
Comparing differential entropies across coordinates as absolute quantitiesuse a common reference measure or relative entropy
Reading the Cramér–Rao bound without its regularity and bias assumptionsstate the model, estimator class, and sample conditions
Reporting Monte Carlo digits without a standard error or autocorrelation analysisestimate uncertainty and effective sample size
Treating quantum state update as ordinary conditioning alonespecify the quantum instrument and its disturbance

Let XX be uniform on [−1,1][-1,1] and let Y=X2Y=X^2. Find the density of YY.

Solution

For 0<y<10\lt y\lt1, the inverse branches are x±=±yx_\pm=\pm\sqrt y, and ∣d(x2)/dx∣=2y\lvert d(x^2)/dx\rvert=2\sqrt y on either branch. Since fX=1/2f_X=1/2,

fY(y)=2(1/22y)=12y,0<y<1.f_Y(y) =2\left(\frac{1/2}{2\sqrt y}\right) =\frac{1}{2\sqrt y}, \qquad 0\lt y\lt1.

It vanishes elsewhere, and ∫01fY(y) dy=1\int_0^1f_Y(y)\,dy=1.

A condition occurs in 1%1\% of a population. A test has 99%99\% sensitivity and 95%95\% specificity. Find the probability that a person with a positive result has the condition.

Solution

Let CC denote the condition and ++ a positive result. Then

P(C∣+)=0.99(0.01)0.99(0.01)+0.05(0.99)=0.00990.0594=16.\mathbb P(C\mid+) =\frac{0.99(0.01)} {0.99(0.01)+0.05(0.99)} =\frac{0.0099}{0.0594} =\frac16.

Despite high sensitivity, false positives from the much larger unaffected population dominate unless the base rate is included.

3. Gaussian moments from the characteristic function

Section titled “3. Gaussian moments from the characteristic function”

For χ(t)=exp⁡(iμt−σ2t2/2)\chi(t)=\exp(i\mu t-\sigma^2t^2/2), recover E[X]\mathbb E[X] and Var⁡(X)\operatorname{Var}(X).

Solution

The moment identities give

E[X]=1iχ′(0)=μ,\mathbb E[X] =\frac{1}{i}\chi'(0) =\mu,

and

E[X2]=−χ′′(0)=μ2+σ2.\mathbb E[X^2] =-\chi''(0) =\mu^2+\sigma^2.

Therefore Var⁡(X)=E[X2]−μ2=σ2\operatorname{Var}(X)=\mathbb E[X^2]-\mu^2=\sigma^2.

An independent-sample Monte Carlo estimate has standard error s/Ns/\sqrt N. By what factor must NN increase to reduce the standard error by a factor of ten?

Solution

Because the error scales as N−1/2N^{-1/2},

1Nnew=1101Nold,\frac{1}{\sqrt{N_{\mathrm{new}}}} =\frac1{10} \frac{1}{\sqrt{N_{\mathrm{old}}}},

so Nnew=100NoldN_{\mathrm{new}}=100N_{\mathrm{old}}. Correlations would require replacing NN by an effective sample size.

  • P. Billingsley, Probability and Measure, 3rd ed., Wiley, 1995.
  • G. Casella and R. L. Berger, Statistical Inference, 2nd ed., Duxbury, 2002.
  • T. M. Cover and J. A. Thomas, Elements of Information Theory, 2nd ed., Wiley, 2006.
  • A. S. Holevo, Probabilistic and Statistical Aspects of Quantum Theory, 2nd ed., Edizioni della Normale, 2011.
  • C. P. Robert and G. Casella, Monte Carlo Statistical Methods, 2nd ed., Springer, 2004.
  • A. W. van der Vaart, Asymptotic Statistics, Cambridge University Press, 1998.