Probability, Statistics, and Information
Probability theory supplies the language of outcomes, distributions, averages, fluctuations, inference, and sampling error. Quantum mechanics uses that language every time a measurement is fixed, but it does not reduce all observables to random variables on one universal classical sample space. This chapter develops the classical machinery first and marks the quantum boundary only after the machinery is secure.
The chapter has four strands: probability models and random variables; moments and conditioning; information and estimation; and Monte Carlo computation. The final comparison page explains which parts carry directly into a quantum experiment and which parts are changed by amplitudes, noncommutativity, and measurement disturbance.
From outcomes to distributions
Section titled “From outcomes to distributions”A probability space is a triple
where is the outcome space, is a collection of measurable events, and is a normalized countably additive measure. A random variable is a measurable map
Its probability distribution is the pushforward measure
The distribution can often be used without retaining the microscopic outcome . This separation among outcome, event, random variable, and value prevents several common category errors.
For a discrete variable, probabilities are masses . For a continuous variable with density relative to ,
A density is probability per unit measure and generally has physical units. The value is not the probability of the exact point . Probability Spaces, Light Version, Random Variables, and Probability Densities form the foundational route.
Coordinate changes need Jacobians
Section titled “Coordinate changes need Jacobians”If and is differentiable with simple inverse branches, then
The sum over branches is essential when is not one-to-one. In several dimensions, the absolute determinant of the inverse Jacobian replaces . A density’s numerical value and units change under reparameterization even though probabilities of corresponding events do not.
This same principle explains why and use different measures and why radial densities acquire volume factors. The Born rule itself remains canonical in Probability and the Born Rule.
Expectations, variance, and covariance
Section titled “Expectations, variance, and covariance”For an integrable function ,
For real-valued variables, the mean is . Variance and covariance are
and
Variance measures mean-square spread, not general unpredictability. Zero covariance does not imply independence except under additional assumptions, notably for jointly Gaussian variables. Covariance matrices must be positive semidefinite because
Use Expectation Values for integration and linearity, then Variance and Covariance for fluctuations and correlations. Quantum expectation values and uncertainty relations have separate canonical homes in Core Formalism.
Conditioning and Bayes inference
Section titled “Conditioning and Bayes inference”For ,
Bayes’ rule reverses a conditional probability:
For a parameter and observed data ,
The likelihood is not a normalized density over until combined with a prior and normalized. Base rates can dominate a seemingly accurate diagnostic when the target event is rare.
Conditional Probability owns conditioning of events and densities. Bayes’ Rule owns inference and parameter updates. Quantum state update may have a conditioning aspect, but a quantum instrument also specifies outcome-dependent physical transformation; it is not determined by Bayes’ rule alone.
Characteristic functions and Gaussians
Section titled “Characteristic functions and Gaussians”The characteristic function
is the Fourier transform of a probability distribution under the probability convention. When the required moments exist,
For independent variables, , turning convolution of distributions into multiplication. Characteristic functions always exist for real random variables because , even when ordinary moments do not.
A one-dimensional Gaussian with mean and variance has density
and characteristic function
Multivariate Gaussians are controlled by a mean vector and positive-semidefinite covariance matrix, with special care needed for singular covariance. Characteristic Functions and Gaussian Distributions supply the detailed route.
Entropy and relative entropy
Section titled “Entropy and relative entropy”For a discrete distribution , Shannon entropy is
The logarithm base sets the unit: base two gives bits and the natural logarithm gives nats. Entropy quantifies average uncertainty under a specified distribution; it is not a universal measure of disorder detached from a model.
For a continuous density,
is differential entropy. It depends on coordinates and reference measure, can be negative, and is not the direct continuous analogue of nonnegative discrete entropy.
Relative entropy compares two distributions:
with when assigns positive mass where assigns zero mass. It is generally asymmetric and is not a metric. Unlike differential entropy alone, relative entropy is invariant under smooth one-to-one coordinate changes when both densities transform against the same reference measure.
Entropy and Relative Entropy own the classical definitions. Von Neumann entropy, quantum relative entropy, and entanglement entropy remain in the density-operator and composite-system volumes.
Fisher information and estimation
Section titled “Fisher information and estimation”For a regular parameterized density , the score is
and the Fisher information is
Under standard regularity assumptions, the score has zero mean. For an unbiased estimator based on independent observations, the Cramér–Rao bound gives
The bound requires its stated assumptions and need not be attained at finite sample size. Fisher information is local in parameter space and transforms as a metric tensor under smooth reparameterization. Fisher Information develops the classical result and links onward to quantum estimation without duplicating quantum Fisher information.
Monte Carlo estimates
Section titled “Monte Carlo estimates”For independent samples with finite variance, the sample mean
is unbiased for and has
The standard error therefore scales as , not . Correlated Markov-chain samples have a smaller effective sample size determined by autocorrelation. Importance sampling can reduce variance by sampling from a proposal that covers the important regions, but support mismatch or highly variable weights can make an estimator unstable.
Monte Carlo Basics owns estimators, uncertainty, importance sampling, and autocorrelation cautions. Numerical implementation, convergence tests, and benchmark design remain in Numerical Mathematics.
Classical and quantum probability
Section titled “Classical and quantum probability”For a fixed quantum state and POVM ,
is an ordinary classical probability distribution over the outcome labels . Once the experiment is specified, classical expectation, likelihood, entropy, and sampling tools apply to its recorded outcomes.
The difference appears when one asks for one joint classical model covering all possible quantum measurements.
| Classical probability | Quantum experiment |
|---|---|
| one sample space and event algebra | each measurement supplies an outcome algebra; incompatible measurements need not share a joint distribution |
| random variable on the sample space | self-adjoint operator or POVM specifying an outcome law |
| law of total probability for exclusive alternatives | amplitudes can interfere before probabilities are formed |
| conditioning changes information about a fixed model | an instrument can also disturb the physical state |
| joint variables exist by construction | joint measurability imposes nontrivial compatibility conditions |
Commuting projective observables admit a common spectral measure and behave like jointly distributed classical variables for that context. Noncommuting observables generally do not. Classical Probability versus Quantum Probability gives the careful comparison; the Born rule, sequential measurement, and state update remain canonical in Core Formalism.
Page map
Section titled “Page map”| Page | Central question |
|---|---|
| Probability Spaces, Light Version | What are outcomes, events, measures, and distributions? |
| Random Variables | How does a measurable map turn outcomes into values? |
| Probability Densities | How is probability represented relative to a continuous measure? |
| Expectation Values | How are probability-weighted averages defined and manipulated? |
| Variance and Covariance | How are spread and linear dependence quantified? |
| Conditional Probability | How does restricting to known information change a distribution? |
| Bayes’ Rule | How are likelihood and prior combined into a posterior? |
| Characteristic Functions | How does Fourier analysis encode distributions and independent sums? |
| Gaussian Distributions | Why do means and covariances completely determine Gaussian laws? |
| Entropy | How is average information or uncertainty quantified? |
| Relative Entropy | How is one distribution compared with another? |
| Fisher Information | How much local parameter sensitivity does a statistical model contain? |
| Monte Carlo Basics | How are expectations estimated from samples with controlled uncertainty? |
| Classical Probability versus Quantum Probability | Which classical rules apply to fixed measurements, and where does quantum structure exceed them? |
Suggested routes
Section titled “Suggested routes”- Born-rule calculations: probability spaces random variables densities expectation variance, then Probability and the Born Rule.
- Inference and tomography: conditioning Bayes’ rule Fisher information, followed by the measurement and estimation pages for the experiment at hand.
- Information theory: entropy relative entropy classical-versus-quantum probability, then Entropy Overview.
- Stochastic computation: Gaussian distributions expectation and variance Monte Carlo, then Error Estimates and Convergence Tests.
Common mistakes
Section titled “Common mistakes”| Mistake | Correction |
|---|---|
| Treating a density value as a point probability | integrate the density over an event and track its units |
| Changing variables without a Jacobian or inverse branches | transform both density and measure and sum over all preimages |
| Assuming zero covariance implies independence | this requires additional structure, such as joint Gaussianity |
| Reversing into without a prior | apply Bayes’ rule and normalize |
| Comparing differential entropies across coordinates as absolute quantities | use a common reference measure or relative entropy |
| Reading the Cramér–Rao bound without its regularity and bias assumptions | state the model, estimator class, and sample conditions |
| Reporting Monte Carlo digits without a standard error or autocorrelation analysis | estimate uncertainty and effective sample size |
| Treating quantum state update as ordinary conditioning alone | specify the quantum instrument and its disturbance |
Exercises
Section titled “Exercises”1. Density under a many-to-one map
Section titled “1. Density under a many-to-one map”Let be uniform on and let . Find the density of .
Solution
For , the inverse branches are , and on either branch. Since ,
It vanishes elsewhere, and .
2. A base-rate calculation
Section titled “2. A base-rate calculation”A condition occurs in of a population. A test has sensitivity and specificity. Find the probability that a person with a positive result has the condition.
Solution
Let denote the condition and a positive result. Then
Despite high sensitivity, false positives from the much larger unaffected population dominate unless the base rate is included.
3. Gaussian moments from the characteristic function
Section titled “3. Gaussian moments from the characteristic function”For , recover and .
Solution
The moment identities give
and
Therefore .
4. Monte Carlo cost
Section titled “4. Monte Carlo cost”An independent-sample Monte Carlo estimate has standard error . By what factor must increase to reduce the standard error by a factor of ten?
Solution
Because the error scales as ,
so . Correlations would require replacing by an effective sample size.
References
Section titled “References”- P. Billingsley, Probability and Measure, 3rd ed., Wiley, 1995.
- G. Casella and R. L. Berger, Statistical Inference, 2nd ed., Duxbury, 2002.
- T. M. Cover and J. A. Thomas, Elements of Information Theory, 2nd ed., Wiley, 2006.
- A. S. Holevo, Probabilistic and Statistical Aspects of Quantum Theory, 2nd ed., Edizioni della Normale, 2011.
- C. P. Robert and G. Casella, Monte Carlo Statistical Methods, 2nd ed., Springer, 2004.
- A. W. van der Vaart, Asymptotic Statistics, Cambridge University Press, 1998.