Skip to content

Classical Probability versus Quantum Probability

Quantum mechanics does not reject ordinary probability theory. Once a state and a measurement have been specified, the outcomes form an ordinary probability experiment: there is a sample space of possible results, events are sets of results, and the Born rule assigns probabilities.

The difference is structural. Classical probability often begins with one underlying sample space on which many random variables are defined simultaneously. Quantum mechanics attaches probabilities to measurement contexts. Compatible measurements can share a joint distribution, but incompatible sharp observables generally cannot be represented as ordinary random variables on one universal sample space while preserving all quantum predictions.

This page is the comparison layer. The basic vocabulary is in Probability Spaces, Light Version, and the physics rule is in Born Rule.

In classical probability one starts with a probability space

(Ω,F,P).(\Omega,\mathcal F,\mathbb P).

A random variable is a measurable function

X:Ω→R.X:\Omega\to\mathbb R.

If XX and YY are both defined on the same Ω\Omega, then they automatically have a joint distribution. For value sets Δ\Delta and Γ\Gamma,

P(X∈Δ, Y∈Γ)=P({ω∈Ω:X(ω)∈Δ, Y(ω)∈Γ}).\mathbb P(X\in\Delta,\ Y\in\Gamma) = \mathbb P \bigl( \{\omega\in\Omega: X(\omega)\in\Delta,\ Y(\omega)\in\Gamma\} \bigr).

The marginals are recovered from the joint distribution:

P(X∈Δ)=P(X∈Δ, Y∈R),\mathbb P(X\in\Delta) = \mathbb P(X\in\Delta,\ Y\in\mathbb R),

and conditional probabilities, covariance, and correlation coefficients are meaningful because the joint probability law exists.

This setup supports a familiar picture: a single underlying outcome ω\omega carries definite values of many quantities at once, even if the observer does not know them.

Quantum Measurements Produce Ordinary Probabilities

Section titled “Quantum Measurements Produce Ordinary Probabilities”

For a projective measurement with projectors {Pa}\{P_a\} and a normalized pure state ∣ψ⟩\lvert\psi\rangle, the Born rule gives

p(a)=⟨ψ∣Pa∣ψ⟩.p(a) = \langle\psi\rvert P_a\lvert\psi\rangle.

For a density operator ρ\rho,

p(a)=Tr⁡(ρPa).p(a) = \operatorname{Tr}(\rho P_a).

These numbers are ordinary probabilities:

p(a)≥0,∑ap(a)=1.p(a)\ge0, \qquad \sum_a p(a)=1.

For this specified measurement, one may take the sample space to be the outcome set

ΩA={a}.\Omega_A=\{a\}.

The reported outcome is then a classical random variable on that experiment’s outcome space. In this limited but essential sense, quantum probability is still probability.

The same statement holds for generalized measurements. If {Ea}\{E_a\} is a POVM, then

p(a)=Tr⁡(ρEa),Ea≥0,∑aEa=I.p(a) = \operatorname{Tr}(\rho E_a), \qquad E_a\ge0, \qquad \sum_a E_a=I.

The special quantum content is not the additivity of probabilities after the measurement is fixed. It is how the probability measure is generated and how different measurement contexts relate to one another.

Compatible Observables and Joint Distributions

Section titled “Compatible Observables and Joint Distributions”

When sharp observables are compatible, quantum mechanics does allow a joint distribution.

Let

A=∑aaPa,B=∑bbQbA=\sum_a aP_a, \qquad B=\sum_b bQ_b

be two discrete observables. If every projector PaP_a commutes with every projector QbQ_b, then

PaQb=QbPaP_aQ_b=Q_bP_a

and PaQbP_aQ_b is itself a projector onto the joint event “outcome aa for AA and outcome bb for BB.” The joint probability is

p(a,b)=Tr⁡(ρPaQb).p(a,b) = \operatorname{Tr}(\rho P_aQ_b).

For a pure state,

p(a,b)=⟨ψ∣PaQb∣ψ⟩=∥PaQb∣ψ⟩∥2.p(a,b) = \langle\psi\rvert P_aQ_b\lvert\psi\rangle = \lVert P_aQ_b\lvert\psi\rangle\rVert^2.

Marginalizing gives the separate Born probabilities:

∑bp(a,b)=Tr⁡(ρPa),∑ap(a,b)=Tr⁡(ρQb).\sum_b p(a,b) = \operatorname{Tr}(\rho P_a), \qquad \sum_a p(a,b) = \operatorname{Tr}(\rho Q_b).

This is the quantum version of an ordinary joint distribution. The finite-dimensional criterion and its caveats are developed in Compatible Observables.

For noncommuting sharp observables, the product PaQbP_aQ_b generally is not a projector onto a symmetric joint event. The expression

Tr⁡(ρPaQb)\operatorname{Tr}(\rho P_aQ_b)

need not behave like a classical joint probability, and there is generally no single joint distribution whose marginals reproduce all the Born distributions for all incompatible observables.

Ordered measurements do have probabilities, but the order is part of the experiment. Measuring AA and then BB gives, for ideal projective measurements,

p(a then b)=Tr⁡(QbPaρPaQb).p(a\text{ then }b) = \operatorname{Tr}(Q_bP_a\rho P_aQ_b).

Reversing the order gives

p(b then a)=Tr⁡(PaQbρQbPa).p(b\text{ then }a) = \operatorname{Tr}(P_aQ_b\rho Q_bP_a).

These are probabilities for two different experiments. They are not two ways of reading the same underlying classical joint table. The operational treatment is in Sequential Measurements.

Consider a spin-1/21/2 state prepared as ∣+z⟩\lvert +z\rangle. A measurement of σz\sigma_z gives

p(z=+)=1,p(z=−)=0.p(z=+)=1, \qquad p(z=-)=0.

A measurement of σx\sigma_x on the same prepared state gives

p(x=+)=12,p(x=−)=12.p(x=+)=\frac12, \qquad p(x=-)=\frac12.

It is tempting to imagine that the particle simply had a definite zz value and an unknown xx value all along. That picture cannot be promoted without care into one global classical model for all spin directions.

If one first measures σx\sigma_x and then measures σz\sigma_z, the intermediate measurement changes the state used for the second probability. The ordered joint probabilities include

p(x=+ then z=+)=14,p(x=− then z=+)=14.p(x=+\text{ then }z=+) = \frac14, \qquad p(x=-\text{ then }z=+) = \frac14.

Thus, after an unrecorded sharp σx\sigma_x measurement, the later probability of z=+z=+ is 1/21/2, not the original value 11. This example shows why measurement context and order matter. It is not, by itself, a proof of contextuality; a classical invasive measurement can also disturb a system. Contextuality theorems impose sharper assumptions.

Classical alternatives that are mutually exclusive add as probabilities. If A1A_1 and A2A_2 are disjoint alternatives leading to an event DD, then

P(D)=P(D∩A1)+P(D∩A2)\mathbb P(D) = \mathbb P(D\cap A_1) + \mathbb P(D\cap A_2)

when the alternatives are part of the same classical probability space.

Quantum alternatives can add as amplitudes when no measurement record distinguishes them. If two alternatives contribute amplitudes α1\alpha_1 and α2\alpha_2 to the same final outcome, then

p=∣α1+α2∣2=∣α1∣2+∣α2∣2+2Re⁡(α1∗α2).\begin{aligned} p &= \lvert\alpha_1+\alpha_2\rvert^2\\ &= \lvert\alpha_1\rvert^2 + \lvert\alpha_2\rvert^2 + 2\operatorname{Re}(\alpha_1^*\alpha_2). \end{aligned}

The last term is the interference term. It can be positive, negative, or zero.

If a which-alternative record is physically available and remains correlated with the alternatives, the interference term is suppressed in the probabilities accessible to the observed subsystem. That transition is part of the Decoherence Preview, while the local rule “add amplitudes before taking squared moduli” is introduced in Probability Amplitudes.

A measurement context is the full compatible arrangement used to define an outcome: the observable, any commuting observables measured with it, the POVM or projective decomposition, and the experimental arrangement that realizes the measurement.

A noncontextual hidden-variable model tries to assign outcomes to observables in a way that is independent of which compatible context is used to measure them. For projectors, the rough idea is to assign each projector a value

v(P)∈{0,1}v(P)\in\{0,1\}

while preserving the rule that exactly one mutually exclusive outcome in a complete projective measurement occurs.

Kochen–Specker-type theorems show that, in Hilbert spaces of dimension at least three, such noncontextual value assignments cannot reproduce the full projective structure under the theorem’s assumptions. Bell-type theorems show different no-go constraints on local hidden-variable models for entangled systems. These are settled mathematical results under stated assumptions; their philosophical interpretation is a separate matter.

The practical lesson for this page is modest: do not treat all quantum observables as ordinary random variables on one hidden sample space unless the extra model and its assumptions have been stated.

Classical Limit and Effective Classicality

Section titled “Classical Limit and Effective Classicality”

Many quantum situations admit excellent effective classical descriptions. A narrow wave packet may follow approximately classical equations for a while. A macroscopic pointer can have robust, nearly exclusive records. Decoherence can make interference between coarse alternatives inaccessible for practical purposes.

Effective classicality is not the same as a universal classical sample space for every observable. It is a regime-dependent approximation in which selected variables, coarse grainings, or records behave classically enough for the question being asked.

This distinction is useful in statistical mechanics, measurement theory, and quantum information: classical probabilities can describe records, ignorance, and ensembles, while the underlying quantum state still carries phase relations, noncommuting observables, and entanglement structure.

  • Saying quantum probabilities are “not real probabilities.” For a fixed measurement, they are ordinary probabilities.
  • Treating amplitudes as probabilities. Probabilities come from squared moduli or trace formulas.
  • Adding probabilities when indistinguishable quantum alternatives require adding amplitudes.
  • Assigning a joint probability distribution to noncommuting observables without specifying a joint measurement or an additional model.
  • Confusing measurement disturbance with the whole content of incompatibility or contextuality.
  • Reading a density operator only as classical ignorance. Mixed states can arise from classical preparation uncertainty, entanglement with an environment, or both.
  • Treating a quasiprobability representation as an ordinary probability distribution when it takes negative or otherwise nonclassical values.
  • A. N. Kolmogorov, Foundations of the Theory of Probability, 2nd English ed., Chelsea, 1956.
  • W. Feller, An Introduction to Probability Theory and Its Applications, Volume I, 3rd ed., Wiley, 1968.
  • J. von Neumann, Mathematical Foundations of Quantum Mechanics, Princeton University Press, 1955.
  • P. A. M. Dirac, The Principles of Quantum Mechanics, 4th ed., Oxford University Press, 1958.
  • J. S. Bell, “On the Einstein Podolsky Rosen Paradox,” Physics Physique Fizika 1, 195-200, 1964.
  • S. Kochen and E. P. Specker, “The Problem of Hidden Variables in Quantum Mechanics,” Journal of Mathematics and Mechanics 17, 59-87, 1967.
  • A. Peres, Quantum Theory: Concepts and Methods, Kluwer, 1995.
  • L. E. Ballentine, Quantum Mechanics: A Modern Development, 2nd ed., World Scientific, 2014.
  • M. A. Nielsen and I. L. Chuang, Quantum Computation and Quantum Information, Cambridge University Press, 2010.
  1. Let XX and YY be classical random variables on the same probability space. Show that the joint distribution marginalizes to the distribution of XX.
Solution

For a value set Δ\Delta,

P(X∈Δ)=P(X∈Δ, Y∈R)=P({ω:X(ω)∈Δ, Y(ω)∈R}).\begin{aligned} \mathbb P(X\in\Delta) &= \mathbb P(X\in\Delta,\ Y\in\mathbb R)\\ &= \mathbb P \bigl( \{\omega:X(\omega)\in\Delta,\ Y(\omega)\in\mathbb R\} \bigr). \end{aligned}

Since Y(ω)∈RY(\omega)\in\mathbb R for every outcome where YY is defined, this is just the event {ω:X(ω)∈Δ}\{\omega:X(\omega)\in\Delta\}.

  1. Suppose PP and QQ are commuting projectors. Show that PQPQ is a projector and explain why ⟨ψ∣PQ∣ψ⟩≥0\langle\psi\rvert PQ\lvert\psi\rangle\ge0 for every state ∣ψ⟩\lvert\psi\rangle.
Solution

Since P2=PP^2=P, Q2=QQ^2=Q, and PQ=QPPQ=QP,

(PQ)2=PQPQ=P2Q2=PQ.(PQ)^2 = PQPQ = P^2Q^2 = PQ.

Also,

(PQ)†=Q†P†=QP=PQ.(PQ)^\dagger = Q^\dagger P^\dagger = QP = PQ.

Thus PQPQ is an orthogonal projector. Therefore

⟨ψ∣PQ∣ψ⟩=∥PQ∣ψ⟩∥2≥0.\langle\psi\rvert PQ\lvert\psi\rangle = \lVert PQ\lvert\psi\rangle\rVert^2 \ge0.
  1. Let two indistinguishable alternatives contribute amplitudes α1\alpha_1 and α2\alpha_2. Expand ∣α1+α2∣2\lvert\alpha_1+\alpha_2\rvert^2 and state when the classical sum of probabilities is recovered.
Solution

The expansion is

∣α1+α2∣2=∣α1∣2+∣α2∣2+2Re⁡(α1∗α2).\lvert\alpha_1+\alpha_2\rvert^2 = \lvert\alpha_1\rvert^2 + \lvert\alpha_2\rvert^2 + 2\operatorname{Re}(\alpha_1^*\alpha_2).

The classical sum is recovered when the interference term vanishes or when a physical which-alternative record makes the cross term inaccessible in the observed probabilities.

  1. A spin-1/21/2 system starts in ∣+z⟩\lvert +z\rangle. A sharp σx\sigma_x measurement is performed and the outcome is ignored. What is the probability of obtaining z=+z=+ in a later sharp σz\sigma_z measurement?
Solution

The first measurement gives ∣+x⟩\lvert +x\rangle or ∣−x⟩\lvert -x\rangle with probability 1/21/2 each. From either xx eigenstate, the later σz\sigma_z measurement gives z=+z=+ with probability 1/21/2. Therefore

p(z=+ after ignored x measurement)=12⋅12+12⋅12=12.p(z=+\text{ after ignored }x\text{ measurement}) = \frac12\cdot\frac12 + \frac12\cdot\frac12 = \frac12.

This differs from measuring σz\sigma_z directly on ∣+z⟩\lvert +z\rangle, which gives z=+z=+ with probability 11.

  1. Why do two separate marginal distributions for noncommuting observables not automatically define a joint distribution?
Solution

Marginal distributions give probabilities for two separately specified experiments. A joint distribution is stronger: it assigns probabilities to simultaneous value pairs and must marginalize consistently. For noncommuting sharp quantum observables, there may be no measurement-independent joint event corresponding to “value aa of AA and value bb of BB.” One needs a compatible joint measurement, an ordered sequential experiment, an unsharp POVM construction, or an additional hidden-variable model with stated assumptions.