Skip to content

Conditional Probability

Conditional probability is the probability of one event after another event is known to have occurred.

If BB has nonzero probability, then the conditional probability of AA given BB is

P(A∣B)=P(A∩B)P(B).\mathbb P(A\mid B) = \frac{\mathbb P(A\cap B)}{\mathbb P(B)}.

The event BB has become the new reference population. Conditioning is not a new probability rule added to the axioms; it is ordinary probability restricted and renormalized on the event being treated as known.

Quantum mechanics uses conditional probabilities constantly, especially in sequential measurements. The caution, developed in Classical Probability versus Quantum Probability, is that quantum state update is not merely classical conditioning on pre-existing values of all observables.

Let (Ω,F,P)(\Omega,\mathcal F,\mathbb P) be a probability space. For events A,B∈FA,B\in\mathcal F with P(B)>0\mathbb P(B)>0,

P(A∣B)=P(A∩B)P(B).\mathbb P(A\mid B) = \frac{\mathbb P(A\cap B)}{\mathbb P(B)}.

This formula has three immediate consequences:

P(B∣B)=1,\mathbb P(B\mid B)=1, P(Ac∣B)=1−P(A∣B),\mathbb P(A^c\mid B) = 1-\mathbb P(A\mid B),

and, for disjoint events A1,A2,…A_1,A_2,\ldots,

P(⋃iAi∣B)=∑iP(Ai∣B).\mathbb P \left( \bigcup_i A_i \mid B \right) = \sum_i \mathbb P(A_i\mid B).

For fixed BB, conditional probability is itself a probability measure on the same event collection.

Rearranging the definition gives the product rule:

P(A∩B)=P(A∣B)P(B).\mathbb P(A\cap B) = \mathbb P(A\mid B)\mathbb P(B).

Equally,

P(A∩B)=P(B∣A)P(A),\mathbb P(A\cap B) = \mathbb P(B\mid A)\mathbb P(A),

when P(A)>0\mathbb P(A)>0.

This symmetry is the starting point for Bayes’ Rule. The full inference-focused treatment is a separate page; here the product rule is enough for sequential-probability calculations.

Suppose B1,…,BnB_1,\ldots,B_n are mutually exclusive and exhaustive events:

Bi∩Bj=∅(i≠j),⋃i=1nBi=Ω.B_i\cap B_j=\emptyset \quad (i\ne j), \qquad \bigcup_{i=1}^n B_i=\Omega.

If each P(Bi)>0\mathbb P(B_i)>0, then

P(A)=∑iP(A∣Bi)P(Bi).\mathbb P(A) = \sum_i \mathbb P(A\mid B_i)\mathbb P(B_i).

This is the law of total probability. It says that an unconditional probability can be recovered by averaging conditional probabilities over the cases being conditioned on.

In quantum language, a similar distinction appears between selective and nonselective measurement: keeping an outcome produces a conditional description, while forgetting outcomes averages over them.

For a random variable XX, conditioning on an event BB gives a conditional distribution:

P(X∈Δ∣B)=P({X∈Δ}∩B)P(B).\mathbb P(X\in\Delta\mid B) = \frac{ \mathbb P(\{X\in\Delta\}\cap B) }{ \mathbb P(B) }.

The conditional expectation, when it exists, is the expectation computed using this conditional distribution.

For a discrete variable with values xix_i,

E[X∣B]=∑ixi P(X=xi∣B).\mathbb E[X\mid B] = \sum_i x_i\,\mathbb P(X=x_i\mid B).

For a continuous variable with conditional density fX∣B(x)f_{X\mid B}(x),

E[X∣B]=∫−∞∞xfX∣B(x) dx.\mathbb E[X\mid B] = \int_{-\infty}^{\infty} x f_{X\mid B}(x)\,dx.

For two continuous random variables with joint density fX,Y(x,y)f_{X,Y}(x,y), the marginal density of YY is

fY(y)=∫−∞∞fX,Y(x,y) dx.f_Y(y) = \int_{-\infty}^{\infty} f_{X,Y}(x,y)\,dx.

When fY(y)>0f_Y(y)>0, the conditional density of XX given Y=yY=y is

fX∣Y(x∣y)=fX,Y(x,y)fY(y).f_{X\mid Y}(x\mid y) = \frac{f_{X,Y}(x,y)}{f_Y(y)}.

It is normalized in xx:

∫−∞∞fX∣Y(x∣y) dx=1.\int_{-\infty}^{\infty} f_{X\mid Y}(x\mid y)\,dx = 1.

The joint density factors as

fX,Y(x,y)=fX∣Y(x∣y)fY(y).f_{X,Y}(x,y) = f_{X\mid Y}(x\mid y)f_Y(y).

This is the continuous analogue of the product rule.

For continuous YY, the event Y=yY=y usually has probability zero:

P(Y=y)=0.\mathbb P(Y=y)=0.

So fX∣Y(x∣y)f_{X\mid Y}(x\mid y) is not obtained by directly dividing by P(Y=y)\mathbb P(Y=y). It is defined through densities, limits, or more advanced regular conditional probability language.

Operationally, conditioning on Y=yY=y means an idealized version of conditioning on a very narrow bin

y−Δy2≤Y≤y+Δy2y-\frac{\Delta y}{2} \le Y\le y+\frac{\Delta y}{2}

and taking the narrow-bin limit when it exists.

This distinction is important in continuous quantum measurements and scattering, where exact continuous outcomes are idealizations and physical detectors have finite resolution.

Events AA and BB are independent if

P(A∩B)=P(A)P(B).\mathbb P(A\cap B) = \mathbb P(A)\mathbb P(B).

If P(B)>0\mathbb P(B)>0, this is equivalent to

P(A∣B)=P(A).\mathbb P(A\mid B) = \mathbb P(A).

For random variables, independence means that their joint distribution factors. For continuous variables,

fX,Y(x,y)=fX(x)fY(y).f_{X,Y}(x,y) = f_X(x)f_Y(y).

Independence implies zero covariance when the second moments exist, but zero covariance does not imply independence; see Variance and Covariance.

For a partition BiB_i with positive probabilities,

E[X]=∑iE[X∣Bi]P(Bi),\mathbb E[X] = \sum_i \mathbb E[X\mid B_i]\mathbb P(B_i),

when the expectations exist.

For continuous conditioning,

E[X]=∫−∞∞E[X∣Y=y] fY(y) dy.\mathbb E[X] = \int_{-\infty}^{\infty} \mathbb E[X\mid Y=y]\,f_Y(y)\,dy.

This is the expectation-value version of total probability. It is a common way to separate a calculation into cases.

Classical Conditioning versus Quantum State Update

Section titled “Classical Conditioning versus Quantum State Update”

Classical conditioning updates probabilities after information is learned:

P(A)⟶P(A∣B).\mathbb P(A) \longrightarrow \mathbb P(A\mid B).

The event space is the same; one is restricting attention to outcomes where BB occurred.

In an ideal quantum projective measurement, the selective state update after outcome aa is

ρ⟶PaρPaTr⁡(ρPa).\rho \longrightarrow \frac{P_a\rho P_a}{\operatorname{Tr}(\rho P_a)}.

The denominator is the probability of the outcome:

p(a)=Tr⁡(ρPa).p(a)=\operatorname{Tr}(\rho P_a).

This resembles conditioning because an outcome is selected and the state is renormalized. But it is not simply conditioning on pre-existing values of all observables. The measurement context matters, the state used for later measurements can change, and noncommuting observables generally do not share one classical joint distribution.

The formal quantum rule is developed in State Update Rule. Ordered probabilities are developed in Sequential Measurements.

For projective measurements with projectors PaP_a followed by QbQ_b, the conditional probability of bb after obtaining aa is

p(b∣a)=Tr⁡(QbPaρPa)Tr⁡(ρPa),p(b\mid a) = \frac{ \operatorname{Tr}(Q_bP_a\rho P_a) }{ \operatorname{Tr}(\rho P_a) },

when p(a)>0p(a)>0. The ordered joint probability is

p(a then b)=p(b∣a)p(a).p(a\ \text{then}\ b) = p(b\mid a)p(a).

The product rule still appears. What is special is the quantum rule assigning the post-outcome state used to compute the later conditional probability.

For conditional states of subsystems after a measurement on another subsystem, see Conditional States.

  • Writing P(A∣B)\mathbb P(A\mid B) when P(B)=0\mathbb P(B)=0 without specifying a limiting or density meaning.
  • Treating P(A∣B)\mathbb P(A\mid B) and P(B∣A)\mathbb P(B\mid A) as interchangeable.
  • Forgetting that conditional probabilities must still sum or integrate to one over the conditioned sample space.
  • Computing covariance or conditional probability without a joint distribution.
  • Treating independence as the same thing as zero covariance.
  • Treating quantum state update as ordinary classical conditioning on hidden pre-existing values.
  • Forgetting that nonselective measurement averages over outcomes rather than conditioning on one outcome.
  • W. Feller, An Introduction to Probability Theory and Its Applications, Volume I, 3rd ed., Wiley, 1968.
  • P. Billingsley, Probability and Measure, 3rd ed., Wiley, 1995.
  • R. Durrett, Probability: Theory and Examples, 5th ed., Cambridge University Press, 2019.
  • A. Peres, Quantum Theory: Concepts and Methods, Kluwer, 1995.
  • M. A. Nielsen and I. L. Chuang, Quantum Computation and Quantum Information, Cambridge University Press, 2010.
  1. A fair die is rolled. Let AA be the event “the result is at least 4” and BB the event “the result is even.” Compute P(A∣B)\mathbb P(A\mid B).
Solution

The even outcomes are

B={2,4,6}.B=\{2,4,6\}.

Among these, the outcomes at least 44 are {4,6}\{4,6\}. Therefore

P(A∣B)=23.\mathbb P(A\mid B) = \frac{2}{3}.
  1. Suppose P(B)=0.2\mathbb P(B)=0.2 and P(A∣B)=0.7\mathbb P(A\mid B)=0.7. Find P(A∩B)\mathbb P(A\cap B).
Solution

Use the product rule:

P(A∩B)=P(A∣B)P(B)=0.7⋅0.2=0.14.\mathbb P(A\cap B) = \mathbb P(A\mid B)\mathbb P(B) = 0.7\cdot0.2 = 0.14.
  1. Let X,YX,Y have joint density fX,Y(x,y)=2f_{X,Y}(x,y)=2 on the triangle 0<x<y<10\lt x\lt y\lt1 and zero elsewhere. Find fY(y)f_Y(y) and fX∣Y(x∣y)f_{X\mid Y}(x\mid y).
Solution

For 0<y<10\lt y\lt1,

fY(y)=∫0y2 dx=2y.f_Y(y) = \int_0^y 2\,dx = 2y.

Therefore, for 0<x<y<10\lt x\lt y\lt1,

fX∣Y(x∣y)=22y=1y.f_{X\mid Y}(x\mid y) = \frac{2}{2y} = \frac1y.

It is uniform on the interval 0<x<y0\lt x\lt y.

  1. If AA and BB are independent and P(B)>0\mathbb P(B)>0, show that P(A∣B)=P(A)\mathbb P(A\mid B)=\mathbb P(A).
Solution

Independence gives

P(A∩B)=P(A)P(B).\mathbb P(A\cap B) = \mathbb P(A)\mathbb P(B).

Thus

P(A∣B)=P(A∩B)P(B)=P(A).\mathbb P(A\mid B) = \frac{\mathbb P(A\cap B)}{\mathbb P(B)} = \mathbb P(A).
  1. In an ideal projective measurement with density operator ρ\rho, why does the selective update divide by Tr⁡(ρPa)\operatorname{Tr}(\rho P_a)?
Solution

The unnormalized post-outcome operator is PaρPaP_a\rho P_a. Its trace is

Tr⁡(PaρPa)=Tr⁡(ρPa),\operatorname{Tr}(P_a\rho P_a) = \operatorname{Tr}(\rho P_a),

using cyclicity of the trace and Pa2=PaP_a^2=P_a. This trace is the probability of the selected outcome. Dividing by it normalizes the conditional state to trace one.