Skip to content

Bayes' Rule

Bayes’ rule converts probabilities of data given hypotheses into probabilities of hypotheses given data.

For events HH and DD with nonzero probabilities,

P(H∣D)=P(D∣H)P(H)P(D).\mathbb P(H\mid D) = \frac{\mathbb P(D\mid H)\mathbb P(H)}{\mathbb P(D)}.

The left side is the posterior probability of HH after learning data DD. The factor P(D∣H)\mathbb P(D\mid H) is the likelihood of the data under HH. The factor P(H)\mathbb P(H) is the prior probability before learning DD. The denominator P(D)\mathbb P(D) normalizes the result.

In quantum mechanics, Bayes’ rule is used for classical inference about unknown preparations, parameters, devices, and models. It should not be confused with the quantum state-update rule after a measurement outcome.

The product rule says

P(H∩D)=P(D∣H)P(H)\mathbb P(H\cap D) = \mathbb P(D\mid H)\mathbb P(H)

and also

P(H∩D)=P(H∣D)P(D).\mathbb P(H\cap D) = \mathbb P(H\mid D)\mathbb P(D).

Equating the two expressions gives

P(H∣D)=P(D∣H)P(H)P(D).\mathbb P(H\mid D) = \frac{\mathbb P(D\mid H)\mathbb P(H)}{\mathbb P(D)}.

The algebra is simple. The conceptual step is remembering which conditional probability is being asked for.

For mutually exclusive and exhaustive hypotheses H1,…,HnH_1,\ldots,H_n,

P(Hi∣D)=P(D∣Hi)P(Hi)∑jP(D∣Hj)P(Hj).\mathbb P(H_i\mid D) = \frac{ \mathbb P(D\mid H_i)\mathbb P(H_i) }{ \sum_j \mathbb P(D\mid H_j)\mathbb P(H_j) }.

The denominator is the total probability of the data:

P(D)=∑jP(D∣Hj)P(Hj).\mathbb P(D) = \sum_j \mathbb P(D\mid H_j)\mathbb P(H_j).

This quantity is often called the evidence or marginal likelihood. It is not optional; it is what makes the posterior probabilities add to one.

In a Bayesian calculation, the common terms are:

  • Prior: P(H)\mathbb P(H), the probability assigned before the data DD are included.
  • Likelihood: P(D∣H)\mathbb P(D\mid H), the probability of the data under the hypothesis.
  • Evidence: P(D)\mathbb P(D), the total probability assigned to the data by the model family.
  • Posterior: P(H∣D)\mathbb P(H\mid D), the updated probability after including the data.

The likelihood is a function of the hypothesis once the observed data are fixed. It is not, by itself, a normalized probability distribution over hypotheses.

For two hypotheses H1H_1 and H2H_2, Bayes’ rule can be written as an odds update:

P(H1∣D)P(H2∣D)=P(D∣H1)P(D∣H2)P(H1)P(H2).\frac{\mathbb P(H_1\mid D)}{\mathbb P(H_2\mid D)} = \frac{\mathbb P(D\mid H_1)}{\mathbb P(D\mid H_2)} \frac{\mathbb P(H_1)}{\mathbb P(H_2)}.

The first factor on the right is the likelihood ratio. It says how strongly the data favor H1H_1 over H2H_2. The second factor is the prior odds.

Thus

posterior odds=likelihood ratio×prior odds.\text{posterior odds} = \text{likelihood ratio} \times \text{prior odds}.

This form is often the cleanest way to see how evidence accumulates.

Suppose a rare condition has prior probability

P(H)=0.01.\mathbb P(H)=0.01.

A test has sensitivity

P(+∣H)=0.99\mathbb P(+\mid H)=0.99

and false-positive probability

P(+∣Hc)=0.05.\mathbb P(+\mid H^c)=0.05.

After a positive test, the posterior is

P(H∣+)=0.99⋅0.010.99⋅0.01+0.05⋅0.99.\mathbb P(H\mid +) = \frac{0.99\cdot0.01}{ 0.99\cdot0.01+0.05\cdot0.99 }.

The numerator is 0.00990.0099, while the denominator is 0.05940.0594, so

P(H∣+)=16≈0.167.\mathbb P(H\mid +) = \frac16 \approx 0.167.

The test was accurate, but the condition was rare. Ignoring the base rate would give a wildly overconfident conclusion.

For a continuous parameter θ\theta, Bayes’ rule becomes a density formula:

π(θ∣D)=L(D∣θ)π(θ)Z.\pi(\theta\mid D) = \frac{ \mathcal L(D\mid\theta)\pi(\theta) }{ Z }.

Here:

Z=∫L(D∣θ)π(θ) dθZ = \int \mathcal L(D\mid\theta)\pi(\theta)\,d\theta

is the evidence, π(θ)\pi(\theta) is the prior density, and π(θ∣D)\pi(\theta\mid D) is the posterior density.

The posterior density must integrate to one:

∫π(θ∣D) dθ=1.\int \pi(\theta\mid D)\,d\theta = 1.

As with all probability densities, the value of the density at one exact parameter is not itself a probability.

If data points d1,…,dNd_1,\ldots,d_N are conditionally independent given θ\theta, then the likelihood factors:

L(D∣θ)=∏k=1Np(dk∣θ).\mathcal L(D\mid\theta) = \prod_{k=1}^N p(d_k\mid\theta).

Equivalently, the log-likelihood is a sum:

log⁡L(D∣θ)=∑k=1Nlog⁡p(dk∣θ).\log \mathcal L(D\mid\theta) = \sum_{k=1}^N \log p(d_k\mid\theta).

This is why repeated measurements can rapidly concentrate a posterior distribution, provided the model is appropriate and the data are genuinely informative.

In quantum inference, the Born rule supplies likelihoods for classical data.

If a state model ρ(θ)\rho(\theta) is measured by a POVM with effects EdE_d, then

p(d∣θ)=Tr⁡[ρ(θ)Ed].p(d\mid\theta) = \operatorname{Tr}[\rho(\theta)E_d].

For repeated conditionally independent outcomes D=(d1,…,dN)D=(d_1,\ldots,d_N),

L(D∣θ)=∏k=1NTr⁡[ρ(θ)Edk].\mathcal L(D\mid\theta) = \prod_{k=1}^N \operatorname{Tr}[\rho(\theta)E_{d_k}].

Bayes’ rule then updates a posterior over the classical parameter θ\theta:

π(θ∣D)∝L(D∣θ)π(θ).\pi(\theta\mid D) \propto \mathcal L(D\mid\theta)\pi(\theta).

The symbol ∝\propto means “proportional to”; the evidence normalizes the posterior.

This is common in quantum parameter estimation, calibration, and tomography. The quantum probabilities come from states and measurement operators, but the inference over θ\theta is ordinary Bayesian inference. For the local sensitivity and Cramér–Rao viewpoint, see Fisher Information.

In state tomography, the unknown object may be a density operator ρ\rho itself. A Bayesian treatment assigns a prior density over allowed density operators and updates it using measurement data:

π(ρ∣D)∝L(D∣ρ)π(ρ).\pi(\rho\mid D) \propto \mathcal L(D\mid\rho)\pi(\rho).

For counts ndn_d of outcomes dd from one fixed POVM, a multinomial likelihood has the form

L(D∣ρ)∝∏d[Tr⁡(ρEd)]nd.\mathcal L(D\mid\rho) \propto \prod_d \left[ \operatorname{Tr}(\rho E_d) \right]^{n_d}.

This posterior is a state of knowledge about an unknown preparation or device. It is not the same object as the post-measurement quantum state of one system after a single outcome. The latter is governed by a measurement instrument or state-update rule.

Bayesian Update versus Quantum State Update

Section titled “Bayesian Update versus Quantum State Update”

Bayesian updating changes a probability distribution over hypotheses:

π(θ)⟶π(θ∣D).\pi(\theta) \longrightarrow \pi(\theta\mid D).

Selective quantum measurement update changes the quantum state assigned after an outcome in a specified measurement model:

ρ⟶MdρMd†Tr⁡(MdρMd†).\rho \longrightarrow \frac{M_d\rho M_d^\dagger}{ \operatorname{Tr}(M_d\rho M_d^\dagger) }.

Both updates include conditioning and normalization. They answer different questions. Bayesian updating asks, “Which model or parameter is plausible after the data?” Quantum state update asks, “What state should be assigned to the system after this measurement outcome, given the measurement model?”

The distinction becomes essential when the same observed data are used to infer an unknown source, calibrate a detector, or predict a later measurement on the same system.

  • Confusing P(D∣H)\mathbb P(D\mid H) with P(H∣D)\mathbb P(H\mid D).
  • Ignoring prior probabilities or base rates.
  • Forgetting the evidence denominator.
  • Treating a likelihood as a normalized posterior.
  • Assigning zero prior probability to a hypothesis and then expecting data to revive it.
  • Comparing continuous posterior densities by point height alone without considering volume.
  • Confusing Bayesian inference about an unknown quantum preparation with quantum state update after a measurement.
  • Treating tomography estimates as exact states without reporting uncertainty, model assumptions, and measurement design.
  • E. T. Jaynes, Probability Theory: The Logic of Science, Cambridge University Press, 2003.
  • D. S. Sivia and J. Skilling, Data Analysis: A Bayesian Tutorial, 2nd ed., Oxford University Press, 2006.
  • A. Gelman, J. B. Carlin, H. S. Stern, D. B. Dunson, A. Vehtari, and D. B. Rubin, Bayesian Data Analysis, 3rd ed., CRC Press, 2013.
  • M. A. Nielsen and I. L. Chuang, Quantum Computation and Quantum Information, Cambridge University Press, 2010.
  • C. W. Helstrom, Quantum Detection and Estimation Theory, Academic Press, 1976.
  1. A box is chosen at random. Box 1 is chosen with probability 0.70.7 and contains a red ball with probability 0.20.2. Box 2 is chosen with probability 0.30.3 and contains a red ball with probability 0.80.8. If a red ball is observed, what is the posterior probability that Box 2 was chosen?
Solution

Let H2H_2 be “Box 2” and RR be “red.” Then

P(H2∣R)=P(R∣H2)P(H2)P(R∣H1)P(H1)+P(R∣H2)P(H2).\mathbb P(H_2\mid R) = \frac{ \mathbb P(R\mid H_2)\mathbb P(H_2) }{ \mathbb P(R\mid H_1)\mathbb P(H_1) + \mathbb P(R\mid H_2)\mathbb P(H_2) }.

Substitute:

P(H2∣R)=0.8⋅0.30.2⋅0.7+0.8⋅0.3=0.240.38=1219.\mathbb P(H_2\mid R) = \frac{0.8\cdot0.3}{ 0.2\cdot0.7+0.8\cdot0.3 } = \frac{0.24}{0.38} = \frac{12}{19}.
  1. Two hypotheses have prior probabilities P(H1)=0.6\mathbb P(H_1)=0.6 and P(H2)=0.4\mathbb P(H_2)=0.4. The likelihoods for data DD are P(D∣H1)=0.1\mathbb P(D\mid H_1)=0.1 and P(D∣H2)=0.3\mathbb P(D\mid H_2)=0.3. Compute the posterior odds H2H_2 to H1H_1.
Solution

The prior odds H2H_2 to H1H_1 are

0.40.6=23.\frac{0.4}{0.6} = \frac23.

The likelihood ratio is

0.30.1=3.\frac{0.3}{0.1}=3.

Thus the posterior odds are

3⋅23=2.3\cdot\frac23=2.

So H2H_2 is twice as likely as H1H_1 after observing DD.

  1. A qubit source is modeled by a parameter qq, the probability of outcome 00 in a computational-basis measurement. If NN independent measurements produce n0n_0 zeros and n1n_1 ones, write the likelihood for qq.
Solution

The probability of a zero is qq and the probability of a one is 1−q1-q. Ignoring the combinatorial factor independent of qq, the likelihood is

L(D∣q)∝qn0(1−q)n1.\mathcal L(D\mid q) \propto q^{n_0}(1-q)^{n_1}.
  1. A POVM has effects E1,E2E_1,E_2. A model state is ρ(θ)\rho(\theta). Write the likelihood for observing outcome sequence (1,2,1)(1,2,1) under conditional independence.
Solution

The Born probabilities are

p(i∣θ)=Tr⁡[ρ(θ)Ei].p(i\mid\theta) = \operatorname{Tr}[\rho(\theta)E_i].

Thus

L(D∣θ)=Tr⁡[ρ(θ)E1] Tr⁡[ρ(θ)E2] Tr⁡[ρ(θ)E1].\mathcal L(D\mid\theta) = \operatorname{Tr}[\rho(\theta)E_1]\, \operatorname{Tr}[\rho(\theta)E_2]\, \operatorname{Tr}[\rho(\theta)E_1].
  1. Why is a Bayesian posterior over ρ\rho in tomography not the same thing as the post-measurement state of a single quantum system?
Solution

The posterior over ρ\rho is a probability distribution over possible source or model states after data are observed. It describes uncertainty about an unknown preparation or device. A post-measurement quantum state is the state assigned to a particular system after a specified measurement outcome and measurement instrument. Both involve conditioning, but they refer to different objects.