Skip to content

Local Hidden-Variable Models

A local hidden-variable model is a conditional-probability model for a Bell experiment. It explains correlations between distant outcomes by a shared variable λ\lambda in their common past, while each local response depends only on the setting at that wing and on λ\lambda. With measurement settings x,yx,y and outcomes a,ba,b, its defining form is

p(a,b∣x,y)=∫Λdμ(λ) pA(a∣x,λ)pB(b∣y,λ).p(a,b\mid x,y) = \int_\Lambda d\mu(\lambda)\, p_A(a\mid x,\lambda) p_B(b\mid y,\lambda).

This formula is stronger than operational no-signaling and more precise than the phrase local realism. It includes stochastic as well as deterministic responses. Bell inequalities test this factorized class together with the assumption that the settings are statistically independent of λ\lambda; they do not test every theory containing unobserved variables.

Background assumed. Familiarity with conditional probability, marginalization, and the law of total probability is assumed. No internal page is required.

Consider many trials with the following recorded variables:

  • x∈Xx\in X: Alice’s measurement setting;
  • y∈Yy\in Y: Bob’s measurement setting;
  • a∈Aa\in A: Alice’s outcome;
  • b∈Bb\in B: Bob’s outcome.

The observed object is the family of conditional distributions p(a,b∣x,y)p(a,b\mid x,y). A hidden-variable model introduces a measurable space Λ\Lambda and a variable λ∈Λ\lambda\in\Lambda. Depending on the proposed theory, λ\lambda may encode a source state, common randomness, additional physical properties, or relevant variables in the overlap of the two measurements’ past light cones. Calling it hidden only means that the recorded data do not condition on its complete value.

For each fixed xx and λ\lambda, Alice has a normalized response function,

pA(a∣x,λ)≥0,∑apA(a∣x,λ)=1,p_A(a\mid x,\lambda)\geq0, \qquad \sum_a p_A(a\mid x,\lambda)=1,

and Bob has the analogous pB(b∣y,λ)p_B(b\mid y,\lambda). The measure μ\mu is normalized. Discrete models replace the integral by a sum.

The factorized expression then says two things at once. First, the common variable may correlate the outcomes after it is averaged over. Second, once λ\lambda is specified, neither local response requires the remote setting or the remote outcome.

The compact integral hides assumptions that should be kept separate.

All settings, outcomes, and hidden variables belong to one probabilistic model with nonnegative normalized probabilities. This does not require that all four possible measurements are actually performed on one trial. It does require one model that supplies the response distribution for whichever setting is selected.

The distribution of the shared variable is taken to be independent of the chosen settings:

dμ(λ∣x,y)=dμ(λ).d\mu(\lambda\mid x,y)=d\mu(\lambda).

When the relevant conditional probabilities have full support, this is equivalent to p(x,y∣λ)=p(x,y)p(x,y\mid\lambda)=p(x,y). It is also called setting independence, freedom of choice, or no conspiracy. The mathematical condition, not a metaphysical theory of free will, is what an inequality uses. Relaxing it changes the model class and therefore changes the bound.

Conditioned on a complete specification λ\lambda of the proposed common causes, the outcomes factorize:

p(a,b∣x,y,λ)=pA(a∣x,λ)pB(b∣y,λ).p(a,b\mid x,y,\lambda) = p_A(a\mid x,\lambda) p_B(b\mid y,\lambda).

This is often called factorizability or Bell locality. Bell’s later term local causality motivates the screening-off structure: events in the remote wing should add no predictive information once the local setting and an adequate specification of the relevant past are given. Whether a particular theory’s λ\lambda is adequate is a physical question, not a purely notational one.

The theoretical distribution refers to the same ensemble whose frequencies are analyzed. If detection, coincidence matching, or postselection depends on settings and hidden variables, conditioning on the retained trials can alter the effective distribution. This is an empirical sampling issue, not an additional step in the algebraic theorem.

These assumptions should not be compressed into an undefined word such as realism. Predetermined outcomes are convenient representatives of the local set, but they need not be postulated separately. Nor does the model assume that observed outcomes are uncorrelated.

Factorization split into two independences

Section titled “Factorization split into two independences”

Two conditional-independence conditions expose different parts of local factorization.

Parameter independence says that, at fixed λ\lambda, the probability of a local outcome does not depend on the remote setting:

p(a∣x,y,λ)=p(a∣x,λ),p(b∣x,y,λ)=p(b∣y,λ).\begin{aligned} p(a\mid x,y,\lambda)&=p(a\mid x,\lambda),\\ p(b\mid x,y,\lambda)&=p(b\mid y,\lambda). \end{aligned}

Outcome independence says that, at fixed settings and λ\lambda, learning the remote outcome does not change the local outcome probability:

p(a∣b,x,y,λ)=p(a∣x,y,λ),p(b∣a,x,y,λ)=p(b∣x,y,λ).\begin{aligned} p(a\mid b,x,y,\lambda)&=p(a\mid x,y,\lambda),\\ p(b\mid a,x,y,\lambda)&=p(b\mid x,y,\lambda). \end{aligned}

Factorization implies both conditions. Conversely, wherever the conditionals are defined, the chain rule together with parameter and outcome independence gives

p(a,b∣x,y,λ)=p(a∣b,x,y,λ)p(b∣x,y,λ)=p(a∣x,λ)p(b∣y,λ).\begin{aligned} p(a,b\mid x,y,\lambda) &=p(a\mid b,x,y,\lambda)p(b\mid x,y,\lambda)\\ &=p(a\mid x,\lambda)p(b\mid y,\lambda). \end{aligned}

Thus their conjunction is equivalent to factorization, up to null events. The split is useful for diagnosing a model, but the product form is the clean assumption used in a Bell-inequality derivation.

Assume factorization and measurement independence. Marginalizing Bob’s outcome gives

p(a∣x,y)=∑b∫Λdμ(λ) pA(a∣x,λ)pB(b∣y,λ)=∫Λdμ(λ) pA(a∣x,λ),\begin{aligned} p(a\mid x,y) &= \sum_b\int_\Lambda d\mu(\lambda)\, p_A(a\mid x,\lambda)p_B(b\mid y,\lambda)\\ &= \int_\Lambda d\mu(\lambda)\,p_A(a\mid x,\lambda), \end{aligned}

which is independent of yy. Similarly, p(b∣x,y)p(b\mid x,y) is independent of xx. Every Bell-local model in this sense is therefore operationally no-signaling.

The converse is false. No-signaling constrains only the observed marginals; it does not require a common-cause decomposition of the joint distribution. Quantum correlations, and even stronger hypothetical correlations, can have setting-independent marginals while failing local factorization.

A shared random bit is the simplest counterexample to the claim that every correlation is nonlocal. Let λ∈{−1,+1}\lambda\in\{-1,+1\} with equal probabilities, and let both parties report it regardless of their settings:

pA(a∣x,λ)=δa,λ,pB(b∣y,λ)=δb,λ.p_A(a\mid x,\lambda)=\delta_{a,\lambda}, \qquad p_B(b\mid y,\lambda)=\delta_{b,\lambda}.

The observed distribution is

p(a,b∣x,y)={12,a=b,0,a≠b.p(a,b\mid x,y) = \begin{cases} \tfrac12,&a=b,\\ 0,&a\ne b. \end{cases}

Both marginals are uniform, yet the correlation ⟨ab⟩=1\langle ab\rangle=1 is perfect. The common variable explains it without either response depending on the remote setting. What Bell inequalities constrain is not correlation by itself but whether one local model can reproduce the correlations for several alternative settings.

Stochastic models and deterministic refinements

Section titled “Stochastic models and deterministic refinements”

For finitely many outcomes, every stochastic local response can be made deterministic by enlarging the hidden variable. Introduce independent random numbers rA,rBr_A,r_B, uniform on [0,1][0,1], and define

λ~=(λ,rA,rB).\widetilde\lambda=(\lambda,r_A,r_B).

Order Alice’s outcomes and partition [0,1][0,1] into intervals whose lengths are pA(a∣x,λ)p_A(a\mid x,\lambda). Let Ax(λ~)A_x(\widetilde\lambda) be the label of the interval containing rAr_A. Then

Pr⁡[Ax(λ~)=a∣λ]=pA(a∣x,λ).\Pr[A_x(\widetilde\lambda)=a\mid\lambda] = p_A(a\mid x,\lambda).

Construct ByB_y in the same way from rBr_B. Averaging the deterministic responses over the enlarged variable reproduces the original stochastic model exactly. The added seeds remain local and independent of the settings.

This refinement explains why proofs often begin with predetermined response tables. For finite Bell scenarios, deterministic strategies are the extreme points of the local correlation set; stochastic local models are their convex mixtures. Determinism is a mathematical refinement of the factorized model, not an extra conclusion about nature.

In the binary two-setting scenario, a deterministic strategy assigns one value to each of A0,A1,B0,B1A_0,A_1,B_0,B_1. There are

22×22=162^2\times2^2=16

such response tables. Their convex hull is a polytope. Linear Bell inequalities are supporting half-spaces of this local polytope; the CHSH inequality is its standard correlation witness.

Proposed explanationBell-local model?Reason
shared random bit emitted by the sourceyesaveraging a common cause preserves factorization
stochastic local detector responseyesresponse probabilities need not be 0 or 1
Alice’s response explicitly depends on Bob’s settingnoparameter independence fails
outcomes remain dependent after complete λ\lambda is fixednooutcome independence fails
λ\lambda is correlated with the later setting pairnot under measurement independencethe standard setting-independence premise fails
nonlocal hidden-variable dynamicsnohidden variables alone do not imply locality
an arbitrary no-signaling distributionnot necessarilyno-signaling is weaker than factorization

The table makes the logical target precise. A violation of a Bell inequality rules out the intersection of ordinary probability, measurement independence, local factorization, and the data-selection assumptions needed to identify the tested distribution. It does not rule out all completions of quantum mechanics, all stochastic theories, or all hidden-variable theories.

Calling factorization mere absence of signaling. Bell factorization is a conditional common-cause structure. No-signaling constrains only observable marginals and is strictly weaker.

Treating determinism as an independent Bell premise. Finite stochastic local models can be determinized by adding local random seeds. The decisive premises are the common distribution, setting independence, and local factorization.

Assuming correlated outcomes violate outcome independence. Unconditional outcomes may be strongly correlated because both depend on λ\lambda. Outcome independence asks whether dependence remains after the proposed complete common cause is fixed.

Saying Bell experiments refute hidden variables. They constrain local hidden-variable models under their stated assumptions. A hidden-variable theory can lie outside that class by being nonlocal or by relaxing another premise.

Equating measurement independence with human free will. The inequality uses a statistical condition on settings and λ\lambda. How an experiment justifies that condition is a separate causal and empirical question.

1. Normalize and marginalize a local model

Section titled “1. Normalize and marginalize a local model”

Starting from the factorized model, prove that ∑a,bp(a,b∣x,y)=1\sum_{a,b}p(a,b\mid x,y)=1 and that Alice’s marginal is independent of yy.

Solution

Normalization follows from the normalized response functions:

∑a,bp(a,b∣x,y)=∫dμ(λ)[∑apA(a∣x,λ)][∑bpB(b∣y,λ)]=∫dμ(λ)=1.\begin{aligned} \sum_{a,b}p(a,b\mid x,y) &= \int d\mu(\lambda) \left[\sum_a p_A(a\mid x,\lambda)\right] \left[\sum_b p_B(b\mid y,\lambda)\right]\\ &=\int d\mu(\lambda)=1. \end{aligned}

For the marginal, summing only over bb removes Bob’s normalized response:

p(a∣x,y)=∫dμ(λ) pA(a∣x,λ),p(a\mid x,y) = \int d\mu(\lambda)\,p_A(a\mid x,\lambda),

which contains no yy. Measurement independence is used when the same dμ(λ)d\mu(\lambda) is employed for every setting pair.

For the shared-bit model above, compute p(a∣x)p(a\mid x), p(b∣y)p(b\mid y), and ⟨ab⟩\langle ab\rangle. Explain why the outcomes are correlated even though the conditional distribution factorizes at fixed λ\lambda.

Solution

Each sign of λ\lambda occurs with probability 1/21/2, so

p(a=+1∣x)=p(a=−1∣x)=12p(a=+1\mid x)=p(a=-1\mid x)=\tfrac12

and similarly for Bob. Because a=b=λa=b=\lambda on every trial, ab=1ab=1 and ⟨ab⟩=1\langle ab\rangle=1. At fixed λ\lambda the outcomes are deterministic, so

p(a,b∣x,y,λ)=δa,λδb,λ.p(a,b\mid x,y,\lambda) = \delta_{a,\lambda}\delta_{b,\lambda}.

The observed dependence appears only after the common cause is averaged over.

Use the chain rule to show that parameter independence plus outcome independence implies factorization wherever the conditional probabilities are defined.

Solution

The chain rule gives

p(a,b∣x,y,λ)=p(a∣b,x,y,λ)p(b∣x,y,λ).p(a,b\mid x,y,\lambda) = p(a\mid b,x,y,\lambda)p(b\mid x,y,\lambda).

Outcome independence removes bb from Alice’s first factor. Parameter independence then removes the remote setting from each factor, yielding

p(a,b∣x,y,λ)=p(a∣x,λ)p(b∣y,λ).p(a,b\mid x,y,\lambda) = p(a\mid x,\lambda)p(b\mid y,\lambda).

The reverse implication follows by summing the product distribution and then conditioning, except on zero-probability events.

Alice has outcomes {−1,+1}\{-1,+1\} with pA(+1∣x,λ)=qx(λ)p_A(+1\mid x,\lambda)=q_x(\lambda). Construct a deterministic response using one uniform random number rA∈[0,1]r_A\in[0,1] and recover the stated probability.

Solution

Define

Ax(λ,rA)={+1,0≤rA<qx(λ),−1,qx(λ)≤rA≤1.A_x(\lambda,r_A) = \begin{cases} +1,&0\leq r_A<q_x(\lambda),\\ -1,&q_x(\lambda)\leq r_A\leq1. \end{cases}

The interval producing +1+1 has Lebesgue measure qx(λ)q_x(\lambda), so averaging over the uniform seed gives the required response probability. The seed is local because Bob’s setting and outcome do not enter the definition.

Model I sets a=f(x,y,λ)a=f(x,y,\lambda) and b=g(y,λ)b=g(y,\lambda) while keeping μ(λ)\mu(\lambda) setting independent. Model II uses local responses but chooses λ=(x,y)\lambda=(x,y) on every trial. Which standard assumption fails in each case?

Solution

In Model I, Alice’s response depends on Bob’s setting, so parameter independence and hence local factorization fail. Measurement independence may still hold.

In Model II, the response functions can be locally factorized after λ\lambda is fixed, but p(λ∣x,y)p(\lambda\mid x,y) is concentrated on λ=(x,y)\lambda=(x,y). It therefore differs for different setting pairs, so measurement independence fails. Neither model belongs to the standard local hidden-variable class.

  • J. S. Bell, “On the Einstein Podolsky Rosen Paradox,” Physics Physique Fizika 1, 195–200, 1964, doi:10.1103/PhysicsPhysiqueFizika.1.195.
  • J. S. Bell, “The Theory of Local Beables,” Epistemological Letters 9, 11–24, 1976; reprinted in Dialectica 39, 85–96, 1985.
  • N. Brunner, D. Cavalcanti, S. Pironio, V. Scarani, and S. Wehner, “Bell Nonlocality,” Reviews of Modern Physics 86, 419–478, 2014, doi:10.1103/RevModPhys.86.419.
  • J. F. Clauser and M. A. Horne, “Experimental Consequences of Objective Local Theories,” Physical Review D 10, 526–535, 1974, doi:10.1103/PhysRevD.10.526.
  • A. Fine, “Hidden Variables, Joint Probability, and the Bell Inequalities,” Physical Review Letters 48, 291–295, 1982, doi:10.1103/PhysRevLett.48.291.
  • J. P. Jarrett, “On the Physical Significance of the Locality Conditions in the Bell Arguments,” Noûs 18, 569–589, 1984, doi:10.2307/2214873.