Skip to content

Shadow Tomography

Shadow tomography estimates selected properties of an unknown quantum state without reconstructing every density-matrix entry. In the experimentally common classical-shadow protocol, each fresh copy of a state ρ\rho is measured in a randomly chosen basis. The basis choice and outcome are converted into an independent classical snapshot ρ^(r)\widehat\rho^{(r)} satisfying

E ⁣[ρ^(r)]=ρ.\mathbb E\!\left[\widehat\rho^{(r)}\right]=\rho.

For an observable OO, the scalar

o^(r)=Tr⁡ ⁣(Oρ^(r))\widehat o^{(r)} = \operatorname{Tr}\!\left( O\widehat\rho^{(r)} \right)

is therefore an unbiased single-shot estimator of ⟨O⟩ρ=Tr⁡(Oρ)\langle O\rangle_\rho=\operatorname{Tr}(O\rho). The same stored snapshots can be queried for many observables. Under a suitable measurement ensemble, the number of state preparations needed for MM simultaneous estimates grows only logarithmically with MM.

The important qualifier is suitable. The cost is controlled by a measurement-dependent variance, usually summarized by a shadow norm. It can be small for local observables and enormous for nonlocal ones, or the reverse, depending on the randomizing circuit. Classical shadows do not make every property of an arbitrary many-qubit state easy.

This page is the canonical home for classical-shadow acquisition, inversion, sample-complexity guarantees, ensemble choice, statistical aggregation, noise calibration, and reporting. State Tomography owns complete state reconstruction and physical estimators. Process Tomography owns channels. POVMs owns generalized measurement theory, while Variance and Covariance and Monte Carlo Basics own the classical sampling concepts used below.

The phrase has two related but distinct technical meanings.

The original problem asks for estimates of MM known two-outcome measurements E1,…,EME_1,\ldots,E_M on a DD-dimensional state:

pj=Tr⁡(Ejρ),0⪯Ej⪯I.p_j=\operatorname{Tr}(E_j\rho), \qquad 0\preceq E_j\preceq I.

The goal is to return p^j\widehat p_j such that, with high probability,

∣p^j−pj∣≤ϵfor every j∈{1,…,M}.\left| \widehat p_j-p_j \right| \leq \epsilon \qquad \text{for every }j\in\{1,\ldots,M\}.

Aaronson showed that copy complexity can be polylogarithmic in both MM and DD. His original bound was

O~ ⁣(log⁡4M log⁡Dϵ4),\widetilde O\!\left( \frac{\log^4 M\,\log D}{\epsilon^4} \right),

where the tilde suppresses additional logarithmic factors. This is an information-theoretic result. Its measurement procedure is adaptive, can use collective quantum operations, and is not the randomized single-copy protocol usually implemented in laboratories.

Classical shadows use a fixed ensemble of randomized measurements on individual copies, followed by explicit classical inversion. The data are reusable because the target observables need not enter the acquisition rule. The price is that the copy complexity depends on how well the chosen ensemble sees each target.

Both formulations explain how many predictions can require far fewer copies than full tomography. They should not be cited as the same algorithm or as having the same resource model. The rest of this page concerns classical shadows unless stated otherwise.

Why Property Estimation Can Beat Reconstruction

Section titled “Why Property Estimation Can Beat Reconstruction”

A density operator on a dd-dimensional Hilbert space has d2−1d^2-1 real parameters. For nn qubits, d=2nd=2^n, so unrestricted reconstruction faces a 4n−14^n-1-parameter model before computational and calibration costs are counted.

Many scientific questions do not require all those parameters. Examples include:

  • the energy of a known Hamiltonian;
  • a list of few-body correlators;
  • fidelity with a known pure target;
  • all reduced states on subsystems of bounded size;
  • a collection of entanglement witnesses;
  • observables used to classify a phase or validate a prepared state.

Each is a projection of ρ\rho onto a much smaller family of operator directions. A measurement protocol can preserve those directions accurately while discarding most of the state. That is the source of the saving.

The saving does not violate parameter counting. A classical shadow is not a uniformly accurate approximation to ρ\rho in trace norm. If a later query asks for a direction that the measurement ensemble sampled only rarely, the estimator has large variance. Demanding accurate answers for a sufficiently rich family eventually recovers the cost of tomography.

Let UU be drawn from a declared ensemble U\mathcal U. After applying UU, a computational-basis measurement returns bit string bb. Define the rotated-back rank-one projector

XU,b=U†∣b⟩⟨b∣U.X_{U,b} = U^\dagger |b\rangle\langle b|U.

For a state ρ\rho, the conditional Born probability is

p(b∣U,ρ)=Tr⁡(XU,bρ).p(b|U,\rho) = \operatorname{Tr}(X_{U,b}\rho).

The ensemble and measurement define a linear channel on operators,

M(A)=EU∼U∑bTr⁡(XU,bA) XU,b.\mathcal M(A) = \mathbb E_{U\sim\mathcal U} \sum_b \operatorname{Tr}(X_{U,b}A)\, X_{U,b}.

This channel says which operator information survives randomization, measurement, and rotation back into the laboratory frame. When M\mathcal M is invertible on the operator subspace of interest, one snapshot is

ρ^(U,b)=M−1(XU,b).\widehat\rho(U,b) = \mathcal M^{-1}(X_{U,b}).

The inverse is generally not a physical quantum channel. Accordingly, ρ^\widehat\rho need not be positive semidefinite. Its purpose is statistical, not ontological.

Averaging first over outcomes and then over random unitaries gives

EU,b ⁣[ρ^(U,b)]=EU∑bp(b∣U,ρ) M−1(XU,b)=M−1 ⁣[M(ρ)]=ρ.\begin{aligned} \mathbb E_{U,b}\!\left[ \widehat\rho(U,b) \right] &= \mathbb E_U \sum_b p(b|U,\rho)\, \mathcal M^{-1}(X_{U,b}) \\ &= \mathcal M^{-1} \!\left[ \mathcal M(\rho) \right] \\ &= \rho. \end{aligned}

Consequently,

EU,bTr⁡ ⁣[Oρ^(U,b)]=Tr⁡(Oρ).\mathbb E_{U,b} \operatorname{Tr} \!\left[ O\widehat\rho(U,b) \right] = \operatorname{Tr}(O\rho).

This derivation exposes every essential assumption: the prepared state is the same ρ\rho on each repetition, the implemented measurement channel equals the one inverted in software, and the inverse exists on the queried directions.

Classical-shadow workflow from randomized measurements through inverse-channel snapshots to many observable queries

Each state copy produces a basis-and-outcome record. Inverting the declared measurement channel creates a reusable snapshot bank, but the uncertainty of each later query still depends on how well the acquisition ensemble covered that observable.

The most accessible qubit protocol chooses an independent basis Pi∈{X,Y,Z}P_i\in\{X,Y,Z\} for every qubit, measures the corresponding eigenvalue si∈{+1,−1}s_i\in\{+1,-1\}, and records (Pi,si)(P_i,s_i).

For one qubit, the observed projector is

ΠP,s=I+sP2.\Pi_{P,s} = \frac{I+sP}{2}.

Uniformly choosing XX, YY, or ZZ shrinks every Bloch-vector component by a factor of three:

M(A)=A+Tr⁡(A)I3.\mathcal M(A) = \frac{A+\operatorname{Tr}(A)I}{3}.

Applying the inverse to an observed projector gives

ρ^=3ΠP,s−I=I+3sP2.\widehat\rho = 3\Pi_{P,s}-I = \frac{I+3sP}{2}.

Its trace is one, but its eigenvalues are 22 and −1-1. A negative eigenvalue is not a reconstruction failure; it is what permits linear unbiased inversion from incomplete one-shot data.

For independent local choices, the measurement channel factorizes. The nn-qubit snapshot is

ρ^=⨂i=1n(3ΠPi,si−I).\widehat\rho = \bigotimes_{i=1}^n \left( 3\Pi_{P_i,s_i}-I \right).

One should almost never construct this 2n×2n2^n\times2^n matrix explicitly. The basis labels and outcome bits already form a compact representation from which local queries can be evaluated.

Let

W=⨂i=1nWi,Wi∈{I,X,Y,Z},W = \bigotimes_{i=1}^n W_i, \qquad W_i\in\{I,X,Y,Z\},

and let S={i:Wi≠I}S=\{i:W_i\neq I\} be its support, with weight w=∣S∣w=|S|. For one local Pauli snapshot,

w^=Tr⁡(Wρ^)=3w(∏i∈Ssi)1 ⁣[Pi=Wi for all i∈S].\widehat w = \operatorname{Tr}(W\widehat\rho) = 3^w \left( \prod_{i\in S}s_i \right) \mathbf 1 \!\left[ P_i=W_i\ \text{for all }i\in S \right].

The indicator is nonzero only when every randomly selected basis matches the target string on its support. This happens with probability 3−w3^{-w}. On matching rounds, the product of outcomes has conditional mean Tr⁡(Wρ)\operatorname{Tr}(W\rho). Therefore

E(w^)=Tr⁡(Wρ).\mathbb E(\widehat w) = \operatorname{Tr}(W\rho).

The second moment is state independent:

E(w^2)=3−w 9w=3w,\mathbb E(\widehat w^2) = 3^{-w}\,9^w = 3^w,

and hence

Var⁡(w^)=3w−Tr⁡(Wρ)2.\operatorname{Var}(\widehat w) = 3^w - \operatorname{Tr}(W\rho)^2.

This exact result captures both the power and the limitation of local shadows. The cost does not depend on the total number of qubits, but it grows exponentially with the number of qubits touched by the observable.

For a Pauli Hamiltonian

H=∑αcαWα,H=\sum_\alpha c_\alpha W_\alpha,

the same snapshot estimates every covered term. The one-shot energy estimator is

h^=∑αcαw^α.\widehat h = \sum_\alpha c_\alpha\widehat w_\alpha.

Covariances between terms matter. Adding the individual variances as if all terms were measured on independent data can understate or overstate the actual energy uncertainty.

At the other extreme, choose a uniformly random nn-qubit Clifford unitary and then measure in the computational basis. For dimension d=2nd=2^n, the associated measurement channel is

M(A)=A+Tr⁡(A)Id+1,\mathcal M(A) = \frac{ A+\operatorname{Tr}(A)I }{d+1},

so the snapshot is

ρ^=(d+1)U†∣b⟩⟨b∣U−I.\widehat\rho = (d+1)U^\dagger |b\rangle\langle b|U-I.

Although the Hilbert space is exponentially large, a Clifford UU and the resulting stabilizer state can be stored and manipulated with tableau methods. Stabilizer Simulation develops that representation.

For a centered observable

O0=O−Tr⁡(O)dI,O_0 = O-\frac{\operatorname{Tr}(O)}{d}I,

global Clifford shadows obey a useful bound of the form

∥O0∥sh2≤3Tr⁡(O02).\|O_0\|_{\mathrm{sh}}^2 \leq 3\operatorname{Tr}(O_0^2).

This favors observables with small Hilbert–Schmidt norm, such as a rank-one projector onto a known pure target. It does not favor a full-weight Pauli operator, for which Tr⁡(W2)=d\operatorname{Tr}(W^2)=d. Global Clifford circuits also require nonlocal scrambling whose depth, noise, and compilation cost may erase their copy-complexity advantage on actual hardware.

There is no universally best shadow.

EnsembleQuantum acquisitionNatural targetsPrincipal cost
Independent local Pauli or single-qubit CliffordBasis rotation and readout on each qubitFew-body Pauli correlators, local Hamiltonians, bounded-size marginalsVariance grows as 3w3^w for a weight-ww Pauli
Global CliffordEntangling Clifford randomization followed by readoutLow-rank global observables and pure-state fidelitiesCircuit depth, noise, and possible exponential cost for local full-rank operators
Biased or derandomized product measurementsTarget-aware local basis scheduleA known list of Pauli terms with uneven importanceReuse becomes target dependent; coverage must remain explicit
Shallow random Clifford circuitsTunable-depth local scramblingIntermediate-scale observablesReconstruction and shadow-norm evaluation can require nontrivial classical contraction
Fermionic Gaussian ensemblesRandom number-preserving or Gaussian mode transformationsFermionic reduced density matricesSpecialized controls and estimators

The decision variable is the total experiment, not copies alone. It includes state-preparation time, randomizing-gate error, readout, calibration, classical preprocessing, query cost, and the desired confidence statement.

The identity component of OO is known exactly because Tr⁡ρ=1\operatorname{Tr}\rho=1, so define the centered operator O0O_0 as above. For a fixed shadow protocol, one useful worst-case definition is

∥O0∥sh2=sup⁡σEρ^∼σ∣Tr⁡ ⁣(O0ρ^)∣2,\|O_0\|_{\mathrm{sh}}^2 = \sup_{\sigma} \mathbb E_{\widehat\rho\sim\sigma} \left| \operatorname{Tr} \!\left( O_0\widehat\rho \right) \right|^2,

where the supremum runs over density operators σ\sigma. The precise norm or seminorm used in a theorem can vary with the estimator and ensemble, but its role is stable: it controls the largest single-shot second moment.

For MM fixed observables O1,…,OMO_1,\ldots,O_M, suppose

B≥max⁡1≤j≤M∥(Oj)0∥sh2.B \geq \max_{1\leq j\leq M} \|(O_j)_0\|_{\mathrm{sh}}^2.

A median-of-means classical-shadow estimator then achieves

∣o^j−Tr⁡(Ojρ)∣≤ϵ\left| \widehat o_j - \operatorname{Tr}(O_j\rho) \right| \leq\epsilon

for every jj with failure probability at most δ\delta, using

N=O ⁣[Bϵ2log⁡ ⁣(2Mδ)]N = O\!\left[ \frac{B}{\epsilon^2} \log\!\left( \frac{2M}{\delta} \right) \right]

independent snapshots. Constants depend on the aggregation theorem and conventions.

The logarithm in MM comes from requiring one dataset to succeed for all queries simultaneously. It does not remove the factor BB, the cost of implementing the ensemble, or the classical work needed to evaluate MM answers.

Partition N=KLN=KL snapshots into KK groups of size LL. For observable OjO_j, form

o‾j(k)=1L∑r=(k−1)L+1kLTr⁡ ⁣(Ojρ^(r)),\overline o_j^{(k)} = \frac{1}{L} \sum_{r=(k-1)L+1}^{kL} \operatorname{Tr} \!\left( O_j\widehat\rho^{(r)} \right),

then return

o^j=median⁡{o‾j(1),…,o‾j(K)}.\widehat o_j = \operatorname{median} \left\{ \overline o_j^{(1)}, \ldots, \overline o_j^{(K)} \right\}.

The group mean reduces variance; the median suppresses the influence of the heavy tails created by inverse-channel weights. Typically LL scales as B/ϵ2B/\epsilon^2 and KK as log⁡(M/δ)\log(M/\delta).

This theorem presumes a finite query family fixed independently of the observed shot noise. The acquired shadow can certainly be reused for later questions, but unlimited data-dependent searching requires a holdout set, an adaptive data-analysis argument, or a new confidence statement.

For

∣Φ+⟩=∣00⟩+∣11⟩2,|\Phi^+\rangle = \frac{|00\rangle+|11\rangle}{\sqrt2},

the target projector has Pauli expansion

∣Φ+⟩⟨Φ+∣=14(I⊗I+X⊗X−Y⊗Y+Z⊗Z).|\Phi^+\rangle\langle\Phi^+| = \frac14 \left( I\otimes I +X\otimes X -Y\otimes Y +Z\otimes Z \right).

Thus the fidelity with an arbitrary two-qubit preparation ρ\rho is

FΦ+=14[1+⟨X⊗X⟩−⟨Y⊗Y⟩+⟨Z⊗Z⟩].F_{\Phi^+} = \frac14 \left[ 1 +\langle X\otimes X\rangle -\langle Y\otimes Y\rangle +\langle Z\otimes Z\rangle \right].

One bank of local Pauli shadows estimates all three correlators. For each snapshot define x^\widehat x, y^\widehat y, and z^\widehat z by the Pauli-string rule, then

F^(r)=14(1+x^(r)−y^(r)+z^(r))\widehat F^{(r)} = \frac14 \left( 1+\widehat x^{(r)} -\widehat y^{(r)} +\widehat z^{(r)} \right)

is unbiased. A single value can lie below zero or above one because it is a Monte Carlo contribution, not a physical fidelity. Clipping each contribution to [0,1][0,1] would generally introduce bias. Physical range constraints belong to the interpretation of the final interval, not to an undocumented nonlinear modification of the raw estimator.

This example also shows why observable structure matters. Only three weight-two strings are needed; reconstructing all fifteen two-qubit Pauli coordinates would answer a broader question at additional cost.

For a subset AA of kk qubits,

ρA=12k∑WA∈Pk⟨WA⟩WA.\rho_A = \frac{1}{2^k} \sum_{W_A\in\mathcal P_k} \langle W_A\rangle W_A.

Local Pauli shadows can estimate all Pauli coordinates for every kk-qubit subset. The number of requested correlators is at most

Mk≤(nk)4k.M_k \leq \binom{n}{k}4^k.

Since each has weight at most kk, the simultaneous coordinate-wise sample scaling is

N=O ⁣[3kϵ2(klog⁡n+k+log⁡1δ)].N = O\!\left[ \frac{3^k}{\epsilon^2} \left( k\log n +k +\log\frac{1}{\delta} \right) \right].

For fixed kk, this grows only logarithmically with the total system size. But coordinate-wise error ϵ\epsilon is not the same as trace-distance error of every ρA\rho_A. Combining 4k4^k uncertain coordinates into a matrix norm requires a separate error propagation and can add substantial kk dependence. Reduced Density Operators owns the physical interpretation of these marginals.

Uniformly sampling XX, YY, and ZZ is robust and reusable, but it wastes shots when the target list is strongly anisotropic. Suppose qubit ii chooses basis PP with probability pi(P)>0p_i(P)>0. The local snapshot becomes

ρ^i=12[I+sipi(Pi)Pi].\widehat\rho_i = \frac12 \left[ I +\frac{s_i}{p_i(P_i)}P_i \right].

For a target Pauli string WW with support SS, the estimator is nonzero on matching rounds and has second moment

E(w^2)=∏i∈S1pi(Wi).\mathbb E(\widehat w^2) = \prod_{i\in S} \frac{1}{p_i(W_i)}.

Increasing the probability of important bases lowers their variance. Setting a needed probability to zero makes that operator direction unidentifiable.

Locally biased shadows optimize probabilities against a known Hamiltonian or observable list. Derandomized methods construct an explicit basis schedule whose coverage objective is at least as good as an associated random schedule. These methods can reduce shot counts dramatically, but the resulting dataset is less universal. The target list, term weights, scheduling algorithm, and coverage counts become part of the reported estimand.

The elementary snapshot identity is linear in ρ\rho. It directly covers Tr⁡(Oρ)\operatorname{Tr}(O\rho), including fidelity with a known pure state because

F(ρ,∣ψ⟩)=⟨ψ∣ρ∣ψ⟩.F(\rho,|\psi\rangle) = \langle\psi|\rho|\psi\rangle.

Nonlinear quantities require independent-snapshot combinations. For example, purity can be estimated without finite-sample self-pairing bias by the U-statistic

P^=1N(N−1)∑r≠sTr⁡ ⁣(ρ^(r)ρ^(s)).\widehat{\mathcal P} = \frac{1}{N(N-1)} \sum_{r\neq s} \operatorname{Tr} \!\left( \widehat\rho^{(r)} \widehat\rho^{(s)} \right).

Independence and unbiasedness give

ETr⁡ ⁣(ρ^(r)ρ^(s))=Tr⁡(ρ2),r≠s.\mathbb E \operatorname{Tr} \!\left( \widehat\rho^{(r)} \widehat\rho^{(s)} \right) = \operatorname{Tr}(\rho^2), \qquad r\neq s.

Higher-degree polynomial functionals use tuples of distinct snapshots. Their variance and computation can be much worse than for linear observables. Entropy estimation is additionally unstable near small eigenvalues and needs its own approximation, continuity, or structural assumptions. One should not apply the linear-observable log⁡M\log M theorem unchanged to purity, entropy, negativity, or mixed-state Uhlmann fidelity.

Suppose the implemented randomizing gates and detector define M~\widetilde{\mathcal M}, while software applies the ideal inverse M−1\mathcal M^{-1}. Then

E(ρ^ideal)=M−1M~(ρ),\mathbb E(\widehat\rho_{\mathrm{ideal}}) = \mathcal M^{-1} \widetilde{\mathcal M}(\rho),

not ρ\rho. The bias in an observable is

bias⁡(O)=Tr⁡ ⁣[O(M−1M~−I)(ρ)].\operatorname{bias}(O) = \operatorname{Tr} \!\left[ O \left( \mathcal M^{-1}\widetilde{\mathcal M} -\mathcal I \right)(\rho) \right].

Randomization can simplify certain noise channels, and calibration experiments can estimate correction factors. Robust-shadow protocols exploit this fact. Their guarantees still depend on assumptions such as stationary, gate-independent or suitably twirled noise, a trusted calibration preparation, and an invertible calibrated channel.

Inversion amplifies weakly observed directions. If a calibrated channel eigenvalue is small, dividing by it can remove bias while greatly increasing variance and sensitivity to calibration error. A correction should therefore report both the corrected central value and the induced sampling overhead. Measurement Tomography owns full detector reconstruction; a shadow calibration is usually a narrower model.

If shot rr prepares ρr\rho_r, an otherwise ideal shadow estimator targets the average state

ρ‾=1N∑r=1Nρr.\overline\rho = \frac1N \sum_{r=1}^N \rho_r.

That average may be scientifically meaningful, but it is not evidence that one stationary state existed. Correlated drift also invalidates confidence formulas that treat all snapshots as independent.

Useful defenses include:

  • randomizing and interleaving measurement settings in time;
  • recording timestamps, calibration identifiers, leakage flags, and shot exclusions;
  • forming blocks that exceed the observed correlation time;
  • comparing early, middle, and late shadow estimates;
  • reserving repeated settings or observables as held-out checks;
  • repeating the experiment under independently prepared calibrations.

Leakage creates a separate issue: binary outcomes may no longer describe a qubit POVM. Postselecting leaked shots changes the estimand to a conditional state and must be declared.

A local Pauli snapshot needs only nn basis labels and nn outcome bits, plus metadata. For a weight-ww Pauli query, evaluation touches only those ww positions. A global Clifford snapshot can be stored as a circuit seed or stabilizer tableau rather than a dense matrix.

Compact acquisition does not guarantee compact analysis:

  • evaluating an explicit list of MM observables generally costs at least enough time to write MM answers;
  • a dense observable may have exponentially many Pauli terms;
  • nonlinear U-statistics naively require O(N2)O(N^2) snapshot pairs;
  • uncertainty analysis over many correlated targets can dominate point estimation;
  • reconstructing every local marginal can produce a large output even when the shot count is modest.

Store the primitive basis-and-outcome records whenever possible. Materializing dense snapshot matrices loses the protocol’s main computational advantage.

  1. Define the query contract. List the target observables or function class, additive or relative precision, confidence level, subsystem, time window, and whether future exploratory queries are anticipated.
  2. Choose an ensemble against the targets. Bound or estimate the relevant shadow norms and include randomizing-circuit depth, noise, calibration, and postprocessing in the comparison.
  3. Verify coverage. Check invertibility on every target direction. For product measurements, audit basis probabilities and realized match counts.
  4. Calibrate the implemented measurement channel. Record the calibration states, assumptions, uncertainty, condition numbers, and validity interval.
  5. Acquire randomized, timestamped shots. Preserve unitary or basis seeds, outcomes, shot ordering, heralds, leakage information, and exclusions.
  6. Construct snapshots transparently. Version the inverse map and test it on simulated states and calibration data with known expectations.
  7. Aggregate with the declared estimator. Keep mean, median-of-means, clipping, shrinkage, and physical projection distinct; they have different bias and uncertainty.
  8. Validate beyond the queried list. Use held-out observables, repeated blocks, direct-measurement comparisons, and residual checks for drift or model failure.
  9. Report the full resource ledger. Include state copies, circuit depth, two-qubit gates, discarded shots, calibration shots, classical runtime, memory, and query count.
  10. Publish reproducible artifacts. Reproducible Notebooks explains the required environment, provenance, and clean-execution record.

A reusable shadow dataset should identify:

  • the physical state-preparation circuit and its parameter version;
  • qubit or mode ordering and all basis conventions;
  • the random-unitary ensemble and sampling distribution;
  • random seeds or the complete per-shot measurement schedule;
  • raw outcomes, timestamps, heralds, and leakage treatment;
  • calibration circuits, fitted channel, assumptions, and uncertainty;
  • the exact inverse or estimator implementation;
  • target observables and whether they were prespecified or exploratory;
  • aggregation method, group partition, confidence procedure, and multiple-query correction;
  • held-out tests, exclusions, software versions, and hardware calibration identifiers.

Without this record, the phrase “classical shadow” does not specify a reproducible statistical experiment.

“The number of measurements is independent of system size”

Section titled ““The number of measurements is independent of system size””

Only under a target family whose shadow norm remains controlled. Local Pauli shadows pay 3w3^w for Pauli weight ww; global Clifford shadows pay according to a different operator geometry.

Individual snapshots can have negative eigenvalues and large operator norm. Projecting each one to the positive cone changes the estimator and usually introduces bias.

Quoting only the logarithm in the number of observables

Section titled “Quoting only the logarithm in the number of observables”

The complete scaling includes the shadow norm, ϵ−2\epsilon^{-2}, confidence, implementation cost, and classical query cost.

A simultaneous guarantee covers a declared finite family. Searching the same data for whichever observable looks most anomalous creates selection bias.

The result estimates M−1M~(ρ)\mathcal M^{-1}\widetilde{\mathcal M}(\rho) unless the implemented channel is calibrated or otherwise justified.

Twirling can simplify noise under assumptions; it does not certify those assumptions or erase drift, leakage, and gate dependence.

Applying a linear theorem to nonlinear quantities

Section titled “Applying a linear theorem to nonlinear quantities”

Purity, entropy, and mixed-state fidelity require distinct estimators and error analyses.

A deep global randomizer and a local basis change are not equivalent experimental resources.

Many observables evaluated on the same snapshots are statistically correlated. That correlation matters for Hamiltonian sums, witnesses, and fitted models.

Calling targeted prediction full tomography

Section titled “Calling targeted prediction full tomography”

A shadow may answer a large query family accurately while remaining poor in trace distance as an approximation to the entire state.

  • State Tomography explains what is gained and paid for when the goal changes from selected properties to a physical full-state estimate.
  • Certification of Entanglement uses shadow estimates for simultaneous witness queries while retaining the separability bound, ensemble calibration, and selection correction.
  • Randomized Benchmarking uses random gate sequences to estimate an ensemble decay; its randomization, estimand, and statistical hierarchy differ from randomized-measurement shadows.
  • Why Benchmarking Is Hard places shadow-derived observables inside a broader validation contract.
  • Density Operators for Quantum Information develops the state representation whose linear functionals are being estimated.
  • Pauli Matrices supplies the operator basis used by local qubit shadows.
  • Fidelity distinguishes the pure-target linear overlap from nonlinear mixed-state fidelity.
  • Metrics for Quantum Hardware discusses what a measured observable or fidelity does and does not establish about a processor.

Let M\mathcal M be invertible only on an operator subspace S\mathcal S. Show that the snapshot estimator is unbiased for every observable O∈SO\in\mathcal S, provided the inverse is interpreted on S\mathcal S. Explain why no estimator built from the same data can identify a direction in ker⁡M\ker\mathcal M without an additional assumption.

Solution

Let MS−1\mathcal M_{\mathcal S}^{-1} denote the inverse restricted to S\mathcal S. For the component ρS\rho_{\mathcal S} visible to the measurement,

E(ρ^S)=MS−1M(ρ)=ρS.\mathbb E(\widehat\rho_{\mathcal S}) = \mathcal M_{\mathcal S}^{-1} \mathcal M(\rho) = \rho_{\mathcal S}.

If O∈SO\in\mathcal S, its expectation depends only on this visible component, so

ETr⁡ ⁣(Oρ^S)=Tr⁡(Oρ).\mathbb E \operatorname{Tr} \!\left( O\widehat\rho_{\mathcal S} \right) = \operatorname{Tr}(O\rho).

Now let A∈ker⁡MA\in\ker\mathcal M. States ρ\rho and ρ+tA\rho+tA, whenever both are physical, produce the same measurement statistics because

M(ρ+tA)=M(ρ).\mathcal M(\rho+tA) = \mathcal M(\rho).

No data-only estimator can distinguish them. Identifying AA requires another measurement, a structural prior, or a model constraint that rules out one of the states.

Write

ρ=12(I+rxX+ryY+rzZ).\rho = \frac12 \left( I+r_xX+r_yY+r_zZ \right).

Average the post-measurement rotated-back projector over uniformly random XX, YY, and ZZ bases. Derive M(ρ)\mathcal M(\rho) and verify that 3ΠP,s−I3\Pi_{P,s}-I is the inverse-channel snapshot.

Solution

For a fixed basis PP, averaging the observed projector over outcomes dephases ρ\rho in that basis:

∑sTr⁡(ΠP,sρ) ΠP,s=12(I+rPP).\sum_s \operatorname{Tr}(\Pi_{P,s}\rho)\, \Pi_{P,s} = \frac12 \left( I+r_P P \right).

Averaging over the three bases gives

M(ρ)=13∑P=X,Y,Z12(I+rPP)=12[I+13(rxX+ryY+rzZ)]=ρ+I3.\begin{aligned} \mathcal M(\rho) &= \frac13 \sum_{P=X,Y,Z} \frac12 \left( I+r_P P \right) \\ &= \frac12 \left[ I+\frac13 \left( r_xX+r_yY+r_zZ \right) \right] \\ &= \frac{\rho+I}{3}. \end{aligned}

More generally,

M(A)=A+Tr⁡(A)I3.\mathcal M(A) = \frac{A+\operatorname{Tr}(A)I}{3}.

Because Tr⁡ΠP,s=1\operatorname{Tr}\Pi_{P,s}=1,

M−1(ΠP,s)=3ΠP,s−I.\mathcal M^{-1}(\Pi_{P,s}) = 3\Pi_{P,s}-I.

Its ensemble average is ρ\rho, as required.

For a weight-ww Pauli string WW, derive the local-shadow estimator and show

Var⁡(w^)=3w−⟨W⟩2.\operatorname{Var}(\widehat w) = 3^w-\langle W\rangle^2.

What changes if ww increases while the total number of qubits remains fixed?

Solution

Every target basis must match, which occurs with probability 3−w3^{-w}. The estimator is zero otherwise. On a match it equals

3w∏i∈Ssi.3^w\prod_{i\in S}s_i.

The conditional mean of the outcome product is ⟨W⟩\langle W\rangle, so the unconditional mean is ⟨W⟩\langle W\rangle. Squaring removes all signs:

E(w^2)=3−w(3w)2=3w.\mathbb E(\widehat w^2) = 3^{-w}(3^w)^2 = 3^w.

Subtracting the squared mean gives the stated variance. Increasing ww by one multiplies the leading second moment by three, regardless of how many idle qubits lie outside the support. The difficulty is observable weight, not total register size by itself.

For an ideal ∣Φ+⟩|\Phi^+\rangle preparation, find the expectations of X⊗XX\otimes X, Y⊗YY\otimes Y, and Z⊗ZZ\otimes Z. Insert them into the shadow fidelity estimator. Explain why an individual Monte Carlo contribution need not lie in [0,1][0,1].

Solution

The Bell state is stabilized by X⊗XX\otimes X and Z⊗ZZ\otimes Z, while Y⊗Y=−(X⊗X)(Z⊗Z)Y\otimes Y=-(X\otimes X)(Z\otimes Z). Hence

⟨X⊗X⟩=1,⟨Y⊗Y⟩=−1,⟨Z⊗Z⟩=1.\langle X\otimes X\rangle=1, \qquad \langle Y\otimes Y\rangle=-1, \qquad \langle Z\otimes Z\rangle=1.

Therefore

FΦ+=14[1+1−(−1)+1]=1.F_{\Phi^+} = \frac14 \left[ 1+1-(-1)+1 \right] = 1.

A single weight-two Pauli estimator is either zero or ±9\pm9. The corresponding single-shot linear combination can therefore lie far outside [0,1][0,1]. Only its expectation equals a physical fidelity. Range projection is nonlinear and changes the estimator’s bias.

5. Scaling for all bounded-size correlators

Section titled “5. Scaling for all bounded-size correlators”

Estimate the number MkM_k of Pauli correlators needed to describe every kk-qubit marginal of an nn-qubit state. Use the local-Pauli shadow bound to obtain the dependence of the required snapshots on nn, kk, ϵ\epsilon, and δ\delta for simultaneous coordinate-wise accuracy.

Solution

There are (nk)\binom nk subsets of size kk and at most 4k4^k Pauli strings per subset, so

Mk≤(nk)4k.M_k\leq\binom nk4^k.

Every requested string has weight at most kk, so B≤3kB\leq3^k. Substituting into the simultaneous-prediction bound gives

N=O ⁣[3kϵ2log⁡ ⁣(2(nk)4kδ)].N = O\!\left[ \frac{3^k}{\epsilon^2} \log\!\left( \frac{2\binom nk4^k}{\delta} \right) \right].

Using log⁡(nk)≤klog⁡(en/k)\log\binom nk\leq k\log(en/k),

N=O ⁣[3kϵ2(klog⁡enk+k+log⁡1δ)].N = O\!\left[ \frac{3^k}{\epsilon^2} \left( k\log\frac{en}{k} +k +\log\frac1\delta \right) \right].

This controls each Pauli coordinate. A matrix-norm guarantee for every marginal requires propagating the coordinate errors and is stronger.

For one target string WW and independent basis probabilities pi(P)p_i(P), show that the second moment is

∏i∈Spi(Wi)−1.\prod_{i\in S}p_i(W_i)^{-1}.

Why is choosing pi(Wi)=1p_i(W_i)=1 not a universal improvement?

Solution

The target bases match with probability

qW=∏i∈Spi(Wi).q_W = \prod_{i\in S}p_i(W_i).

On a match, the inverse-probability estimator has magnitude qW−1q_W^{-1}. Therefore

E(w^2)=qWqW−2=∏i∈Spi(Wi)−1.\mathbb E(\widehat w^2) = q_W q_W^{-2} = \prod_{i\in S}p_i(W_i)^{-1}.

Setting each needed probability to one minimizes variance for this one string. It assigns zero probability to incompatible bases, however, making other Pauli directions invisible. It is optimal only for the narrow query contract in which those other directions are irrelevant.

In a local Pauli experiment, suppose the recorded eigenvalue is flipped with known probability q<1/2q<1/2, independently of the state and chosen basis. Show the bias of the ideal estimator for a target containing one measured Pauli factor. Give an unbiased correction and its variance-amplification factor.

Solution

If s′s' is the recorded value, then

E(s′∣s)=(1−q)s+q(−s)=(1−2q)s.\mathbb E(s'|s) = (1-q)s+q(-s) = (1-2q)s.

Every affected Pauli factor therefore multiplies the target expectation by 1−2q1-2q. For one affected factor,

E(w^ideal)=(1−2q)⟨W⟩.\mathbb E(\widehat w_{\mathrm{ideal}}) = (1-2q)\langle W\rangle.

An unbiased correction is

w^corr=w^ideal1−2q.\widehat w_{\mathrm{corr}} = \frac{\widehat w_{\mathrm{ideal}}}{1-2q}.

Its second moment, and hence its leading variance scale, is amplified by

1(1−2q)2.\frac{1}{(1-2q)^2}.

Calibration uncertainty in qq adds another contribution not included in this conditional calculation.

Show that the U-statistic

P^=1N(N−1)∑r≠sTr⁡ ⁣(ρ^(r)ρ^(s))\widehat{\mathcal P} = \frac{1}{N(N-1)} \sum_{r\neq s} \operatorname{Tr} \!\left( \widehat\rho^{(r)} \widehat\rho^{(s)} \right)

is unbiased for Tr⁡(ρ2)\operatorname{Tr}(\rho^2). Why does including terms with r=sr=s generally introduce a finite-sample bias?

Solution

For r≠sr\neq s, the snapshots are independent. Thus

ETr⁡ ⁣(ρ^(r)ρ^(s))=Tr⁡ ⁣[E(ρ^(r))E(ρ^(s))]=Tr⁡(ρ2).\begin{aligned} \mathbb E \operatorname{Tr} \!\left( \widehat\rho^{(r)} \widehat\rho^{(s)} \right) &= \operatorname{Tr} \!\left[ \mathbb E(\widehat\rho^{(r)}) \mathbb E(\widehat\rho^{(s)}) \right] \\ &= \operatorname{Tr}(\rho^2). \end{aligned}

Every ordered pair has the same expectation, so their average is unbiased. For a self-pair,

ETr⁡ ⁣[(ρ^(r))2]\mathbb E \operatorname{Tr} \!\left[ (\widehat\rho^{(r)})^2 \right]

depends on the snapshot second moment and is generally not Tr⁡(ρ2)\operatorname{Tr}(\rho^2). Including the NN self-pairs therefore adds a bias of order 1/N1/N unless it is explicitly corrected.