Variational Quantum Algorithms
Short Definition
Section titled “Short Definition”A variational quantum algorithm (VQA) optimizes a classical parameter vector using quantities estimated from a parameterized quantum circuit. In a common form,
and the task is encoded in an objective such as
The quantum processor prepares states and samples measurements; a classical routine constructs cost or gradient estimates and proposes the next parameters. The loop is repeated until a declared stopping rule is met.
That description identifies an architecture, not a complexity advantage. Whether a VQA is useful depends on the representational quality of the circuit family, the statistical and hardware cost of estimating the objective, the geometry of the optimization landscape, the classical competition, and the strength of the validation evidence. A small circuit or a decreasing training curve alone establishes none of these.
This page is the canonical home for the general hybrid loop, parameterized circuit design, finite-sample objective estimation, quantum gradients, optimization geometry, barren plateaus, and VQA-wide validation practice. Application-specific pages own their additional structure: VQE owns the Rayleigh–Ritz energy problem, QAOA owns alternating cost-and-mixer ansatzes, and Quantum Machine Learning owns data encoding, generalization, and test-set semantics. Hybrid Quantum Simulation owns variational real- and imaginary-time projection, its trajectory error budget, and its relation to digital–analog control and quantum embedding.
The Computational Contract
Section titled “The Computational Contract”Before choosing an optimizer or running a circuit, specify the problem as an input-output contract.
Problem definition
Section titled “Problem definition”A generic instance contains:
- classical data describing a Hamiltonian, graph, target state, data set, differential equation, or other task;
- an initial state or state-preparation procedure;
- a parameter domain ;
- a family of implementable circuits ;
- an objective, constraints, and a rule for estimating them;
- a classical update procedure and a resource budget.
The output may be a parameter vector, a prepared quantum state, an estimated observable, or a classical decision derived from measurements. These are not interchangeable. If the useful output is a quantum state, success must be tested through operational observables or a task-specific certificate. If the useful output is a classical number, its uncertainty and total acquisition cost belong in the output definition.
Success criterion
Section titled “Success criterion”An optimization trace ending at is not itself a success criterion. A defensible criterion has the form
for a loss-like quality measure , or an analogous lower-tail statement for a score. The probability must state what is random: shots, circuit instances, device drift, initialization, optimizer randomness, problem instances, or all of them.
Training and validation should be distinct whenever selection has occurred. The same noisy evaluations used to choose the best iteration cannot provide an unbiased assessment of that iteration. At minimum, re-evaluate the selected parameters with fresh shots. For a claimed family-level result, also use held- out problem instances or a prespecified benchmark distribution.
Input model and cost model
Section titled “Input model and cost model”The circuit may receive its instance through gate parameters, an oracle, a prepared quantum state, or an explicitly compiled operator. Data loading is part of the algorithm. So are:
- circuit compilation and routing;
- state preparation and reset;
- the number of distinct circuits;
- shots, rejected trials, and measurement settings;
- classical preprocessing, optimization, and postprocessing;
- calibration and error-mitigation overhead;
- repeated seeds and failed runs.
A query count can be mathematically useful while omitting nearly all wall-clock cost. Report both the abstract model and the implemented boundary. The Resource Estimation Tools page develops the broader accounting discipline.
The Hybrid Loop
Section titled “The Hybrid Loop”The basic loop is adaptive:
declare objective, circuit family, constraints, and budgetinitialize parameters and optimizer staterepeat compile or bind the requested circuits execute a prespecified measurement allocation estimate objective, constraints, gradients, and uncertainty update optimizer state and propose new parametersuntil a stopping or budget rule firesvalidate the selected result with fresh dataThe hybrid loop is an adaptive experiment. Shot allocations, calibration context, seeds, mitigation, and stopping decisions affect the final estimator. Independent validation and a versioned result record are part of the algorithmic workflow, not optional presentation layers.
The distinction between symbolic parameters and executables matters. Changing may only bind rotation angles in a fixed compiled template, or it may trigger new synthesis, routing, pulse generation, or calibration choices. The latter changes both cost and noise. A reproducible experiment records which operations were fixed across the loop and which were regenerated.
The optimizer is also part of the adaptive measurement policy. It determines which parameter points are queried and often how many shots are spent there. Consequently, one cannot generally treat all observations collected during training as independent, identically distributed samples from one fixed experiment.
Objective, Estimand, and Estimator
Section titled “Objective, Estimand, and Estimator”Three objects are easy to conflate:
- the ideal objective defined by the mathematical problem;
- the implemented estimand defined by the compiled circuit and physical experiment;
- the finite-data estimator computed from observed records.
For an ideal pure-state objective,
Compilation approximation and noise produce a physical state , so the hardware estimand is
where denotes time-dependent context. A sample mean estimates , not automatically . The useful decomposition is
The first term is statistical error. The second includes coherent miscalibration, decoherence, readout error, drift, compilation approximation, and any mismatch between the intended and realized observable. More shots reduce the first term but need not reduce the second.
Objectives built from many observables
Section titled “Objectives built from many observables”Suppose
With separately estimated terms,
If term receives independent shots and its single-shot variance is , then, ignoring covariance between settings,
For fixed total shots , the continuous optimum is
Uniform allocation is therefore not generally optimal. In practice, is unknown and can be learned from pilot data. Commuting terms can sometimes share measurement settings, but grouping introduces covariance and changes circuit count, basis-change depth, and statistical analysis. The measurement design must be optimized as a whole rather than term by term.
Nonlinear postprocessing
Section titled “Nonlinear postprocessing”Ratios, logarithms, normalization by estimated quantities, postselected conditional means, and mitigation inverses can be biased at finite sample size. For a nonlinear map ,
in general. Bootstrap intervals can help with irregular estimators, but only if resampling respects the experimental hierarchy. Analytic propagation, likelihood methods, or simulation-based calibration may be preferable.
Parameterized Circuit Families
Section titled “Parameterized Circuit Families”An ansatz is the map
Its reachable set is a manifold or stratified subset of state space, often of far lower dimension than the full Hilbert space. The key question is not whether the family can approximate arbitrary states, but whether it can reach useful states for the target instance at acceptable depth and can be searched with available precision.
Problem-inspired circuits
Section titled “Problem-inspired circuits”Problem-inspired ansatzes use generators, symmetries, conserved quantities, or known limiting states from the application. A layered form is
If the generators commute with a symmetry and the initial state lies in one symmetry sector, then the circuit remains there:
This can exclude invalid states and reduce the search space. It can also exclude the solution if the sector, encoding, or assumed symmetry is wrong. Constraints should therefore be justified from the problem, not imposed only because they simplify optimization.
Hardware-efficient circuits
Section titled “Hardware-efficient circuits”Hardware-efficient ansatzes alternate native single-qubit rotations and entangling layers chosen to match connectivity. They can minimize compiled depth for a given number of layers, but the phrase does not mean efficient optimization or favorable asymptotic scaling. Deep, highly random families can be expressive while producing concentrated costs and gradients. Native gates can also drift, crosstalk, and vary strongly across qubits.
Adaptive and growing circuits
Section titled “Adaptive and growing circuits”An adaptive method adds generators or layers according to measured information. This can keep early circuits shallow and allocate parameters where they appear useful. It also creates a nested selection process: each proposed operator, gradient estimate, and stopping decision consumes measurements. Comparisons must count the rejected candidates and all intermediate circuits, not only the final compact ansatz.
Expressibility is not the objective
Section titled “Expressibility is not the objective”One expressibility diagnostic compares the distribution of state overlaps generated by the ansatz with the Haar distribution. Such metrics can rank circuit families, but closeness to Haar-random states is not synonymous with task performance. An ansatz may be:
- too restricted to represent the target;
- expressive enough for the target but easy to train;
- highly expressive yet redundant and statistically untrainable;
- expressive in ideal simulation but ineffective after noise.
The desired family is task sufficient, symmetry appropriate, compilable, and trainable under the actual estimator. Maximum generic expressibility is usually neither necessary nor desirable.
Exact Quantum Gradients
Section titled “Exact Quantum Gradients”For
a derivative can often be written as a linear combination of expectation values evaluated by shifted circuits. This avoids differentiating sampled bit strings as though they were deterministic floating-point outputs.
Two-eigenvalue parameter-shift rule
Section titled “Two-eigenvalue parameter-shift rule”Consider one gate
where the Hermitian generator has eigenvalues . Holding all other gates fixed,
For
one has and therefore
This identity is exact for the ideal expectation function. Its finite-shot implementation remains a noisy gradient estimator. If one parameter controls several gates, the chain rule contributes once per occurrence unless a generalized rule exploits the full Fourier spectrum.
Generators with more than two distinct eigenvalues generally require more shifted evaluations or stochastic rules. Hardware noise can also make the implemented function differ from the ideal gate model. Parameter shift is not a universal two-circuit recipe independent of generator, compilation, and noise.
Gradient-estimator variance
Section titled “Gradient-estimator variance”If the two shifted costs are estimated independently with equal shot count and variance , then the Pauli-rotation rule gives
Using common random numbers is not usually literal on hardware, because measurement outcomes from distinct quantum circuits are independent. Correlated calibration blocks or paired execution schedules can nevertheless reduce drift mismatch, and covariance must then be included explicitly.
For parameters, a naive full parameter-shift gradient can require order objective evaluations per iteration, each of which may contain many measurement settings. This multiplicative cost is often more important than the nominal depth of one ansatz circuit.
Finite differences
Section titled “Finite differences”The centered finite difference
has truncation bias of order for a smooth objective, while its sampling variance scales as order . Making smaller reduces truncation bias but amplifies shot noise, calibration quantization, and drift. A finite-difference step therefore requires an explicit bias-variance choice.
Simultaneous perturbation
Section titled “Simultaneous perturbation”Simultaneous perturbation stochastic approximation (SPSA) draws a random direction , often with independent entries in , and uses two cost evaluations:
Its circuit-evaluation count does not grow linearly with per stochastic gradient estimate, but the estimate can have high variance and requires careful gain schedules. “Two evaluations” is not “two shots”: each cost may still require many circuits and measurements.
Classical Optimization
Section titled “Classical Optimization”No optimizer is uniformly best for VQAs. The relevant regime combines noisy function evaluations, periodic parameters, constraints, heterogeneous curvature, device drift, expensive batches, and sometimes discrete ansatz choices.
First-order stochastic methods
Section titled “First-order stochastic methods”An elementary update is
Stochastic-gradient methods can converge with unbiased finite-variance estimators under regularity and step-size assumptions. VQA gradients may be estimated by sampling shots, observable terms, data records, or combinations of these. Such convergence statements do not guarantee reaching the global minimum of a nonconvex objective, and they do not rescue exponentially small signals.
Momentum and adaptive coordinate scaling can improve empirical behavior, but their hyperparameters and stopping rules are part of the algorithm. Tuning them on the reported instances without accounting for selection exaggerates performance.
Derivative-free methods
Section titled “Derivative-free methods”Nelder–Mead, Powell-type searches, coordinate methods, Bayesian optimization, and trust-region model building can be useful when gradients are unavailable or very noisy. Their decisions still depend on resolving cost differences. In a barren plateau those differences can be exponentially suppressed, so choosing a gradient-free optimizer does not by itself remove the information problem.
Curvature and trust regions
Section titled “Curvature and trust regions”Quasi-Newton and trust-region methods use local models to adapt steps to curvature and uncertainty. They can reduce iteration count but may require additional evaluations, stable Hessian approximations, or bounds on model error. A noisy optimizer should reject or shrink a step when the predicted improvement is not statistically resolved, rather than treating every fluctuation as a descent direction.
Quantum Geometry and Natural Gradient
Section titled “Quantum Geometry and Natural Gradient”The same physical state can have many parameter descriptions, and equal Euclidean parameter steps need not represent equal changes in state. For a normalized pure-state family, the Fubini–Study metric is
It defines the local line element
For pure states, the quantum Fisher information matrix is under a common convention. Because authors differ by this factor, an implementation should state which matrix it uses.
Quantum natural gradient chooses a step by solving
The regularizer stabilizes directions in which is singular or poorly estimated. This update is invariant under smooth reparameterizations in the ideal geometric limit, unlike ordinary Euclidean gradient descent. Practical implementations often use diagonal or block-diagonal approximations.
Geometry is not free. Estimating can require additional circuits and shots, and inverting an ill-conditioned matrix amplifies estimator noise. The comparison that matters is total cost to a validated answer, not iterations to a low training objective.
Redundant parameters
Section titled “Redundant parameters”If a nonzero vector satisfies
then the corresponding infinitesimal parameter displacement does not change the physical state to second order in the metric. Such null directions can come from repeated generators, circuit identities, symmetries, or state-dependent stabilizers. They create flat directions without necessarily creating a many-qubit barren plateau.
The effective local dimension is
When , reporting only the number of trainable parameters overstates the local capacity. Pruning redundant coordinates, tying parameters, or choosing a better chart can improve conditioning while leaving the reachable states unchanged locally.
Barren Plateaus
Section titled “Barren Plateaus”A barren plateau is a scaling phenomenon, not merely a plot that looks flat. For an ensemble of instances or random initializations, a common diagnostic is
for some as the number of qubits grows. Typical gradients then become exponentially small. Resolving their sign or relative size with incoherent sampling can require exponentially many measurements.
This definition is about a sequence of problem sizes and a specified sampling ensemble. A small gradient at one point may instead indicate a local optimum, a saddle, a symmetry, a redundant parameter, an insensitive observable, or a poor coordinate scale. Conversely, visibly nonzero gradients on eight qubits do not establish polynomial trainability.
Randomness and unitary designs
Section titled “Randomness and unitary designs”For sufficiently expressive randomly initialized circuits, parts of the circuit can approximate unitary 2-designs. Under the assumptions of the standard results, concentration of measure makes expectation gradients exponentially small. Depth, connectivity, gate ensemble, initial state, and observable all affect when that regime appears.
The lesson is not that every deep circuit has a barren plateau. It is that generic expressibility can erase the structured signal the optimizer needs. Random initialization should be treated as a substantive algorithmic choice.
Global and local costs
Section titled “Global and local costs”The support of the measured observable matters. Global costs can concentrate even for shallow circuits. In particular models, replacing a global fidelity- like objective with a sum of local costs changes exponential gradient scaling to polynomial scaling while preserving the same exact optimum.
That is not a universal license to localize every objective. A surrogate local cost must remain operational: low surrogate loss should imply success on the original task, with a quantitative bound if possible. Otherwise trainability is purchased by optimizing the wrong quantity.
Noise-induced plateaus
Section titled “Noise-induced plateaus”Noise contracts distinguishable states and can suppress both objective variation and gradients. For local unital Pauli noise, rigorous results show exponential gradient decay with circuit depth; when depth grows linearly with , this becomes exponential in system size. Later analyses extend the picture to important nonunital channels and identify noise-induced limit sets as well as plateaus.
This mechanism differs from random-circuit concentration. A carefully initialized problem-inspired ansatz can still lose trainable signal as noisy depth grows. Error mitigation may reduce bias, but its sampling overhead and stability must be included before claiming that it restores trainability.
Gradient-free optimization does not evade concentration
Section titled “Gradient-free optimization does not evade concentration”Derivative-free methods compare nearby or modeled cost values. In a barren plateau, those differences are suppressed along with gradients. Resolving them still demands high precision. Changing optimizers can improve finite-size behavior, but it does not remove an information-theoretic signal-to-noise bottleneck.
Worst-case hardness
Section titled “Worst-case hardness”Training broad classes of variational algorithms is NP-hard in the worst case, even for some systems whose underlying quantum dynamics are classically tractable. This result rules out a universal efficient classical optimizer under standard complexity assumptions. It does not prove that every physical instance, ansatz, or initialization is hard. Complexity statements must retain their quantifiers.
Designing for Trainability
Section titled “Designing for Trainability”There is no single cure, but several design principles can improve the available signal.
Use structure before capacity
Section titled “Use structure before capacity”Choose initial states, generators, symmetries, and connectivity from the problem. Verify that the target is reachable in the chosen sector. Add depth only when a shallower family fails a prespecified approximation test.
Problem inspiration is not itself a proof against barren plateaus. The dynamical Lie algebra generated by the ansatz, the initial state, and the observable can still produce concentration. Numerical gradient scaling over a range of sizes is useful evidence, but a narrow range should not be fitted to an asymptotic law with false confidence.
Initialize near a meaningful map
Section titled “Initialize near a meaningful map”Options include:
- parameters corresponding to a known limiting solution;
- identity-block initialization, where new gate blocks initially compose to the identity;
- parameter transfer from a nearby instance or smaller system;
- layerwise growth, with new layers initialized not to perturb the current state;
- restricted randomization around a structured point.
These methods can preserve initial gradients. They can also bias the search toward one basin, so multiple structured seeds and transparent failure reporting remain important.
Match the cost to the information scale
Section titled “Match the cost to the information scale”Prefer observables whose changes are resolvable at the available shot budget. Normalize weighted sums so that one large coefficient does not dominate optimizer scaling. If constraints are imposed through penalties,
the penalties change curvature and measurement cost. A coefficient too small permits invalid states; one too large can make the landscape ill-conditioned. Report the chosen scale and a constraint-violation curve.
Grow precision with optimization
Section titled “Grow precision with optimization”Early iterations may need only the direction of a large improvement. Late iterations may require enough shots to distinguish small changes. An adaptive policy can increase shots according to estimated variance, gradient norm, or trust-region acceptance.
For an estimated component with standard error , a simple signal diagnostic is
This ratio is not a universal stopping test, especially after selecting the largest component, but it exposes updates driven mainly by measurement noise. Prespecify how the shot budget changes when is small.
Noise, Drift, and Mitigation
Section titled “Noise, Drift, and Mitigation”On hardware, optimization targets the time-dependent implemented objective . If calibration changes during the loop, the optimizer sees a moving surface:
The first term on the right is intended optimization; the last term is drift. Interleaving reference circuits, randomizing execution order, and blocking paired parameter-shift circuits can help distinguish them. A single calibration snapshot before a long run is rarely enough to establish stationarity.
Noise can have several qualitatively different effects:
- bias: the minimum value and minimizer move;
- variance: shot-to-shot outcomes broaden;
- smoothing: high-frequency landscape features are attenuated;
- false structure: coherent errors introduce new parameter dependence;
- leakage or loss: the effective state space and retention probability change;
- nonstationarity: the objective changes during training.
Error mitigation changes the estimator. Readout correction, probabilistic error cancellation, zero-noise extrapolation, symmetry verification, and postselection have different assumptions and overheads. A mitigated objective can have lower bias and much larger variance. Optimization on mitigated values also creates adaptive selection effects. Report both raw and mitigated traces, the mitigation model, calibration data, effective sample overhead, and failure rules. The symmetry-verification specialist owns its projector, check, estimand, covariance, acceptance, and validation contract; this page retains hybrid-loop, gradient, trainability, optimizer, and algorithmic evidence ownership. See Noise in Quantum Information for the canonical noise taxonomy.
A One-Qubit Worked Example
Section titled “A One-Qubit Worked Example”Take
The exact state is
so
The parameter-shift rule returns
Each measurement is a random variable with
Near , the gradient is small because this point is a maximum of the cost, not because of a barren plateau: there is no system-size scaling here. Near , both shifted circuits have deterministic ideal outcomes, but hardware noise can reintroduce variance and bias. Even this trivial example separates exact calculus, shot statistics, and physical implementation.
Stopping and Validation
Section titled “Stopping and Validation”Common stopping rules include:
- a fixed total resource budget;
- a small estimated gradient with uncertainty accounted for;
- no statistically resolved improvement over a window;
- a trust-region radius below a threshold;
- satisfaction of an independently meaningful target;
- convergence of several seeds to compatible validated results.
Stopping when the observed cost first crosses a threshold is biased toward favorable noise fluctuations. Validate the selected point with fresh measurements and report the selection rule.
Minimum validation stack
Section titled “Minimum validation stack”A mature VQA study should include:
- exact or high-accuracy classical simulation at small sizes;
- analytically solvable limiting cases;
- comparison of ideal, sampled, noisy, and hardware objectives;
- fresh-shot re-evaluation of selected parameters;
- multiple seeds with all-run and failure statistics;
- at least one strong classical method under a comparable input-output contract;
- scaling in problem size, depth, parameter count, shots, and total time;
- raw and mitigated results when mitigation is used;
- checks of constraints, conserved quantities, and task-specific observables.
The Quantum Circuit Simulation page covers simulator roles, while Algorithmic Benchmarking owns broader benchmark design.
Do not validate only the training objective
Section titled “Do not validate only the training objective”A low objective can fail to imply useful output because the ansatz optimizes a surrogate, the estimator is biased, the selected point exploits noise, or the observable leaves important degrees of freedom unconstrained. Report task-level quantities that were not directly optimized. Examples include held-out observables, symmetry checks, approximation ratios, residual norms, or independent energy estimates.
Resource Accounting
Section titled “Resource Accounting”Let iteration request circuit configurations with shot counts . A basic quantum acquisition count is
This still omits rejected shots, resets, calibration, mitigation, queueing, and latency. An end-to-end decomposition is
For a batch-access device, the number of quantum–classical round trips may dominate. Report:
- number of objective, gradient, metric, and validation evaluations;
- distinct executable circuits and total shots;
- circuit depth before and after routing;
- two-qubit gate count and measurement settings;
- optimizer iterations, seeds, restarts, and failed jobs;
- classical compute time and memory;
- QPU access time and end-to-end elapsed time;
- mitigation and calibration overhead;
- cost to the accepted result, not only cost of the winning run.
An empirical quantum advantage claim additionally requires a scaling study, a well-defined classical comparator, and uncertainty over representative instances. The Verification of Quantum Advantage page owns that stronger evidentiary standard.
The Quantum Algorithms and Complexity chapter guide places the complete hybrid ledger and evidence classification inside a matched algorithm claim before any speedup conclusion.
Classical Comparison
Section titled “Classical Comparison”The appropriate comparator depends on the output contract. Possibilities include exact diagonalization, tensor networks, Monte Carlo, local search, convex relaxations, classical variational families, or domain-specific heuristics. “Classically hard in general” is not a benchmark.
Use the same:
- instance distribution and preprocessing;
- accuracy or task-quality target;
- success probability;
- inclusion or exclusion of warm starts;
- tuning budget and number of seeds;
- hardware and wall-clock boundary, or a clearly separated asymptotic model.
Small VQA instances are often deliberately classically simulable so that the experiment can be validated. That is scientifically useful, but it is not evidence of computational advantage. Conversely, failure to beat a classical method at small size does not prove asymptotic uselessness. State exactly which claim the evidence supports.
Common Mistakes
Section titled “Common Mistakes”Calling every parameterized circuit a variational algorithm
Section titled “Calling every parameterized circuit a variational algorithm”A parameterized circuit becomes a VQA only when paired with a defined objective, estimation protocol, update rule, and output criterion. A circuit diagram alone is an ansatz.
Equating shallow depth with low total cost
Section titled “Equating shallow depth with low total cost”Training may require thousands of circuit variants, measurement groups, and adaptive round trips. Depth is one resource coordinate.
Reporting the ideal objective as the hardware target
Section titled “Reporting the ideal objective as the hardware target”The device samples an implemented, noisy estimand. Separate statistical error from physical bias.
Treating parameter shift as noiseless differentiation
Section titled “Treating parameter shift as noiseless differentiation”The identity is exact for the modeled expectation function; its experimental estimate has shot noise, drift, and model mismatch.
Maximizing expressibility
Section titled “Maximizing expressibility”Approaching Haar-random behavior can worsen concentration and erase useful structure. Seek task-sufficient capacity.
Diagnosing a barren plateau from one flat trace
Section titled “Diagnosing a barren plateau from one flat trace”Barren plateaus are scaling statements over a specified ensemble. Check gradient distributions across sizes and rule out ordinary stationarity, symmetry, redundancy, and estimator resolution.
Hiding unsuccessful seeds
Section titled “Hiding unsuccessful seeds”Best-of-many selection changes both expected performance and total cost. Report the full run distribution and selection rule.
Comparing optimizer iteration counts
Section titled “Comparing optimizer iteration counts”One iteration may use two circuits, shifted evaluations, a metric tensor, or a large Bayesian batch. Compare acquisition and end-to-end resources.
Validating on adaptively reused measurements
Section titled “Validating on adaptively reused measurements”Fresh data are needed after choosing parameters, ansatz depth, mitigation settings, or the best checkpoint.
Inferring advantage from a decreasing cost
Section titled “Inferring advantage from a decreasing cost”A descending training curve shows that one optimizer extracted some signal on one tested regime. Advantage requires a separate comparative scaling claim.
Where Variational Algorithms Are Used
Section titled “Where Variational Algorithms Are Used”VQAs appear in ground- and excited-state estimation, combinatorial optimization, dynamical simulation, linear-system objectives, state preparation, compilation, error correction, quantum machine learning, and metrology. These applications share hybrid optimization but differ in what their objectives certify.
The connection to Optimal Control is especially close: both optimize parameterized quantum evolution. Control usually treats physical pulse fields and dynamical constraints as primary, whereas circuit VQAs often treat a gate-level ansatz and an algorithmic objective as primary. The mathematical tools overlap, but their canonical engineering boundaries differ.
For reproducibility, archive the complete adaptive history, executable artifacts, environment, seeds, raw records, and analysis. The Reproducible Notebooks and Reporting Standards pages specify the broader artifact and provenance requirements.
Negative Results and Limitations places barren-plateau theorems, empirical trainability failures, classical reversals, and still-open utility claims in their proper logical categories.
Research Status
Section titled “Research Status”The hybrid variational architecture, parameter-shift calculus, and several barren-plateau mechanisms are established. Which structured ansatzes remain trainable and useful at scientifically relevant scale is not settled. Promising initialization, local-cost, adaptive-ansatz, geometry-aware, and shot-frugal methods have strong results in particular regimes, but no universal method removes representation error, nonconvex optimization, hardware noise, and measurement cost simultaneously.
Near-term demonstrations should therefore be described as empirical studies under explicit size, device, and budget conditions. Claims of scalability need measured or proved scaling; claims of utility need a task-level comparator; claims of advantage need end-to-end evidence. The absence of a known large-scale advantage is not a proof that all VQAs fail, and a successful small instance is not evidence that the obstacles disappear.
References
Section titled “References”- A. Peruzzo et al., “A variational eigenvalue solver on a photonic quantum processor,” Nature Communications 5, 4213 (2014), doi:10.1038/ncomms5213.
- J. R. McClean, J. Romero, R. Babbush, and A. Aspuru-Guzik, “The theory of variational hybrid quantum-classical algorithms,” New Journal of Physics 18, 023023 (2016), doi:10.1088/1367-2630/18/2/023023.
- J. Preskill, “Quantum computing in the NISQ era and beyond,” Quantum 2, 79 (2018), doi:10.22331/q-2018-08-06-79.
- K. Mitarai, M. Negoro, M. Kitagawa, and K. Fujii, “Quantum circuit learning,” Physical Review A 98, 032309 (2018), doi:10.1103/PhysRevA.98.032309.
- M. Schuld, V. Bergholm, C. Gogolin, J. Izaac, and N. Killoran, “Evaluating analytic gradients on quantum hardware,” Physical Review A 99, 032331 (2019), doi:10.1103/PhysRevA.99.032331.
- M. Cerezo et al., “Variational quantum algorithms,” Nature Reviews Physics 3, 625–644 (2021), doi:10.1038/s42254-021-00348-9.
- K. Bharti et al., “Noisy intermediate-scale quantum algorithms,” Reviews of Modern Physics 94, 015004 (2022), doi:10.1103/RevModPhys.94.015004.
- D. Wierichs, J. Izaac, C. Wang, and C. Y.-Y. Lin, “General parameter-shift rules for quantum gradients,” Quantum 6, 677 (2022), doi:10.22331/q-2022-03-30-677.
- L. Banchi and G. E. Crooks, “Measuring analytic gradients of general quantum evolution with the stochastic parameter shift rule,” Physical Review A 103, 052414 (2021), doi:10.1103/PhysRevA.103.052414.
- R. Sweke et al., “Stochastic gradient descent for hybrid quantum-classical optimization,” Quantum 4, 314 (2020), doi:10.22331/q-2020-08-31-314.
- J. M. Kübler, A. Arrasmith, L. Cincio, and P. J. Coles, “An adaptive optimizer for measurement-frugal variational algorithms,” Quantum 4, 263 (2020), doi:10.22331/q-2020-05-11-263.
- J. Stokes, J. Izaac, N. Killoran, and G. Carleo, “Quantum natural gradient,” Quantum 4, 269 (2020), doi:10.22331/q-2020-05-25-269.
- B. van Straaten and B. Koczor, “Measurement cost of metric-aware variational quantum algorithms,” PRX Quantum 2, 030324 (2021), doi:10.1103/PRXQuantum.2.030324.
- S. Sim, P. D. Johnson, and A. Aspuru-Guzik, “Expressibility and entangling capability of parameterized quantum circuits for hybrid quantum-classical algorithms,” Advanced Quantum Technologies 2, 1900070 (2019), doi:10.1002/qute.201900070.
- T. Haug, K. Bharti, and M. S. Kim, “Capacity and quantum geometry of parametrized quantum circuits,” PRX Quantum 2, 040309 (2021), doi:10.1103/PRXQuantum.2.040309.
- Z. Holmes, K. Sharma, M. Cerezo, and P. J. Coles, “Connecting ansatz expressibility to gradient magnitudes and barren plateaus,” PRX Quantum 3, 010313 (2022), doi:10.1103/PRXQuantum.3.010313.
- J. R. McClean, S. Boixo, V. N. Smelyanskiy, R. Babbush, and H. Neven, “Barren plateaus in quantum neural network training landscapes,” Nature Communications 9, 4812 (2018), doi:10.1038/s41467-018-07090-4.
- M. Cerezo, A. Sone, T. Volkoff, L. Cincio, and P. J. Coles, “Cost function dependent barren plateaus in shallow parametrized quantum circuits,” Nature Communications 12, 1791 (2021), doi:10.1038/s41467-021-21728-w.
- S. Wang et al., “Noise-induced barren plateaus in variational quantum algorithms,” Nature Communications 12, 6961 (2021), doi:10.1038/s41467-021-27045-6.
- P. Singkanipa and D. A. Lidar, “Beyond unital noise in variational quantum algorithms: noise-induced barren plateaus and limit sets,” Quantum 9, 1617 (2025), doi:10.22331/q-2025-01-30-1617.
- A. Arrasmith, M. Cerezo, P. Czarnik, L. Cincio, and P. J. Coles, “Effect of barren plateaus on gradient-free optimization,” Quantum 5, 558 (2021), doi:10.22331/q-2021-10-05-558.
- M. Larocca et al., “Diagnosing barren plateaus with tools from quantum optimal control,” Quantum 6, 824 (2022), doi:10.22331/q-2022-09-29-824.
- E. Grant, L. Wossnig, M. Ostaszewski, and M. Benedetti, “An initialization strategy for addressing barren plateaus in parametrized quantum circuits,” Quantum 3, 214 (2019), doi:10.22331/q-2019-12-09-214.
- A. Skolik, J. R. McClean, M. Mohseni, P. van der Smagt, and M. Leib, “Layerwise learning for quantum neural networks,” Quantum Machine Intelligence 3, 5 (2021), doi:10.1007/s42484-020-00036-4.
- L. Bittel and M. Kliesch, “Training variational quantum algorithms is NP-hard,” Physical Review Letters 127, 120502 (2021), doi:10.1103/PhysRevLett.127.120502.
- J. C. Spall, “Multivariate stochastic approximation using a simultaneous perturbation gradient approximation,” IEEE Transactions on Automatic Control 37, 332–341 (1992), doi:10.1109/9.119632.
- S. Endo, Z. Cai, S. C. Benjamin, and X. Yuan, “Hybrid quantum-classical algorithms and quantum error mitigation,” Journal of the Physical Society of Japan 90, 032001 (2021), doi:10.7566/JPSJ.90.032001.
Exercises
Section titled “Exercises”1. Derive the parameter-shift rule
Section titled “1. Derive the parameter-shift rule”Let
Show directly that an expectation is a trigonometric polynomial of the form , and derive the two-point shift rule.
Solution
Because ,
Substituting this expression and its adjoint around the observable produces terms proportional to , , and . Double-angle identities therefore give
Its derivative is
Meanwhile,
Thus the shifted expression equals exactly.
2. Allocate shots between observables
Section titled “2. Allocate shots between observables”An objective is
Pilot measurements estimate single-shot standard deviations and . Allocate shots to minimize the variance when the terms are measured independently.
Solution
The continuous optimum obeys
The weights are
Therefore
and
Integer rounding can be adjusted to preserve the total. The derivation assumes independent settings and known stable variances; grouping or drift changes the allocation problem.
3. Separate bias from variance
Section titled “3. Separate bias from variance”A simulator predicts . Hardware repetitions give with standard error . Explain what increasing the shot count by a factor of can and cannot establish.
Solution
Under independent stationary sampling, the standard error falls by a factor of , to about . This measures the hardware estimand more precisely. It does not force that estimand toward . The observed difference of may be physical bias from noise, compilation, readout, or drift. More shots distinguish a stable discrepancy more sharply but do not diagnose or remove it.
4. Identify a redundant parameter
Section titled “4. Identify a redundant parameter”Consider
acting on an arbitrary state. Find a null direction of the state-space metric.
Solution
Because rotations about the same axis compose,
Only the sum matters. The displacement
leaves the unitary exactly unchanged, so is a null direction of the Fubini–Study metric for every input state. The two-coordinate model has at most one effective local dimension.
5. Diagnose a flat trace
Section titled “5. Diagnose a flat trace”A ten-qubit run shows nearly constant cost for 50 iterations. List four checks needed before calling this a barren plateau.
Solution
Check whether:
- gradients and cost differences are below the estimator’s resolution;
- the initialization is already near a stationary point;
- symmetries or redundant parameters make the chosen directions inactive;
- calibration drift or optimizer step sizes obscure genuine changes.
Then study gradient distributions over initializations and a sequence of problem sizes. A barren-plateau claim requires exponential scaling under a specified ensemble, not one flat optimization history.
6. Count a naive gradient batch
Section titled “6. Count a naive gradient batch”An objective has 40 independently measured groups. A circuit has 120 parameters, each appearing in one Pauli rotation. A full parameter-shift gradient uses 2,000 shots per group and shifted point. How many shots does one gradient batch use?
Solution
There are two shifted points per parameter, so the number of group evaluations is
At 2,000 shots each,
This excludes the unshifted objective, metric estimation, calibration, validation, rejected trials, and repeated seeds. The example shows why circuit depth alone is a poor proxy for VQA cost.
7. Audit best-of-seed reporting
Section titled “7. Audit best-of-seed reporting”A study runs 30 seeds, validates only the seed with the lowest noisy training cost, and reports that validation score as typical performance. What is wrong, and what should be reported?
Solution
The chosen seed is an order statistic selected after noisy evaluation. Validating that seed with fresh data removes reuse bias for its score, but it does not make the seed typical. Report all seeds, convergence and failure rates, the distribution of fresh validation scores, the prespecified selection rule, and the total cost of 30 runs. If the operational workflow intentionally uses best-of-30, report its performance as such and include all 30-run cost.
8. Design a stopping rule under shot noise
Section titled “8. Design a stopping rule under shot noise”Propose a stopping rule for a noisy VQA that is more defensible than “stop after three iterations without a lower observed cost.”
Solution
One option is to maintain a trust region and estimate the improvement of the candidate over the incumbent using paired execution blocks. Stop when, for a prespecified number of accepted or attempted steps, the upper confidence bound on improvement is below a practical threshold and the remaining shot budget cannot resolve a smaller target improvement. Also stop at a fixed global budget. Re-evaluate the selected incumbent with fresh shots afterward.
The exact interval method and threshold must be declared in advance, and the analysis must account for repeated adaptive comparisons. The rule distinguishes “no detectable useful improvement” from “three noisy values happened not to decrease.”