Skip to content

Reproducible Notebooks

A reproducible quantum-information notebook is a versioned executable argument that connects a declared physical or algorithmic model to numerical outputs through visible conventions, code, parameters, environment metadata, and validation tests. The notebook is not just its final plot, and successful execution is not by itself evidence that the calculation is correct.

A useful abstraction is

A=(S,I,E,M,V,O,P),\mathcal A = \left( S,I,E,M,V,O,P \right),

where SS is source code and narrative, II the inputs, EE the execution environment, MM the mathematical and physical model, VV the validation record, OO the outputs, and PP the provenance. A reader should be able to inspect every component needed for the claim the artifact supports.

This page owns the quantum-information notebook inventory: stable planned filenames, links to canonical physics pages, minimum computations, admission gates, and the current state of each artifact. It does not redefine general notebook policy. Notebook Index owns the sitewide metadata contract, Reproducibility Status owns status labels, Environments owns dependency policy, and Validation Tests owns the general test taxonomy. Individual concept and algorithm pages remain the canonical homes for the physics.

As of the review date above, the repository does not contain a committed notebooks/quantum-information/ directory. Every filename in the catalogs below is therefore planned. None is downloadable, none has been cleanly executed in a recorded environment, and none may support a numerical claim.

The distinction is basic but important:

planned≠present,present≠executed,executed≠validated,validated once≠currently reproduced.\begin{aligned} \text{planned} &\neq \text{present}, \\ \text{present} &\neq \text{executed}, \\ \text{executed} &\neq \text{validated}, \\ \text{validated once} &\neq \text{currently reproduced}. \end{aligned}

The committed notebook family under notebooks/wave-mechanics-canonical-systems/ provides an implementation precedent. Its successful checks do not transfer to quantum circuits, stochastic sampling, variational optimization, or decoders. Notebook-style pages under Computational Notebooks provide useful open-system contracts, but a page describing a notebook is not evidence that an .ipynb artifact exists or runs.

Inventory state answers, “What artifact is present?” Reproducibility status answers, “What evidence has been recorded for that artifact?” Keep the two axes separate.

Inventory stateMeaningMay support a claim?
plannedfilename and admission contract are reserved; no artifact existsno
present_unreviewedsource exists but clean execution and checks have not been reviewedno
candidateclean execution succeeds and the validation record is under reviewnot yet
admittedsource, environment, checks, and provenance meet this index contractonly under its separate reproducibility label

Once a notebook is admitted, assign the sitewide status reproduced, reproduced_with_warnings, needs_update, broken, or archived. An admitted notebook can later become needs_update after dependency or convention drift. Conversely, a planned filename cannot be called conceptual_only: that label describes a real explanatory artifact, not a proposal.

Different research communities use “repeatability,” “reproducibility,” and “replicability” differently. The inventory uses the labels defined by the reference library rather than silently changing meaning. When independent reimplementation matters, state the actors, artifacts, environment, and tolerance explicitly.

The first family reserves a flat, readable directory:

notebooks/quantum-information/
basic-qubit-states.ipynb
single-two-qubit-gates.ipynb
bell-teleportation.ipynb
grover-search.ipynb
phase-estimation.ipynb
shor-period-finding-toy.ipynb
hamiltonian-simulation-toy.ipynb
vqe-h2.ipynb
qaoa-small-graphs.ipynb
noise-channels-fidelity.ipynb
randomized-benchmarking-simulation.ipynb
stabilizer-syndrome-extraction.ipynb
surface-code-toy-decoder.ipynb
qkd-finite-key-toy.ipynb
ramsey-fisher-information.ipynb
classical-shadows.ipynb

These names are semantic and stable. A later reorganization may add helper modules, data, or environment files without renaming cited notebooks. Once a path is used by a page, figure, release, or external citation, preserve it or publish an explicit migration record.

One notebook should make one computational argument. Shared circuit construction, channel utilities, decoders, and statistical estimators belong in tested modules once they are reused. A notebook may orchestrate those modules and expose checks, but it should not become the only place where a large software system exists.

Evidence chain from canonical physics through notebook execution and validation to cited outputs

A notebook becomes citable evidence only when the model, source, environment, clean execution, physics checks, outputs, and provenance remain connected. Failed validation returns the artifact to revision rather than being hidden by a new plot.

The source notebook is only one node in this chain. A complete release bundle should make the following relations traversable:

  • the notebook links to each canonical physics, convention, and formula page;
  • the environment record identifies an interpreter, direct dependencies, and a reconstruction mechanism;
  • the execution record identifies the source revision, command, runtime class, platform qualifications, and date;
  • each reported result points to a validation test and predeclared tolerance;
  • generated figures or tables point back to source data and the producing notebook or script;
  • the documentation page records which artifact revision it cites.

A cryptographic digest can detect changed bytes, but it cannot establish physical correctness. Hashes protect identity; validation protects a declared claim.

Every candidate notebook should expose the following material near its beginning.

State one concrete question and the level of evidence sought. “Explore Grover’s algorithm” is too broad. “Verify the exact success curve for one and two marked states, then compare finite-shot estimates with binomial uncertainty” is testable.

Link to the canonical pages that define the state, circuit, channel, code, algorithm, or benchmark. The notebook may summarize enough theory to be read, but it must not become a competing derivation.

Record at least:

  • tensor-factor and basis ordering;
  • the mapping between qubit labels, state-vector axes, and displayed bitstrings;
  • gate matrices, rotation signs, and global-phase policy;
  • measurement-bit and classical-register ordering;
  • units, scaled quantities, and time convention;
  • channel representation and vectorization convention;
  • boundary conditions or truncations when a Hilbert space is finite;
  • whether exact arithmetic, floating point, sampling, or hardware data is being used.

The integer index of a basis state is not a convention-free object. If

k=∑j=0n−12jbj,k = \sum_{j=0}^{n-1}2^j b_j,

then b0b_0 is the least-significant bit of kk. A package may nevertheless draw qubit zero at the top, display bitstrings in the reverse order, or order tensor factors differently. The notebook must state each map instead of asking a plot to reveal it.

Name the algorithm and its approximation regime. State whether the result comes from exact state-vector propagation, a stabilizer tableau, a tensor network, a density matrix, trajectories, finite shots, a variational optimizer, a decoder, or a remote processor.

Record direct dependency versions and a lock or reconstruction recipe. Also record the kernel, operating-system or accelerator qualification when relevant, thread and precision settings, input data, and random-number generator. A seed makes a pseudorandom stream repeatable; it does not turn one sample into a statistical error analysis.

Declare checks before examining the headline output. Store machine-readable summary values before rendering figures. Every plot should expose labels, units, sampling counts, and uncertainty where applicable.

The notebook should end with an evidence horizon:

  • which identities or benchmarks passed;
  • which parameter region was tested;
  • which numerical and statistical errors were bounded;
  • which claims remain outside the calculation;
  • which warnings or failures were retained.

Generic “runs without an exception” tests miss the errors most likely to corrupt a quantum-information calculation.

For a state vector, begin with normalization:

ϵnorm=∣⟨ψ∣ψ⟩−1∣.\epsilon_{\mathrm{norm}} = \left| \langle\psi|\psi\rangle-1 \right|.

When comparing pure states, use a phase-insensitive quantity such as

F=∣⟨ψref∣ψ⟩∣2,ϵψ=1−F.F = \left| \langle\psi_{\mathrm{ref}}|\psi\rangle \right|^2, \qquad \epsilon_\psi = \sqrt{1-F}.

Elementwise equality can fail for physically identical states that differ by a global phase. The same issue occurs for gates. For unitary matrices UU and VV, first test unitarity, then compare their induced operations modulo a common phase:

ϵunitary=∥U†U−I∥,ϵU,V=min⁡ϕ∈[0,2π)∥U−eiϕV∥∥U∥.\begin{aligned} \epsilon_{\mathrm{unitary}} &= \lVert U^\dagger U-I\rVert, \\ \epsilon_{U,V} &= \min_{\phi\in[0,2\pi)} \frac{\lVert U-e^{i\phi}V\rVert} {\lVert U\rVert}. \end{aligned}

If controlled versions are compared, a phase that was global on the target operation can become a relative phase. The notebook must test the actual controlled circuit rather than discarding every phase automatically.

For a claimed quantum channel E\mathcal E, test more than trace preservation. With a declared Choi convention, complete positivity and trace preservation require

J(E)⪰0,Tr⁡outJ(E)=Iin.J(\mathcal E)\succeq0, \qquad \operatorname{Tr}_{\mathrm{out}}J(\mathcal E) =I_{\mathrm{in}}.

Report the minimum Choi eigenvalue and the norm of the partial-trace residual before applying any numerical projection. Channel-composition order, idle-noise placement, measurement noise, leakage, and conditional branches must match the timeline described by the canonical model. Noise Simulation owns the broader method-selection and convergence workflow.

When an exact probability distribution p(x)p(x) is available, compare it with the computed distribution before adding sampling. A useful discrepancy is total variation distance,

DTV(p,q)=12∑x∣p(x)−q(x)∣.D_{\mathrm{TV}}(p,q) = \frac{1}{2} \sum_x \left|p(x)-q(x)\right|.

For NshotN_{\mathrm{shot}} independent Bernoulli trials with success probability pp, the standard deviation of the sample proportion is

σp^=p(1−p)Nshot.\sigma_{\hat p} = \sqrt{ \frac{p(1-p)} {N_{\mathrm{shot}}} }.

This formula is a model-based sampling scale, not a universal pass threshold. Use a declared confidence interval, account for multiple comparisons, and separate shot noise from simulator approximation or device drift.

Algorithm notebooks should test the mathematical shape predicted by the canonical derivation, not only one favorable instance. For Grover search with MM marked items among NN candidates,

sin⁡2θ=MN,Pr=sin⁡2 ⁣((2r+1)θ).\sin^2\theta = \frac{M}{N}, \qquad P_r = \sin^2\!\left((2r+1)\theta\right).

The notebook should compare the full iteration curve with this expression, include M=0M=0 and overshoot behavior, and count oracle calls separately from high-level iterations. Grover Search owns the derivation and query-complexity statement.

For phase estimation, use exactly representable dyadic phases as unit tests, then study off-grid phases and finite precision. Report eigenstate overlap, controlled-unitary cost, output-bit convention, circular phase error, and success window. Quantum Phase Estimation owns the precision guarantees and resource interpretation.

Variational notebooks need a classically solvable reference. In an ideal state-vector calculation with a normalized ansatz,

E(θ)=⟨ψ(θ)∣H∣ψ(θ)⟩≥E0.E(\boldsymbol\theta) = \langle\psi(\boldsymbol\theta)| H |\psi(\boldsymbol\theta)\rangle \geq E_0.

Report all declared initialization seeds, optimizer termination reasons, function evaluations, and energy errors. A single successful seed is not an optimizer benchmark, while a noisy estimate slightly below E0E_0 is not automatically a physical violation; its uncertainty and estimator bias must be examined.

For stabilizer generators SiS_i, first check commutation and code-space dimension. If a Pauli error EE is applied, the binary syndrome bit may be defined by

SiE=(−1)siESi,si∈{0,1}.S_iE = (-1)^{s_i} ES_i, \qquad s_i\in\{0,1\}.

The notebook must state whether a syndrome value of one means an anticommuting error, a detector event, or a package-specific boolean. Test all single-qubit errors for a toy code before estimating logical failure rates.

For RR independent trials and KK logical failures, the naive estimator is

p^L=KR.\widehat p_{\mathrm L} = \frac{K}{R}.

When K=0K=0, reporting pL=0p_{\mathrm L}=0 is unjustified. Under the independent Bernoulli model, the approximate one-sided 95% upper limit 3/R3/R is a useful diagnostic. Production studies should use a stated interval method, assess correlations, and use rare-event methods when direct sampling is inefficient.

Planned Catalog: States, Gates, and Protocols

Section titled “Planned Catalog: States, Gates, and Protocols”

All entries in this and the following catalogs are currently planned.

Canonical filenamePhysics ownerMinimum computationAdmission gate
basic-qubit-states.ipynbBloch Sphere for Quantum Informationstate vectors, density matrices, Bloch coordinates, basis measurementsnormalization; positivity; $
single-two-qubit-gates.ipynbSingle-Qubit Gates and Multi-Qubit Gatesgate matrices, tensor placement, controlled and exchange operationsunitarity; basis truth tables; phase-aware identities; ordering fixtures
bell-teleportation.ipynbQuantum Teleportationexact branches and finite-shot circuit for arbitrary input statesfour equiprobable outcomes; unit branch fidelity after correction; unconditioned Bob state I/2I/2

The first two notebooks should be deliberately small. Their main role is to make conventions executable and to supply fixtures reused by every later artifact. A visually attractive Bloch sphere does not compensate for a wrong density-matrix convention or reversed bitstring.

Let the unknown state be

∣ψ⟩=α∣0⟩+β∣1⟩,∣α∣2+∣β∣2=1.|\psi\rangle = \alpha|0\rangle+\beta|1\rangle, \qquad |\alpha|^2+|\beta|^2=1.

Under a declared circuit convention, let Alice’s Bell-measurement outcomes be a,b∈{0,1}a,b\in\{0,1\} and Bob’s conditional pre-correction state be

∣ψab⟩∼XbZa∣ψ⟩.|\psi_{ab}\rangle \sim X^b Z^a|\psi\rangle.

The exact-branch test should verify pab=1/4p_{ab}=1/4 and unit fidelity after the declared correction for random normalized inputs and the basis fixtures ∣0⟩|0\rangle, ∣1⟩|1\rangle, ∣+⟩|+\rangle, and ∣+i⟩|+i\rangle. Before conditioning on Alice’s classical record, Bob’s reduced state must be I/2I/2. Together, these checks expose qubit ordering, classical-bit ordering, correction order, global phase, and the no-signaling average.

Planned Catalog: Algorithms and Simulation

Section titled “Planned Catalog: Algorithms and Simulation”
Canonical filenamePhysics ownerMinimum computationAdmission gate
grover-search.ipynbGrover Searchexact amplitudes and sampled searches for small NN and several MManalytic PrP_r curve; edge cases; oracle-call ledger; bit-order fixtures
phase-estimation.ipynbQuantum Phase Estimationexact and finite-shot phase histograms with configurable precisiondyadic exactness; normalized off-grid distribution; circular-error and success-window checks
shor-period-finding-toy.ipynbShor Algorithmorder finding for small coprime bases and classical post-processingmodular-arithmetic truth table; order check; continued-fraction recovery; explicit failed draws
hamiltonian-simulation-toy.ipynbWhat Is Quantum Simulation?exact evolution and a product formula for a small noncommuting Hamiltonianunitarity; exact comparison; time-step convergence; observable and state error
vqe-h2.ipynbplanned VQE canonical page; simulation overview abovea declared small-basis H2_2 Hamiltonian, exact diagonalization, ansatz, and optimizer ensembleHamiltonian provenance; variational bound; all-seed report; parameter and shot convergence
qaoa-small-graphs.ipynbQAOA; simulation overview aboveexact objective landscape and optimizer runs on named small graphsbrute-force optimum; approximation ratio convention; all-seed distribution; graph-instance record

Toy does not mean evidentially casual. The small size is what permits stronger checks: complete truth tables, exact diagonalization, exhaustive optima, and cross-method comparisons. A Shor notebook must display unsuccessful random bases and post-processing failures rather than presenting one hand-selected run as the algorithm’s success probability.

The VQE artifact must either embed a small Hamiltonian with traceable provenance or generate it through a separately versioned chemistry workflow. The notebook should not silently download changing data. The QAOA artifact must store each graph explicitly and distinguish the best sampled bitstring, expected objective, approximation ratio, and optimizer performance.

Planned Catalog: Noise, Benchmarking, and Codes

Section titled “Planned Catalog: Noise, Benchmarking, and Codes”
Canonical filenamePhysics ownerMinimum computationAdmission gate
noise-channels-fidelity.ipynbCommon Noise Models and Noise SimulationKraus, Choi, and sampled forms of standard one- and two-qubit channelscomplete positivity; trace preservation; limiting cases; representation agreement; fidelity convention
randomized-benchmarking-simulation.ipynbRandomized Benchmarking; Metrics for Quantum Hardwaregenerated Clifford sequences, inverse checks, noisy survival data, and fitted decaysequence inversion; seed ensemble; confidence interval; fit residuals; known injected-error recovery
stabilizer-syndrome-extraction.ipynbStabilizer Formalismtoy code generators, encoded states, Pauli errors, syndrome table, and recoverycommutation; code dimension; exhaustive single-error syndromes; logical-operator checks
surface-code-toy-decoder.ipynbSurface Codesmall code geometry, detector events, a named decoder, and logical-failure trialsboundary and detector conventions; hand-worked fixtures; decoder determinism; confidence interval
classical-shadows.ipynbShadow Tomography; Quantum Measurement as Estimationsmall-state randomized measurements and estimators for declared observablesexact expectation comparison; independent seeds; confidence coverage; observable-set declaration

Randomized benchmarking and decoder notebooks are stochastic benchmark artifacts, not merely circuit demonstrations. Preserve the generated sequence or an unambiguous seed and generator version. Fit uncertainty, model mismatch, and between-sequence variation must remain visible.

A surface-code toy decoder should not claim a threshold from one code distance. Its role is to make geometry, detector conventions, error injection, matching or decoding, logical-observable tracking, and confidence intervals inspectable on small fixtures. Large-scale threshold studies need a separate benchmark design and computational budget.

Planned Catalog: Communication and Metrology

Section titled “Planned Catalog: Communication and Metrology”
Canonical filenamePhysics ownerMinimum computationAdmission gate
qkd-finite-key-toy.ipynbQuantum Key Distributiontransparent BB84 sampling, sifting, parameter estimation, and toy key-length accountingbasis and disclosure ledger; abort cases; confidence parameters; no production-security claim
ramsey-fisher-information.ipynbQuantum Measurement as Estimation and Standard Quantum Limitexact Ramsey probabilities, finite shots, likelihood, estimator, and Fisher informationanalytic probability and Fisher information; estimator bias; confidence coverage; phase-wrap handling

The QKD notebook is pedagogical. A small simulator cannot establish implementation security, composable finite-key security, side-channel resistance, or production randomness quality. It must name every idealization and link security statements to the canonical protocol page.

For a simple Ramsey fringe

p(1∣φ)=1−Vcos⁡φ2,p(1|\varphi) = \frac{1-\mathcal V\cos\varphi}{2},

the classical Fisher information for one binary observation is

F(φ)=[∂φp(1∣φ)]2p(1∣φ)[1−p(1∣φ)].F(\varphi) = \frac{ \left[\partial_\varphi p(1|\varphi)\right]^2 }{ p(1|\varphi)\left[1-p(1|\varphi)\right] }.

The notebook should compare this analytic expression with a numerical derivative away from singular parameterizations, then test estimator bias and interval coverage over repeated synthetic datasets. Plotting a likelihood for one sample is not a coverage study.

Core notebooks should prefer transparent numerical dependencies and explicit matrices for small fixtures. A notebook may offer adapters for common quantum SDKs, but the conceptual result should not depend on one vendor’s bit order, cloud service, account, transpiler default, or calibration archive.

Use three layers when an SDK adds value:

  1. a tool-neutral reference fixture;
  2. an adapter that translates the fixture into the SDK;
  3. a comparison cell that returns both results to the same canonical convention.

Record the package version and backend options. Do not call a result portable because two interfaces use the same gate names. Their matrix conventions, implicit swaps, noise placement, measurement representation, and compiler passes may differ.

Remote hardware runs belong in dated result bundles. The notebook should run without credentials in a simulator-only mode, while a separate opt-in cell records provider, backend identifier, calibration time, compilation artifact, job identifier, shot count, queue exclusions, and returned data. A provider job URL is useful provenance but not an archival copy.

A candidate notebook should have machine-readable metadata equivalent to:

notebook:
path: notebooks/quantum-information/grover-search.ipynb
runtime: python
runtime_class: short
packages:
numpy: '<exact version>'
scipy: '<exact version>'
matplotlib: '<exact version>'
environment: '<lockfile or environment recipe>'
canonical_targets:
- quantum-information/grover-search
source_revision: '<commit or release identifier>'
source_sha256: '<digest>'
tested: '<YYYY-MM-DD>'
status: '<sitewide reproducibility label>'

The notebook should also emit or accompany a compact validation record:

validation:
artifact: grover-search.ipynb
test_id: qi-grover-success-curve
model:
search_space: 32
marked_items: 1
method: exact-state-vector
tolerance:
absolute_probability_error: 1.0e-12
observed:
maximum_absolute_probability_error: '<measured value>'
result: '<passed, warning, failed, or skipped>'
environment_hash: '<digest>'

Placeholders are not valid release metadata. They show the required fields without pretending that a planned notebook has run. Tolerances must be declared from arithmetic, conditioning, approximation order, or statistical coverage, not copied mechanically across notebook families.

A release candidate must execute from a fresh kernel, from first cell to last, using only declared inputs. Restart-and-run-all catches hidden state such as a variable created in a deleted or out-of-order cell. Automated execution can enforce this property, but it should fail on unexpected errors rather than save a superficially complete notebook.

Assign one runtime class:

Runtime classIntended useValidation cadence
shortdeterministic fixtures and small sampled examplesevery relevant change
mediumparameter sweeps or moderate stochastic ensemblesscheduled or affected changes
longexpensive optimization, decoding, or convergence studiesrelease and periodic validation
specializedaccelerator, cluster, proprietary, or hardware-backed workqualified environment with saved summary evidence

Ordinary documentation builds should not depend on long or specialized notebooks. Extract small deterministic checks into tests, retain versioned reference summaries, and run expensive artifacts on a declared schedule. Cached output must be labeled with its source and environment revision; cache presence is not a passing test.

A minimum continuous-check sequence is:

  1. validate notebook structure and metadata;
  2. create or restore the declared environment;
  3. execute in a clean working directory with network access disabled unless explicitly required;
  4. run deterministic and statistical validation cells;
  5. compare machine-readable diagnostics with declared tolerances;
  6. regenerate figures and tables;
  7. verify links and artifact digests;
  8. publish a run report even when a test fails.

Dependency updates trigger a notebook review. Update one layer at a time, inspect changed numerical outputs, and record the reason for accepting any new baseline. Regenerating a golden file before diagnosing the difference turns a regression test into a formatting exercise.

What Reproduction Does and Does Not Establish

Section titled “What Reproduction Does and Does Not Establish”

A clean rerun can establish that a declared source and environment reproduce specified outputs within tolerance. It does not automatically establish:

  • that the physical model is appropriate for a device;
  • that a simulator represents unmodeled noise or drift;
  • that an asymptotic quantum speedup survives input, compilation, correction, and readout costs;
  • that a variational optimizer works on larger or different instances;
  • that a finite-key toy model provides implementation security;
  • that a small decoder benchmark establishes a threshold;
  • that an artifact is independently replicated from a separate implementation.

The evidential label belongs to the exact question tested. Claims, Hype, and Evidence Standards provides the broader vocabulary for distinguishing a simulation, benchmark, experiment, estimate, projection, and application claim.

  • Treating saved output as proof that the notebook still runs.
  • Executing cells out of order and leaving hidden kernel state.
  • Using a fixed seed as a substitute for uncertainty quantification.
  • Comparing state vectors elementwise without accounting for global phase.
  • Comparing bitstrings before reconciling qubit and classical-register order.
  • Sampling a tiny circuit without first checking exact probabilities.
  • Testing trace preservation but not complete positivity for a channel.
  • Reporting zero logical failures as zero logical error probability.
  • Selecting the best optimizer seed and hiding the run distribution.
  • Letting a remote backend, mutable dataset, or network request change an otherwise undocumented input.
  • Updating tolerances or golden outputs without diagnosing the discrepancy.
  • Citing an SDK tutorial as the canonical statement of an algorithm.
  • Calling a planned path, notebook-style page, or rendered figure a reproduced artifact.

A file named phase-estimation.ipynb has been committed. It opens successfully, but no one has executed it from a fresh environment and it contains no validation record. What inventory and reproducibility status should it receive?

Solution

Its inventory state is present_unreviewed. It should not receive a positive reproducibility status and cannot support a numerical claim. Opening a notebook checks neither hidden state nor dependency reconstruction, execution, conventions, or physics.

A two-qubit notebook prepares the state that its author calls ∣01⟩|01\rangle. The state-vector backend reports its only nonzero amplitude at integer index two. Is this necessarily wrong?

Solution

No. Under k=b0+2b1k=b_0+2b_1, index two corresponds to b1b0=10b_1b_0=10 when displayed most-significant bit first. It may represent a state called ∣01⟩|01\rangle under a tensor-factor convention that lists qubit zero first. The notebook is ambiguous until it declares the tensor order, index map, qubit labels, and display order. A useful fixture prepares each computational basis state and records all four mappings.

A simulator returns V=−iUV=-iU for a target one-qubit gate UU. The elementwise difference is large. What should be tested, and when might the phase still matter?

Solution

Test that both matrices are unitary and minimize ∥U−eiϕV∥\lVert U-e^{i\phi}V\rVert over a common phase. As isolated one-qubit operations, they induce the same channel. The phase can matter after forming a controlled operation or combining branches, because a phase that was global on one block can become relative to another block. Test the complete circuit being claimed.

An exact circuit predicts a success probability p=0.36p=0.36, and a sampled run uses Nshot=10 000N_{\mathrm{shot}}=10\,000. Estimate the binomial standard deviation of the measured proportion and explain why “within one standard deviation” is not a complete release rule.

Solution

The model-based standard deviation is

σp^=0.36(0.64)10 000≈0.0048.\sigma_{\hat p} = \sqrt{\frac{0.36(0.64)}{10\,000}} \approx 0.0048.

A release rule should declare a confidence level or hypothesis test before the run, account for repeated tests, and include non-sampling errors. A one-standard-deviation interval has limited coverage, while simulator approximation, correlated shots, or drift can invalidate the independent binomial model.

A noisy-circuit notebook verifies Tr⁡E(ρ)=Tr⁡ρ\operatorname{Tr}\mathcal E(\rho)=\operatorname{Tr}\rho for several random states. Why is that insufficient, and what small-system test should be added?

Solution

Trace preservation does not imply complete positivity, and random state tests do not exhaust the operator space. Construct the Choi matrix using a declared normalization, check its minimum eigenvalue against a justified tolerance, and verify the required partial trace equals the input identity. Also compare Kraus, superoperator, and direct-action forms on fixed fixtures when those representations are claimed equivalent.

A toy decoder observes no logical failures in R=50 000R=50\,000 independent trials. What should be reported instead of p^L=0\widehat p_{\mathrm L}=0?

Solution

Report the count, trial number, sampling assumptions, and a confidence bound. The rule-of-three diagnostic gives an approximate one-sided 95% upper limit

pL≲350 000=6×10−5.p_{\mathrm L} \lesssim \frac{3}{50\,000} = 6\times10^{-5}.

For a release, use a stated interval method and check whether trials are independent. If the target rate is much smaller, direct Monte Carlo has not resolved it and a rare-event method or many more trials are needed.

Twenty VQE initializations were run, but the notebook plots only the lowest energy and reports that the optimizer is reliable. What evidence is missing?

Solution

Report every predeclared seed, initial parameters, final energy, error from exact diagonalization, evaluations, termination reason, runtime, and any failures. Summarize the distribution and sensitivity to shots, optimizer settings, and ansatz depth. The minimum of twenty trials estimates a best-found result, not the probability that a new run succeeds.

A documentation figure has a source notebook link and the notebook has a passing validation cell. Name four additional records needed for a durable claim.

Solution

Useful records include the exact notebook revision or digest, environment lock or digest, input-data version, clean-execution command and date, machine-readable diagnostics and tolerances, source data used by the plot, figure digest, runtime/platform qualification, and the reproducibility status. The page should identify which artifact revision it cites.

Notebook formats, clean execution, dependency recording, version control, physics-aware testing, and artifact review are mature practices. Their application to quantum information is technically demanding because tensor ordering, global phase, stochastic sampling, channel conventions, compiler defaults, calibration drift, rare logical failures, and optimization variability can each produce plausible but wrong output.

The software ecosystem is active. Quantum SDK interfaces, simulator backends, compiler passes, hardware services, accelerator stacks, and environment tools change faster than the underlying algorithms. The durable object is therefore not one package-specific tutorial. It is a small canonical problem, explicit conventions, a tool-neutral reference, versioned adapters, and validation evidence strong enough to reveal drift.

No notebook in this planned family is currently evidence. The first useful milestone is not sixteen colorful files; it is one short artifact that passes the complete contract from a fresh environment and establishes the pattern for the rest.

  • How to Use Computational Notebooks explains how a reader should inspect models, units, convergence, benchmarks, and plots.
  • Notebook Index provides the sitewide inventory and promotion contract.
  • Reproducibility Status defines the maintenance labels used after an artifact exists.
  • Validation Tests provides reusable smoke, shape, Hermiticity, normalization, conservation, analytic-target, convergence, regression, and stochastic checks.
  • Environments owns dependency locks, reconstruction instructions, platform qualifications, and archival environment records.
  • Code Style owns notebook organization and the boundary between narrative orchestration and reusable modules.
  • Quantum Circuit Simulation develops output-aware simulation methods and exact cross-checks.
  • Stabilizer Simulation develops tableau, Pauli-frame, detector, decoder, and logical-rate workflows.
  • Noise Simulation develops density, trajectory, structured, and finite-memory simulations with convergence and uncertainty.
  • Resource Estimation Tools gives the analogous versioned-evidence contract for logical and physical resource forecasts.
  • Why Benchmarking Is Hard defines the task, implementation, reference, uncertainty, and scope contract that a benchmark notebook must execute.
  • Reporting Standards defines the broader human-readable and machine-actionable report manifest that links claims, systems, executables, records, analysis, uncertainty, resources, and restricted artifacts.
  1. Project Jupyter, “The Notebook file format,” nbformat 5.11 documentation, reviewed 2026-08-10, official documentation.
  2. Project Jupyter, “Executing notebooks,” nbclient documentation, reviewed 2026-08-10, official documentation.
  3. A. Rule et al., “Ten simple rules for writing and sharing computational analyses in Jupyter Notebooks,” PLOS Computational Biology 15, e1007007 (2019), doi:10.1371/journal.pcbi.1007007.
  4. G. K. Sandve, A. Nekrutenko, J. Taylor, and E. Hovig, “Ten simple rules for reproducible computational research,” PLOS Computational Biology 9, e1003285 (2013), doi:10.1371/journal.pcbi.1003285.
  5. G. Wilson et al., “Best practices for scientific computing,” PLOS Biology 12, e1001745 (2014), doi:10.1371/journal.pbio.1001745.
  6. G. Wilson et al., “Good enough practices in scientific computing,” PLOS Computational Biology 13, e1005510 (2017), doi:10.1371/journal.pcbi.1005510.
  7. M. D. Wilkinson et al., “The FAIR Guiding Principles for scientific data management and stewardship,” Scientific Data 3, 160018 (2016), doi:10.1038/sdata.2016.18.
  8. T. Kluyver et al., “Jupyter Notebooks — a publishing format for reproducible computational workflows,” in Positioning and Power in Academic Publishing, 87–90 (2016), doi:10.3233/978-1-61499-649-1-87.
  9. Project Jupyter et al., “Binder 2.0 — reproducible, interactive, shareable environments for science at scale,” Proceedings of the 17th Python in Science Conference, 113–120 (2018), doi:10.25080/Majora-4af1f417-011.
  10. Association for Computing Machinery, “Artifact Review and Badging,” version 1.1 policy, reviewed 2026-08-10, official policy.
  11. A. W. Cross et al., “OpenQASM 3: A broader and deeper quantum assembly language,” ACM Transactions on Quantum Computing 3, article 12 (2022), doi:10.1145/3505636.
  12. M. A. Nielsen and I. L. Chuang, Quantum Computation and Quantum Information, 10th anniversary ed., Cambridge University Press (2010), doi:10.1017/CBO9780511976667.