Skip to content

Dirac’s Transformation Theory

Matrix mechanics and wave mechanics first looked like different theories. Matrix mechanics used arrays of transition quantities; wave mechanics used differential equations for wavefunctions. Dirac’s transformation theory helped reveal a deeper unity: both are representations of the same quantum structure.

In modern language, a state is not identical with any one column vector or wavefunction. A representation is a way of writing the state relative to a chosen basis or complete set of observables. Dirac’s transformation theory made that idea central.

By 1926, quantum mechanics had several successful languages:

  • Heisenberg’s matrix mechanics emphasized transition quantities and noncommutative products.
  • Born and Jordan’s formulation clarified the matrix algebra and canonical commutators.
  • Schrödinger’s wave mechanics described systems by wavefunctions and differential equations.
  • Born’s probability interpretation connected amplitudes to observed frequencies.

The urgent problem was not only to show that these languages gave the same spectra in examples. It was to understand why they were the same theory.

Dirac’s contribution was to treat transformations between descriptions as part of the theory itself. Instead of asking whether the matrix or wave picture was more real, transformation theory asks how to pass between representations without changing the underlying physical state.

In finite-dimensional modern notation, let {∣a⟩}\{\lvert a\rangle\} and {∣b⟩}\{\lvert b\rangle\} be two orthonormal bases. A state ∣ψ⟩\lvert\psi\rangle has components

ψa=⟨a∣ψ⟩,ψb=⟨b∣ψ⟩.\psi_a = \langle a\vert\psi\rangle, \qquad \psi_b = \langle b\vert\psi\rangle.

The overlap between bases is the transformation amplitude

Sba=⟨b∣a⟩.S_{ba} = \langle b\vert a\rangle.

Then

ψb=∑a⟨b∣a⟩ψa.\psi_b = \sum_a \langle b\vert a\rangle \psi_a.

The same idea works for continuous representations, with sums replaced by integrals. Position and momentum wavefunctions are related by a transformation kernel:

ψ(x)=⟨x∣ψ⟩,ϕ(p)=⟨p∣ψ⟩.\psi(x) = \langle x\vert\psi\rangle, \qquad \phi(p) = \langle p\vert\psi\rangle.

With a common convention,

⟨x∣p⟩=12πℏeipx/ℏ,\langle x\vert p\rangle = \frac{1}{\sqrt{2\pi\hbar}} e^{ipx/\hbar},

so

ψ(x)=∫−∞∞⟨x∣p⟩ϕ(p) dp.\psi(x) = \int_{-\infty}^{\infty} \langle x\vert p\rangle \phi(p)\,dp.

The kernel is not a new physical wave by itself. It is a transformation amplitude between two representations.

Abstract quantum state connected to matrix components, wavefunction representation, and transformation kernels between bases

Dirac’s transformation viewpoint separates the abstract state from its representations. Matrix components, wavefunctions, and overlap kernels are different ways of expressing the same underlying quantum object.

Transformation theory points toward the modern state-observable language. A state can be represented in the eigenbasis of an observable AA, in the eigenbasis of another observable BB, or in a continuous position or momentum representation. The physical state is the invariant object behind those coordinate descriptions.

For a discrete nondegenerate observable AA, the modern spectral form is

A=∑aa ∣a⟩⟨a∣.A = \sum_a a\, \lvert a\rangle\langle a\rvert.

The amplitude for finding outcome aa in state ∣ψ⟩\lvert\psi\rangle is

⟨a∣ψ⟩,\langle a\vert\psi\rangle,

and the probability is

P(a)=∣⟨a∣ψ⟩∣2.P(a) = \lvert\langle a\vert\psi\rangle\rvert^2.

For another observable BB, the relevant amplitudes are instead ⟨b∣ψ⟩\langle b\vert\psi\rangle. The transformation amplitudes ⟨b∣a⟩\langle b\vert a\rangle tell how the two descriptions are related.

This is why representation changes are not optional bookkeeping. They are how quantum mechanics expresses incompatible measurement contexts, alternative bases, and equivalent formulations.

Dirac’s later bra-ket notation packages the transformation viewpoint compactly:

  • a ket ∣ψ⟩\lvert\psi\rangle denotes the state vector;
  • a bra ⟨a∣\langle a\rvert extracts an amplitude in the aa representation;
  • an inner product ⟨a∣ψ⟩\langle a\vert\psi\rangle is a component or transition amplitude;
  • an outer product ∣a⟩⟨a∣\lvert a\rangle\langle a\rvert is a projector in the nondegenerate discrete case.

Historically, the notation and the conceptual framework matured over time. It is still useful to read the notation as a compact summary of transformation theory: kets are not column vectors until a basis is chosen, and bras are not merely decorative row symbols. They encode the dual pairing that produces amplitudes.

For the working convention, see Bra-Ket Notation. For the mathematical translation, see Dirac Notation as Linear Algebra.

Transformation theory also explains why later operator and field-theory languages could be representation-flexible. In operator quantum mechanics, the central objects are states, observables, amplitudes, and transformations. A coordinate-space wavefunction, an energy-basis column vector, and a momentum-space amplitude are different representations, not different physical theories.

This becomes especially important in more advanced settings:

  • the Heisenberg picture moves time dependence into operators;
  • scattering theory studies transition amplitudes between asymptotic states;
  • path integrals compute kernels such as ⟨xf,tf∣xi,ti⟩\langle x_f,t_f\vert x_i,t_i\rangle;
  • quantum field theory organizes states and operators in Fock-space and field representations.

The historical bridge should not be overstated. Dirac’s transformation theory did not by itself provide all later mathematical rigor or all quantum-field-theoretic machinery. But it gave a durable conceptual grammar: choose a representation when useful, transform when necessary, and keep the representation-independent object in view.

Transformation theory clarified equivalence and notation, but it did not remove every subtlety. Continuous bases require generalized eigenvectors and distributions. Unbounded operators require domains. Measurements require probability rules and state-update models. Infinite systems and quantum fields can have inequivalent representations.

Those caveats are not failures of the transformation idea. They are reminders that the clean finite-dimensional formulas are ideal guides, not the whole mathematical story.

  • Treating the wavefunction as the state itself rather than a representation of the state.
  • Treating a matrix as the operator itself without specifying the basis.
  • Thinking transformation theory is only about time evolution. It is also about changing representation.
  • Forgetting that ⟨b∣a⟩\langle b\vert a\rangle is an amplitude, not a classical conditional probability.
  • Applying finite-dimensional basis formulas to continuous spectra without distributional care.
  • Reading modern bra-ket notation back into every 1920s paper as if it was already standardized.
  • P. A. M. Dirac, “The fundamental equations of quantum mechanics,” Proceedings of the Royal Society A 109, 642-653, 1925, DOI: 10.1098/rspa.1925.0150.
  • P. A. M. Dirac, “The physical interpretation of the quantum dynamics,” Proceedings of the Royal Society A 113, 621-641, 1927, DOI: 10.1098/rspa.1927.0012.
  • P. A. M. Dirac, The Principles of Quantum Mechanics, 4th ed., Oxford University Press, 1958.
  • B. L. van der Waerden, ed., Sources of Quantum Mechanics, Dover, 1968.
  • M. Jammer, The Conceptual Development of Quantum Mechanics, 2nd ed., American Institute of Physics, 1989.
  • J. Mehra and H. Rechenberg, The Historical Development of Quantum Theory, Springer, 1982-2001.
  1. Let {∣a⟩}\{\lvert a\rangle\} and {∣b⟩}\{\lvert b\rangle\} be orthonormal bases with Sba=⟨b∣a⟩S_{ba}=\langle b\vert a\rangle. Show that ψb=∑aSbaψa\psi_b=\sum_a S_{ba}\psi_a.
Solution

Insert the identity in the aa basis:

I=∑a∣a⟩⟨a∣.I = \sum_a \lvert a\rangle\langle a\rvert.

Then

ψb=⟨b∣ψ⟩=∑a⟨b∣a⟩⟨a∣ψ⟩=∑aSbaψa.\psi_b = \langle b\vert\psi\rangle = \sum_a \langle b\vert a\rangle \langle a\vert\psi\rangle = \sum_a S_{ba}\psi_a.
  1. Using ⟨x∣p⟩=(2πℏ)−1/2eipx/ℏ\langle x\vert p\rangle=(2\pi\hbar)^{-1/2}e^{ipx/\hbar}, write ψ(x)\psi(x) in terms of ϕ(p)\phi(p).
Solution

The position-space wavefunction is obtained by inserting the momentum representation:

ψ(x)=∫−∞∞⟨x∣p⟩ϕ(p) dp=12πℏ∫−∞∞eipx/ℏϕ(p) dp.\psi(x) = \int_{-\infty}^{\infty} \langle x\vert p\rangle \phi(p)\,dp = \frac{1}{\sqrt{2\pi\hbar}} \int_{-\infty}^{\infty} e^{ipx/\hbar} \phi(p)\,dp.
  1. Why is ⟨b∣a⟩\langle b\vert a\rangle not a classical conditional probability?
Solution

It is generally complex and can interfere with other amplitudes. A probability is obtained only after the relevant amplitudes have been combined and squared according to the Born rule. The quantity ⟨b∣a⟩\langle b\vert a\rangle is therefore a transformation amplitude, not an ordinary conditional probability.

  1. Explain why calling ψ(x)\psi(x) “the state” can be misleading.
Solution

ψ(x)=⟨x∣ψ⟩\psi(x)=\langle x\vert\psi\rangle is the position representation of the abstract state ∣ψ⟩\lvert\psi\rangle. The same state can also be represented by momentum-space amplitudes, energy-basis components, or another basis. Calling ψ(x)\psi(x) the state hides the basis dependence and can make representation changes look like physical changes.