---
title: Commutators and the Generalized Uncertainty Principle
module: The Formalism of Quantum Mechanics
moduleNumber: 4
lessonNumber: 5
order: 405
summary: >
  The commutator of two observables measures the obstruction to sharing an
  eigenbasis, and it bounds how sharply both can be known at once. We derive the
  generalized uncertainty relation from the Schwarz inequality, recover the
  position–momentum bound as a special case, characterize the minimum-uncertainty
  states that saturate it as Gaussians, and give the energy–time relation its
  correct reading as a lifetime bound rather than a commutator relation.
topics: [The Formalism of Quantum Mechanics]
sources:
  - book: Griffiths & Schroeter
    ref: "Ch. 3; §3.5 The Uncertainty Principle, §3.5.2 The Energy–Time Uncertainty Principle"
  - book: Sakurai & Napolitano
    ref: "Ch. 1; §1.4 The Uncertainty Relation (Schwarz inequality derivation)"
  - book: Shankar
    ref: "Ch. 9 — The Heisenberg Uncertainty Relations"
draft: false
---

The [canonical commutator](/quantum-mechanics/formalism/position-momentum-and-continuous-spectra)
$[\hat x,\hat p]=i\hbar$ says position and momentum have no common eigenbasis. This
lesson makes that qualitative statement quantitative: for any two observables the
product of their spreads in a state is bounded below by the expectation of their
commutator. The bound is exact, saturated by identifiable states, and it applies to
every pair of incompatible observables, not only position and momentum. The
derivation is a direct application of the Cauchy–Schwarz inequality from the
[first lesson](/quantum-mechanics/formalism/hilbert-space-and-dirac-notation).

## The commutator as obstruction

Two observables can be simultaneously sharp only if the state is a common
eigenvector, and [commuting operators share an eigenbasis](/quantum-mechanics/formalism/observables-hermitian-operators-and-eigenvalues).
When $[\hat A,\hat B]\neq 0$ no state is a joint eigenstate, so at least one of the
two spreads is nonzero in every state. The commutator quantifies that residual
disturbance. Two algebraic facts about it are used repeatedly.

- **The commutator of Hermitian operators is anti-Hermitian.** If
  $\hat A^\dagger=\hat A$ and $\hat B^\dagger=\hat B$, then
  $[\hat A,\hat B]^\dagger = (\hat A\hat B-\hat B\hat A)^\dagger = \hat B\hat A-\hat A\hat B = -[\hat A,\hat B]$.
  Its expectation value is therefore purely imaginary.
- **The anticommutator is Hermitian.** The symmetric product
  $\{\hat A,\hat B\}=\hat A\hat B+\hat B\hat A$ satisfies $\{\hat A,\hat B\}^\dagger=\{\hat A,\hat B\}$,
  so its expectation value is real. Together the two decompose the product
  $\hat A\hat B=\tfrac12\{\hat A,\hat B\}+\tfrac12[\hat A,\hat B]$ into Hermitian and
  anti-Hermitian parts, real and imaginary expectations respectively.

$$
% caption: Because incompatible observables share no eigenbasis, the order of two
% measurements matters: preparing a sharp value of one and then measuring the other
% leaves the first no longer sharp, the physical face of a nonzero commutator.
\begin{tikzpicture}[>=stealth, font=\footnotesize, scale=1.0,
  box/.style={draw, minimum width=20mm, minimum height=10mm, align=center}]
\definecolor{acc}{HTML}{4A6FA5}
\node[box] (s0) at (0,0) {sharp A};
\node[box] (s1) at (3.6,0) {measure B};
\node[box, draw=acc, text=acc] (s2) at (7.4,0) {A now spread};
\draw[->, acc, thick] (s0) -- (s1);
\draw[->, acc, thick] (s1) -- (s2);
\node[anchor=north, black!70] at (3.7,-0.85) {order matters when A and B do not commute};
\end{tikzpicture}
$$

## The generalized uncertainty relation

Fix a normalized state $\ket{\psi}$ and define the mean-subtracted operators, whose
spreads are the variances,

$$
\Delta\hat A = \hat A - \langle\hat A\rangle, \qquad
\sigma_A^2 = \braket{\psi|(\Delta\hat A)^2|\psi} = \lVert \Delta\hat A\ket{\psi}\rVert^2.
$$

Both $\Delta\hat A$ and $\Delta\hat B$ are Hermitian, and they share the commutator
of the originals, $[\Delta\hat A,\Delta\hat B]=[\hat A,\hat B]$, since subtracting
constants does not change a commutator.

> **Theorem (Generalized uncertainty principle).** For any two observables and any
> state,
> $$
> \sigma_A^2\,\sigma_B^2 \;\ge\; \left(\frac{1}{2i}\braket{[\hat A,\hat B]}\right)^2
> \;=\; \frac14\bigl\lvert\braket{[\hat A,\hat B]}\bigr\rvert^2,
> $$
> equivalently $\sigma_A\sigma_B\ge\tfrac12\bigl\lvert\braket{[\hat A,\hat B]}\bigr\rvert$.

> **Proof.** Introduce the vectors
> $\ket{f}=\Delta\hat A\ket{\psi}$ and $\ket{g}=\Delta\hat B\ket{\psi}$, so
> $\sigma_A^2=\braket{f|f}$ and $\sigma_B^2=\braket{g|g}$. The Cauchy–Schwarz
> inequality gives
> $$\sigma_A^2\sigma_B^2 = \braket{f|f}\braket{g|g} \ge \lvert\braket{f|g}\rvert^2.$$
> For any complex number $z$, $\lvert z\rvert^2 \ge (\operatorname{Im} z)^2 = \bigl(\tfrac{1}{2i}(z-z^\ast)\bigr)^2$.
> Apply this with $z=\braket{f|g}$. Because $\Delta\hat A,\Delta\hat B$ are
> Hermitian,
> $$
> \braket{f|g} = \braket{\psi|\Delta\hat A\,\Delta\hat B|\psi},
> \qquad \braket{f|g}^\ast = \braket{g|f} = \braket{\psi|\Delta\hat B\,\Delta\hat A|\psi},
> $$
> so their difference is the commutator's expectation,
> $$
> \braket{f|g} - \braket{f|g}^\ast = \braket{\psi|[\Delta\hat A,\Delta\hat B]|\psi}
> = \braket{[\hat A,\hat B]}.
> $$
> Hence $\lvert\braket{f|g}\rvert^2 \ge \bigl(\tfrac{1}{2i}\braket{[\hat A,\hat B]}\bigr)^2$,
> and chaining the two inequalities gives the result.

The bound keeps only the imaginary part of $\braket{f|g}$; discarding the real part,
which is $\tfrac12\braket{\{\Delta\hat A,\Delta\hat B\}}$, is why generic states do
not saturate the relation. The stronger Robertson–Schrödinger inequality retains
that term,

$$
\sigma_A^2\sigma_B^2 \ge \left(\tfrac12\braket{\{\Delta\hat A,\Delta\hat B\}}\right)^2
+ \left(\tfrac{1}{2i}\braket{[\hat A,\hat B]}\right)^2,
$$

and reduces to the Heisenberg form when the symmetric correlation vanishes.

$$
% caption: Schwarz-inequality geometry: the two mean-subtracted vectors have a
% bracket whose magnitude is bounded by the product of their lengths, and the
% imaginary part of that bracket carries the commutator, giving the uncertainty
% bound.
\begin{tikzpicture}[>=stealth, font=\footnotesize, scale=1.0]
\definecolor{acc}{HTML}{4A6FA5}
\coordinate (O) at (0,0);
\coordinate (F) at (4.0,0);
\coordinate (G) at (2.3,2.3);
\draw[->, acc, very thick] (O) -- (F) node[right] {$f$};
\draw[->, acc, very thick] (O) -- (G) node[above] {$g$};
\draw[black, dashed] (G) -- (2.3,0);
\draw[->, black, thick] (O) -- (2.3,0) node[midway, below, black!70] {real part};
\draw[->, black, thick] (2.3,0) -- (G) node[midway, right, black!70] {imag part};
\node[black!70, anchor=west] at (4.2,0.9) {length product bounds the bracket};
\end{tikzpicture}
$$

## Position and momentum

The archetype recovers Heisenberg's original relation. With $\hat A=\hat x$,
$\hat B=\hat p$, and $[\hat x,\hat p]=i\hbar$, the commutator expectation is the
constant $\braket{i\hbar}=i\hbar$, so

$$
\sigma_x\sigma_p \ge \frac12\lvert i\hbar\rvert = \frac{\hbar}{2}.
$$

Position and momentum cannot both have vanishing spread in any state: driving
$\sigma_x\to 0$ forces $\sigma_p\to\infty$. The bound is a property of the state
space, not of the apparatus — no cleverness in measurement evades it, because it
constrains the spreads of outcomes over an ensemble of identically prepared
systems, before any single measurement is made. The physical readings of the
relation collected in the [matter-waves treatment](/quantum-mechanics/matter-waves/the-uncertainty-principle)
— zero-point energy, the size of the hydrogen atom, natural line widths — are all
consequences of this one inequality applied to a confined particle.

The bound is not always a fixed constant. When the commutator is itself an operator,
the right side depends on the state. Angular momentum is the standard case: with
$[\hat J_x,\hat J_y]=i\hbar\hat J_z$, the theorem gives

$$
\sigma_{J_x}\,\sigma_{J_y} \ge \frac\hbar2\bigl\lvert\braket{\hat J_z}\bigr\rvert.
$$

A state with $\braket{\hat J_z}=0$ places no lower bound on the transverse spreads,
which is why an eigenstate of $\hat J_z$ with eigenvalue zero can have arbitrarily
small $\sigma_{J_x}$ and $\sigma_{J_y}$ simultaneously, whereas a state aligned along
$z$ ($\braket{\hat J_z}=j\hbar$) forces a large transverse spread — the algebraic
origin of the "vector-model" cone on which angular momentum cannot point exactly
along an axis.

## Minimum-uncertainty states

The states that saturate $\sigma_x\sigma_p=\hbar/2$ can be found from the two
equality conditions inside the proof. Saturation requires both the Cauchy–Schwarz
inequality and the discard of the real part to be equalities.

- **Schwarz equality** demands $\ket{g}=c\ket{f}$ for some complex $c$, i.e.
  $\Delta\hat p\ket{\psi}=c\,\Delta\hat x\ket{\psi}$.
- **Vanishing real part** demands $\braket{\{\Delta\hat x,\Delta\hat p\}}=0$, which
  with the proportionality forces $c$ to be purely imaginary, $c=i\lambda$ with
  $\lambda$ real.

In the position representation, with $\langle\hat x\rangle=x_0$ and
$\langle\hat p\rangle=p_0$, the condition
$\bigl(\hat p-p_0\bigr)\psi = i\lambda\bigl(\hat x-x_0\bigr)\psi$ becomes a
first-order differential equation,

$$
-i\hbar\,\frac{\d\psi}{\d x} - p_0\psi = i\lambda(x-x_0)\psi,
$$

whose solution is a Gaussian modulated by a plane wave,

$$
\psi(x) = N\,\exp\!\left[-\frac{\lambda}{2\hbar}(x-x_0)^2\right]\exp\!\left[\frac{i p_0 x}{\hbar}\right],
\qquad \lambda>0.
$$

A state saturates the uncertainty bound if and only if it is a Gaussian wave
packet. Their width
is set by $\lambda$: $\sigma_x^2=\hbar/2\lambda$ and $\sigma_p^2=\hbar\lambda/2$, so
$\sigma_x\sigma_p=\hbar/2$ for every $\lambda$. This is why the Gaussian recurs as
the "most classical" state — it is the closest a quantum state comes to a sharp
point in phase space — and it is the ground state of the
[harmonic oscillator](/quantum-mechanics/wave-mechanics-1d/operators-expectation-values-and-the-harmonic-oscillator)
and the [coherent state](/quantum-mechanics/oscillator-and-symmetry/coherent-and-squeezed-states)
for exactly this reason.

$$
% caption: A Gaussian packet saturates the uncertainty bound: its position and
% momentum spreads multiply to exactly $\hbar/2$, and the error box in phase space
% has the smallest area the theory permits.
\begin{tikzpicture}[>=stealth, font=\footnotesize, scale=1.0]
\definecolor{acc}{HTML}{4A6FA5}
\draw[black, ->] (0,0) -- (5.0,0) node[right, black!70] {$x$};
\draw[black, ->] (0,0) -- (0,3.8) node[above, black!70] {$p$};
% minimum-uncertainty box
\draw[acc, very thick] (1.4,1.2) rectangle (2.9,2.7);
\fill[acc!12] (1.4,1.2) rectangle (2.9,2.7);
\draw[acc, very thick] (1.4,1.2) rectangle (2.9,2.7);
\draw[black, <->] (1.4,0.85) -- (2.9,0.85);
\node[anchor=north, black!70] at (2.15,0.85) {width in x};
\draw[black, <->] (3.25,1.2) -- (3.25,2.7);
\node[anchor=west, black!70] at (3.3,1.95) {width in p};
\node[acc, anchor=south west] at (2.9,2.7) {least area};
\end{tikzpicture}
$$

## The energy–time relation

The relation $\Delta E\,\Delta t\gtrsim\hbar/2$ resembles the position–momentum
bound but is not an instance of the generalized theorem, because time is not an
operator in quantum mechanics — there is no $\hat t$ to commute with $\hat H$. Its
correct meaning comes from the rate of change of observables. For any observable
$\hat Q$ with no explicit time dependence, the Heisenberg equation of motion gives

$$
\frac{\d\langle\hat Q\rangle}{\d t} = \frac{i}{\hbar}\braket{[\hat H,\hat Q]},
$$

derived in the [time-evolution lesson](/quantum-mechanics/formalism/time-evolution-schrodinger-and-heisenberg-pictures).
Applying the generalized uncertainty principle to $\hat H$ and $\hat Q$,

$$
\sigma_H\,\sigma_Q \ge \frac12\bigl\lvert\braket{[\hat H,\hat Q]}\bigr\rvert
= \frac\hbar2\left\lvert\frac{\d\langle\hat Q\rangle}{\d t}\right\rvert.
$$

Define $\Delta t = \sigma_Q\bigl/\lvert\d\langle\hat Q\rangle/\d t\rvert$, the time
for the expectation of $\hat Q$ to shift by one standard deviation — the time scale
over which the state changes appreciably as registered by $\hat Q$. Then, writing
$\Delta E=\sigma_H$,

$$
\;\Delta E\,\Delta t \ge \frac{\hbar}{2}\;
$$

> **Definition (Energy–time uncertainty).** $\Delta t$ is not a spread in a time
> measurement but the characteristic time for the system to evolve into a
> distinguishable state, and $\Delta E$ is the spread in energy. A state of sharp
> energy ($\Delta E=0$) is stationary and never changes, so $\Delta t=\infty$; a
> short-lived state must have a correspondingly broad energy distribution.

The relation quantifies the natural line width of a decaying state: an excited level
with lifetime $\tau$ has an energy uncertainty $\Delta E\sim\hbar/\tau$, broadening
its emission line by $\Delta\nu\sim 1/(2\pi\tau)$. A state decaying as
$\lvert c(t)\rvert^2=e^{-t/\tau}$ has amplitude $c(t)\propto e^{-iE_0 t/\hbar}e^{-t/2\tau}$,
whose Fourier transform is a **Lorentzian** in energy,

$$
\lvert c(E)\rvert^2 \propto \frac{1}{(E-E_0)^2 + (\hbar/2\tau)^2},
$$

with full width at half maximum $\Gamma=\hbar/\tau$. The finite lifetime and the
spectral width are Fourier conjugates, exactly as position width and momentum width
are. It also explains why a truly stationary state ($\tau\to\infty$) has an exactly
defined energy, and why rapid processes require access to a wide band of energies.
The interpretation is the recurring pitfall: the relation is about how fast states
evolve, not about an uncertainty principle between two simultaneously measured
quantities.

$$
% caption: A finite state lifetime and its energy width are Fourier conjugates: a
% long-lived level is sharp in energy, a short-lived one broad, with the Lorentzian
% width set by $\Gamma = \hbar/\tau$.
\begin{tikzpicture}[>=stealth, font=\footnotesize, scale=1.0]
\definecolor{acc}{HTML}{4A6FA5}
% long lifetime -> narrow line
\draw[black, ->] (0,0) -- (3.6,0) node[right, black!70] {$E$};
\draw[black, ->] (0,0) -- (0,2.8);
\draw[acc, very thick] plot[domain=0.1:3.5, samples=140]
  (\x, {2.4/(1 + 40*(\x-1.8)^2)});
\node[acc, anchor=south] at (1.8,2.3) {long life};
\node[anchor=north, black!70] at (1.8,-0.15) {narrow};
% short lifetime -> broad line
\begin{scope}[xshift=5.2cm]
\draw[black, ->] (0,0) -- (3.6,0) node[right, black!70] {$E$};
\draw[black, ->] (0,0) -- (0,2.8);
\draw[acc, very thick] plot[domain=0.1:3.5, samples=140]
  (\x, {2.4/(1 + 3.0*(\x-1.8)^2)});
\node[acc, anchor=south] at (1.8,2.3) {short life};
\node[anchor=north, black!70] at (1.8,-0.15) {broad};
\end{scope}
\end{tikzpicture}
$$

> **Worked example.** An excited atomic state decays with mean lifetime
> $\tau=1.6\ \text{ns}$ (a typical allowed optical transition). Taking
> $\Delta t\approx\tau$, the energy width is
> $$
> \Delta E \approx \frac{\hbar}{2\tau}
> = \frac{1.055\times10^{-34}\ \text{J s}}{2(1.6\times10^{-9}\ \text{s})}
> \approx 3.3\times10^{-26}\ \text{J} \approx 2.1\times10^{-7}\ \text{eV}.
> $$
> The corresponding frequency width is
> $\Delta\nu\approx 1/(2\pi\tau)\approx 1.0\times10^{8}\ \text{Hz}=100\ \text{MHz}$,
> the natural linewidth. Against an optical transition frequency near
> $5\times10^{14}\ \text{Hz}$ this is a fractional width of about
> $2\times10^{-7}$, small but the irreducible floor beneath Doppler and collisional
> broadening.

The commutator has now been developed from an obstruction to a quantitative limit.
The final lesson of the module puts the commutator in charge of dynamics: an
observable that commutes with the Hamiltonian is conserved, and the algebra of
commutators generates [time evolution](/quantum-mechanics/formalism/time-evolution-schrodinger-and-heisenberg-pictures)
itself.

[^griffiths-unc]: **Griffiths & Schroeter**, _Introduction to Quantum Mechanics_ 3rd ed., §3.5 — the generalized uncertainty principle from the Schwarz inequality, minimum-uncertainty Gaussians, and §3.5.2 the energy–time relation interpreted through the evolution rate of an observable. Cambridge, 2018.
[^sakurai-unc]: **Sakurai & Napolitano**, _Modern Quantum Mechanics_ 3rd ed., §1.4 — the Schwarz inequality derivation of the uncertainty relation and the role of the anti-Hermitian commutator. Cambridge, 2021.
