7.3 Sampling in practice#
Now that we understand the theorem, how should we choose \(f_s\) for digital audio? The upper limit of human hearing is roughly 20 kHz (more on this in Chapter 15). Treating \(f_{\max} = 20\) kHz, the sampling theorem tells us we want
But there is a subtlety. Sound in the natural world routinely contains frequency content above 20 kHz, even if we cannot hear it. If we simply sample such a sound at 40 kHz, that inaudible high-frequency content will alias down into the audible range and corrupt what we hear. We must remove it before sampling, while it is still separable.
The fix is an anti-aliasing filter: a filter applied to the continuous signal, before sampling, that removes frequency content above the Nyquist frequency. With everything above \(f_s/2\) stripped away, the signal is genuinely bandlimited and sampling is safe.
Fig. 40 An anti-aliasing filter removes content above the Nyquist frequency before sampling. The signal’s energy above 20 kHz (hatched) would otherwise alias into the audible band. Removing it first keeps the sampled signal clean.#
This explains the standard audio sample rates of 44.1 kHz and 48 kHz. Both comfortably exceed the 40 kHz minimum, and they were chosen for two additional reasons:
They leave a little headroom above the 40 kHz minimum to accommodate the fact that real anti-aliasing filters cannot cut off perfectly sharply at exactly 20 kHz.
They are convenient integer multiples of common video frame rates (like 50 and 60 Hz), which simplified building data formats that interleave video with its accompanying audio.