import matplotlib
if not hasattr(matplotlib.RcParams, "_get"):
matplotlib.RcParams._get = dict.get
7.4 Quantization and decibels#
Sampling is only half the battle in converting continuous sound to digital audio. Recall from Chapter 1 that we must also quantize the real-valued samples so they can be stored in a finite number of bits. Using signed pulse-code modulation with a bit depth of \(b\), we round each amplitude to its nearest representable integer:
Here is the catch. While audio can be perfectly reconstructed from real-valued samples under the Nyquist condition, quantization is a fundamentally destructive operation. Rounding is many-to-one: distinct amplitudes that round to the same integer become indistinguishable. Quantization introduces an irreversible error, heard as quantization noise.
Fig. 41 Quantizing a sine wave at two bit depths. With \(b = 2\) bits (4 levels, left) the staircase is coarse and the error is large. With \(b = 4\) bits (16 levels, right) the error shrinks. Each additional bit doubles the number of levels, halving the error.#
How much noise does quantization add, and how many bits do we need to make it inaudible? Before answering, experiment with the effect for yourself. The interactive below (from Chapter 1) quantizes a sine wave at any bit depth: drag \(b\) and watch the staircase coarsen as levels are removed.
Drag \(b\): the red staircase is the sine after quantization to \(2^b\) levels, and the gold trace is the error left behind. The audio card underneath always plays the current bit depth.
Answering our question quantitatively requires a way to reason about amplitude the way our ears do, which brings us to a short but essential detour.
A detour: amplitude perception and the decibel#
Human hearing spans an astonishing range of sound pressures. The quietest audible sound corresponds to a pressure fluctuation of about 20 μPa (the threshold of hearing), while the onset of pain occurs around 20 Pa (the threshold of pain). That is a factor of one million: six orders of magnitude between the softest and loudest sounds we can handle. This enormous range is precisely why our perception of loudness is roughly logarithmic rather than linear. A logarithmic response lets us hear both rustling leaves and a roaring engine without any adjustment.
To reason about such a wide range, we need a logarithmic unit. That unit is the decibel (dB). A decibel is one tenth of a bel, and it is fundamentally defined over power \(P\) (the rate at which sound energy is transmitted) relative to a reference power \(P_0\):
In computer music we usually work with amplitude \(a\), which is proportional to pressure, rather than power. Since power is proportional to the square of amplitude (\(P \propto a^2\)), the square becomes a factor of 2 outside the logarithm:
Definition 22 (Decibel (amplitude))
The level of an amplitude \(a\) relative to a reference amplitude \(a_0\), expressed in decibels (dB), is
The decibel is a relative unit: it always compares an amplitude \(a\) to some reference \(a_0\). Two conventions for that reference are common:
dBFS (decibels relative to full scale) uses \(a_0 = 1\), the maximum amplitude before clipping. So \(\text{dBFS} = 20\log_{10}(a)\). Because amplitudes are at most 1, dBFS values are normally negative, and a value of 0 dBFS means the signal is right at the clipping point.
dB SPL (sound pressure level) uses the physical reference \(p_0 = 20\) μPa, the threshold of hearing, so \(\text{dB SPL} = 20\log_{10}(p/p_0)\). This grounds the decibel in real-world pressure. It matters less for us in this book, but it connects our unitless amplitudes back to physical sound.
Because the decibel is logarithmic, multiplying an amplitude corresponds to adding decibels. Two relationships are worth committing to memory:
Multiplying or dividing an amplitude by 10 is a change of \(\pm 20\) dB (since \(20\log_{10}(10) = 20\)).
Multiplying or dividing an amplitude by 2 is a change of about \(\pm 6\) dB (since \(20\log_{10}(2) \approx 6\)).
With these, we can quickly estimate the dynamic range of human hearing: six orders of magnitude is \(6 \times 20 = 120\) dB. In practice, ambient background noise usually limits the usable range to something closer to 100 dB.
To calibrate your ear to the scale, here is the same 440 Hz sine tone at a ladder of levels, each 20 dB (a factor of 10 in amplitude) below the last:
A 440 Hz sine at five levels, each 20 dB quieter than the previous (a tenfold drop in amplitude). Even the faintest, at \(-86\) dBFS, is audible on most systems, hinting at the wide dynamic range our ears command. Be careful not to turn your volume up to hear the quiet ones, or the loud ones may surprise you.
Pyquist provides helpers to convert between amplitudes and dBFS:
import pyquist as pq
pq.helper.db_to_amplitude(-6.0) # ~0.501 (halving amplitude)
pq.helper.db_to_amplitude(-20.0) # ~0.1 (one tenth amplitude)
pq.helper.amplitude_to_db(0.125) # ~-18.06 dB
How many bits are enough?#
We can now quantify quantization noise in perceptual terms. With \(b\) bits spanning the full-scale range \([-1, 1]\), the spacing between adjacent quantization levels is roughly \(1/2^{b-1}\). If we assume each sample falls at a random point between two levels, the typical rounding error is about half that spacing, or roughly \(1/2^b\).
The key consequence follows immediately. Each time we add one bit, we double the number of levels, which halves the quantization error. And halving an amplitude, as we just learned, is a reduction of about 6 dB. Therefore:
Important
Each additional bit of depth reduces quantization noise by about 6 dB, and so buys about 6 dB of dynamic range.
This gives us a simple rule for choosing a bit depth. At \(b = 16\) bits, we get about \(16 \times 6 = 96\) dB of dynamic range. That is close to the roughly 100 dB practical limit of human hearing, which is exactly why 16 bits per sample (“CD quality”) is enough for transparent audio. It is also conveniently a multiple of 8 bits, aligning with computer word sizes. Professional workflows sometimes use 24 bits to leave extra headroom during editing, but 16 bits is perceptually sufficient for final playback.