import matplotlib
if not hasattr(matplotlib.RcParams, "_get"):
matplotlib.RcParams._get = dict.get
1.2 From analog to digital#
Computers cannot store an analog signal \(x(t)\) directly. The function takes real-valued inputs and produces real-valued outputs, so even a one-second clip carries an infinite amount of information. To bring sound into the digital world, we have to approximate \(x(t)\) with a finite amount of data. The pipeline that performs this approximation is called analog-to-digital conversion (ADC).
Transforming this continuous sound to digital audio involves discretizing both time and amplitude:
Sampling in time: measure the signal amplitude at discrete, evenly spaced points known as samples.
Quantizing in amplitude: latch each amplitude to its nearest neighbor in a finite set of amplitude values.
Sampling#
Definition 1 (Sampling)
To sample a continuous signal means to measure or evaluate it at a sequence of discrete time points, uniformly spaced at some interval \(\Delta t\).
We call \(\Delta t\) the sampling period; its units are \(\frac{\text{seconds}}{\text{sample}}\). Its reciprocal \(f_s\), in units of \(\frac{\text{samples}}{\text{second}}\), is called the sample rate. The sample rate represents the number of samples captured per second, and the units reveal that \(f_s = 1 / \Delta t\). Sample rates of 44,100 Hz and 48,000 Hz are common values of \(f_s\) in practice; that is, digital audio usually involves tens of thousands of samples per second.
We index samples by an integer \(n\) and adopt the convention
so \(x[0]\) is the signal at time \(t = 0\), \(x[1]\) is its value at time \(t = 1 \cdot \Delta t\), and so on. Continuous-time signals get parentheses (\(x(t)\)); discrete-time sample sequences get square brackets (\(x[n]\)). This distinction will matter throughout the book. You should grow very accustomed to converting between \(\text{samples}\) and \(\text{seconds}\) by dividing or multiplying by \(f_s\).
After sampling, an infinite continuous function has been replaced by a finite ordered sequence of real numbers. Specifically, for some duration \(T\), \(x\) is now a array of numbers of length \(T \cdot f_s\), i.e., \(x \in \mathbb{R}^{T \cdot f_s}\). But the values \(x[n]\) are still real-valued, and we still cannot store real numbers exactly.
Quantization#
Sampling shrank time from a continuum to a finite grid, but we have an analogous problem in amplitude. The values \(x[n] \in \mathbb{R}\) are still real-valued, and a computer cannot store an arbitrary real number exactly.
Definition 2 (Quantization)
To quantize a sample is to latch its amplitude to a nearby element of a finite set of amplitude values.
A common quantization convention in digital audio is signed pulse-code modulation (PCM). We pick an integer bit depth \(b\) and define
as the set of \(2^b\) integers representable in \(b\) bits using two’s complement. We then map each amplitude \(x[n] \in [-1, 1]\) to its quantized integer counterpart by scaling and truncating:
For example, at \(b = 16\) (“CD quality”), \(\mathbb{Z}_{16}\) contains the \(2^{16} = 65{,}536\) integers between \(-32{,}768\) and \(32{,}767\), and amplitudes of \(\{-1.0, 0.0, 1.0\}\) correspond to integers \(\{-32767, 0, 32767\}\) respectively (\(-32768\) is unused).
Quantization is lossy: any two amplitudes that floor to the same integer become indistinguishable in \(\hat{x}[n]\). We will study and quantify the impacts of amplitude quantization when we cover quantization and decibels in more detail.
The interactive below quantizes a sine wave at any bit depth. Drag \(b\) and watch the staircase coarsen.
Drag \(b\): the red staircase is the sine after quantization to \(2^b\) levels, and the gold trace is the error left behind. The audio card underneath always plays the current bit depth.
A signal sampled at \(f_s\) samples per second and quantized to \(b\) bits per sample has a bitrate
For so-called “CD-quality” audio (\(f_s = 44{,}100\), \(b = 16\)), that is \(44{,}100 \cdot 16 = 705{,}600 \left[\frac{\text{bits}}{\text{seconds}}\right]\). To get a more intuitive sense of file size, we can convert to kilobytes per second by chaining the standard relationships \(8\,\text{bits} = 1\,\text{byte}\) and \(1000\,\text{bytes} = 1\,\text{kilobyte}\):
A three-minute song therefore occupies roughly \(88 \cdot 180 \approx 16\) megabytes on disk in this uncompressed form.
Most music is stored and reproduced in stereo, meaning there are two arrays or channels (one for each of our ears) that allow us to perceive basic music spatialization. This doubles the storage size, resulting in \(1{,}411{,}200 \left[\frac{\text{bits}}{\text{seconds}}\right]\) for stereo CD-quality audio. Note that, unless otherwise specified, we are usually referring to mono (single channel) digital audio in this text.
Digital audio is just an array of numbers!#
The punchline here is that, when stored on disk in formats like WAV, digital audio is basically just an array of numbers (samples) together with the sample rate.
When stored on disk, these numbers are usually integers. Why integers and not floats? A 32-bit floating-point number reserves a large fraction of its 32 bits for representing very large and very small magnitudes, i.e., values far outside \([-1, 1]\) that audio simply never uses. The audible range \([-1, 1]\) is a thin sliver of float’s representable range, so most of those bits go to waste on every sample. Integer PCM, by contrast, packs every bit into uniform amplitude resolution inside \([-1, 1]\), giving more precision per bit of storage.
When synthesizing or manipulating samples in memory, the conventions differ. When you write computer music programs, you’ll almost always manipulate \(x[n]\) as a floating-point number in \([-1, 1]\) for arithmetic convenience: mixing, filtering, and synthesis all involve multiplication, addition, and transcendental functions that are awkward and lossy in integer space. Quantization typically only enters the picture at the boundary, when reading samples from a sound file or writing them out.