10. Frame-based Processing

10. Frame-based Processing#

So far we have studied two extremes of how a computer handles time. When we studied sampling in Chapter 7, we saw that music audio is usually sampled at more than \(40{,}000\) times per second, fast enough to capture the highest frequencies we can hear. When we studied the Fourier transform in Chapter 5 and its practical cousin the DFT in Chapter 8, we did the opposite: we integrated across all of time to produce a single summary of a sound’s frequency content, in effect measuring it just once no matter how long it was (a “rate” of \(0\) measurements per second).

Most phenomena in music live between these two extremes. The attack of a plucked string lasts about a hundredth of a second, a four-on-the-floor kick drum at 120 BPM lands twice a second, a pianist playing Bach’s Prelude in C plays around five notes a second, and the world’s fastest drummer can manage twenty strokes a second. None of these needs the microsecond precision of individual samples, but all of them are lost to the time integration of a global Fourier transform.

Table 3 The rate at which things happen in music, from a single Fourier measurement to individual samples. The musically interesting middle (blue) is what this chapter is about.#

Phenomenon

Interval

Rate

Fourier transform (whole recording)

\(\red{\infty}\)

\(\red{0}\) Hz

Kick drum at 120 BPM

\(\blue{500}\) ms

\(\blue{2}\) Hz

Melody (Bach, ~5 notes/sec)

\(\blue{200}\) ms

\(\blue{5}\) Hz

World’s fastest drummer

\(\blue{50}\) ms

\(\blue{20}\) Hz

Instrument attack

\(\blue{10}\) ms

\(\blue{100}\) Hz

Audio samples

\(\red{0.023}\) ms

\(\red{44{,}100}\) Hz

How do we process phenomena that happen at these intermediate, musically intuitive rates, say tens to hundreds of times per second? The answer is frame-based processing, a family of techniques that aggregate audio samples into chunks called frames and then analyze or manipulate those frames. It is the foundation for techniques like granular synthesis, the spectrogram, time stretching, and real-time audio software. We will examine all of those techniques in this chapter.

Throughout the chapter we will this recording of a jazz trio as a running example:

Eight seconds of a jazz trio, which we will slice, scramble, stretch, and analyze throughout this chapter. 725677 by draganov89, License: Attribution NonCommercial 4.0.