Tutorials / T1
MFCC and PLP from scratch: how machines first learned to listen
Build the two classic speech features step by step in NumPy, watch one frame of real speech travel through both pipelines, and see what noise does to them. The front end every recogniser used before deep learning, and the subject of my first paper.
Why start here
Before Whisper and wav2vec, every speech recogniser began with a hand-designed front end that turned audio into a short vector per 10 ms. Two recipes ruled for thirty years and were the defaults in Kaldi: MFCC (Davis & Mermelstein, 1980) and PLP (Hermansky, 1990). Comparing them in noise was my engineering thesis and my first paper.
One frame, two pipelines
| Step | MFCC | PLP |
|---|---|---|
| Frequency warping | 40 triangular Mel filters | 21 Bark critical bands |
| Loudness | logarithm | equal-loudness curve, then cube root |
| Smoothing | cosine transform (DCT) | all-pole model (LPC, order 12) |
| Output | 13 cepstral coefficients | 13 cepstral coefficients |
Two models of the ear
Both warp frequency the way the ear does: fine resolution below 1 kHz, coarse above. PLP also models how loudness is perceived, and samples the spectrum in 1-Bark steps: 21 bands cover 0–8 kHz.
Checked against references
- MFCC matches
librosato within 10−8 once window, padding and Mel conventions are aligned. - PLP is checked piece by piece: Levinson-Durbin against SciPy’s Toeplitz solver (10−15) and the cepstrum recursion against a brute-force FFT cepstrum (10−16). The widely used
spafepackage departs from Hermansky’s paper in several steps; the notebook shows where, rather than forcing a match.
What noise does
In the 2019 paper we trained 16 Kaldi recognisers on Spanish speech and decoded 3.6 hours of test audio with real noise. The two features tie in quiet conditions; MFCC holds up better as noise grows.
The notebook runs a much smaller experiment: it adds white and babble noise from 20 to 0 dB and measures how far each feature drifts from its clean version.
The three front ends degrade almost in parallel, within 5% of each other. That is not a contradiction of the paper: feature distortion is not word error rate. Without an acoustic model, the notebook can show how features move, not which one a recogniser would prefer.
From MFCC to Whisper
Whisper’s input is the first half of the MFCC pipeline: an 80-band log-Mel spectrogram at 25 ms / 10 ms, which the notebook reproduces with a correlation of 0.94. Deep models dropped the final cosine transform: it existed to decorrelate features for Gaussian models, and a network learns its own mixing. The next two tutorials start from exactly this representation: fine-tuning Whisper and adapting it with LoRA.
References: Davis & Mermelstein, 1980; Hermansky, JASA 1990; Hermansky & Morgan, IEEE TSAP 1994; Povey et al., Kaldi, 2011; Panayotov et al., LibriSpeech, 2015; Ramírez Sánchez, Montalvo Bereau & Calvo de Lara, RIELAC 2019.