Working papers · speech, health & trustworthy AIVigo, 2026

Tutorials / T1

MFCC and PLP from scratch: how machines first learned to listen

In one sentence

Build the two classic speech features step by step in NumPy, watch one frame of real speech travel through both pipelines, and see what noise does to them. The front end every recogniser used before deep learning, and the subject of my first paper.

Why start here

Before Whisper and wav2vec, every speech recogniser began with a hand-designed front end that turned audio into a short vector per 10 ms. Two recipes ruled for thirty years and were the defaults in Kaldi: MFCC (Davis & Mermelstein, 1980) and PLP (Hermansky, 1990). Comparing them in noise was my engineering thesis and my first paper.

One frame, two pipelines

waveform · 5.9 s · LibriSpeech power spectrum · 0–8 kHz Mel · 40 bandsMFCC · 13 Bark · 21 bandsPLP · 13 MFCC over timePLP over time →→→→
Figure 1. A real LibriSpeech utterance, frame by frame. Each 25 ms frame becomes a power spectrum, which splits into two paths: a Mel filterbank and a cosine transform give 13 MFCC; Bark critical bands, loudness compression and an all-pole model give 13 PLP coefficients. Below, both matrices print over time. All values come from the notebook.
Table 1. The same five steps, two views of hearing.
StepMFCCPLP
Frequency warping40 triangular Mel filters21 Bark critical bands
Loudnesslogarithmequal-loudness curve, then cube root
Smoothingcosine transform (DCT)all-pole model (LPC, order 12)
Output13 cepstral coefficients13 cepstral coefficients

Two models of the ear

Mel · 40 triangles (MFCC) Bark · 21 critical bands (PLP) 0 kHz2 kHz4 kHz6 kHz8 kHz
Figure 2. The filters, on the same 0–8 kHz axis. Mel triangles are narrow and dense at low frequencies; Bark filters follow Hermansky's critical-band masking curve, with a gentle low side and a steep high side.

Both warp frequency the way the ear does: fine resolution below 1 kHz, coarse above. PLP also models how loudness is perceived, and samples the spectrum in 1-Bark steps: 21 bands cover 0–8 kHz.

Checked against references

  • MFCC matches librosa to within 10−8 once window, padding and Mel conventions are aligned.
  • PLP is checked piece by piece: Levinson-Durbin against SciPy’s Toeplitz solver (10−15) and the cepstrum recursion against a brute-force FFT cepstrum (10−16). The widely used spafe package departs from Hermansky’s paper in several steps; the notebook shows where, rather than forcing a match.

What noise does

In the 2019 paper we trained 16 Kaldi recognisers on Spanish speech and decoded 3.6 hours of test audio with real noise. The two features tie in quiet conditions; MFCC holds up better as noise grows.

MFCC RASTA-PLP 0%10%20%30%40% Clean 20.8% 21.1% Indoor 15–25 dB 22.2% 22.9% Outdoor 15–25 dB 21.7% 22.9% Indoor 5–15 dB 28.8% 31.9% Outdoor 5–15 dB 30.7% 34.4%
Figure 3. From the paper: word error rate of the combined Kaldi system with each feature. Lower is better.

The notebook runs a much smaller experiment: it adds white and babble noise from 20 to 0 dB and measures how far each feature drifts from its clean version.

MFCCPLPRASTA-PLP white noise 0.60.81.01.2 20 dB15 dB10 dB5 dB0 dB babble noise 0.60.81.01.2 20 dB15 dB10 dB5 dB0 dB
Figure 4. How far each feature moves from its clean version as noise grows (RMS difference after normalisation; 1.41 would mean unrelated). Ten utterances, CPU, seconds.

The three front ends degrade almost in parallel, within 5% of each other. That is not a contradiction of the paper: feature distortion is not word error rate. Without an acoustic model, the notebook can show how features move, not which one a recogniser would prefer.

From MFCC to Whisper

Whisper’s input is the first half of the MFCC pipeline: an 80-band log-Mel spectrogram at 25 ms / 10 ms, which the notebook reproduces with a correlation of 0.94. Deep models dropped the final cosine transform: it existed to decorrelate features for Gaussian models, and a network learns its own mixing. The next two tutorials start from exactly this representation: fine-tuning Whisper and adapting it with LoRA.

References: Davis & Mermelstein, 1980; Hermansky, JASA 1990; Hermansky & Morgan, IEEE TSAP 1994; Povey et al., Kaldi, 2011; Panayotov et al., LibriSpeech, 2015; Ramírez Sánchez, Montalvo Bereau & Calvo de Lara, RIELAC 2019.