Working papers · speech, health & trustworthy AIVigo, 2026

Research / Robust speech recognition and spoken search / [13]

MFCC or PLP? What happens to speech recognition when the room gets noisy

Evaluación de Rasgos Acústicos para el Reconocimiento Automático del Habla en Escenarios Ruidosos usando Kaldi. Revista Ingeniería Electrónica, Automática y Comunicaciones (RIELAC) 40(3), 2019.

In one sentence

Across 16 Kaldi systems trained on Spanish speech, MFCC and RASTA-PLP perform alike in quiet conditions, but MFCC loses less accuracy as noise grows, while RASTA-PLP decodes faster.

16Kaldi systems
18real noise types
3.6 hSpanish test speech
−3.6 ptsWER with MFCC, noisy outdoors

The problem

Before deep learning, every speech recogniser started by turning audio into a compact acoustic feature. Two recipes dominated: MFCC (Davis & Mermelstein, 1980), built on a Mel filterbank, and PLP (Hermansky, 1990), built on models of human hearing and linear prediction. There was no consensus on which one holds up better when speech is recorded in noise.

What we did

We trained eight Kaldi recognisers of increasing complexity, from monophone HMM-GMMs to SGMMs, a DNN and a system combination, each twice: once with MFCC and once with RASTA-PLP, with identical settings otherwise. The training data was spontaneous Spanish speech from the TC-STAR corpus. We then decoded 3.6 hours of test speech clean and with 18 real-world noises from the DEMAND database, mixed at 5–15 dB and 15–25 dB SNR, indoors and outdoors.

MFCC RASTA-PLP 0%10%20%30%40% Clean 20.8% 21.1% Indoor 15–25 dB 22.2% 22.9% Outdoor 15–25 dB 21.7% 22.9% Indoor 5–15 dB 28.8% 31.9% Outdoor 5–15 dB 30.7% 34.4%
Figure 1. Word error rate of the combined Kaldi system, from quiet to noisiest scenario. The two features tie in clean speech; the gap opens as noise grows. Lower is better.

What we found

In clean speech the choice barely matters: for systems with the same acoustic model, the WER differences stay under one point, with one exception. As noise grows, MFCC systems lose less accuracy, and the gap widens with every step down in SNR, reaching 3.6 points for the combined system and 7.6 points for a triphone system in the noisiest outdoor scenario. RASTA-PLP had its own advantage: decoding was faster for every system, 50 minutes less over the test set for the tri3 model.

My contribution

This was my engineering thesis. I designed and ran all experiments: trained the 16 Kaldi systems, built the noisy test sets with FaNT and DEMAND, and analysed WER and decoding times.

Try it yourself

The tutorial MFCC and PLP from scratch builds both features step by step on real speech.

Why it matters

Feature choice is a trade-off, not a free lunch: robustness for noisy deployments versus speed for real-time decoding. The same question, which representation of sound survives the real world, runs through all my later work on voice as a clinical signal.

Cite

Ramírez Sánchez J.M., Montalvo Bereau A.R., Calvo de Lara J.R. Evaluación de Rasgos Acústicos para el Reconocimiento Automático del Habla en Escenarios Ruidosos usando Kaldi. Revista Ingeniería Electrónica, Automática y Comunicaciones (RIELAC) 40(3), 2019.