Tutorials / T3
PEFT and LoRA on Whisper: a downstream classifier vs. an embedding extractor
Freeze Whisper, inject LoRA adapters in the right places, and train under 2% of the parameters, either for an end-to-end classifier or for embeddings you feed to your own classifier.
Why this tutorial
Speech models are large; clinical audio datasets are small. Fine-tuning every weight is expensive and erases what the model already knows. With PEFT the backbone stays frozen and only small adapters are trained. Parameter-efficient fine-tuning of audio transformers for clinical screening is the subject of my doctoral thesis; this notebook distils the basic mechanics.
The notebook builds two things from the same idea:
- A classifier: Whisper + LoRA + a new head, returning logits.
- An embedding extractor: Whisper + LoRA, no head, feeding any external classifier.
LoRA in one line
A pretrained weight matrix W0 stays frozen, and LoRA learns a low-rank update next to it:
W = W0 + BA, B ∈ ℝd×r, A ∈ ℝr×k, r ≪ min(d, k)
Only A and B are trained, and r is typically between 4 and 32.
The key design choice is target_modules: which layers receive an adapter.
The design decision: where to put the adapters
| Part A: downstream classifier | Part B: embedding extractor | |
|---|---|---|
| Model class | WhisperForAudioClassification | WhisperModel.get_encoder() |
LoRA target_modules | q_proj, v_proj | q_proj, v_proj, fc1, fc2 |
modules_to_save | projector, classifier | none |
| What you keep | encoder + LoRA + head | encoder + LoRA only |
With a head (A), the adapters only need to shift what the encoder attends to: attention is enough. Without a head (B), the embedding itself must move, so the feed-forward layers that reshape it get adapters too.
classifier_lora_config = LoraConfig(
r=8,
lora_alpha=16,
lora_dropout=0.05,
target_modules=["q_proj", "v_proj"], # encoder attention only
modules_to_save=["projector", "classifier"], # new head -> trained in full
)
In Part B a throwaway linear probe pushes a training signal into the adapters and is then discarded.
Results
Measured on an 8-core CPU without a GPU, with openai/whisper-tiny and synthetic tones as data.
| Part A: classifier | Part B: extractor | |
|---|---|---|
| Trainable parameters | 148,226 (1.75%) | 86,016 (1.04%) |
| Saved adapter | 0.60 MB | 0.35 MB |
| Base model in fp32 | 8.3 M parameters, about 33.2 MB | |
| Training time (3 epochs) | 86 s | 98 s |
| Validation accuracy | 0.50 → 1.00 | 1.00 before and after |
The adapter is about 55 times smaller than the base model, and reloading it on a clean backbone reproduces the same accuracy.
Part B is already at 1.00 before adaptation because the synthetic classes differ only in pitch. The mechanism is what transfers: on real tasks such as healthy vs. pathological voice, off-the-shelf embeddings are rarely aligned with the task.
What to take away
- LoRA is not a single recipe. Adapting attention is enough to move a decision boundary; reshaping a representation calls for the feed-forward layers too.
modules_to_saveis for new, randomly initialised layers. A low-rank adapter on random weights makes no sense.- Whisper always pads audio to a 30-second window, so compute per clip is constant regardless of clip length. Plan batch sizes with that in mind.
- The whole flow runs on a laptop CPU with
whisper-tiny. Moving to real data only changes the data-loading section.
References: Hu et al., LoRA, 2021 (arXiv:2106.09685); Radford et al., Whisper, 2022 (arXiv:2212.04356); Hugging Face PEFT.