Working papers · speech, health & trustworthy AIVigo, 2026

Tutorials / T3

PEFT and LoRA on Whisper: a downstream classifier vs. an embedding extractor

In one sentence

Freeze Whisper, inject LoRA adapters in the right places, and train under 2% of the parameters, either for an end-to-end classifier or for embeddings you feed to your own classifier.

Why this tutorial

Speech models are large; clinical audio datasets are small. Fine-tuning every weight is expensive and erases what the model already knows. With PEFT the backbone stays frozen and only small adapters are trained. Parameter-efficient fine-tuning of audio transformers for clinical screening is the subject of my doctoral thesis; this notebook distils the basic mechanics.

The notebook builds two things from the same idea:

  1. A classifier: Whisper + LoRA + a new head, returning logits.
  2. An embedding extractor: Whisper + LoRA, no head, feeding any external classifier.

LoRA in one line

A pretrained weight matrix W0 stays frozen, and LoRA learns a low-rank update next to it:

W = W0 + BA,  B ∈ ℝd×r,  A ∈ ℝr×k,  r ≪ min(d, k)

Only A and B are trained, and r is typically between 4 and 32.

W₀ · frozenBA · rank 2ΔW = B·A +×= trainable 64 of 320 parameters (20%) step 24 / 24
Figure 1. LoRA training, animated. The frozen weights W0 (screened) never change. Only the thin matrices B and A (ink) learn. B starts at zero, so the update ΔW = B·A grows from nothing. Here rank 2 trains 20% of the parameters; on Whisper-tiny it is 1.75%.

The key design choice is target_modules: which layers receive an adapter.

The design decision: where to put the adapters

Part A: downstream classifierPart B: embedding extractor
Model classWhisperForAudioClassificationWhisperModel.get_encoder()
LoRA target_modulesq_proj, v_projq_proj, v_proj, fc1, fc2
modules_to_saveprojector, classifiernone
What you keepencoder + LoRA + headencoder + LoRA only

With a head (A), the adapters only need to shift what the encoder attends to: attention is enough. Without a head (B), the embedding itself must move, so the feed-forward layers that reshape it get adapters too.

classifier_lora_config = LoraConfig(
    r=8,
    lora_alpha=16,
    lora_dropout=0.05,
    target_modules=["q_proj", "v_proj"],          # encoder attention only
    modules_to_save=["projector", "classifier"],  # new head -> trained in full
)

In Part B a throwaway linear probe pushes a training signal into the adapters and is then discarded.

Results

Measured on an 8-core CPU without a GPU, with openai/whisper-tiny and synthetic tones as data.

Table 1. What each mode trains and stores.
Part A: classifierPart B: extractor
Trainable parameters148,226 (1.75%)86,016 (1.04%)
Saved adapter0.60 MB0.35 MB
Base model in fp328.3 M parameters, about 33.2 MB
Training time (3 epochs)86 s98 s
Validation accuracy0.50 → 1.001.00 before and after

The adapter is about 55 times smaller than the base model, and reloading it on a clean backbone reproduces the same accuracy.

Part B is already at 1.00 before adaptation because the synthetic classes differ only in pitch. The mechanism is what transfers: on real tasks such as healthy vs. pathological voice, off-the-shelf embeddings are rarely aligned with the task.

What to take away

  • LoRA is not a single recipe. Adapting attention is enough to move a decision boundary; reshaping a representation calls for the feed-forward layers too.
  • modules_to_save is for new, randomly initialised layers. A low-rank adapter on random weights makes no sense.
  • Whisper always pads audio to a 30-second window, so compute per clip is constant regardless of clip length. Plan batch sizes with that in mind.
  • The whole flow runs on a laptop CPU with whisper-tiny. Moving to real data only changes the data-loading section.

References: Hu et al., LoRA, 2021 (arXiv:2106.09685); Radford et al., Whisper, 2022 (arXiv:2212.04356); Hugging Face PEFT.