Tutorials / T2
Fine-tuning Whisper with Optuna, cross-validation and MLflow
A complete, reproducible fine-tuning loop for Whisper on real speech, with grouped cross-validation, Optuna pruning, nested MLflow runs and a locked test set, ending in an explicit decision against a zero-shot baseline.
Why this tutorial
Most fine-tuning notebooks show one run and a number going down. This one builds the full loop a real project needs: hyperparameter search without leakage, traceable runs, and an honest check against the model you started with. It runs openai/whisper-tiny on real English speech, on a laptop CPU, in about nine minutes.
The data and the splits
The notebook uses hf-internal-testing/librispeech_asr_dummy: 73 utterances of read English from LibriSpeech (CC BY 4.0), one speaker, three book chapters. The chapters give the grouping.
| Phase | Data | Decision it allows |
|---|---|---|
| Baseline | test chapter (15 clips) | Zero-shot WER of the untouched model, measured first |
| Hyperparameter search | 2 train chapters (58 clips), leave-one-chapter-out | Pick hyperparameters by mean validation loss |
| Refit | all 58 train clips | Fit the candidate model |
| Test | test chapter, scored once | Keep the fine-tuned model, or keep the base model |
Folds are grouped, never random. With clinical data you group by patient, so the same voice is never on both sides of a split.
Practical detail: the notebook decodes FLAC with soundfile, avoiding the torchcodec/FFmpeg dependency of recent datasets.
The details that make it trustworthy
- A fresh model per fold, so no fold contaminates the next.
- Loss to select, WER to report: loss needs no text generation inside the loop.
- Pruning per fold, reported with two lines:
trial.report(partial_mean_loss, step=fold_index + 1)
if trial.should_prune():
mlflow.set_tag('optuna.state', 'pruned')
raise optuna.TrialPruned()
- A fair WER: Whisper’s English normalizer makes “Mr.” and “MISTER” the same word.
- Readable runs: MLflow nests
tutorial → hpo → trial → fold, plusrefit; the decision is logged on the parent. - Nothing sensitive logged: parameters and aggregate metrics only. No audio, transcripts or weights.
Results
| Run | Learning rate | Epochs | Fold losses | Mean CV loss | Test WER |
|---|---|---|---|---|---|
| Zero-shot baseline | – | – | – | – | 9.7% |
| Trial 0 | 8.5 × 10−6 | 2 | 2.052, 1.744 | 1.898 | – |
| Trial 1 | 7.1 × 10−6 | 1 | 2.559, 2.592 | 2.575 | – |
| Trial 2 (selected) | 2.1 × 10−5 | 2 | 1.501, 1.495 | 1.498 | – |
| Refit on all train | 2.1 × 10−5 | 2 | – | – | 11.5% |
The search worked as intended: it found the best of the three configurations by validation loss. But the fine-tuned model is worse on the locked test chapter than the model it started from, so the pipeline logs the decision keep-base.
That is the lesson, not a bug. Whisper was pre-trained on 680,000 hours of speech; six minutes from one speaker cannot improve on that with full fine-tuning. The pipeline exists to catch exactly this.
Moving to real projects
- Swap the data loader, keep the rest. Group folds by speaker or patient.
- Lock
testbefore the search and score it once. Never choose parameters or epochs from it. - Always log a zero-shot or previous-model baseline next to the candidate, and make the keep-or-replace decision explicit.
- With little data, prefer parameter-efficient adaptation: see PEFT and LoRA on Whisper.
- Log dataset and checkpoint revisions, manifest hashes and the resolved configuration; never audio, transcripts or health information.
References: Panayotov et al., LibriSpeech, ICASSP 2015; Optuna; MLflow Tracking; Whisper in Transformers.