Working papers · speech, health & trustworthy AIVigo, 2026

Tutorials / T2

Fine-tuning Whisper with Optuna, cross-validation and MLflow

In one sentence

A complete, reproducible fine-tuning loop for Whisper on real speech, with grouped cross-validation, Optuna pruning, nested MLflow runs and a locked test set, ending in an explicit decision against a zero-shot baseline.

Why this tutorial

Most fine-tuning notebooks show one run and a number going down. This one builds the full loop a real project needs: hyperparameter search without leakage, traceable runs, and an honest check against the model you started with. It runs openai/whisper-tiny on real English speech, on a laptop CPU, in about nine minutes.

The data and the splits

The notebook uses hf-internal-testing/librispeech_asr_dummy: 73 utterances of read English from LibriSpeech (CC BY 4.0), one speaker, three book chapters. The chapters give the grouping.

Table 1. Which data each phase sees, and what it is allowed to decide.
PhaseDataDecision it allows
Baselinetest chapter (15 clips)Zero-shot WER of the untouched model, measured first
Hyperparameter search2 train chapters (58 clips), leave-one-chapter-outPick hyperparameters by mean validation loss
Refitall 58 train clipsFit the candidate model
Testtest chapter, scored onceKeep the fine-tuned model, or keep the base model

Folds are grouped, never random. With clinical data you group by patient, so the same voice is never on both sides of a split.

Practical detail: the notebook decodes FLAC with soundfile, avoiding the torchcodec/FFmpeg dependency of recent datasets.

The details that make it trustworthy

triallrfold 1 lossfold 2 lossresult T0 8e-6 2.0 1.8 mean 1.90 T1 7e-6 2.6 2.6 mean 2.60 T2 2e-5 1.5 1.5 ■ BEST 1.50 T3 3e-6 2.9 × skipped PRUNED T4 1.5e-5 1.7 1.6 mean 1.65 T5 4e-6 2.4 × skipped PRUNED
Figure 1. How the search spends its budget, animated with illustrative losses. Each trial trains one fresh model per fold. After fold 1, Optuna's median pruner stops any trial worse than the median of the completed ones, so fold 2 is never paid for. The lowest mean loss wins.
  • A fresh model per fold, so no fold contaminates the next.
  • Loss to select, WER to report: loss needs no text generation inside the loop.
  • Pruning per fold, reported with two lines:
trial.report(partial_mean_loss, step=fold_index + 1)
if trial.should_prune():
    mlflow.set_tag('optuna.state', 'pruned')
    raise optuna.TrialPruned()
  • A fair WER: Whisper’s English normalizer makes “Mr.” and “MISTER” the same word.
  • Readable runs: MLflow nests tutorial → hpo → trial → fold, plus refit; the decision is logged on the parent.
  • Nothing sensitive logged: parameters and aggregate metrics only. No audio, transcripts or weights.

Results

Table 2. Search and refit, 3 trials × 2 chapter folds, CPU. Lower is better.
RunLearning rateEpochsFold lossesMean CV lossTest WER
Zero-shot baseline––––9.7%
Trial 08.5 × 10−622.052, 1.7441.898–
Trial 17.1 × 10−612.559, 2.5922.575–
Trial 2 (selected)2.1 × 10−521.501, 1.4951.498–
Refit on all train2.1 × 10−52––11.5%
051015 Zero-shot 9.7 Fine-tuned 11.5
Figure 2. Word error rate (%) on the locked test chapter, after Whisper's English text normalizer. Solid bar: the model the pipeline keeps.

The search worked as intended: it found the best of the three configurations by validation loss. But the fine-tuned model is worse on the locked test chapter than the model it started from, so the pipeline logs the decision keep-base.

That is the lesson, not a bug. Whisper was pre-trained on 680,000 hours of speech; six minutes from one speaker cannot improve on that with full fine-tuning. The pipeline exists to catch exactly this.

Moving to real projects

  • Swap the data loader, keep the rest. Group folds by speaker or patient.
  • Lock test before the search and score it once. Never choose parameters or epochs from it.
  • Always log a zero-shot or previous-model baseline next to the candidate, and make the keep-or-replace decision explicit.
  • With little data, prefer parameter-efficient adaptation: see PEFT and LoRA on Whisper.
  • Log dataset and checkpoint revisions, manifest hashes and the resolved configuration; never audio, transcripts or health information.

References: Panayotov et al., LibriSpeech, ICASSP 2015; Optuna; MLflow Tracking; Whisper in Transformers.