Build the two classic speech features step by step in NumPy, watch one frame of real speech travel through both pipelines, and see what noise does to them. The front end every recogniser used before deep learning, and the subject of my first paper.
A complete, reproducible fine-tuning loop for Whisper on real speech, with grouped cross-validation, Optuna pruning, nested MLflow runs and a locked test set, ending in an explicit decision against a zero-shot baseline.
Freeze Whisper, inject LoRA adapters in the right places, and train under 2% of the parameters, either for an end-to-end classifier or for embeddings you feed to your own classifier.
Open a safety guard on a laptop CPU: read its verdict as a probability, watch the decision form layer by layer, and see why a harmful-content guard lets prompt injection hidden in documents straight through.