Working papers · speech, health & trustworthy AIVigo, 2026

Researcher · ML engineer_

What the voice reveals, and how to keep AI that listens safe

José Manuel Ramírez Sánchez1

1Multimedia Technologies Group (GTM), atlanTTic Research Center, Universidade de Vigo

0000-0003-4700-6592 josemspeechtech

Plate I. The author, Atkinson 1-bit dither.

Abstract

I am a speech and language technology researcher finishing my PhD at the Universidade de Vigo, and an engineer who ships. I build models that read health from the voice: Long COVID, respiratory disease, emotional distress and suicide risk. In industry, I delivered the voice classifier for the COPERIA screening platform at Bahía Software, and I now advise Balidea on agentic AI systems for the Galician Health Service (SERGAS). My focus now is making generative AI that talks with vulnerable people aligned, robust and hard to misuse.

Keywords — speech biomarkers; multimodal learning; self-supervised speech representations; AI safety; alignment; safeguards for generative AI; mental-health AI.

§1 Research lines

Table 1. Lines of work, the modalities involved and the papers behind each.
LineModalityRefs.
Voice as a clinical signal [1][4][6]
Mental-health AI and conversational safeguards [2][3][5]
Robust speech recognition and spoken search [9][10][11][12][13][14]
Language technology for under-resourced languages [7][8]
Alignment and safeguards for generative AI (current) —

§2 Selected papers

  1. [1]

    A short walk makes voice-based screening more reliable: coughs and vowels recorded after exercise reveal post-COVID sequelae better than recordings at rest.

    Vera-López Á. et al. · Frontiers in Medicine · 2026

    Journal Read summary DOI
    Full citation

    Vera-López Á., Tilves-Santiago D., Ramírez-Sánchez J.M., Docío-Fernández L., García-Mateo C., Bustillo-Casado M., García-Caballero A.A. Improving respiratory disease detection through SSL-enhanced acoustic analysis and exercise-rest measurements. Frontiers in Medicine, 2026.

  2. [3]

    A chat agent designed to spot suicide risk factors during a live conversation and to respond safely when it finds them.

    Ramírez Sánchez J.M. et al. · IWSDS · 2025

    Conference Link
    Full citation

    Ramírez Sánchez J.M., Manso Vázquez M., García-Mateo C., Docío-Fernández L., Fernández-Iglesias M.J., Gómez-Gómez B., Pinal B., Brañas A., García-Caballero A. Design of a conversational agent to support people on suicide risk. IWSDS 2025 (International Workshop on Spoken Dialogue Systems Technology), 2025.

  3. [6]

    The first study to test whether voice recordings alone can tell Long COVID patients apart from healthy people. Coughs after exertion worked best.

    Ramírez Sánchez J.M. et al. · JMIR Preprints · 2023

    Preprint Read summary DOI
    Full citation

    Ramírez Sánchez J.M., Docío-Fernández L., García Mateo C., Bustillo-Casado M., García-Caballero A.A. Identifying patients with Long COVID: study of the effectiveness of voice signal analysis and machine learning. JMIR Preprints, 2023.

  4. [13]

    MFCC or PLP? In quiet rooms it barely matters; in noise, MFCC keeps speech recognition working better, while PLP decodes faster. My first paper.

    Ramírez Sánchez J.M. et al. · Revista Ingeniería Electrónica · 2019

    Journal Read summary Link
    Full citation

    Ramírez Sánchez J.M., Montalvo Bereau A.R., Calvo de Lara J.R. Evaluación de Rasgos Acústicos para el Reconocimiento Automático del Habla en Escenarios Ruidosos usando Kaldi. Revista Ingeniería Electrónica, Automática y Comunicaciones (RIELAC) 40(3), 2019.

All papers (14) →

§3 Tutorials

MFCC and PLP from scratch: how machines first learned to listen

Build the two classic speech features step by step in NumPy, watch one frame of real speech travel through both pipelines, and see what noise does to them. The front end every recogniser used before deep learning, and the subject of my first paper.

NumPySciPylibrosaWhisper

Fine-tuning Whisper with Optuna, cross-validation and MLflow

A complete, reproducible fine-tuning loop for Whisper on real speech, with grouped cross-validation, Optuna pruning, nested MLflow runs and a locked test set, ending in an explicit decision against a zero-shot baseline.

WhisperTransformersOptunaMLflowLibriSpeech

§4 Writing