# José Manuel Ramírez Sánchez > Speech and language technology researcher, PhD candidate at the Universidade de Vigo (atlanTTic Research Center, Multimedia Technologies Group, Spain). He builds models that read health from the voice (Long COVID, respiratory disease, emotional distress and suicide risk) and now focuses on alignment and safeguards for generative AI that interacts with vulnerable people. He combines academic research with production engineering: at Bahía Software he delivered the voice classifier for the COPERIA screening platform, and he is currently an external advisor to Balidea on agentic AI systems for the Galician Health Service (SERGAS). For the European Commission, he was the Galician language representative in the EU-funded European Language Equality (ELE) project (2021–2022) and author of its Language Report for Galician. Identity: ORCID 0000-0003-4700-6592. Publishes as "José Manuel Ramírez Sánchez", "J. M. Ramírez Sánchez" or "Jose M. Ramirez"; GitHub user JMasr. Earlier work (2017–2020) was done at CENATAV, Havana, Cuba. Papers in medicine or surgery signed "J. Ramírez" belong to other people with a similar name. Profiles: [ORCID](https://orcid.org/0000-0003-4700-6592) · [Google Scholar](https://scholar.google.com/citations?user=lcQpcR8AAAAJ) · [LinkedIn](https://www.linkedin.com/in/josemspeechtech/) · [GitHub](https://github.com/JMasr) · [OpenAlex](https://openalex.org/A5100602668) Doctoral thesis: "Clinical Screening from Voice via Parameter-Efficient Fine-Tuning of Audio Transformers", supervised by Carmen García-Mateo and Laura Docío-Fernández (Universidade de Vigo); expected defence December 2026. Open to: research roles in AI safety, alignment and health AI. ## Pages - [Home](https://jmramirez.engineer/): profile, research lines and selected papers - [Research](https://jmramirez.engineer/research/): all 14 publications with plain-language summaries - [Tutorials](https://jmramirez.engineer/tutorials/): runnable audio machine-learning notebooks - [Writing](https://jmramirez.engineer/writing/): essays and analyses - [About](https://jmramirez.engineer/about/): career timeline - [Full text for LLMs](https://jmramirez.engineer/llms-full.txt): every summary and write-up in one Markdown file ## Research lines - Voice as a clinical signal (3 papers) - Mental-health AI and conversational safeguards (3 papers) - Robust speech recognition and spoken search (6 papers) - Language technology for under-resourced languages (2 papers) - Alignment and safeguards for generative AI (current) ## Publications - [A six-minute walk makes the voice a better diagnostic signal](https://jmramirez.engineer/research/respiratory-ssl-exercise-2026/) (2026, Journal): A short walk makes voice-based screening more reliable: coughs and vowels recorded after exercise reveal post-COVID sequelae better than recordings at rest. - [VisIA-Q: A cross-sectional psychometric and demographic dataset of adolescents at high-risk for suicide](https://doi.org/10.5281/zenodo.20703908) (2026, Dataset): An open psychometric and demographic dataset of adolescents at high risk of suicide, released to support reproducible mental-health AI. - [Design of a conversational agent to support people on suicide risk](https://aclanthology.org/people/jose-manuel-ramirez-sanchez/) (2025, Conference): A chat agent designed to spot suicide risk factors during a live conversation and to respond safely when it finds them. - [Dataset-Driven Voice Biomarker Validation: Advancing Clinical Standards with the Enhanced V3 Framework](https://doi.org/10.17979/spudc.9788497498913.38) (2024, Conference): Voice biomarkers need clinical-grade validation: an extension of the V3 framework, focused on the data life cycle of clinical trials, and an open toolkit to check voice health datasets. - [VisIA Project: design of an automated AI-based emotional distress and suicide risk detection system](https://doi.org/10.21437/iberspeech.2024-56) (2024, Conference): The plan for VisIA: a multimodal system that detects emotional distress and a speech-based conversational agent built on clinical guidelines, for non-invasive suicide-risk assessment in young adults. - [Can the voice alone identify Long COVID patients?](https://jmramirez.engineer/research/long-covid-voice-2023/) (2023, Preprint): The first study to test whether voice recordings alone can tell Long COVID patients apart from healthy people. Coughs after exertion worked best. - [Language Report Galician](https://doi.org/10.1007/978-3-031-28819-7_17) (2023, Chapter): The state of speech and language technology for Galician, written as the Galician representative in the EU-funded European Language Equality project. - [Galician's Language Technologies in the Digital Age](https://doi.org/10.21437/iberspeech.2022-5) (2022, Conference): Where Galician stands in speech and language technology, and which resources are still missing. - [The Multi-Domain International Search on Speech 2020 ALBAYZIN Evaluation: Overview, Systems, Results, Discussion and Post-Evaluation Analyses](https://doi.org/10.3390/app11188519) (2021, Journal): Results of an international challenge on finding spoken words and phrases in audio from several domains. - [GELABERT: herramienta para la detección automática de términos hablados](https://jmramirez.engineer/research/) (2020, Conference): A tool that finds spoken terms in audio recordings. - [A Survey of the Effects of Data Augmentation for Automatic Speech Recognition Systems](https://doi.org/10.1007/978-3-030-33904-3_63) (2019, Conference): Which data augmentation techniques actually help speech recognisers. The author's most-cited paper. - [ALBAYZIN 2018 spoken term detection evaluation: a multi-domain international evaluation in Spanish](https://doi.org/10.1186/s13636-019-0159-7) (2019, Journal): Results of the ALBAYZIN 2018 international evaluation on detecting spoken terms in Spanish audio. - [MFCC or PLP? What happens to speech recognition when the room gets noisy](https://jmramirez.engineer/research/kaldi-noisy-features-2019/) (2019, Journal): MFCC or PLP? In quiet rooms it barely matters; in noise, MFCC keeps speech recognition working better, while PLP decodes faster. The author's first paper. - [Cenatav Voice Group System for Albayzin 2018 Search on Speech Evaluation](https://doi.org/10.21437/IberSPEECH.2018-53) (2018, Conference): CENATAV's system for the Albayzin 2018 search-on-speech challenge. ## Tutorials - [MFCC and PLP from scratch: how machines first learned to listen](https://jmramirez.engineer/tutorials/mfcc-plp-from-scratch/): Build the two classic speech features step by step in NumPy, watch one frame of real speech travel through both pipelines, and see what noise does to them. The front end every recogniser used before deep learning, and the subject of my first paper. Notebook: https://github.com/JMasr/ai-tutorials/blob/main/en/mfcc-plp-from-scratch.ipynb - [Fine-tuning Whisper with Optuna, cross-validation and MLflow](https://jmramirez.engineer/tutorials/whisper-optuna-mlflow/): A complete, reproducible fine-tuning loop for Whisper on real speech, with grouped cross-validation, Optuna pruning, nested MLflow runs and a locked test set, ending in an explicit decision against a zero-shot baseline. Notebook: https://github.com/JMasr/ai-tutorials/blob/main/en/whisper-optuna-mlflow.ipynb - [PEFT and LoRA on Whisper: a downstream classifier vs. an embedding extractor](https://jmramirez.engineer/tutorials/whisper-peft-lora/): Freeze Whisper, inject LoRA adapters in the right places, and train under 2% of the parameters, either for an end-to-end classifier or for embeddings you feed to your own classifier. Notebook: https://github.com/JMasr/ai-tutorials/blob/main/en/whisper-peft-lora.ipynb - [Inside an LLM guard: the verdict token, the logit lens and the injection blind spot](https://jmramirez.engineer/tutorials/llm-guard-internals/): Open a safety guard on a laptop CPU: read its verdict as a probability, watch the decision form layer by layer, and see why a harmful-content guard lets prompt injection hidden in documents straight through. Notebook: https://github.com/JMasr/ai-tutorials/blob/main/en/llm-guard-internals.ipynb ## Writing - [Inside four LLM guards: what the weights say about safeguards for agentic systems](https://jmramirez.engineer/writing/inside-four-llm-guards/) (2026-09-29): I audited four candidate guards for an agentic assistant layer by layer. They all decide in one token and late in the network, and the two dedicated guards miss prompt injection hidden in documents. A small generalist catches almost everything if it writes its justification before the verdict. - [How LLM guards decide: a tour inside four safety models](https://jmramirez.engineer/writing/how-llm-guards-decide/) (2026-09-29): Where the policy lives, which layers read it, what the router does and why the last block can overturn a verdict: a tour through the internals of four LLM guards, measured on the weights they actually run.