All projects were conducted in collaboration with Pr. Najim Dehak, Dr. Jesús Antonio Villalba López, and Dr. Laureano Moro-Velázquez.
Spoken conversation analysis, speaker characterization, and emotional state modeling using audio foundation models and multimodal large language models.
Related publications
- Scaling Multi-Talker ASR with Speaker-Agnostic Activity Streams — Xiluo He et al., IEEE ICASSP, 2026
- Speaker Verification with Speech-Aware LLMs: Evaluation and Augmentation — Thomas Thebaud et al., Interspeech, 2026
- Beyond Transcripts: Iterative Peer-Editing with Audio Unlocks High-Quality Human Summaries of Conversational Speech — Kaavya Chaparala et al., Arxiv Preprint, 2026
- Demographic Attributes Prediction from Speech Using WavLM Embeddings — Yuchen Yang et al., CISS, 2025
- Enhancing Dialogue Annotation with Speaker Characteristics Leveraging a Frozen LLM — Thomas Thebaud et al., IEEE ASRU, 2025
- Paired by the Teacher: Turning Unpaired Data into High-Fidelity Pairs for Low-Resource Text Generation — Yen-Ju Lu et al., EMNLP, 2025
- Spoken DialogSum: An Emotion-Rich Conversational Dataset for Spoken Dialogue Summarization — Yen-Ju Lu et al., IEEE ASRU, 2025
Evaluation of speech-to-speech large language models in conversational settings.
Related publications
- StanceBench: A Benchmark for Audio LLM-Based Interpersonal Stance Evaluation from Speech — Yuzhe Wang et al., Interspeech, 2026
- TRACE: Temporal Relationship-Aware Conversational Entrainment Detection in Dyadic Speech — Sathvik Manikantan Napa Ugandhar et al., Arxiv Preprint, 2026
- Reference-Based Prosody and Rhythm Evaluation for Spoken Dialogue Systems — Ashish Hallur et al., Arxiv Preprint, 2026
- TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue — Hao Zhang et al., Arxiv Preprint, 2026
- SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models — Thomas Thebaud et al., Arxiv Preprint, 2026
Voice conversion for speech anonymization.
With Pr. Dani Byrd and Pr. Shrikanth Narayanan.
Related publications
- MaskVCT: Masked Voice Codec Transformer for Zero-Shot Voice Conversion with Increased Controllability via Multiple Guidances — Junhyeok Lee et al., IEEE ICASSP, 2026
- CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech — Helin Wang et al., IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2026
- Reconstruct! Don't Encode: Self-Supervised Representation Reconstruction Loss for High-Intelligibility and Low-Latency Streaming Neural Audio Codec — Junhyeok Lee et al., Arxiv Preprint, 2026
- Can LLMs Help Localize Fake Words in Partially Fake Speech? — Lin Zhang et al., Arxiv Preprint, 2026
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions — Thomas Thebaud et al., Arxiv Preprint, 2026
- Vox-Profile: A Speech Foundation Model Benchmark for Characterizing Diverse Speaker and Speech Traits — Tiantian Feng et al., Arxiv Preprint, 2025
- Rhythm Features for Speaker Identification — Nick Mehlman et al., Arxiv Preprint, 2025
- PPX-Anon: Prosody, Pitch and X-Vectors for De-Anonymization; Our Submission to the Voice Attacker Challenge 2024 — Thomas Thebaud et al., SPSC, 2025
- The SHADOW Team Submission to the ASVspoof 2024 Challenge — Jesus Antonio Villalba et al., ASVspoof Workshop, 2024
NeuroLogical signals for neurodegenerative disease detection and severity assessment using speech, eye movements, handwriting, and gait.
With Dr. Ankur Butala.
Related publications
- Detecting Neurodegenerative Diseases Using Frame-Level Handwriting Embeddings — Sarah Laouedj et al., IEEE ICASSP, 2025
- Interpretable Features for the Assessment of Neurodegenerative Diseases through Handwriting Analysis — Thomas Thebaud et al., IEEE Journal of Biomedical and Health Informatics, 2025
- Advancing Alzheimer's Disease Detection via Multimodal Fusion of Speech and Eye Movement Data — Anna Favaro et al., IEEE EMBC, 2025
- Sampling Rate Guidelines for Eye-Tracking in Neurological Disorders: A Task-Specific Multimetric Analysis — Yuzhe Wang et al., IEEE EMBC, 2025
- Multimodal Characterization of Alzheimer's Disease Using Speech, Eye Movement, and Handwriting — Laureano Moro-Velazquez et al., Alzheimer's & Dementia, 2024
- Cognitive Assessment through Writing Tasks — Casey Chen et al., Alzheimer's & Dementia, 2024
- Analyzing Attention Focus in the Cookie Theft Picture Description Task Using Word Alignment — Anna Favaro et al., Alzheimer's & Dementia, 2024
- Dynamics of Handwriting for Cognitive Assessment — Gabrielle Chavez et al., Alzheimer's & Dementia, 2024
- Consistent Explainable Features for Alzheimer's Disease Assessment through Handwriting — Thomas Thebaud et al., Alzheimer's & Dementia, 2024
- Discovering Invariant Patterns of Cognitive Decline via an Automated Analysis of the Cookie Thief Picture Description Task — Anna Favaro et al., Odyssey, 2024
- Exploring the Complementary Nature of Speech and Eye Movements for Profiling Neurological Disorders — Yuzhe Wang et al., Interspeech, 2024
- Unveiling Early Signs of Parkinson's Disease via a Longitudinal Analysis of Celebrity Speech Recordings — Anna Favaro et al., npj Parkinson's Disease, 2024
- Technology and Dementia Preconference — Laureano Moro-Velazquez et al., Alzheimer's & Dementia, 2024
- Interpretable Speech Features vs. DNN Embeddings: What to Use in the Automatic Assessment of Parkinson's Disease in Multi-Lingual Scenarios — Anna Favaro et al., Computers in Biology and Medicine, 2023
- Do Phonatory Features Display Robustness to Characterize Parkinsonian Speech across Corpora? — Anna Favaro et al., Interspeech, 2023
- Evaluation of Interpretable Speech Biomarkers for Monitoring Alzheimer's Disease and Mild Cognitive Impairment Progression — Anna Favaro et al., Alzheimer's & Dementia, 2023
- Handwriting Characteristics Analysis for Alzheimer's Disease and Mild Cognitive Impairments Assessment — Thomas Thebaud et al., Alzheimer's & Dementia, 2023
Creation of portable equipment to record numerical biomarkers for neurodegenerative diseases from speech, eye movements, and handwriting.
With Dr. Ankur Butala.
Related publications
Research on unsupervised source separation in noisy and far-field conditions.
Related publications
- SoloSpeech: Enhancing Intelligibility and Quality in Target Speech Extraction through a Cascaded Generative Pipeline — Helin Wang et al., IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2026
- DiT-Flow: Speech Enhancement Robust to Multiple Distortions Based on Flow Matching in Latent Space and Diffusion Transformers — Tianyu Cao et al., Arxiv Preprint, 2026
- ReFESS-QI: Reference-Free Evaluation for Speech Separation with Joint Quality and Intelligibility Scoring — Ari Frummer et al., Arxiv Preprint, 2025
- Noise-Robust Speech Separation with Fast Generative Correction — Helin Wang et al., Interspeech, 2024
Improving the robustness of image, speech, and video processing algorithms against adversarial and poisoning attacks.
With Pr. Sanjeev Khudanpur.
Related publications
Reverse engineering adversarial attacks against speech processing systems.
Related publications
Study of biases in speech processing algorithms, measurement of these biases, and development of balancing solutions.
With Pr. Zsuzsanna Fagyal and Pr. Mark Hasegawa-Johnson.
Related publications
- FaiST: A Benchmark Dataset for Fairness in Speech Technology — Maliha Jahan et al., ISCA, 2025
- Unveiling Performance Bias in ASR Systems: A Study on Gender, Age, Accent, and More — Maliha Jahan et al., IEEE ICASSP, 2025
- Finding Spoken Identifications: Using GPT-4 Annotation for an Efficient and Fast Dataset Creation Pipeline — Maliha Jahan et al., LREC-COLING, 2024
- Model-Based Fairness Metric for Speaker Verification — Maliha Jahan et al., IEEE ASRU, 2023
Research collaboration with École de technologie supérieure (ÉTS) in Montréal, Canada.
With Pr. Patrick Cardinal.
Related publications
Research collaboration with AGH University of Science and Technology in Kraków, Poland.
With Pr. Konrad Kowalczyk.
Related publications
Research collaboration with the University of Texas at Dallas, USA.
With Dr. Berrak Sisman and Pr. Carlos Busso.
Related publications
Research collaboration with the Università Politecnica delle Marche, Italy.
With Dr. Lucia Migliorelli, Dr. Leonardo Gabrielli, and Pr. Stefano Squartini.
Related publications