Skip to main navigation Skip to search Skip to main content

Improving Dysarthria Assessment Through Voice Conversion in Low-Resource Settings

  • Emily Chimbo*
  • , Felipe Grijalva
  • , Karen Rosero
  • *Corresponding author for this work
  • Universidad San Francisco de Quito
  • Carnegie Mellon University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

The development of automatic dysarthria assessment systems is often limited by the scarcity of labeled pathological speech, particularly in low-resource clinical environments. In this work, we explore the use of generative voice conversion to address this bottleneck. We propose a pipeline based on speech enhancement and voice conversion to transfer dysarthric vocal traits onto healthy utterances to generate realistic pathological speech. A total of 4,082 synthetic samples were generated and enhanced to improve quality while preserving pathological prosody. A dual-branch CNN trained under four experimental setups showed that combining real and synthetic data improved the classification accuracy from 77.52% to 98.36%, while using only enhanced synthetic data reached 97.10%. These results support the use of voice conversion to expand clinical datasets and reduce dependence on real patient recordings.

Original languageEnglish
Title of host publicationETCM 2025 - 9th Ecuador Technical Chapters Meeting
PublisherInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronic)9798331552640
DOIs
StatePublished - 2025
Event9th Ecuador Technical Chapters Meeting, ETCM 2025 - Quito, Ecuador
Duration: 21 Oct 202524 Oct 2025

Publication series

NameETCM 2025 - 9th Ecuador Technical Chapters Meeting

Conference

Conference9th Ecuador Technical Chapters Meeting, ETCM 2025
Country/TerritoryEcuador
CityQuito
Period21/10/2524/10/25

Keywords

  • Dysarthria
  • Generative AI
  • Smart Healthcare
  • Speech Pathology
  • Voice Conversion

Fingerprint

Dive into the research topics of 'Improving Dysarthria Assessment Through Voice Conversion in Low-Resource Settings'. Together they form a unique fingerprint.

Cite this