Ir directamente a la navegación principal Ir directamente a la búsqueda Ir directamente al contenido principal

Fine-Tuning Wav2Vec2 for Low-Resource Kichwa Automatic Speech Recognition

  • Universidad San Francisco de Quito
  • Carnegie Mellon University

Producción científica: Capítulo del libro/informe/acta de congresoContribución a la conferenciarevisión exhaustiva

Resumen

In recent years, advancements in artificial intelligence (AI) have significantly accelerated the development of natural language processing and automatic speech recognition (ASR) systems for high-resource languages, raising concerns about the marginalization of ancestral and underrepresented languages. In this context, this work explores the fine-tuning of the Wav2Vec 2.0 model, developed by Meta AI, for ASR in Kichwa-a low-resource language spoken in the Ecuadorian Andes. The training process utilized two datasets totaling approximately 8 hours of audio, segmented into clips ranging from 1.5 to 5 seconds, with manually aligned transcriptions created using ELAN software. Fine-tuning was performed using the Connectionist Temporal Classification (CTC) loss function. After multiple experiments, a two-tailed Wilcoxon signed-rank test revealed no statistically significant improvement when applying SpecAugment. The best-performing model, trained without data augmentation, achieved promising results on the test set: a Word Error Rate (WER) of 0.262, a Character Error Rate (CER) of 0.120, and a Match Error Rate (MER) of 0.401. These findings d emonstrate the viability of adapting pre-trained self-supervised models to low-resource settings and underscore the potential of ASR technologies to support greater linguistic inclusivity in artificial intelligence.

Idioma originalInglés
Título de la publicación alojadaETCM 2025 - 9th Ecuador Technical Chapters Meeting
EditorialInstitute of Electrical and Electronics Engineers Inc.
ISBN (versión digital)9798331552640
DOI
EstadoPublicada - 2025
Evento9th Ecuador Technical Chapters Meeting, ETCM 2025 - Quito, Ecuador
Duración: 21 oct 202524 oct 2025

Serie de la publicación

NombreETCM 2025 - 9th Ecuador Technical Chapters Meeting

Conferencia

Conferencia9th Ecuador Technical Chapters Meeting, ETCM 2025
País/TerritorioEcuador
CiudadQuito
Período21/10/2524/10/25

Huella

Profundice en los temas de investigación de 'Fine-Tuning Wav2Vec2 for Low-Resource Kichwa Automatic Speech Recognition'. En conjunto forman una huella única.

Citar esto