SMILE: Speech Meta In-Context Learning for Low-Resource Language Automatic Speech Recognition
Journal
ASRU 2025 - 2025 IEEE Automatic Speech Recognition and Understanding Workshop
Start Page
1
End Page
7
ISBN (of the container)
979-833154426-3
ISBN
[9798331544263]
Date Issued
2025-12-06
Author(s)
Hsu, Ming-Hao
Abstract
Automatic Speech Recognition (ASR) models demonstrate outstanding performance on high-resource languages but face significant challenges when applied to low-resource languages due to limited training data and insufficient cross-lingual generalization. Existing adaptation strategies, such as shallow fusion, data augmentation, and direct fine-tuning, either rely on external resources, suffer computational inefficiencies, or fail in test-time adaptation scenarios. To address these limitations, we introduce Speech Meta In-Context LEarning (SMILE), an innovative framework that combines meta-learning with speech in-context learning (SICL). SMILE leverages metatraining from high-resource languages to enable robust, few-shot generalization to low-resource languages without explicit fine-tuning on the target domain. Extensive experiments on the ML-SUPERB benchmark show that SMILE consistently outperforms baseline methods, significantly reducing character and word error rates in training-free few-shot multilingual ASR tasks.
Event(s)
2025 IEEE Automatic Speech Recognition and Understanding Workshop, ASRU 2025
Subjects
Automatic Speech Recognition
In-Context Learning
Inference-time Adaptation
Low-Resource Language
Meta Learning
Publisher
Institute of Electrical and Electronics Engineers (IEEE)
Type
conference paper
