Similarity-based Accent Recognition with Continuous and Discrete Self-supervised Speech Representations
Part Of
ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing
Start Page
1
End Page
5
ISBN
[9798350368741]
Date Issued
2025-04-06
Author(s)
Abstract
The primary challenge in accent recognition lies in data scarcity due to the high diversity of accents, which make the collection of large-scale training data for each accent almost impossible in practice. To overcome this challenge, we propose a simple solution that leverages both continuous and discrete feature representations from pretrained speech self-supervised learning (SSL) models. Our model is simplified to a linear projection layer and a set of trainable accent class embeddings. Cosine similarity between the accent embeddings and the latent features of an audio sample is used to predict its accent class. This approach enables the model to access features that contain rich accent-related information while reducing the risk of model overfitting. Our method provides a practical and efficient way to tackle accent recognition, especially in low-resource scenarios. Experimental results on English accent recognition show that our best model achieves an accuracy of 84.0% on the AESRC 2020 dataset and an Unweighted Average Recall (UAR) of 50.0% on the VCTK corpus, setting new state-of-the-art results on both datasets.
Event(s)
2025 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2025
Publisher
IEEE
Type
conference paper
