HighRateMOS: Sampling-Rate Aware Modeling for Speech Quality Assessment
Journal
ASRU 2025 - 2025 IEEE Automatic Speech Recognition and Understanding Workshop
Start Page
1
End Page
4
ISBN (of the container)
979-833154426-3
ISBN
[9798331544263]
Date Issued
2025-12-06
Author(s)
Ren, Wenze
Lin, Yi-Cheng
Huang, Wen-Chin
Zezario, Ryandhimas E.
Fu, Szu-Wei
Huang, Sung-Feng
Cooper, Erica
Wu, Haibin
Wei, Hung-Yu
Wang, Hsin-Min
Tsao, Yu
Abstract
Modern speech quality prediction models are trained on audio data resampled to a specific sampling rate. When tested on audio with a higher sampling rate, these models can produce biased scores. We present HighRateMOS, the first non-intrusive mean opinion score (MOS) model that explicitly considers sampling rate. HighRateMOS ensembles three model variants that exploit the following information: (i) a learnable embedding of speech sampling rate, (ii) Wav2vec 2.0 selfsupervised embeddings, (iii) multi-scale CNN spectral features, and (iv) MFCC features. In AudioMOS 2025 Track 3, HighRateMOS ranked first in five of eight metrics. Our experiments confirm that modeling sampling rate leads to more robust and sampling-rate-agnostic speech quality predictions.
Event(s)
2025 IEEE Automatic Speech Recognition and Understanding Workshop, ASRU 2025
Subjects
MOS
Sampling rate
Speech quality assessment
Publisher
Institute of Electrical and Electronics Engineers(IEEE)
Type
conference paper
