MMMOS: Multi-domain Multi-axis Audio Quality Assessment
Journal
ASRU 2025 - 2025 IEEE Automatic Speech Recognition and Understanding Workshop
Start Page
1
End Page
4
ISBN (of the container)
979-833154426-3
ISBN
[9798331544263]
Date Issued
2025-12-06
Author(s)
Abstract
Accurate audio quality estimation is essential for developing and evaluating audio generation, retrieval, and enhancement systems. Existing non-intrusive assessment models predict a single Mean Opinion Score (MOS) for speech, merging diverse perceptual factors and failing to generalize beyond speech. We propose MMMOS, a no-reference, multidomain audio quality assessment system that estimates four orthogonal axes: Production Quality, Production Complexity, Content Enjoyment, and Content Usefulness across speech, music, and environmental sounds. MMMOS fuses frame-level embeddings from three pretrained encoders (WavLM, MuQ, and M2D) and evaluates three aggregation strategies with four loss functions. By ensembling the top eight models, MMMOS shows a 2 0 - 3 0 % reduction in mean squared error and a 4-5% increase in Kendall's τ versus baseline, gains first place in six of eight Production Complexity metrics, and ranks among the top three on 17 of 32 challenge metrics.
Event(s)
2025 IEEE Automatic Speech Recognition and Understanding Workshop, ASRU 2025
Subjects
Audio Quality assessment
Mean opinion score (MOS)
Publisher
Institute of Electrical and Electronics Engineers(IEEE)
Type
conference paper
