Improving Contextual Biasing in Chinese ASR via Multimodal Large Language Models
Journal
IEEE Access
Journal Volume
14
Start Page
29706
End Page
29728
ISSN
2169-3536
Date Issued
2026-02-23
Author(s)
Abstract
This study addresses the challenge in Chinese automatic speech recognition (ASR) systems of accurately recognizing proper nouns such as place names, personal names, song titles, and movie or TV show titles, which is often hindered by tonal features and abundant homophones. We propose an integrated contextual biasing framework centered on multimodal large language models (MLLMs) to enhance the system’s context awareness and task adaptability. The core of this framework is an intent-driven dynamic contextual biasing mechanism: first, a fine-tuned MLLM performs end-to-end intent recognition, achieving an 81.82% relative error rate reduction compared to the unfine-tuned model and a 66.71% reduction relative to a cascaded model; subsequently, based on the highly accurate intent predictions, context-relevant keyword prompts are dynamically generated to guide speech recognition. Models fine-tuned using this strategy demonstrate significant improvements in both character error rate (CER) and keyword error rate (KER), with a 41.48% relative error reduction in KER. To address the cold-start problem, we also develop an automated data generation pipeline that requires only a domain-specific list of proper nouns to generate natural sentences using a small language model, followed by speech synthesis to produce training audio. Experiments show that models fine-tuned with synthetic data achieve a 41.91% relative error reduction in keyword recognition, nearly matching the performance of models trained on real annotated data. Overall, this work provides an innovative framework for contextual biasing in Chinese ASR and demonstrates, through open-source code and evaluation standards, the potential of multimodal large language models in integrating speech understanding and recognition tasks.
Subjects
automated data generation
automatic speech recognition
Chinese ASR
contextual biasing
intent recognition
multimodal large language models
Publisher
Institute of Electrical and Electronics Engineers (IEEE)
Type
journal article
