DeSTA2.5-Audio: Toward General-Purpose Large Audio Language Model With Self-Generated Cross-Modal Alignment
Journal
IEEE Transactions on Audio, Speech and Language Processing
Journal Volume
34
Start Page
2062
End Page
2076
ISSN
29984173
Date Issued
2026
Author(s)
Lu, Ke-Han
Chen, Zhehuai
Fu, Szu-Wei
Yang, Chao-Han Huck
Huang, Sung-Feng
Yang, Chih-Kai
Yu, Chee-En
Chen, Wei-Chih
Huang, Chien-yu
Lin, Yu-Xiang
Fu, Chi-An
Kuan, Chun-Yi
Ren, Wenze
Chen, Xuanjun
Huang, Wei-Ping
Hu, En-Pei
Lin, Tzu-Quan
Wu, Yuan-Kuei
Huang, Kuan-Po
Huang, Hsiao-Ying
Chou, Huang-Cheng
Chang, Kai-Wei
Chiang, Cheng-Han
Ginsburg, Boris
Abstract
We introduce DeSTA2.5-Audio, a general-purpose Large Audio Language Model (LALM) designed for robust auditory perception and instruction-following. Recent LALMs augment Large Language Models (LLMs) with auditory capabilities by training on large-scale audio-instruction datasets. However, existing LALMs have often suffered from the catastrophic forgetting of the LLM's original abilities. Therefore, balancing knowledge retention and audio perception has become a critical challenge. To address this, we revisit the data construction pipeline and propose a self-generated cross-modal alignment strategy in which the backbone LLM generates its own training targets, named DeSTA. This approach aims at preserving the LLM's native language proficiency thereby enabling zero-shot generalization without task-specific tuning. We construct DeSTA-AQA5M, a large-scale, task-agnostic dataset containing 5 million training samples derived from 7,000 hours of audio spanning 50 diverse datasets, including speech, environmental sounds, and music. DeSTA2.5-Audio achieves state-of-the-art or competitive performance across a wide range of audio-language benchmarks, including Dynamic-SUPERB, MMAU, SAKURA, Speech-IFEval, and VoiceBench. Comprehensive comparative studies demonstrate that our self-generated strategy outperforms existing training strategies. Our findings underscore the importance of carefully designed data construction in LALM development and offer practical insights for building robust, general-purpose LALMs.
Subjects
Cross-modal alignment
dataset construction
instruction-tuning
large audio language model
Publisher
Institute of Electrical and Electronics Engineers Inc.
Type
journal article
