Toward a Unified Perspective on Parameter-Efficient Fine Tuning for Speaker Verification
Journal
IEEE Transactions on Audio, Speech and Language Processing
Journal Volume
34
Start Page
2276
End Page
2289
ISSN
2998-4173
Date Issued
2026-04-08
Author(s)
Abstract
Fine-tuning pre-trained Transformer models (PTMs) for speech tasks in a parameter-efficient fine-tuning (PEFT) manner can optimize memory usage while leveraging rich representations from large-scale unlabeled data. Although PEFT is effective, the interconnections between various PEFT methods are notfully understood. This paper analyzes state-of-the-art PEFT methods and introduces a unified frameworkto clarify their interrelationships. Specifically, we employ a dynamic prompts tuning strategy that selects optimal prompts from a predefined pool, ensuring each prompt is fine-tuned by its closely matched speaker. The goal is to cluster the prompts in the pool according to speaker traits, improving speaker predictionin the downstream classifier while preserving the flexibility of the pre-trained Transformers. Additionally, we integrate the mixture-of-experts (MoE) adapter into the Transformer encoders, enabling the fine-tunedPTM to select the most relevant output for extracting task-specific information. Furthermore, we improve existing PEFT techniques by incorporating spectral information from pre-trained weight matrices into the LoRA-based fine-tuning process. Extensive experiments on VoxCeleb, CN-Celeb, and CU-MARVEL demonstrate that the proposed method offers a memory- and computation-efficient solution for fine-tuning pre-trained Transformers.
Subjects
LoRA
MoE
parameter-efficient fine-tuning
pre-trained transformer model
prompt tuning
Speaker verification
Publisher
Institute of Electrical and Electronics Engineers (IEEE)
Type
journal article
