Towards Generalized Source Tracing for Codec-Based Deepfake Speech
Journal
ASRU 2025 - 2025 IEEE Automatic Speech Recognition and Understanding Workshop
Start Page
1
End Page
8
ISBN (of the container)
979-833154426-3
ISBN
[9798331544263]
Date Issued
2025-12-06
Author(s)
Abstract
Recent attempts at source tracing for codecbased deepfake speech (CodecFake), generated by neural audio codec-based speech generation (CoSG) models, have exhibited suboptimal performance. However, how to train source tracing models using simulated CoSG data while maintaining strong performance on real CoSG-generated audio remains an open challenge. In this paper, we show that models trained solely on codec-resynthesized data tend to overfit to non-speech regions and struggle to generalize to unseen content. To mitigate these challenges, we introduce the Semantic-Acoustic Source Tracing Network (SASTNet), which jointly leverages Whisper for semantic feature encoding and Wav2vec2 with AudioMAE for acoustic feature encoding. Our proposed SASTNet achieves state-of-theart performance on the CoSG test set of CodecFake+ dataset, demonstrating its effectiveness for reliable source tracing.
Event(s)
2025 IEEE Automatic Speech Recognition and Understanding Workshop, ASRU 2025
Subjects
Anti-spoofing
audio deepfake detection
explainability
neural audio codec
source tracing
Publisher
Institute of Electrical and Electronics Engineers (IEEE)
Type
conference paper
