VoiceNoNG: Robust High-Quality Speech Editing Model without Hallucinations
Journal
Interspeech 2025
Series/Report No.
Proceedings of the Annual Conference of the International Speech Communication Association Interspeech
Start Page
3469
End Page
3473
ISSN
2308457X
Date Issued
2025-08-17
Author(s)
Huang, Sung-Feng
Kuo, Heng-Cheng
Chen, Zhehuai
Yang, Xuesong
Ku, Pin-Jui
Jukic, Ante
Yang, Huck
Tsao, Yu
Fu, Szu-Wei
Abstract
Voicebox and VoiceCraft are the current most representative models for non-autoregressive and autoregressive speech editing, respectively. Although both of them can generate high-quality speech edits, we identify their limitations: Voicebox is not good at editing speech with background audio, while VoiceCraft suffers from the hallucination-like problem. To maintain speech quality for varying audio scenarios and address the hallucination issue, we introduce VoiceNoNG, which combines the strengths of both model frameworks. VoiceNoNG utilizes a latent flow-matching framework to model the pre-quantization features of a neural codec. The vector quantizer in the neural codec provides additional robustness against minor prediction errors from the editing model, which enables VoiceNoNG to achieve state-of-the-art performance in both objective and subjective evaluations under diverse audio conditions. Audio examples and ablations can be found on the demo page and supplementary material. © 2025 International Speech Communication Association.
Event(s)
26th Interspeech Conference 2025
Subjects
flow-matching model
neural speech editing
speech infilling
Publisher
ISCA
Type
conference paper
