SpeechDPR: End-To-End Spoken Passage Retrieval For Open-Domain Spoken Question Answering
Part Of
ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing
Start Page
12476
End Page
12480
ISSN
15206149
Date Issued
2024-04-14
Author(s)
Chyi-Jiunn Lin
Guan-Ting Lin
Yung-Sung Chuang
Wei-Lun Wu
Shang-Wen Li
Abdelrahman Mohamed
DOI
10.1109/ICASSP48485.2024.10448210
Abstract
Spoken Question Answering (SQA) is essential for machines to reply to user’s question by finding the answer span within a given spoken passage. SQA has been previously achieved without ASR to avoid recognition errors and Out-of-Vocabulary (OOV) problems. However, the real-world problem of Open-domain SQA (openSQA), in which the machine needs to first retrieve passages that possibly contain the answer from a spoken archive in addition, was never considered. This paper proposes the first known end-to-end frame-work, Speech Dense Passage Retriever (SpeechDPR), for the retrieval component of the openSQA problem. SpeechDPR learns a sentence-level semantic representation by distilling knowledge from the cascading model of unsupervised ASR (UASR) and text dense retriever (TDR). No manually transcribed speech data is needed. Initial experiments showed performance comparable to the cascading model of UASR and TDR, and significantly better when UASR was poor, verifying this approach is more robust to speech recognition errors.
Event(s)
2024 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2024
Publisher
IEEE
Type
conference paper
