Supervised Spoken Document Summarization Based on Structured Support Vector Machine with Utterance Clusters as Hidden Variables
Journal
Proceedings of the Annual Conference of the International Speech Communication Association, INTERSPEECH
Pages
2728-2732
ISSN
2308457X
Date Issued
2013-08
Author(s)
Abstract
This paper presents a supervised approach for extractive summarization of spoken document considering utterance clusters in the documents as hidden variables. Utterances in important clusters may be jointly included in the summary, while those in less important clusters may be excluded as a whole. The summaries are therefore selected based on not only the conventional principle of including the most important utterances and minimizing the redundancy but also the hidden cluster structure in the document. The cluster structure of the documents is not known but can be inferred from the documents, and the summaries can be jointly obtained by the structured SVM learned from the training examples. Encouraging results were obtained on a lecture corpus in the preliminary experiments.
Event(s)
14th Annual Conference of the International Speech Communication Association, INTERSPEECH 2013
Subjects
Hidden variables
Speech summarization
Structured SVM
Type
conference paper
