Robust Unsupervised Topic-Based Language Model Adaptation
Date Issued
2007
Date
2007
Author(s)
Heidel, Aaron
DOI
en-US
Abstract
We present a novel topic mixture-based language model adaptation approach that uses Latent Dirichlet Allocation (LDA). We use Probabilistic Latent Semantic Analysis (PLSA) to automatically cluster a heterogeneous training corpus, and then train an LDA model using the resultant topic-document assignments. Using this LDA model, we construct fine-grained topic-specific corpora at the utterance level, which we use to train topic language models. These topic LMs are interpolated with a background language model during language model adaptation under an N-best rescoring framework. We describe several techniques for hardening LDA topic inference to first-pass recognition errors, and demonstrate the effectiveness of metadata-based segmentation when combined with show-level language model adaptation. Good improvements over state-of-the-art schemes were obtained in experiments on multi-genre GALE Project data in Mandarin Chinese.
Subjects
語音辨識
語言模型調適
潛藏語意分析
片段分割
非監督式調適
speech recognition
language model adaptation, topic modeling,story segmentation, unsupervised adaptation
Type
thesis
File(s)![Thumbnail Image]()
Loading...
Name
ntu-96-R94922037-1.pdf
Size
23.31 KB
Format
Adobe PDF
Checksum
(MD5):f559bc21a0d9f3988d13cdf84fdff370
