A discriminative HMM/N-gram-based retrieval approach for mandarin spoken documents
Resource
ACM Transactions on Asian Language Information Processing 3 (2): 128-145
Journal
ACM Transactions on Asian Language Information Processing
Journal Volume
3
Journal Issue
2
Pages
128-145
Date Issued
2004
Date
2004
Author(s)
Abstract
In recent years, statistical modeling approaches have steadily gained in popularity in the field of information retrieval. This article presents an HMM/N-gram-based retrieval approach for Mandarin spoken documents. The underlying characteristics and the various structures of this approach were extensively investigated and analyzed. The retrieval capabilities were verified by tests with word- and syllable-level indexing features and comparisons to the conventional vector-space model approach. To further improve the discrimination capabilities of the HMMs, both the expectation-maximization (EM) and minimum classification error (MCE) training algorithms were introduced in training. Fusion of information via indexing word- and syllable-level features was also investigated. The spoken document retrieval experiments were performed on the Topic Detection and Tracking Corpora (TDT-2 and TDT-3). Very encouraging retrieval performance was obtained.
Subjects
Hidden Markov models; Mandarin spoken documents; Syllable-level Indexing features
SDGs
Other Subjects
Hidden Markov models; Mandarin spoken documents; Statistical modeling; Syllable-level indexing features; Error analysis; Information retrieval; Mathematical models; Pattern recognition; Probability; Problem solving; Speech recognition; Statistical methods; Indexing (of information)
Type
journal article
File(s)![Thumbnail Image]()
Loading...
Name
15.pdf
Size
292.66 KB
Format
Adobe PDF
Checksum
(MD5):0711d1dcf3d54249de097a1b74809858
