Special speech recognition approaches for the highly confusing Mandarin syllables based on hidden Markov models
Journal
Computer Speech and Language
Journal Volume
5
Journal Issue
2
Pages
181-201
Date Issued
1991
Author(s)
Abstract
In this paper, several special speech recognition approaches based on hidden Markov models (HMMs) are presented for the highly confusing Mandarin syllables by considering the characteristics of the vocabulary. This is because there are totally 408 syllables (disregarding the tones) in Mandarin speech, and it is believed that correct recognition of these syllables is the key to the development of a Mandarin dictation machine which recognizes Mandarin speech with very large vocabulary and unlimited texts. However, accurate recognition of these 408 syllables is very difficult because there exist 38 confusing sets among them, each of which has at most 19 very confusing syllables. Direct application of conventional standard approaches of HMMs to these syllables gives recognition rates in the order of only 70-80%, thus several special approaches are proposed in this paper to provide better performance. Some of these approaches concentrate on the training algorithms for the HMMs, including the two-pass training, the revised two-pass training, the three-pass training, and the revised three-pass training approaches. The basic idea is to protrude the very short initial parts (initial consonant parts) and de-emphasize the final parts (vowel or diphthong parts but including possible medial or nasal ending) of the syllables such that the confusing syllables can be better distinguished. Also, some other approaches concentrate on the recognition phase of the HMMs, including the state duration bounds and the two-stage search strategy, which can also improve the recognition rates and/or speeds. A special hardware is also implemented to complete all the recognition operations in real-time, on which all the approaches discussed in this paper can be applied. © 1991.
SDGs
Type
journal article
