Integrating Gabor and Pitch Features in Tandem Systems for Mandarin Speech Recognition
Date Issued
2011
Date
2011
Author(s)
Li, Shang-Wen
Abstract
In conventional speech recognition, we use MFCC features to extract speech information in waveform. We further train statistic models with these features for decoding. However, MFCC features retain only the information within a short time span. Recently, many researches focus on extracting long-term information from speech signal or the variation in spectral, temporal or spectro-temporal modulation frequency, and these studies achieve significant performance improvement.
Here, we utilize Gabor filters to extract Gabor features, which are abundant in spectro-temporal information. An MLP is trained for learning the variation of Gabor features among different phonemes. The outputs of MLP are Gabor posteriors. We use Tandem system to integrate Gabor and MFCC posteriors and achieve better performance in our speech recognition system. Furthermore, we estimate posteriors more accurately by clustered hierarchical MLP, which emphasize on the classification of error-prone phoneme pairs. Thus, we obtain even better recognition performance. Finally, we add pitch features while MLP training and adopt tonal acoustic units. With these modifications, we significantly improve the performance in Mandarin large vocabulary broadcast news recognition.
Subjects
speech recognition
feature extraction
Tandem system
Type
thesis
File(s)![Thumbnail Image]()
Loading...
Name
ntu-100-R98942035-1.pdf
Size
23.32 KB
Format
Adobe PDF
Checksum
(MD5):87c131a1a2650385bf948ebb2db96b74
