Minimum phone error model training on merged acoustic units for transcribing bilingual code-switched speech
Journal
2012 8th International Symposium on Chinese Spoken Language Processing, ISCSLP 2012
Pages
320-324
Date Issued
2012
Author(s)
Abstract
This paper proposes to perform Minimum Phone Error (MPE) model training on merged acoustic units for transcribing Mandarin-English code-switched lectures with highly imbalanced language distribution. Some of the acoustic events in Mandarin and English may have very similar characteristics, so the states or Gaussian mixtures representing them can be merged with identical shared parameters. When MPE is performed afterwards, these merged identical states or Gaussian mixtures can form a compact acoustic unit set. In this way MPE can better discriminate the acoustic units of both languages, because similar units are merged while distinct units are differentiated. Significant improvements in recognition accuracy were observed in the preliminary experiments on real-world bilingual code-switched lecture corpus recorded at National Taiwan University.
SDGs
Type
conference paper
