PTT Corpus: Construction and Applications
Date Issued
2014
Date
2014
Author(s)
Liu, Tsun-Jui
Abstract
In recent years, corpus-based and corpus-driven studies are getting considerable attentions. In Taiwan Mandarin, two of the most widely used corpora are Academia Sinica Balanced Corpus (Chen et al., 1996) and Chinese Gigawords (Huang et al., 2005). However, both of the corpora have some limitations on the source of the data, and they have not updated for some time, which makes it difficult to collect more recent examples of language uses. Therefore, the aim of this thesis attempts to establish a dynamic corpus, PTT Corpus, which can automatically collect, update and process data from PTT (批踢 踢), and provide the applications with a user-friendly interface for researchers. Corpora are segmented with Jseg, a Chinese segmentator trained with data from Sinica Corpus, and part-of-speech (POS) tagged by Brill Tagger (Brill, 1992), a POS tagger trained with data trained on the 9999 sentences in the Sinica Treebank (Chen et al., 1999). PTT Corpus provides a web interface with several applications, including Concordancer, Collocation extractor, Emoticon Detector, etc. To conclude, establishing PTT Corpus may be of importance in enriching the source of modern corpora, providing useful corpus tools, simplifying the analysis of recent language uses and changes in linguists in Taiwan Mandarin.
Subjects
PTT
dynamic corpus
Taiwan Mandarin
Type
thesis
File(s)![Thumbnail Image]()
Loading...
Name
ntu-103-R99142008-1.pdf
Size
23.54 KB
Format
Adobe PDF
Checksum
(MD5):bdb55de990f9a0f1e862d3975190875e
