https://scholars.lib.ntu.edu.tw/handle/123456789/413147
標題: | Development of a web-scale Chinese word N-gram corpus with parts of speech information | 作者: | Yu C.-H. Tang Y.-J. Chen H.-H. |
關鍵字: | ClueWeb09;Encoding detection;Part-of-speech n-grams | 公開日期: | 2012 | 起(迄)頁: | 320-324 | 來源出版物: | 8th International Conference on Language Resources and Evaluation, LREC 2012 | 摘要: | Web provides a large-scale corpus for researchers to study the language usages in real world. Developing a web-scale corpus needs not only a lot of computation resources, but also great efforts to handle the large variations in the web texts, such as character encoding in processing Chinese web texts. In this paper, we aim to develop a web-scale Chinese word N-gram corpus with parts of speech information called NTU PN-Gram corpus using the ClueWeb09 dataset. We focus on the character encoding and some Chinese-specific issues. The statistics about the dataset is reported. We will make the resulting corpus a public available resource to boost the Chinese language processing. |
URI: | https://scholars.lib.ntu.edu.tw/handle/123456789/413147 | ISBN: | 9782951740877 |
顯示於: | 資訊工程學系 |
在 IR 系統中的文件,除了特別指名其著作權條款之外,均受到著作權保護,並且保留所有的權利。