Automatic Training Corpora Acquisition for Document Classification
Date Issued
2008
Date
2008
Author(s)
Chi, Chun-Yi
Abstract
Document classification is a typical problem in several fields for many years. However, most previous work has the assumptions that the corpora can be explicitly-labeled and well-classified. In this work, we will concentrate on automatic acquisition of training data in good quality. We propose mining approaches to collect training data from given unlabeled corpus or the web, and our proposed approaches are fully automatic which is only needed to construct classes by humans in advance. n our work, the concept of class name can be captured by comparing with other classes, which is the common concept among classes. Moreover, we can discover discriminative concepts iteratively within each class. In this way, by finding common concepts and discriminative concepts, we can acquire training data of high quality. The evaluation gives empirical evidence that the classifiers thus created have promising accuracy. In a word, the automatic acquisition of training data in good quality by our proposed methods is the primary contributions of this work.
Subjects
Document Classification
Training Data
Type
thesis
File(s)![Thumbnail Image]()
Loading...
Name
ntu-97-R95922085-1.pdf
Size
23.32 KB
Format
Adobe PDF
Checksum
(MD5):eabe4f6c2f20834bab1e7584de2c2488
