Efficient Training for Imbalance Data and Threshold Analysis for Self-training
Date Issued
2009
Date
2009
Author(s)
Chang, Chun-Min
Abstract
There are two methods proposed to address classification problems of imbalanced data. First, we propose a method that has smaller parameter space and more performance when using self-training. We train confidence thresholds for each classifier using labeled data to identify high confident data, and label them pseudo labels for re-train. Through this scheme we get less training time for parameters and get better performance. Second, we proposed an efficient training method for imbalanced data. We start with down-sampling and using a method like bootstrap. The model will approximate the model of up-sampling. Using less training data leads to less training time. We do experiments on KDDCUP 2008 data. The result shows that our threshold-based self-training has better performance and the approximated model has the same performance as up-sampling but cost only 0.75 times training time of up-sampling.
Subjects
semi-supervised learning
self-training
imbalanced data
kddcup 08
Type
thesis
File(s)![Thumbnail Image]()
Loading...
Name
ntu-98-R96944012-1.pdf
Size
23.32 KB
Format
Adobe PDF
Checksum
(MD5):673c90f33240f857972537cb67db5476
