Oversampling to overcome overfitting: Exploring the relationship between data set composition, molecular descriptors, and predictive modeling methods
Resource
J. Chem. Inf. Model., 53, 958-971
Journal
Journal of Chemical Information and Modeling
Journal Volume
53
Journal Issue
4
Pages
958-971
Date Issued
2013
Author(s)
Abstract
The traditional biological assay is very time-consuming, and thus the ability to quickly screen large numbers of compounds against a specific biological target is appealing. To speed up the biological evaluation of compounds, high-throughput screening is widely used in the fields of biomedical, biological information, and drug discovery. The research presented in this study focuses on the use of support vector machines, a machine learning method, various classes of molecular descriptors, and different sampling techniques to overcome overfitting to classify compounds for cytotoxicity with respect to the Jurkat cell line. The cell cytotoxicity data set is imbalanced (a few active compounds and very many inactive compounds), and the ability of the predictive modeling methods is adversely affected in these situations. Commonly imbalanced data sets are overfit with respect to the dominant classified end point; in this study the models routinely overfit toward inactive (noncytotoxic) compounds when the imbalance was substantial. Support vector machine (SVM) models were used to probe the proficiency of different classes of molecular descriptors and oversampling ratios. The SVM models were constructed from 4D-FPs, MOE (1D, 2D, and 21/2D), noNP+MOE, and CATS2D trial descriptors pools and compared to the predictive abilities of CATS2D-based random forest models. Compared to previous results in the literature, the SVM models built from oversampled data sets exhibited better predictive abilities for the training and external test sets. © 2013 American Chemical Society.
SDGs
Other Subjects
Cell culture; Decision trees; Drug products; Predictive analytics; Support vector machines; Biological evaluation; Biological information; High throughput screening; Imbalanced Data-sets; Machine learning methods; Molecular descriptors; Predictive abilities; Predictive modeling; Learning systems; cytotoxin; article; cell survival; chemistry; drug development; drug effect; high throughput screening; human; leukemia cell line; molecular library; predictive value; quantitative structure activity relation; statistical model; support vector machine; Cell Survival; Cytotoxins; Drug Discovery; High-Throughput Screening Assays; Humans; Jurkat Cells; Models, Statistical; Predictive Value of Tests; Quantitative Structure-Activity Relationship; Small Molecule Libraries; Support Vector Machines
Type
journal article
File(s)![Thumbnail Image]()
Loading...
Name
Oversampling to Overcome Overfitting Exploring the relationship between data set composition, molecular descriptors, and predictive modeling methods.pdf
Size
2 MB
Format
Adobe PDF
Checksum
(MD5):252a7d5929880db8af3a11cedcab3f3d
