Benchmarking binary classification models on data sets with different degrees of imbalance

作者：Ligang Zhou, Kin Keung Lai

摘要

In practice, there are many binary classification problems, such as credit risk assessment, medical testing for determining if a patient has a certain disease or not, etc. However, different problems have different characteristics that may lead to different difficulties of the problem. One important characteristic is the degree of imbalance of two classes in data sets. For data sets with different degrees of imbalance, are the commonly used binary classification methods still feasible? In this study, various binary classification models, including traditional statistical methods and newly emerged methods from artificial intelligence, such as linear regression, discriminant analysis, decision tree, neural network, support vector machines, etc., are reviewed, and their performance in terms of the measure of classification accuracy and area under Receiver Operating Characteristic (ROC) curve are tested and compared on fourteen data sets with different imbalance degrees. The results help to select the appropriate methods for problems with different degrees of imbalance.

论文关键词：binary classification, area under Receiver Operating Characteristic (ROC) curve, classification accuracy, degrees of imbalance

论文评审过程：

论文官网地址：https://doi.org/10.1007/s11704-009-0027-1