상세 보기
초록
Class distribution of unbalanced data is an important part of the digital world and is a significant part of cybersecurity. Abnormalactivity of unbalanced data should be found and problems solved. Although a system capable of tracking patterns in all transactionsis needed, machine learning with disproportionate data, which typically has abnormal patterns, can ignore and degrade performancefor minority layers, and predictive models can be inaccurately biased. In this paper, we predict target variables and improve accuracyby combining estimates using Synthetic Minority Oversampling Technique (SMOTE) and Light GBM algorithms as an approach to addressunbalanced datasets. Experimental results were compared with logistic regression, decision tree, KNN, Random Forest, and XGBoostalgorithms. The performance was similar in accuracy and reproduction rate, but in precision, two algorithms performed at Random Forest80.76% and Light GBM 97.16%, and in F1-score, Random Forest 84.67% and Light GBM 91.96%. As a result of this experiment, it wasconfirmed that Light GBM's performance was similar without deviation or improved by up to 16% compared to five algorithms.
키워드
- 제목
- SMOTE와 Light GBM 기반의 불균형 데이터 개선 기법
- 제목 (타언어)
- Imbalanced Data Improvement Techniques Based on SMOTE and Light GBM
- 저자
- 한영진; 조인휘
- 발행일
- 2022-12
- 권
- 11
- 호
- 12
- 페이지
- 445 ~ 452