상세 보기
Windows 악성코드 패밀리 데이터셋 구축 및 분류 실험
- 김태영;
- 최두섭;
- 임을규
초록
Malware family classification is a critical task for enhancing the efficiency of threat analysis and enabling rapid response strategies. However, accurate classification remains challenging due to behavioral similarities among different families and the ambiguous boundaries between variants. Moreover, most previous studies rely on outdated datasets, limiting their ability to reflect the latest trends in malware. To address these issues, this study constructs a new dataset of 3,357 Windows malware samples collected in 2024, with high label reliability ensured through cross-verification. Using this dataset, we applied a hybrid feature approach that combines static and dynamic features to a Random Forest model, achieving a maximum classification accuracy of 92.14%. An analysis of misclassified samples revealed that classification errors were often caused by shared API call sequences among certain malware families, leading to confusion, or by premature termination of malware execution, which hindered the collection of sufficient dynamic information. Based on these findings, we suggest the need for more sophisticated behavior-based feature extraction and improvements to the dynamic analysis environment to prevent early termination. This study is expected to make a practical contribution to enhancing the accuracy and reliability of future malware detection systems.
키워드
- 제목
- Windows 악성코드 패밀리 데이터셋 구축 및 분류 실험
- 제목 (타언어)
- Windows Malware Family Dataset Construction and Classification
- 저자
- 김태영; 최두섭; 임을규
- 발행일
- 2025-09
- 유형
- Y
- 저널명
- 정보처리학회 논문지
- 권
- 14
- 호
- 9
- 페이지
- 651 ~ 661