Classifying Malicious Documents on the Basis of Plain-Text Features: Problem, Solution, and Experiences

Citations

WEB OF SCIENCE

4
Citations

SCOPUS

5

초록

Cyberattacks widely occur by using malicious documents. A malicious document is an electronic document containing malicious codes along with some plain-text data that is human-readable. In this paper, we propose a novel framework that takes advantage of such plaintext data to determine whether a given document is malicious. We extracted plaintext features from the corpus of electronic documents and utilized them to train a classification model for detecting malicious documents. Our extensive experimental results with different combinations of three well-known vectorization strategies and three popular classification methods on five types of electronic documents demonstrate that our framework provides high prediction accuracy in detecting malicious documents.

키워드

malwaremalicious documentclassificationtext analysisNEURAL-NETWORKS
제목
Classifying Malicious Documents on the Basis of Plain-Text Features: Problem, Solution, and Experiences
저자
Hong, JiwonJeong, DonghoKim, Sang-Wook
DOI
10.3390/app12084088
발행일
2022-04
유형
Article
저널명
APPLIED SCIENCES-BASEL
12
8
페이지
1 ~ 13

파일 다운로드