XMP: A CROSS-ATTENTION MULTI-SCALE PERFORMER FOR FILE FRAGMENT CLASSIFICATION

Citations

WEB OF SCIENCE

4
Citations

SCOPUS

6

초록

File fragment classification (FFC) is the task of identifying the file type given a small fraction of binary data, and serves a crucial role in digital forensics and cybersecurity. Recent studies have adopted convolutional neural networks (CNNs) for this problem, significantly improving the accuracy over the traditional methods relying on handcrafted features. In this paper, we aim to expand on the recent performance gain by better leveraging the large amount of digital files available for training. We propose to achieve this by employing a Transformer encoder-based network known for its weak inductive bias suited for large-scale training. Our model, XMP, is inspired by the CrossViT architecture for image recognition and utilizes multi-scale self and cross-attentions between CNN features extracted from the byte n-grams of input binary data. Experimental results on the latest public dataset show XMP achieving state-of-the-art accuracies in almost all scenarios without need for additional preprocessing of binary data such as bit shifting, demonstrating the effectiveness of the Transformer-based architecture for FFC. The benefit of each proposed component is assessed through ablation study. Our code is available at github.com/pank40/xmp.

키워드

file fragment classificationTransformermulti-scale attentioncross-attentionperformerConvolutional Neural NetworkBinary DataImage RecognitionConvolutional Neural Network FeaturesFile TypeInductive BiasBit-shiftModel PerformanceTraining DataSupport Vector MachineData AugmentationAttention MechanismMultilayer PerceptronOutput FeatureHyperparameter TuningMulti-scale FeaturesLarge-scale FeaturesMulti-scale Feature ExtractionPosition EmbeddingSmall-scale FeaturesBytes Of DataAttention MatrixKey Matrix
제목
XMP: A CROSS-ATTENTION MULTI-SCALE PERFORMER FOR FILE FRAGMENT CLASSIFICATION
저자
Park, Jeong GyuLiu, SisungHong, Je Hyeong
DOI
10.1109/ICASSP48485.2024.10447626
발행일
2024-04
유형
Proceedings Paper
저널명
ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings
페이지
4505 ~ 4509