Convolutional Approach to learning Temporal Feature Effectively in Transformer

트랜스포머의 효과적인 시간 특징 정보 학습을 위한 합성곱 기법

초록

In the video classification task, a well-performing deep learning model is likely to extract proper temporal features to classify the data. However, we found out several problems of attention-based TimeSformer[2] related to extracting temporal features, and replaced the time attention module in the TimeSformer with the 3D convolution module for better temporal feature processing. Through several experiments and visualization results, we demonstrate that the 3D convolution module can extract more accurate temporal features of video data than the time-attention module.

제목
Convolutional Approach to learning Temporal Feature Effectively in Transformer
제목 (타언어)
트랜스포머의 효과적인 시간 특징 정보 학습을 위한 합성곱 기법
저자
Park, Hae Sung Jung, Hyuck ChulChoi, Yong Suk
발행일
2022-12
유형
Proceeding
저널명
2023 한국소프트웨어종합학술대회 (KSC 2023)
페이지
517 ~ 519