비전 트랜스포머를 통한 Full Transformer 비디오 캡셔닝 모델 제안

A Full Transformer Video Captioning Model via Vision Transformer

초록

In the field of computer vision, a method to use a transformer model is emerging. Active research is underway on performance improvement using transformers not only for images but also for video data. ViViT proposed learning the temporal and spatial information of the video with two types of transformers. However, in the case of the ViViT and other using ViT models, including the first proposed ViT, only the learned CLS Token is used, and the remaining Patch sequence is not considered. In this study, various methods using the information of Patch Sequence learned through Self-attention are experimented. Thereafter, a video captioning task is performed using the corresponding method as a feature extraction network. Performance evaluation is conducted through four metrics, and the captioning results for the MSVD dataset of the final proposal network are close to those of the SOTA models in the BLEU-4, METEOR, and ROUGE-L metrics, even though only the appearance feature is used.

키워드

비전 트랜스포머비디오 캡셔닝딥러닝유니버셜 트랜스포머ViTvideo captioningdeep learninguniversal transformer
제목
비전 트랜스포머를 통한 Full Transformer 비디오 캡셔닝 모델 제안
제목 (타언어)
A Full Transformer Video Captioning Model via Vision Transformer
저자
임희주최용석
DOI
10.5626/KTCP.2023.29.8.378
발행일
2023-08
저널명
정보과학회 컴퓨팅의 실제 논문지
29
8
페이지
378 ~ 383