상세 보기
비전 트랜스포머를 통한 Full Transformer 비디오 캡셔닝 모델 제안
- 임희주;
- 최용석
초록
In the field of computer vision, a method to use a transformer model is emerging. Active research is underway on performance improvement using transformers not only for images but also for video data. ViViT proposed learning the temporal and spatial information of the video with two types of transformers. However, in the case of the ViViT and other using ViT models, including the first proposed ViT, only the learned CLS Token is used, and the remaining Patch sequence is not considered. In this study, various methods using the information of Patch Sequence learned through Self-attention are experimented. Thereafter, a video captioning task is performed using the corresponding method as a feature extraction network. Performance evaluation is conducted through four metrics, and the captioning results for the MSVD dataset of the final proposal network are close to those of the SOTA models in the BLEU-4, METEOR, and ROUGE-L metrics, even though only the appearance feature is used.
키워드
- 제목
- 비전 트랜스포머를 통한 Full Transformer 비디오 캡셔닝 모델 제안
- 제목 (타언어)
- A Full Transformer Video Captioning Model via Vision Transformer
- 저자
- 임희주; 최용석
- 발행일
- 2023-08
- 권
- 29
- 호
- 8
- 페이지
- 378 ~ 383