Video Captioning with Vision Transformer

비전 트랜스포머를 이용한 비디오 캡셔닝

초록

Video captioning not only needs sequence-to-sequence captioning model but also the convolutional Neutral Network (CNN) backbones to extract features such as appearance, motion, and object features. In this paper, we propose a full transformer video captioning model without several CNNs. Moreover, we compare our captioning model to other models that use several features.

제목
Video Captioning with Vision Transformer
제목 (타언어)
비전 트랜스포머를 이용한 비디오 캡셔닝
저자
Hee Ju Im Ho Seok Ahn 최용석
발행일
2021-12
유형
Proceeding
저널명
한국소프트웨어종합학술대회 (KSC 2021)
페이지
1 ~ 3