상세 보기
Video Captioning with Vision Transformer
비전 트랜스포머를 이용한 비디오 캡셔닝
- Hee Ju Im ;
- Ho Seok Ahn ;
- 최용석
초록
Video captioning not only needs sequence-to-sequence captioning model but also the convolutional Neutral Network (CNN) backbones to extract features such as appearance, motion, and object features. In this paper, we propose a full transformer video captioning model without several CNNs. Moreover, we compare our captioning model to other models that use several features.
- 제목
- Video Captioning with Vision Transformer
- 제목 (타언어)
- 비전 트랜스포머를 이용한 비디오 캡셔닝
- 저자
- Hee Ju Im ; Ho Seok Ahn ; 최용석
- 발행일
- 2021-12
- 유형
- Proceeding
- 저널명
- 한국소프트웨어종합학술대회 (KSC 2021)
- 페이지
- 1 ~ 3