Dense but Efficient VideoQA for Intricate Compositional Reasoning

Citations

WEB OF SCIENCE

5
Citations

SCOPUS

5

초록

It is well known that most of the conventional video question answering (VideoQA) datasets consist of easy questions requiring simple reasoning processes. However, long videos inevitably contain complex and compositional semantic structures along with the spatio-temporal axis, which requires a model to understand the compositional structures inherent in the videos. In this paper, we suggest a new compositional VideoQA method based on transformer architecture with a deformable attention mechanism to address the complex VideoQA tasks. The deformable attentions are introduced to sample a subset of informative visual features from the dense visual feature map to cover a temporally long range of frames efficiently. Furthermore, the dependency structure within the complex question sentences is also combined with the language embeddings to readily understand the relations among question words. Extensive experiments and ablation studies show that the suggested dense but efficient model outperforms other baselines.

키워드

Algorithms: Video recognition and understanding (tracking, action recognition, etc.)Vision + language and/or other modalitiesComputer visionAction recognitionAlgorithm: video recognition and understanding (tracking, action recognition, etc.Compositional reasoningQuestion AnsweringSimple++Video recognitionVideo understandingVision + language and/or other modalityVisual featureSemantics
제목
Dense but Efficient VideoQA for Intricate Compositional Reasoning
저자
Lee, JihyeonKang, WooyoungKim, Eun Sol
DOI
10.1109/WACV56688.2023.00117
발행일
2023-01
유형
Proceedings Paper
저널명
2023 IEEE/CVF WINTER CONFERENCE ON APPLICATIONS OF COMPUTER VISION (WACV)
페이지
1114 ~ 1123