시공간 그래프 랜덤워크를 활용한 비디오 의미구조 이해

Understanding Video Semantic Structure with Spatiotemporal Graph Random Walk

초록

Understanding a long video focuses on finding various semantic units present in the video and interpreting complex relationships among them. Conventional approaches utilize models based on CNNs or transformers to encode contextual information for short clips and then consider temporal relationships among them. However, such approaches struggle to capture complex relationships among smaller semantic units within video clips. In this paper, we present video inputs using a spatiotemporal graph with objects as vertices and relative space-time information between objects as edges, to explicitly express relationships among these semantic units. Additionally, we proposed a novel method to represent major semantic units as compositions of smaller units using high-order relationship information obtained by spatiotemporal random walks on the graph. Through experiments on CATER dataset, which involved complex actions of multiple objects, we demonstrated that our approach exhibited effective semantic unit capturing capabilities.

키워드

video understandingcompositional learningspatiotemporal graphrandom walksemantic unit비디오 이해구성적 학습시공간 그래프랜덤워크의미단위
제목
시공간 그래프 랜덤워크를 활용한 비디오 의미구조 이해
제목 (타언어)
Understanding Video Semantic Structure with Spatiotemporal Graph Random Walk
저자
윤호영김민서김은솔
DOI
10.5626/JOK.2024.51.9.801
발행일
2024-09
저널명
정보과학회논문지
51
9
페이지
801 ~ 806