LiveCap: Live Video Captioning with Sequential Encoding Network

Citations

SCOPUS

0

초록

Today, video captioning frameworks are very useful in places such as video surveillance systems. Most of these systems require real-time captioning, however existing video captioning frameworks still have some limitations in live video. Specifically, they require the whole video to describe. In this paper, we propose LiveCap, a framework for generating sentences corresponding to the current scene in real time from live video. LiveCap consists of three modules: sequential encoding network, captioning network, and context gating network. Our framework accumulates context for sequentially given video segments (sequential encoding network) and generates sentences based on it (captioning network). Furthermore, the context gating network controls the flow between the two networks to determine when to generate sentences. We train and test LiveCap on the ActivityNet Captions dataset and verify that LiveCap generates fluent and coherent captions in live video.

키워드

Encoding (symbols)Network codingSecurity systemsStatistical testsVideo signal processingReal time systemscurrentEncodingsFluentsLive videoNetwork-controlReal- timeSentence-basedVideo segmentsVideo surveillance systems
제목
LiveCap: Live Video Captioning with Sequential Encoding Network
저자
Choi, WangyuYoon, Jungwon
DOI
10.1109/ICTC55196.2022.9952747
발행일
2022-10
유형
Conference Paper
저널명
International Conference on ICT Convergence
2022-October
페이지
1894 ~ 1896