Semantic-Aware Dynamic Parameter for Video Inpainting Transformer

Citations

WEB OF SCIENCE

3
Citations

SCOPUS

6

초록

Recent learning-based video inpainting approaches have achieved considerable progress. However, they still cannot fully utilize semantic information within the video frames and predict improper scene layout, failing to restore clear object boundaries for mixed scenes. To mitigate this problem, we introduce a new transformer-based video inpainting technique that can exploit semantic information within the input and considerably improve reconstruction quality. In this study, we use the mixture-of-experts scheme and train multiple experts to handle mixed scenes, including various semantics. We leverage these multiple experts and produce locally (token-wise) different network parameters to achieve semantic-aware inpainting results. Extensive experiments on YouTube-VOS and DAVIS benchmark datasets demonstrate that, compared with existing conventional video inpainting approaches, the proposed method has superior performance in synthesizing visually pleasing videos with much clearer semantic structures and textures.

제목
Semantic-Aware Dynamic Parameter for Video Inpainting Transformer
저자
Lee, EunhyeYoo, JinsuYang, YunjeongBaik, SungyongKim, Tae Hyun
DOI
10.1109/ICCV51070.2023.01190
발행일
2023-10
유형
Proceedings Paper
저널명
2023 IEEE/CVF INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV 2023)
페이지
12903 ~ 12912