시각적 관계 예측을 위한 계산 효율적인 조합적 전이 표현 학습법

Efficient Compositional Translation Embedding for Visual Relationship Detection

초록

Scene graphs are widely used to express high-order visual relationships between objects present in an image. To generate the scene graph automatically, we propose an algorithm that detects visual relationships between objects and predicts the relationship as a predicate. Inspired by the well-known knowledge graph embedding method TransR, we present the CompTransR algorithm that i) defines latent relational subspaces considering the compositional perspective of visual relationships and ii) encodes predicate representations by applying transitive constraints between the object representations in each subspace. Our proposed model not only reduces computational complexity but also outperformed previous state-of-the-art performance in predicate detection tasks in three benchmark datasets: VRD, VG200, and VrR-VG. We also showed that a scene graph could be applied to the image-caption retrieval task, which is one of the high-level visual reasoning tasks, and the scene graph generated by our model increased retrieval performance.

키워드

장면 그래프 생성시각적 관계 예측이미지 캡션 검색전이 표현scene graph generationvisual relationship detectionimage caption retrievaltranslation embedding
제목
시각적 관계 예측을 위한 계산 효율적인 조합적 전이 표현 학습법
제목 (타언어)
Efficient Compositional Translation Embedding for Visual Relationship Detection
저자
허유정김은솔최우석온경운장병탁
DOI
10.5626/JOK.2022.49.7.544
발행일
2022-07
저널명
정보과학회논문지
49
7
페이지
544 ~ 554