외부 지식 그래프 결합을 위한 그래프 변환기 알고리즘

A New Graph Transformer Algorithm for Leveraging External Knowledge Graph

초록

Visual Commonsense Reasoning(VCR) presents a more challenging problem compared to Visual Question Answering(VQA), which primarily requires understanding visual characteristics and relationships among objects within an image. In addition to the question itself, VCR necessitates a contextual comprehension of the scene and general commonsense knowledge. This paper proposes a knowledge graph construction and graph transformer learning algorithm to integrate knowledge related to general commonsense from external knowledge systems. In our proposed model, knowledge is retrieved from ConceptNet, an external knowledge system, based on the given modality information to construct a knowledge graph. This knowledge graph, along with images and text, is used as input for the graph transformer as a unified token without distinguishing between nodes and edges during training. To demonstrate the superiority of our proposed model, we conduct experiments using the VCR dataset and compare the improved performance with baseline models.

키워드

external knowledge baseknowledge graphgraph transformercommonsense reasoningmulti-modal learning외부 지식 체계지식 그래프그래프 변환기일반 상식 추론다중 양상 학습
제목
외부 지식 그래프 결합을 위한 그래프 변환기 알고리즘
제목 (타언어)
A New Graph Transformer Algorithm for Leveraging External Knowledge Graph
저자
안경환김은솔
DOI
10.5626/KTCP.2024.30.11.588
발행일
2024-11
저널명
정보과학회 컴퓨팅의 실제 논문지
30
11
페이지
588 ~ 593