GNN-Transformer Task Planning Enhanced with Semantic-Driven Data Augmentation

  • Jeong, Soojin
  • Byeon, Seongwan
  • Kim, Sangwoo
  • Kwon, HyeokJun
  • Oh, Yoonseon
Citations

SCOPUS

1

초록

Natural language is the most intuitive means for humans to interact with robots, making task planning based on natural language commands a longstanding area of research. Large language models (LLMs) have significantly improved task planning by enhancing understanding of language and common sense. However, current methods still face several challenges: they lack a deep understanding of physical environments, their performance relies heavily on prompt examples, LLMs are oversized and not customized for specific tasks, and the planning costs remain high. To overcome these issues, we introduce the GNN-Transformer Task Planner (GTTP), designed to predict task-level actions by leveraging the semantic environment and incorporating historical state data. The GTTP architecture is scalable through the use of GNN layers, while transformer layers facilitate understanding task progression. In addition, our model uses a text encoder to embed environments, allowing it to be trained on simulated datasets and applied directly in real-world scenarios. We also propose an automated data generation method that includes semantic augmentation, planning verification, and instruction generation via LLM. This method enables the collection of 14k instruction-annotated tasks in the VirtualHome environment with minimal human effort. The model has been validated across diverse scenes containing up to 715 objects, achieving significantly higher success rates compared to baseline models. It has also been successfully deployed on a physical mobile manipulator, demonstrating its practical applicability and effectiveness.

키워드

Deep learningVirtual environments
제목
GNN-Transformer Task Planning Enhanced with Semantic-Driven Data Augmentation
저자
Jeong, SoojinByeon, SeongwanKim, SangwooKwon, HyeokJunOh, Yoonseon
DOI
10.1609/aaai.v39i14.33598
발행일
2025-04
유형
Conference paper
저널명
Proceedings of the AAAI Conference on Artificial Intelligence
39
14
페이지
14585 ~ 14593