상세 보기
From Language to Grasp: Object Retrieval and Grasping Through Explicit and Implicit Linguistic Commands
- Yoon, Dongmin;
- Cha, Seonghun;
- Oh, Yoonseon
SCOPUS
1초록
In human-centered environments, assistive robots are required to understand verbal commands to retrieve and grasp objects within complex scenes. We propose a novel Language Understanding Object Retrieval module (LUOR) by fine-tuning the CLIP text encoder to enhance robot manipulators' understanding of both explicit and implicit natural language commands. A new dataset with 712 verb-object pairs is created for training. This dataset includes 78 verbs associated with 244 ImageNet classes, providing a comprehensive range of scenarios. Additionally, 336 verb-object pairs cover 54 verbs for 138 ObjectNet classes, further expanding the model's applicability. Experimental results demonstrate that LUOR outperforms existing baselines in both accuracy and efficiency, particularly in handling implicit commands. The integrated system with the Multi-Task Detection module (MTD) shows strong performance in real-world robotic applications using a Panda Franka manipulator. These findings confirm the practical applicability of our approach and suggest potential for further improvements in robotic grasping and manipulation tasks.
키워드
- 제목
- From Language to Grasp: Object Retrieval and Grasping Through Explicit and Implicit Linguistic Commands
- 저자
- Yoon, Dongmin; Cha, Seonghun; Oh, Yoonseon
- 발행일
- 2024-10
- 유형
- Conference paper
- 저널명
- International Conference on Control, Automation and Systems
- 페이지
- 1565 ~ 1566