From Language to Grasp: Object Retrieval and Grasping Through Explicit and Implicit Linguistic Commands

Citations

SCOPUS

1

초록

In human-centered environments, assistive robots are required to understand verbal commands to retrieve and grasp objects within complex scenes. We propose a novel Language Understanding Object Retrieval module (LUOR) by fine-tuning the CLIP text encoder to enhance robot manipulators' understanding of both explicit and implicit natural language commands. A new dataset with 712 verb-object pairs is created for training. This dataset includes 78 verbs associated with 244 ImageNet classes, providing a comprehensive range of scenarios. Additionally, 336 verb-object pairs cover 54 verbs for 138 ObjectNet classes, further expanding the model's applicability. Experimental results demonstrate that LUOR outperforms existing baselines in both accuracy and efficiency, particularly in handling implicit commands. The integrated system with the Multi-Task Detection module (MTD) shows strong performance in real-world robotic applications using a Panda Franka manipulator. These findings confirm the practical applicability of our approach and suggest potential for further improvements in robotic grasping and manipulation tasks.

키워드

grasp detectionmulti-modal learningRobotic object retrievalAdversarial machine learningContent based retrievalContrastive LearningIndustrial robotsLinguisticsModular robotsMulti-task learningNatural language processing systemsObject detectionObject recognitionRobot applicationsRobot learning
제목
From Language to Grasp: Object Retrieval and Grasping Through Explicit and Implicit Linguistic Commands
저자
Yoon, DongminCha, SeonghunOh, Yoonseon
DOI
10.23919/ICCAS63016.2024.10773029
발행일
2024-10
유형
Conference paper
저널명
International Conference on Control, Automation and Systems
페이지
1565 ~ 1566