Multi-modal dataset and fusion network for simultaneous semantic segmentation of on-road dynamic objects

  • Cho, Jieun
  • Ha, Jinsu
  • Song, Hamin
  • Jang, Sungmoon
  • Jo, Kichun
Citations

WEB OF SCIENCE

3
Citations

SCOPUS

4

초록

An accurate and robust perception system is essential for autonomous vehicles to interact with various dynamic objects on the road. By applying semantic segmentation techniques to the data from the camera sensor and light detection and ranging sensor, dynamic objects can be classified at pixel and point levels respectively. However, there are challenges when using a single sensor, especially under adverse lighting conditions or with sparse point densities. To address these challenges, this paper proposes a network for simultaneous point cloud and image semantic segmentation based on sensor fusion. The proposed network adopts a modal-specific architecture to fully leverage the characteristics of sensor data and achieves geometrically accurate matching through the image, point, and voxel feature fusion module. Additionally, we introduce the dataset that provides semantic labels for synchronized images and point clouds. Experimental results show that the proposed fusion approach outperforms uni-modal based methods and demonstrates robust performance even in challenging real-world scenarios. The dataset is publicly available at https://github.com/ailab-konkuk/Multi-Modal-Dataset.

키워드

Autonomous drivingDeep learningPerceptionSemantic segmentationSensor fusionImage matchingSensor data fusion
제목
Multi-modal dataset and fusion network for simultaneous semantic segmentation of on-road dynamic objects
저자
Cho, JieunHa, JinsuSong, HaminJang, SungmoonJo, Kichun
DOI
10.1016/j.engappai.2025.110024
발행일
2025-03
유형
Article
저널명
Engineering Applications of Artificial Intelligence
143
페이지
1 ~ 11