跳到主要导航 跳到搜索 跳到主要内容

Towards efficient multi-modal 3D object detection: Homogeneous sparse fuse network

  • Beijing Institute of Technology
  • Nanyang Technological University

科研成果: 期刊稿件文章同行评审

摘要

LiDAR-only 3D detection methods struggle with the sparsity of point clouds. To overcome this issue, multi-modal methods have been proposed, but their fusion is a challenge due to the heterogeneous representation of images and point clouds. This paper proposes a novel multi-modal framework, Homogeneous Sparse Fusion (HS-Fusion), which generates pseudo point clouds from depth completion. The proposed framework introduces a 3D foreground-aware middle extractor that efficiently extracts high-responding foreground features from sparse point cloud data. This module can be integrated into existing sparse convolutional neural networks. Furthermore, the proposed homogeneous attentive fusion enables cross-modality consistency fusion. Finally, the proposed HS-Fusion can simultaneously combine 2D image features and 3D geometric features of pseudo point clouds using multi-representation feature extraction. The proposed network has been found to attain better performance on the 3D object detection benchmarks. In particular, the proposed model demonstrates a 4.02% improvement in accuracy compared to the pure model. Moreover, its inference speed surpasses that of other models, thus further validating the efficacy of HS-Fusion.

源语言英语
期刊论文编号124945
期刊Expert Systems with Applications
256
DOI
出版状态已出版 - 5 12月 2024

学术指纹

探究 'Towards efficient multi-modal 3D object detection: Homogeneous sparse fuse network' 的科研主题。它们共同构成独一无二的学术指纹。

引用此