跳到主要导航 跳到搜索 跳到主要内容

M3Detection: Multi-Frame Multi-Level Feature Fusion for Multi-Modal 3-D Object Detection With Camera and 4-D Imaging Radar

  • Beijing Institute of Technology
  • Ministry of Education in China
  • Beijing Racobit Electronic Information Technology Co.Ltd

科研成果: 期刊稿件文章同行评审

摘要

Recent advances in 4D imaging radar have enabled robust perception in adverse weather, while camera sensors provide dense semantic information. Fusing these complementary modalities has great potential for accurate and cost-effective 3D perception. However, most existing camera-radar fusion methods are limited to single-frame inputs, capturing only a partial view of the scene. The incomplete scene information, compounded by image degradation and 4D radar sparsity, hinders overall detection performance. In contrast, multi-frame fusion offers richer spatial-temporal information but faces two challenges: achieving robust and effective object feature fusion across frames and modalities, and mitigating the computational cost of redundant feature extraction. Consequently, we propose M3Detection, a unified multi-frame 3D object detection framework that performs multi-level feature fusion on multi-modal data from camera and 4D imaging radar. In contrast to conventional architectures, our framework leverages intermediate features from the baseline detector and employs the tracker to produce reference trajectories, improving computational efficiency and providing richer information for second-stage. In the second stage, to address tracking uncertainties and enable fine-grained modeling, we design a global-level inter-object feature aggregation module (GOA) guided by radar information to align global features across candidate proposals and a local-level inter-grid feature aggregation module (LGA) that expands local features along the reference trajectories to enhance fine-grained object representation. The aggregated features are then processed by a trajectory-level multi-frame spatial-temporal fusion module (MSTF) to encode cross-frame interactions and enhance temporal representation. Extensive experiments on the View-of-Delft, TJ4DRadSet, and OmniHD-Scenes datasets demonstrate that M3Detection achieves state-of-the-art 3D detection performance, validating its effectiveness in multi-frame detection with camera-4D imaging radar fusion.

源语言英语
期刊IEEE Transactions on Intelligent Transportation Systems
DOI
出版状态已接受/待刊 - 2026
已对外发布

指纹

探究 'M3Detection: Multi-Frame Multi-Level Feature Fusion for Multi-Modal 3-D Object Detection With Camera and 4-D Imaging Radar' 的科研主题。它们共同构成独一无二的指纹。

引用此