Skip to main navigation Skip to search Skip to main content

M3Detection: Multi-Frame Multi-Level Feature Fusion for Multi-Modal 3-D Object Detection With Camera and 4-D Imaging Radar

  • Beijing Institute of Technology
  • Ministry of Education in China
  • Beijing Racobit Electronic Information Technology Co.Ltd

Research output: Contribution to journalArticlepeer-review

Abstract

Recent advances in 4D imaging radar have enabled robust perception in adverse weather, while camera sensors provide dense semantic information. Fusing these complementary modalities has great potential for accurate and cost-effective 3D perception. However, most existing camera-radar fusion methods are limited to single-frame inputs, capturing only a partial view of the scene. The incomplete scene information, compounded by image degradation and 4D radar sparsity, hinders overall detection performance. In contrast, multi-frame fusion offers richer spatial-temporal information but faces two challenges: achieving robust and effective object feature fusion across frames and modalities, and mitigating the computational cost of redundant feature extraction. Consequently, we propose M3Detection, a unified multi-frame 3D object detection framework that performs multi-level feature fusion on multi-modal data from camera and 4D imaging radar. In contrast to conventional architectures, our framework leverages intermediate features from the baseline detector and employs the tracker to produce reference trajectories, improving computational efficiency and providing richer information for second-stage. In the second stage, to address tracking uncertainties and enable fine-grained modeling, we design a global-level inter-object feature aggregation module (GOA) guided by radar information to align global features across candidate proposals and a local-level inter-grid feature aggregation module (LGA) that expands local features along the reference trajectories to enhance fine-grained object representation. The aggregated features are then processed by a trajectory-level multi-frame spatial-temporal fusion module (MSTF) to encode cross-frame interactions and enhance temporal representation. Extensive experiments on the View-of-Delft, TJ4DRadSet, and OmniHD-Scenes datasets demonstrate that M3Detection achieves state-of-the-art 3D detection performance, validating its effectiveness in multi-frame detection with camera-4D imaging radar fusion.

Original languageEnglish
JournalIEEE Transactions on Intelligent Transportation Systems
DOIs
Publication statusAccepted/In press - 2026
Externally publishedYes

Keywords

  • 3D object detection
  • 4D imaging radar
  • autonomous driving
  • camera
  • computer vision
  • multi-frame detection

Fingerprint

Dive into the research topics of 'M3Detection: Multi-Frame Multi-Level Feature Fusion for Multi-Modal 3-D Object Detection With Camera and 4-D Imaging Radar'. Together they form a unique fingerprint.

Cite this