TY - JOUR
T1 - M3Detection
T2 - Multi-Frame Multi-Level Feature Fusion for Multi-Modal 3-D Object Detection With Camera and 4-D Imaging Radar
AU - Li, Xiaozhi
AU - Di, Huijun
AU - Li, Jian
AU - Liu, Feng
AU - Liang, Wei
N1 - Publisher Copyright:
© 2000-2011 IEEE.
PY - 2026
Y1 - 2026
N2 - Recent advances in 4D imaging radar have enabled robust perception in adverse weather, while camera sensors provide dense semantic information. Fusing these complementary modalities has great potential for accurate and cost-effective 3D perception. However, most existing camera-radar fusion methods are limited to single-frame inputs, capturing only a partial view of the scene. The incomplete scene information, compounded by image degradation and 4D radar sparsity, hinders overall detection performance. In contrast, multi-frame fusion offers richer spatial-temporal information but faces two challenges: achieving robust and effective object feature fusion across frames and modalities, and mitigating the computational cost of redundant feature extraction. Consequently, we propose M3Detection, a unified multi-frame 3D object detection framework that performs multi-level feature fusion on multi-modal data from camera and 4D imaging radar. In contrast to conventional architectures, our framework leverages intermediate features from the baseline detector and employs the tracker to produce reference trajectories, improving computational efficiency and providing richer information for second-stage. In the second stage, to address tracking uncertainties and enable fine-grained modeling, we design a global-level inter-object feature aggregation module (GOA) guided by radar information to align global features across candidate proposals and a local-level inter-grid feature aggregation module (LGA) that expands local features along the reference trajectories to enhance fine-grained object representation. The aggregated features are then processed by a trajectory-level multi-frame spatial-temporal fusion module (MSTF) to encode cross-frame interactions and enhance temporal representation. Extensive experiments on the View-of-Delft, TJ4DRadSet, and OmniHD-Scenes datasets demonstrate that M3Detection achieves state-of-the-art 3D detection performance, validating its effectiveness in multi-frame detection with camera-4D imaging radar fusion.
AB - Recent advances in 4D imaging radar have enabled robust perception in adverse weather, while camera sensors provide dense semantic information. Fusing these complementary modalities has great potential for accurate and cost-effective 3D perception. However, most existing camera-radar fusion methods are limited to single-frame inputs, capturing only a partial view of the scene. The incomplete scene information, compounded by image degradation and 4D radar sparsity, hinders overall detection performance. In contrast, multi-frame fusion offers richer spatial-temporal information but faces two challenges: achieving robust and effective object feature fusion across frames and modalities, and mitigating the computational cost of redundant feature extraction. Consequently, we propose M3Detection, a unified multi-frame 3D object detection framework that performs multi-level feature fusion on multi-modal data from camera and 4D imaging radar. In contrast to conventional architectures, our framework leverages intermediate features from the baseline detector and employs the tracker to produce reference trajectories, improving computational efficiency and providing richer information for second-stage. In the second stage, to address tracking uncertainties and enable fine-grained modeling, we design a global-level inter-object feature aggregation module (GOA) guided by radar information to align global features across candidate proposals and a local-level inter-grid feature aggregation module (LGA) that expands local features along the reference trajectories to enhance fine-grained object representation. The aggregated features are then processed by a trajectory-level multi-frame spatial-temporal fusion module (MSTF) to encode cross-frame interactions and enhance temporal representation. Extensive experiments on the View-of-Delft, TJ4DRadSet, and OmniHD-Scenes datasets demonstrate that M3Detection achieves state-of-the-art 3D detection performance, validating its effectiveness in multi-frame detection with camera-4D imaging radar fusion.
KW - 3D object detection
KW - 4D imaging radar
KW - autonomous driving
KW - camera
KW - computer vision
KW - multi-frame detection
UR - https://www.scopus.com/pages/publications/105040537495
U2 - 10.1109/TITS.2026.3695618
DO - 10.1109/TITS.2026.3695618
M3 - Article
AN - SCOPUS:105040537495
SN - 1524-9050
JO - IEEE Transactions on Intelligent Transportation Systems
JF - IEEE Transactions on Intelligent Transportation Systems
ER -