Abstract
Propelled by advances in autonomous driving, accurate vehicle perception is essential for reliable decision-making. Compared with LiDAR, 4D millimeter-wave radar (4D MMW) provides stronger robustness under adverse conditions, lower deployment cost, and direct velocity measurements, motivating increasing research. However, its sparse point clouds limit scene description. This work presents Velo3DF (3D Velocity-aware Multimodal Fusion Network), a radar-camera fusion model that incorporates 3D velocity vectors into 4D MMW-based 3D detection. To mitigate positional redundancy and better utilize velocity information, a lightweight Velocity-Component-Aware module is introduced, compatible with backbones such as PointPillars. For cross-modal integration, an Alignment-Fusion module performs geometric alignment and deep fusion of image and point-cloud features. A single-modality variant, Velo3DF-R, contains 0.702M parameters and runs at 74.02 FPS while maintaining competitive accuracy.
| Original language | English |
|---|---|
| Journal | IEEE Transactions on Multimedia |
| DOIs | |
| Publication status | Accepted/In press - 2026 |
Keywords
- 3D Object Detection
- 4DMMW Radar
- Multimodal Fusion
- Vehicle Perception
Fingerprint
Dive into the research topics of 'Velocity First? Rethinking 3D Object Detection with 4D Millimeter Wave Radar'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver