Abstract
Fusion-based 3D object detection is critical for autonomous driving. However, existing multimodal fusion methods often suffer from an insufficient receptive field, inefficient cross-modal interaction, and weak generalization to rare or irregular objects. We present BEV-SA, a novel bird’s eye view–based fusion framework that overcomes these limitations through two key modules. The state-space BEV interaction (SSB) module serializes BEV features along a Hilbert curve and models them using a structured state space duality (SSD) mechanism, enabling global context modeling, precise modality alignment, and linear computational complexity. The accelerated separable fusion (ASF) module employs depthwise separable convolutions to efficiently merge aligned BEV features, reducing inference time by 12.5% compared to conventional fusion without sacrificing accuracy. We also introduce the anomalous-shaped commercial vehicle (ASC) dataset, a benchmark focusing on large and irregular commercial vehicles in challenging environments. Experiments on the nuScenes benchmark show that BEV-SA achieves state-of-the-art performance, improving NDS by 0.24% and mAP by 2.73%, while maintaining high efficiency. On ASC, BEV-SA demonstrates remarkable generalization to rare and irregular objects.
| Original language | English |
|---|---|
| Article number | 133029 |
| Journal | Expert Systems with Applications |
| Volume | 330 |
| DOIs | |
| Publication status | Published - 1 Dec 2026 |
Keywords
- 3D object detection
- Autonomous driving
- Depthwise separable convolutions
- LiDAR-camera fusion
- Structured state space duality
Fingerprint
Dive into the research topics of 'Global context alignment and separable fusion for generalizable multi-modal 3D object detection'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver