TY - JOUR
T1 - Global context alignment and separable fusion for generalizable multi-modal 3D object detection
AU - Tang, Yingjuan
AU - He, Hongwen
AU - Wu, Jingda
AU - Shen, Yongpeng
AU - Wang, Yong
AU - Wu, Yifan
N1 - Publisher Copyright:
© 2026 Published by Elsevier Ltd.
PY - 2026/12/1
Y1 - 2026/12/1
N2 - Fusion-based 3D object detection is critical for autonomous driving. However, existing multimodal fusion methods often suffer from an insufficient receptive field, inefficient cross-modal interaction, and weak generalization to rare or irregular objects. We present BEV-SA, a novel bird’s eye view–based fusion framework that overcomes these limitations through two key modules. The state-space BEV interaction (SSB) module serializes BEV features along a Hilbert curve and models them using a structured state space duality (SSD) mechanism, enabling global context modeling, precise modality alignment, and linear computational complexity. The accelerated separable fusion (ASF) module employs depthwise separable convolutions to efficiently merge aligned BEV features, reducing inference time by 12.5% compared to conventional fusion without sacrificing accuracy. We also introduce the anomalous-shaped commercial vehicle (ASC) dataset, a benchmark focusing on large and irregular commercial vehicles in challenging environments. Experiments on the nuScenes benchmark show that BEV-SA achieves state-of-the-art performance, improving NDS by 0.24% and mAP by 2.73%, while maintaining high efficiency. On ASC, BEV-SA demonstrates remarkable generalization to rare and irregular objects.
AB - Fusion-based 3D object detection is critical for autonomous driving. However, existing multimodal fusion methods often suffer from an insufficient receptive field, inefficient cross-modal interaction, and weak generalization to rare or irregular objects. We present BEV-SA, a novel bird’s eye view–based fusion framework that overcomes these limitations through two key modules. The state-space BEV interaction (SSB) module serializes BEV features along a Hilbert curve and models them using a structured state space duality (SSD) mechanism, enabling global context modeling, precise modality alignment, and linear computational complexity. The accelerated separable fusion (ASF) module employs depthwise separable convolutions to efficiently merge aligned BEV features, reducing inference time by 12.5% compared to conventional fusion without sacrificing accuracy. We also introduce the anomalous-shaped commercial vehicle (ASC) dataset, a benchmark focusing on large and irregular commercial vehicles in challenging environments. Experiments on the nuScenes benchmark show that BEV-SA achieves state-of-the-art performance, improving NDS by 0.24% and mAP by 2.73%, while maintaining high efficiency. On ASC, BEV-SA demonstrates remarkable generalization to rare and irregular objects.
KW - 3D object detection
KW - Autonomous driving
KW - Depthwise separable convolutions
KW - LiDAR-camera fusion
KW - Structured state space duality
UR - https://www.scopus.com/pages/publications/105040813714
U2 - 10.1016/j.eswa.2026.133029
DO - 10.1016/j.eswa.2026.133029
M3 - Article
AN - SCOPUS:105040813714
SN - 0957-4174
VL - 330
JO - Expert Systems with Applications
JF - Expert Systems with Applications
M1 - 133029
ER -