Skip to main navigation Skip to search Skip to main content

Global context alignment and separable fusion for generalizable multi-modal 3D object detection

  • Yingjuan Tang
  • , Hongwen He*
  • , Jingda Wu
  • , Yongpeng Shen
  • , Yong Wang
  • , Yifan Wu
  • *Corresponding author for this work
  • Zhengzhou University of Light Industry
  • Beijing Institute of Technology
  • The University of Hong Kong

Research output: Contribution to journalArticlepeer-review

Abstract

Fusion-based 3D object detection is critical for autonomous driving. However, existing multimodal fusion methods often suffer from an insufficient receptive field, inefficient cross-modal interaction, and weak generalization to rare or irregular objects. We present BEV-SA, a novel bird’s eye view–based fusion framework that overcomes these limitations through two key modules. The state-space BEV interaction (SSB) module serializes BEV features along a Hilbert curve and models them using a structured state space duality (SSD) mechanism, enabling global context modeling, precise modality alignment, and linear computational complexity. The accelerated separable fusion (ASF) module employs depthwise separable convolutions to efficiently merge aligned BEV features, reducing inference time by 12.5% compared to conventional fusion without sacrificing accuracy. We also introduce the anomalous-shaped commercial vehicle (ASC) dataset, a benchmark focusing on large and irregular commercial vehicles in challenging environments. Experiments on the nuScenes benchmark show that BEV-SA achieves state-of-the-art performance, improving NDS by 0.24% and mAP by 2.73%, while maintaining high efficiency. On ASC, BEV-SA demonstrates remarkable generalization to rare and irregular objects.

Original languageEnglish
Article number133029
JournalExpert Systems with Applications
Volume330
DOIs
Publication statusPublished - 1 Dec 2026

Keywords

  • 3D object detection
  • Autonomous driving
  • Depthwise separable convolutions
  • LiDAR-camera fusion
  • Structured state space duality

Fingerprint

Dive into the research topics of 'Global context alignment and separable fusion for generalizable multi-modal 3D object detection'. Together they form a unique fingerprint.

Cite this