跳到主要导航 跳到搜索 跳到主要内容

Global context alignment and separable fusion for generalizable multi-modal 3D object detection

  • Yingjuan Tang
  • , Hongwen He*
  • , Jingda Wu
  • , Yongpeng Shen
  • , Yong Wang
  • , Yifan Wu
  • *此作品的通讯作者
  • Zhengzhou University of Light Industry
  • Beijing Institute of Technology
  • The University of Hong Kong

科研成果: 期刊稿件文章同行评审

摘要

Fusion-based 3D object detection is critical for autonomous driving. However, existing multimodal fusion methods often suffer from an insufficient receptive field, inefficient cross-modal interaction, and weak generalization to rare or irregular objects. We present BEV-SA, a novel bird’s eye view–based fusion framework that overcomes these limitations through two key modules. The state-space BEV interaction (SSB) module serializes BEV features along a Hilbert curve and models them using a structured state space duality (SSD) mechanism, enabling global context modeling, precise modality alignment, and linear computational complexity. The accelerated separable fusion (ASF) module employs depthwise separable convolutions to efficiently merge aligned BEV features, reducing inference time by 12.5% compared to conventional fusion without sacrificing accuracy. We also introduce the anomalous-shaped commercial vehicle (ASC) dataset, a benchmark focusing on large and irregular commercial vehicles in challenging environments. Experiments on the nuScenes benchmark show that BEV-SA achieves state-of-the-art performance, improving NDS by 0.24% and mAP by 2.73%, while maintaining high efficiency. On ASC, BEV-SA demonstrates remarkable generalization to rare and irregular objects.

源语言英语
文章编号133029
期刊Expert Systems with Applications
330
DOI
出版状态已出版 - 1 12月 2026

指纹

探究 'Global context alignment and separable fusion for generalizable multi-modal 3D object detection' 的科研主题。它们共同构成独一无二的指纹。

引用此