TY - JOUR
T1 - BEVMamba
T2 - Time Sequence Dense Bird’s-Eye-View Perception Modeling With State Space Model
AU - Liu, Xiao
AU - Zhong, Jiaru
AU - Sun, Chao
N1 - Publisher Copyright:
© 2000-2011 IEEE.
PY - 2025
Y1 - 2025
N2 - BEV-based 3D perception with multi-frame images input is crucial for autonomous driving. However, current methods for temporal BEV perception fail to fully utilize long sequence features because of local fusion or high complexity. Recently, Mamba, a powerful temporal modeling network with linear complexity, has shown exceptional performance in various 2D vision tasks, but its application to 3D perception tasks remains unexplored. Therefore, this paper proposes a general BEV perception backbone named BEVMamba, which is the first work to leverage State Space Model for 3D perception. Built upon the BEVFormer, to adapt Mamba for 3D perception we first add Hybrid Positional Encoding to the BEV features, enabling the networks to be aware of their spatial-temporal position. In the Temporal SSM block, the proposed 3D Factorized Scan ensures that historical BEV features are enriched with global temporal-spatial information. Subsequently, the Spatial-Temporal Corridor Fusion aggregates all BEV features in a physically meaningful manner, achieving precise feature fusion. The reliable BEV features obtained by BEVMamba are used for various perception tasks, including 3D object detection and 3D occupancy prediction. Results on the nuScenes and Occ-3D nuScenes datasets show that BEVMamba outperforms its baseline BEVFormer in both dense and sparse perception tasks and demonstrates competitive performance compared to other methods, highlighting the potential of Mamba in 3D perception tasks.
AB - BEV-based 3D perception with multi-frame images input is crucial for autonomous driving. However, current methods for temporal BEV perception fail to fully utilize long sequence features because of local fusion or high complexity. Recently, Mamba, a powerful temporal modeling network with linear complexity, has shown exceptional performance in various 2D vision tasks, but its application to 3D perception tasks remains unexplored. Therefore, this paper proposes a general BEV perception backbone named BEVMamba, which is the first work to leverage State Space Model for 3D perception. Built upon the BEVFormer, to adapt Mamba for 3D perception we first add Hybrid Positional Encoding to the BEV features, enabling the networks to be aware of their spatial-temporal position. In the Temporal SSM block, the proposed 3D Factorized Scan ensures that historical BEV features are enriched with global temporal-spatial information. Subsequently, the Spatial-Temporal Corridor Fusion aggregates all BEV features in a physically meaningful manner, achieving precise feature fusion. The reliable BEV features obtained by BEVMamba are used for various perception tasks, including 3D object detection and 3D occupancy prediction. Results on the nuScenes and Occ-3D nuScenes datasets show that BEVMamba outperforms its baseline BEVFormer in both dense and sparse perception tasks and demonstrates competitive performance compared to other methods, highlighting the potential of Mamba in 3D perception tasks.
KW - 3d object detection
KW - 3d occupancy prediction
KW - Autonomous driving
KW - bird’s-eye-view perception
KW - state space model
KW - time sequence modeling
UR - https://www.scopus.com/pages/publications/105009482055
U2 - 10.1109/TITS.2025.3569706
DO - 10.1109/TITS.2025.3569706
M3 - Article
AN - SCOPUS:105009482055
SN - 1524-9050
VL - 26
SP - 13466
EP - 13476
JO - IEEE Transactions on Intelligent Transportation Systems
JF - IEEE Transactions on Intelligent Transportation Systems
IS - 9
ER -