TY - JOUR
T1 - TriQuery-BEV
T2 - Enhancing 3D Perception for Autonomous Driving with Temporal Query Filtering and Uncertainty-Aware Fusion
AU - Dong, Junyi
AU - Chen, Xuemei
AU - Liu, Zemin
N1 - Publisher Copyright:
© 2026 by the authors.
PY - 2026/5
Y1 - 2026/5
N2 - Existing BEV perception methods unify multi-view information in a bird’s-eye-view coordinate system, yet their performance in dynamic traffic scenes remains limited by three major error sources: depth-noise amplification during image-to-BEV lifting, representation discontinuity caused by time-varying occlusion and visibility, and temporal drift induced by recursive fusion of historical BEV features. To address these issues while preserving computational tractability, we propose TriQuery-BEV, a modular enhancement framework over BEVFormer that improves BEV query modeling from the perspectives of geometric ambiguity, occlusion robustness, and temporal consistency. The proposed framework integrates three components: Query Mask (QM) for structured regularization in the BEV query space, depth-modulated hybrid positional encoding (DM-HPE) for geometry-aware positional representation, and a Temporal Query Filter (TQF) for uncertainty-aware temporal fusion. Experiments on the nuScenes benchmark demonstrate consistent improvements over BEVFormer across different model scales. TriQuery-BEV improves the nuScenes detection score (NDS) and mean average precision (mAP) by 5.4%/6.4% under the Tiny (ResNet-50) setting and by 6.0%/6.5% under the Base (ResNet-101) setting. It also reduces key true-positive error metrics, including mean translation error (mATE) by 2.9%/6.5%, mean orientation error (mAOE) by 5.7%/8.3%, and mean velocity error (mAVE) by 7.3%/15.0% for Tiny/Base, respectively. Extensive ablations further verify the effectiveness of DM-HPE, TQF, and QM, confirming improved robustness, geometric accuracy, and temporal consistency in highly dynamic environments.
AB - Existing BEV perception methods unify multi-view information in a bird’s-eye-view coordinate system, yet their performance in dynamic traffic scenes remains limited by three major error sources: depth-noise amplification during image-to-BEV lifting, representation discontinuity caused by time-varying occlusion and visibility, and temporal drift induced by recursive fusion of historical BEV features. To address these issues while preserving computational tractability, we propose TriQuery-BEV, a modular enhancement framework over BEVFormer that improves BEV query modeling from the perspectives of geometric ambiguity, occlusion robustness, and temporal consistency. The proposed framework integrates three components: Query Mask (QM) for structured regularization in the BEV query space, depth-modulated hybrid positional encoding (DM-HPE) for geometry-aware positional representation, and a Temporal Query Filter (TQF) for uncertainty-aware temporal fusion. Experiments on the nuScenes benchmark demonstrate consistent improvements over BEVFormer across different model scales. TriQuery-BEV improves the nuScenes detection score (NDS) and mean average precision (mAP) by 5.4%/6.4% under the Tiny (ResNet-50) setting and by 6.0%/6.5% under the Base (ResNet-101) setting. It also reduces key true-positive error metrics, including mean translation error (mATE) by 2.9%/6.5%, mean orientation error (mAOE) by 5.7%/8.3%, and mean velocity error (mAVE) by 7.3%/15.0% for Tiny/Base, respectively. Extensive ablations further verify the effectiveness of DM-HPE, TQF, and QM, confirming improved robustness, geometric accuracy, and temporal consistency in highly dynamic environments.
KW - 3D perception
KW - bird’s-eye view perception
KW - temporal query fusion
UR - https://www.scopus.com/pages/publications/105040140619
U2 - 10.3390/s26102934
DO - 10.3390/s26102934
M3 - Article
C2 - 42197742
AN - SCOPUS:105040140619
SN - 1424-8220
VL - 26
JO - Sensors
JF - Sensors
IS - 10
M1 - 2934
ER -