TY - JOUR
T1 - DINO-CoDT
T2 - Multi-Class Collaborative Detection and VFM-Guided Tracking
AU - He, Xunjie
AU - Lee, Christina Dao Wen
AU - Wang, Meiling
AU - Yuan, Chengran
AU - Huang, Zefan
AU - Yue, Yufeng
AU - Ang, Marcelo H.
N1 - Publisher Copyright:
© 2000-2011 IEEE.
PY - 2026
Y1 - 2026
N2 - Collaborative perception is critical for advancing environmental understanding by extending perceptual coverage and enhancing resilience to sensor failures. However, most existing works (both 3D detection and tracking) focus primarily on the vehicle category, lacking robust multi-class solutions. This gap restricts their applicability in real-world scenarios, which involve a diverse range of objects with distinct appearances and motion characteristics. To overcome these limitations, we propose a multi-class collaborative detection and tracking framework tailored for diverse road users. We first present a detector with a global spatial attention fusion (GSAF) module, enhancing local-global feature learning for objects of varying sizes. Next, we introduce a tracklet re-identification (REID) module that leverages visual semantics with a vision foundation model to effectively reduce ID switch errors, in cases of erroneous mismatches involving small objects such as pedestrians. We further design a velocity-based adaptive tracklet management (VATM) module that adjusts the tracking interval dynamically based on object motion. Extensive experiments on the V2X-Real and OPV2V datasets demonstrate that our approach consistently outperforms state-of-the-art methods. In particular, on the pedestrian class of the V2X-Real dataset, it achieves a 1.9% improvement in mAP@0.3 for detection and a 0.79% gain in sAMOTA for tracking.
AB - Collaborative perception is critical for advancing environmental understanding by extending perceptual coverage and enhancing resilience to sensor failures. However, most existing works (both 3D detection and tracking) focus primarily on the vehicle category, lacking robust multi-class solutions. This gap restricts their applicability in real-world scenarios, which involve a diverse range of objects with distinct appearances and motion characteristics. To overcome these limitations, we propose a multi-class collaborative detection and tracking framework tailored for diverse road users. We first present a detector with a global spatial attention fusion (GSAF) module, enhancing local-global feature learning for objects of varying sizes. Next, we introduce a tracklet re-identification (REID) module that leverages visual semantics with a vision foundation model to effectively reduce ID switch errors, in cases of erroneous mismatches involving small objects such as pedestrians. We further design a velocity-based adaptive tracklet management (VATM) module that adjusts the tracking interval dynamically based on object motion. Extensive experiments on the V2X-Real and OPV2V datasets demonstrate that our approach consistently outperforms state-of-the-art methods. In particular, on the pedestrian class of the V2X-Real dataset, it achieves a 1.9% improvement in mAP@0.3 for detection and a 0.79% gain in sAMOTA for tracking.
KW - Collaborative perception
KW - multi-class perception
KW - multi-object tracking
KW - pedestrian REID
UR - https://www.scopus.com/pages/publications/105041973823
U2 - 10.1109/TITS.2026.3697384
DO - 10.1109/TITS.2026.3697384
M3 - Article
AN - SCOPUS:105041973823
SN - 1524-9050
JO - IEEE Transactions on Intelligent Transportation Systems
JF - IEEE Transactions on Intelligent Transportation Systems
ER -