Abstract
Collaborative perception is critical for advancing environmental understanding by extending perceptual coverage and enhancing resilience to sensor failures. However, most existing works (both 3D detection and tracking) focus primarily on the vehicle category, lacking robust multi-class solutions. This gap restricts their applicability in real-world scenarios, which involve a diverse range of objects with distinct appearances and motion characteristics. To overcome these limitations, we propose a multi-class collaborative detection and tracking framework tailored for diverse road users. We first present a detector with a global spatial attention fusion (GSAF) module, enhancing local-global feature learning for objects of varying sizes. Next, we introduce a tracklet re-identification (REID) module that leverages visual semantics with a vision foundation model to effectively reduce ID switch errors, in cases of erroneous mismatches involving small objects such as pedestrians. We further design a velocity-based adaptive tracklet management (VATM) module that adjusts the tracking interval dynamically based on object motion. Extensive experiments on the V2X-Real and OPV2V datasets demonstrate that our approach consistently outperforms state-of-the-art methods. In particular, on the pedestrian class of the V2X-Real dataset, it achieves a 1.9% improvement in mAP@0.3 for detection and a 0.79% gain in sAMOTA for tracking.
| Original language | English |
|---|---|
| Journal | IEEE Transactions on Intelligent Transportation Systems |
| DOIs | |
| Publication status | Accepted/In press - 2026 |
| Externally published | Yes |
Keywords
- Collaborative perception
- multi-class perception
- multi-object tracking
- pedestrian REID
Fingerprint
Dive into the research topics of 'DINO-CoDT: Multi-Class Collaborative Detection and VFM-Guided Tracking'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver