Skip to main navigation Skip to search Skip to main content

DINO-CoDT: Multi-Class Collaborative Detection and VFM-Guided Tracking

  • Xunjie He
  • , Christina Dao Wen Lee
  • , Meiling Wang
  • , Chengran Yuan
  • , Zefan Huang
  • , Yufeng Yue*
  • , Marcelo H. Ang
  • *Corresponding author for this work
  • Beijing Institute of Technology
  • National University of Singapore

Research output: Contribution to journalArticlepeer-review

Abstract

Collaborative perception is critical for advancing environmental understanding by extending perceptual coverage and enhancing resilience to sensor failures. However, most existing works (both 3D detection and tracking) focus primarily on the vehicle category, lacking robust multi-class solutions. This gap restricts their applicability in real-world scenarios, which involve a diverse range of objects with distinct appearances and motion characteristics. To overcome these limitations, we propose a multi-class collaborative detection and tracking framework tailored for diverse road users. We first present a detector with a global spatial attention fusion (GSAF) module, enhancing local-global feature learning for objects of varying sizes. Next, we introduce a tracklet re-identification (REID) module that leverages visual semantics with a vision foundation model to effectively reduce ID switch errors, in cases of erroneous mismatches involving small objects such as pedestrians. We further design a velocity-based adaptive tracklet management (VATM) module that adjusts the tracking interval dynamically based on object motion. Extensive experiments on the V2X-Real and OPV2V datasets demonstrate that our approach consistently outperforms state-of-the-art methods. In particular, on the pedestrian class of the V2X-Real dataset, it achieves a 1.9% improvement in mAP@0.3 for detection and a 0.79% gain in sAMOTA for tracking.

Original languageEnglish
JournalIEEE Transactions on Intelligent Transportation Systems
DOIs
Publication statusAccepted/In press - 2026
Externally publishedYes

Keywords

  • Collaborative perception
  • multi-class perception
  • multi-object tracking
  • pedestrian REID

Fingerprint

Dive into the research topics of 'DINO-CoDT: Multi-Class Collaborative Detection and VFM-Guided Tracking'. Together they form a unique fingerprint.

Cite this