CrossRay3D: Geometry and Distribution Guidance for Efficient Multimodal 3D Detection

  • Huiming Yang
  • , Wenzhuo Liu
  • , Yicheng Qiao
  • , Lei Yang
  • , Xianzhu Zeng
  • , Li Wang
  • , Zhiwei Li*
  • , Zijian Zeng
  • , Zhiying Jiang*
  • , Huaping Liu
  • , Kunfeng Wang
  • *Corresponding author for this work

Research output: Contribution to journalArticlepeer-review

Abstract

The sparse cross-modality detector offers more advantages than its counterpart, the Bird’s-Eye-View (BEV) detector, particularly in terms of adaptability for downstream tasks and computational cost savings. However, existing sparse detectors overlook the quality of token representation, leaving it with a sub-optimal foreground quality and limited performance. In this paper, we identify that the geometric structure preserved and the class distribution are the key to improving the performance of the sparse detector, and propose a Sparse Selector (SS). The core module of SS is Ray-Aware Supervision (RAS), which preserves rich geometric information during the training stage, and Class-Balanced Supervision, which adaptively reweights the salience of class semantics, ensuring that tokens associated with small objects are retained during token sampling. Thereby, outperforming other sparse multi-modal detectors in the representation of tokens. Additionally, we design Ray Positional Encoding (Ray PE) to address the distribution differences between the LiDAR modality and the image. Finally, we integrate the aforementioned module into an end-to-end sparse multi-modality detector, dubbed CrossRay3D. Experiments show that, on the challenging nuScenes benchmark, CrossRay3D achieves state-of-the-art performance with 72.4% mAP and 74.7% NDS, while running 1.84x faster than other leading methods. Moreover, CrossRay3D demonstrates strong robustness even in scenarios where LiDAR or camera data are partially or entirely missing.

Original languageEnglish
JournalIEEE Transactions on Intelligent Transportation Systems
DOIs
Publication statusAccepted/In press - 2026
Externally publishedYes

Keywords

  • 3D object detection
  • Computer vision
  • sparse detector

Fingerprint

Dive into the research topics of 'CrossRay3D: Geometry and Distribution Guidance for Efficient Multimodal 3D Detection'. Together they form a unique fingerprint.

Cite this