Skip to main navigation Skip to search Skip to main content

3D-MLV: Single Stage 3D Visual Grounding Using Multi-scale Local Voting

  • Qi A
  • , Sanyuan Zhao*
  • , Ken Yang
  • *Corresponding author for this work
  • Beijing Institute of Technology
  • Space Star Technology CO.LTD

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

3D Visual Grounding (3DVG) involves identifying corresponding objects within a three-dimensional point cloud based on natural language descriptions. Most existing approaches focus on two-stage methods, but their performance is heavily dependent on the quality of the object detector. In contrast, single-stage methods directly perform cross-modal inference from the point cloud, bypassing the need for object detectors and preserving surrounding environmental information during the point-cloud filtering process. However, single-stage methods remain relatively underexplored. The main challenges include: (1) the difficulty in aligning point cloud features with language features across modalities, and (2) the inability of traditional Transformer structures to effectively model the local relationships between objects of different sizes within the 3D scene. To address these challenges, we propose 3D Multi-scale Local Voting (3D-MLV), a single-stage 3DVG method that employs a multi-scale local voting mechanism. This method leverages the encoder of a 3D object detector for deep feature encoding. To achieve effective cross-modal alignment, we design an optimization framework that incrementally incorporates language information into the point cloud feature vector. Additionally, we introduce a Transformer-based multi-scale local voting mechanism for seed point selection. Unlike traditional global attention, this mechanism focuses on local information around key points and encodes multi-scale contextual features through multi-head attention. Experimental results demonstrate that the 3D-MLV approach significantly enhances the performance of single-stage 3DVG.

Original languageEnglish
Title of host publicationComputer Science and Education. AI Technology Frontiers - 19th International Conference, ICCSE 2025, Proceedings
EditorsWenxing Hong, Binyue Cui, Yang Weng, Chao Li
PublisherSpringer Science and Business Media Deutschland GmbH
Pages266-286
Number of pages21
ISBN (Print)9789819572533
DOIs
Publication statusPublished - 2026
Event19th International Conference on Computer Science and Education, ICCSE 2025 - Osaka and Fukui, Japan
Duration: 19 Aug 202524 Aug 2025

Publication series

NameCommunications in Computer and Information Science
Volume2760 CCIS
ISSN (Print)1865-0929
ISSN (Electronic)1865-0937

Conference

Conference19th International Conference on Computer Science and Education, ICCSE 2025
Country/TerritoryJapan
CityOsaka and Fukui
Period19/08/2524/08/25

Keywords

  • 3D visual grounding
  • multi-scale local voting
  • point clouds comprehension

Fingerprint

Dive into the research topics of '3D-MLV: Single Stage 3D Visual Grounding Using Multi-scale Local Voting'. Together they form a unique fingerprint.

Cite this