跳到主要导航 跳到搜索 跳到主要内容

3D-MLV: Single Stage 3D Visual Grounding Using Multi-scale Local Voting

  • Beijing Institute of Technology
  • Space Star Technology CO.LTD

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

3D Visual Grounding (3DVG) involves identifying corresponding objects within a three-dimensional point cloud based on natural language descriptions. Most existing approaches focus on two-stage methods, but their performance is heavily dependent on the quality of the object detector. In contrast, single-stage methods directly perform cross-modal inference from the point cloud, bypassing the need for object detectors and preserving surrounding environmental information during the point-cloud filtering process. However, single-stage methods remain relatively underexplored. The main challenges include: (1) the difficulty in aligning point cloud features with language features across modalities, and (2) the inability of traditional Transformer structures to effectively model the local relationships between objects of different sizes within the 3D scene. To address these challenges, we propose 3D Multi-scale Local Voting (3D-MLV), a single-stage 3DVG method that employs a multi-scale local voting mechanism. This method leverages the encoder of a 3D object detector for deep feature encoding. To achieve effective cross-modal alignment, we design an optimization framework that incrementally incorporates language information into the point cloud feature vector. Additionally, we introduce a Transformer-based multi-scale local voting mechanism for seed point selection. Unlike traditional global attention, this mechanism focuses on local information around key points and encodes multi-scale contextual features through multi-head attention. Experimental results demonstrate that the 3D-MLV approach significantly enhances the performance of single-stage 3DVG.

源语言英语
主期刊名Computer Science and Education. AI Technology Frontiers - 19th International Conference, ICCSE 2025, Proceedings
编辑Wenxing Hong, Binyue Cui, Yang Weng, Chao Li
出版商Springer Science and Business Media Deutschland GmbH
266-286
页数21
ISBN(印刷版)9789819572533
DOI
出版状态已出版 - 2026
活动19th International Conference on Computer Science and Education, ICCSE 2025 - Osaka and Fukui, 日本
期限: 19 8月 202524 8月 2025

出版系列

姓名Communications in Computer and Information Science
2760 CCIS
ISSN(印刷版)1865-0929
ISSN(电子版)1865-0937

会议

会议19th International Conference on Computer Science and Education, ICCSE 2025
国家/地区日本
Osaka and Fukui
时期19/08/2524/08/25

指纹

探究 '3D-MLV: Single Stage 3D Visual Grounding Using Multi-scale Local Voting' 的科研主题。它们共同构成独一无二的指纹。

引用此