跳到主要导航 跳到搜索 跳到主要内容

Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation

  • Ziyu Zhu
  • , Xilin Wang
  • , Yixuan Li
  • , Zhuofan Zhang
  • , Xiaojian Ma
  • , Yixin Chen
  • , Baoxiong Jia
  • , Wei Liang
  • , Qian Yu
  • , Zhidong Deng
  • , Siyuan Huang
  • , Qing Li
  • Tsinghua University
  • BIGAI
  • Beihang University
  • Beijing Institute of Technology

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Embodied scene understanding requires not only comprehending visual-spatial information that has been observed but also determining where to explore next in the 3D physical world. Existing 3D Vision-Language (3D-VL) models primarily focus on grounding objects in static observations from 3D reconstruction, such as meshes and point clouds, but lack the ability to actively perceive and explore their environment. To address this limitation, we introduce Move to Understand (MTU3D), a unified framework that integrates active perception with 3D vision-language learning, enabling embodied agents to effectively explore and understand their environment. This is achieved by three key innovations: 1) Online query-based representation learning, enabling direct spatial memory construction from RGB-D frames, eliminating the need for explicit 3D reconstruction. 2) A unified objective for grounding and ex-ploring, which represents unexplored locations as frontier queries and jointly optimizes object grounding and frontier selection. 3) End-to-end trajectory learning that combines Vision-Language-Exploration pre-training over a million diverse trajectories collected from both simulated and realworld RGB-D sequences. Extensive evaluations across various embodied navigation and question-answering benchmarks show that MTU3D outperforms state-of-the-art reinforcement learning and modular navigation approaches by 14%, 23%, 9%, and 2% in success rate on HM3D-OVON, GOAT-Bench, SG3D, and A-EQA, respectively. MTU3D's versatility enables navigation using diverse input modalities, including categories, language descriptions, and reference images. The deployment on a real robot demonstrates MTU3D's effectiveness in handling real-world data. These findings highlight the importance of bridging visual grounding and exploration for embodied intelligence.

源语言英语
主期刊名Proceedings - 2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025
出版商Institute of Electrical and Electronics Engineers Inc.
8120-8132
页数13
ISBN(电子版)9798331587758
DOI
出版状态已出版 - 2025
活动2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025 - Honolulu, 美国
期限: 19 10月 202523 10月 2025

丛书

姓名Proceedings of the IEEE International Conference on Computer Vision
ISSN(印刷版)1550-5499
ISSN(电子版)2380-7504

会议

会议2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025
国家/地区美国
Honolulu
时期19/10/2523/10/25

学术指纹

探究 'Move to Understand a 3D Scene: Bridging Visual Grounding and Exploration for Efficient and Versatile Embodied Navigation' 的科研主题。它们共同构成独一无二的学术指纹。

引用此