跳到主要导航 跳到搜索 跳到主要内容

Diffusion-Head VLA Model for Adaptive UAV Vision-Language Navigation

  • Beijing Institute of Technology

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Visual-Language-Action (VLA) models demonstrate strong semantic understanding and action modeling capabilities in UAV Visual-Language Navigation (VLN) tasks, yet they perform ineffectively in complex scenarios, such as long-distance and multi-dimensional continuous motion tasks. To address this, we propose a Diffusion-Head VLA model with an Adaptive Action Policy (AAP), enabling continuous and fine-grained action generation. This strategy allows the model to leverage high-level semantic guidance from Vision-Language Models (VLMs) while producing continuous and fine-grained control via the diffusion head. The Adaptive Action Policy dynamically integrates the actions generated by the VLM and the diffusion-based action head, thereby improving the UAV's performance. The History-Aware Directional Perturbation (HADP) mechanism injects perturbations aligned with historical action trajectories, enhancing both action continuity and diversity, which leads to more stable and effective training. Experimental results compared with baselines show that our approach improves UAV navigation performance.

源语言英语
主期刊名38th Chinese Control and Decision Conference, CCDC 2026
出版商Institute of Electrical and Electronics Engineers Inc.
6256-6261
页数6
ISBN(电子版)9798331550707
DOI
出版状态已出版 - 2026
已对外发布
活动38th Chinese Control and Decision Conference, CCDC 2026 - Nanjing, 中国
期限: 15 5月 202618 5月 2026

丛书

姓名38th Chinese Control and Decision Conference, CCDC 2026

会议

会议38th Chinese Control and Decision Conference, CCDC 2026
国家/地区中国
Nanjing
时期15/05/2618/05/26

学术指纹

探究 'Diffusion-Head VLA Model for Adaptive UAV Vision-Language Navigation' 的科研主题。它们共同构成独一无二的学术指纹。

引用此