Skip to main navigation Skip to search Skip to main content

Diffusion-Head VLA Model for Adaptive UAV Vision-Language Navigation

  • Beijing Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Visual-Language-Action (VLA) models demonstrate strong semantic understanding and action modeling capabilities in UAV Visual-Language Navigation (VLN) tasks, yet they perform ineffectively in complex scenarios, such as long-distance and multi-dimensional continuous motion tasks. To address this, we propose a Diffusion-Head VLA model with an Adaptive Action Policy (AAP), enabling continuous and fine-grained action generation. This strategy allows the model to leverage high-level semantic guidance from Vision-Language Models (VLMs) while producing continuous and fine-grained control via the diffusion head. The Adaptive Action Policy dynamically integrates the actions generated by the VLM and the diffusion-based action head, thereby improving the UAV's performance. The History-Aware Directional Perturbation (HADP) mechanism injects perturbations aligned with historical action trajectories, enhancing both action continuity and diversity, which leads to more stable and effective training. Experimental results compared with baselines show that our approach improves UAV navigation performance.

Original languageEnglish
Title of host publication38th Chinese Control and Decision Conference, CCDC 2026
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages6256-6261
Number of pages6
ISBN (Electronic)9798331550707
DOIs
Publication statusPublished - 2026
Externally publishedYes
Event38th Chinese Control and Decision Conference, CCDC 2026 - Nanjing, China
Duration: 15 May 202618 May 2026

Publication series

Name38th Chinese Control and Decision Conference, CCDC 2026

Conference

Conference38th Chinese Control and Decision Conference, CCDC 2026
Country/TerritoryChina
CityNanjing
Period15/05/2618/05/26

Keywords

  • Adaptive Action Policy
  • Aerial Vision-Language Navigation
  • Diffusion-Head
  • Visual-Language-Action model

Fingerprint

Dive into the research topics of 'Diffusion-Head VLA Model for Adaptive UAV Vision-Language Navigation'. Together they form a unique fingerprint.

Cite this