TY - GEN
T1 - Diffusion-Head VLA Model for Adaptive UAV Vision-Language Navigation
AU - Lin, Tanhui
AU - Fei, Qing
AU - Geng, Qingbo
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - Visual-Language-Action (VLA) models demonstrate strong semantic understanding and action modeling capabilities in UAV Visual-Language Navigation (VLN) tasks, yet they perform ineffectively in complex scenarios, such as long-distance and multi-dimensional continuous motion tasks. To address this, we propose a Diffusion-Head VLA model with an Adaptive Action Policy (AAP), enabling continuous and fine-grained action generation. This strategy allows the model to leverage high-level semantic guidance from Vision-Language Models (VLMs) while producing continuous and fine-grained control via the diffusion head. The Adaptive Action Policy dynamically integrates the actions generated by the VLM and the diffusion-based action head, thereby improving the UAV's performance. The History-Aware Directional Perturbation (HADP) mechanism injects perturbations aligned with historical action trajectories, enhancing both action continuity and diversity, which leads to more stable and effective training. Experimental results compared with baselines show that our approach improves UAV navigation performance.
AB - Visual-Language-Action (VLA) models demonstrate strong semantic understanding and action modeling capabilities in UAV Visual-Language Navigation (VLN) tasks, yet they perform ineffectively in complex scenarios, such as long-distance and multi-dimensional continuous motion tasks. To address this, we propose a Diffusion-Head VLA model with an Adaptive Action Policy (AAP), enabling continuous and fine-grained action generation. This strategy allows the model to leverage high-level semantic guidance from Vision-Language Models (VLMs) while producing continuous and fine-grained control via the diffusion head. The Adaptive Action Policy dynamically integrates the actions generated by the VLM and the diffusion-based action head, thereby improving the UAV's performance. The History-Aware Directional Perturbation (HADP) mechanism injects perturbations aligned with historical action trajectories, enhancing both action continuity and diversity, which leads to more stable and effective training. Experimental results compared with baselines show that our approach improves UAV navigation performance.
KW - Adaptive Action Policy
KW - Aerial Vision-Language Navigation
KW - Diffusion-Head
KW - Visual-Language-Action model
UR - https://www.scopus.com/pages/publications/105043901991
U2 - 10.1109/CCDC69976.2026.11560596
DO - 10.1109/CCDC69976.2026.11560596
M3 - Conference contribution
AN - SCOPUS:105043901991
T3 - 38th Chinese Control and Decision Conference, CCDC 2026
SP - 6256
EP - 6261
BT - 38th Chinese Control and Decision Conference, CCDC 2026
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 38th Chinese Control and Decision Conference, CCDC 2026
Y2 - 15 May 2026 through 18 May 2026
ER -