TY - GEN
T1 - Predictive Beamforming in Low-Altitude Wireless Networks
T2 - 2026 IEEE International Conference on Communications, ICC 2026
AU - Zhao, Xiaotong
AU - Cui, Yuanhao
AU - Yuan, Weijie
AU - Jia, Ziye
AU - Liu, Heng
AU - Xing, Chengwen
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - Accurate beam prediction is essential for maintaining reliable links and high spectral efficiency in dynamic Low-Altitude Wireless Networks (LAWN). However, existing approaches often fail to capture the deep correlations across heterogeneous sensing modalities, limiting their adaptability in complex three-dimensional (3D) environments. To overcome these challenges, we propose a multi-modal predictive beamforming method based on a cross-attention fusion mechanism that jointly leverages visual and structured sensor data. The proposed model utilizes a Convolutional Neural Network (CNN) to learn multi-scale spatial feature hierarchies from visual images and a Transformer encoder to capture cross-dimensional dependencies within sensor data. Then, a cross-attention fusion module is introduced to integrate complementary information between the two modalities, generating a unified and discriminative representation for accurate beam prediction. Through experimental evaluations conducted on a real-world dataset, our method reaches 79.7% Top-1 accuracy and 99.3% Top-3 accuracy, surpassing the baseline method by 4.4%-23.2% across Top-1 to Top-5 metrics. These results verify that multi-modal cross-attention fusion is effective for intelligent beam selection in dynamic LAWN.
AB - Accurate beam prediction is essential for maintaining reliable links and high spectral efficiency in dynamic Low-Altitude Wireless Networks (LAWN). However, existing approaches often fail to capture the deep correlations across heterogeneous sensing modalities, limiting their adaptability in complex three-dimensional (3D) environments. To overcome these challenges, we propose a multi-modal predictive beamforming method based on a cross-attention fusion mechanism that jointly leverages visual and structured sensor data. The proposed model utilizes a Convolutional Neural Network (CNN) to learn multi-scale spatial feature hierarchies from visual images and a Transformer encoder to capture cross-dimensional dependencies within sensor data. Then, a cross-attention fusion module is introduced to integrate complementary information between the two modalities, generating a unified and discriminative representation for accurate beam prediction. Through experimental evaluations conducted on a real-world dataset, our method reaches 79.7% Top-1 accuracy and 99.3% Top-3 accuracy, surpassing the baseline method by 4.4%-23.2% across Top-1 to Top-5 metrics. These results verify that multi-modal cross-attention fusion is effective for intelligent beam selection in dynamic LAWN.
KW - CNN
KW - Cross-Attention
KW - Low-Altitude Wireless Networks
KW - Predictive Beamforming
KW - Transformer
UR - https://www.scopus.com/pages/publications/105045412671
U2 - 10.1109/ICC59461.2026.11588255
DO - 10.1109/ICC59461.2026.11588255
M3 - Conference contribution
AN - SCOPUS:105045412671
T3 - IEEE International Conference on Communications
BT - ICC 2026 - IEEE International Conference on Communications, Proceedings
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 24 May 2026 through 28 May 2026
ER -