TY - GEN
T1 - Multi-task Perception Model for Unmanned Systems in Urban Environments
AU - Bai, Chen
AU - Wang, Shaojie
AU - Zhang, Jianye
AU - Wu, Weichao
N1 - Publisher Copyright:
© Beijing HIWING Scientific and Technological Information Institute 2026.
PY - 2026
Y1 - 2026
N2 - In complex dynamic environments for unmanned systems, achieving efficient and robust multi-task perception is crucial for enhancing environmental understanding and decision-making capabilities. This paper presents a unified multi-task perception framework capable of simultaneously addressing five key perception tasks: depth estimation, pose estimation, optical flow estimation, motion segmentation, and semantic segmentation. The framework employs a shared encoder architecture to improve inter-task synergy through unified feature representations, while integrating both optical flow-based self-supervised depth estimation for dynamic scenes and a Mask2Former-based semantic segmentation model to enhance geometric perception and semantic understanding. By leveraging multi-task collaborative learning, our approach combines spatiotemporal consistency constraints with global semantic information from segmentation to jointly optimize depth estimation and motion segmentation in dynamic scenarios.
AB - In complex dynamic environments for unmanned systems, achieving efficient and robust multi-task perception is crucial for enhancing environmental understanding and decision-making capabilities. This paper presents a unified multi-task perception framework capable of simultaneously addressing five key perception tasks: depth estimation, pose estimation, optical flow estimation, motion segmentation, and semantic segmentation. The framework employs a shared encoder architecture to improve inter-task synergy through unified feature representations, while integrating both optical flow-based self-supervised depth estimation for dynamic scenes and a Mask2Former-based semantic segmentation model to enhance geometric perception and semantic understanding. By leveraging multi-task collaborative learning, our approach combines spatiotemporal consistency constraints with global semantic information from segmentation to jointly optimize depth estimation and motion segmentation in dynamic scenarios.
KW - Depth estimation
KW - Environment sensing
KW - Multi task
KW - Semantic segmentation
UR - https://www.scopus.com/pages/publications/105043000262
U2 - 10.1007/978-981-95-7660-9_43
DO - 10.1007/978-981-95-7660-9_43
M3 - Conference contribution
AN - SCOPUS:105043000262
SN - 9789819576593
T3 - Lecture Notes in Electrical Engineering
SP - 467
EP - 476
BT - Proceedings of 5th 2025 International Conference on Autonomous Unmanned Systems, ICAUS - Volume 7
A2 - Xie, Shaorong
A2 - Niu, Yifeng
A2 - Fu, Wenxing
A2 - Qu, Yi
PB - Springer Science and Business Media Deutschland GmbH
T2 - 5th International Conference on Autonomous Unmanned Systems, ICAUS 2025
Y2 - 17 October 2025 through 19 October 2025
ER -