TY - GEN
T1 - Self-Supervised Monocular Visual Odometry Based on Multi-View Spatio-Temporal Feature Fusion
AU - Liu, Jiaqi
AU - Xiao, Zhuoling
AU - Yan, Bo
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - Learning-based visual odometry (VO) estimates the ego-motion of a camera by leveraging consistent pixel movements between pairs of consecutive image frames. Unlike most existing VOs that only concentrate on a single imaging plane, our method, MultiSTVO, utilizes the physical prior of camera motions to mine and fuse enhanced temporal motion features from multiple views. To simultaneously focus on pixel movements under various views, the Multi-View Spatio-Temporal Feature Enhancement is proposed to fully extract features sensitive to different observation views. In addition, the graph attention network-based Pose Graph Attention Refinement is also designed to achieve refined selection and fusion. This meticulous integration process helps to extract substantial and robust motion information to realize the full utilization of motion features. Experiments on KITTI / Málaga demonstrate the promising performance of MultiSTVO. Compared with state-of-the-art other methods, it achieves improvements of 28.9% and 43.1% in translation and rotation evaluations, respectively.
AB - Learning-based visual odometry (VO) estimates the ego-motion of a camera by leveraging consistent pixel movements between pairs of consecutive image frames. Unlike most existing VOs that only concentrate on a single imaging plane, our method, MultiSTVO, utilizes the physical prior of camera motions to mine and fuse enhanced temporal motion features from multiple views. To simultaneously focus on pixel movements under various views, the Multi-View Spatio-Temporal Feature Enhancement is proposed to fully extract features sensitive to different observation views. In addition, the graph attention network-based Pose Graph Attention Refinement is also designed to achieve refined selection and fusion. This meticulous integration process helps to extract substantial and robust motion information to realize the full utilization of motion features. Experiments on KITTI / Málaga demonstrate the promising performance of MultiSTVO. Compared with state-of-the-art other methods, it achieves improvements of 28.9% and 43.1% in translation and rotation evaluations, respectively.
KW - graph attention networks
KW - multi-view features
KW - temporal information aggregation
KW - Visual odometry
UR - https://www.scopus.com/pages/publications/105043431067
U2 - 10.1109/ISCAS66217.2026.11561951
DO - 10.1109/ISCAS66217.2026.11561951
M3 - Conference contribution
AN - SCOPUS:105043431067
T3 - Proceedings - IEEE International Symposium on Circuits and Systems
SP - 2223
EP - 2227
BT - ISCAS 2026 - 2026 IEEE International Symposium on Circuits and Systems
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2026 IEEE International Symposium on Circuits and Systems, ISCAS 2026
Y2 - 24 May 2026 through 27 May 2026
ER -