Skip to main navigation Skip to search Skip to main content

Self-Supervised Monocular Visual Odometry Based on Multi-View Spatio-Temporal Feature Fusion

  • Jiaqi Liu
  • , Zhuoling Xiao*
  • , Bo Yan
  • *Corresponding author for this work
  • University of Electronic Science and Technology of China

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Learning-based visual odometry (VO) estimates the ego-motion of a camera by leveraging consistent pixel movements between pairs of consecutive image frames. Unlike most existing VOs that only concentrate on a single imaging plane, our method, MultiSTVO, utilizes the physical prior of camera motions to mine and fuse enhanced temporal motion features from multiple views. To simultaneously focus on pixel movements under various views, the Multi-View Spatio-Temporal Feature Enhancement is proposed to fully extract features sensitive to different observation views. In addition, the graph attention network-based Pose Graph Attention Refinement is also designed to achieve refined selection and fusion. This meticulous integration process helps to extract substantial and robust motion information to realize the full utilization of motion features. Experiments on KITTI / Málaga demonstrate the promising performance of MultiSTVO. Compared with state-of-the-art other methods, it achieves improvements of 28.9% and 43.1% in translation and rotation evaluations, respectively.

Original languageEnglish
Title of host publicationISCAS 2026 - 2026 IEEE International Symposium on Circuits and Systems
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages2223-2227
Number of pages5
ISBN (Electronic)9798331577698
DOIs
Publication statusPublished - 2026
Externally publishedYes
Event2026 IEEE International Symposium on Circuits and Systems, ISCAS 2026 - Shanghai, China
Duration: 24 May 202627 May 2026

Publication series

NameProceedings - IEEE International Symposium on Circuits and Systems
ISSN (Print)0271-4310

Conference

Conference2026 IEEE International Symposium on Circuits and Systems, ISCAS 2026
Country/TerritoryChina
CityShanghai
Period24/05/2627/05/26

Keywords

  • graph attention networks
  • multi-view features
  • temporal information aggregation
  • Visual odometry

Fingerprint

Dive into the research topics of 'Self-Supervised Monocular Visual Odometry Based on Multi-View Spatio-Temporal Feature Fusion'. Together they form a unique fingerprint.

Cite this