Skip to main navigation Skip to search Skip to main content

E2Depth: Efficient Self-Supervised Surround-View Depth Estimation With Explicit Geometric Enhancement

  • Sheng Zhang
  • , Juan Li*
  • , Chang Liu
  • , Chang Liu
  • , Jie Li
  • , Dongxiao Yang
  • *Corresponding author for this work
  • Beijing Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Surround-view depth estimation is critical for robotic perception and autonomous driving. Existing methods rely on implicit cross-view feature learning, which incurs a high computational cost. In this letter, we propose E^2Depth, an efficient self-supervised surround-view depth estimation framework that explicitly leverages geometric and signal priors at each pipeline stage. First, hierarchical volumetric fusion is performed to align multi-level features across cameras with controlled memory consumption. Second, wavelet-domain edge enhancement is introduced to recover sharper depth boundaries without external supervision. Finally, an explicit pose estimation network guided by noisy motion priors is designed to stabilize training and improve depth scale reliability. Extensive experiments on the DDAD and nuScenes benchmarks demonstrate that E^2Depth achieves a favorable trade-off between accuracy and efficiency, supporting practical deployment on real-world platforms.

Original languageEnglish
Pages (from-to)9080-9087
Number of pages8
JournalIEEE Robotics and Automation Letters
Volume11
Issue number8
DOIs
Publication statusPublished - 2026
Externally publishedYes

Keywords

  • Surround-view system
  • depth estimation
  • self-supervised learning

Fingerprint

Dive into the research topics of 'E2Depth: Efficient Self-Supervised Surround-View Depth Estimation With Explicit Geometric Enhancement'. Together they form a unique fingerprint.

Cite this