Abstract
Surround-view depth estimation is critical for robotic perception and autonomous driving. Existing methods rely on implicit cross-view feature learning, which incurs a high computational cost. In this letter, we propose E^2Depth, an efficient self-supervised surround-view depth estimation framework that explicitly leverages geometric and signal priors at each pipeline stage. First, hierarchical volumetric fusion is performed to align multi-level features across cameras with controlled memory consumption. Second, wavelet-domain edge enhancement is introduced to recover sharper depth boundaries without external supervision. Finally, an explicit pose estimation network guided by noisy motion priors is designed to stabilize training and improve depth scale reliability. Extensive experiments on the DDAD and nuScenes benchmarks demonstrate that E^2Depth achieves a favorable trade-off between accuracy and efficiency, supporting practical deployment on real-world platforms.
| Original language | English |
|---|---|
| Pages (from-to) | 9080-9087 |
| Number of pages | 8 |
| Journal | IEEE Robotics and Automation Letters |
| Volume | 11 |
| Issue number | 8 |
| DOIs | |
| Publication status | Published - 2026 |
| Externally published | Yes |
Keywords
- Surround-view system
- depth estimation
- self-supervised learning
Fingerprint
Dive into the research topics of 'E2Depth: Efficient Self-Supervised Surround-View Depth Estimation With Explicit Geometric Enhancement'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver