Abstract
This research proposes FFDCFormer, a novel framework for depth completion that reconstructs dense depth maps from sparse or incomplete inputs guided by color images, while maintaining both local detail preservation and global structural consistency. FFDCFormer is designed as a progressive feature-fusion encoder-decoder architecture that couples neighborhood attention-based vision Transformer with convolutional attention mechanisms, enabling joint preservation of local details and global structure. By hierarchically integrating multi-scale features, it captures fine-grained local structures and gradually aggregates global semantics. A depth refinement network is then appended to enhance cross-region geometric consistency. Experiments on the NYUv2 dataset demonstrate that FFDCFormer achieves state-of-the-art performance across multiple evaluation metrics. Furthermore, an occlusion study is conducted as a new evaluation paradigm, confirming the robustness of FFDCFormer under structured sparsity and underscoring its effectiveness and practicality in real-world scenarios.
| Original language | English |
|---|---|
| Pages (from-to) | 2474-2479 |
| Number of pages | 6 |
| Journal | Youth Academic Annual Conference of Chinese Association of Automation, YAC |
| Issue number | 2026 |
| DOIs | |
| Publication status | Published - 2026 |
| Externally published | Yes |
| Event | 41st Youth Academic Annual Conference of Chinese Association of Automation, YAC 2026 - Changsha, China Duration: 8 May 2026 → 10 May 2026 |
Keywords
- convolutional attention
- Depth completion
- neighborhood attention
- spatial propagation network
- vision transformer
Fingerprint
Dive into the research topics of 'FFDCFormer: A Progressive Feature Fusion Transformer Framework for Depth Completion'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver