Skip to main navigation Skip to search Skip to main content

FFDCFormer: A Progressive Feature Fusion Transformer Framework for Depth Completion

  • Beijing Institute of Technology

Research output: Contribution to journalConference articlepeer-review

Abstract

This research proposes FFDCFormer, a novel framework for depth completion that reconstructs dense depth maps from sparse or incomplete inputs guided by color images, while maintaining both local detail preservation and global structural consistency. FFDCFormer is designed as a progressive feature-fusion encoder-decoder architecture that couples neighborhood attention-based vision Transformer with convolutional attention mechanisms, enabling joint preservation of local details and global structure. By hierarchically integrating multi-scale features, it captures fine-grained local structures and gradually aggregates global semantics. A depth refinement network is then appended to enhance cross-region geometric consistency. Experiments on the NYUv2 dataset demonstrate that FFDCFormer achieves state-of-the-art performance across multiple evaluation metrics. Furthermore, an occlusion study is conducted as a new evaluation paradigm, confirming the robustness of FFDCFormer under structured sparsity and underscoring its effectiveness and practicality in real-world scenarios.

Original languageEnglish
Pages (from-to)2474-2479
Number of pages6
JournalYouth Academic Annual Conference of Chinese Association of Automation, YAC
Issue number2026
DOIs
Publication statusPublished - 2026
Externally publishedYes
Event41st Youth Academic Annual Conference of Chinese Association of Automation, YAC 2026 - Changsha, China
Duration: 8 May 202610 May 2026

Keywords

  • convolutional attention
  • Depth completion
  • neighborhood attention
  • spatial propagation network
  • vision transformer

Fingerprint

Dive into the research topics of 'FFDCFormer: A Progressive Feature Fusion Transformer Framework for Depth Completion'. Together they form a unique fingerprint.

Cite this