Abstract
Recent advances in transformer architectures have demonstrated strong performance across diverse computer vision tasks, including multiview 3D reconstruction. However, when handling a large number of input views, conventional transformers face challenges of efficiency and representational complexity due to the abundance of image tokens and the richness of visual content. This results in high computational overhead and difficulty in modeling meaningful inter-view relationships. Existing approaches often mitigate these issues by reducing token counts or discarding cross-view attention, which typically degrades performance. To address these limitations, we propose CCTNet, a transformer architecture based on a divide-and-conquer strategy. At its core, we introduce the context coupling token (CCT) mechanism, which aggregates tokens from all views and partitions them into multiple groups. Each group integrates tokens sampled across all perspectives, yielding a holistic representation that preserves inter-view diversity. This design enables our encoder to capture inter-view dependencies through grouped attention, while standard self-attention layers retain intra-view feature learning. Furthermore, a proactive upsampling decoder is employed to efficiently generate accurate voxel outputs. Extensive experiments on the ShapeNet dataset demonstrate that CCTNet achieves an IoU of 0.778 with 20 views, matching or surpassing prior methods such as UMIFormer (0.778) and GARNet+ (0.742), while requiring significantly fewer computations (4.5 GFLOPs versus 11.7 GFLOPs for K-means clustering). These results highlight that CCTNet delivers both superior accuracy and efficiency for multiview 3D reconstruction.
| Original language | English |
|---|---|
| Article number | 033001 |
| Journal | Journal of Electronic Imaging |
| Volume | 35 |
| Issue number | 3 |
| DOIs | |
| Publication status | Published - 1 May 2026 |
| Externally published | Yes |
Keywords
- 3D reconstruction
- ShapeNet
- multi-view learning
- token clustering
- transformers
Fingerprint
Dive into the research topics of 'Context-coupled token clustering for multiview 3D reconstruction'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver