跳到主要导航 跳到搜索 跳到主要内容

Context-coupled token clustering for multiview 3D reconstruction

  • Sheik Md Rajib
  • , Ronghua Li*
  • , Yuanyi Fan
  • , Jingchuan Wang
  • , Ran Li
  • , Han Zhang
  • , Md Masum Billa Shagar
  • *此作品的通讯作者
  • Dalian Jiaotong University
  • Shanghai Jiao Tong University
  • Shenyang Railway Science and Technology Research Institute
  • College of Data Science

科研成果: 期刊稿件文章同行评审

摘要

Recent advances in transformer architectures have demonstrated strong performance across diverse computer vision tasks, including multiview 3D reconstruction. However, when handling a large number of input views, conventional transformers face challenges of efficiency and representational complexity due to the abundance of image tokens and the richness of visual content. This results in high computational overhead and difficulty in modeling meaningful inter-view relationships. Existing approaches often mitigate these issues by reducing token counts or discarding cross-view attention, which typically degrades performance. To address these limitations, we propose CCTNet, a transformer architecture based on a divide-and-conquer strategy. At its core, we introduce the context coupling token (CCT) mechanism, which aggregates tokens from all views and partitions them into multiple groups. Each group integrates tokens sampled across all perspectives, yielding a holistic representation that preserves inter-view diversity. This design enables our encoder to capture inter-view dependencies through grouped attention, while standard self-attention layers retain intra-view feature learning. Furthermore, a proactive upsampling decoder is employed to efficiently generate accurate voxel outputs. Extensive experiments on the ShapeNet dataset demonstrate that CCTNet achieves an IoU of 0.778 with 20 views, matching or surpassing prior methods such as UMIFormer (0.778) and GARNet+ (0.742), while requiring significantly fewer computations (4.5 GFLOPs versus 11.7 GFLOPs for K-means clustering). These results highlight that CCTNet delivers both superior accuracy and efficiency for multiview 3D reconstruction.

源语言英语
期刊论文编号033001
期刊Journal of Electronic Imaging
35
3
DOI
出版状态已出版 - 1 5月 2026
已对外发布

学术指纹

探究 'Context-coupled token clustering for multiview 3D reconstruction' 的科研主题。它们共同构成独一无二的学术指纹。

引用此