TY - JOUR
T1 - Context-coupled token clustering for multiview 3D reconstruction
AU - Md Rajib, Sheik
AU - Li, Ronghua
AU - Fan, Yuanyi
AU - Wang, Jingchuan
AU - Li, Ran
AU - Zhang, Han
AU - Shagar, Md Masum Billa
N1 - Publisher Copyright:
© 2026 SPIE and IS&T.
PY - 2026/5/1
Y1 - 2026/5/1
N2 - Recent advances in transformer architectures have demonstrated strong performance across diverse computer vision tasks, including multiview 3D reconstruction. However, when handling a large number of input views, conventional transformers face challenges of efficiency and representational complexity due to the abundance of image tokens and the richness of visual content. This results in high computational overhead and difficulty in modeling meaningful inter-view relationships. Existing approaches often mitigate these issues by reducing token counts or discarding cross-view attention, which typically degrades performance. To address these limitations, we propose CCTNet, a transformer architecture based on a divide-and-conquer strategy. At its core, we introduce the context coupling token (CCT) mechanism, which aggregates tokens from all views and partitions them into multiple groups. Each group integrates tokens sampled across all perspectives, yielding a holistic representation that preserves inter-view diversity. This design enables our encoder to capture inter-view dependencies through grouped attention, while standard self-attention layers retain intra-view feature learning. Furthermore, a proactive upsampling decoder is employed to efficiently generate accurate voxel outputs. Extensive experiments on the ShapeNet dataset demonstrate that CCTNet achieves an IoU of 0.778 with 20 views, matching or surpassing prior methods such as UMIFormer (0.778) and GARNet+ (0.742), while requiring significantly fewer computations (4.5 GFLOPs versus 11.7 GFLOPs for K-means clustering). These results highlight that CCTNet delivers both superior accuracy and efficiency for multiview 3D reconstruction.
AB - Recent advances in transformer architectures have demonstrated strong performance across diverse computer vision tasks, including multiview 3D reconstruction. However, when handling a large number of input views, conventional transformers face challenges of efficiency and representational complexity due to the abundance of image tokens and the richness of visual content. This results in high computational overhead and difficulty in modeling meaningful inter-view relationships. Existing approaches often mitigate these issues by reducing token counts or discarding cross-view attention, which typically degrades performance. To address these limitations, we propose CCTNet, a transformer architecture based on a divide-and-conquer strategy. At its core, we introduce the context coupling token (CCT) mechanism, which aggregates tokens from all views and partitions them into multiple groups. Each group integrates tokens sampled across all perspectives, yielding a holistic representation that preserves inter-view diversity. This design enables our encoder to capture inter-view dependencies through grouped attention, while standard self-attention layers retain intra-view feature learning. Furthermore, a proactive upsampling decoder is employed to efficiently generate accurate voxel outputs. Extensive experiments on the ShapeNet dataset demonstrate that CCTNet achieves an IoU of 0.778 with 20 views, matching or surpassing prior methods such as UMIFormer (0.778) and GARNet+ (0.742), while requiring significantly fewer computations (4.5 GFLOPs versus 11.7 GFLOPs for K-means clustering). These results highlight that CCTNet delivers both superior accuracy and efficiency for multiview 3D reconstruction.
KW - 3D reconstruction
KW - ShapeNet
KW - multi-view learning
KW - token clustering
KW - transformers
UR - https://www.scopus.com/pages/publications/105043822463
U2 - 10.1117/1.JEI.35.3.033001
DO - 10.1117/1.JEI.35.3.033001
M3 - Article
AN - SCOPUS:105043822463
SN - 1017-9909
VL - 35
JO - Journal of Electronic Imaging
JF - Journal of Electronic Imaging
IS - 3
M1 - 033001
ER -