TY - JOUR
T1 - A Latent Diffusion Model With Spatial Attributes for 3D Reconstruction in Biplanar X-ray Images
AU - Bian, Jingyi
AU - Xiao, Deqiang
AU - Zhang, Teng
AU - Shao, Long
AU - Ai, Danni
AU - Fan, Jingfan
AU - Fu, Tianyu
AU - Lin, Yucong
AU - Wang, Yuanyuan
AU - Song, Hong
AU - Wang, Junqiang
AU - Yang, Jian
N1 - Publisher Copyright:
© 1991-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - Computed tomography (CT) imaging provides essential 3D anatomical information for diagnosis and treatment planning, yet its clinical use is often limited by radiation exposure and equipment constraints. Given the accessibility and low radiation of biplanar X-ray imaging, 3D CT reconstruction from two orthogonal projections is an appealing alternative. However, the inherent 2D-to-3D ambiguity under extreme view sparsity makes it an ill-posed inverse problem, often resulting in geometric inconsistencies and degraded anatomical fidelity. To address this challenge, we propose a Spatial Attribute-aware Latent Diffusion Model (SA-LDM) that encodes deterministic X-ray geometry as structured spatial priors to guide 3D reconstruction. A Cross-View Feature Interaction (CVFI) module establishes semantic correspondences between orthogonal projections, while Multi-Scale Attention (MSA) layers extract hierarchical anatomical cues to enhance structural consistency. A Spatial Attribute-aware Feature Aggregation (SAFA) module further integrates spatial geometric attributes for geometry-aware multi-view fusion. Operating in a compact latent space with a Discrete Wavelet Domain (DWD) loss, SA-LDM accelerates convergence and preserves fine anatomical details while maintaining high computational efficiency. Through this design, SA-LDM reconstructs anatomically coherent and detail-preserving 3D volumes from only two orthogonal X-rays, effectively mitigating the geometric inconsistencies and structural distortions inherent to sparse-view reconstruction. Extensive experiments on lung, pelvis, and knee datasets demonstrate state-of-the-art performance across both quantitative and perceptual metrics, indicating that SA-LDM is able to generate anatomically faithful 3D reconstructions under extreme view sparsity.
AB - Computed tomography (CT) imaging provides essential 3D anatomical information for diagnosis and treatment planning, yet its clinical use is often limited by radiation exposure and equipment constraints. Given the accessibility and low radiation of biplanar X-ray imaging, 3D CT reconstruction from two orthogonal projections is an appealing alternative. However, the inherent 2D-to-3D ambiguity under extreme view sparsity makes it an ill-posed inverse problem, often resulting in geometric inconsistencies and degraded anatomical fidelity. To address this challenge, we propose a Spatial Attribute-aware Latent Diffusion Model (SA-LDM) that encodes deterministic X-ray geometry as structured spatial priors to guide 3D reconstruction. A Cross-View Feature Interaction (CVFI) module establishes semantic correspondences between orthogonal projections, while Multi-Scale Attention (MSA) layers extract hierarchical anatomical cues to enhance structural consistency. A Spatial Attribute-aware Feature Aggregation (SAFA) module further integrates spatial geometric attributes for geometry-aware multi-view fusion. Operating in a compact latent space with a Discrete Wavelet Domain (DWD) loss, SA-LDM accelerates convergence and preserves fine anatomical details while maintaining high computational efficiency. Through this design, SA-LDM reconstructs anatomically coherent and detail-preserving 3D volumes from only two orthogonal X-rays, effectively mitigating the geometric inconsistencies and structural distortions inherent to sparse-view reconstruction. Extensive experiments on lung, pelvis, and knee datasets demonstrate state-of-the-art performance across both quantitative and perceptual metrics, indicating that SA-LDM is able to generate anatomically faithful 3D reconstructions under extreme view sparsity.
KW - CT Volume Reconstruction
KW - Diffusion Model
KW - Multi Modality
KW - Spatial Attribute
UR - https://www.scopus.com/pages/publications/105043941093
U2 - 10.1109/TCSVT.2026.3709462
DO - 10.1109/TCSVT.2026.3709462
M3 - Article
AN - SCOPUS:105043941093
SN - 1051-8215
JO - IEEE Transactions on Circuits and Systems for Video Technology
JF - IEEE Transactions on Circuits and Systems for Video Technology
ER -