TY - JOUR
T1 - Disentangling structure from modality
T2 - An attention-based framework for CBCT-to-CT synthesis
AU - Guo, Shuanshuan
AU - Mao, Shufan
AU - Shi, Tianrui
AU - Lu, Shanfu
AU - Chen, Jiayu
AU - You, Rui
AU - Tang, Tianmin
AU - Yan, Ziye
AU - Li, Jianwu
AU - Zhou, Jianhua
N1 - Publisher Copyright:
© 2026 Elsevier Ltd
PY - 2026/10/15
Y1 - 2026/10/15
N2 - Cone-beam CT (CBCT) is essential for daily image guidance in radiotherapy; nevertheless, image degradation due to scatter and noise limits its quantitative application. Current deep learning techniques frequently compromise anatomical accuracy by indiscriminately altering all latent characteristics. We present an attention-guided generative adversarial network (AGD-GAN) that maintains anatomical structures by learning to provide a pixel-level spatial mask. This mask, based on a ResNet-50 encoder and concurrent attention modules, distinctly differentiates modality-invariant anatomy from modality-specific appearance. A compact transformation network maps solely the appearance features to the target domain, while the isolated anatomical features are transmitted unaltered to a StyleGAN2-based decoder. This distinction is maintained by a latent feature-consistency loss that offers direct structural oversight. Assessed using the public SynthRAD2023 benchmark (five-fold cross-validation, 1080 paired volumes), our approach attained a mean absolute HU error of 42.72, surpassing robust GAN and diffusion baselines. The robustness was validated in an external retrospective cohort of 25 patients. Our methodology provides a reliable and efficient solution for generating high-fidelity synthetic CT by developing an interpretable mask of the preserved anatomy, thereby directly enhancing adaptive radiotherapy operations.
AB - Cone-beam CT (CBCT) is essential for daily image guidance in radiotherapy; nevertheless, image degradation due to scatter and noise limits its quantitative application. Current deep learning techniques frequently compromise anatomical accuracy by indiscriminately altering all latent characteristics. We present an attention-guided generative adversarial network (AGD-GAN) that maintains anatomical structures by learning to provide a pixel-level spatial mask. This mask, based on a ResNet-50 encoder and concurrent attention modules, distinctly differentiates modality-invariant anatomy from modality-specific appearance. A compact transformation network maps solely the appearance features to the target domain, while the isolated anatomical features are transmitted unaltered to a StyleGAN2-based decoder. This distinction is maintained by a latent feature-consistency loss that offers direct structural oversight. Assessed using the public SynthRAD2023 benchmark (five-fold cross-validation, 1080 paired volumes), our approach attained a mean absolute HU error of 42.72, surpassing robust GAN and diffusion baselines. The robustness was validated in an external retrospective cohort of 25 patients. Our methodology provides a reliable and efficient solution for generating high-fidelity synthetic CT by developing an interpretable mask of the preserved anatomy, thereby directly enhancing adaptive radiotherapy operations.
KW - CBCT-to-CT synthesis
KW - Cone-beam CT
KW - Feature disentanglement
KW - Generative adversarial networks
KW - Medical image translation
UR - https://www.scopus.com/pages/publications/105042575829
U2 - 10.1016/j.bspc.2026.110832
DO - 10.1016/j.bspc.2026.110832
M3 - Article
AN - SCOPUS:105042575829
SN - 1746-8094
VL - 126
JO - Biomedical Signal Processing and Control
JF - Biomedical Signal Processing and Control
M1 - 110832
ER -