TY - JOUR
T1 - Cascaded diffusion model and segment anything model for medical image synthesis
AU - Pang, Haowen
AU - Hong, Xiaoming
AU - Zhang, Peng
AU - Zhu, Pengli
AU - Yao, Shun
AU - An, Fengping
AU - Yang, Guoyuan
AU - Liu, Tiantian
AU - Qiu, Anqi
AU - Ye, Chuyang
AU - Yan, Tianyi
N1 - Publisher Copyright:
© 2026 Elsevier Ltd
PY - 2026/12
Y1 - 2026/12
N2 - Multi-modal medical images provide complementary information essential for comprehensive diagnosis. However, constraints such as scanning time and patient safety often limit the availability of certain modalities. To address the problem, the synthesis of missing image modalities from available source modalities has become a promising approach. Although the recent Diffusion Model (DM) has improved the synthesis quality, existing methods frequently struggle to accurately synthesize regions with pathological abnormalities. In this work, to enhance the synthesis robustness, we propose a framework cascading the DM with the Segment Anything Model (SAM), which is referred to as DM-SAM, and it integrates the strengths of DM and SAM via uncertainty-guided prompt generation and multi-level prompt interaction. Specifically, in DM-SAM, prompts are first generated to highlight regions with high uncertainty in DM outputs that are likely abnormal, and they are used to guide SAM-based refinement of DM outputs. The guidance is achieved with the Uncertainty-Guided Cross-Attention (UGCA) module, where prompts serve as queries that selectively attend to relevant regions of both the source input image and initially synthesized image given by DM, thereby enabling more targeted and context-aware synthesis refinement. The guidance of the prompts is further enhanced with a multi-level interaction mechanism, where different levels of features are taken into account to accommodate high-level and low-level information. Moreover, for multi-modal source inputs, we propose Multi-UGCA that integrates features from any number of source modalities. Extensive experiments on three public datasets and 19 synthesis tasks show that DM-SAM significantly outperforms existing generic synthesis methods, particularly in abnormal regions.
AB - Multi-modal medical images provide complementary information essential for comprehensive diagnosis. However, constraints such as scanning time and patient safety often limit the availability of certain modalities. To address the problem, the synthesis of missing image modalities from available source modalities has become a promising approach. Although the recent Diffusion Model (DM) has improved the synthesis quality, existing methods frequently struggle to accurately synthesize regions with pathological abnormalities. In this work, to enhance the synthesis robustness, we propose a framework cascading the DM with the Segment Anything Model (SAM), which is referred to as DM-SAM, and it integrates the strengths of DM and SAM via uncertainty-guided prompt generation and multi-level prompt interaction. Specifically, in DM-SAM, prompts are first generated to highlight regions with high uncertainty in DM outputs that are likely abnormal, and they are used to guide SAM-based refinement of DM outputs. The guidance is achieved with the Uncertainty-Guided Cross-Attention (UGCA) module, where prompts serve as queries that selectively attend to relevant regions of both the source input image and initially synthesized image given by DM, thereby enabling more targeted and context-aware synthesis refinement. The guidance of the prompts is further enhanced with a multi-level interaction mechanism, where different levels of features are taken into account to accommodate high-level and low-level information. Moreover, for multi-modal source inputs, we propose Multi-UGCA that integrates features from any number of source modalities. Extensive experiments on three public datasets and 19 synthesis tasks show that DM-SAM significantly outperforms existing generic synthesis methods, particularly in abnormal regions.
KW - Diffusion model
KW - Medical image synthesis
KW - Segment anything model
UR - https://www.scopus.com/pages/publications/105041628270
U2 - 10.1016/j.patcog.2026.114148
DO - 10.1016/j.patcog.2026.114148
M3 - Article
AN - SCOPUS:105041628270
SN - 0031-3203
VL - 180
JO - Pattern Recognition
JF - Pattern Recognition
M1 - 114148
ER -