Skip to main navigation Skip to search Skip to main content

Cascaded diffusion model and segment anything model for medical image synthesis

  • Haowen Pang
  • , Xiaoming Hong
  • , Peng Zhang
  • , Pengli Zhu
  • , Shun Yao
  • , Fengping An
  • , Guoyuan Yang
  • , Tiantian Liu
  • , Anqi Qiu*
  • , Chuyang Ye
  • , Tianyi Yan
  • *Corresponding author for this work
  • Beijing Institute of Technology
  • Hong Kong Polytechnic University
  • Sun Yat-Sen University
  • Shanxi University
  • Johns Hopkins University

Research output: Contribution to journalArticlepeer-review

Abstract

Multi-modal medical images provide complementary information essential for comprehensive diagnosis. However, constraints such as scanning time and patient safety often limit the availability of certain modalities. To address the problem, the synthesis of missing image modalities from available source modalities has become a promising approach. Although the recent Diffusion Model (DM) has improved the synthesis quality, existing methods frequently struggle to accurately synthesize regions with pathological abnormalities. In this work, to enhance the synthesis robustness, we propose a framework cascading the DM with the Segment Anything Model (SAM), which is referred to as DM-SAM, and it integrates the strengths of DM and SAM via uncertainty-guided prompt generation and multi-level prompt interaction. Specifically, in DM-SAM, prompts are first generated to highlight regions with high uncertainty in DM outputs that are likely abnormal, and they are used to guide SAM-based refinement of DM outputs. The guidance is achieved with the Uncertainty-Guided Cross-Attention (UGCA) module, where prompts serve as queries that selectively attend to relevant regions of both the source input image and initially synthesized image given by DM, thereby enabling more targeted and context-aware synthesis refinement. The guidance of the prompts is further enhanced with a multi-level interaction mechanism, where different levels of features are taken into account to accommodate high-level and low-level information. Moreover, for multi-modal source inputs, we propose Multi-UGCA that integrates features from any number of source modalities. Extensive experiments on three public datasets and 19 synthesis tasks show that DM-SAM significantly outperforms existing generic synthesis methods, particularly in abnormal regions.

Original languageEnglish
Article number114148
JournalPattern Recognition
Volume180
DOIs
Publication statusPublished - Dec 2026

Keywords

  • Diffusion model
  • Medical image synthesis
  • Segment anything model

Fingerprint

Dive into the research topics of 'Cascaded diffusion model and segment anything model for medical image synthesis'. Together they form a unique fingerprint.

Cite this