跳到主要导航 跳到搜索 跳到主要内容

Cascaded diffusion model and segment anything model for medical image synthesis

  • Haowen Pang
  • , Xiaoming Hong
  • , Peng Zhang
  • , Pengli Zhu
  • , Shun Yao
  • , Fengping An
  • , Guoyuan Yang
  • , Tiantian Liu
  • , Anqi Qiu*
  • , Chuyang Ye
  • , Tianyi Yan
  • *此作品的通讯作者
  • Beijing Institute of Technology
  • Hong Kong Polytechnic University
  • Sun Yat-Sen University
  • Shanxi University
  • Johns Hopkins University

科研成果: 期刊稿件文章同行评审

摘要

Multi-modal medical images provide complementary information essential for comprehensive diagnosis. However, constraints such as scanning time and patient safety often limit the availability of certain modalities. To address the problem, the synthesis of missing image modalities from available source modalities has become a promising approach. Although the recent Diffusion Model (DM) has improved the synthesis quality, existing methods frequently struggle to accurately synthesize regions with pathological abnormalities. In this work, to enhance the synthesis robustness, we propose a framework cascading the DM with the Segment Anything Model (SAM), which is referred to as DM-SAM, and it integrates the strengths of DM and SAM via uncertainty-guided prompt generation and multi-level prompt interaction. Specifically, in DM-SAM, prompts are first generated to highlight regions with high uncertainty in DM outputs that are likely abnormal, and they are used to guide SAM-based refinement of DM outputs. The guidance is achieved with the Uncertainty-Guided Cross-Attention (UGCA) module, where prompts serve as queries that selectively attend to relevant regions of both the source input image and initially synthesized image given by DM, thereby enabling more targeted and context-aware synthesis refinement. The guidance of the prompts is further enhanced with a multi-level interaction mechanism, where different levels of features are taken into account to accommodate high-level and low-level information. Moreover, for multi-modal source inputs, we propose Multi-UGCA that integrates features from any number of source modalities. Extensive experiments on three public datasets and 19 synthesis tasks show that DM-SAM significantly outperforms existing generic synthesis methods, particularly in abnormal regions.

源语言英语
文章编号114148
期刊Pattern Recognition
180
DOI
出版状态已出版 - 12月 2026

指纹

探究 'Cascaded diffusion model and segment anything model for medical image synthesis' 的科研主题。它们共同构成独一无二的指纹。

引用此