跳到主要导航 跳到搜索 跳到主要内容

VLAlignSeg: Vision-language alignment for few-shot medical image segmentation

  • Haipeng Qiao
  • , Yanmei Zhang*
  • , Xibin Jia*
  • , Chao Fan
  • , Yang Li
  • *此作品的通讯作者
  • Beijing Institute of Technology
  • Beijing University of Technology

科研成果: 期刊稿件文章同行评审

摘要

Existing few-shot medical image segmentation methods suffer from insufficient feature extraction capabilities, as they focus solely on mining visual features while neglecting the extraction of clinical prior textual information. To address this, we propose VLAlignSeg, a few-shot medical image segmentation algorithm based on Vision-Language alignment. The overall framework of the segmentation network is based on the prototype learning paradigm, leveraging the open-world semantic library of Large Language Models (LLMs) to generate clinical prior texts, and incorporating a Text-Visual Fusion Module (TVFM) to achieve deep integration of semantic prompts and image features. Furthermore, to fully exploit intra-class information within foreground regions and enhance the boundary sensitivity of prototype representation, we propose a Boundary Enhancement and Contrastive Learning (BCEL) prototype representation module. To achieve fine-grained segmentation of foreground regions, we present a Two-Stage Boundary Perception Mask Correction (TSBPM) module, which first performs preliminary mask correction via SAM and then conducts secondary calibration through a Boundary Enhancement Module (BEM). Additionally, to suppress model over-segmentation, we adopt an over-segmentation penalty strategy. Validation experiments on three public datasets demonstrate that the proposed method achieves superior segmentation accuracy compared to existing algorithms. The core code of VLAlignSeg is available at https://github.com/QYAOYang/VLAlignSeg-.

源语言英语
文章编号132972
期刊Expert Systems with Applications
329
DOI
出版状态已出版 - 1 11月 2026
已对外发布

指纹

探究 'VLAlignSeg: Vision-language alignment for few-shot medical image segmentation' 的科研主题。它们共同构成独一无二的指纹。

引用此