Skip to main navigation Skip to search Skip to main content

VLAlignSeg: Vision-language alignment for few-shot medical image segmentation

  • Haipeng Qiao
  • , Yanmei Zhang*
  • , Xibin Jia*
  • , Chao Fan
  • , Yang Li
  • *Corresponding author for this work
  • Beijing Institute of Technology
  • Beijing University of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Existing few-shot medical image segmentation methods suffer from insufficient feature extraction capabilities, as they focus solely on mining visual features while neglecting the extraction of clinical prior textual information. To address this, we propose VLAlignSeg, a few-shot medical image segmentation algorithm based on Vision-Language alignment. The overall framework of the segmentation network is based on the prototype learning paradigm, leveraging the open-world semantic library of Large Language Models (LLMs) to generate clinical prior texts, and incorporating a Text-Visual Fusion Module (TVFM) to achieve deep integration of semantic prompts and image features. Furthermore, to fully exploit intra-class information within foreground regions and enhance the boundary sensitivity of prototype representation, we propose a Boundary Enhancement and Contrastive Learning (BCEL) prototype representation module. To achieve fine-grained segmentation of foreground regions, we present a Two-Stage Boundary Perception Mask Correction (TSBPM) module, which first performs preliminary mask correction via SAM and then conducts secondary calibration through a Boundary Enhancement Module (BEM). Additionally, to suppress model over-segmentation, we adopt an over-segmentation penalty strategy. Validation experiments on three public datasets demonstrate that the proposed method achieves superior segmentation accuracy compared to existing algorithms. The core code of VLAlignSeg is available at https://github.com/QYAOYang/VLAlignSeg-.

Original languageEnglish
Article number132972
JournalExpert Systems with Applications
Volume329
DOIs
Publication statusPublished - 1 Nov 2026
Externally publishedYes

Keywords

  • Boundary enhancement
  • Contrastive learning
  • Few-shot medical image segmentation
  • Vision-language alignment

Fingerprint

Dive into the research topics of 'VLAlignSeg: Vision-language alignment for few-shot medical image segmentation'. Together they form a unique fingerprint.

Cite this