TY - JOUR
T1 - VLAlignSeg
T2 - Vision-language alignment for few-shot medical image segmentation
AU - Qiao, Haipeng
AU - Zhang, Yanmei
AU - Jia, Xibin
AU - Fan, Chao
AU - Li, Yang
N1 - Publisher Copyright:
© 2026 Elsevier Ltd. All rights are reserved, including those for text and data mining, AI training, and similar technologies.
PY - 2026/11/1
Y1 - 2026/11/1
N2 - Existing few-shot medical image segmentation methods suffer from insufficient feature extraction capabilities, as they focus solely on mining visual features while neglecting the extraction of clinical prior textual information. To address this, we propose VLAlignSeg, a few-shot medical image segmentation algorithm based on Vision-Language alignment. The overall framework of the segmentation network is based on the prototype learning paradigm, leveraging the open-world semantic library of Large Language Models (LLMs) to generate clinical prior texts, and incorporating a Text-Visual Fusion Module (TVFM) to achieve deep integration of semantic prompts and image features. Furthermore, to fully exploit intra-class information within foreground regions and enhance the boundary sensitivity of prototype representation, we propose a Boundary Enhancement and Contrastive Learning (BCEL) prototype representation module. To achieve fine-grained segmentation of foreground regions, we present a Two-Stage Boundary Perception Mask Correction (TSBPM) module, which first performs preliminary mask correction via SAM and then conducts secondary calibration through a Boundary Enhancement Module (BEM). Additionally, to suppress model over-segmentation, we adopt an over-segmentation penalty strategy. Validation experiments on three public datasets demonstrate that the proposed method achieves superior segmentation accuracy compared to existing algorithms. The core code of VLAlignSeg is available at https://github.com/QYAOYang/VLAlignSeg-.
AB - Existing few-shot medical image segmentation methods suffer from insufficient feature extraction capabilities, as they focus solely on mining visual features while neglecting the extraction of clinical prior textual information. To address this, we propose VLAlignSeg, a few-shot medical image segmentation algorithm based on Vision-Language alignment. The overall framework of the segmentation network is based on the prototype learning paradigm, leveraging the open-world semantic library of Large Language Models (LLMs) to generate clinical prior texts, and incorporating a Text-Visual Fusion Module (TVFM) to achieve deep integration of semantic prompts and image features. Furthermore, to fully exploit intra-class information within foreground regions and enhance the boundary sensitivity of prototype representation, we propose a Boundary Enhancement and Contrastive Learning (BCEL) prototype representation module. To achieve fine-grained segmentation of foreground regions, we present a Two-Stage Boundary Perception Mask Correction (TSBPM) module, which first performs preliminary mask correction via SAM and then conducts secondary calibration through a Boundary Enhancement Module (BEM). Additionally, to suppress model over-segmentation, we adopt an over-segmentation penalty strategy. Validation experiments on three public datasets demonstrate that the proposed method achieves superior segmentation accuracy compared to existing algorithms. The core code of VLAlignSeg is available at https://github.com/QYAOYang/VLAlignSeg-.
KW - Boundary enhancement
KW - Contrastive learning
KW - Few-shot medical image segmentation
KW - Vision-language alignment
UR - https://www.scopus.com/pages/publications/105040629342
U2 - 10.1016/j.eswa.2026.132972
DO - 10.1016/j.eswa.2026.132972
M3 - Article
AN - SCOPUS:105040629342
SN - 0957-4174
VL - 329
JO - Expert Systems with Applications
JF - Expert Systems with Applications
M1 - 132972
ER -