TY - JOUR
T1 - CEAP
T2 - Checkpoint Ensemble-based Active Learning and Pseudo-Labeling for annotation-efficient MLLM domain adaptation
AU - Zhao, Shuai
AU - Huang, Heyan
AU - Li, Xinge
N1 - Publisher Copyright:
© 2026 Published by Elsevier Ltd.
PY - 2027/1
Y1 - 2027/1
N2 - Adapting Multimodal Large Language Models (MLLMs) to specialized domains faces a significant annotation bottleneck, where acquiring high-quality labeled data is often expensive and time-consuming. To address this, we propose CEAP (Checkpoint Ensemble-based Active Learning and Pseudo-Labeling), an annotation-efficient framework for medical MLLM domain adaptation. CEAP repurposes naturally saved checkpoints from each LoRA fine-tuning run as an ensemble (constructed from five uniformly-spaced checkpoints per run), enabling uncertainty estimation through prediction disagreement and reliability assessment through prediction consensus without training multiple models from scratch. Our framework comprises two synergistic components: CEAL (Checkpoint Ensemble-based Active Learning) identifies high-uncertainty samples through ensemble disagreement for strategic annotation, while CEPL (Checkpoint Ensemble-based Pseudo-Labeling) generates reliable pseudo-labels for high-consensus samples to augment training data. We conduct experiments on medical visual question answering benchmarks, which represent a challenging testbed for annotation-efficient domain adaptation. Experimental results demonstrate CEAP’s effectiveness in reducing annotation requirements while maintaining competitive performance. On the SLAKE benchmark, CEAP-MiniCPM-V achieves 82.02% open-set recall and 83.89% closed-set accuracy using only 20% task-specific supervised fine-tuning (SFT) data, outperforming random sampling baseline (76.90% recall/79.09% accuracy) while approaching the fully-supervised upper bound (85.08% recall/87.26% accuracy). Compared to the state-of-the-art methods like LLaVA-Med, CEAP-MiniCPM-V achieves competitive performance (82.02% vs. 83.08% on open-set, and 83.89% vs. 85.34% on closed-set) with 80% reduction in annotation requirements, 99.8% fewer trainable parameters (17M vs. 13,000M), and 95% reduction in training time (7.5 vs. 160 GPU hours). Unlike existing methods that demand multi-GPU clusters and extensive domain-specific instruction data, CEAP operates on a single consumer-grade RTX 4090 GPU without any domain-specific data, making high-performance medical MLLM domain adaptation accessible for resource-constrained institutions and researchers.
AB - Adapting Multimodal Large Language Models (MLLMs) to specialized domains faces a significant annotation bottleneck, where acquiring high-quality labeled data is often expensive and time-consuming. To address this, we propose CEAP (Checkpoint Ensemble-based Active Learning and Pseudo-Labeling), an annotation-efficient framework for medical MLLM domain adaptation. CEAP repurposes naturally saved checkpoints from each LoRA fine-tuning run as an ensemble (constructed from five uniformly-spaced checkpoints per run), enabling uncertainty estimation through prediction disagreement and reliability assessment through prediction consensus without training multiple models from scratch. Our framework comprises two synergistic components: CEAL (Checkpoint Ensemble-based Active Learning) identifies high-uncertainty samples through ensemble disagreement for strategic annotation, while CEPL (Checkpoint Ensemble-based Pseudo-Labeling) generates reliable pseudo-labels for high-consensus samples to augment training data. We conduct experiments on medical visual question answering benchmarks, which represent a challenging testbed for annotation-efficient domain adaptation. Experimental results demonstrate CEAP’s effectiveness in reducing annotation requirements while maintaining competitive performance. On the SLAKE benchmark, CEAP-MiniCPM-V achieves 82.02% open-set recall and 83.89% closed-set accuracy using only 20% task-specific supervised fine-tuning (SFT) data, outperforming random sampling baseline (76.90% recall/79.09% accuracy) while approaching the fully-supervised upper bound (85.08% recall/87.26% accuracy). Compared to the state-of-the-art methods like LLaVA-Med, CEAP-MiniCPM-V achieves competitive performance (82.02% vs. 83.08% on open-set, and 83.89% vs. 85.34% on closed-set) with 80% reduction in annotation requirements, 99.8% fewer trainable parameters (17M vs. 13,000M), and 95% reduction in training time (7.5 vs. 160 GPU hours). Unlike existing methods that demand multi-GPU clusters and extensive domain-specific instruction data, CEAP operates on a single consumer-grade RTX 4090 GPU without any domain-specific data, making high-performance medical MLLM domain adaptation accessible for resource-constrained institutions and researchers.
KW - Active learning
KW - Checkpoint ensemble
KW - Domain adaptation
KW - Low-Rank Adaptation (LoRA)
KW - Multimodal Large Language Models
KW - Pseudo-labeling
UR - https://www.scopus.com/pages/publications/105045233929
U2 - 10.1016/j.ipm.2026.105062
DO - 10.1016/j.ipm.2026.105062
M3 - Article
AN - SCOPUS:105045233929
SN - 0306-4573
VL - 64
JO - Information Processing and Management
JF - Information Processing and Management
IS - 1
M1 - 105062
ER -