Abstract
Adapting Multimodal Large Language Models (MLLMs) to specialized domains faces a significant annotation bottleneck, where acquiring high-quality labeled data is often expensive and time-consuming. To address this, we propose CEAP (Checkpoint Ensemble-based Active Learning and Pseudo-Labeling), an annotation-efficient framework for medical MLLM domain adaptation. CEAP repurposes naturally saved checkpoints from each LoRA fine-tuning run as an ensemble (constructed from five uniformly-spaced checkpoints per run), enabling uncertainty estimation through prediction disagreement and reliability assessment through prediction consensus without training multiple models from scratch. Our framework comprises two synergistic components: CEAL (Checkpoint Ensemble-based Active Learning) identifies high-uncertainty samples through ensemble disagreement for strategic annotation, while CEPL (Checkpoint Ensemble-based Pseudo-Labeling) generates reliable pseudo-labels for high-consensus samples to augment training data. We conduct experiments on medical visual question answering benchmarks, which represent a challenging testbed for annotation-efficient domain adaptation. Experimental results demonstrate CEAP’s effectiveness in reducing annotation requirements while maintaining competitive performance. On the SLAKE benchmark, CEAP-MiniCPM-V achieves 82.02% open-set recall and 83.89% closed-set accuracy using only 20% task-specific supervised fine-tuning (SFT) data, outperforming random sampling baseline (76.90% recall/79.09% accuracy) while approaching the fully-supervised upper bound (85.08% recall/87.26% accuracy). Compared to the state-of-the-art methods like LLaVA-Med, CEAP-MiniCPM-V achieves competitive performance (82.02% vs. 83.08% on open-set, and 83.89% vs. 85.34% on closed-set) with 80% reduction in annotation requirements, 99.8% fewer trainable parameters (17M vs. 13,000M), and 95% reduction in training time (7.5 vs. 160 GPU hours). Unlike existing methods that demand multi-GPU clusters and extensive domain-specific instruction data, CEAP operates on a single consumer-grade RTX 4090 GPU without any domain-specific data, making high-performance medical MLLM domain adaptation accessible for resource-constrained institutions and researchers.
| Original language | English |
|---|---|
| Article number | 105062 |
| Journal | Information Processing and Management |
| Volume | 64 |
| Issue number | 1 |
| DOIs | |
| Publication status | Published - Jan 2027 |
Keywords
- Active learning
- Checkpoint ensemble
- Domain adaptation
- Low-Rank Adaptation (LoRA)
- Multimodal Large Language Models
- Pseudo-labeling
Fingerprint
Dive into the research topics of 'CEAP: Checkpoint Ensemble-based Active Learning and Pseudo-Labeling for annotation-efficient MLLM domain adaptation'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver