Cluster optimized batch mode active learning sample selection method

Zhonghai He*, Zhichao Xia, Yinzhi Du, Xiaofang Zhang

*此作品的通讯作者

科研成果: 期刊稿件文章同行评审

摘要

Active learning for selecting representative samples submitted to labeling can save model development costs. However, the performance of single-sample selection in each iteration is compromised by a heavy computational burden and low efficiency in reference measurements, issues that can be addressed through batch mode active learning. The sample redundancy in batch mode active learning has long been a challenge. To overcome the shortcomings, a batch mode sample selection method that takes representativeness, diversity, and informativeness into account is proposed, called Gaussian Process Cluster Optimized Active Learning (GPCOAL). Firstly, the Gaussian process is utilized to obtain the variance (information) of each sample. Subsequently, K-means clustering is performed to ensure diversity, and the sample with largest silhouettes is selected from each cluster to ensure representativeness. Finally, the Gaussian process variance and the silhouettes of each sample are integrated to select the most suitable samples within each cluster. Experimental validation is conducted on spectroscopic datasets to illustrate the effectiveness of the GPCOAL sample selection method.

源语言英语
文章编号105746
期刊Infrared Physics and Technology
145
DOI
出版状态已出版 - 3月 2025

引用此

He, Z., Xia, Z., Du, Y., & Zhang, X. (2025). Cluster optimized batch mode active learning sample selection method. Infrared Physics and Technology, 145, 文章 105746. https://doi.org/10.1016/j.infrared.2025.105746