TY - JOUR
T1 - Online learning a binary classifier for improving Google image search results
AU - Wan, Yu Chai
AU - Liu, Xia Bi
AU - Han, Fei Fei
AU - Tong, Kun Qi
AU - Liu, Yu
N1 - Publisher Copyright:
© 2014 Acta Automatica Sinica. All rights reserved.
PY - 2014/8/1
Y1 - 2014/8/1
N2 - It is promising to improve web image search results through exploiting the results' visual contents for learning a binary classifier which is used to refine the results' relevance degrees to the given query. This paper proposes an algorithm framework as a solution to this problem and investigates the key issue of training data selection under the framework. The training data selection process is divided into two stages: initial selection for triggering the classifier learning and dynamic selection in the iterations of classifier learning. We investigate two main ways of initial training data selection, including clustering based and ranking based, and compare automatic training data selection schemes with manual manner. Furthermore, support vector machines and the max-min pseudo-probability (MMP) based Bayesian classifier are employed to support image classification, respectively. By varying these factors in the framework, we implement eight algorithms and tested them on keyword based image search results from Google search engine. The experimental results confirm that how to select the training data from noisy search results is really a key issue in the problem considered in this paper and show that the proposed algorithm is effective to improve Google search results, especially at top ranks, thus is helpful to reduce the user labor in finding the desired images by browsing the ranking in depth. Even so, it is still worth meditative to make automatic training data selection scheme better towards perfect human annotation.
AB - It is promising to improve web image search results through exploiting the results' visual contents for learning a binary classifier which is used to refine the results' relevance degrees to the given query. This paper proposes an algorithm framework as a solution to this problem and investigates the key issue of training data selection under the framework. The training data selection process is divided into two stages: initial selection for triggering the classifier learning and dynamic selection in the iterations of classifier learning. We investigate two main ways of initial training data selection, including clustering based and ranking based, and compare automatic training data selection schemes with manual manner. Furthermore, support vector machines and the max-min pseudo-probability (MMP) based Bayesian classifier are employed to support image classification, respectively. By varying these factors in the framework, we implement eight algorithms and tested them on keyword based image search results from Google search engine. The experimental results confirm that how to select the training data from noisy search results is really a key issue in the problem considered in this paper and show that the proposed algorithm is effective to improve Google search results, especially at top ranks, thus is helpful to reduce the user labor in finding the desired images by browsing the ranking in depth. Even so, it is still worth meditative to make automatic training data selection scheme better towards perfect human annotation.
KW - Content-based image retrieval (CBIR)
KW - Image classifier learning
KW - Image search engine
KW - Search results improvement
KW - Training data selection
UR - http://www.scopus.com/inward/record.url?scp=84907092169&partnerID=8YFLogxK
U2 - 10.1004/SP.J.1004.2014.01699
DO - 10.1004/SP.J.1004.2014.01699
M3 - Article
AN - SCOPUS:84907092169
SN - 0254-4156
VL - 40
SP - 1699
EP - 1708
JO - Zidonghua Xuebao/Acta Automatica Sinica
JF - Zidonghua Xuebao/Acta Automatica Sinica
IS - 8
ER -