跳到主要导航 跳到搜索 跳到主要内容

Active learning for cross language text categorization

  • Beijing Institute of Technology

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Cross Language Text Categorization (CLTC) is the task of assigning class labels to documents written in a target language (e.g. Chinese) while the system is trained using labeled examples in a source language (e.g. English). With the technique of CLTC, we can build classifiers for multiple languages employing the existing training data in only one language, therefore avoid the cost of preparing training data for each individual language. One challenge for CLTC is the culture differences between languages, which causes the classifier trained on the source language doesn't perform well on the target language. In this paper, we propose an active learning algorithm for CLTC, which takes full advantage of both labeled data in the source language and unlabeled data in the target language. The classifier first learns the classification knowledge from the source language, and then learns the cultural dependent knowledge from the target language. In addition, we extend our algorithm to double viewed form by considering the source and target language as two views of the classification problem. Experiments show that our algorithm can effectively improve the cross language classification performance.

源语言英语
主期刊名Advances in Knowledge Discovery and Data Mining - 16th Pacific-Asia Conference, PAKDD 2012, Proceedings
出版商Springer Verlag
195-206
页数12
版本PART 1
ISBN(印刷版)9783642302169
DOI
出版状态已出版 - 2012
已对外发布
活动16th Pacific-Asia Conference on Advances in Knowledge Discovery and Data Mining, PAKDD 2012 - Kuala Lumpur, 马来西亚
期限: 29 5月 20121 6月 2012

出版系列

姓名Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
编号PART 1
7301 LNAI
ISSN(印刷版)0302-9743
ISSN(电子版)1611-3349

会议

会议16th Pacific-Asia Conference on Advances in Knowledge Discovery and Data Mining, PAKDD 2012
国家/地区马来西亚
Kuala Lumpur
时期29/05/121/06/12

指纹

探究 'Active learning for cross language text categorization' 的科研主题。它们共同构成独一无二的指纹。

引用此