跳到主要导航 跳到搜索 跳到主要内容

Construction of Uighur-Chinese parallel corpus

  • J. L. Song
  • , L. Dai
  • Beijing Institute of Technology

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

Uighur-Chinese parallel corpus is an important foundation of Uighur-Chinese cross-language information processing. As a corpus of minority language, its construction is relatively more difficult. In this paper, we discuss issues related to the construction. We firstly introduce the selection of corpus resources. Second, in order to accelerate the construction and improve the quality of the corpus, we develop an assistant construction system based on webpage content extraction and text duplication removal, etc. By using this system, we build a Uighur-Chinese parallel corpus with approximately 300,000 sentence pairs and a moderate size of dictionary of person name and place name. Finally, to evaluate the corpus, we build a demo Uighur-Chinese statistical translation system to explore the corpus. The result preliminarily verifies its effectiveness.

源语言英语
主期刊名Multimedia, Communication and Computing Application - Proceedings of the International Conference on Multimedia, Communication and Computing Application, MCCA 2014
编辑Ally Leung
出版商CRC Press/Balkema
353-356
页数4
ISBN(印刷版)9781138027756
DOI
出版状态已出版 - 2015
已对外发布
活动International Conference on Multimedia, Communication and Computing Application, MCCA 2014 - Xiamen, 中国
期限: 15 10月 201416 10月 2014

丛书

姓名Multimedia, Communication and Computing Application - Proceedings of the International Conference on Multimedia, Communication and Computing Application, MCCA 2014

会议

会议International Conference on Multimedia, Communication and Computing Application, MCCA 2014
国家/地区中国
Xiamen
时期15/10/1416/10/14

学术指纹

探究 'Construction of Uighur-Chinese parallel corpus' 的科研主题。它们共同构成独一无二的学术指纹。

引用此