Improving Parallel Corpus Quality for Chinese-Vietnamese Statistical Machine Translation

Huu Anh Tran; Yuhang Guo; Ping Jian; Shumin Shi; Heyan Huang

doi:10.15918/j.jbit1004-0579.201827.0116

Improving Parallel Corpus Quality for Chinese-Vietnamese Statistical Machine Translation

Huu Anh Tran, Yuhang Guo, Ping Jian, Shumin Shi, Heyan Huang^*

^*此作品的通讯作者

计算机学院

Beijing Institute of Technology

科研成果: 期刊稿件 › 文章 › 同行评审

3 引用（Scopus）

摘要

The performance of a machine translation system heavily depends on the quantity and quality of the bilingual language resource. However, getting a parallel corpus, which has a large scale and is of high quality, is a very difficult task especially for low resource languages such as Chinese-Vietnamese. Fortunately, multilingual user generated contents (UGC), such as bilingual movie subtitles, provide us access to automatic construction of the parallel corpus. Although the amount of UGC parallel corpora can be considerable, the original corpus is not suitable for statistical machine translation (SMT) systems. The corpus may contain translation errors, sentence mismatching, free translations, etc. To improve the quality of the bilingual corpus for SMT systems, three filtering methods are proposed: sentence length difference, the semantic of sentence pairs, and machine learning. Experiments are conducted on the Chinese to Vietnamese translation corpus. Experimental results demonstrate that all the three methods effectively improve the corpus quality, and the machine translation performance (BLEU score) can be improved by 1.32.

源语言	英语
页（从-至）	127-136
页数	10
期刊	Journal of Beijing Institute of Technology (English Edition)
卷	27
期	1
DOI	https://doi.org/10.15918/j.jbit1004-0579.201827.0116
出版状态	已出版 - 1 3月 2018

访问文件

10.15918/j.jbit1004-0579.201827.0116

其它文件与链接

链接到 Scopus 的出版物

引用此

@article{3e671c73715241a387c224a564db209c,

title = "Improving Parallel Corpus Quality for Chinese-Vietnamese Statistical Machine Translation",

abstract = "The performance of a machine translation system heavily depends on the quantity and quality of the bilingual language resource. However, getting a parallel corpus, which has a large scale and is of high quality, is a very difficult task especially for low resource languages such as Chinese-Vietnamese. Fortunately, multilingual user generated contents (UGC), such as bilingual movie subtitles, provide us access to automatic construction of the parallel corpus. Although the amount of UGC parallel corpora can be considerable, the original corpus is not suitable for statistical machine translation (SMT) systems. The corpus may contain translation errors, sentence mismatching, free translations, etc. To improve the quality of the bilingual corpus for SMT systems, three filtering methods are proposed: sentence length difference, the semantic of sentence pairs, and machine learning. Experiments are conducted on the Chinese to Vietnamese translation corpus. Experimental results demonstrate that all the three methods effectively improve the corpus quality, and the machine translation performance (BLEU score) can be improved by 1.32.",

keywords = "Bilingual movie subtitles, Chinese-Vietnamese translation, Low resource languages, Machine translation, Parallel corpus filtering",

author = "Tran, {Huu Anh} and Yuhang Guo and Ping Jian and Shumin Shi and Heyan Huang",

note = "Publisher Copyright: {\textcopyright} 2018 Editorial Department of Journal of Beijing Institute of Technology.",

year = "2018",

month = mar,

day = "1",

doi = "10.15918/j.jbit1004-0579.201827.0116",

language = "English",

volume = "27",

pages = "127--136",

journal = "Journal of Beijing Institute of Technology (English Edition)",

issn = "1004-0579",

publisher = "Beijing Institute of Technology",

number = "1",

}

TY - JOUR

T1 - Improving Parallel Corpus Quality for Chinese-Vietnamese Statistical Machine Translation

AU - Tran, Huu Anh

AU - Guo, Yuhang

AU - Jian, Ping

AU - Shi, Shumin

AU - Huang, Heyan

PY - 2018/3/1

Y1 - 2018/3/1

N2 - The performance of a machine translation system heavily depends on the quantity and quality of the bilingual language resource. However, getting a parallel corpus, which has a large scale and is of high quality, is a very difficult task especially for low resource languages such as Chinese-Vietnamese. Fortunately, multilingual user generated contents (UGC), such as bilingual movie subtitles, provide us access to automatic construction of the parallel corpus. Although the amount of UGC parallel corpora can be considerable, the original corpus is not suitable for statistical machine translation (SMT) systems. The corpus may contain translation errors, sentence mismatching, free translations, etc. To improve the quality of the bilingual corpus for SMT systems, three filtering methods are proposed: sentence length difference, the semantic of sentence pairs, and machine learning. Experiments are conducted on the Chinese to Vietnamese translation corpus. Experimental results demonstrate that all the three methods effectively improve the corpus quality, and the machine translation performance (BLEU score) can be improved by 1.32.

AB - The performance of a machine translation system heavily depends on the quantity and quality of the bilingual language resource. However, getting a parallel corpus, which has a large scale and is of high quality, is a very difficult task especially for low resource languages such as Chinese-Vietnamese. Fortunately, multilingual user generated contents (UGC), such as bilingual movie subtitles, provide us access to automatic construction of the parallel corpus. Although the amount of UGC parallel corpora can be considerable, the original corpus is not suitable for statistical machine translation (SMT) systems. The corpus may contain translation errors, sentence mismatching, free translations, etc. To improve the quality of the bilingual corpus for SMT systems, three filtering methods are proposed: sentence length difference, the semantic of sentence pairs, and machine learning. Experiments are conducted on the Chinese to Vietnamese translation corpus. Experimental results demonstrate that all the three methods effectively improve the corpus quality, and the machine translation performance (BLEU score) can be improved by 1.32.

KW - Bilingual movie subtitles

KW - Chinese-Vietnamese translation

KW - Low resource languages

KW - Machine translation

KW - Parallel corpus filtering

UR - http://www.scopus.com/inward/record.url?scp=85046162747&partnerID=8YFLogxK

U2 - 10.15918/j.jbit1004-0579.201827.0116

DO - 10.15918/j.jbit1004-0579.201827.0116

M3 - Article

AN - SCOPUS:85046162747

SN - 1004-0579

VL - 27

SP - 127

EP - 136

JO - Journal of Beijing Institute of Technology (English Edition)

JF - Journal of Beijing Institute of Technology (English Edition)

IS - 1

ER -

Improving Parallel Corpus Quality for Chinese-Vietnamese Statistical Machine Translation

摘要

访问文件

其它文件与链接

指纹

引用此