摘要
Despite the impressive capabilities of large language models (LLMs) in general machine translation tasks, their performance in domain-specific translation remains limited due to inadequate domain adaptation and lack of specialized knowledge. We propose a two-stage framework that enhances general LLMs to enable the model to learn domain-specific translation features. In the first stage, we enhanced the model’s capabilities in professional terminology alignment and domain-specific stylistic expression by designing two translator-like instructions and RAG-based example-selected methods. In the second stage, we constructed a preference dataset from domain-specific translation data aligned through GPT alignment and employed the Direct Preference Optimization (DPO) algorithm to further optimize the domain translation capabilities of the Large Language Model. The results on three different datasets show that our method not only achieves the highest BLEU, ChrF, and COMET scores in the WMT23 Terminology Translation Task, but also surpasses GPT-4 in COMET scores on both general and domain-specific translation tasks, particularly in the subtitle and literary domains. Demonstrating superior terminological consistency and domain awareness by simulating professional translators.
| 源语言 | 英语 |
|---|---|
| 文章编号 | 20250132 |
| 期刊 | Data Intelligence |
| 卷 | 8 |
| 期 | 2 |
| DOI | |
| 出版状态 | 已出版 - 1 6月 2026 |
学术指纹
探究 'Instruction-Guided Alignment to Simulate Translator Preferences in Domain-Specific Machine Translation' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver