Skip to main navigation Skip to search Skip to main content

Instruction-Guided Alignment to Simulate Translator Preferences in Domain-Specific Machine Translation

  • Xuan Zhao
  • , Chong Feng*
  • , Danjie Han
  • , Ge Shi
  • , Shuanghong Huang
  • , Haojie Xu
  • , Fangnuan Han
  • *Corresponding author for this work
  • Beijing Institute of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Despite the impressive capabilities of large language models (LLMs) in general machine translation tasks, their performance in domain-specific translation remains limited due to inadequate domain adaptation and lack of specialized knowledge. We propose a two-stage framework that enhances general LLMs to enable the model to learn domain-specific translation features. In the first stage, we enhanced the model’s capabilities in professional terminology alignment and domain-specific stylistic expression by designing two translator-like instructions and RAG-based example-selected methods. In the second stage, we constructed a preference dataset from domain-specific translation data aligned through GPT alignment and employed the Direct Preference Optimization (DPO) algorithm to further optimize the domain translation capabilities of the Large Language Model. The results on three different datasets show that our method not only achieves the highest BLEU, ChrF, and COMET scores in the WMT23 Terminology Translation Task, but also surpasses GPT-4 in COMET scores on both general and domain-specific translation tasks, particularly in the subtitle and literary domains. Demonstrating superior terminological consistency and domain awareness by simulating professional translators.

Original languageEnglish
Article number20250132
JournalData Intelligence
Volume8
Issue number2
DOIs
Publication statusPublished - 1 Jun 2026

Keywords

  • Direct Preference Optimization
  • Domain-specific translation
  • Instruction tuning
  • Knowledge injection
  • Large language model

Fingerprint

Dive into the research topics of 'Instruction-Guided Alignment to Simulate Translator Preferences in Domain-Specific Machine Translation'. Together they form a unique fingerprint.

Cite this