TY - GEN
T1 - Diagnose-Then-Optimize
T2 - 21st China Conference on Machine Translation, CCMT 2025
AU - Zhao, Xuan
AU - Feng, Chong
AU - Xu, Haojie
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2026.
PY - 2026
Y1 - 2026
N2 - While Large Language Models (LLMs) have shown promising capabilities in machine translation, their outputs often lack controllability and the ability to leverage error correction for translation improvement. To address this, we propose a two-stage framework, Diagnose-Then-Optimize (DTO), for structured translation quality enhancement. In the first stage, we fine-tune the large language model using human-annotated error data, enabling it to leverage translation error information for translation correction. In the second stage, we construct a preference dataset using response comparisons evaluated by ChatGPT, focusing on error correctness, correction effectiveness. We apply Direct Preference Optimization (DPO) to refine the model’s output behaviors based on these preferences. Our method demonstrates strong post-editing capabilities, consistently improving translation quality across WMT23 different systems’ outputs. The most significant gains are observed in English-Chinese, highlighting the model’s effectiveness in correcting diverse and complex translation errors. Experiments on WMT23 datasets across English–German, English–Russian, and English–Chinese demonstrate that DTO consistently improves the base LLaMA-3-8B, outperforming large-scale machine translation models such as NLLB_Greedy and Aya-23-35B in COMET scores. Our results highlight the effectiveness of combining structured error supervision with preference-driven fine-tuning, offering a robust and interpretable solution for controllable translation correction.
AB - While Large Language Models (LLMs) have shown promising capabilities in machine translation, their outputs often lack controllability and the ability to leverage error correction for translation improvement. To address this, we propose a two-stage framework, Diagnose-Then-Optimize (DTO), for structured translation quality enhancement. In the first stage, we fine-tune the large language model using human-annotated error data, enabling it to leverage translation error information for translation correction. In the second stage, we construct a preference dataset using response comparisons evaluated by ChatGPT, focusing on error correctness, correction effectiveness. We apply Direct Preference Optimization (DPO) to refine the model’s output behaviors based on these preferences. Our method demonstrates strong post-editing capabilities, consistently improving translation quality across WMT23 different systems’ outputs. The most significant gains are observed in English-Chinese, highlighting the model’s effectiveness in correcting diverse and complex translation errors. Experiments on WMT23 datasets across English–German, English–Russian, and English–Chinese demonstrate that DTO consistently improves the base LLaMA-3-8B, outperforming large-scale machine translation models such as NLLB_Greedy and Aya-23-35B in COMET scores. Our results highlight the effectiveness of combining structured error supervision with preference-driven fine-tuning, offering a robust and interpretable solution for controllable translation correction.
KW - Direct Preference Optimization
KW - Large Language Models
KW - Machine Translation
UR - https://www.scopus.com/pages/publications/105043997932
U2 - 10.1007/978-981-92-0199-0_2
DO - 10.1007/978-981-92-0199-0_2
M3 - Conference contribution
AN - SCOPUS:105043997932
SN - 9789819201983
T3 - Communications in Computer and Information Science
SP - 18
EP - 31
BT - Machine Translation - 21st China Conference, CCMT 2025, Proceedings
A2 - Xu, Jin'an
A2 - Tu, Zhaopeng
A2 - Chen, Kehai
A2 - Guo, Yuhang
PB - Springer Science and Business Media Deutschland GmbH
Y2 - 26 September 2025 through 28 September 2025
ER -