TY - JOUR
T1 - Enhancing End-to-end Multilingual Medical Speech Translation via Terminology Injection Mechanism
AU - Huang, Shuanghong
AU - Feng, Chong
AU - Liu, Xia
AU - Xu, Jinlei
AU - Zhao, Xuan
AU - Shi, Ge
AU - Guo, Yuhang
AU - Gao, Yulong
N1 - Publisher Copyright:
© 2026 Science Press. All rights reserved.
PY - 2026/6/1
Y1 - 2026/6/1
N2 - Speech-to-text translation (S2TT) in the medical domain presents significant challenges due to the complexity of medical terminology and the scarcity of high-quality multilingual data. To address these issues, we propose MMST (Multilingual Medical Speech-to-text Translation), a novel framework that integrates domain knowledge into S2TT. Initially, we develop a multilingual medical terminology dictionary utilizing a large language model to extract terms from multilingual medical corpora. Following this, we create MMST, which consists of two main components: a two-stage training approach that integrates general pretraining with fine-tuning specific to the medical domain, and a terminology injection mechanism that embeds target terms into translation prompts, directing the generation process during training. Experiments on a many-to-many multilingual medical dataset demonstrate that MMST consistently outperforms the strong QwenAudio baseline, achieving an average BLEU improvement of +6.97 and higher BERTScore, particularly on terminology-rich and low-resource language pairs.
AB - Speech-to-text translation (S2TT) in the medical domain presents significant challenges due to the complexity of medical terminology and the scarcity of high-quality multilingual data. To address these issues, we propose MMST (Multilingual Medical Speech-to-text Translation), a novel framework that integrates domain knowledge into S2TT. Initially, we develop a multilingual medical terminology dictionary utilizing a large language model to extract terms from multilingual medical corpora. Following this, we create MMST, which consists of two main components: a two-stage training approach that integrates general pretraining with fine-tuning specific to the medical domain, and a terminology injection mechanism that embeds target terms into translation prompts, directing the generation process during training. Experiments on a many-to-many multilingual medical dataset demonstrate that MMST consistently outperforms the strong QwenAudio baseline, achieving an average BLEU improvement of +6.97 and higher BERTScore, particularly on terminology-rich and low-resource language pairs.
UR - https://www.scopus.com/pages/publications/105043570140
U2 - 10.3724/2096-7004.di.2025.0112
DO - 10.3724/2096-7004.di.2025.0112
M3 - Article
AN - SCOPUS:105043570140
SN - 2096-7004
VL - 8
JO - Data Intelligence
JF - Data Intelligence
IS - 2
M1 - 20250112
ER -