TY - GEN
T1 - DRA
T2 - 14th National CCF Conference on Natural Language Processing and Chinese Computing, NLPCC 2025
AU - Wu, Haiming
AU - Nie, Zhinie
AU - Ji, Songkun
AU - Song, Dawei
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2026.
PY - 2026
Y1 - 2026
N2 - Chinese Spelling Check (CSC), a foundational task in natural language processing and Chinese Computing, aims to detect and correct misspelled characters in Chinese texts. However, existing CSC methods suffer from the challenge of domain adaptation, for which the supervised learning approaches require large amounts of labeled data. LLM-based methods alleviate this problem through in-context learning (ICL) in the few-shot setting, but struggle to generalize across domain-specific tasks due to a lack of domain knowledge. To address these limitations, this paper proposes a novel dual retrieval architecture (DRA) for domain-specific CSC. Unlike existing LLM-based methods that rely on task examples, DRA integrates two core components: (1) a robust retriever to extract contextually relevant domain knowledge from external document corpora, and (2) an example retriever to provide correction pattern guidance. In the presence of input sentences with misspelled characters, the robust retriever mitigates retrieval errors via two synergistic strategies: (i) multi-modal modeling of phonetic, glyphic, and semantic features to link characters with plausible candidates; (ii) confusion-set augmented training to enhance robustness against error patterns. Extensive experiments on three domain-specific CSC benchmarks (LAW, MED, and ODW) demonstrate the effectiveness of DRA. It achieves correction F1 scores of 86.6%, 77.2%, and 93.1%, surpassing previous state-of-the-art methods by significant margins. Ablation studies confirm the critical role of the robust retriever in enhancing contextual accuracy and reducing dependency on domain-specific annotations.
AB - Chinese Spelling Check (CSC), a foundational task in natural language processing and Chinese Computing, aims to detect and correct misspelled characters in Chinese texts. However, existing CSC methods suffer from the challenge of domain adaptation, for which the supervised learning approaches require large amounts of labeled data. LLM-based methods alleviate this problem through in-context learning (ICL) in the few-shot setting, but struggle to generalize across domain-specific tasks due to a lack of domain knowledge. To address these limitations, this paper proposes a novel dual retrieval architecture (DRA) for domain-specific CSC. Unlike existing LLM-based methods that rely on task examples, DRA integrates two core components: (1) a robust retriever to extract contextually relevant domain knowledge from external document corpora, and (2) an example retriever to provide correction pattern guidance. In the presence of input sentences with misspelled characters, the robust retriever mitigates retrieval errors via two synergistic strategies: (i) multi-modal modeling of phonetic, glyphic, and semantic features to link characters with plausible candidates; (ii) confusion-set augmented training to enhance robustness against error patterns. Extensive experiments on three domain-specific CSC benchmarks (LAW, MED, and ODW) demonstrate the effectiveness of DRA. It achieves correction F1 scores of 86.6%, 77.2%, and 93.1%, surpassing previous state-of-the-art methods by significant margins. Ablation studies confirm the critical role of the robust retriever in enhancing contextual accuracy and reducing dependency on domain-specific annotations.
KW - Domain Chinese spelling check
KW - Dual retrieval architecture
KW - Retrieval-augmented generation
KW - Robust retriever
UR - https://www.scopus.com/pages/publications/105046007894
U2 - 10.1007/978-981-95-3346-6_19
DO - 10.1007/978-981-95-3346-6_19
M3 - Conference contribution
AN - SCOPUS:105046007894
SN - 9789819533459
T3 - Lecture Notes in Computer Science
SP - 250
EP - 262
BT - Natural Language Processing and Chinese Computing - 14th National CCF Conference, NLPCC 2025, Proceedings
A2 - Mao, Xian-Ling
A2 - Ren, Zhaochun
A2 - Yang, Muyun
PB - Springer Science and Business Media Deutschland GmbH
Y2 - 7 August 2025 through 9 August 2025
ER -