Skip to main navigation Skip to search Skip to main content

DRA: A Dual Retrieval Architecture for Domain Chinese Spelling Check

  • Haiming Wu
  • , Zhinie Nie
  • , Songkun Ji
  • , Dawei Song*
  • *Corresponding author for this work
  • Beijing Institute of Technology
  • Beijing Information Science & Technology University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Chinese Spelling Check (CSC), a foundational task in natural language processing and Chinese Computing, aims to detect and correct misspelled characters in Chinese texts. However, existing CSC methods suffer from the challenge of domain adaptation, for which the supervised learning approaches require large amounts of labeled data. LLM-based methods alleviate this problem through in-context learning (ICL) in the few-shot setting, but struggle to generalize across domain-specific tasks due to a lack of domain knowledge. To address these limitations, this paper proposes a novel dual retrieval architecture (DRA) for domain-specific CSC. Unlike existing LLM-based methods that rely on task examples, DRA integrates two core components: (1) a robust retriever to extract contextually relevant domain knowledge from external document corpora, and (2) an example retriever to provide correction pattern guidance. In the presence of input sentences with misspelled characters, the robust retriever mitigates retrieval errors via two synergistic strategies: (i) multi-modal modeling of phonetic, glyphic, and semantic features to link characters with plausible candidates; (ii) confusion-set augmented training to enhance robustness against error patterns. Extensive experiments on three domain-specific CSC benchmarks (LAW, MED, and ODW) demonstrate the effectiveness of DRA. It achieves correction F1 scores of 86.6%, 77.2%, and 93.1%, surpassing previous state-of-the-art methods by significant margins. Ablation studies confirm the critical role of the robust retriever in enhancing contextual accuracy and reducing dependency on domain-specific annotations.

Original languageEnglish
Title of host publicationNatural Language Processing and Chinese Computing - 14th National CCF Conference, NLPCC 2025, Proceedings
EditorsXian-Ling Mao, Zhaochun Ren, Muyun Yang
PublisherSpringer Science and Business Media Deutschland GmbH
Pages250-262
Number of pages13
ISBN (Print)9789819533459
DOIs
Publication statusPublished - 2026
Externally publishedYes
Event14th National CCF Conference on Natural Language Processing and Chinese Computing, NLPCC 2025 - Urumqi, China
Duration: 7 Aug 20259 Aug 2025

Publication series

NameLecture Notes in Computer Science
Volume16103 LNAI
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference14th National CCF Conference on Natural Language Processing and Chinese Computing, NLPCC 2025
Country/TerritoryChina
CityUrumqi
Period7/08/259/08/25

Keywords

  • Domain Chinese spelling check
  • Dual retrieval architecture
  • Retrieval-augmented generation
  • Robust retriever

Fingerprint

Dive into the research topics of 'DRA: A Dual Retrieval Architecture for Domain Chinese Spelling Check'. Together they form a unique fingerprint.

Cite this