Skip to main navigation Skip to search Skip to main content

Enhancing Front-door Defense with Causal Effect Maximization in Large Language Models

  • Beijing Institute of Technology
  • University of Chinese Academy of Sciences

Research output: Contribution to journalConference articlepeer-review

Abstract

Large language models (LLMs) face threats by different backdoor attacks from specific words, sentences, syntactic structures used to affect the behaviour of the LLM. Though causal inference-based defenses (such as front-door adjustment) offer a strong theory for trigger-independent defenses, the exsited method are hindered by high computational complexity and suboptimal candidate selection strategies. This paper presents a defense framework that maximizes Individual Treatment Effects (ITE) to alleviate these issues. Compared to state-of-the-art causal defenses using front-door adjustment to prove ITEs, our defense has higher practical defense capability and higher computational efficiency while still having strong theoretical support. We empirically show the success of our approach over existing causal defenses on various public datasets, achieving an average reduction in the attack success rate of 6.4% while maintaining comparable clean accuracy. Our proposed scheme requires only knowledge of the external inputs to the blackbox model, thus also providing an efficient and practical backdoor defense to defend LLMs in the real environment.

Original languageEnglish
Pages (from-to)3778-3783
Number of pages6
JournalProceedings of the International Conference on Computer Supported Cooperative Work in Design, CSCWD
Issue number2026
DOIs
Publication statusPublished - 2026
Externally publishedYes
Event29th International Conference on Computer Supported Cooperative Work in Design, CSCWD 2026 - Fuzhou, China
Duration: 13 May 202615 May 2026

Keywords

  • Backdoor Defense
  • Causal Inference
  • Front-door Adjustment
  • Individual Treatment Effect
  • Large Language Models

Fingerprint

Dive into the research topics of 'Enhancing Front-door Defense with Causal Effect Maximization in Large Language Models'. Together they form a unique fingerprint.

Cite this