跳到主要导航 跳到搜索 跳到主要内容

Enhancing Front-door Defense with Causal Effect Maximization in Large Language Models

  • Beijing Institute of Technology
  • University of Chinese Academy of Sciences

科研成果: 期刊稿件会议文章同行评审

摘要

Large language models (LLMs) face threats by different backdoor attacks from specific words, sentences, syntactic structures used to affect the behaviour of the LLM. Though causal inference-based defenses (such as front-door adjustment) offer a strong theory for trigger-independent defenses, the exsited method are hindered by high computational complexity and suboptimal candidate selection strategies. This paper presents a defense framework that maximizes Individual Treatment Effects (ITE) to alleviate these issues. Compared to state-of-the-art causal defenses using front-door adjustment to prove ITEs, our defense has higher practical defense capability and higher computational efficiency while still having strong theoretical support. We empirically show the success of our approach over existing causal defenses on various public datasets, achieving an average reduction in the attack success rate of 6.4% while maintaining comparable clean accuracy. Our proposed scheme requires only knowledge of the external inputs to the blackbox model, thus also providing an efficient and practical backdoor defense to defend LLMs in the real environment.

源语言英语
页(从-至)3778-3783
页数6
期刊Proceedings of the International Conference on Computer Supported Cooperative Work in Design, CSCWD
2026
DOI
出版状态已出版 - 2026
已对外发布
活动29th International Conference on Computer Supported Cooperative Work in Design, CSCWD 2026 - Fuzhou, 中国
期限: 13 5月 202615 5月 2026

学术指纹

探究 'Enhancing Front-door Defense with Causal Effect Maximization in Large Language Models' 的科研主题。它们共同构成独一无二的学术指纹。

引用此