摘要
Large language models (LLMs) face threats by different backdoor attacks from specific words, sentences, syntactic structures used to affect the behaviour of the LLM. Though causal inference-based defenses (such as front-door adjustment) offer a strong theory for trigger-independent defenses, the exsited method are hindered by high computational complexity and suboptimal candidate selection strategies. This paper presents a defense framework that maximizes Individual Treatment Effects (ITE) to alleviate these issues. Compared to state-of-the-art causal defenses using front-door adjustment to prove ITEs, our defense has higher practical defense capability and higher computational efficiency while still having strong theoretical support. We empirically show the success of our approach over existing causal defenses on various public datasets, achieving an average reduction in the attack success rate of 6.4% while maintaining comparable clean accuracy. Our proposed scheme requires only knowledge of the external inputs to the blackbox model, thus also providing an efficient and practical backdoor defense to defend LLMs in the real environment.
| 源语言 | 英语 |
|---|---|
| 页(从-至) | 3778-3783 |
| 页数 | 6 |
| 期刊 | Proceedings of the International Conference on Computer Supported Cooperative Work in Design, CSCWD |
| 期 | 2026 |
| DOI | |
| 出版状态 | 已出版 - 2026 |
| 已对外发布 | 是 |
| 活动 | 29th International Conference on Computer Supported Cooperative Work in Design, CSCWD 2026 - Fuzhou, 中国 期限: 13 5月 2026 → 15 5月 2026 |
学术指纹
探究 'Enhancing Front-door Defense with Causal Effect Maximization in Large Language Models' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver