Abstract
Large language models (LLMs) face threats by different backdoor attacks from specific words, sentences, syntactic structures used to affect the behaviour of the LLM. Though causal inference-based defenses (such as front-door adjustment) offer a strong theory for trigger-independent defenses, the exsited method are hindered by high computational complexity and suboptimal candidate selection strategies. This paper presents a defense framework that maximizes Individual Treatment Effects (ITE) to alleviate these issues. Compared to state-of-the-art causal defenses using front-door adjustment to prove ITEs, our defense has higher practical defense capability and higher computational efficiency while still having strong theoretical support. We empirically show the success of our approach over existing causal defenses on various public datasets, achieving an average reduction in the attack success rate of 6.4% while maintaining comparable clean accuracy. Our proposed scheme requires only knowledge of the external inputs to the blackbox model, thus also providing an efficient and practical backdoor defense to defend LLMs in the real environment.
| Original language | English |
|---|---|
| Pages (from-to) | 3778-3783 |
| Number of pages | 6 |
| Journal | Proceedings of the International Conference on Computer Supported Cooperative Work in Design, CSCWD |
| Issue number | 2026 |
| DOIs | |
| Publication status | Published - 2026 |
| Externally published | Yes |
| Event | 29th International Conference on Computer Supported Cooperative Work in Design, CSCWD 2026 - Fuzhou, China Duration: 13 May 2026 → 15 May 2026 |
Keywords
- Backdoor Defense
- Causal Inference
- Front-door Adjustment
- Individual Treatment Effect
- Large Language Models
Fingerprint
Dive into the research topics of 'Enhancing Front-door Defense with Causal Effect Maximization in Large Language Models'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver