跳到主要导航 跳到搜索 跳到主要内容

Detection and Defense Against Backdoor Attacks in Large Language Models Based on Repeated Words Analysis

  • Beijing Institute of Technology
  • Ministry of Education in China
  • China University of Political Science and Law
  • Information Engineering University
  • University of Electronic Science and Technology of China

科研成果: 期刊稿件文章同行评审

摘要

Backdoor attacks pose significant threats to the security and reliability of large language models (LLMs). Existing approaches to backdoor detection often struggle with accurately identifying complex and stealthy triggers, especially in diverse and large-scale datasets, leading to gaps in defense effectiveness. This paper proposes a novel approach to detect and defend against such attacks by analyzing repeated patterns in input data. By identifying repeated words that frequently appear in malicious inputs, the proposed approach effectively locates backdoor triggers and mitigates their impact on LLMs. The method leverages semantic clustering and recursive optimization to enhance detection precision and ensure minimal disruption to benign outputs. Experimental results based on a real-world movie review dataset demonstrate the accuracy, robustness, and efficiency of this approach in detecting backdoor attacks and enhancing model security.

源语言英语
页(从-至)700-712
页数13
期刊Chinese Journal of Electronics
35
2
DOI
出版状态已出版 - 1 3月 2026
已对外发布

指纹

探究 'Detection and Defense Against Backdoor Attacks in Large Language Models Based on Repeated Words Analysis' 的科研主题。它们共同构成独一无二的指纹。

引用此