跳到主要导航 跳到搜索 跳到主要内容

Stealing supervised fine tuning samples: A data extraction attack driven by token modulation and loss ratio

  • Beijing Institute of Technology

科研成果: 期刊稿件文章同行评审

摘要

Extraction attacks aim to steal Supervised Fine-Tuning (SFT) samples from fine-tuned Large Language Models (LLMs), posing severe threats to commercial assets and user privacy. Mainstream attack methods typically induce fine-tuned LLMs to generate candidate samples and then select one as the extracted sample. However, existing approaches struggle with two primary challenges. First, generated candidates often suffer from structural and stylistic deviations. Because fine-tuned models retain strong pre-training priors, their outputs tend to drift toward general corpora rather than the exact SFT format. Second, extracting the correct sample is hindered by semantic overlap. Current methods rely on semantic similarity for screening, which evaluates semantic equivalence rather than exact data memorization, leading to high mis-selection rates. To address these issues, a Token Modulation and Cross-Model Loss Ratio-Guided SFT Sample Extraction Attack method named TCEA is proposed. To overcome structural drift, TCEA employs a contrastive text generation module. By superimposing the logits difference between the fine-tuned and pre-trained models during decoding, this method dynamically modulates token probabilities to amplify SFT-specific features. Furthermore, to ensure accurate sample selection, TCEA evaluates the cross-model conditional loss ratio. This explicitly quantifies the fitting discrepancy between models, verifying verbatim text memorization rather than mere semantic similarity. Experimental results demonstrate that TCEA surpasses existing state-of-the-art methods across multi-domain datasets. Ultimately, by effectively quantifying cross-model fitting discrepancies, TCEA achieves high- accuracy extraction of SFT samples, demonstrating the genuine severity of this privacy threat.

源语言英语
文章编号116311
期刊Knowledge-Based Systems
348
DOI
出版状态已出版 - 3 8月 2026

指纹

探究 'Stealing supervised fine tuning samples: A data extraction attack driven by token modulation and loss ratio' 的科研主题。它们共同构成独一无二的指纹。

引用此