Abstract
Extraction attacks aim to steal Supervised Fine-Tuning (SFT) samples from fine-tuned Large Language Models (LLMs), posing severe threats to commercial assets and user privacy. Mainstream attack methods typically induce fine-tuned LLMs to generate candidate samples and then select one as the extracted sample. However, existing approaches struggle with two primary challenges. First, generated candidates often suffer from structural and stylistic deviations. Because fine-tuned models retain strong pre-training priors, their outputs tend to drift toward general corpora rather than the exact SFT format. Second, extracting the correct sample is hindered by semantic overlap. Current methods rely on semantic similarity for screening, which evaluates semantic equivalence rather than exact data memorization, leading to high mis-selection rates. To address these issues, a Token Modulation and Cross-Model Loss Ratio-Guided SFT Sample Extraction Attack method named TCEA is proposed. To overcome structural drift, TCEA employs a contrastive text generation module. By superimposing the logits difference between the fine-tuned and pre-trained models during decoding, this method dynamically modulates token probabilities to amplify SFT-specific features. Furthermore, to ensure accurate sample selection, TCEA evaluates the cross-model conditional loss ratio. This explicitly quantifies the fitting discrepancy between models, verifying verbatim text memorization rather than mere semantic similarity. Experimental results demonstrate that TCEA surpasses existing state-of-the-art methods across multi-domain datasets. Ultimately, by effectively quantifying cross-model fitting discrepancies, TCEA achieves high- accuracy extraction of SFT samples, demonstrating the genuine severity of this privacy threat.
| Original language | English |
|---|---|
| Article number | 116311 |
| Journal | Knowledge-Based Systems |
| Volume | 348 |
| DOIs | |
| Publication status | Published - 3 Aug 2026 |
Keywords
- Data extraction attack
- LLMs security
- Privacy
Fingerprint
Dive into the research topics of 'Stealing supervised fine tuning samples: A data extraction attack driven by token modulation and loss ratio'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver