TY - JOUR
T1 - Stealing supervised fine tuning samples
T2 - A data extraction attack driven by token modulation and loss ratio
AU - Pi, Jiawei
AU - Luo, Senlin
AU - Pan, Limin
AU - Xu, Chengke
AU - Zhang, Ji
N1 - Publisher Copyright:
© 2026
PY - 2026/8/3
Y1 - 2026/8/3
N2 - Extraction attacks aim to steal Supervised Fine-Tuning (SFT) samples from fine-tuned Large Language Models (LLMs), posing severe threats to commercial assets and user privacy. Mainstream attack methods typically induce fine-tuned LLMs to generate candidate samples and then select one as the extracted sample. However, existing approaches struggle with two primary challenges. First, generated candidates often suffer from structural and stylistic deviations. Because fine-tuned models retain strong pre-training priors, their outputs tend to drift toward general corpora rather than the exact SFT format. Second, extracting the correct sample is hindered by semantic overlap. Current methods rely on semantic similarity for screening, which evaluates semantic equivalence rather than exact data memorization, leading to high mis-selection rates. To address these issues, a Token Modulation and Cross-Model Loss Ratio-Guided SFT Sample Extraction Attack method named TCEA is proposed. To overcome structural drift, TCEA employs a contrastive text generation module. By superimposing the logits difference between the fine-tuned and pre-trained models during decoding, this method dynamically modulates token probabilities to amplify SFT-specific features. Furthermore, to ensure accurate sample selection, TCEA evaluates the cross-model conditional loss ratio. This explicitly quantifies the fitting discrepancy between models, verifying verbatim text memorization rather than mere semantic similarity. Experimental results demonstrate that TCEA surpasses existing state-of-the-art methods across multi-domain datasets. Ultimately, by effectively quantifying cross-model fitting discrepancies, TCEA achieves high- accuracy extraction of SFT samples, demonstrating the genuine severity of this privacy threat.
AB - Extraction attacks aim to steal Supervised Fine-Tuning (SFT) samples from fine-tuned Large Language Models (LLMs), posing severe threats to commercial assets and user privacy. Mainstream attack methods typically induce fine-tuned LLMs to generate candidate samples and then select one as the extracted sample. However, existing approaches struggle with two primary challenges. First, generated candidates often suffer from structural and stylistic deviations. Because fine-tuned models retain strong pre-training priors, their outputs tend to drift toward general corpora rather than the exact SFT format. Second, extracting the correct sample is hindered by semantic overlap. Current methods rely on semantic similarity for screening, which evaluates semantic equivalence rather than exact data memorization, leading to high mis-selection rates. To address these issues, a Token Modulation and Cross-Model Loss Ratio-Guided SFT Sample Extraction Attack method named TCEA is proposed. To overcome structural drift, TCEA employs a contrastive text generation module. By superimposing the logits difference between the fine-tuned and pre-trained models during decoding, this method dynamically modulates token probabilities to amplify SFT-specific features. Furthermore, to ensure accurate sample selection, TCEA evaluates the cross-model conditional loss ratio. This explicitly quantifies the fitting discrepancy between models, verifying verbatim text memorization rather than mere semantic similarity. Experimental results demonstrate that TCEA surpasses existing state-of-the-art methods across multi-domain datasets. Ultimately, by effectively quantifying cross-model fitting discrepancies, TCEA achieves high- accuracy extraction of SFT samples, demonstrating the genuine severity of this privacy threat.
KW - Data extraction attack
KW - LLMs security
KW - Privacy
UR - https://www.scopus.com/pages/publications/105041447194
U2 - 10.1016/j.knosys.2026.116311
DO - 10.1016/j.knosys.2026.116311
M3 - Article
AN - SCOPUS:105041447194
SN - 0950-7051
VL - 348
JO - Knowledge-Based Systems
JF - Knowledge-Based Systems
M1 - 116311
ER -