TY - JOUR
T1 - An epoch-weighted privacy budget allocation framework for fine-tuning large language models
AU - Zhang, Peisheng
AU - Wang, Naiyu
AU - Wu, Longfei
AU - Zhu, Liehuang
AU - Guan, Zhitao
N1 - Publisher Copyright:
© Science China Press 2026.
PY - 2026/4
Y1 - 2026/4
N2 - Fine-tuning large language models (LLMs) has demonstrated outstanding performance across various downstream tasks. However, fine-tuning LLMs on private training data poses significant privacy risks, as adversaries can employ various attack methods to extract sensitive information during training. Existing differential privacy frameworks for LLMs often rely on a uniform privacy budget allocation strategy, which neglects the increasing sensitivity of model parameters to noise perturbations as optimization progresses. To address these issues, we propose an epoch-weighted privacy budget allocation framework for fine-tuning LLMs (EW-FT), which incorporates epoch factors into the privacy budget allocation algorithm and utilizes stacked autoencoders to mitigate the curse of dimensionality. By injecting less noise into the forward hidden embeddings of more sensitive fine-tuning epochs, EW-FT achieves more targeted local differential privacy perturbations. Extensive experiments on three downstream tasks demonstrate that, while maintaining the same level of privacy protection, our EW-FT achieves higher model accuracy compared with state-of-the-art techniques.
AB - Fine-tuning large language models (LLMs) has demonstrated outstanding performance across various downstream tasks. However, fine-tuning LLMs on private training data poses significant privacy risks, as adversaries can employ various attack methods to extract sensitive information during training. Existing differential privacy frameworks for LLMs often rely on a uniform privacy budget allocation strategy, which neglects the increasing sensitivity of model parameters to noise perturbations as optimization progresses. To address these issues, we propose an epoch-weighted privacy budget allocation framework for fine-tuning LLMs (EW-FT), which incorporates epoch factors into the privacy budget allocation algorithm and utilizes stacked autoencoders to mitigate the curse of dimensionality. By injecting less noise into the forward hidden embeddings of more sensitive fine-tuning epochs, EW-FT achieves more targeted local differential privacy perturbations. Extensive experiments on three downstream tasks demonstrate that, while maintaining the same level of privacy protection, our EW-FT achieves higher model accuracy compared with state-of-the-art techniques.
KW - differential privacy
KW - distributed scenario
KW - fine-tuning
KW - large language model
KW - utility improvement
UR - https://www.scopus.com/pages/publications/105033452785
U2 - 10.1007/s11432-025-4580-x
DO - 10.1007/s11432-025-4580-x
M3 - Article
AN - SCOPUS:105033452785
SN - 1674-733X
VL - 69
JO - Science China Information Sciences
JF - Science China Information Sciences
IS - 4
M1 - 142106
ER -