TY - JOUR
T1 - A surrogate ensemble fusion framework for robust watermark extraction in black-box models
AU - Zhang, Zhao
AU - Luo, Senlin
AU - Xing, Fengtong
AU - Pan, Limin
N1 - Publisher Copyright:
© 2026 Elsevier B.V.
PY - 2026/12
Y1 - 2026/12
N2 - Watermarking is widely used to protect the intellectual property of deep neural networks (DNNs), yet its robustness can be undermined by black-box extraction attacks in which adversaries attempt to recover embedded watermarks through query-only access to deployed models. Existing watermark extraction methods often suffer from unstable trigger generation, sensitivity to distribution shifts, and limited adaptability to heterogeneous model architectures. In particular, random perturbations may introduce excessive noise that weakens watermark activation, while direct extraction from trigger samples is vulnerable to embedding misalignment, reducing extraction fidelity. To address these challenges, we propose TSGM, a trigger-sample and surrogate-model-based framework for black-box watermark extraction. The trigger sample generation module constructs reliable watermark-activating queries using surrogate-guided adversarial perturbations combined with entropy and confidence evaluation, followed by dimensionality reduction to capture representative perturbation patterns. The heterogeneous surrogate construction module trains multiple architectures using the generated triggers and integrates them into a cascaded surrogate model to reproduce watermark-related behaviors. Finally, watermark extraction is performed by localizing watermark-sensitive neurons and decoding watermark information from parameter deviations within the surrogate model. Extensive experiments on benchmark datasets demonstrate that TSGM consistently outperforms state-of-the-art baselines in extraction success rate and watermark fidelity, confirming its robustness across diverse black-box scenarios.
AB - Watermarking is widely used to protect the intellectual property of deep neural networks (DNNs), yet its robustness can be undermined by black-box extraction attacks in which adversaries attempt to recover embedded watermarks through query-only access to deployed models. Existing watermark extraction methods often suffer from unstable trigger generation, sensitivity to distribution shifts, and limited adaptability to heterogeneous model architectures. In particular, random perturbations may introduce excessive noise that weakens watermark activation, while direct extraction from trigger samples is vulnerable to embedding misalignment, reducing extraction fidelity. To address these challenges, we propose TSGM, a trigger-sample and surrogate-model-based framework for black-box watermark extraction. The trigger sample generation module constructs reliable watermark-activating queries using surrogate-guided adversarial perturbations combined with entropy and confidence evaluation, followed by dimensionality reduction to capture representative perturbation patterns. The heterogeneous surrogate construction module trains multiple architectures using the generated triggers and integrates them into a cascaded surrogate model to reproduce watermark-related behaviors. Finally, watermark extraction is performed by localizing watermark-sensitive neurons and decoding watermark information from parameter deviations within the surrogate model. Extensive experiments on benchmark datasets demonstrate that TSGM consistently outperforms state-of-the-art baselines in extraction success rate and watermark fidelity, confirming its robustness across diverse black-box scenarios.
KW - Intellectual property protection
KW - Sample generation
KW - Surrogate model construction
KW - Watermark extraction trigger
UR - https://www.scopus.com/pages/publications/105042482274
U2 - 10.1016/j.inffus.2026.104547
DO - 10.1016/j.inffus.2026.104547
M3 - Article
AN - SCOPUS:105042482274
SN - 1566-2535
VL - 136
JO - Information Fusion
JF - Information Fusion
M1 - 104547
ER -