TY - JOUR
T1 - TransMark
T2 - Lossless high-capacity watermarking for LLMs via neuron permutation invariances
AU - Ye, Pei Gen
AU - Chen, Zhuorong
AU - Sang, Haiwei
AU - Zheng, Jun
N1 - Publisher Copyright:
Copyright © 2026. Published by Elsevier B.V.
PY - 2027/1
Y1 - 2027/1
N2 - Large language models (LLMs) require robust ownership verification to protect intellectual property, yet existing watermarking methods inevitably degrade model performance or lack resilience to attacks. We propose TransMark, a watermarking framework that rearranges certain neurons in a transformer’s feed-forward layers to embed a high-capacity, verifiable watermark. This design does not alter any model parameters or demand additional training steps. It leverages row and column permutations in feed-forward sub-layers, which are strict mathematical symmetries that leave the model’s function intact. We offer three strategies for choosing which neurons to swap: (i) a baseline method that pairs neurons based on norm and directional dissimilarity, (ii) a sensitivity-based strategy that minimizes changes to model outputs by focusing on less sensitive neurons, and (iii) a gradient-based approach that applies a small set of calibration texts to find neurons with the smallest gradient importance. We also include a Hamming-based error correction component to improve reliability. Our experiments confirm that TransMark preserves model accuracy and withstands quantization, random noise, and minor fine-tuning, all while embedding a large amount of information with no performance drop.
AB - Large language models (LLMs) require robust ownership verification to protect intellectual property, yet existing watermarking methods inevitably degrade model performance or lack resilience to attacks. We propose TransMark, a watermarking framework that rearranges certain neurons in a transformer’s feed-forward layers to embed a high-capacity, verifiable watermark. This design does not alter any model parameters or demand additional training steps. It leverages row and column permutations in feed-forward sub-layers, which are strict mathematical symmetries that leave the model’s function intact. We offer three strategies for choosing which neurons to swap: (i) a baseline method that pairs neurons based on norm and directional dissimilarity, (ii) a sensitivity-based strategy that minimizes changes to model outputs by focusing on less sensitive neurons, and (iii) a gradient-based approach that applies a small set of calibration texts to find neurons with the smallest gradient importance. We also include a Hamming-based error correction component to improve reliability. Our experiments confirm that TransMark preserves model accuracy and withstands quantization, random noise, and minor fine-tuning, all while embedding a large amount of information with no performance drop.
KW - Copyright protection
KW - Large language model
KW - Neuron swapping
KW - Permutation invariances
KW - Watermarking
UR - https://www.scopus.com/pages/publications/105040922821
U2 - 10.1016/j.csi.2026.104180
DO - 10.1016/j.csi.2026.104180
M3 - Article
AN - SCOPUS:105040922821
SN - 0920-5489
VL - 99
JO - Computer Standards and Interfaces
JF - Computer Standards and Interfaces
M1 - 104180
ER -