Abstract
Large language models (LLMs) require robust ownership verification to protect intellectual property, yet existing watermarking methods inevitably degrade model performance or lack resilience to attacks. We propose TransMark, a watermarking framework that rearranges certain neurons in a transformer’s feed-forward layers to embed a high-capacity, verifiable watermark. This design does not alter any model parameters or demand additional training steps. It leverages row and column permutations in feed-forward sub-layers, which are strict mathematical symmetries that leave the model’s function intact. We offer three strategies for choosing which neurons to swap: (i) a baseline method that pairs neurons based on norm and directional dissimilarity, (ii) a sensitivity-based strategy that minimizes changes to model outputs by focusing on less sensitive neurons, and (iii) a gradient-based approach that applies a small set of calibration texts to find neurons with the smallest gradient importance. We also include a Hamming-based error correction component to improve reliability. Our experiments confirm that TransMark preserves model accuracy and withstands quantization, random noise, and minor fine-tuning, all while embedding a large amount of information with no performance drop.
| Original language | English |
|---|---|
| Article number | 104180 |
| Journal | Computer Standards and Interfaces |
| Volume | 99 |
| DOIs | |
| Publication status | Published - Jan 2027 |
| Externally published | Yes |
Keywords
- Copyright protection
- Large language model
- Neuron swapping
- Permutation invariances
- Watermarking
Fingerprint
Dive into the research topics of 'TransMark: Lossless high-capacity watermarking for LLMs via neuron permutation invariances'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver