TY - JOUR
T1 - StegaFusion
T2 - Steganography for information hiding and fusion in multimodality
AU - Xu, Zihao
AU - Xu, Dawei
AU - Li, Zihan
AU - Hu, Juan
AU - Zheng, Baokun
AU - Zhang, Chuan
AU - Zhu, Liehuang
N1 - Publisher Copyright:
Copyright © 2026. Published by Elsevier B.V.
PY - 2026/7
Y1 - 2026/7
N2 - Current generative steganography techniques have attracted considerable attention due to their security. However, different platforms and social environments exhibit varying preferred modalities, and existing generative steganography techniques are often restricted to a single modality. Inspired by advancements in inpainting techniques, we observe that the inpainting process is inherently generative. Moreover, cross-modal inpainting minimally perturbs unchanged regions and shares a consistent masking-and-fill procedure. Based on these insights, we introduce StegaFusion, a novel framework for unifying multimodal generative steganography. StegaFusion leverages shared generation seeds and conditional information, which enables the receiver to deterministically reconstruct the reference content. The receiver then performs differential analysis on the inpainting-generated stego content to extract the secret message. Compared to traditional unimodal methods, StegaFusion enhances controllability, security, compatibility, and interpretability without requiring additional model training. To the best of our knowledge, StegaFusion is the first framework to formalize and unify cross-modal generative steganography, offering wide applicability. Extensive qualitative and quantitative experiments demonstrate the superior performance of StegaFusion in terms of controllability, security, and cross-modal compatibility.
AB - Current generative steganography techniques have attracted considerable attention due to their security. However, different platforms and social environments exhibit varying preferred modalities, and existing generative steganography techniques are often restricted to a single modality. Inspired by advancements in inpainting techniques, we observe that the inpainting process is inherently generative. Moreover, cross-modal inpainting minimally perturbs unchanged regions and shares a consistent masking-and-fill procedure. Based on these insights, we introduce StegaFusion, a novel framework for unifying multimodal generative steganography. StegaFusion leverages shared generation seeds and conditional information, which enables the receiver to deterministically reconstruct the reference content. The receiver then performs differential analysis on the inpainting-generated stego content to extract the secret message. Compared to traditional unimodal methods, StegaFusion enhances controllability, security, compatibility, and interpretability without requiring additional model training. To the best of our knowledge, StegaFusion is the first framework to formalize and unify cross-modal generative steganography, offering wide applicability. Extensive qualitative and quantitative experiments demonstrate the superior performance of StegaFusion in terms of controllability, security, and cross-modal compatibility.
KW - Inpainting
KW - Multimodality
KW - Steganography
UR - https://www.scopus.com/pages/publications/105028978575
U2 - 10.1016/j.inffus.2026.104150
DO - 10.1016/j.inffus.2026.104150
M3 - Article
AN - SCOPUS:105028978575
SN - 1566-2535
VL - 131
JO - Information Fusion
JF - Information Fusion
M1 - 104150
ER -