TY - GEN
T1 - LayoutGD
T2 - 16th ACM International Conference on Multimedia Retrieval, ICMR 2026
AU - Zhou, Xudong
AU - Li, Guozheng
AU - Liu, Chi Harold
N1 - Publisher Copyright:
© 2026 Copyright held by the owner/author(s).
PY - 2026/6/15
Y1 - 2026/6/15
N2 - Content-aware layout generation is crucial in poster design for automatically arranging layout elements. With the data scarcity problem, existing methods mainly employ retrieval augmentation or leverage Large Language Models (LLMs). However, these approaches still face persistent issues such as element overlap, misalignment, and high resource consumption (especially for LLMs). Additionally, these methods ignore and hardly handle layout generation based on pre-existing text canvases (i.e., text-rich images), which are common in real-world poster design. To address these issues, we propose LayoutGD, a graph-based diffusion method that aims to optimize overlap and misalignment while simultaneously enhancing performance on text-rich images. Our method represents all layout elements and image patches as independent nodes, constructs graphs based on specific topologies, and applies Graph Neural Networks (GNNs) to capture their high-dimensional spatial relationships. Furthermore, LayoutGD can process both text-clean and text-rich canvases in a unified framework, benefiting from our Enhance-Branch architecture. Extensive experiments demonstrate that our method achieves the state-of-the-art on various benchmarks. To further validate our performance and facilitate future research in text-rich canvas layout generation, we also construct a challenging text-rich dataset named TRich5001, which contains a wide variety of pre-existing text images from the real-world.
AB - Content-aware layout generation is crucial in poster design for automatically arranging layout elements. With the data scarcity problem, existing methods mainly employ retrieval augmentation or leverage Large Language Models (LLMs). However, these approaches still face persistent issues such as element overlap, misalignment, and high resource consumption (especially for LLMs). Additionally, these methods ignore and hardly handle layout generation based on pre-existing text canvases (i.e., text-rich images), which are common in real-world poster design. To address these issues, we propose LayoutGD, a graph-based diffusion method that aims to optimize overlap and misalignment while simultaneously enhancing performance on text-rich images. Our method represents all layout elements and image patches as independent nodes, constructs graphs based on specific topologies, and applies Graph Neural Networks (GNNs) to capture their high-dimensional spatial relationships. Furthermore, LayoutGD can process both text-clean and text-rich canvases in a unified framework, benefiting from our Enhance-Branch architecture. Extensive experiments demonstrate that our method achieves the state-of-the-art on various benchmarks. To further validate our performance and facilitate future research in text-rich canvas layout generation, we also construct a challenging text-rich dataset named TRich5001, which contains a wide variety of pre-existing text images from the real-world.
KW - Content-Aware Layout
KW - Diffusion
KW - Graph
KW - Graph Neural Network
KW - Poster Design
UR - https://www.scopus.com/pages/publications/105043288121
U2 - 10.1145/3805622.3810729
DO - 10.1145/3805622.3810729
M3 - Conference contribution
AN - SCOPUS:105043288121
T3 - ICMR 2026 - Proceedings of the 16th ACM International Conference on Multimedia Retrieval
SP - 1832
EP - 1841
BT - ICMR 2026 - Proceedings of the 16th ACM International Conference on Multimedia Retrieval
PB - Association for Computing Machinery, Inc
Y2 - 16 June 2026 through 19 June 2026
ER -