KeyBoxGAN: enhancing 2D object detection through annotated and editable image synthesis

Yashuo Bai; Yong Song; Fei Dong; Xu Li; Ya Zhou; Yizhao Liao; Jinxiang Huang; Xin Yang

doi:10.1007/s40747-025-01817-9

KeyBoxGAN: enhancing 2D object detection through annotated and editable image synthesis

Yashuo Bai, Yong Song^*, Fei Dong, Xu Li, Ya Zhou, Yizhao Liao, Jinxiang Huang, Xin Yang

^*Corresponding author for this work

Research output: Contribution to journal › Article › peer-review

Abstract

Sample augmentation, especially sample generation is conducive for addressing the challenge of training robust image and video object detection models based on the deep learning. Still, the existing methods lack sample editing capability and suffer from annotation work. This paper proposes an image sample generation method based on key box points detection and Generative adversarial network (GAN), named as KeyBoxGAN, to make image sample generation labeled and editable. KeyBoxGAN firstly predefines key box points positions, embeddings which control the objects’ positions and then the corresponding masks are generated according to Mahalanobis–Gaussuan heatmaps and Swin Transformer-SPADE generator to control objects’ generation regions, as well as the background generation. This adaptive and precisely supervised image generation method disentangles object position and appearance, enables image editable and self-labeled abilities. The experiments show KeyBoxGAN surpasses DCGAN, StyleGAN2 and DDPM in objective assessments, including Inception Distance (FID), Inception Score (IS), and Multi-Scale Structural Similarity Index (MS-SSIM), as well as in subjective evaluations by showing better visual quality. Moreover, the editable and self-labeled image generation capabilities make it a valuable tool in addressing challenges like occlusion, deformation, and varying environmental conditions in the 2D object detection.

Original language	English
Article number	186
Journal	Complex and Intelligent Systems
Volume	11
Issue number	4
DOIs	https://doi.org/10.1007/s40747-025-01817-9
Publication status	Published - Apr 2025

Keywords

Controllable image generation
Data augmentation
GANs
Image editing
Swin Transformer

Access to Document

10.1007/s40747-025-01817-9

Cite this

Bai, Y., Song, Y., Dong, F., Li, X., Zhou, Y., Liao, Y., Huang, J., & Yang, X. (2025). KeyBoxGAN: enhancing 2D object detection through annotated and editable image synthesis. Complex and Intelligent Systems, 11(4), Article 186. https://doi.org/10.1007/s40747-025-01817-9

@article{1c1b4a786ed34c0bbd7089659b0d7cd1,

title = "KeyBoxGAN: enhancing 2D object detection through annotated and editable image synthesis",

abstract = "Sample augmentation, especially sample generation is conducive for addressing the challenge of training robust image and video object detection models based on the deep learning. Still, the existing methods lack sample editing capability and suffer from annotation work. This paper proposes an image sample generation method based on key box points detection and Generative adversarial network (GAN), named as KeyBoxGAN, to make image sample generation labeled and editable. KeyBoxGAN firstly predefines key box points positions, embeddings which control the objects{\textquoteright} positions and then the corresponding masks are generated according to Mahalanobis–Gaussuan heatmaps and Swin Transformer-SPADE generator to control objects{\textquoteright} generation regions, as well as the background generation. This adaptive and precisely supervised image generation method disentangles object position and appearance, enables image editable and self-labeled abilities. The experiments show KeyBoxGAN surpasses DCGAN, StyleGAN2 and DDPM in objective assessments, including Inception Distance (FID), Inception Score (IS), and Multi-Scale Structural Similarity Index (MS-SSIM), as well as in subjective evaluations by showing better visual quality. Moreover, the editable and self-labeled image generation capabilities make it a valuable tool in addressing challenges like occlusion, deformation, and varying environmental conditions in the 2D object detection.",

keywords = "Controllable image generation, Data augmentation, GANs, Image editing, Swin Transformer",

author = "Yashuo Bai and Yong Song and Fei Dong and Xu Li and Ya Zhou and Yizhao Liao and Jinxiang Huang and Xin Yang",

note = "Publisher Copyright: {\textcopyright} The Author(s) 2025.",

year = "2025",

month = apr,

doi = "10.1007/s40747-025-01817-9",

language = "English",

volume = "11",

journal = "Complex and Intelligent Systems",

issn = "2199-4536",

publisher = "Springer International Publishing AG",

number = "4",

}

TY - JOUR

T1 - KeyBoxGAN

T2 - enhancing 2D object detection through annotated and editable image synthesis

AU - Bai, Yashuo

AU - Song, Yong

AU - Dong, Fei

AU - Li, Xu

AU - Zhou, Ya

AU - Liao, Yizhao

AU - Huang, Jinxiang

AU - Yang, Xin

N1 - Publisher Copyright: © The Author(s) 2025.

PY - 2025/4

Y1 - 2025/4

N2 - Sample augmentation, especially sample generation is conducive for addressing the challenge of training robust image and video object detection models based on the deep learning. Still, the existing methods lack sample editing capability and suffer from annotation work. This paper proposes an image sample generation method based on key box points detection and Generative adversarial network (GAN), named as KeyBoxGAN, to make image sample generation labeled and editable. KeyBoxGAN firstly predefines key box points positions, embeddings which control the objects’ positions and then the corresponding masks are generated according to Mahalanobis–Gaussuan heatmaps and Swin Transformer-SPADE generator to control objects’ generation regions, as well as the background generation. This adaptive and precisely supervised image generation method disentangles object position and appearance, enables image editable and self-labeled abilities. The experiments show KeyBoxGAN surpasses DCGAN, StyleGAN2 and DDPM in objective assessments, including Inception Distance (FID), Inception Score (IS), and Multi-Scale Structural Similarity Index (MS-SSIM), as well as in subjective evaluations by showing better visual quality. Moreover, the editable and self-labeled image generation capabilities make it a valuable tool in addressing challenges like occlusion, deformation, and varying environmental conditions in the 2D object detection.

AB - Sample augmentation, especially sample generation is conducive for addressing the challenge of training robust image and video object detection models based on the deep learning. Still, the existing methods lack sample editing capability and suffer from annotation work. This paper proposes an image sample generation method based on key box points detection and Generative adversarial network (GAN), named as KeyBoxGAN, to make image sample generation labeled and editable. KeyBoxGAN firstly predefines key box points positions, embeddings which control the objects’ positions and then the corresponding masks are generated according to Mahalanobis–Gaussuan heatmaps and Swin Transformer-SPADE generator to control objects’ generation regions, as well as the background generation. This adaptive and precisely supervised image generation method disentangles object position and appearance, enables image editable and self-labeled abilities. The experiments show KeyBoxGAN surpasses DCGAN, StyleGAN2 and DDPM in objective assessments, including Inception Distance (FID), Inception Score (IS), and Multi-Scale Structural Similarity Index (MS-SSIM), as well as in subjective evaluations by showing better visual quality. Moreover, the editable and self-labeled image generation capabilities make it a valuable tool in addressing challenges like occlusion, deformation, and varying environmental conditions in the 2D object detection.

KW - Controllable image generation

KW - Data augmentation

KW - GANs

KW - Image editing

KW - Swin Transformer

UR - http://www.scopus.com/inward/record.url?scp=105000012065&partnerID=8YFLogxK

U2 - 10.1007/s40747-025-01817-9

DO - 10.1007/s40747-025-01817-9

M3 - Article

AN - SCOPUS:105000012065

SN - 2199-4536

VL - 11

JO - Complex and Intelligent Systems

JF - Complex and Intelligent Systems

IS - 4

M1 - 186

ER -

KeyBoxGAN: enhancing 2D object detection through annotated and editable image synthesis

Abstract

Keywords

Access to Document

Other files and links

Fingerprint

Cite this