TY - JOUR
T1 - High-fidelity tabular data synthesis by quantile-based distribution harmonization under extreme class imbalance
AU - Zhou, Jinjie
AU - Luo, Senlin
AU - Pan, Limin
AU - Yang, Zongyuan
AU - Yang, Xiaonan
AU - Xu, Zehao
N1 - Publisher Copyright:
© 2026 Elsevier B.V.
PY - 2026/10/9
Y1 - 2026/10/9
N2 - Synthetic tabular data provides a practical way to support deep learning applications while reducing direct exposure of sensitive records. Mixed-type variables are commonly found in datasets with extreme class imbalance. Existing generative models tend to be biased toward discrete components due to they are concentrated and easily distinguishable. This bias can lead to the underfitting of relatively sparse continuous components, thereby reducing the diversity of synthetic data. Additionally, logarithmic sampling, which is commonly used to mitigate class imbalance, shifts the posterior distributions of continuous variables and may disrupt the conditional dependencies between features and labels, consequently reducing the utility of the synthetic data. To mitigate these issues, this paper proposes Quantile-based Distribution Harmonization GAN (QDHGAN) for imbalanced tabular data synthesis. QDHGAN distinguishes discrete components and partitions continuous components into distribution regions according to the cumulative distribution characteristics of mixed-type variables. It then maps these components into uniform distributions through index replacement, reducing density differences among components and alleviating generator bias. Furthermore, QDHGAN evenly segments the index range and uses discrepancies between synthetic and sampled distributions within each segment to guide the generator, compensating for distribution shifts and preserving feature-label dependencies. Experimental results demonstrate that QDHGAN achieves state-of-the-art performance. It can handle asymmetric distributions commonly found in imbalanced datasets, including mixed-type, long-tail, and skewed multi-mode distributions.
AB - Synthetic tabular data provides a practical way to support deep learning applications while reducing direct exposure of sensitive records. Mixed-type variables are commonly found in datasets with extreme class imbalance. Existing generative models tend to be biased toward discrete components due to they are concentrated and easily distinguishable. This bias can lead to the underfitting of relatively sparse continuous components, thereby reducing the diversity of synthetic data. Additionally, logarithmic sampling, which is commonly used to mitigate class imbalance, shifts the posterior distributions of continuous variables and may disrupt the conditional dependencies between features and labels, consequently reducing the utility of the synthetic data. To mitigate these issues, this paper proposes Quantile-based Distribution Harmonization GAN (QDHGAN) for imbalanced tabular data synthesis. QDHGAN distinguishes discrete components and partitions continuous components into distribution regions according to the cumulative distribution characteristics of mixed-type variables. It then maps these components into uniform distributions through index replacement, reducing density differences among components and alleviating generator bias. Furthermore, QDHGAN evenly segments the index range and uses discrepancies between synthetic and sampled distributions within each segment to guide the generator, compensating for distribution shifts and preserving feature-label dependencies. Experimental results demonstrate that QDHGAN achieves state-of-the-art performance. It can handle asymmetric distributions commonly found in imbalanced datasets, including mixed-type, long-tail, and skewed multi-mode distributions.
KW - Class imbalance
KW - Data synthesis
KW - GAN
KW - Tabular data
UR - https://www.scopus.com/pages/publications/105046432234
U2 - 10.1016/j.knosys.2026.116694
DO - 10.1016/j.knosys.2026.116694
M3 - Article
AN - SCOPUS:105046432234
SN - 0950-7051
VL - 351
JO - Knowledge-Based Systems
JF - Knowledge-Based Systems
M1 - 116694
ER -