Abstract
Synthetic tabular data provides a practical way to support deep learning applications while reducing direct exposure of sensitive records. Mixed-type variables are commonly found in datasets with extreme class imbalance. Existing generative models tend to be biased toward discrete components due to they are concentrated and easily distinguishable. This bias can lead to the underfitting of relatively sparse continuous components, thereby reducing the diversity of synthetic data. Additionally, logarithmic sampling, which is commonly used to mitigate class imbalance, shifts the posterior distributions of continuous variables and may disrupt the conditional dependencies between features and labels, consequently reducing the utility of the synthetic data. To mitigate these issues, this paper proposes Quantile-based Distribution Harmonization GAN (QDHGAN) for imbalanced tabular data synthesis. QDHGAN distinguishes discrete components and partitions continuous components into distribution regions according to the cumulative distribution characteristics of mixed-type variables. It then maps these components into uniform distributions through index replacement, reducing density differences among components and alleviating generator bias. Furthermore, QDHGAN evenly segments the index range and uses discrepancies between synthetic and sampled distributions within each segment to guide the generator, compensating for distribution shifts and preserving feature-label dependencies. Experimental results demonstrate that QDHGAN achieves state-of-the-art performance. It can handle asymmetric distributions commonly found in imbalanced datasets, including mixed-type, long-tail, and skewed multi-mode distributions.
| Original language | English |
|---|---|
| Article number | 116694 |
| Journal | Knowledge-Based Systems |
| Volume | 351 |
| DOIs | |
| Publication status | Published - 9 Oct 2026 |
Keywords
- Class imbalance
- Data synthesis
- GAN
- Tabular data
Fingerprint
Dive into the research topics of 'High-fidelity tabular data synthesis by quantile-based distribution harmonization under extreme class imbalance'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver