EDBT 2026 Demo / reviewers in the wild / expert
Alex X. Wang
dblp:334/7221
· DBLP profile ↗
10ranked-venue papers
10as first author
10since 2021 · last 2026
0000-0002-3691-8652ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 6 · 6 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TTVAE: Transformer-Based Generative Modeling for Tabular Data Generation (Abstract Reprint)abstractTabular data synthesis presents unique challenges, with Transformer models remaining underexplored despite the applications of Variational Autoencoders and Generative Adversarial Networks. To address this gap, we propose the Transformer-based Tabular Variational AutoEncoder (TTVAE), leveraging the attention mechanism for capturing complex data distributions. The inclusion of the attention mechanism enables our model to understand complex relationships among heterogeneous features, a task often difficult for traditional methods. TTVAE facilitates the integration of interpolation within the latent space during the data generation process. Specifically, TTVAE is trained once, establishing a low-dimensional representation of real data, and then various latent interpolation methods can efficiently generate synthetic latent points. Through extensive experiments on diverse datasets, TTVAE consistently achieves state-of-the-art performance, highlighting its adaptability across different feature types and data sizes. This innovative approach, empowered by the attention mechanism and the integration of interpolation, addresses the complex challenges of tabular data synthesis, establishing TTVAE as a powerful solution. Alex X. Wang, Binh P. Nguyen |
AAAI | 1 |
| 2025 | Generative AI for Tabular Data Synthesis
Alex X. Wang, Binh P. Nguyen, Colin R. Simpson |
PAKDD (4) | 1 |
| 2025 | TTVAE: Transformer-based generative modeling for tabular data generation
Alex X. Wang, Binh P. Nguyen |
Artif. Intell. | 1 |
| 2025 | Blending is all you need: Data-centric ensemble synthetic data
Alex X. Wang, Colin R. Simpson, Binh P. Nguyen |
Inf. Sci. | 1 |
| 2025 | CTVAE: Contrastive Tabular Variational Autoencoder for imbalance dataabstractAbstract Class imbalance, where datasets often lack sufficient samples for minority classes, is a persistent challenge in machine learning. Existing solutions often generate synthetic data to mitigate this issue, but they typically struggle with complex data distributions, primarily because they focus on oversampling the minority class while neglecting the relationships with the majority class. To overcome these limitations, we propose the Contrastive Tabular Variational Autoencoder (CTVAE), which integrates conditional Variational Autoencoders with contrastive learning techniques. CTVAE excels at generating high-quality synthetic samples that capture the intricate data distributions of both minority and majority classes. Additionally, it can be seamlessly integrated with variants of the Synthetic Minority Oversampling Technique (SMOTE) for enhanced effectiveness. Experimental results demonstrate that CTVAE substantially improves classification performance on imbalanced datasets, offering a more robust and holistic solution to the class imbalance problem. Alex X. Wang, Minh Quang Le, Huu-Thanh Duong, Bay Nguyen Van, Binh P. Nguyen |
Knowl. Inf. Syst. | 1 |
| 2025 | Deterministic Autoencoder using Wasserstein loss for tabular data generationabstractTabular data generation is a complex task due to its distinctive characteristics and inherent complexities. While Variational Autoencoders have been adapted from the computer vision domain for tabular data synthesis, their reliance on non-deterministic latent space regularization introduces limitations. The stochastic nature of Variational Autoencoders can contribute to collapsed posteriors, yielding suboptimal outcomes and limiting control over the latent space. This characteristic also constrains the exploration of latent space interpolation. To address these challenges, we present the Tabular Wasserstein Autoencoder (TWAE), leveraging the deterministic encoding mechanism of Wasserstein Autoencoders. This characteristic facilitates a deterministic mapping of inputs to latent codes, enhancing the stability and expressiveness of our model's latent space. This, in turn, enables seamless integration with shallow interpolation mechanisms like the synthetic minority over-sampling technique (SMOTE) within the data generation process via deep learning. Specifically, TWAE is trained once to establish a low-dimensional representation of real data, and various latent interpolation methods efficiently generate synthetic latent points, achieving a balance between accuracy and efficiency. Extensive experiments consistently demonstrate TWAE's superiority, showcasing its versatility across diverse feature types and dataset sizes. This innovative approach, combining WAE principles with shallow interpolation, effectively leverages SMOTE's advantages, establishing TWAE as a robust solution for complex tabular data synthesis. Alex X. Wang, Binh P. Nguyen |
Neural Networks | 1 |
| 2024 | Comparative Analysis of Oversampling Techniques and Deep Learning for Imbalanced Tabular Data
Alex X. Wang, Colin R. Simpson, Binh P. Nguyen |
TENCON | 1 |
| 2024 | Enhancing public research on citizen data: An empirical investigation of data synthesis using Statistics New Zealand's Integrated Data InfrastructureabstractThe Integrated Data Infrastructure (IDI) in New Zealand is a critical asset that integrates citizen data from various public and private organizations for population-level analyses. However, access restrictions within the IDI environment present challenges for fully utilizing its potential. This study examines synthetic data as a potential solution, offering a comprehensive framework for generating customizable and easily implementable synthetic data. The evaluation of multiple data synthesis algorithms considers statistical similarity, machine learning utility, and privacy concerns. The findings reveal that distance-based algorithms, like SMOTE, strike a balance between accuracy and computational cost, making them suitable for IDI. The study also identifies the need for a clear release guide for micro-level synthetic data and proposes exploring a fully automatic data evaluation pipeline in future research. Additionally, the study highlights opportunities enabled by synthetic data, such as familiarization with administrative datasets, reproducibility of studies, pilot analyses, and enhanced cross-domain collaboration. Overall, the proposed framework and findings offer valuable insights and guidance for synthetic data projects within the IDI, advancing synthetic data privacy research and facilitating reproducibility, collaboration, and data sharing in the IDI ecosystem. Alex X. Wang, Stefanka S. Chukova, Andrew Sporle, Barry J. Milne, Colin R. Simpson, Binh P. Nguyen |
Inf. Process. Manag. | 1 |
| 2023 | Ensemble k-nearest neighbors based on centroid displacementabstractk-nearest neighbors (k-NN) is a well-known classification algorithm that is widely used in different domains. Despite its simplicity, effectiveness and robustness, k-NN is limited by the use of the Euclidean distance as the similarity metric, the arbitrarily selected neighborhood size k, the computational challenge of high-dimensional data, and the use of the simple majority voting rule in class determination. We sought to address the last issue and proposed the Centroid Displacement-based k-NN algorithm, where centroid displacement is used for class determination. This paper presents a simple yet efficient variant of our previous work, named Ensemble Centroid Displacement-based k-NN, which leverages the homogeneity of the nearest neighbors of test instances. Extensive experiments on various real and synthetic datasets were conducted to show the effectiveness and robustness of the proposed algorithm. Our experimental results demonstrate that the proposed algorithm is able to enhance the classification performance of the standard k-NN algorithm and its variants and also improve the computational efficiency. The performance of our algorithm was consistent and robust for both balanced and imbalanced datasets. Alex X. Wang, Stefanka S. Chukova, Binh P. Nguyen |
Inf. Sci. | 1 |
| 2022 | Implementation and Analysis of Centroid Displacement-Based k-Nearest Neighbors
Alex X. Wang, Stefanka S. Chukova, Binh P. Nguyen |
ADMA (1) | 1 |