Colin R. Simpson

dblp:254/4736 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2026
0000-0002-5194-8083ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Edge-updating graph neural networks for modeling feature interactions in tabular data
abstract
We proposed a message-passing graph neural network (GNN) based on graph isomorphism network (GIN) for learning on tabular data. Fully-connected, unweighted feature graphs were constructed from tabular data using contextual feature encodings for numerical and categorical features. A classification node was added to the feature graph to represent the entire graph during inference. Feature interactions were modeled using the proposed architecture, which used a neural network to learn edge attributes while leveraging residual connections for both node and edge updates to alleviate oversmoothing, a commonly-faced problem in GNNs. Our model was evaluated on 12 publicly available datasets and achieved the best mean rank among 6 tabular deep learning and GNN models. Furthermore, since gradient-boosted decision trees are considered to be state-of-the-art models for tabular data, we compared our model to XGBoost and CatBoost and found that our model outperformed both models in all datasets when using default hyperparameters and achieved the best results in 8 datasets when using tuned hyperparameters. A further comparison was made against 5 commonly used or recently-proposed GNNs to investigate the effectiveness of our model, in which our model achieved the top result in all datasets.
Pimwipa Charuthamrong, Colin R. Simpson, Binh P. Nguyen
Neural Networks2
2025 Generative AI for Tabular Data Synthesis
Alex X. Wang, Binh P. Nguyen, Colin R. Simpson
PAKDD (4)3
2025 Blending is all you need: Data-centric ensemble synthetic data
Alex X. Wang, Colin R. Simpson, Binh P. Nguyen
Inf. Sci.2
2024 Comparative Analysis of Oversampling Techniques and Deep Learning for Imbalanced Tabular Data
Alex X. Wang, Colin R. Simpson, Binh P. Nguyen
TENCON2
2024 Enhancing public research on citizen data: An empirical investigation of data synthesis using Statistics New Zealand's Integrated Data Infrastructure
abstract
The Integrated Data Infrastructure (IDI) in New Zealand is a critical asset that integrates citizen data from various public and private organizations for population-level analyses. However, access restrictions within the IDI environment present challenges for fully utilizing its potential. This study examines synthetic data as a potential solution, offering a comprehensive framework for generating customizable and easily implementable synthetic data. The evaluation of multiple data synthesis algorithms considers statistical similarity, machine learning utility, and privacy concerns. The findings reveal that distance-based algorithms, like SMOTE, strike a balance between accuracy and computational cost, making them suitable for IDI. The study also identifies the need for a clear release guide for micro-level synthetic data and proposes exploring a fully automatic data evaluation pipeline in future research. Additionally, the study highlights opportunities enabled by synthetic data, such as familiarization with administrative datasets, reproducibility of studies, pilot analyses, and enhanced cross-domain collaboration. Overall, the proposed framework and findings offer valuable insights and guidance for synthetic data projects within the IDI, advancing synthetic data privacy research and facilitating reproducibility, collaboration, and data sharing in the IDI ecosystem.
Alex X. Wang, Stefanka S. Chukova, Andrew Sporle, Barry J. Milne, Colin R. Simpson, Binh P. Nguyen
Inf. Process. Manag.5