Dang Nguyen 0002

dblp:78/4136-2 · DBLP profile ↗
← Back
13ranked-venue papers in the field
7as first author
8since 2021 · last 2025
0000-0002-0401-988XORCID · verified

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 10 (5 first)Knowledge Engineering, Semantic Web & Information Systems · 2 (2 first)Database Systems & Data Management · 1
YearPublicationVenuePosition
2025 Bidirectional Diffusion Bridge Models
abstract
Diffusion bridges have shown potential in paired image-to-image (I2I) translation tasks. However, existing methods are limited by their unidirectional nature, requiring separate models for forward and reverse translations. This not only doubles the computational cost but also restricts their practicality. In this work, we introduce the Bidirectional Diffusion Bridge Model (BDBM), a scalable approach that facilitates bidirectional translation between two coupled distributions using a single network. BDBM leverages the Chapman-Kolmogorov Equation for bridges, enabling it to model data distribution shifts across timesteps in both forward and backward directions by exploiting the interchangeability of the initial and target timesteps within this framework. Notably, when the marginal distribution given endpoints is Gaussian, BDBM's transition kernels in both directions possess analytical forms, allowing for efficient learning with a single network. We demonstrate the connection between BDBM and existing bridge methods, such as Doob's h-transform and variational approaches, and highlight its advantages. Extensive experiments on high-resolution I2I translation tasks demonstrate that BDBM not only enables bidirectional translation with minimal additional cost but also outperforms state-of-the-art bridge models. Our source code is available at https://github.com/kvmduc/BDBM.
Duc Kieu, Kien Do, Toan Nguyen 0004, Dang Nguyen 0002, Thin Nguyen
KDD (2)4
2024 Generating Realistic Tabular Data with Large Language Models
abstract
While most generative models show achievements in image data generation, few are developed for tabular data generation. Recently, due to success of large language models (LLM) in diverse tasks, they have also been used for tabular data generation. However, these methods do not capture the correct correlation between the features and the target variable, hindering their applications in downstream predictive tasks. To address this problem, we propose a LLM-based method with three important improvements to correctly capture the ground-truth feature-class correlation in the real data. First, we propose a novel permutation strategy for the input data in the fine-tuning phase. Second, we propose a feature-conditional sampling approach to generate synthetic samples. Finally, we generate the labels by constructing prompts based on the generated samples to query our fine-tuned LLM. Our extensive experiments show that our method significantly outperforms 10 SOTA baselines on 20 datasets in downstream tasks. It also produces highly realistic synthetic samples in terms of quality and diversity. More importantly, classifiers trained with our synthetic data can even compete with classifiers trained with the original data on half of the benchmark datasets, which is a significant achievement in tabular data generation.
Dang Nguyen 0002, Sunil Gupta 0001, Kien Do, Thin Nguyen, Svetha Venkatesh
ICDM1
2024 Improving Diversity in Black-Box Few-Shot Knowledge Distillation
Tri-Nhan Vo, Dang Nguyen 0002, Kien Do, Sunil Gupta 0001
ECML/PKDD (2)2
2022 Efficient Classification with Counterfactual Reasoning and Active Learning
Azhar Mohammed, Dang Nguyen 0002, Bao Duong, Thin Nguyen
ACIIDS (1)2
2022 Handling Missing Data with Markov Boundary
Azhar Mohammed, Dang Nguyen 0002, Bao Duong, Melanie Nichols, Thin Nguyen
ADMA (1)2
2021 Knowledge Distillation with Distribution Mismatch
Dang Nguyen 0002, Sunil Gupta 0001, Trong Nguyen, Santu Rana, Phuoc Nguyen, Truyen Tran 0001, Ky Le, Shannon Ryan, Svetha Venkatesh
ECML/PKDD (2)1
2021 Fast Conditional Network Compression Using Bayesian HyperNetworks
Phuoc Nguyen, Truyen Tran 0001, Ky Le, Sunil Gupta 0001, Santu Rana, Dang Nguyen 0002, Trong Nguyen, Shannon Ryan, Svetha Venkatesh
ECML/PKDD (3)6
2021 Fairness improvement for black-box classifiers with Gaussian process
Dang Nguyen 0002, Sunil Gupta 0001, Santu Rana, Alistair Shilton, Svetha Venkatesh
Inf. Sci.1
2020 Bayesian Optimization with Missing Inputs
Phuc Luong, Dang Nguyen 0002, Sunil Gupta 0001, Santu Rana, Svetha Venkatesh
ECML/PKDD (2)2
2018 Trans2Vec: Learning Transaction Embedding via Items and Frequent Itemsets
Dang Nguyen 0002, Tu Dinh Nguyen, Wei Luo 0001, Svetha Venkatesh
PAKDD (3)1
2018 Sqn2Vec: Learning Sequence Representation via Sequential Patterns with a Gap Constraint
Dang Nguyen 0002, Wei Luo 0001, Tu Dinh Nguyen, Svetha Venkatesh, Dinh Q. Phung
ECML/PKDD (2)1
2018 Learning Graph Representation via Frequent Subgraphs
abstract
We propose a novel approach to learn distributed representation for graph data. Our idea is to combine a recently introduced neural document embedding model with a traditional pattern mining technique, by treating a graph as a document and frequent subgraphs as atomic units for the embedding process. Compared to the latest graph embedding methods, our proposed method offers three key advantages: fully unsupervised learning, entire-graph embedding, and edge label leveraging. We demonstrate our method on several datasets in comparison with a comprehensive list of up-to-date state-of-the-art baselines where we show its advantages for both classification and clustering tasks.
Dang Nguyen 0002, Wei Luo 0001, Tu Dinh Nguyen, Svetha Venkatesh, Dinh Q. Phung
SDM1
2015 A novel method for constrained class association rule mining
Dang Nguyen 0002, Loan T. T. Nguyen, Bay Vo, Tzung-Pei Hong
Inf. Sci.1