EDBT 2026 Demo / reviewers in the wild / expert
Kien Do
dblp:185/0836
· DBLP profile ↗
8ranked-venue papers in the field
3as first author
5since 2021 · last 2025
0000-0002-0119-122XORCID · verified
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 8 (3 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Bidirectional Diffusion Bridge ModelsabstractDiffusion bridges have shown potential in paired image-to-image (I2I) translation tasks. However, existing methods are limited by their unidirectional nature, requiring separate models for forward and reverse translations. This not only doubles the computational cost but also restricts their practicality. In this work, we introduce the Bidirectional Diffusion Bridge Model (BDBM), a scalable approach that facilitates bidirectional translation between two coupled distributions using a single network. BDBM leverages the Chapman-Kolmogorov Equation for bridges, enabling it to model data distribution shifts across timesteps in both forward and backward directions by exploiting the interchangeability of the initial and target timesteps within this framework. Notably, when the marginal distribution given endpoints is Gaussian, BDBM's transition kernels in both directions possess analytical forms, allowing for efficient learning with a single network. We demonstrate the connection between BDBM and existing bridge methods, such as Doob's h-transform and variational approaches, and highlight its advantages. Extensive experiments on high-resolution I2I translation tasks demonstrate that BDBM not only enables bidirectional translation with minimal additional cost but also outperforms state-of-the-art bridge models. Our source code is available at https://github.com/kvmduc/BDBM. Duc Kieu, Kien Do, Toan Nguyen 0004, Dang Nguyen 0002, Thin Nguyen |
KDD (2) | 2 |
| 2025 | Defense Against Multi-target Multi-trigger Backdoor Attacks
Haripriya Harikumar, Santu Rana, Kien Do, Sunil Gupta 0001, Wei Zong, Willy Susilo, Svetha Venkatesh |
PAKDD (6) | 3 |
| 2024 | Generating Realistic Tabular Data with Large Language ModelsabstractWhile most generative models show achievements in image data generation, few are developed for tabular data generation. Recently, due to success of large language models (LLM) in diverse tasks, they have also been used for tabular data generation. However, these methods do not capture the correct correlation between the features and the target variable, hindering their applications in downstream predictive tasks. To address this problem, we propose a LLM-based method with three important improvements to correctly capture the ground-truth feature-class correlation in the real data. First, we propose a novel permutation strategy for the input data in the fine-tuning phase. Second, we propose a feature-conditional sampling approach to generate synthetic samples. Finally, we generate the labels by constructing prompts based on the generated samples to query our fine-tuned LLM. Our extensive experiments show that our method significantly outperforms 10 SOTA baselines on 20 datasets in downstream tasks. It also produces highly realistic synthetic samples in terms of quality and diversity. More importantly, classifiers trained with our synthetic data can even compete with classifiers trained with the original data on half of the benchmark datasets, which is a significant achievement in tabular data generation. Dang Nguyen 0002, Sunil Gupta 0001, Kien Do, Thin Nguyen, Svetha Venkatesh |
ICDM | 3 |
| 2024 | Improving Diversity in Black-Box Few-Shot Knowledge Distillation
Tri-Nhan Vo, Dang Nguyen 0002, Kien Do, Sunil Gupta 0001 |
ECML/PKDD (2) | 3 |
| 2023 | Causal Inference via Style Transfer for Out-of-distribution GeneralisationabstractOut-of-distribution (OOD) generalisation aims to build a model that can generalise well on an unseen target domain using knowledge from multiple source domains. To this end, the model should seek the causal dependence between inputs and labels, which may be determined by the semantics of inputs and remain invariant across domains. However, statistical or non-causal methods often cannot capture this dependence and perform poorly due to not considering spurious correlations learnt from model training via unobserved confounders. A well-known existing causal inference method like back-door adjustment cannot be applied to remove spurious correlations as it requires the observation of confounders. In this paper, we propose a novel method that effectively deals with hidden confounders by successfully implementing front-door adjustment (FA). FA requires the choice of a mediator, which we regard as the semantic information of images that helps access the causal mechanism without the need for observing confounders. Further, we propose to estimate the combination of the mediator with other observed images in the front-door formula via style transfer algorithms. Our use of style transfer to estimate FA is novel and sensible for OOD generalisation, which we justify by extensive experimental results on widely used benchmark datasets. Toan Nguyen 0004, Kien Do, Duc Thanh Nguyen, Bao Duong, Thin Nguyen |
KDD | 2 |
| 2019 | Graph Transformation Policy Network for Chemical Reaction PredictionabstractWe address a fundamental problem in chemistry known as chemical reaction product prediction. Our main insight is that the input reactant and reagent molecules can be jointly represented as a graph, and the process of generating product molecules from reactant molecules can be formulated as a sequence of graph transformations. To this end, we propose Graph Transformation Policy Network (GTPN) - a novel generic method that combines the strengths of graph neural networks and reinforcement learning to learn reactions directly from data with minimal chemical knowledge. Compared to previous methods, GTPN has some appealing properties such as: end-to-end learning, and making no assumption about the length or the order of graph transformations. In order to guide model search through the complex discrete space of sets of bond changes effectively, we extend the standard policy gradient loss by adding useful constraints. Evaluation results show that GTPN improves the top-1 accuracy over the current state-of-the-art method by about 3% on the large USPTO dataset. Kien Do, Truyen Tran 0001, Svetha Venkatesh |
KDD | 1 |
| 2018 | Energy-based anomaly detection for mixed data
Kien Do, Truyen Tran 0001, Svetha Venkatesh |
Knowl. Inf. Syst. | 1 |
| 2016 | Outlier Detection on Mixed-Type Data: An Energy-Based Approach
Kien Do, Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh |
ADMA | 1 |