EDBT 2026 Demo / reviewers in the wild / expert
Thin Nguyen
dblp:77/8172
· DBLP profile ↗
37ranked-venue papers in the field
14as first author
13since 2021 · last 2025
0000-0003-3467-8963ORCID · verified
Domains — venue-derived; a paper can count in several
Data Mining & Knowledge Discovery · 24 (4 first)Information Retrieval & Web Search · 12 (10 first)Database Systems & Data Management · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Neural Autoregressive Flows for Markov Boundary LearningabstractRecovering Markov boundary-the minimal set of variables that maximizes predictive performance for a response variable-is crucial in many applications. While recent advances improve upon traditional constraint-based techniques by scoring local causal structures, they still rely on nonparametric estimators and heuristic searches, lacking theoretical guarantees for reliability. This paper investigates a framework for efficient Markov boundary discovery by integrating conditional entropy from information theory as a scoring criterion. We design a novel masked autoregressive network to capture complex dependencies. A parallelizable greedy search strategy in polynomial time is proposed, supported by analytical evidence. We also discuss how initializing a graph with learned Markov boundaries accelerates the convergence of causal discovery. Comprehensive evaluations on real-world and synthetic datasets demonstrate the scalability and superior performance of our method in both Markov boundary discovery and causal discovery tasks. Bao Duong, Viet Huynh, Thin Nguyen |
ICDM | 4 |
| 2025 | Bidirectional Diffusion Bridge ModelsabstractDiffusion bridges have shown potential in paired image-to-image (I2I) translation tasks. However, existing methods are limited by their unidirectional nature, requiring separate models for forward and reverse translations. This not only doubles the computational cost but also restricts their practicality. In this work, we introduce the Bidirectional Diffusion Bridge Model (BDBM), a scalable approach that facilitates bidirectional translation between two coupled distributions using a single network. BDBM leverages the Chapman-Kolmogorov Equation for bridges, enabling it to model data distribution shifts across timesteps in both forward and backward directions by exploiting the interchangeability of the initial and target timesteps within this framework. Notably, when the marginal distribution given endpoints is Gaussian, BDBM's transition kernels in both directions possess analytical forms, allowing for efficient learning with a single network. We demonstrate the connection between BDBM and existing bridge methods, such as Doob's h-transform and variational approaches, and highlight its advantages. Extensive experiments on high-resolution I2I translation tasks demonstrate that BDBM not only enables bidirectional translation with minimal additional cost but also outperforms state-of-the-art bridge models. Our source code is available at https://github.com/kvmduc/BDBM. Duc Kieu, Kien Do, Toan Nguyen 0004, Dang Nguyen 0002, Thin Nguyen |
KDD (2) | 5 |
| 2025 | Amortized Conditional Independence Testing
Bao Duong, Nu Hoang, Thin Nguyen |
PAKDD (1) | 3 |
| 2025 | Clustering-Based Meta Bayesian Optimization with Theoretical Guarantee
Viet Huynh, Binh Tran, Tri Pham, Tin Huynh, Thin Nguyen |
PAKDD (3) | 6 |
| 2024 | Generating Realistic Tabular Data with Large Language ModelsabstractWhile most generative models show achievements in image data generation, few are developed for tabular data generation. Recently, due to success of large language models (LLM) in diverse tasks, they have also been used for tabular data generation. However, these methods do not capture the correct correlation between the features and the target variable, hindering their applications in downstream predictive tasks. To address this problem, we propose a LLM-based method with three important improvements to correctly capture the ground-truth feature-class correlation in the real data. First, we propose a novel permutation strategy for the input data in the fine-tuning phase. Second, we propose a feature-conditional sampling approach to generate synthetic samples. Finally, we generate the labels by constructing prompts based on the generated samples to query our fine-tuned LLM. Our extensive experiments show that our method significantly outperforms 10 SOTA baselines on 20 datasets in downstream tasks. It also produces highly realistic synthetic samples in terms of quality and diversity. More importantly, classifiers trained with our synthetic data can even compete with classifiers trained with the original data on half of the benchmark datasets, which is a significant achievement in tabular data generation. Dang Nguyen 0002, Sunil Gupta 0001, Kien Do, Thin Nguyen, Svetha Venkatesh |
ICDM | 4 |
| 2024 | Robust Estimation of Causal Heteroscedastic Noise ModelsabstractDistinguishing the cause and effect from bivariate observational data is the foundational problem that finds applications in many scientific disciplines. One solution to this problem is assuming that cause and effect are generated from a structural causal model, enabling identification of the causal direction after estimating the model in each direction. The heteroscedastic noise model is a type of structural causal model where the cause can contribute to both the mean and variance of the noise. Current methods for estimating heteroscedas-tic noise models choose the Gaussian likelihood as the optimization objective which can be suboptimal and unstable when the data has a non-Gaussian distribution. To address this limitation, we propose a novel approach to estimating this model with Student's t-distribution, which is known for its robustness in accounting for sampling variability with smaller sample sizes and extreme values without significantly altering the overall distribution shape. This adaptability is beneficial for capturing the parameters of the noise distribution in het-eroscedastic noise models. Our empirical evaluations demonstrate that our estimators are more robust and achieve better overall performance across synthetic and real benchmarks. Quang-Duy Tran, Bao Duong, Phuoc Nguyen, Thin Nguyen |
SDM | 4 |
| 2024 | Normalizing flows for conditional independence testingabstractAbstract Detecting conditional independencies plays a key role in several statistical and machine learning tasks, especially in causal discovery algorithms, yet it remains a highly challenging problem due to dimensionality and complex relationships presented in data. In this study, we introduce LCIT (Latent representation-based Conditional Independence Test)—a novel method for conditional independence testing based on representation learning. Our main contribution involves a hypothesis testing framework in which to test for the independence between X and Y given Z, we first learn to infer the latent representations of target variables X and Y that contain no information about the conditioning variable Z. The latent variables are then investigated for any significant remaining dependencies, which can be performed using a conventional correlation test. Moreover, LCIT can also handle discrete and mixed-type data in general by converting discrete variables into the continuous domain via variational dequantization. The empirical evaluations show that LCIT outperforms several state-of-the-art baselines consistently under different evaluation metrics, and is able to adapt really well to both nonlinear, high-dimensional, and mixed data settings on a diverse collection of synthetic and real data sets. Bao Duong, Thin Nguyen |
Knowl. Inf. Syst. | 2 |
| 2024 | Constraining acyclicity of differentiable Bayesian structure learning with topological orderingabstractAbstract Distributional estimates in Bayesian approaches in structure learning have advantages compared to the ones performing point estimates when handling epistemic uncertainty. Differentiable methods for Bayesian structure learning have been developed to enhance the scalability of the inference process and are achieving optimistic outcomes. However, in the differentiable continuous setting, constraining the acyclicity of learned graphs emerges as another challenge. Various works utilize post-hoc penalization scores to impose this constraint which cannot assure acyclicity. The topological ordering of the variables is one type of prior knowledge that contains valuable information about the acyclicity of a directed graph. In this work, we propose a framework to guarantee the acyclicity of inferred graphs by integrating the information from the topological ordering into the inference process. Our integration framework does not interfere with the differentiable inference process while being able to strictly assure the acyclicity of learned graphs and reduce the inference complexity. Our extensive empirical experiments on both synthetic and real data have demonstrated the effectiveness of our approach with preferable results compared to related Bayesian approaches. Quang-Duy Tran, Phuoc Nguyen, Bao Duong, Thin Nguyen |
Knowl. Inf. Syst. | 4 |
| 2023 | Differentiable Bayesian Structure Learning with Acyclicity AssuranceabstractScore-based approaches in the structure learning task are thriving because of their scalability. Continuous relaxation has been the key reason for this advancement. Despite achieving promising outcomes, most of these methods are still struggling to ensure that the graphs generated from the latent space are acyclic by minimizing a defined score. There has also been another trend of permutation-based approaches, which concern the search for the topological ordering of the variables in the directed acyclic graph in order to limit the search space of the graph. In this study, we propose an alternative approach for strictly constraining the acyclicty of the graphs with an integration of the knowledge from the topological orderings. Our approach can reduce inference complexity while ensuring the structures of the generated graphs to be acyclic. Our empirical experiments with simulated and real-world data show that our approach can outperform related Bayesian score-based approaches. Quang-Duy Tran, Phuoc Nguyen, Bao Duong, Thin Nguyen |
ICDM | 4 |
| 2023 | Causal Inference via Style Transfer for Out-of-distribution GeneralisationabstractOut-of-distribution (OOD) generalisation aims to build a model that can generalise well on an unseen target domain using knowledge from multiple source domains. To this end, the model should seek the causal dependence between inputs and labels, which may be determined by the semantics of inputs and remain invariant across domains. However, statistical or non-causal methods often cannot capture this dependence and perform poorly due to not considering spurious correlations learnt from model training via unobserved confounders. A well-known existing causal inference method like back-door adjustment cannot be applied to remove spurious correlations as it requires the observation of confounders. In this paper, we propose a novel method that effectively deals with hidden confounders by successfully implementing front-door adjustment (FA). FA requires the choice of a mediator, which we regard as the semantic information of images that helps access the causal mechanism without the need for observing confounders. Further, we propose to estimate the combination of the mediator with other observed images in the front-door formula via style transfer algorithms. Our use of style transfer to estimate FA is novel and sensible for OOD generalisation, which we justify by extensive experimental results on widely used benchmark datasets. Toan Nguyen 0004, Kien Do, Duc Thanh Nguyen, Bao Duong, Thin Nguyen |
KDD | 5 |
| 2022 | Efficient Classification with Counterfactual Reasoning and Active Learning
Azhar Mohammed, Dang Nguyen 0002, Bao Duong, Thin Nguyen |
ACIIDS (1) | 4 |
| 2022 | Handling Missing Data with Markov Boundary
Azhar Mohammed, Dang Nguyen 0002, Bao Duong, Melanie Nichols, Thin Nguyen |
ADMA (1) | 5 |
| 2022 | Conditional Independence Testing via Latent Representation LearningabstractDetecting conditional independencies plays a key role in several statistical and machine learning tasks, especially in causal discovery algorithms, yet it remains a highly challenging problem due to dimensionality and complex relationships presented in data. In this study, we introduce LCIT (Latent representation based Conditional Independence Test) -a novel method for conditional independence testing based on representation learning. Our main contribution involves a hypothesis testing framework in which to test for the independence between X and Y given Z, we first learn to infer the latent representations of target variables X and Y that contain no information about the conditioning variable Z. The latent variables are then investigated for any significant remaining dependencies, which can be performed using a conventional correlation test. The empirical evaluations show that LCIT outperforms several state-of-the-art baselines consistently under different evaluation metrics, and is able to adapt really well to both non-linear and high-dimensional settings on a diverse collection of synthetic and real data sets. Bao Duong, Thin Nguyen |
ICDM | 2 |
| 2020 | Computational Methods for Predicting Autism Spectrum Disorder from Gene Expression Data
Junpeng Zhang 0001, Thin Nguyen, Buu Minh Thanh Truong, Lin Liu 0003, Jiuyong Li, Thuc Duy Le |
ADMA | 2 |
| 2018 | Differentially Private Prescriptive AnalyticsabstractPrivacy preservation is important. Prescriptive analytics is a method to extract corrective actions to avoid undesirable outcomes. We propose a privacy preserving prescriptive analytics algorithm to protect the data used during the construction of the prescriptive analytics algorithm. We use differential privacy mechanism to achieve strong privacy guarantee. Differential privacy mechanism requires computation of sensitivity: maximum change in the output between two training datasets, which is differed by only one instance. The main challenge we addressed is the computation of sensitivity of the prescription vector. In absence of any analytical form, we construct a nested global optimization problem to compute the sensitivity. We solve the optimization problem using constrained Bayesian optimization, as the nested structure makes the objective function expensive. We demonstrate our algorithm on two real world datasets and observe that the prescription vectors remains useful even after making them private. Haripriya Harikumar, Santu Rana, Sunil Gupta 0001, Thin Nguyen, M. R. Kaimal 0001, Svetha Venkatesh |
ICDM | 4 |
| 2018 | Prescriptive Analytics Through Constrained Bayesian Optimization
Haripriya Harikumar, Santu Rana, Sunil Gupta 0001, Thin Nguyen, M. R. Kaimal 0001, Svetha Venkatesh |
PAKDD (1) | 4 |
| 2018 | Jointly Predicting Affective and Mental Health Scores Using Deep Neural Networks of Visual Cues on the Web
Van Nguyen 0002, Thin Nguyen, Mark E. Larsen, Bridianne O'Dea, Duc Thanh Nguyen, Trung Le 0001, Dinh Q. Phung, Svetha Venkatesh, Helen Christensen |
WISE (2) | 3 |
| 2017 | Animal Recognition and Identification with Deep Convolutional Neural Networks for Automated Wildlife MonitoringabstractEfficient and reliable monitoring of wild animals in their natural habitats is essential to inform conservation and management decisions. Automatic covert cameras or "camera traps" are being an increasingly popular tool for wildlife monitoring due to their effectiveness and reliability in collecting data of wildlife unobtrusively, continuously and in large volume. However, processing such a large volume of images and videos captured from camera traps manually is extremely expensive, time-consuming and also monotonous. This presents a major obstacle to scientists and ecologists to monitor wildlife in an open environment. Leveraging on recent advances in deep learning techniques in computer vision, we propose in this paper a framework to build automated animal recognition in the wild, aiming at an automated wildlife monitoring system. In particular, we use a single-labeled dataset from Wildlife Spotter project, done by citizen scientists, and the state-of-the-art deep convolutional neural network architectures, to train a computational system capable of filtering animal images and identifying species automatically. Our experimental results achieved an accuracy at 96.6% for the task of detecting images containing animal, and 90.4% for identifying the three most common species among the set of images of wild animals taken in South-central Victoria, Australia, demonstrating the feasibility of building fully automated wildlife observation. This, in turn, can therefore speed up research findings, construct more efficient citizen sciencebased monitoring systems and subsequent management decisions, having the potential to make significant impacts to the world of ecology and trap camera images analysis. Sarah J. Maclagan, Tu Dinh Nguyen, Thin Nguyen, Paul Flemons, Kylie Andrews, Euan G. Ritchie, Dinh Q. Phung |
DSAA | 4 |
| 2017 | Estimating Support Scores of Autism Communities in Large-Scale Web Information Systems
Thin Nguyen, Svetha Venkatesh, Dinh Q. Phung |
WISE (1) | 1 |
| 2016 | Understanding Behavioral Differences Between Short and Long-Term Drinking Abstainers from Social Media
Haripriya Harikumar, Thin Nguyen, Sunil Gupta 0001, Santu Rana, M. R. Kaimal 0001, Svetha Venkatesh |
ADMA | 2 |
| 2016 | Extracting Key Challenges in Achieving Sobriety Through Shared Subspace Learning
Haripriya Harikumar, Thin Nguyen, Santu Rana, Sunil Gupta 0001, M. R. Kaimal 0001, Svetha Venkatesh |
ADMA | 2 |
| 2016 | Textual Cues for Online Depression in Community and Personal Settings
Thin Nguyen, Svetha Venkatesh, Dinh Q. Phung |
ADMA | 1 |
| 2016 | Discriminative Cues for Different Stages of Smoking Cessation in Online Community
Thin Nguyen, Ron Borland, John Yearwood, Hua-Hie Yong, Svetha Venkatesh, Dinh Q. Phung |
WISE (2) | 1 |
| 2016 | Large-Scale Stylistic Analysis of Formality in Academia and Social Media
Thin Nguyen, Svetha Venkatesh, Dinh Q. Phung |
WISE (2) | 1 |
| 2015 | Nonparametric discovery of online mental health-related communitiesabstractPeople are increasingly using social media, especially online communities, to discuss mental health issues and seek supports. Understanding topics, interaction, sentiment and clustering structures of these communities informs important aspects of mental health. It can potentially add knowledge to the underlying cognitive dynamics, mood swings patterns, shared interests, and interaction. There has been growing research interest in analyzing online mental health communities; however sentiment analysis of these communities has been largely under-explored. This study presents an analysis of online Live Journal communities with and without mental health-related conditions including depression and autism. Latent topics for mood tags, affective words, and generic words in the content of the posts made in these communities were learned using nonparametric topic modelling. These representations were then input into a nonparametric clustering to discover meta-groups among the communities. The best performance results can be achieved on clustering communities with latent mood-based representation for such communities. The study also found significant differences in usage latent topics for mood tags and affective features between online communities with and without affective disorders. The findings reveal useful insights into hyper-group detection of online mental health-related communities. Bo Dao, Thin Nguyen, Svetha Venkatesh, Dinh Q. Phung |
DSAA | 2 |
| 2015 | Differentiating Sub-groups of Online Depression-Related Communities Using Textual Cues
Thin Nguyen, Bridianne O'Dea, Mark E. Larsen, Dinh Q. Phung, Svetha Venkatesh, Helen Christensen |
WISE (2) | 1 |
| 2014 | Analysis of circadian rhythms from online communities of individuals with affective disordersabstractThe circadian system regulates 24 hour rhythms in biological creatures. It impacts mood regulation. The disruptions of circadian rhythms cause destabilization in individuals with affective disorders, such as depression and bipolar disorders. Previous work has examined the role of the circadian system on effects of light interactions on mood-related systems, the effects of light manipulation on brain, the impact of chronic stress on rhythms. However, such studies have been conducted in small, preselected populations. The deluge of data is now changing the landscape of research practice. The unprecedented growth of social media data allows one to study individual behavior across large and diverse populations. In particular, individuals with affective disorders from online communities have not been examined rigorously. In this paper, we aim to use social media as a sensor to identify circadian patterns for individuals with affective disorders in online communities.We use a large scale study cohort of data collecting from online affective disorder communities. We analyze changes in hourly, daily, weekly and seasonal affect of these clinical groups in contrast with control groups of general communities. By comparing the behaviors between the clinical groups and the control groups, our findings show that individuals with affective disorders show a significant distinction in their circadian rhythms across the online activity. The results shed light on the potential of using social media for identifying diurnal individual variation in affective state, providing key indicators and risk factors for noninvasive wellbeing monitoring and prediction. Bo Dao, Thin Nguyen, Svetha Venkatesh, Dinh Q. Phung |
DSAA | 2 |
| 2014 | Effect of Mood, Social Connectivity and Age in Online Depression Community via Topic and Linguistic Analysis
Bo Dao, Thin Nguyen, Dinh Q. Phung, Svetha Venkatesh |
WISE (1) | 2 |
| 2014 | Affective, Linguistic and Topic Patterns in Online Autism Communities
Thin Nguyen, Thi V. Duong, Dinh Q. Phung, Svetha Venkatesh |
WISE (2) | 1 |
| 2014 | iPoll: Automatic Polling Using Online Search
Thin Nguyen, Dinh Q. Phung, Wei Luo 0001, Truyen Tran 0001, Svetha Venkatesh |
WISE (1) | 1 |
| 2014 | Mood sensing from social media texts and its applications
Thin Nguyen, Dinh Q. Phung, Brett Adams, Svetha Venkatesh |
Knowl. Inf. Syst. | 1 |
| 2013 | Online Social Capital: Mood, Topical and Psycholinguistic Analysis
Thin Nguyen, Bo Dao, Dinh Q. Phung, Svetha Venkatesh, Michael Berk |
ICWSM | 1 |
| 2013 | Event extraction using behaviors of sentiment signals and burst structure in social media
Thin Nguyen, Dinh Q. Phung, Brett Adams, Svetha Venkatesh |
Knowl. Inf. Syst. | 1 |
| 2012 | A Sentiment-Aware Approach to Community Formation in Social Media
Thin Nguyen, Dinh Q. Phung, Brett Adams, Svetha Venkatesh |
ICWSM | 1 |
| 2011 | Towards Discovery of Influence and Personality Traits through Social Link Prediction
Thin Nguyen, Dinh Q. Phung, Brett Adams, Svetha Venkatesh |
ICWSM | 1 |
| 2011 | Prediction of Age, Sentiment, and Connectivity from Social Media Text
Thin Nguyen, Dinh Q. Phung, Brett Adams, Svetha Venkatesh |
WISE | 1 |
| 2010 | Classification and Pattern Discovery of Mood in Weblogs
Thin Nguyen, Dinh Q. Phung, Brett Adams, Truyen Tran 0001, Svetha Venkatesh |
PAKDD (2) | 1 |