VLDB 2026 Research / reviewers in the wild / expert
Viet Huynh
dblp:161/2718
· DBLP profile ↗
20ranked-venue papers
8as first author
9since 2021 · last 2026
0000-0003-1308-1164ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 7 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Efficient and Effective Evaluator for Text2SQL Models on Unseen and Unlabeled DataabstractRecent advances in large language models have strengthened Text2SQL systems that translate natural language questions into database queries. A persistent deployment challenge is to assess a newly trained Text2SQL system on an unseen and unlabeled dataset when no verified answers are available. This situation arises frequently because database content and structure evolve, privacy policies slow manual review, and carefully written SQL labels are costly and time-consuming. Without timely evaluation, organizations cannot approve releases or detect failures early. FusionSQL addresses this gap by working with any Text2SQL models and estimating accuracy without reference labels, allowing teams to measure quality on unseen and unlabeled datasets. It analyzes patterns in the system's own outputs to characterize how the target dataset differs from the material used during training. FusionSQL supports pre-release checks, continuous monitoring of new databases, and detection of quality decline. Experiments across diverse application settings and question types show that FusionSQL closely follows actual accuracy and reliably signals emerging issues. Our code is available at https://github.com/phkhanhtrinh23/FusionSQL. Trinh Pham, Thanh Tam Nguyen, Viet Huynh, Hongzhi Yin, Nguyen Quoc Viet Hung |
ICDE | 3 |
| 2026 | ADCC-Bench: A Benchmark Framework for Anomaly Detection in Cryptocurrency Transactions
Tan-Gia-Quoc Pham, Huynh Quoc Khanh, Nguyen Anh Khoa, Viet Huynh, Van-Hau Pham, Phan The Duy |
IWCMC | 4 |
| 2025 | Neural Autoregressive Flows for Markov Boundary LearningabstractRecovering Markov boundary-the minimal set of variables that maximizes predictive performance for a response variable-is crucial in many applications. While recent advances improve upon traditional constraint-based techniques by scoring local causal structures, they still rely on nonparametric estimators and heuristic searches, lacking theoretical guarantees for reliability. This paper investigates a framework for efficient Markov boundary discovery by integrating conditional entropy from information theory as a scoring criterion. We design a novel masked autoregressive network to capture complex dependencies. A parallelizable greedy search strategy in polynomial time is proposed, supported by analytical evidence. We also discuss how initializing a graph with learned Markov boundaries accelerates the convergence of causal discovery. Comprehensive evaluations on real-world and synthetic datasets demonstrate the scalability and superior performance of our method in both Markov boundary discovery and causal discovery tasks. Bao Duong, Viet Huynh, Thin Nguyen |
ICDM | 3 |
| 2025 | Clustering-Based Meta Bayesian Optimization with Theoretical Guarantee
Viet Huynh, Binh Tran, Tri Pham, Tin Huynh, Thin Nguyen |
PAKDD (3) | 2 |
| 2023 | Rapid Identification of Protein Formulations with Bayesian OptimisationabstractProtein formulation is a critical aspect of the pharmaceutical industry which aims to improve the efficacy and the safety of the active drug ingredients during the storage, transportation and administration of the drug. Buffer screening is the first stage of this formulation process that selects the promising combinations of buffer and excipients that can help maintain both the stability and efficacy of the drug. In this paper, we propose an interactive Bayesian Optimisation approach that streamlines the buffer screening process and reduces the number of experiments needed to identify an optimal combination of buffer and excipients. Our approach employs two novel formulations of the (multi-buffer) optimisation problem: (i) one that unifies all buffers into a single Bayesian Optimisation framework, and (ii) the other that performs meta-learning to aggregate important excipient information over multiple buffers, in order to predict the most promising buffer and excipients combination to sample next. Our experimental results show that the proposed approach can identify an optimal combination of buffer and excipients while minimising the number of experiments required, and demonstrate the potential of using Bayesian Optimisation to enhance the protein formulation process. Viet Huynh, Buser Say, Peter Vogel, Lucy Cao, Geoffrey I. Webb, Aldeida Aleti |
ICMLA | 1 |
| 2021 | Neural Topic Model via Optimal Transport
He Zhao 0001, Dinh Q. Phung, Viet Huynh, Trung Le 0001, Wray L. Buntine |
ICLR | 3 |
| 2021 | Optimal Transport for Deep Generative Models: State of the Art and Research ChallengesabstractOptimal transport has a long history in mathematics which was proposed by Gaspard Monge in the eighteenth century (Monge, 1781). However, until recently, advances in optimal transport theory pave the way for its use in the AI community, particularly for formulating deep generative models. In this paper, we provide a comprehensive overview of the literature in the field of deep generative models using optimal transport theory with an aim of providing a systematic review as well as outstanding problems and more importantly, open research opportunities to use the tools from the established optimal transport theory in the deep generative model domain. Viet Huynh, Dinh Q. Phung, He Zhao 0001 |
IJCAI | 1 |
| 2021 | Topic Modelling Meets Deep Neural Networks: A SurveyabstractTopic modelling has been a successful technique for text analysis for almost twenty years. When topic modelling met deep neural networks, there emerged a new and increasingly popular research area, neural topic models, with nearly a hundred models developed and a wide range of applications in neural language understanding such as text generation, summarisation and language models. There is a need to summarise research developments and discuss open problems and future directions. In this paper, we provide a focused yet comprehensive overview of neural topic models for interested researchers in the AI community, so as to facilitate them to navigate and innovate in this fast-growing research area. To the best of our knowledge, ours is the first review on this specific topic. He Zhao 0001, Dinh Q. Phung, Viet Huynh, Lan Du 0002, Wray L. Buntine |
IJCAI | 3 |
| 2021 | On efficient multilevel Clustering via Wasserstein distancesabstractWe propose a novel approach to the problem of multilevel clustering, which aims to simultaneously partition data in each group and discover grouping patterns among groups in a potentially large hierarchically structured corpus of data. Our method involves a joint optimization formulation over several spaces of discrete probability measures, which are endowed with Wasserstein distance metrics. We propose several variants of this problem, which admit fast optimization algorithms, by exploiting the connection to the problem of finding Wasserstein barycenters. Consistency properties are established for the estimates of both local and global clusters. Finally, experimental results with both synthetic and real data are presented to demonstrate the flexibility and scalability of the proposed approach. Viet Huynh, Nhat Ho, Nhan Dam, XuanLong Nguyen, Mikhail Yurochkin, Hung Hai Bui, Dinh Q. Phung |
J. Mach. Learn. Res. | 1 |
| 2020 | Stein Variational Gradient Descent with Variance ReductionabstractProbabilistic inference is a common and important task in statistical machine learning. The recently proposed Stein variational gradient descent (SVGD) is a generic Bayesian inference method that has been shown to be successfully applied in a wide range of contexts, especially in dealing with large datasets, where existing probabilistic inference methods have been known to be ineffective. In a large-scale data setting, SVGD employs the mini-batch strategy but its mini-batch estimator has large variance, hence compromising its estimation quality in practice. To this end, we propose in this paper a generic SVGD-based inference method that can significantly reduce the variance of mini-batch estimator when working with large datasets. Our experiments on 14 datasets show that the proposed method enjoys substantial and consistent improvements compared with baseline methods in binary classification task and its pseudo-online learning setting, and regression task. Furthermore, our framework is generic and applicable to a wide range of probabilistic inference problems such as in Bayesian neural networks and Markov random fields. Nhan Dam, Trung Le 0001, Viet Huynh, Dinh Q. Phung |
IJCNN | 3 |
| 2020 | OptiGAN: Generative Adversarial Networks for Goal Optimized Sequence GenerationabstractOne of the challenging problems in sequence generation tasks is the optimized generation of sequences with specific desired goals. Current sequential generative models mainly generate sequences to closely mimic the training data, without direct optimization of desired goals or properties specific to the task. We introduce OptiGAN, a generative model that incorporates both Generative Adversarial Networks (GAN) and Reinforcement Learning (RL) to optimize desired goal scores using policy gradients. We apply our model to text and real-valued sequence generation, where our model is able to achieve higher desired scores out-performing GAN and RL baselines, while not sacrificing output sample diversity. Mahmoud Hossam, Trung Le 0001, Viet Huynh, Michael Papasimeon, Dinh Q. Phung |
IJCNN | 3 |
| 2020 | OTLDA: A Geometry-aware Optimal Transport Approach for Topic ModelingabstractWe present an optimal transport framework for learning topics from textual data. While the celebrated Latent Dirichlet allocation (LDA) topic model and its variants have been applied to many disciplines, they mainly focus on word-occurrences and neglect to incorporate semantic regularities in language. Even though recent works have tried to exploit the semantic relationship between words to bridge this gap, however, these models which are usually extensions of LDA or Dirichlet Multinomial mixture (DMM) are tailored to deal effectively with either regular or short documents. The optimal transport distance provides an appealing tool to incorporate the geometry of word semantics into it. Moreover, recent developments on efficient computation of optimal transport distance also promote its application in topic modeling. In this paper we ground on optimal transport theory to naturally exploit the geometric structures of semantically related words in embedding spaces which leads to more interpretable learned topics. Comprehensive experiments illustrate that the proposed framework outperforms competitive approaches in terms of topic coherence on assorted text corpora which include both long and short documents. The representation of learned topic also leads to better accuracy on classification downstream tasks, which is considered as an extrinsic evaluation. Viet Huynh, He Zhao 0001, Dinh Q. Phung |
NeurIPS | 1 |
| 2019 | Probabilistic Multilevel Clustering via Composite Transportation DistanceabstractWe propose a novel probabilistic approach to multilevel clustering problems based on composite transportation distance, which is a variant of transportation distance where the underlying metric is Kullback-Leibler divergence. Our method involves solving a joint optimization problem over spaces of probability measures to simultaneously discover grouping structures within groups and among groups. By exploiting the connection of our method to the problem of finding composite transportation barycenters, we develop fast and efficient optimization algorithms even for potentially large-scale multilevel datasets. Finally, we present experimental results with both synthetic and real data to demonstrate the efficiency and scalability of the proposed approach. Nhat Ho, Viet Huynh, Dinh Q. Phung, Michael I. Jordan |
AISTATS | 2 |
| 2017 | Forward-Backward Smoothing for Hidden Markov Models of Point Pattern DataabstractThis paper considers a discrete-time sequential latent model for point pattern data, specifically a hidden Markov model (HMM) where each observation is an instantiation of a random finite set (RFS). This so-called RFS-HMM is worthy of investigation since point pattern data are ubiquitous in artificial intelligence and data science. We address the three basic problems typically encountered in such a sequential latent model, namely likelihood computation, hidden state inference, and parameter estimation. Moreover, we develop algorithms for solving these problems including forward-backward smoothing for likelihood computation and hidden state inference, and expectation-maximisation for parameter estimation. Simulation studies are used to demonstrate key properties of RFS-HMM, whilst real data in the domain of human dynamics are used to demonstrate its applicability. Nhan Dam, Dinh Q. Phung, Ba-Ngu Vo, Viet Huynh |
DSAA | 4 |
| 2017 | Multilevel Clustering via Wasserstein MeansabstractWe propose a novel approach to the problem of multilevel clustering, which aims to simultaneously partition data in each group and discover grouping patterns among groups in a potentially large hierarchically structured corpus of data. Our method involves a joint optimization formulation over several spaces of discrete probability measures, which are endowed with Wasserstein distance metrics. We propose a number of variants of this problem, which admit fast optimization algorithms, by exploiting the connection to the problem of finding Wasserstein barycenters. Consistency properties are established for the estimates of both local and global clusters. Finally, experiment results with both synthetic and real data are presented to demonstrate the flexibility and scalability of the proposed approach. Nhat Ho, XuanLong Nguyen, Mikhail Yurochkin, Hung Hai Bui, Viet Huynh, Dinh Q. Phung |
ICML | 5 |
| 2017 | Supervised Restricted Boltzmann Machines
Tu Dinh Nguyen, Dinh Q. Phung, Viet Huynh, Trung Le 0001 |
UAI | 3 |
| 2017 | Streaming clustering with Bayesian nonparametric models
Viet Huynh, Dinh Q. Phung |
Neurocomputing | 1 |
| 2016 | Scalable Nonparametric Bayesian Multilevel Clustering
Viet Huynh, Dinh Q. Phung, Svetha Venkatesh, XuanLong Nguyen, Matthew Hoffman 0001, Hung Hai Bui |
UAI | 1 |
| 2015 | Streaming Variational Inference for Dirichlet Process Mixtures
Viet Huynh, Dinh Q. Phung, Svetha Venkatesh |
ACML | 1 |
| 2015 | Learning Conditional Latent Structures from Multiple Data Sources
Viet Huynh, Dinh Q. Phung, XuanLong Nguyen, Svetha Venkatesh, Hung Hai Bui |
PAKDD (1) | 1 |