Khoat Than

dblp:118/4726 · DBLP profile ↗
← Back
39ranked-venue papers
5as first author
22since 2021 · last 2026
0000-0001-8615-2854ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 5 first-author · 19 since 2021Databases, data management, data science and information retrieval · 12 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Using Synthetic Data to Estimate the True Error is Theoretically and Practically Doable
Thanh Hai Hoang, Duy-Tung Nguyen, Hung The Tran, Khoat Than
Mach. Learn.4
2026 Generalized ordered Wasserstein distance for sequential data
Tung Doan 0001, Hiep To, Thi-Hong Vuong, Muriel Visani, Atsuhiro Takasu, Khoat Than
Pattern Recognit.6
2026 Enhancing visual feature attribution via weighted integrated gradients
Kien Tran Duc Tuan, Nguyen Trong Tam, Nguyen Hoang Son, Khoat Than
Pattern Recognit. Lett.4
2025 Boosting Multiple Views for pretrained-based Continual Learning
abstract
Recent research has shown that Random Projection (RP) can effectively improve the performance of pre-trained models in Continual learning (CL). The authors hypothesized that using RP to map features onto a higher-dimensional space can make them more linearly separable. In this work, we theoretically analyze the role of RP and present its benefits for improving the model’s generalization ability in each task and facilitating CL overall. Additionally, we take this result to the next level by proposing a Multi-View Random Projection scheme for a stronger ensemble classifier. In particular, we train a set of linear experts, among which diversity is encouraged based on the principle of AdaBoost, which was initially very challenging to apply to CL. Moreover, we employ a task-based adaptive backbone with distinct prompts dedicated to each task for better representation learning. To properly select these task-specific components and mitigate potential feature shifts caused by misprediction, we introduce a simple yet effective technique called the self-improvement process. Experimentally, our method consistently outperforms state-of-the-art baselines across a wide range of datasets.
Quyen Tran, Tung Lam Tran, Khanh Doan, Toan Tran 0003, Dinh Q. Phung, Khoat Than, Trung Le 0001
ICLR6
2025 DPaI: Differentiable Pruning at Initialization with Node-Path Balance Principle
abstract
Pruning at Initialization (PaI) is a technique in neural network optimization characterized by the proactive elimination of weights before the network's training on designated tasks. This innovative strategy potentially reduces the costs for training and inference, significantly advancing computational efficiency. A key factor leading to PaI's effectiveness is that it considers the saliency of weights in an untrained network, and prioritizes the trainability and optimization potential of the pruned subnetworks. Recent methods can effectively prevent the formation of hard-to-optimize networks, e.g. through iterative adjustments at each network layer. However, this way often results in large-scale discrete optimization problems, which could make PaI further challenging. This paper introduces a novel method, called DPaI, that involves a differentiable optimization of the pruning mask. DPaI adopts a dynamic and adaptable pruning process, allowing easier optimization processes and better solutions. More importantly, our differentiable formulation enables readily use of the existing rich body of efficient gradient-based methods for PaI. Our empirical results demonstrate that DPaI significantly outperforms current state-of-the-art PaI methods on various architectures, such as Convolutional Neural Networks and Vision-Transformers. Code is available at https://github.com/QuanNguyen-Tri/DPaI.git
Lichuan Xiang, Quan Nguyen-Tri, Lan-Cuong Nguyen, Khoat Than, Long Tran-Thanh, Hongkai Wen 0001
ICLR5
2025 Provably Improving Generalization of Few-shot models with Synthetic Data
abstract
Few-shot image classification remains challenging due to the scarcity of labeled training examples. Augmenting them with synthetic data has emerged as a promising way to alleviate this issue, but models trained on synthetic samples often face performance degradation due to the inherent gap between real and synthetic distributions. To address this limitation, we develop a theoretical framework that quantifies the impact of such distribution discrepancies on supervised learning, specifically in the context of image classification. More importantly, *our framework suggests practical ways to generate good synthetic samples and to train a predictor with high generalization ability*. Building upon this framework, we propose a novel theoretical-based algorithm that integrates prototype learning to optimize both data partitioning and model training, effectively bridging the gap between real few-shot data and synthetic data. Extensive experiments results show that our approach demonstrates superior performance compared to state-of-the-art methods, outperforming them across multiple datasets.
Lan-Cuong Nguyen, Quan Nguyen-Tri, Bang Tran Khanh, Dung D. Le, Long Tran-Thanh, Khoat Than
ICML6
2025 Out-of-vocabulary handling and topic quality control strategies in streaming topic models
Ngo Van Linh 0001, Ha Bang Ban, Khoat Than
Neurocomputing5
2025 Gentle local robustness implies generalization
Khoat Than, Dat Phan, Giang Vu
Mach. Learn.1
2024 On Inference Stability for Diffusion Models
abstract
Denoising Probabilistic Models (DPMs) represent an emerging domain of generative models that excel in generating diverse and high-quality images. However, most current training methods for DPMs often neglect the correlation between timesteps, limiting the model's performance in generating images effectively. Notably, we theoretically point out that this issue can be caused by the cumulative estimation gap between the predicted and the actual trajectory. To minimize that gap, we propose a novel sequence-aware loss that aims to reduce the estimation gap to enhance the sampling quality. Furthermore, we theoretically show that our proposed loss function is a tighter upper bound of the estimation loss in comparison with the conventional loss in DPMs. Experimental results on several benchmark datasets including CIFAR10, CelebA, and CelebA-HQ consistently show a remarkable improvement of our proposed method regarding the image generalization quality measured by FID and Inception Score compared to several DPM baselines. Our code and pre-trained checkpoints are available at https://github.com/VinAIResearch/SA-DPM.
Viet Nguyen, Giang Vu, Tung Nguyen Thanh, Khoat Than
AAAI4
2024 Robust Visual Reinforcement Learning by Prompt Tuning
Tung Tran 0006, Khoat Than, Danilo Vasconcellos Vargas
ACCV (9)2
2024 Partial ordered Wasserstein distance for sequential data
Tung Doan 0001, Tuan Phan, Phu Nguyen, Khoat Than, Muriel Visani, Atsuhiro Takasu
Neurocomputing4
2024 Continual variational dropout: a view of auxiliary local variables in continual learning
Ngo Van Linh 0001, Thien Huu Nguyen, Khoat Than
Mach. Learn.5
2023 Unsupervised image segmentation with robust virtual class contrast
Kien Do, Truong Vu, Khoat Than
Pattern Recognit. Lett.4
2023 Dynamic Transformation of Prior Knowledge Into Bayesian Models for Data Streams
abstract
We consider how to effectively use prior knowledge when learning a Bayesian model from streaming environments where the data come endlessly and sequentially. This problem is highly important in the era of data explosion and rich sources of valuable external knowledge such as pre-trained models, ontologies, Wikipedia, etc. We show that some existing approaches can forget any knowledge very fast. We then propose a novel framework that enables to incorporate the prior knowledge of different forms into a base Bayesian model for data streams. Our framework subsumes some existing popular models for time-series/dynamic data. Extensive experiments show that our framework outperforms existing methods with a large margin. In particular, our framework can help Bayesian models generalize well on extremely short text while other methods overfit. An implementation of our framework is available athttp://github.com/bachtranxuan/TPS.
Tran Xuan Bach, Nguyen Duc Anh, Ngo Van Linh 0001, Khoat Than
IEEE Trans. Knowl. Data Eng.4
2022 Auxiliary Local Variables for Improving Regularization/Prior Approach in Continual Learning
Ngo Van Linh 0001, Khoat Than
PAKDD (1)4
2022 Reducing Catastrophic Forgetting in Neural Networks via Gaussian Mixture Approximation
Hoang Phan, Anh Phan Tuan, Ngo Van Linh 0001, Khoat Than
PAKDD (1)5
2022 A graph convolutional topic model for short and noisy text streams
Ngo Van Linh 0001, Tran Xuan Bach, Khoat Than
Neurocomputing3
2022 Balancing stability and plasticity when learning topic models from short and noisy text streams
Trung Mai, Ngo Van Linh 0001, Khoat Than
Neurocomputing5
2022 From implicit to explicit feedback: A deep neural network for modeling sequential behaviours and long-short term preferences of online users
Quyen Tran, Lam Tran, Linh Chu Hai, Ngo Van Linh 0001, Khoat Than
Neurocomputing5
2022 Adaptive infinite dropout for noisy and sparse data streams
Ngo Van Linh 0001, Khoat Than
Mach. Learn.5
2021 Structured Dropout Variational Inference for Bayesian Neural Networks
abstract
Approximate inference in Bayesian deep networks exhibits a dilemma of how to yield high fidelity posterior approximations while maintaining computational efficiency and scalability. We tackle this challenge by introducing a novel variational structured approximation inspired by the Bayesian interpretation of Dropout regularization. Concretely, we focus on the inflexibility of the factorized structure in Dropout posterior and then propose an improved method called Variational Structured Dropout (VSD). VSD employs an orthogonal transformation to learn a structured representation on the variational Gaussian noise with plausible complexity, and consequently induces statistical dependencies in the approximate posterior. Theoretically, VSD successfully addresses the pathologies of previous Variational Dropout methods and thus offers a standard Bayesian justification. We further show that VSD induces an adaptive regularization term with several desirable properties which contribute to better generalization. Finally, we conduct extensive experiments on standard benchmarks to demonstrate the effectiveness of VSD over state-of-the-art variational methods on predictive accuracy, uncertainty estimation, and out-of-distribution detection.
Khoat Than, Nhat Ho
NeurIPS4
2021 Boosting prior knowledge in streaming variational Bayes
Ngo Van Linh 0001, Nguyen Kim Anh, Canh Hao Nguyen, Khoat Than
Neurocomputing5
2020 Predictive Coding for Locally-Linear Control
abstract
High-dimensional observations and unknown dynamics are major challenges when applying optimal control to many real-world decision making tasks. The Learning Controllable Embedding (LCE) framework addresses these challenges by embedding the observations into a lower dimensional latent space, estimating the latent dynamics, and then performing control directly in the latent space. To ensure the learned latent dynamics are predictive of next-observations, all existing LCE approaches decode back into the observation space and explicitly perform next-observation prediction—a challenging high-dimensional task that furthermore introduces a large number of nuisance parameters (i.e., the decoder) which are discarded during control. In this paper, we propose a novel information-theoretic LCE approach and show theoretically that explicit next-observation prediction can be replaced with predictive coding. We then use predictive coding to develop a decoder-free LCE model whose latent dynamics are amenable to locally-linear control. Extensive experiments on benchmark tasks show that our model reliably learns a controllable latent space that leads to superior performance when compared with state-of-the-art LCE baselines.
Yinlam Chow, Khoat Than, Mohammad Ghavamzadeh, Stefano Ermon, Hung H. Bui
ICML5
2020 Bag of biterms modeling for short texts
Anh Phan Tuan, Tran Xuan Bach, Thien Huu Nguyen, Ngo Van Linh 0001, Khoat Than
Knowl. Inf. Syst.5
2019 Employing the Correspondence of Relations and Connectives to Identify Implicit Discourse Relations via Label Embeddings
abstract
It has been shown that implicit connectives can be exploited to improve the performance of the models for implicit discourse relation recognition (IDRR).An important property of the implicit connectives is that they can be accurately mapped into the discourse relations conveying their functions.In this work, we explore this property in a multi-task learning framework for IDRR in which the relations and the connectives are simultaneously predicted, and the mapping is leveraged to transfer knowledge between the two prediction tasks via the embeddings of relations and connectives.We propose several techniques to enable such knowledge transfer that yield the state-of-the-art performance for IDRR on several settings of the benchmark dataset (i.e., the Penn Discourse Treebank dataset).
Linh The Nguyen, Ngo Van Linh 0001, Khoat Than, Thien Huu Nguyen
ACL (1)3
2019 From Implicit to Explicit Feedback: A deep neural network for modeling the sequential behavior of online users
abstract
We demonstrate the advantages of taking into account multiple types of behavior in recommendation systems. Intuitively, each user has to do some \textbf{implicit} actions (e.g., click) before making an \textbf{explicit} decision (e.g., purchase). Previous works showed that implicit and explicit feedback has distinct properties to make a useful recommendation. However, these works exploit implicit and explicit behavior separately and therefore ignore the semantic of interaction between users and items. In this paper, we propose a novel model namely \textit{Implicit to Explicit (ITE)} which directly models the order of user actions. Furthermore, we present an extended version of ITE, namely \textit{Implicit to Explicit with Side information (ITE-Si)}, which incorporates side information to enrich the representations of users and items. The experimental results show that both ITE and ITE-Si outperform existing recommendation systems and also demonstrate the effectiveness of side information in two large scale datasets.
Anh Phan Tuan, Nhat Nguyen Trong, Duong Bui Trong, Ngo Van Linh 0001, Khoat Than
ACML5
2019 Infinite Dropout for training Bayesian models from data streams
abstract
The ability to continuously train Bayesian models in streaming environments is highly important in the era of big data. However, it has to face the famous stability-plasticity dilemma and the problem of noisy and sparse data. We propose a novel and easy-to-implement framework, called Infinite Dropout (iDropout), to address these challenges. iDropout has an easy mechanism to balance between old and new information, which allows models to trade off stability against plasticity. Thanks to the ability to reduce overfitting and the ensemble property of Dropout, our framework obtains better generalization, thus effectively handles undesirable effects of noise and sparsity. Further, iDropout is able to adapt quickly to abnormal changes in data streams. We theoretically analyze the equivalence of Dropout in iDropout to a regularizer, well applied to a much larger context than what was known before. Extensive experiments show that iDropout significantly outperforms the state-of-the-art baselines.
Van-Son Nguyen, Duc-Tung Nguyen, Ngo Van Linh 0001, Khoat Than
IEEE BigData4
2019 Eliminating overfitting of probabilistic topic models on short and noisy text: The role of dropout
Cuong Ha, Van-Dang Tran, Ngo Van Linh 0001, Khoat Than
Int. J. Approx. Reason.4
2018 A Fast Algorithm for Posterior Inference with Latent Dirichlet Allocation
Bui Thi-Thanh-Xuan, Vu Van-Tu, Atsuhiro Takasu, Khoat Than
ACIIDS (2)4
2018 Collaborative Topic Model for Poisson distributed ratings
Hoa M. Le, Son Ta Cong, Quyen Pham The, Ngo Van Linh 0001, Khoat Than
Int. J. Approx. Reason.5
2017 Sparse Stochastic Inference with Regularization
Tung Doan 0001, Khoat Than
PAKDD (1)2
2017 Keeping Priors in Streaming Bayesian Learning
Ngo Van Linh 0001, Nguyen Kim Anh, Khoat Than
PAKDD (2)4
2017 An effective and interpretable method for document classification
Ngo Van Linh 0001, Nguyen Kim Anh, Khoat Than, Chien Nguyen Dang
Knowl. Inf. Syst.3
2016 Enabling Hierarchical Dirichlet Processes to Work Better for Short Texts at Large Scale
Khai Mai, Sang Mai, Ngo Van Linh 0001, Khoat Than
PAKDD (2)5
2015 Effective and Interpretable Document Classification Using Distinctly Labeled Dirichlet Process Mixture Models of von Mises-Fisher Distributions
Ngo Van Linh 0001, Nguyen Kim Anh, Khoat Than, Nguyen Nguyen Tat
DASFAA (2)3
2014 Dual online inference for latent Dirichlet allocation
Khoat Than, Tung Doan 0001
ACML1
2014 Modeling the diversity and log-normality of data
abstract
We investigate two important properties of real data: diversity and log-normality. Log-normality accounts for the fact that data follow the lognormal distribution, whereas diversity measures variations of the attributes in the data. To our knowledge,
Khoat Than
Intell. Data Anal.1
2014 An effective framework for supervised dimension reduction
Khoat Than, Duy Khuong Nguyen
Neurocomputing1
2012 Fully Sparse Topic Models
Khoat Than
ECML/PKDD (1)1