Tu Dinh Nguyen

dblp:129/2628 · DBLP profile ↗
← Back
38ranked-venue papers
9as first author
2since 2021 · last 2022
0000-0002-9281-8613ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 6 first-author · 1 since 2021Databases, data management, data science and information retrieval · 15 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-authorApplied, interdisciplinary, general and emerging computing · 2Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
14 papers
Probabilistic and Bayesian machine learning · 28% Generative modeling · 28% Learning theory · 13%

Topics — the 30 heaviest of 37, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
generative adversarial network
1.752019
Learning Generative Adversarial Networks from Multiple Data Sources · IJCAI 2019
Three-Player Wasserstein GAN via Amortised Duality · IJCAI 2019
Geometric Enclosing Networks · IJCAI 2018
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference › variational inference › particle-based variational inference
stein variational gradient descent
0.922022
Robust Variational Learning for Multiclass Kernel Models With Stein Refinement · IEEE Trans. Knowl. Data Eng. 2022
Robust Bayesian Kernel Machine via Stein Variational Gradient Descent for Big Data · KDD 2018
Machine learning › Learning theory › online learning
online kernel learning
0.832017
Approximation Vector Machines for Large-scale Online Learning · J. Mach. Learn. Res. 2017
Large-scale Online Kernel Learning with Random Feature Reparameterization · IJCAI 2017
Dual Space Gradient Descent for Online Learning · NIPS 2016
Machine learning › Learning theory
online learning
0.832017
Approximation Vector Machines for Large-scale Online Learning · J. Mach. Learn. Res. 2017
GoGP: Fast Online Regression with Gaussian Processes · ICDM 2017
Dual Space Gradient Descent for Online Learning · NIPS 2016
Machine learning › Kernel, tree and ensemble methods › kernel methods
kernel learning
0.612022
Robust Variational Learning for Multiclass Kernel Models With Stein Refinement · IEEE Trans. Knowl. Data Eng. 2022
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference
0.612022
Robust Variational Learning for Multiclass Kernel Models With Stein Refinement · IEEE Trans. Knowl. Data Eng. 2022
Machine learning › Generative modeling › diffusion model
conditional generation
0.412019
Learning Generative Adversarial Networks from Multiple Data Sources · IJCAI 2019
Machine learning › Generative modeling › diffusion model › controllable generation
constrained generation
0.412019
Learning Generative Adversarial Networks from Multiple Data Sources · IJCAI 2019
Machine learning › Optimization for machine learning
optimal transport
0.412019
Three-Player Wasserstein GAN via Amortised Duality · IJCAI 2019
Computer vision › Video understanding and tracking
video anomaly detection
0.412019
Robust Anomaly Detection in Videos Using Multilevel Representations · AAAI 2019
Machine learning › Generative modeling › generative adversarial network
Wasserstein GAN
0.412019
Three-Player Wasserstein GAN via Amortised Duality · IJCAI 2019
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference › bayesian nonparametric model
bayesian kernel models
0.312018
Robust Bayesian Kernel Machine via Stein Variational Gradient Descent for Big Data · KDD 2018
Machine learning › Kernel, tree and ensemble methods › kernel methods › kernel machines
large-scale kernel methods
0.312018
Robust Bayesian Kernel Machine via Stein Variational Gradient Descent for Big Data · KDD 2018
Machine learning › Representation and self-supervised learning › representation learning › dimensionality reduction
manifold learning
0.312018
Geometric Enclosing Networks · IJCAI 2018
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
scalable bayesian inference
0.312018
Robust Bayesian Kernel Machine via Stein Variational Gradient Descent for Big Data · KDD 2018
Machine learning › Generative modeling
variational autoencoder
0.312018
Geometric Enclosing Networks · IJCAI 2018
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
gaussian process
0.312017
GoGP: Fast Online Regression with Gaussian Processes · ICDM 2017
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › gaussian process
gaussian process regression
0.312017
GoGP: Fast Online Regression with Gaussian Processes · ICDM 2017
Machine learning › Kernel, tree and ensemble methods
large-scale kernel learning
0.312017
Large-scale Online Kernel Learning with Random Feature Reparameterization · IJCAI 2017
Machine learning › Generative modeling › generative adversarial network
mode collapse mitigation
0.312017
Dual Discriminator Generative Adversarial Nets · NIPS 2017
Machine learning › Kernel, tree and ensemble methods › kernel methods › kernel approximation
random features
0.312017
Large-scale Online Kernel Learning with Random Feature Reparameterization · IJCAI 2017
Machine learning › Efficient and distributed learning › model compression
sparsity
0.312017
Approximation Vector Machines for Large-scale Online Learning · J. Mach. Learn. Res. 2017
Machine learning › Optimization for machine learning
stochastic gradient descent
0.312017
Large-scale Online Kernel Learning with Random Feature Reparameterization · IJCAI 2017
Machine learning › Optimization for machine learning › gradient-based optimization
gradient descent
0.212016
Dual Space Gradient Descent for Online Learning · NIPS 2016
Computer vision › Image recognition and object detection › image classification › large-scale image classification
large-scale classification
0.212016
One-Pass Logistic Regression for Label-Drift and Large-Scale Classification on Distributed Systems · ICDM 2016
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › regression › generalized linear model
logistic regression
0.212016
One-Pass Logistic Regression for Label-Drift and Large-Scale Classification on Distributed Systems · ICDM 2016
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model
0.212015
Tensor-Variate Restricted Boltzmann Machines · AAAI 2015
Machine learning › Probabilistic and Bayesian machine learning › boltzmann machine
restricted boltzmann machine
0.212015
Tensor-Variate Restricted Boltzmann Machines · AAAI 2015
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
posterior inference
0.212022
Robust Variational Learning for Multiclass Kernel Models With Stein Refinement · IEEE Trans. Knowl. Data Eng. 2022
Machine learning › Generative modeling › generative adversarial network
conditional GAN
0.112019
Robust Anomaly Detection in Videos Using Multilevel Representations · AAAI 2019

Methods — techniques the papers use, named apart from their topics

stein variational gradient descent · 0.9geometric optimization · 0.6variational inference · 0.6recurrent neural network · 0.6convergence analysis · 0.5min-max optimization · 0.4kantorovich duality · 0.4denoising autoencoder · 0.4conditional generative adversarial network · 0.4amortized optimization · 0.4
YearPublicationVenuePosition
2022 Robust Variational Learning for Multiclass Kernel Models With Stein Refinement
abstract
Kernel-based models have a strong generalization ability, but most, including SVM, are vulnerable to the curse of kernelization. Moreover, their predictive performance is sensitive to hyperparameter tuning, which demands high computational resources. These problems render kernel methods problematic when dealing with large-scale datasets. To this end, we first formulate the optimization problem in a kernel-based learning setting as a posterior inference problem, and then develop a rich family of Recurrent Neural Network-based variational inference techniques. Unlike existing literature, which stops at the variational distribution and uses it as the surrogate for the true posterior distribution, here we further leverage Stein Variational Gradient Descent to further bring the variational distribution closer to the true posterior, we refer to this step asStein Refinement. Putting these altogether, we arrive at a robust and efficient variational learning method for multiclass kernel machines with extremely accurate approximation. Moreover, our formulation enables efficient learning of kernel parameters and hyperparameters which robustifies the proposed method against data uncertainties. The extensive experiments show that without tuning any parameter on modest quantities of data our method obtains comparable accuracy to LIBSVM, a well-known implementation of SVM, and outperforms other baselines, while being able to seamlessly scale with large-scale datasets.
Trung Le 0001, Tu Dinh Nguyen, Geoffrey I. Webb, Dinh Q. Phung
IEEE Trans. Knowl. Data Eng.3
2021 Quaternion Graph Neural Networks
abstract
Recently, graph neural networks (GNNs) have become an important and active research direction in deep learning. It is worth noting that most of the existing GNN-based methods learn graph representations within the Euclidean vector space. Beyond the Euclidean space, learning representation and embeddings in hyper-complex space have also shown to be a promising and effective approach. To this end, we propose Quaternion Graph Neural Networks (QGNN) to learn graph representations within the Quaternion space. As demonstrated, the Quaternion space, a hyper-complex vector space, provides highly meaningful computations and analogical calculus through Hamilton product compared to the Euclidean and complex vector spaces. Our QGNN obtains state-of-the-art results on a range of benchmark datasets for graph classification and node classification. Besides, regarding knowledge graphs, our QGNN-based embedding model achieves state-of-the-art results on three new and challenging benchmark datasets for knowledge graph completion. Our code is available at: \url{https://github.com/daiquocnguyen/QGNN}.
Dai Quoc Nguyen, Tu Dinh Nguyen, Dinh Q. Phung
ACML2
2020 A Capsule Network-based Model for Learning Node Embeddings
abstract
In this paper, we focus on learning low-dimensional embeddings for nodes in graph-structured data. To achieve this, we propose Caps2NE -- a new unsupervised embedding model leveraging a network of two capsule layers. Caps2NE induces a routing process to aggregate feature vectors of context neighbors of a given target node at the first capsule layer, then feed these features into the second capsule layer to infer a plausible embedding for the target node. Experimental results show that our proposed Caps2NE obtains state-of-the-art performances on benchmark datasets for the node classification task. Our code is available at: https://github.com/daiquocnguyen/Caps2NE.
Dai Quoc Nguyen, Tu Dinh Nguyen, Dat Quoc Nguyen, Dinh Q. Phung
CIKM2
2020 A Self-attention Network Based Node Embedding Model
Dai Quoc Nguyen, Tu Dinh Nguyen, Dinh Q. Phung
ECML/PKDD (3)2
2019 Robust Anomaly Detection in Videos Using Multilevel Representations
abstract
Detecting anomalies in surveillance videos has long been an important but unsolved problem. In particular, many existing solutions are overly sensitive to (often ephemeral) visual artifacts in the raw video data, resulting in false positives and fragmented detection regions. To overcome such sensitivity and to capture true anomalies with semantic significance, one natural idea is to seek validation from abstract representations of the videos. This paper introduces a framework of robust anomaly detection using multilevel representations of both intensity and motion data. The framework consists of three main components: 1) representation learning using Denoising Autoencoders, 2) level-wise representation generation using Conditional Generative Adversarial Networks, and 3) consolidating anomalous regions detected at each representation level. Our proposed multilevel detector shows a significant improvement in pixel-level Equal Error Rate, namely 11.35%, 12.32% and 4.31% improvement in UCSD Ped 1, UCSD Ped 2 and Avenue datasets respectively. In addition, the model allowed us to detect mislabeled anomalies in the UCDS Ped 1.
Hung Vu, Tu Dinh Nguyen, Trung Le 0001, Wei Luo 0001, Dinh Q. Phung
AAAI2
2019 Three-Player Wasserstein GAN via Amortised Duality
abstract
We propose a new formulation for learning generative adversarial networks (GANs) using optimal transport cost (the general form of Wasserstein distance) as the objective criterion to measure the dissimilarity between target distribution and learned distribution. Our formulation is based on the general form of the Kantorovich duality which is applicable to optimal transport with a wide range of cost functions that are not necessarily metric. To make optimising this duality form amenable to gradient-based methods, we employ a function that acts as an amortised optimiser for the innermost optimisation problem. Interestingly, the amortised optimiser can be viewed as a mover since it strategically shifts around data points. The resulting formulation is a sequential min-max-min game with 3 players: the generator, the critic, and the mover where the new player, the mover, attempts to fool the critic by shifting the data around. Despite involving three players, we demonstrate that our proposed formulation can be trained reasonably effectively via a simple alternative gradient learning strategy. Compared with the existing Lipschitz-constrained formulations of Wasserstein GAN on CIFAR-10, our model yields significantly better diversity scores than weight clipping and comparable performance to gradient penalty method.
Nhan Dam, Quan Hoang, Trung Le 0001, Tu Dinh Nguyen, Hung Hai Bui, Dinh Q. Phung
IJCAI4
2019 Learning Generative Adversarial Networks from Multiple Data Sources
abstract
Generative Adversarial Networks (GANs) are a powerful class of deep generative models. In this paper, we extend GAN to the problem of generating data that are not only close to a primary data source but also required to be different from auxiliary data sources. For this problem, we enrich both GANs' formulations and applications by introducing pushing forces that thrust generated samples away from given auxiliary data sources. We term our method Push-and-Pull GAN (P2GAN). We conduct extensive experiments to demonstrate the merit of P2GAN in two applications: generating data with constraints and addressing the mode collapsing problem. We use CIFAR-10, STL-10, and ImageNet datasets and compute Fréchet Inception Distance to evaluate P2GAN's effectiveness in addressing the mode collapsing problem. The results show that P2GAN outperforms the state-of-the-art baselines. For the problem of generating data with constraints, we show that P2GAN can successfully avoid generating specific features such as black hair.
Trung Le 0001, Quan Hoang, Hung Vu, Tu Dinh Nguyen, Hung Hai Bui, Dinh Q. Phung
IJCAI4
2019 GoGP: scalable geometric-based Gaussian process for online regression
Trung Le 0001, Vu Nguyen 0001, Tu Dinh Nguyen, Dinh Q. Phung
Knowl. Inf. Syst.4
2018 Clustering Induced Kernel Learning
abstract
Learning rich and expressive kernel functions is a challenging task in kernel-based supervised learning. Multiple kernel learning (MKL) approach addresses this problem by combining a mixed variety of kernels and letting the optimization solver choose the most appropriate combination. However, most of existing methods are parametric in the sense that they require a predefined list of kernels. Hence, there appears a substantial trade-off between computation and the modeling risk of not being able to explore more expressive and suitable kernel functions. Moreover, current existing approaches to combine kernels cannot exploit clustering structure carried in data, especially when data are heterogeneous. In this work, we present a new framework that leverages Bayesian nonparametric models (i.e, automatically grow kernel functions) with multiple kernel learning to develop a new framework that enjoys the nonparametric flavor in the context of multiple kernel learning. In particular, we propose Clustering Induced Kernel Learning (CIK) method that can automatically discover clustering structure from the data and train a single kernel machine to fit data in each discovered cluster simultaneously. The outcome of our proposed method includes both clustering analysis and multiple kernel classifier for a given dataset. We conduct extensive experiments on several benchmark datasets. The experimental results show that our method can improve classification and clustering performance when datasets have complex clustering structure with different preferred kernels.
Nhan Dam, Trung Le 0001, Tu Dinh Nguyen, Dinh Q. Phung
ACML4
2018 Batch Normalized Deep Boltzmann Machines
abstract
Training Deep Boltzmann Machines (DBMs) is a challenging task in deep generative model studies. The careless training usually leads to a divergence or a useless model. We discover that this phenomenon is due to the change of DBM layers’ input signals during model parameter updates, similar to other deterministic deep networks such as Convolutional Neural Networks (CNNs) or Recurrent Neural Networks (RNNs). The change of layers’ input distributions not only complicates the learning process but also causes redundant neurons that simply imitate the others’ behaviors. Although this phenomenon can be coped using batch normalization in deep learning, integrating this technique into the probabilistic network of DBMs is a challenging problem since it has to satisfy two conditions of energy function and conditional probabilities. In this paper, we introduce Batch Normalized Deep Boltzmann Machines (BNDBMs) that meet both aforementioned conditions and successfully combine batch normalization and DBMs into the same framework. However, unlike CNNs, due to the probabilistic nature of DBMs, training DBMs with batch normalization has some differences: i) fixing shift parameters $\bnshift$ but learning scale parameters $\bnscale$; ii) avoiding normalizing the first hidden layer and iii) maintaining multiple pairs of population means and variances per neuron rather than one pair in CNNs. We observe that our proposed BNDBMs can stabilize the input signals of network layers and facilitate the training process as well as improve the model quality. More interestingly, BNDBMs can be trained successfully without pretraining, which is usually a mandatory step in most existing DBMs. The experimental results in MNIST, Fashion-MNIST and Caltech 101 Silhouette datasets show that our BNDBMs outperform DBMs and centered DBMs in terms of feature representation and classification accuracy ($3.98%$ and $5.84%$ average improvement for pretraining and no pretraining respectively).
Hung Vu, Tu Dinh Nguyen, Trung Le 0001, Wei Luo 0001, Dinh Q. Phung
ACML2
2018 MGAN: Training Generative Adversarial Nets with Multiple Generators
Quan Hoang, Tu Dinh Nguyen, Trung Le 0001, Dinh Q. Phung
ICLR (Poster)2
2018 Bayesian Multi-Hyperplane Machine for Pattern Recognition
abstract
Current existing multi-hyperplane machine approach deals with high-dimensional and complex datasets by approximating the input data region using a parametric mixture of hyperplanes. Consequently, this approach requires an excessively time-consuming parameter search to find the set of optimal hyper-parameters. Another serious drawback of this approach is that it is often suboptimal since the optimal choice for the hyper-parameter is likely to lie outside the searching space due to the space discretization step required in grid search. To address these challenges, we propose in this paper BAyesian Multi-hyperplane Machine (BAMM). Our approach departs from a Bayesian perspective, and aims to construct an alternative probabilistic view in such a way that its maximum-a-posteriori (MAP) estimation reduces exactly to the original optimization problem of a multi-hyperplane machine. This view allows us to endow prior distributions over hyper-parameters and augment auxiliary variables to efficiently infer model parameters and hyper-parameters via Markov chain Monte Carlo (MCMC) method. We then employ a Stochastic Gradient Descent (SGD) framework to scale our model up with ever-growing large datasets. Extensive experiments demonstrate the capability of our proposed method in learning the optimal model without using any parameter tuning, and in achieving comparable accuracies compared with the state-of-art baselines; in the meantime our model can seamlessly handle with large-scale datasets.
Trung Le 0001, Tu Dinh Nguyen, Dinh Q. Phung
ICPR3
2018 Geometric Enclosing Networks
abstract
Training model to generate data has increasingly attracted research attention and become important in modern world applications. We propose in this paper a new geometry-based optimization approach to address this problem. Orthogonal to current state-of-the-art density-based approaches, most notably VAE and GAN, we present a fresh new idea that borrows the principle of minimal enclosing ball to train a generator G\left(\bz\right) in such a way that both training and generated data, after being mapped to the feature space, are enclosed in the same sphere. We develop theory to guarantee that the mapping is bijective so that its inverse from feature space to data space results in expressive nonlinear contours to describe the data manifold, hence ensuring data generated are also lying on the data manifold learned from training data. Our model enjoys a nice geometric interpretation, hence termed Geometric Enclosing Networks (GEN), and possesses some key advantages over its rivals, namely simple and easy-to-control optimization formulation, avoidance of mode collapsing and efficiently learn data manifold representation in a completely unsupervised manner. We conducted extensive experiments on synthesis and real-world datasets to illustrate the behaviors, strength and weakness of our proposed GEN, in particular its ability to handle multi-modal data and quality of generated data.
Trung Le 0001, Hung Vu, Tu Dinh Nguyen, Dinh Q. Phung
IJCAI3
2018 Robust Bayesian Kernel Machine via Stein Variational Gradient Descent for Big Data
abstract
Kernel methods are powerful supervised machine learning models for their strong generalization ability, especially on limited data to effectively generalize on unseen data. However, most kernel methods, including the state-of-the-art LIBSVM, are vulnerable to the curse of kernelization, making them infeasible to apply to large-scale datasets. This issue is exacerbated when kernel methods are used in conjunction with a grid search to tune their kernel parameters and hyperparameters which brings in the question of model robustness when applied to real datasets. In this paper, we propose a robust Bayesian Kernel Machine (BKM) - a Bayesian kernel machine that exploits the strengths of both the Bayesian modelling and kernel methods. A key challenge for such a formulation is the need for an efficient learning algorithm. To this end, we successfully extended the recent Stein variational theory for Bayesian inference for our proposed model, resulting in fast and efficient learning and prediction algorithms. Importantly our proposed BKM is resilient to the curse of kernelization, hence making it applicable to large-scale datasets and robust to parameter tuning, avoiding the associated expense and potential pitfalls with current practice of parameter tuning. Our extensive experimental results on 12 benchmark datasets show that our BKM without tuning any parameter can achieve comparable predictive performance with the state-of-the-art LIBSVM and significantly outperforms other baselines, while obtaining significantly speedup in terms of the total training time compared with its rivals
Trung Le 0001, Tu Dinh Nguyen, Dinh Q. Phung, Geoffrey I. Webb
KDD3
2018 Trans2Vec: Learning Transaction Embedding via Items and Frequent Itemsets
Dang Nguyen 0002, Tu Dinh Nguyen, Wei Luo 0001, Svetha Venkatesh
PAKDD (3)2
2018 Sqn2Vec: Learning Sequence Representation via Sequential Patterns with a Gap Constraint
Dang Nguyen 0002, Wei Luo 0001, Tu Dinh Nguyen, Svetha Venkatesh, Dinh Q. Phung
ECML/PKDD (2)3
2018 Learning Graph Representation via Frequent Subgraphs
abstract
We propose a novel approach to learn distributed representation for graph data. Our idea is to combine a recently introduced neural document embedding model with a traditional pattern mining technique, by treating a graph as a document and frequent subgraphs as atomic units for the embedding process. Compared to the latest graph embedding methods, our proposed method offers three key advantages: fully unsupervised learning, entire-graph embedding, and edge label leveraging. We demonstrate our method on several datasets in comparison with a comprehensive list of up-to-date state-of-the-art baselines where we show its advantages for both classification and clustering tasks.
Dang Nguyen 0002, Wei Luo 0001, Tu Dinh Nguyen, Svetha Venkatesh, Dinh Q. Phung
SDM3
2017 Animal Recognition and Identification with Deep Convolutional Neural Networks for Automated Wildlife Monitoring
abstract
Efficient and reliable monitoring of wild animals in their natural habitats is essential to inform conservation and management decisions. Automatic covert cameras or "camera traps" are being an increasingly popular tool for wildlife monitoring due to their effectiveness and reliability in collecting data of wildlife unobtrusively, continuously and in large volume. However, processing such a large volume of images and videos captured from camera traps manually is extremely expensive, time-consuming and also monotonous. This presents a major obstacle to scientists and ecologists to monitor wildlife in an open environment. Leveraging on recent advances in deep learning techniques in computer vision, we propose in this paper a framework to build automated animal recognition in the wild, aiming at an automated wildlife monitoring system. In particular, we use a single-labeled dataset from Wildlife Spotter project, done by citizen scientists, and the state-of-the-art deep convolutional neural network architectures, to train a computational system capable of filtering animal images and identifying species automatically. Our experimental results achieved an accuracy at 96.6% for the task of detecting images containing animal, and 90.4% for identifying the three most common species among the set of images of wild animals taken in South-central Victoria, Australia, demonstrating the feasibility of building fully automated wildlife observation. This, in turn, can therefore speed up research findings, construct more efficient citizen sciencebased monitoring systems and subsequent management decisions, having the potential to make significant impacts to the world of ecology and trap camera images analysis.
Sarah J. Maclagan, Tu Dinh Nguyen, Thin Nguyen, Paul Flemons, Kylie Andrews, Euan G. Ritchie, Dinh Q. Phung
DSAA3
2017 GoGP: Fast Online Regression with Gaussian Processes
abstract
One of the most current challenging problems in Gaussian process regression (GPR) is to handle large-scale datasets and to accommodate an online learning setting where data arrive irregularly on the fly. In this paper, we introduce a novel online Gaussian process model that could scale with massive datasets. Our approach is formulated based on alternative representation of the Gaussian process under geometric and optimization views, hence termed geometric-based online GP (GoGP). We developed theory to guarantee that with a good convergence rate our proposed algorithm always produces a (sparse) solution which is close to the true optima to any arbitrary level of approximation accuracy specified a priori. Furthermore, our method is proven to scale seamlessly not only with large-scale datasets, but also to adapt accurately with streaming data. We extensively evaluated our proposed model against state-of-the-art baselines using several large-scale datasets for online regression task. The experimental results show that our GoGP delivered comparable, or slightly better, predictive performance while achieving a magnitude of computational speedup compared with its rivals under online setting. More importantly, its convergence behavior is guaranteed through our theoretical analysis, which is rapid and stable while achieving lower errors.
Trung Le 0001, Vu Nguyen 0001, Tu Dinh Nguyen, Dinh Q. Phung
ICDM4
2017 Large-scale Online Kernel Learning with Random Feature Reparameterization
abstract
A typical online kernel learning method faces two fundamental issues: the complexity in dealing with a huge number of observed data points (a.k.a the curse of kernelization) and the difficulty in learning kernel parameters, which often assumed to be fixed. Random Fourier feature is a recent and effective approach to address the former by approximating the shift-invariant kernel function via Bocher's theorem, and allows the model to be maintained directly in the random feature space with a fixed dimension, hence the model size remains constant w.r.t. data size. We further introduce in this paper the reparameterized random feature (RRF), a random feature framework for large-scale online kernel learning to address both aforementioned challenges. Our initial intuition comes from the so-called "reparameterization trick" [Kingma et al., 2014] to lift the source of randomness of Fourier components to another space which can be independently sampled, so that stochastic gradient of the kernel parameters can be analytically derived. We develop a well-founded underlying theory for our method, including a general way to reparameterize the kernel, and a new tighter error bound on the approximation quality. This view further inspires a direct application of stochastic gradient descent for updating our model under an online learning setting. We then conducted extensive experiments on several large-scale datasets where we demonstrate that our work achieves state-of-the-art performance in both learning efficacy and efficiency.
Tu Dinh Nguyen, Trung Le 0001, Hung Hai Bui, Dinh Q. Phung
IJCAI1
2017 Dual Discriminator Generative Adversarial Nets
abstract
We propose in this paper a novel approach to tackle the problem of mode collapse encountered in generative adversarial network (GAN). Our idea is intuitive but proven to be very effective, especially in addressing some key limitations of GAN. In essence, it combines the Kullback-Leibler (KL) and reverse KL divergences into a unified objective function, thus it exploits the complementary statistical properties from these divergences to effectively diversify the estimated density in capturing multi-modes. We term our method dual discriminator generative adversarial nets (D2GAN) which, unlike GAN, has two discriminators; and together with a generator, it also has the analogy of a minimax game, wherein a discriminator rewards high scores for samples from data distribution whilst another discriminator, conversely, favoring data from the generator, and the generator produces data to fool both two discriminators. We develop theoretical analysis to show that, given the maximal discriminators, optimizing the generator of D2GAN reduces to minimizing both KL and reverse KL divergences between data distribution and the distribution induced from the data generated by the generator, hence effectively avoiding the mode collapsing problem. We conduct extensive experiments on synthetic and real-world large-scale datasets (MNIST, CIFAR-10, STL-10, ImageNet), where we have made our best effort to compare our D2GAN with the latest state-of-the-art GAN's variants in comprehensive qualitative and quantitative evaluations. The experimental results demonstrate the competitive and superior performance of our approach in generating good quality and diverse samples over baselines, and the capability of our method to scale up to ImageNet database.
Tu Dinh Nguyen, Trung Le 0001, Hung Vu, Dinh Q. Phung
NIPS1
2017 Energy-Based Localized Anomaly Detection in Video Surveillance
Hung Vu, Tu Dinh Nguyen, Anthony Travers, Svetha Venkatesh, Dinh Q. Phung
PAKDD (1)2
2017 Supervised Restricted Boltzmann Machines
Tu Dinh Nguyen, Dinh Q. Phung, Viet Huynh, Trung Le 0001
UAI1
2017 Approximation Vector Machines for Large-scale Online Learning
abstract
One of the most challenging problems in kernel online learning is to bound the model size and to promote model sparsity. Sparse models not only improve computation and memory usage, but also enhance the generalization capacity -- a principle that concurs with the law of parsimony. However, inappropriate sparsity modeling may also significantly degrade the performance. In this paper, we propose Approximation Vector Machine (AVM), a model that can simultaneously encourage sparsity and safeguard its risk in compromising the performance. In an online setting context, when an incoming instance arrives, we approximate this instance by one of its neighbors whose distance to it is less than a predefined threshold. Our key intuition is that since the newly seen instance is expressed by its nearby neighbor the optimal performance can be analytically formulated and maintained. We develop theoretical foundations to support this intuition and further establish an analysis for the common loss functions including Hinge, smooth Hinge, and Logistic (i.e., for the classification task) and $\ell_{1}$, $\ell_{2}$, and $\varepsilon$-insensitive (i.e., for the regression task) to characterize the gap between the approximation and optimal solutions. This gap crucially depends on two key factors including the frequency of approximation (i.e., how frequent the approximation operation takes place) and the predefined threshold. We conducted extensive experiments for classification and regression tasks in batch and online modes using several benchmark datasets. The quantitative results show that our proposed AVM obtained comparable predictive performances with current state-of-the-art methods while simultaneously achieving significant computational speed-up due to the ability of the proposed AVM in maintaining the model size.
Trung Le 0001, Tu Dinh Nguyen, Vu Nguyen 0001, Dinh Q. Phung
J. Mach. Learn. Res.2
2016 Multiple Kernel Learning with Data Augmentation
abstract
The motivations of multiple kernel learning (MKL) approach are to increase kernel expressiveness capacity and to avoid the expensive grid search over a wide spectrum of kernels. A large amount of work has been proposed to improve the MKL in terms of the computational cost and the sparsity of the solution. However, these studies still either require an expensive grid search on the model parameters or scale unsatisfactorily with the numbers of kernels and training samples. In this paper, we address these issues by conjoining MKL, Stochastic Gradient Descent (SGD) framework, and data augmentation technique. The pathway of our proposed method is developed as follows. We first develop a maximum-a-posteriori (MAP) view for MKL under a probabilistic setting and described in a graphical model. This view allows us to develop data augmentation technique to make the inference for finding the optimal parameters feasible, as opposed to traditional approach of training MKL via convex optimization techniques. As a result, we can use the standard SGD framework to learn weight matrix and extend the model to support online learning. We validate our method on several benchmark datasets in both batch and online settings. The experimental results show that our proposed method can learn the parameters in a principled way to eliminate the expensive grid search while gaining a significant computational speedup comparing with the state-of-the-art baselines.
Trung Le 0001, Vu Nguyen 0001, Tu Dinh Nguyen, Dinh Q. Phung
ACML4
2016 Nonparametric Budgeted Stochastic Gradient Descent
abstract
One of the most challenging problems in kernel online learning is to bound the model size. Budgeted kernel online learning addresses this issue by bounding the model size to a predefined budget. However, determining an appropriate value for such predefined budget is arduous. In this paper, we propose the Nonparametric Budgeted Stochastic Gradient Descent that allows the model size to automatically grow with data in a principled way. We provide theoretical analysis to show that our framework is guaranteed to converge for a large collection of loss functions (e.g. Hinge, Logistic, L2, L1, and \varepsilon-insensitive) which enables the proposed algorithm to perform both classification and regression tasks without hurting the ideal convergence rate O\left(\frac1T\right) of the standard Stochastic Gradient Descent. We validate our algorithm on the real-world datasets to consolidate the theoretical claims.
Trung Le 0001, Vu Nguyen 0001, Tu Dinh Nguyen, Dinh Q. Phung
AISTATS3
2016 One-Pass Logistic Regression for Label-Drift and Large-Scale Classification on Distributed Systems
abstract
Logistic regression (LR) for classification is the workhorse in industry, where a set of predefined classes is required. The model, however, fails to work in the case where the class labels are not known in advance, a problem we term label-drift classification. Label-drift classification problem naturally occurs in many applications, especially in the context of streaming settings where the incoming data may contain samples categorized with new classes that have not been previously seen. Additionally, in the wave of big data, traditional LR methods may fail due to their expense of running time. In this paper, we introduce a novel variant of LR, namely one-pass logistic regression (OLR) to offer a principled treatment for label-drift and large-scale classifications. To handle largescale classification for big data, we further extend our OLR to a distributed setting for parallelization, termed sparkling OLR (Spark-OLR). We demonstrate the scalability of our proposed methods on large-scale datasets with more than one hundred million data points. The experimental results show that the predictive performances of our methods are comparable orbetter than those of state-of-the-art baselines whilst the executiontime is much faster at an order of magnitude. In addition, the OLR and Spark-OLR are invariant to data shuffling and have no hyperparameter to tune that significantly benefits data practitioners and overcomes the curse of big data cross-validationto select optimal hyperparameters.
Vu Nguyen 0001, Tu Dinh Nguyen, Trung Le 0001, Svetha Venkatesh, Dinh Q. Phung
ICDM2
2016 Distributed data augmented support vector machine on Spark
abstract
Support vector machines (SVMs) are widely-used for classification in machine learning and data mining tasks. However, they traditionally have been applied to small to medium datasets. Recent need to scale up with data size has attracted research attention to develop new methods and implementation for SVM to perform tasks at scale. Distributed SVMs are relatively new and studied recently, but the distributed implementation for SVM with data augmentation has not been developed. This paper introduces a distributed data augmentation implementation for SVM on Apache Spark, a recent advanced and popular platform for distributed computing that has been employed widely in research as well as in industry. We term our implementation sparkling vector machine (SkVM) which supports both classification and regression tasks by scanning through the data exactly once. In addition, we further develop a framework to handle the data with new classes arriving under an online classification setting where new data points can have labels that have not previously seen - a problem we term label-drift classification. We demonstrate the scalability of our proposed method on large-scale datasets with more than one hundred million data points. The experimental results show that the predictive performances of our method are comparable or better than those of baselines whilst the execution time is much faster at an order of magnitude.
Tu Dinh Nguyen, Vu Nguyen 0001, Trung Le 0001, Dinh Q. Phung
ICPR1
2016 Dual Space Gradient Descent for Online Learning
abstract
One crucial goal in kernel online learning is to bound the model size. Common approaches employ budget maintenance procedures to restrict the model sizes using removal, projection, or merging strategies. Although projection and merging, in the literature, are known to be the most effective strategies, they demand extensive computation whilst removal strategy fails to retain information of the removed vectors. An alternative way to address the model size problem is to apply random features to approximate the kernel function. This allows the model to be maintained directly in the random feature space, hence effectively resolve the curse of kernelization. However, this approach still suffers from a serious shortcoming as it needs to use a high dimensional random feature space to achieve a sufficiently accurate kernel approximation. Consequently, it leads to a significant increase in the computational cost. To address all of these aforementioned challenges, we present in this paper the Dual Space Gradient Descent (DualSGD), a novel framework that utilizes random features as an auxiliary space to maintain information from data points removed during budget maintenance. Consequently, our approach permits the budget to be maintained in a simple, direct and elegant way while simultaneously mitigating the impact of the dimensionality issue on learning performance. We further provide convergence analysis and extensively conduct experiments on five real-world datasets to demonstrate the predictive performance and scalability of our proposed method in comparison with the state-of-the-art baselines.
Trung Le 0001, Tu Dinh Nguyen, Vu Nguyen 0001, Dinh Q. Phung
NIPS2
2016 Budgeted Semi-supervised Support Vector Machine
Trung Le 0001, Phuong Duong, Mi Dinh, Tu Dinh Nguyen, Vu Nguyen 0001, Dinh Q. Phung
UAI4
2016 Graph-induced restricted Boltzmann machines for document modeling
Tu Dinh Nguyen, Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh
Inf. Sci.1
2015 Tensor-Variate Restricted Boltzmann Machines
abstract
Restricted Boltzmann Machines (RBMs) are an important class of latent variable models for representing vector data. An under-explored area is multimode data, where each data point is a matrix or a tensor. Standard RBMs applying to such data would require vectorizing matrices and tensors, thus resulting in unnecessarily high dimensionality and at the same time, destroying the inherent higher-order interaction structures. This paper introduces Tensor-variate Restricted Boltzmann Machines (TvRBMs) which generalize RBMs to capture the multiplicative interaction between data modes and the latent variables. TvRBMs are highly compact in that the number of free parameters grows only linear with the number of modes. We demonstrate the capacity of TvRBMs on three real-world applications: handwritten digit classification, face recognition and EEG-based alcoholic diagnosis. The learnt features of the model are more discriminative than the rivals, resulting in better classification performance.
Tu Dinh Nguyen, Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh
AAAI1
2015 Stabilizing Sparse Cox Model Using Statistic and Semantic Structures in Electronic Medical Records
Shivapratap Gopakumar, Tu Dinh Nguyen, Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh
PAKDD (2)2
2015 Learning vector representation of medical objects via EMR-driven nonnegative restricted Boltzmann machines (eNRBM)
Truyen Tran 0001, Tu Dinh Nguyen, Dinh Q. Phung, Svetha Venkatesh
J. Biomed. Informatics2
2015 Stabilizing High-Dimensional Prediction Models Using Feature Graphs
abstract
We investigate feature stability in the context of clinical prognosis derived from high-dimensional electronic medical records. To reduce variance in the selected features that are predictive, we introduce Laplacian-based regularization into a regression model. The Laplacian is derived on a feature graph that captures both the temporal and hierarchic relations between hospital events, diseases, and interventions. Using a cohort of patients with heart failure, we demonstrate better feature stability and goodness-of-fit through feature graph stabilization.
Shivapratap Gopakumar, Truyen Tran 0001, Tu Dinh Nguyen, Dinh Q. Phung, Svetha Venkatesh
IEEE J. Biomed. Health Informatics3
2013 Learning Parts-based Representations with Nonnegative Restricted Boltzmann Machine
abstract
The success of any machine learning system depends critically on effective representations of data. In many cases, especially those in vision, it is desirable that a representation scheme uncovers the parts-based, additive nature of the data. Of current representation learning schemes, restricted Boltzmann machines (RBMs) have proved to be highly effective in unsupervised settings. However, when it comes to parts-based discovery, RBMs do not usually produce satisfactory results. We enhance such capacity of RBMs by introducing nonnegativity into the model weights, resulting in a variant called \emphnonnegative restricted Boltzmann machine (NRBM). The NRBM produces not only controllable decomposition of data into interpretable parts but also offers a way to estimate the intrinsic nonlinear dimensionality of data. We demonstrate the capacity of our model on well-known datasets of handwritten digits, faces and documents. The decomposition quality on images is comparable with or better than what produced by the nonnegative matrix factorisation (NMF), and the thematic features uncovered from text are qualitatively interpretable in a similar manner to that of the latent Dirichlet allocation (LDA). However, the learnt features, when used for classification, are more discriminative than those discovered by both NMF and LDA and comparable with those by RBM.
Tu Dinh Nguyen, Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh
ACML1
2013 Learning sparse latent representation and distance metric for image retrieval
abstract
The performance of image retrieval depends critically on the semantic representation and the distance function used to estimate the similarity of two images. A good representation should integrate multiple visual and textual (e.g., tag) features and offer a step closer to the true semantics of interest (e.g., concepts). As the distance function operates on the representation, they are interdependent, and thus should be addressed at the same time. We propose a probabilistic solution to learn both the representation from multiple feature types and modalities and the distance metric from data. The learning is regularised so that the learned representation and information-theoretic metric will (i) preserve the regularities of the visual/textual spaces, (ii) enhance structured sparsity, (iii) encourage small intra-concept distances, and (iv) keep inter-concept images separated. We demonstrate the capacity of our method on the NUS-WIDE data. For the well-studied 13 animal subset, our method outperforms state-of-the-art rivals. On the subset of single-concept images, we gain 79:5% improvement over the standard nearest neighbours approach on the MAP score, and 45.7% on the NDCG.
Tu Dinh Nguyen, Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh
ICME1
2013 Latent Patient Profile Modelling and Applications with Mixed-Variate Restricted Boltzmann Machine
Tu Dinh Nguyen, Truyen Tran 0001, Dinh Q. Phung, Svetha Venkatesh
PAKDD (1)1