VLDB 2026 Research / reviewers in the wild / expert
Nguyen Quang Uy
dblp:18/1212 · also Quang Uy Nguyen
· DBLP profile ↗
27ranked-venue papers
9as first author
9since 2021 · last 2026
0000-0001-7099-5187ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 8 first-author · 3 since 2021Computer networks · 7 · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Attention-guided bidirectional memory autoencoder for video anomaly detection
Nguyen Quang Uy |
Neurocomputing | 2 |
| 2025 | A Dual-Decoder Variational Auto-Encoder for Anomaly DetectionabstractAnomaly detection aims to identify patterns and events that deviate from the norm. However, current methods struggle to achieve high detection accuracy due to data complexity, i.e., imbalance and particularly hidden features not directly observed in the raw data. To address this, we propose a novel neural network architecture/model, called Dual-Decoder Variational Auto-Encoder (DDVAE) which consists of an encoder and two decoders. The encoder maps the input data into the extended latent space, where the dimensionality is greater than that of the input data, providing more room to capture the relationship among features. The data in the extended latent space of DDVAE is modelled based on the distribution of the normal samples to generate stochastic latent variables before they are fed into the first decoder to reconstruct the input data. After that, reconstructed data at the output of the first decoder are used to reconstruct the extended latent space, in which anomalies produce higher reconstruction errors than normal samples when reconstructing both the input data and the extended latent space. The Area Under the ROC Curve (AUC) obtained by DDVAE is significantly greater than that of conventional non-parametric methods and generative models on eight benchmark anomaly datasets, by up to 3.9%. DDVAE also achieves an average miss detection rate of 16.9%. We observe that DDVAE achieves high AUC in cases of a low rate of anomalies compared to normal samples, Phai Vu Dinh, Diep N. Nguyen, Dinh Thai Hoang, Nguyen Quang Uy, Son Pham Bao, Eryk Dutkiewicz |
WCNC | 4 |
| 2024 | A Deep Learning Approach for Outlier Detection in Heterogeneous/Non-IID DataabstractOutlier/anomaly detection plays a pivotal role in AI applications, e.g., in classification, and intrusion/threat detection in cybersecurity. However, most existing methods face challenges of heterogeneity amongst feature subsets posed by non-independent and identically distributed (non-IID) data. To address this, we propose a novel neural network model called Multiple-Input Auto-Encoder for Anomaly Detection (MIAEAD). MIAEAD assigns an anomaly score to each feature subset of a data sample to indicate its likelihood of being an anomaly. This is done by using the reconstruction error of its sub-encoder as the anomaly score. All sub-encoders are then simultaneously trained using unsupervised learning to determine the anomaly scores of feature subsets. The final Area Under the ROC Curve (AUC) of MIAEAD is determined by the maximum value among the feature subsets. Extensive experiments on eight real-world anomaly datasets from different domains, e.g., health, finance, cybersecurity, and satellite imaging, demonstrate the superior performance of MIAEAD over conventional methods and the state-of-the-art unsupervised models, by up to 4.3% in terms of AUC score. Furthermore, experimental results show that the AUC obtained by MIAEAD is mostly not impacted as the ratio of anomalies to normal samples in the dataset increases. In contrast, the fundamental anomaly detection method using Auto-Encoder (AE) experiences a significant decline in AUC as this ratio increases. We observe that MIAEAD has a high AUC when applied to feature subsets with low heterogeneity based on the coefficient of variation (CV) score. We also prove that the MIAEAD model uses fewer parameters than the AE model. Phai Vu Dinh, Diep N. Nguyen, Dinh Thai Hoang, Nguyen Quang Uy, Trung Hieu Le, Son Pham Bao, Eryk Dutkiewicz |
GLOBECOM | 4 |
| 2024 | Multiple-Input Auto-Encoder for IoT Intrusion Detection Systems with Heterogeneous DataabstractMachine learning is a core component of many Intrusion Detection Systems (IDS) for IoT networks. However, developing a robust machine-learning model for IDSs on IoT networks is very challenging. This is due to the diversity of IoT devices resulting in the inconsistency and high dimensions of data collected from IoT networks. In other words, in IoT environments, the training data for IDSs on IoT networks is often heterogeneous since they are collected from multiple sources with different characteristics. To tackle these problems, this paper proposes a novel neural network architecture called Multiple-Input Auto-Encoder (MIAE). MIAE has multiple sub-encoders that can process multiple input sources with different dimensions. Moreover, the MIAE model is trained in an unsupervised learning mode, and it can transfer the heterogeneous inputs into lower-dimensional representation to facilitate classifiers to distinguish between the normal samples and types of attacks. The experimental results on the three most popular benchmark IDS datasets, i.e., NSLKDD, UNSW-NB15, and IDS2017, show the superior performance of the MIAE over three groups of methods including conventional classifiers, state-of-the-art dimensionality reduction models, and unsupervised representation learning methods for multiple inputs with different dimensions. MIAE combined with the Random Forest (RF) classifier also achieves 96.2% in terms of accuracy in detecting sophisticated attacks, e.g., Slowloris. In addition, the average running time for detecting an attack sample obtained by MIAE combined with RF classifier is only roughly 5E-7 seconds, whilst the model size is lower than 1 MB. This clearly shows the effectiveness of our proposed model when deployed in practice. Phai Vu Dinh, Dinh Thai Hoang, Nguyen Quang Uy, Diep N. Nguyen, Son Pham Bao, Eryk Dutkiewicz |
ICC | 3 |
| 2024 | Constrained Twin Variational Auto-Encoder for Intrusion Detection in IoT SystemsabstractIntrusion detection systems (IDSs) play a critical role in protecting billions of IoT devices from malicious attacks. However, the IDSs for IoT devices face inherent challenges of IoT systems, including the heterogeneity of IoT data/devices, the high dimensionality of training data, and the imbalanced data. Moreover, the deployment of IDSs on IoT systems is challenging, and sometimes impossible, due to the limited resources, such as memory/storage and computing capability of typical IoT devices. To tackle these challenges, this article proposes a novel deep neural network/architecture called constrained twin variational auto-encoder (CTVAE) that can feed classifiers of IDSs with more separable/distinguishable and lower dimensional representation data. Additionally, in comparison to the state-of-the-art neural networks used in IDSs, CTVAE requires less memory/storage and computing power, hence making it more suitable for IoT IDS systems. Extensive experiments with the 11 most popular IoT botnet data sets show that CTVAE can boost around 1% in terms of accuracy and Fscore in detection attack compared to the state-of-the-art machine learning and representation learning methods, whilst the running time for attack detection is lower than$2E{\mathrm{ -}}6$s and the model size is lower than 1 MB. We also further investigate various characteristics of CTVAE in the latent space and in the reconstruction representation to demonstrate its efficacy compared with current well-known methods. Phai Vu Dinh, Nguyen Quang Uy, Dinh Thai Hoang, Diep N. Nguyen, Son Pham Bao, Eryk Dutkiewicz |
IEEE Internet Things J. | 2 |
| 2023 | Deep Generative Learning Models for Cloud Intrusion Detection SystemsabstractIntrusion detection (ID) on the cloud environment has received paramount interest over the last few years. Among the latest approaches, machine learning-based ID methods allow us to discover unknown attacks. However, due to the lack of malicious samples and the rapid evolution of diverse attacks, constructing a cloud ID system (IDS) that is robust to a wide range of unknown attacks remains challenging. In this article, we propose a novel solution to enable robust cloud IDSs using deep neural networks. Specifically, we develop two deep generative models to synthesize malicious samples on the cloud systems. The first model, conditional denoising adversarial autoencoder (CDAAE), is used to generate specific types of malicious samples. The second model (CDAEE-KNN) is a hybrid of CDAAE and the K -nearest neighbor algorithm to generate malicious borderline samples that further improve the accuracy of a cloud IDS. The synthesized samples are merged with the original samples to form the augmented datasets. Three machine learning algorithms are trained on the augmented datasets and their effectiveness is analyzed. The experiments conducted on four popular IDS datasets show that our proposed techniques significantly improve the accuracy of the cloud IDSs compared with the baseline technique and the state-of-the-art approaches. Moreover, our models also enhance the accuracy of machine learning algorithms in detecting some currently challenging distributed denial of service (DDoS) attacks, including low-rate DDoS attacks and application layer DDoS attacks. Ly Vu, Nguyen Quang Uy, Diep N. Nguyen, Dinh Thai Hoang, Eryk Dutkiewicz |
IEEE Trans. Cybern. | 2 |
| 2022 | Balanced Twin Auto-Encoder for IoT Intrusion DetectionabstractIntrusion detection systems (IDSs) provide an ef-fective solution for protecting loT systems. However, due to the massive number of loT devices (in billions) and their heterogeneity, IDSs face challenges posed by the complexity of loT data such as correlation-based features, high dimensions, and imbalance. To address these problems, this paper proposes a novel neural network architecture, called Balanced Twin Auto-Encoder (BTAE) which consists of three components, i.e., an encoder, a hermaphrodite, and a decoder. The encoder of BTAE first aims to transfer the input data into the latent space before data samples (pre-images) are translated into this space by different translation vectors. In addition, the data of the skewed labels are also generated in the latent space to address the problem of imbalanced data in which the number of attack samples is often significantly lower than those of the benign samples. Second, the hermaphrodite component serves as a bridge to move the data from the encoder to the decoder. Third, the decoder tries to copy the distribution of the samples in the latent space. BTAE is trained by a supervised learning technique, and its data representation extracted from the decoder can well distinguish the attack from the normal data. The experiments on five loT botnet datasets show that BTAE outperforms three existing groups of methods, e.g., the typical supervised learning, the well-known sampling, and the state-of-the-art representation learning. In addition, the false alarm rate (FAR) of BTAE applied for loT intrusion detection is less than equal to 1.2%. Phai Vu Dinh, Diep N. Nguyen, Dinh Thai Hoang, Nguyen Quang Uy, Son Pham Bao, Eryk Dutkiewicz |
GLOBECOM | 4 |
| 2022 | Twin Variational Auto-Encoder for Representation Learning in IoT Intrusion DetectionabstractIntrusion detection systems (IDSs) play a pivotal role in defending IoT systems. However, developing a robust and efficient IDS is challenging due to the rapid and continuing evolving of various forms of cyber-attacks as well as a massive number of low-end IoT devices. In this paper, we introduce a novel deep learning architecture based on auto-encoders that allows to develop a robust intrusion detection system. Specifically, we propose a novel neural network architecture called Twin Variational Auto-Encoder (TVAE) for representation learning. TVAE includes a variational Auto-Encoder (VAE) and an Auto-Encoder (AE) that share a common stage where the decoder of the VAE is used as the encoder of the AE. The TVAE is trained in an unsupervised manner to effectively transform the original representation of data at the input of the VAE into a new representation at the output of the AE. In the new representation space, the difference between normal and attack data is more distinguishable. A variant of TVAE, namely Twin Sparse Variational Auto-Encoder (TSVAE) is also introduced by imposing a sparsity constraint on the representation units. The effectiveness of TVAE and TSVAE is evaluated using popular IDS and IoT botnet datasets. The simulation results show that the accuracy of TVAE and TSVAE can achieve the best results on six datasets, which is higher than those of state-of-the-art AE and VAE variants. We also investigate various characteristics of TVAE in the latent space as well as in the data extraction process. Besides applications on the IoT IDS, TVAE can also be applicable to all conventional network IDSs. Phai Vu Dinh, Nguyen Quang Uy, Diep N. Nguyen, Dinh Thai Hoang, Son Pham Bao, Eryk Dutkiewicz |
WCNC | 2 |
| 2022 | Learning Latent Representation for IoT Anomaly DetectionabstractInternet of Things (IoT) has emerged as a cutting-edge technology that is changing human life. The rapid and widespread applications of IoT, however, make cyberspace more vulnerable, especially to IoT-based attacks in which IoT devices are used to launch attack on cyber-physical systems. Given a massive number of IoT devices (in order of billions), detecting and preventing these IoT-based attacks are critical. However, this task is very challenging due to the limited energy and computing capabilities of IoT devices and the continuous and fast evolution of attackers. Among IoT-based attacks, unknown ones are far more devastating as these attacks could surpass most of the current security systems and it takes time to detect them and "cure" the systems. To effectively detect new/unknown attacks, in this article, we propose a novel representation learning method to better predictively "describe" unknown attacks, facilitating supervised learning-based anomaly detection methods. Specifically, we develop three regularized versions of autoencoders (AEs) to learn a latent representation from the input data. The bottleneck layers of these regularized AEs trained in a supervised manner using normal data and known IoT attacks will then be used as the new input features for classification algorithms. We carry out extensive experiments on nine recent IoT datasets to evaluate the performance of the proposed models. The experimental results demonstrate that the new latent representation can significantly enhance the performance of supervised learning methods in detecting unknown IoT attacks. We also conduct experiments to investigate the characteristics of the proposed models and the influence of hyperparameters on their performance. The running time of these models is about 1.3 ms that is pragmatic for most applications. Ly Vu, Van Loi Cao, Nguyen Quang Uy, Diep N. Nguyen, Dinh Thai Hoang, Eryk Dutkiewicz |
IEEE Trans. Cybern. | 3 |
| 2019 | Learning Latent Distribution for Distinguishing Network Traffic in Intrusion Detection SystemabstractWe develop a novel deep learning model, Multidistributed Variational AutoEncoder (MVAE), for the network intrusion detection. To make the traffic more distinguishable, MVAE introduces the label information of data samples into the Kullback-Leibler (KL) term of the loss function of Variational AutoEncoder (VAE). This label information allows MVAEs to force/partition network data samples into different classes with different regions in the latent feature space. As a result, the network traffic samples are more distinguishable in the new representation space (i.e., the latent feature space of MVAE), thereby improving the accuracy in detecting intrusions. To evaluate the efficiency of the proposed solution, we carry out intensive experiments on two popular network intrusion datasets, i.e., NSL-KDD and UNSWNB15 under four conventional classifiers including Gaussian Naive Bayes (GNB), Support Vector Machine (SVM), Decision Tree (DT), and Random Forest (RF). The experimental results demonstrate that our proposed approach can significantly improve the accuracy of intrusion detection algorithms up to 24.6% compared to the original one (using area under the curve metric). Ly Vu, Van Loi Cao, Nguyen Quang Uy, Diep N. Nguyen, Dinh Thai Hoang, Eryk Dutkiewicz |
ICC | 3 |
| 2018 | Semantic tournament selection for genetic programming based on statistical analysis of error vectors
Thi Huong Chu, Nguyen Quang Uy, Michael O'Neill 0001 |
Inf. Sci. | 2 |
| 2016 | Tournament Selection Based on Statistical Test in Genetic Programming
Thi Huong Chu, Nguyen Quang Uy, Michael O'Neill 0001 |
PPSN | 2 |
| 2015 | Transfer learning in Genetic ProgrammingabstractTransfer learning is a process in which a system can apply knowledge and skills learned in previous tasks to novel tasks. This technique has emerged as a new framework to enhance the performance of learning methods in machine learning. Surprisingly, transfer learning has not deservedly received the attention from the Genetic Programming research community. In this paper, we propose several transfer learning methods for Genetic Programming (GP). These methods were implemented by transferring a number of good individuals or sub-individuals from the source to the target problem. They were tested on two families of symbolic regression problems. The experimental results showed that transfer learning methods help GP to achieve better training errors. Importantly, the performance of GP on unseen data when implemented with transfer learning was also considerably improved. Furthermore, the impact of transfer learning to GP code bloat was examined that showed that limiting the size of transferred individuals helps to reduce the code growth problem in GP. Thi Thu Huong Dinh, Thi Huong Chu, Nguyen Quang Uy |
CEC | 3 |
| 2014 | Malware detection using genetic programmingabstractMalware is any software aiming to disrupt computer operation. Malware is also used to gather sensitive information or gain access to private computer systems. This is widely seen as one of the major threats to computer systems nowadays. Traditionally, anti-malware software is based on a signature detection system which keeps updating from the Internet malware database and thus keeping track of known malwares. While this method may be very accurate to detect previously known malwares, it is unable to detect unknown malicious codes. Recently, several machine learning methods have been used for malware detection, achieving remarkable success. In this paper, we propose a method in this strand by using Genetic Programming for detecting malwares. The experiments were conducted with the malwares collected from an updated malware database on the Internet and the results show that Genetic Programming, compared to some other well-known machine learning methods, can produce the best results on both balanced and imbalanced datasets. Thi Anh Le 0002, Thi Huong Chu, Nguyen Quang Uy, Nguyen Xuan Hoai |
CISDA | 3 |
| 2013 | Examining the Diversity Property of Semantic Similarity Based Crossover
Pham Tuan Anh, Nguyen Quang Uy, Nguyen Xuan Hoai, Michael O'Neill 0001 |
EuroGP | 2 |
| 2013 | On the roles of semantic locality of crossover in genetic programming
Nguyen Quang Uy, Nguyen Xuan Hoai, Michael O'Neill 0001, Robert I. McKay, Dao Ngoc Phong |
Inf. Sci. | 1 |
| 2012 | Evolving approximations for the Gaussian Q-function by Genetic Programming with semantic based crossoverabstractThe Gaussian Q-function is of great importance in the field of communications, where the noise is often characterized by the Gaussian distribution. However, no simple exact closed form of the Q-function is known. Consequently, a number of approximations have been proposed over the past several decades. In this paper, we use Genetic Programming with semantic based crossover to approximate the Q-function in two forms: the free and the exponential forms. Using this form, we found approximations in both forms that are more accurate than all previous approximations designed by human experts. Dao Ngoc Phong, Nguyen Quang Uy, Nguyen Xuan Hoai, Robert I. McKay |
IEEE Congress on Evolutionary Computation | 2 |
| 2012 | An Investigation of Fitness Sharing with Semantic and Syntactic Distance Metrics
Nguyen Quang Uy, Nguyen Xuan Hoai, Michael O'Neill 0001, Alexandros Agapitos |
EuroGP | 1 |
| 2012 | Evolving the best known approximation to the Q functionabstractThe Gaussian Q-function is the integral of the tail of the Gaussian distribution; as such, it is important across a vast range of fields requiring stochastic analysis. No elementary closed form is possible, so a number of approximations have been proposed. We use a Genetic Programming (GP) system, Tree Adjoining Grammar Guided GP (TAG3P) with local search operators to evolve approximations of the Q-function in the form given by Benitez [1]. We found more accurate approximations than any previously published. This confirms the practical importance of local search in TAG3P. Dao Ngoc Phong, Nguyen Xuan Hoai, Robert I. McKay, Constantin Siriteanu, Nguyen Quang Uy, Namyong Park 0001 |
GECCO | 5 |
| 2011 | Examining the landscape of semantic similarity based mutationabstractThis paper examines how the semantic locality of a search operator affects the fitness landscape of Genetic Programming (GP). We compare the fitness landscapes of GP search when standard subtree mutation and arecently proposed semantic-based mutation, Semantic Similarity-based Mutation (SSM), are used. The comparison is based on two well-studied fitness landscape measures, namely, the autocorrelation function and information content. The experiments were conducted on a family of symbolic regression problems with increasing degrees of difficulty. The results show that SSM helps to significantly smooth out the fitness landscape of GP compared to standard subtree mutation. This gives an explanation for the better performance of SSM over standard subtree mutation operator. Nguyen Quang Uy, Nguyen Xuan Hoai, Michael O'Neill 0001 |
GECCO | 1 |
| 2010 | Self-adapting semantic sensitivities for Semantic Similarity based CrossoverabstractThis paper presents two methods for self-adapting the semantic sensitivities in a recently proposed semantics-based crossover: Semantic Similarity based Crossover (SSC). The first self-adaptation method is inspired by a self-adaptive method for controlling mutation step size in Evolutionary Strategies (1/5 rule). The design of the second takes into account more of our previous experimental observations, that SSC works well only when a certain portion of events successfully exchange semantically similar subtrees. These two proposed methods are then tested on a number of real-valued symbolic regression problems, their performance being compared with SSC using predetermined sensitivities and with standard crossover. The results confirm the benefits of the second self-adaption method. Nguyen Quang Uy, Robert I. McKay, Michael O'Neill 0001, Nguyen Xuan Hoai |
IEEE Congress on Evolutionary Computation | 1 |
| 2010 | Improving the Generalisation Ability of Genetic Programming with Semantic Similarity based Crossover
Nguyen Quang Uy, Nguyen Thi Hien, Nguyen Xuan Hoai, Michael O'Neill 0001 |
EuroGP | 1 |
| 2010 | Semantics based crossover for boolean problemsabstractThis paper investigates the role of semantic diversity and locality of crossover operators in. Genetic Programming (GP) for Boolean problems. We propose methods for measuring and storing semantics of subtrees in Boolean domains using Trace Semantics, and design several new crossovers on this basis. They can be categorised into two classes depending on their purposes: promoting semantic diversity or improving semantic locality. We test the operators on several well-known Boolean problems, comparing them with Standard GP Crossovers and with the Semantic Driven Crossover of Beadle and Johnson. The experimental results show the positive effects both of promoting semantic diversity, and of improving semantic locality, in crossover operators. They also show that the latter has a greater positive effect on GP performance than the former. Nguyen Quang Uy, Nguyen Xuan Hoai, Michael O'Neill 0001, Robert I. McKay |
GECCO | 1 |
| 2010 | The Role of Syntactic and Semantic Locality of Crossover in Genetic Programming
Nguyen Quang Uy, Nguyen Xuan Hoai, Michael O'Neill 0001, Robert I. McKay |
PPSN (2) | 1 |
| 2009 | Semantic Aware Crossover for Genetic Programming: The Case for Real-Valued Function Regression
Nguyen Quang Uy, Nguyen Xuan Hoai, Michael O'Neill 0001 |
EuroGP | 1 |
| 2007 | Initialising PSO with randomised low-discrepancy sequences: the comparative resultsabstractIn this paper, we investigate the use of some wel-known randomised low-discrepancy sequences (Halton, Sobol, and Faure sequences) for initialising particle swarms. We experimented with the standard global-best particle swarm algorithm for function optimization on some benchmark problems, using randomised low-discrepancy sequences for initialisation, and the results were compared with the same particle swarm algorithm using uniform initialisation with a pseudo-random generator. The results show that, the former initialisation method could help the particle swarm algorithm improve its performance over the latter on the problems tried. Furthermore the comparisons also indicate that the use of different randomised low-discrepancy sequences in the initialisation phase could bring different effects on the performance of PSO. Nguyen Quang Uy, Nguyen Xuan Hoai, Robert I. McKay, Pham Minh Tuan |
IEEE Congress on Evolutionary Computation | 1 |
| 2007 | PSO with randomized low-discrepancy sequencesabstractWe initialize a global-best particle swarm with a Halton sequence, comparing it with uniform initialization on a range of benchmark function optimization problems. We see substantial improvements in performance, particularly with high complexity problems /small populations. Halton initialization yields equivalent performance to uniform initialization with substantially smaller populations. Nguyen Xuan Hoai, Nguyen Quang Uy, Robert I. McKay |
GECCO | 2 |