Phai Vu Dinh

dblp:320/8648 · also Vu Dinh Phai · DBLP profile ↗
← Back
7ranked-venue papers
7as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 7 · 7 first-author · 7 since 2021
YearPublicationVenuePosition
2026 PFAE: Personalized Federated Learning for Anomaly Detection Over Heterogeneous IoT Domains
Phai Vu Dinh, Marwan Krunz, Diep N. Nguyen, Dinh Thai Hoang
INFOCOM1
2025 A Dual-Decoder Variational Auto-Encoder for Anomaly Detection
abstract
Anomaly detection aims to identify patterns and events that deviate from the norm. However, current methods struggle to achieve high detection accuracy due to data complexity, i.e., imbalance and particularly hidden features not directly observed in the raw data. To address this, we propose a novel neural network architecture/model, called Dual-Decoder Variational Auto-Encoder (DDVAE) which consists of an encoder and two decoders. The encoder maps the input data into the extended latent space, where the dimensionality is greater than that of the input data, providing more room to capture the relationship among features. The data in the extended latent space of DDVAE is modelled based on the distribution of the normal samples to generate stochastic latent variables before they are fed into the first decoder to reconstruct the input data. After that, reconstructed data at the output of the first decoder are used to reconstruct the extended latent space, in which anomalies produce higher reconstruction errors than normal samples when reconstructing both the input data and the extended latent space. The Area Under the ROC Curve (AUC) obtained by DDVAE is significantly greater than that of conventional non-parametric methods and generative models on eight benchmark anomaly datasets, by up to 3.9%. DDVAE also achieves an average miss detection rate of 16.9%. We observe that DDVAE achieves high AUC in cases of a low rate of anomalies compared to normal samples,
Phai Vu Dinh, Diep N. Nguyen, Dinh Thai Hoang, Nguyen Quang Uy, Son Pham Bao, Eryk Dutkiewicz
WCNC1
2024 A Deep Learning Approach for Outlier Detection in Heterogeneous/Non-IID Data
abstract
Outlier/anomaly detection plays a pivotal role in AI applications, e.g., in classification, and intrusion/threat detection in cybersecurity. However, most existing methods face challenges of heterogeneity amongst feature subsets posed by non-independent and identically distributed (non-IID) data. To address this, we propose a novel neural network model called Multiple-Input Auto-Encoder for Anomaly Detection (MIAEAD). MIAEAD assigns an anomaly score to each feature subset of a data sample to indicate its likelihood of being an anomaly. This is done by using the reconstruction error of its sub-encoder as the anomaly score. All sub-encoders are then simultaneously trained using unsupervised learning to determine the anomaly scores of feature subsets. The final Area Under the ROC Curve (AUC) of MIAEAD is determined by the maximum value among the feature subsets. Extensive experiments on eight real-world anomaly datasets from different domains, e.g., health, finance, cybersecurity, and satellite imaging, demonstrate the superior performance of MIAEAD over conventional methods and the state-of-the-art unsupervised models, by up to 4.3% in terms of AUC score. Furthermore, experimental results show that the AUC obtained by MIAEAD is mostly not impacted as the ratio of anomalies to normal samples in the dataset increases. In contrast, the fundamental anomaly detection method using Auto-Encoder (AE) experiences a significant decline in AUC as this ratio increases. We observe that MIAEAD has a high AUC when applied to feature subsets with low heterogeneity based on the coefficient of variation (CV) score. We also prove that the MIAEAD model uses fewer parameters than the AE model.
Phai Vu Dinh, Diep N. Nguyen, Dinh Thai Hoang, Nguyen Quang Uy, Trung Hieu Le, Son Pham Bao, Eryk Dutkiewicz
GLOBECOM1
2024 Multiple-Input Auto-Encoder for IoT Intrusion Detection Systems with Heterogeneous Data
abstract
Machine learning is a core component of many Intrusion Detection Systems (IDS) for IoT networks. However, developing a robust machine-learning model for IDSs on IoT networks is very challenging. This is due to the diversity of IoT devices resulting in the inconsistency and high dimensions of data collected from IoT networks. In other words, in IoT environments, the training data for IDSs on IoT networks is often heterogeneous since they are collected from multiple sources with different characteristics. To tackle these problems, this paper proposes a novel neural network architecture called Multiple-Input Auto-Encoder (MIAE). MIAE has multiple sub-encoders that can process multiple input sources with different dimensions. Moreover, the MIAE model is trained in an unsupervised learning mode, and it can transfer the heterogeneous inputs into lower-dimensional representation to facilitate classifiers to distinguish between the normal samples and types of attacks. The experimental results on the three most popular benchmark IDS datasets, i.e., NSLKDD, UNSW-NB15, and IDS2017, show the superior performance of the MIAE over three groups of methods including conventional classifiers, state-of-the-art dimensionality reduction models, and unsupervised representation learning methods for multiple inputs with different dimensions. MIAE combined with the Random Forest (RF) classifier also achieves 96.2% in terms of accuracy in detecting sophisticated attacks, e.g., Slowloris. In addition, the average running time for detecting an attack sample obtained by MIAE combined with RF classifier is only roughly 5E-7 seconds, whilst the model size is lower than 1 MB. This clearly shows the effectiveness of our proposed model when deployed in practice.
Phai Vu Dinh, Dinh Thai Hoang, Nguyen Quang Uy, Diep N. Nguyen, Son Pham Bao, Eryk Dutkiewicz
ICC1
2024 Constrained Twin Variational Auto-Encoder for Intrusion Detection in IoT Systems
abstract
Intrusion detection systems (IDSs) play a critical role in protecting billions of IoT devices from malicious attacks. However, the IDSs for IoT devices face inherent challenges of IoT systems, including the heterogeneity of IoT data/devices, the high dimensionality of training data, and the imbalanced data. Moreover, the deployment of IDSs on IoT systems is challenging, and sometimes impossible, due to the limited resources, such as memory/storage and computing capability of typical IoT devices. To tackle these challenges, this article proposes a novel deep neural network/architecture called constrained twin variational auto-encoder (CTVAE) that can feed classifiers of IDSs with more separable/distinguishable and lower dimensional representation data. Additionally, in comparison to the state-of-the-art neural networks used in IDSs, CTVAE requires less memory/storage and computing power, hence making it more suitable for IoT IDS systems. Extensive experiments with the 11 most popular IoT botnet data sets show that CTVAE can boost around 1% in terms of accuracy and Fscore in detection attack compared to the state-of-the-art machine learning and representation learning methods, whilst the running time for attack detection is lower than$2E{\mathrm{ -}}6$s and the model size is lower than 1 MB. We also further investigate various characteristics of CTVAE in the latent space and in the reconstruction representation to demonstrate its efficacy compared with current well-known methods.
Phai Vu Dinh, Nguyen Quang Uy, Dinh Thai Hoang, Diep N. Nguyen, Son Pham Bao, Eryk Dutkiewicz
IEEE Internet Things J.1
2022 Balanced Twin Auto-Encoder for IoT Intrusion Detection
abstract
Intrusion detection systems (IDSs) provide an ef-fective solution for protecting loT systems. However, due to the massive number of loT devices (in billions) and their heterogeneity, IDSs face challenges posed by the complexity of loT data such as correlation-based features, high dimensions, and imbalance. To address these problems, this paper proposes a novel neural network architecture, called Balanced Twin Auto-Encoder (BTAE) which consists of three components, i.e., an encoder, a hermaphrodite, and a decoder. The encoder of BTAE first aims to transfer the input data into the latent space before data samples (pre-images) are translated into this space by different translation vectors. In addition, the data of the skewed labels are also generated in the latent space to address the problem of imbalanced data in which the number of attack samples is often significantly lower than those of the benign samples. Second, the hermaphrodite component serves as a bridge to move the data from the encoder to the decoder. Third, the decoder tries to copy the distribution of the samples in the latent space. BTAE is trained by a supervised learning technique, and its data representation extracted from the decoder can well distinguish the attack from the normal data. The experiments on five loT botnet datasets show that BTAE outperforms three existing groups of methods, e.g., the typical supervised learning, the well-known sampling, and the state-of-the-art representation learning. In addition, the false alarm rate (FAR) of BTAE applied for loT intrusion detection is less than equal to 1.2%.
Phai Vu Dinh, Diep N. Nguyen, Dinh Thai Hoang, Nguyen Quang Uy, Son Pham Bao, Eryk Dutkiewicz
GLOBECOM1
2022 Twin Variational Auto-Encoder for Representation Learning in IoT Intrusion Detection
abstract
Intrusion detection systems (IDSs) play a pivotal role in defending IoT systems. However, developing a robust and efficient IDS is challenging due to the rapid and continuing evolving of various forms of cyber-attacks as well as a massive number of low-end IoT devices. In this paper, we introduce a novel deep learning architecture based on auto-encoders that allows to develop a robust intrusion detection system. Specifically, we propose a novel neural network architecture called Twin Variational Auto-Encoder (TVAE) for representation learning. TVAE includes a variational Auto-Encoder (VAE) and an Auto-Encoder (AE) that share a common stage where the decoder of the VAE is used as the encoder of the AE. The TVAE is trained in an unsupervised manner to effectively transform the original representation of data at the input of the VAE into a new representation at the output of the AE. In the new representation space, the difference between normal and attack data is more distinguishable. A variant of TVAE, namely Twin Sparse Variational Auto-Encoder (TSVAE) is also introduced by imposing a sparsity constraint on the representation units. The effectiveness of TVAE and TSVAE is evaluated using popular IDS and IoT botnet datasets. The simulation results show that the accuracy of TVAE and TSVAE can achieve the best results on six datasets, which is higher than those of state-of-the-art AE and VAE variants. We also investigate various characteristics of TVAE in the latent space as well as in the data extraction process. Besides applications on the IoT IDS, TVAE can also be applicable to all conventional network IDSs.
Phai Vu Dinh, Nguyen Quang Uy, Diep N. Nguyen, Dinh Thai Hoang, Son Pham Bao, Eryk Dutkiewicz
WCNC1