Van Nguyen 0002

dblp:94/5094-2 · DBLP profile ↗
← Back
28ranked-venue papers
9as first author
19since 2021 · last 2026
0000-0002-5838-3409ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 5 first-author · 5 since 2021Databases, data management, data science and information retrieval · 9 · 2 first-author · 6 since 2021Software engineering, systems software and programming languages · 7 · 1 first-author · 7 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 An overview of evaluation and enhancement methods for code generation by large language models
abstract
Context: Recent advances in Large Language Models (LLMs) have led to the rapid deployment of automated generation tools capable of producing source code. As these models increasingly transition from being experimental tools to established elements of the software development, a critical question arises: to what extent do the models and the code they generate satisfy, or can be made to satisfy, the rigorous, multifaceted quality standards required for professional, real-world engineering? Objective: The primary aim of this study is to find the answer to this question by exploring existing evaluation frameworks and enhancement strategies for LLMs and the code they generate. By examining how generated code quality is currently assessed and improved, we hope to determine if the current research methodologies provide a balanced coverage of the software quality spectrum or if significant disparities exist. Method: We propose a code quality dimension taxonomy adapted from the ISO/IEC 25010 standard, encompassing four principal attributes: Functional Correctness (FC), Security (SE), Performance Efficiency (PE), and Maintainability (MA). Using this framework, we conduct a literature review analysing existing research in evaluation frameworks and enhancement strategies across these dimensions. Results: Our analysis reveals a substantial imbalance in research focus. FC, and increasingly, SE have well-established evaluation frameworks and improvement strategies. In contrast, PE and MA remain significantly underexamined, with few standardised benchmarks and a lack of targeted fine-tuning approaches for these critical software quality dimensions. Conclusion: The survey identifies a pressing need for broader research into PE and MA-oriented evaluation and enhancement. We propose several promising directions: (i) the creation of formal benchmarks; (ii) the development of reinforcement learning techniques leveraging static and dynamic code feedback; and (iii) the use of multi-agent frameworks for iterative, critique-based improvement grounded in verifiable diagnostic artefacts.
Jacob Truong, Van Nguyen 0002, Thanh Thi Nguyen 0001
Inf. Softw. Technol.2
2025 SAFE: A Novel Approach For Software Vulnerability Detection from Enhancing The Capability of Large Language Models
abstract
Software vulnerabilities (SVs) have emerged as a prevalent and crucial concern for safety-critical systems. This has spurred significant advancements in utilizing AI-based methods, including machine learning and deep learning, for software vulnerability detection (SVD). While AI-based methods have shown promising performance in SVD, their effectiveness on real-world, complex, and diverse source code datasets remains limited in practice. To tackle this challenge, in this paper, we propose a novel framework that enhances the capability of large language models to learn and utilize semantic and syntactic relationships from source code data for SVD. As a result, our proposed SAFE approach can enable the acquisition of fundamental knowledge from source code data while adeptly utilizing crucial relationships, i.e., semantic and syntactic associations, to improve the effectiveness of solving the SVD problem. The rigorous and extensive experimental results on three real-world challenging datasets (i.e., Devign, ReVeal, and D2A) demonstrate the superiority of our approach over eight effective and state-of-the-art baselines. In summary, on average, our SAFE approach achieves higher performances from 4.79% to 11.57% for F1-measure and from 16.93% to 26.24% for Recall compared to the baseline methods across all the datasets used.
Van Nguyen 0002, Surya Nepal, Xingliang Yuan, Tingmin Wu, Carsten Rudolph
AsiaCCS1
2025 AI2TALE: An Innovative Information Theory-based Approach for Learning to Localize Phishing Attacks
abstract
Phishing attacks remain a significant challenge for detection, explanation, and defense, despite over a decade of research on both technical and non-technical solutions. AI-based phishing detection methods are among the most effective approaches for defeating phishing attacks, providing predictions on the vulnerability label (i.e., phishing or benign) of data. However, they often lack intrinsic explainability, failing to identify the specific information that triggers the classification. To this end, we propose AI2TALE, an innovative deep learning-based approach for email (the most common phishing medium) phishing attack localization. Our method aims to not only predict the vulnerability label of the email data but also provide the capability to automatically learn and identify the most important and phishing-relevant information (i.e., sentences) in the phishing email data, offering useful and concise explanations for the identified vulnerability. Extensive experiments on seven diverse real-world email datasets demonstrate the capability and effectiveness of our method in selecting crucial information, enabling accurate detection and offering useful and concise explanations (via the most important and phishing-relevant information triggering the classification) for the vulnerability of phishing emails. Notably, our approach outperforms state-of-the-art baselines by 1.5% to 3.5% on average in Label-Accuracy and Cognitive-True-Positive metrics under a weakly supervised setting, where only vulnerability labels are used without requiring ground truth phishing information.
Van Nguyen 0002, Tingmin Wu, Xingliang Yuan, Marthie Grobler, Surya Nepal, Carsten Rudolph
ICLR1
2025 DeepVulMatch: Learning and Matching Latent Vulnerability Representations for Dual-Granularity Vulnerability Detection
abstract
Deep learning (DL) models are widely used to detect software vulnerabilities, but identifying vulnerabilities at the line level remains challenging due to varied coding styles and the spread of vulnerabilities across multiple lines. We observe that vulnerable line embeddings tend to form clusters in the feature space, which can help models capture hidden patterns more effectively. In this article, we propose a novel approach that leverages vector quantization (VQ) and optimal transport (OT) to exploit the clustering characteristics of vulnerable line embeddings and enhance detection performance. Specifically, we extract vulnerable line embeddings from the training data to form a vulnerability collection, which we condense into a compact vulnerability codebook using VQ and OT. Inspired by static analysis tools that rely on pattern matching, our model uses this codebook to match latent vulnerability representations during inference. Our approach also introduces dual-granularity detection, predicting both vulnerable functions and, when a function is predicted vulnerable, identifying the specific vulnerable lines within it. We evaluate our approach against 12 baselines on two large-scale datasets of real-world open-source vulnerabilities. Our method achieves the highest F1 scores at both the function and line levels.
Trung Le 0001, Van Nguyen 0002, Chakkrit Tantithamthavorn, Dinh Q. Phung
IEEE Trans. Reliab.3
2024 Decompose, Enrich, and Extract! Schema-aware Event Extraction using LLMs
abstract
Large Language Models (LLMs) demonstrate significant capabilities in processing natural language data, promising efficient knowledge extraction from diverse textual sources to enhance situational awareness and support decision-making. However, concerns arise due to their susceptibility to hallucination, resulting in contextually inaccurate content. This work focuses on harnessing LLMs for automated Event Extraction, introducing a new method to address hallucination by decomposing the task into Event Detection and Event Argument Extraction. Moreover, the proposed method integrates dynamic schema-aware augmented retrieval examples into prompts tailored for each specific inquiry, thereby extending and adapting advanced prompting techniques such as Retrieval-Augmented Generation. Evaluation findings on prominent event extraction benchmarks and results from a synthesized benchmark illustrate the method’s superior performance compared to baseline approaches.
Fatemeh Shiri, Farhad Moghimifar, Gholamreza Haffari, Yuan-Fang Li, Van Nguyen 0002, John Yoo
FUSION5
2024 AIBugHunter: A Practical tool for predicting, classifying and repairing software vulnerabilities
abstract
Abstract Many Machine Learning(ML)-based approaches have been proposed to automatically detect, localize, and repair software vulnerabilities. While ML-based methods are more effective than program analysis-based vulnerability analysis tools, few have been integrated into modern Integrated Development Environments (IDEs), hindering practical adoption. To bridge this critical gap, we propose in this article AIBugHunter , a novel Machine Learning-based software vulnerability analysis tool for C/C++ languages that is integrated into the Visual Studio Code (VS Code) IDE. AIBugHunter helps software developers to achieve real-time vulnerability detection, explanation, and repairs during programming. In particular, AIBugHunter scans through developers’ source code to (1) locate vulnerabilities, (2) identify vulnerability types, (3) estimate vulnerability severity, and (4) suggest vulnerability repairs. We integrate our previous works (i.e., LineVul and VulRepair) to achieve vulnerability localization and repairs. In this article, we propose a novel multi-objective optimization (MOO)-based vulnerability classification approach and a transformer-based estimation approach to help AIBugHunter accurately identify vulnerability types and estimate severity. Our empirical experiments on a large dataset consisting of 188K+ C/C++ functions confirm that our proposed approaches are more accurate than other state-of-the-art baseline methods for vulnerability classification and estimation. Furthermore, we conduct qualitative evaluations including a survey study and a user study to obtain software practitioners’ perceptions of our AIBugHunter tool and assess the impact that AIBugHunter may have on developers’ productivity in security aspects. Our survey study shows that our AIBugHunter is perceived as useful where 90% of the participants consider adopting our AIBugHunter during their software development. Last but not least, our user study shows that our AIBugHunter can enhance developers’ productivity in combating cybersecurity issues during software development. AIBugHunter is now publicly available in the Visual Studio Code marketplace.
Chakkrit Tantithamthavorn, Trung Le 0001, Yuki Kume, Van Nguyen 0002, Dinh Q. Phung, John C. Grundy
Empir. Softw. Eng.5
2024 Vision Transformer Inspired Automated Vulnerability Repair
abstract
Recently, automated vulnerability repair approaches have been widely adopted to combat increasing software security issues. In particular, transformer-based encoder-decoder models achieve competitive results. Whereas vulnerable programs may only consist of a few vulnerable code areas that need repair, existing AVR approaches lack a mechanism guiding their model to pay more attention to vulnerable code areas during repair generation. In this article, we propose a novel vulnerability repair framework inspired by the Vision Transformer based approaches for object detection in the computer vision domain. Similar to the object queries used to locate objects in object detection in computer vision, we introduce and leverage vulnerability queries (VQs) to locate vulnerable code areas and then suggest their repairs. In particular, we leverage the cross-attention mechanism to achieve the cross-match between VQs and their corresponding vulnerable code areas. To strengthen our cross-match and generate more accurate vulnerability repairs, we propose to learn a novel vulnerability mask (VM) and integrate it into decoders’ cross-attention, which makes our VQs pay more attention to vulnerable code areas during repair generation. In addition, we incorporate our VM into encoders’ self-attention to learn embeddings that emphasize the vulnerable areas of a program. Through an extensive evaluation using the real-world 5,417 vulnerabilities, our approach outperforms all of the automated vulnerability repair baseline methods by 2.68% to 32.33%. Additionally, our analysis of the cross-attention map of our approach confirms the design rationale of our VM and its effectiveness. Finally, our survey study with 71 software practitioners highlights the significance and usefulness of AI-generated vulnerability repairs in the realm of software security. The training code and pre-trained models are available at https://github.com/awsm-research/VQM.
Van Nguyen 0002, Chakkrit Tantithamthavorn, Dinh Q. Phung, Trung Le 0001
ACM Trans. Softw. Eng. Methodol.2
2024 Deep Domain Adaptation With Max-Margin Principle for Cross-Project Imbalanced Software Vulnerability Detection
abstract
Software vulnerabilities (SVs) have become a common, serious, and crucial concern due to the ubiquity of computer software. Many AI-based approaches have been proposed to solve the software vulnerability detection (SVD) problem to ensure the security and integrity of software applications (in both the development and testing phases). However, there are still two open and significant issues for SVD in terms of (i) learning automatic representations to improve the predictive performance of SVD, and (ii) tackling the scarcity of labeled vulnerability datasets that conventionally need laborious labeling effort by experts. In this paper, we propose a novel approach to tackle these two crucial issues. We first exploit the automatic representation learning with deep domain adaptation for SVD. We then propose a novel cross-domain kernel classifier leveraging the max-margin principle to significantly improve the transfer learning process of SVs from imbalanced labeled into imbalanced unlabeled projects. Our approach is the first work that leverages solid body theories of the max-margin principle, kernel methods, and bridging the gap between source and target domains for imbalanced domain adaptation (DA) applied in cross-project SVD . The experimental results on real-world software datasets show the superiority of our proposed method over state-of-the-art baselines. In short, our method obtains a higher performance on F1-measure, one of the most important measures in SVD, from 1.83% to 6.25% compared to the second highest method in the used datasets.
Van Nguyen 0002, Trung Le 0001, Chakkrit Tantithamthavorn, John C. Grundy, Dinh Q. Phung
ACM Trans. Softw. Eng. Methodol.1
2023 ChatGPT for Vulnerability Detection, Classification, and Repair: How Far Are We?
abstract
Large language models (LLMs) like ChatGPT (i.e., gpt-3.5-turbo and gpt-4) exhibited remarkable advancement in a range of software engineering tasks associated with source code such as code review and code generation. In this paper, we undertake a comprehensive study by instructing ChatGPT for four prevalent vulnerability tasks: function and line-level vulnerability prediction, vulnerability classification, severity estimation, and vulnerability repair. We compare ChatGPT with state-of-the-art language models designed for software vulnerability purposes. Through an empirical assessment employing extensive real-world datasets featuring over 190,000 C/C++ functions, we found that ChatGPT achieves limited performance, trailing behind other language models in vulnerability contexts by a significant margin. The experimental outcomes highlight the challenging nature of vulnerability prediction tasks, requiring domain-specific expertise. Despite ChatGPT's substantial model scale, exceeding that of source code-pre-trained language models (e.g., CodeBERT) by a factor of 14,000, the process of fine-tuning remains imperative for ChatGPT to generalize for vulnerability prediction tasks. We publish the studied dataset, experimental prompts for ChatGPT, and experimental results at https://github.com/awsm-research/ChatGPT4Vul.
Chakkrit Tantithamthavorn, Van Nguyen 0002, Trung Le 0001
APSEC3
2023 Bounded Subjective Opinions
abstract
The shifted Dirichlet distribution is utilized to extend the definition of subjective opinions to include upper and lower bounds on belief and base rates. The characteristics of the bounded subjective opinions are examined and contrasted with unbounded subjective opinions.
Michael McDonald 0001, Lance M. Kaplan, Van Nguyen 0002
FUSION3
2023 Few-shot Domain-Adaptative Visually-fused Event Detection from Text
abstract
Incorporating auxiliary modalities such as images into event detection models has attracted increasing interest over the last few years. The complexity of natural language in describing situations has motivated researchers to leverage the related visual context to improve event detection performance. However, current approaches in this area suffer from data scarcity, where a large amount of labelled text-image pairs are required for model training. Furthermore, limited access to the visual context at inference time negatively impacts the performance of such models, which makes them practically ineffective in real-world scenarios. In this paper, we present a novel domain-adaptive visually-fused event detection approach that can be trained on a few labelled image-text paired data points. Specifically, we introduce a visual imaginator method that synthesises images from text in the absence of visual context. Moreover, the imaginator can be customised to a specific domain. In doing so, our model can leverage the capabilities of pre-trained vision-language models and can be trained in a few-shot setting. This also allows for effective inference where only single-modality data (i.e. text) is available. The experimental evaluation on the benchmark M2E2 dataset shows that our model outperforms existing state-of-the-art models, by up to 11 points.
Farhad Moghimifar, Fatemeh Shiri, Gholamreza Haffari, Yuan-Fang Li, Van Nguyen 0002
FUSION5
2023 An Additive Instance-Wise Approach to Multi-class Model Interpretation
Vy Vo, Van Nguyen 0002, Trung Le 0001, Quan Hung Tran, Gholamreza Haffari, Seyit Ahmet Çamtepe, Dinh Q. Phung
ICLR2
2023 Feature-based Learning for Diverse and Privacy-Preserving Counterfactual Explanations
abstract
Interpretable machine learning seeks to understand the reasoning process of complex black-box systems that are long notorious for lack of explainability. One flourishing approach is through counterfactual explanations, which provide suggestions on what a user can do to alter an outcome. Not only must a counterfactual example counter the original prediction from the black-box classifier but it should also satisfy various constraints for practical applications. Diversity is one of the critical constraints that however remains less discussed. While diverse counterfactuals are ideal, it is computationally challenging to simultaneously address some other constraints. Furthermore, there is a growing privacy concern over the released counterfactual data. To this end, we propose a feature-based learning framework that effectively handles the counterfactual constraints and contributes itself to the limited pool of private explanation models. We demonstrate the flexibility and effectiveness of our method in generating diverse counterfactuals of actionability and plausibility. Our counterfactual engine is more efficient than counterparts of the same capacity while yielding the lowest re-identification risks.
Vy Vo, Trung Le 0001, Van Nguyen 0002, He Zhao 0001, Edwin V. Bonilla, Gholamreza Haffari, Dinh Q. Phung
KDD3
2023 VulExplainer: A Transformer-Based Hierarchical Distillation for Explaining Vulnerability Types
abstract
Deep learning-based vulnerability prediction approaches are proposed to help under-resourced security practitioners to detect vulnerable functions. However, security practitioners still do not know what type of vulnerabilities correspond to a given prediction (aka CWE-ID). Thus, a novel approach to explain the type of vulnerabilities for a given prediction is imperative. In this paper, we proposeVulExplainer, an approach to explain the type of vulnerabilities. We representVulExplaineras a vulnerability classification task. However, vulnerabilities have diverse characteristics (i.e., CWE-IDs) and the number of labeled samples in each CWE-ID is highly imbalanced (known as a highly imbalanced multi-class classification problem), which often lead to inaccurate predictions. Thus, we introduce a Transformer-based hierarchical distillation for software vulnerability classification in order to address the highly imbalanced types of software vulnerabilities. Specifically, we split a complex label distribution into sub-distributions based on CWE abstract types (i.e., categorizations that group similar CWE-IDs). Thus, similar CWE-IDs can be grouped and each group will have a more balanced label distribution. We learn TextCNN teachers on each of the simplified distributions respectively, however, they only perform well in their group. Thus, we build a transformer student model to generalize the performance of TextCNN teachers through our hierarchical knowledge distillation framework. Through an extensive evaluation using the real-world 8,636 vulnerabilities, our approach outperforms all of the baselines by 5%–29%. The results also demonstrate that our approach can be applied to Transformer-based architectures such as CodeBERT, GraphCodeBERT, and CodeGPT. Moreover, our method maintains compatibility with any Transformer-based model without requiring any architectural modifications but only adds a special distillation token to the input. These results highlight our significant contributions towards the fundamental and practical problem of explaining software vulnerability.
Van Nguyen 0002, Chakkrit Tantithamthavorn, Trung Le 0001, Dinh Q. Phung
IEEE Trans. Software Eng.2
2022 Paraphrasing Techniques for Maritime QA system
Fatemeh Shiri, Terry Yue Zhuo, Zhuang Li 0001, Shirui Pan, Weiqing Wang 0001, Gholamreza Haffari, Yuan-Fang Li, Van Nguyen 0002
FUSION8
2022 VulRepair: a T5-based automated software vulnerability repair
abstract
As software vulnerabilities grow in volume and complexity, researchers proposed various Artificial Intelligence (AI)-based approaches to help under-resourced security analysts to find, detect, and localize vulnerabilities. However, security analysts still have to spend a huge amount of effort to manually fix or repair such vulnerable functions. Recent work proposed an NMT-based Automated Vulnerability Repair, but it is still far from perfect due to various limitations. In this paper, we propose VulRepair, a T5-based automated software vulnerability repair approach that leverages the pre-training and BPE components to address various technical limitations of prior work. Through an extensive experiment with over 8,482 vulnerability fixes from 1,754 real-world software projects, we find that our VulRepair achieves a Perfect Prediction of 44%, which is 13%-21% more accurate than competitive baseline approaches. These results lead us to conclude that our VulRepair is considerably more accurate than two baseline approaches, highlighting the substantial advancement of NMT-based Automated Vulnerability Repairs. Our additional investigation also shows that our VulRepair can accurately repair as many as 745 out of 1,706 real-world well-known vulnerabilities (e.g., Use After Free, Improper Input Validation, OS Command Injection), demonstrating the practicality and significance of our VulRepair for generating vulnerability repairs, helping under-resourced security analysts on fixing vulnerabilities.
Chakkrit Tantithamthavorn, Trung Le 0001, Van Nguyen 0002, Dinh Q. Phung
ESEC/SIGSOFT FSE4
2022 Cycle class consistency with distributional optimal transport and knowledge distillation for unsupervised domain adaptation
abstract
Unsupervised domain adaptation (UDA) aims to transfer knowledge from a model trained on a labeled source domain to an unlabeled target domain. To this end, we propose in this paper a novel cycle class-consistent model based on optimal transport (OT) and knowledge distillation. The model consists of two agents, a teacher and a student cooperatively working in a cycle process under the guidance of the distributional optimal transport and distillation manner. The OT distance is designed to bridge the gap between the distribution of the target data and a distribution over the source class-conditional distributions. The optimal probability matrix then provides pseudo labels to learn a teacher that achieves a good classification performance on the target domain. Knowledge distillation is performed in the next step in which the teacher distills and transfers its knowledge to the student. And finally, the student produces its prediction for the optimal transport step. This process forms a closed cycle in which the teacher and student networks are simultaneously trained to conduct transfer learning from the source to the target domain. Extensive experiments show that our proposed method outperforms existing methods, especially the class-aware and OT-based ones on benchmark datasets including Office-31, Office-Home, and ImageCLEF-DA.
Tuan Nguyen 0004, Van Nguyen 0002, Trung Le 0001, He Zhao 0001, Quan Hung Tran, Dinh Q. Phung
UAI2
2021 Toward the Automated Construction of Probabilistic Knowledge Graphs for the Maritime Domain
Fatemeh Shiri, Teresa Wang, Shirui Pan, Xiaojun Chang, Yuan-Fang Li, Gholamreza Haffari, Van Nguyen 0002
FUSION7
2021 Information-theoretic Source Code Vulnerability Highlighting
abstract
Software vulnerabilities are a crucial and serious concern in the software industry and computer security. A variety of methods have been proposed to detect vulnerabilities in real-world software. Recent methods based on deep learning approaches for automatic feature extraction have improved software vulnerability identification compared with machine learning approaches based on hand-crafted feature extraction. However, these methods can usually only detect software vulnerabilities at a function or program level, which is much less informative because, out of hundreds (thousands) of code statements in a program or function, only a few core statements contribute to a software vulnerability. This requires us to find a way to detect software vulnerabilities at a fine-grained level. In this paper, we propose a novel method based on the concept of mutual information that can help us to detect and isolate software vulnerabilities at a fine-grained level (i.e., several statements that are highly relevant to a software vulnerability that include the core vulnerable statements) in both unsupervised and semi-supervised contexts. We conduct comprehensive experiments on real-world software projects to demonstrate that our proposed method can detect vulnerabilities at a fine-grained level by identifying several statements that mostly contribute to the vulnerability detection decision.
Van Nguyen 0002, Trung Le 0001, Olivier Y. de Vel, Paul Montague, John C. Grundy, Dinh Q. Phung
IJCNN1
2020 Code Pointer Network for Binary Function Scope Identification
abstract
Function identification is a preliminary step in binary analysis for many extensive applications from malware detection, common vulnerability detection and binary instrumentation to name a few. In this paper, we propose the Code Pointer Network that leverages the underlying idea of a pointer network to efficiently and effectively tackle function scope identification - the hardest and most crucial task in function identification. We establish extensive experiments to compare our proposed method with the deep learning based baseline. Experimental results demonstrate that our proposed method significantly outperforms the state-of-the-art baseline in terms of both predictive performance and running time.
Van Nguyen 0002, Trung Le 0001, Tue Le, Olivier Y. de Vel, Paul Montague, Dinh Q. Phung
IJCNN1
2020 Code Action Network for Binary Function Scope Identification
Van Nguyen 0002, Trung Le 0001, Tue Le, Olivier Y. de Vel, Paul Montague, John C. Grundy, Dinh Q. Phung
PAKDD (1)1
2020 Dual-Component Deep Domain Adaptation: A New Approach for Cross Project Software Vulnerability Detection
Van Nguyen 0002, Trung Le 0001, Olivier Y. de Vel, Paul Montague, John C. Grundy, Dinh Q. Phung
PAKDD (1)1
2019 Deep Domain Adaptation for Vulnerable Code Function Identification
abstract
Due to the ubiquity of computer software, software vulnerability detection (SVD) has become crucial in the software industry and in the field of computer security. Two significant issues in SVD arise when using machine learning, namely: i) how to learn automatic features that can help improve the predictive performance of vulnerability detection and ii) how to overcome the scarcity of labeled vulnerabilities in projects that require the laborious labeling of code by software security experts. In this paper, we address these two crucial concerns by proposing a novel architecture which leverages deep domain adaptation with automatic feature learning for software vulnerability identification. Based on this architecture, we keep the principles and reapply the state-of-the-art deep domain adaptation methods to indicate that deep domain adaptation for SVD is plausible and promising. Moreover, we further propose a novel method named Semi-supervised Code Domain Adaptation Network (SCDAN) that can efficiently utilize and exploit information carried in unlabeled target data by considering them as the unlabeled portion in a semi-supervised learning context. The proposed SCDAN method enforces the clustering assumption, which is a key principle in semi-supervised learning. The experimental results using six real-world software project datasets show that our SCDAN method and the baselines using our architecture have better predictive performance by a wide margin compared with the Deep Code Network (VulDeePecker) method without domain adaptation. Also, the proposed SCDAN significantly outperforms the DIRT-T which to the best of our knowledge is currently the-state-of-the-art method in deep domain adaptation and other baselines.
Van Nguyen 0002, Trung Le 0001, Tue Le, Olivier Y. de Vel, Paul Montague, Lizhen Qu, Dinh Q. Phung
IJCNN1
2018 Jointly Predicting Affective and Mental Health Scores Using Deep Neural Networks of Visual Cues on the Web
Van Nguyen 0002, Thin Nguyen, Mark E. Larsen, Bridianne O'Dea, Duc Thanh Nguyen, Trung Le 0001, Dinh Q. Phung, Svetha Venkatesh, Helen Christensen
WISE (2)2
2016 Fast Kernel-based method for anomaly detection
abstract
Anomaly detection (AD) involves detecting abnormality from normality and has a wide spectrum of applications in reality. Kernel-based methods for AD have been proven robust with diverse data distributions and offering good generalization ability. Stochastic gradient descent (SGD) method has recently emerged as a promising framework to devise ultra-fast learning methods. In this paper, we conjoin the advantages of Kernel-based method and SGD-based method to propose fast learning methods for anomaly detection. We validate the proposed methods on 8 benchmark datasets in UCI repository and KDD cup 1999 dataset. The experimental results show that the proposed methods offer a comparable one-class classification accuracy while simultaneously achieving a significantly computational speed-up.
Trung Le 0001, Van Nguyen 0002, Dat Tran 0001
IJCNN4
2015 Graph-based semi-supervised Support Vector Data Description for novelty detection
abstract
Support Vector Data Description (SVDD) is a well-known supervised learning method for novelty detection purpose. For its classification task, SVDD requires a fully-labeled dataset. Nonetheless, contemporary datasets always consist of a collection of labeled data samples jointly a much larger collection of unlabeled ones. This fact impedes the usage of SVDD in the real-world problems. In this paper, we propose to utilize the information implicated in a spectral graph to leverage SVDD in the context of semi-supervised learning. The theory and experiment evidence that the proposed method is able to efficiently employ the information carried in the spectral graph to not only enhance the generalization ability of SVDD but also enforce the cluster assumption which is crucial for a semi-supervised learning method.
Phuong Duong, Van Nguyen 0002, Mi Dinh, Trung Le 0001, Dat Tran 0001, Wanli Ma 0003
IJCNN2
2014 Kernel-based semi-supervised learning for novelty detection
abstract
One-class Support Vector Machine (OCSVM) is a well-known method for novelty detection. However, OCSVM regards all negative data samples as a common symbol and thereby not being able to utilize the information carried by them. Furthermore, OCSVM requires a fully labeled data set and cannot work efficiently with data set with both labeled and unlabeled data samples which is very popular nowadays. In this paper, we first extend the model of OCSVM to enable efficiently using the negative data samples. We then propose two methods to integrate the semi-supervised learning paradigm to the extended model for novelty detection purpose.
Van Nguyen 0002, Trung Le 0001, Thien Pham, Mi Dinh
IJCNN1
2013 Maximal margin learning vector quantisation
abstract
Kernel Generalised Learning Vector Quantisation (KGLVQ) was proposed to extend Generalised Learning Vector Quantisation into the kernel feature space to deal with complex class boundaries and thus yielded promising performance for complex classification tasks in pattern recognition. However KGLVQ does not follow the maximal margin principle, which is crucial for kernel-based learning methods. In this paper we propose a maximal margin approach (MLVQ) to the KGLVQ algorithm. MLVQ inherits the merits of KGLVQ and also follows the maximal margin principle to improve the generalisation capability. Experiments performed on the well-known data sets available in UCI repository show promising classification results for the proposed method.
Trung Le 0001, Dat Tran 0001, Van Nguyen 0002, Wanli Ma 0003
IJCNN3