Guanjun Lin

dblp:202/6884 · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
9since 2021 · last 2025
0000-0003-3280-1307ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 8 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Computer networks · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author
YearPublicationVenuePosition
2025 Enhancing vulnerability detection efficiency: An exploration of light-weight LLMs with hybrid code features
Guanjun Lin, Huan Mei, Yonghang Tai
J. Inf. Secur. Appl.2
2024 Software vulnerable functions discovery based on code composite feature
Guanjun Lin, Huan Mei, Yonghang Tai, Jun Zhang 0065
J. Inf. Secur. Appl.2
2024 Multi-theme hierarchical monitoring method for wireless sensor networks
Chuiju You, Guanjun Lin, Lili Sun, Shaoyu Zhao
Wirel. Networks2
2023 Detecting vulnerabilities in IoT software: New hybrid model and comprehensive data analysis
Huan Mei, Guanjun Lin, Da Fang, Jun Zhang 0065
J. Inf. Secur. Appl.2
2023 The application of neural network for software vulnerability detection: a review
Yuhui Zhu, Guanjun Lin, Jun Zhang 0065
Neural Comput. Appl.2
2022 Deep Neural Embedding for Software Vulnerability Discovery: Comparison and Optimization
abstract
Due to multitudinous vulnerabilities in sophisticated software programs, the detection performance of existing approaches requires further improvement. Multiple vulnerability detection approaches have been proposed to aid code inspection. Among them, there is a line of approaches that apply deep learning (DL) techniques and achieve promising results. This paper attempts to utilize CodeBERT which is a deep contextualized model as an embedding solution to facilitate the detection of vulnerabilities in C open-source projects. The application of CodeBERT for code analysis allows the rich and latent patterns within software code to be revealed, having the potential to facilitate various downstream tasks such as the detection of software vulnerability. CodeBERT inherits the architecture of BERT, providing a stacked encoder of transformer in a bidirectional structure. This facilitates the learning of vulnerable code patterns which requires long-range dependency analysis. Additionally, the multihead attention mechanism of transformer enables multiple key variables of a data flow to be focused, which is crucial for analyzing and tracing potentially vulnerable data flaws, eventually, resulting in optimized detection performance. To evaluate the effectiveness of the proposed CodeBERT-based embedding solution, four mainstream-embedding methods are compared for generating software code embeddings, including Word2Vec, GloVe, and FastText. Experimental results show that CodeBERT-based embedding outperforms other embedding models on the downstream vulnerability detection tasks. To further boost performance, we proposed to include synthetic vulnerable functions and perform synthetic and real-world data fine tuning to facilitate the model learning of C-related vulnerable code patterns. Meanwhile, we explored the suitable configuration of CodeBERT. The evaluation results show that the model with new parameters outperform some state-of-the-art detection methods in our dataset.
Guanjun Lin, Yonghang Tai, Jun Zhang 0065
Secur. Commun. Networks2
2022 CD-VulD: Cross-Domain Vulnerability Discovery Based on Deep Domain Adaptation
abstract
A major cause of security incidents such as cyber attacks is rooted in software vulnerabilities. These vulnerabilities should ideally be found and fixed before the code gets deployed. Machine learning-based approaches achieve state-of-the-art performance in capturing vulnerabilities. These methods are predominantly supervised. Their prediction models are trained on a set of ground truth data where the training data and test data are assumed to be drawn from the same probability distribution. However, in practice, the test data often differs from the training data in terms of distribution because they are from different projects or they differ in the types of vulnerability. In this article, we present a new system forCrossDomain SoftwareVulnerabilityDiscovery (CD-VulD) using deep learning (DL) and domain adaptation (DA). We employ DL because it has the capacity of automatically constructing high-level abstract feature representations of programs, which are likely of more cross-domain useful than the handcrafted features driven by domain knowledge. The divergence between distributions is reduced by learning cross-domain representations. First, given software program representations, CD-VulD converts them into token sequences and learns the token embeddings for generalization across tokens. Next, CD-VulD employs a deep feature model to build abstract high-level presentations based on those sequences. Then, the metric transfer learning framework (MTLF) technique is employed to learn cross-domain representations by minimizing the distribution divergence between the source domain and the target domain. Finally, the cross-domain representations are used to build a classifier for vulnerability detection. Experimental results show that CD-VulD outperforms the state-of-the-art vulnerability detection approaches by a wide margin. We make the new datasets publicly available so that our work is replicable and can be further improved.
Shigang Liu, Guanjun Lin, Lizhen Qu, Jun Zhang 0010, Olivier Y. de Vel, Paul Montague, Yang Xiang 0001
IEEE Trans. Dependable Secur. Comput.2
2021 Deep neural-based vulnerability discovery demystified: data, model and performance
Guanjun Lin, Leo Yu Zhang, Shang Gao 0003, Yonghang Tai, Jun Zhang 0010
Neural Comput. Appl.1
2021 Software Vulnerability Discovery via Learning Multi-Domain Knowledge Bases
abstract
Machine learning (ML) has great potential in automated code vulnerability discovery. However, automated discovery application driven by off-the-shelf machine learning tools often performs poorly due to the shortage of high-quality training data. The scarceness of vulnerability data is almost always a problem for any developing software project during its early stages, which is referred to as the cold-start problem. This article proposes a framework that utilizes transferable knowledge from pre-existing data sources. In order to improve the detection performance, multiple vulnerability-relevant data sources were selected to form a broader base for learning transferable knowledge. The selected vulnerability-relevant data sources are cross-domain, including historical vulnerability data from different software projects and data from the Software Assurance Reference Database (SARD) consisting of synthetic vulnerability examples and proof-of-concept test cases. To extract the information applicable in vulnerability detection from the cross-domain data sets, we designed a deep-learning-based framework with Long-short Term Memory (LSTM) cells. Our framework combines the heterogeneous data sources to learn unified representations of the patterns of the vulnerable source codes. Empirical studies showed that the unified representations generated by the proposed deep learning networks are feasible and effective, and are transferable for real-world vulnerability detection. Our experiments demonstrated that by leveraging two heterogeneous data sources, the performance of our vulnerability detection outperformed the static vulnerability discovery toolFlawfinder. The findings of this article may stimulate further research in ML-based vulnerability detection using heterogeneous data sources.
Guanjun Lin, Jun Zhang 0010, Wei Luo 0001, Lei Pan 0002, Olivier Y. de Vel, Paul Montague, Yang Xiang 0001
IEEE Trans. Dependable Secur. Comput.1
2020 Software Vulnerability Detection Using Deep Neural Networks: A Survey
abstract
The constantly increasing number of disclosed security vulnerabilities have become an important concern in the software industry and in the field of cybersecurity, suggesting that the current approaches for vulnerability detection demand further improvement. The booming of the open-source software community has made vast amounts of software code available, which allows machine learning and data mining techniques to exploit abundant patterns within software code. Particularly, the recent breakthrough application of deep learning to speech recognition and machine translation has demonstrated the great potential of neural models’ capability of understanding natural languages. This has motivated researchers in the software engineering and cybersecurity communities to apply deep learning for learning and understanding vulnerable code patterns and semantics indicative of the characteristics of vulnerable code. In this survey, we review the current literature adopting deep-learning-/neural-network-based approaches for detecting software vulnerabilities, aiming at investigating how the state-of-the-art research leverages neural techniques for learning and understanding code semantics to facilitate vulnerability discovery. We also identify the challenges in this new field and share our views of potential research directions.
Guanjun Lin, Sheng Wen, Qing-Long Han, Jun Zhang 0010, Yang Xiang 0001
Proc. IEEE1
2020 DeepBalance: Deep-Learning and Fuzzy Oversampling for Vulnerability Detection
abstract
Software vulnerability has long been an important but critical research issue in cybersecurity. Recently, the machine learning (ML)-based approach has attracted increasing interest in the research of software vulnerability detection. However, the detection performance of existing ML-based methods require further improvement. There are two challenges: one is code representation for ML and the other is class imbalance between vulnerable code and nonvulnerable code. To overcome these challenges, this article develops a DeepBalance system, which combines the new ideas of deep code representation learning and fuzzy-based class rebalancing. We design a deep neural network with bidirectional long short-term memory to learn invariant and discriminative code representations from labeled vulnerable and nonvulnerable code. Then, a new fuzzy oversampling method is employed to rebalance the training data by generating synthetic samples for the class of vulnerable code. To evaluate the performance of the new system, we carry out a series of experiments in a real-world ground-truth dataset that consists of the code from the projects of LibTIFF, LibPNG, and FFmpeg. The results show that the proposed new system can significantly improve the vulnerability detection performance. For example, the improvement is 15% in terms of F-measure.
Shigang Liu, Guanjun Lin, Qing-Long Han, Sheng Wen, Jun Zhang 0010, Yang Xiang 0001
IEEE Trans. Fuzzy Syst.2
2020 Neural Model Stealing Attack to Smart Mobile Device on Intelligent Medical Platform
abstract
To date, the Medical Internet of Things (MIoT) technology has been recognized and widely applied due to its convenience and practicality. The MIoT enables the application of machine learning to predict diseases of various kinds automatically and accurately, assisting and facilitating effective and efficient medical treatment. However, the MIoT are vulnerable to cyberattacks which have been constantly advancing. In this paper, we establish a MIoT platform and demonstrate a scenario where a trained Convolutional Neural Network (CNN) model for predicting lung cancer complicated with pulmonary embolism can be attacked. First, we use CNN to build a model to predict lung cancer complicated with pulmonary embolism and obtain high detection accuracy. Then, we build a copycat model using only a small amount of data labeled by the target network, aiming to steal the established prediction model. Experimental results prove that the stolen model can also achieve a relatively high prediction outcome, revealing that the copycat network could successfully copy the prediction performance from the target network to a large extent. This also shows that such a prediction model deployed on MIoT devices can be stolen by attackers, and effective prevention strategies are open questions for researchers.
Liqiang Zhang 0009, Guanjun Lin, Bixuan Gao, Zhibao Qin, Yonghang Tai, Jun Zhang 0065
Wirel. Commun. Mob. Comput.2
2019 Deep Learning-Based Vulnerable Function Detection: A Benchmark
Guanjun Lin, Jun Zhang 0010, Yang Xiang 0001
ICICS1
2018 Cross-Project Transfer Representation Learning for Vulnerable Function Discovery
abstract
Machine learning is now widely used to detect security vulnerabilities in the software, even before the software is released. But its potential is often severely compromised at the early stage of a software project when we face a shortage of high-quality training data and have to rely on overly generic hand-crafted features. This paper addresses this cold-start problem of machine learning, by learning rich features that generalize across similar projects. To reach an optimal balance between feature-richness and generalizability, we devise a data-driven method including the following innovative ideas. First, the code semantics are revealed through serialized abstract syntax trees (ASTs), with tokens encoded by Continuous Bag-of-Words neural embeddings. Next, the serialized ASTs are fed to a sequential deep learning classifier (Bi-LSTM) to obtain a representation indicative of software vulnerability. Finally, the neural representation obtained from existing software projects is then transferred to the new project to enable early vulnerability detection even with a small set of training labels. To validate this vulnerability detection approach, we manually labeled 457 vulnerable functions and collected 30 000+ nonvulnerable functions from six open-source projects. The empirical results confirmed that the trained model is capable of generating representations that are indicative of program vulnerability and is adaptable across multiple projects. Compared with the traditional code metrics, our transfer-learned representations are more effective for predicting vulnerable functions, both within a project and across multiple projects.
Guanjun Lin, Jun Zhang 0010, Wei Luo 0001, Lei Pan 0002, Yang Xiang 0001, Olivier Y. de Vel, Paul Montague
IEEE Trans. Ind. Informatics1
2017 POSTER: Vulnerability Discovery with Function Representation Learning from Unlabeled Projects
abstract
In cybersecurity, vulnerability discovery in source code is a fundamental problem. To automate vulnerability discovery, Machine learning (ML) based techniques has attracted tremendous attention. However, existing ML-based techniques focus on the component or file level detection, and thus considerable human effort is still required to pinpoint the vulnerable code fragments. Using source code files also limit the generalisability of the ML models across projects. To address such challenges, this paper targets at the function-level vulnerability discovery in the cross-project scenario. A function representation learning method is proposed to obtain the high-level and generalizable function representations from the abstract syntax tree (AST). First, the serialized ASTs are used to learn project independence features. Then, a customized bi-directional LSTM neural network is devised to learn the sequential AST representations from the large number of raw features. The new function-level representation demonstrated promising performance gain, using a unique dataset where we manually labeled 6000+ functions from three open-source projects. The results confirm that the huge potential of the new AST-based function representation learning.
Guanjun Lin, Jun Zhang 0010, Wei Luo 0001, Lei Pan 0002, Yang Xiang 0001
CCS1