Tongshuai Wu

dblp:284/9399 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
11since 2021 · last 2024
0000-0003-3839-1660ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Network and information security
1 paper
Blockchain and cryptocurrency security · 50% Systems and software security · 50%
Software engineering, system software, and programming languages
1 paper
Program analysis · 77% Software testing · 23%

Topics — the 3 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Systems and software security › vulnerability discovery › machine-learning-based vulnerability detection
deep learning-based vulnerability detection
0.812024
UltraVCS: Ultra-Fine-Grained Variable-Based Code Slicing for Automated Vulnerability Detection · IEEE Trans. Inf. Forensics Secur. 2024
Blockchain and cryptocurrency security › smart contract security
vulnerability detection
0.812024
UltraVCS: Ultra-Fine-Grained Variable-Based Code Slicing for Automated Vulnerability Detection · IEEE Trans. Inf. Forensics Secur. 2024
Program analysis › static analysis
program slicing
0.812024
UltraVCS: Ultra-Fine-Grained Variable-Based Code Slicing for Automated Vulnerability Detection · IEEE Trans. Inf. Forensics Secur. 2024

Methods — techniques the papers use, named apart from their topics

deep learning · 1.5code slicing · 1.5
YearPublicationVenuePosition
2024 CPMSVD: Cross-Project Multiclass Software Vulnerability Detection Via Fused Deep Feature and Domain Adaptation
abstract
Many deep learning-based approaches have achieved excellent performance for Software Vulnerability Detection(SVD) but the most imperative issue is coping with the scarcity of labeled software vulnerabilities. When employing transfer learning techniques, researchers only detected the presence of vulnerabilities but cannot identify vulnerability types. In this paper, we propose the first system for Cross-Project Multiclass Software Vulnerability Detection (CPMSVD) which incorporates inter-procedure code lines as local feature and detects at the granularity of code snippet. Principles are defined to generate snippet attentions and a deep model is proposed to obtain the fusion representations. We then extend domain adaptation techniques to reduce feature distributions among different projects. Experimental results show that our approach outperforms other state-of-the-art ones.
Gewangzi Du, Tongshuai Wu
ICASSP3
2024 Code Property Graph based Cross-Domain Vulnerability Detection via Deep Fused Feature
abstract
Deep learning is becoming an important means to detect source code vulnerabilities. However, the most severe problem is the compromise of detection performance when there is a scarcity of labeled data. Researchers employed transfer learning skills to solve the problem but existing approaches utilised limited information, which failed to contain various of vulnerability patterns. The code property graph (CPG), which encapsulates abundant syntax and semantic information, is able to accommodate more vulnerability patterns. In this paper, we propose the first CPG-based cross-domain vulnerability detection system which includes an approach to represent the CPG of code snippet into vector. A deep fusion model is devised to generate the fused deep features; Moreover, we extend a metric learning algorithm to reduce data distributions from different domains. Experimental results prove our system is much more effective compared with other state-of-the-art cross-domain approaches.
Gewangzi Du, Tongshuai Wu, Xiong Zheng
ISCAS3
2024 CDNM: Clustering-Based Data Normalization Method For Automated Vulnerability Detection
abstract
Abstract The key to deep learning vulnerability detection framework is pre-processing source code and learning vulnerability features. Traditional source code representation techniques take a complete normalization to user-defined symbols but ignore the semantic information associated with vulnerabilities. The current mainstream vulnerability feature learning model is Recurrent Neural Network (RNN), whose time-series structure determines its insufficient remote information acquisition capability. This paper proposes a new vulnerability detection framework to solve the above problems. We propose a new data normalization method in the source code pre-processing phase. The user-defined symbols are clustered using the unsupervised clustering algorithm K-means. The normalized classification is performed according to the clustering results, which preserves the primary semantic information in the source code and ensures the smoothness of the sample data. In the feature extraction stage, we input the source code after performing text representation into Bidirectional Encoder Representations for Transformers (BERT) for feature automation learning, which enhances semantic information extraction and remote information acquisition. Experimental results show that the vulnerability detection precision of this method is 18.3% higher than that of the current mainstream vulnerability detection framework in the real-world data collected by ourselves. Further, our method improves the precision of the state-of-the-art method by 4.2%.
Tongshuai Wu, Gewangzi Du, Ningning Cui
Comput. J.1
2024 Vulnerability Localization Based On Intermediate Code Representation and Feature Fusion
abstract
Abstract Vulnerability localization can assist security professionals in vulnerability validation and analysis. This study proposes an intelligent vulnerability localization method based on fine-grained program representation and feature fusion. Firstly, we generate efficient fine-grained program representations of the program. This involves transforming the source code into intermediate code. We use abstract syntax tree characteristics to correspond to the points of interest of the intermediate code. We slice the intermediate code file based on the point of interest and program dependency relationships. Subsequently, we use the word2vec model to the vectorization of the intermediate code slices. Then, we propose a vulnerability localization framework based on a feature fusion method, which can better combine the advantages of bidirectional gate recurrent unit and convolutional neural network to capture the syntax and semantics of program representation. Through comparing different program representations, we have discovered that the fine-grained representation based on intermediate code in this study provides a more accurate portrayal of program semantics. By comparing various methods, the proposed feature fusion approach in this paper improves vulnerability localization. We also conducted a visualization display of vulnerability localization. Furthermore, we have validated the effectiveness of this method in localizing vulnerabilities across five common vulnerability types.
Renzheng Wei, Tongshuai Wu, Gewangzi Du
Comput. J.4
2024 UltraVCS: Ultra-Fine-Grained Variable-Based Code Slicing for Automated Vulnerability Detection
abstract
Detecting vulnerabilities in source code using deep learning models is emerging as a valuable research area. The key issue in using deep learning to detect vulnerabilities is the accurate representation. Current approaches for detecting vulnerabilities in C/C++ programs use functions or lines of code as the unit and only consider the basic syntactic structure of vulnerabilities. Unfortunately, functions and lines of code still have vulnerability-unrelated information, which is redundant for vulnerability features and is not conducive to deep learning models to learn accurate vulnerability patterns. This paper deeply analyzes the essential features of vulnerabilities and attacks. Then, we propose a novel variable-based deep learning vulnerability detection method for C/C++ that is more granular than existing function- or line of code-based vulnerability detection methods. Based on the triggering mechanism of vulnerabilities and typical memory attacks, we propose the concepts of key variables and insecure operations; these are used to propose new rules for determining the center point of code slices with more accurate vulnerability features. We propose the first ultra-fine-grained variable-based code slicing (UltraVCS) method by the new center point, which focuses on the vulnerability-related variable. This method removes as much vulnerability-unrelated information as possible to achieve more accurate vulnerability feature extraction. Experiments show that our approach can generate more code slices, achieve more precise vulnerability representation, and perform better vulnerability detection in open-source projects compared to state-of-the-art methods. Furthermore, we have discovered four zero-day vulnerabilities in real-world application scenarios in open-source projects.
Tongshuai Wu, Gewangzi Du, Dan Meng 0002
IEEE Trans. Inf. Forensics Secur.1
2023 Cross Domain on Snippets: BiLSTM-TextCNN based Vulnerability Detection with Domain Adaptation
abstract
Due to the ubiquity of computer software, software vulnerability detection(SVD) problem is essential to protect cyber system from attacks. Recently, deep learning-based vulnerability detection has achieved outstanding performance, relieving experts from tedious task of manually defining vulnerability features as well. However, its detection capability is compromised when facing with the scarcity of labeled data. One possible solution is to leverage training data with adequate labels from other domains, but the data distributions in different domains differ significantly. On the other hand, function level detection is too coarse-grained and not able to capture inter-procedure vulnerability patterns. In this paper, we propose a systematic Snippet-Oriented Cross-Domain Vulnerability Detection Framework with Domain Adaptation, which is the first time to detect cross-project vulnerabilities at a finer granularity than function. Firstly, we generate Code Snippets from 5 real-world projects and 3 types of CWE in NVD and SARD for cross-project and cross-type detection; Secondly, we propose an novel and effective approach to obtain deep features for domain adaptation; Finally, we employ the domain adaptation algorithm on these deep features to reduce the divergence between different domains and get the final result. Experimental results show that our framework outperforms other state-of-the-art approaches.
Gewangzi Du, Tongshuai Wu, Xiong Zheng, Ningning Cui
CSCWD3
2023 Code Property Graph based Vulnerability Type Identification with Fusion Representation
abstract
Deep learning-based vulnerability detection methods have become one of the mainstream methods of vulnerability detection. The vulnerability type information is of great value in helping vulnerability location and vulnerability remediation. This paper proposes a framework for Vulnerability Type Identification based on Code Property Graph with Fusion Representation. First, this paper uses code property graph information. Code property graph(CPG) is a joint data structure that combines Abstract Syntax Trees(AST), Control Flow Graphs (CFG), and Program Dependency Graphs (PDG). We encode CPG information. Secondly, we use Convolutional neural network combined with Recurrent Neural Network(CNN-RNN) and Attention-Based Bidirectional Gate Recurrent Unit (Att-BiGRU) to extract AST and CFG combined with PDG information. We fuse the extracted features to obtain an effective representation. And then, we perform multi-classification to derive the predicted value of the vulnerability type. Finally, we use 59 vulnerabilities with third-level CWE-ID for evaluation. The experiments show that this paper’s code property graph information can better represent the type information of vulnerabilities. Compared with the classical RNNs, our model in this paper has a more accurate identification effect.
Tongshuai Wu, Ningning Cui, Xiong Zheng
CSCWD2
2023 Identify Vulnerability Types: A Cross-Project Multiclass Vulnerability Classification System Based on Deep Domain Adaptation
Gewangzi Du, Tongshuai Wu
ICONIP (6)3
2022 Inductive Vulnerability Detection via Gated Graph Neural Network
abstract
Vulnerability detection is an essential means to ensure the normal operation of various software tools and system security. The Recurrent Neural Networks (RNNs) have achieved remarkable results in vulnerability detection, but the sequence-based code representation has great limitations in feature expression and propagation. In this paper, we propose a fine-grained code vulnerability detection framework based on Gated Graph Neural Network (GGNN). Firstly, we process the source code into fine-grained slices. Secondly, graph embedding of code slices is constructed by clustering neighborhood information. Finally, GGNN is used to learn the syntax and semantic information of vulnerability codes for graph-level classification. Furthermore, we theoretically analyze that GGNN has a strong inductive learning ability. This means that the model requires only a small amount of training data to obtain sufficient advanced features, which is significant for vulnerability detection tasks that are difficult to collect data sets. We carry out conventional experiments and inductive experiments with manually collected data sets, and the results show that the framework is superior to RNNs in vulnerability detection performance. Moreover, our framework performs better than RNNs under inductive conditions.
Tongshuai Wu, Gewangzi Du, Ningning Cui
CSCWD1
2022 Automated Vulnerable Codes Mutation through Deep Learning for Variability Detection
abstract
At present, there are many studies on the automatic generation or mutation of common code dataset, but the research on the mass generation of vulnerable code dataset has received little public attention. Wanting to generate more vulnerable code seems counterproductive, but it is very important to vulnerability detection technology. It can be used to discover the blind spots of vulnerability detection tools through fuzzing testing technology. In particular, the use of machine learning and deep learning techniques for vulnerability detection has highly dataset imbalance problem due to the lack of vulnerable codes, which seriously affects the performance of the vulnerability detection model. In this paper, we propose a new vulnerable code mutation technique called Vuls-Mutation. Based on a generative Sequence-to-Sequence model, our system automatically and continuously mutate to generate new vulnerable code, by changing the control flow or data flow of the potentially tainted data in the existing vulnerable code. Experiments show that the grammatical correctness rate of the mutated code is about 71 % and the true positive rate of the mutated code based on the correct grammar is about 93%. We add this new set of mutation programs to train deep learning vulnerability detection models, and the results show that all indicators are better than the baseline method.
Gewangzi Du, Tongshuai Wu
IJCNN3
2022 BERT-Based Vulnerability Type Identification with Effective Program Representation
Gewangzi Du, Tongshuai Wu, Ningning Cui
WASA (1)3