VLDB 2026 Research / reviewers in the wild / expert
Xiaoxue Wu 0001
dblp:159/4418-1
· DBLP profile ↗
43ranked-venue papers
8as first author
37since 2021 · last 2026
0000-0002-7567-3643ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 33 · 7 first-author · 27 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Security and privacy · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PIONEER: improving the robustness of student models when compressing pre-trained models of code
Xiangyue Liu 0002, Lili Bo, Xiaoxue Wu 0001, Yun Yang 0003, Xiaobing Sun 0001 |
Autom. Softw. Eng. | 4 |
| 2026 | AdvGen-X: Transferability driven adversarial example generation for pre-trained models of code
Xiangyue Liu 0002, Xiaobing Sun 0001, Lili Bo, Bin Li 0006, Xiaoxue Wu 0001, Sicong Cao, Yufei Hu |
Empir. Softw. Eng. | 6 |
| 2025 | Enhancing concurrency vulnerability detection through AST-based static fuzz mutation
Wei Zheng 0006, Peiran Deng, Xiang Chen 0005, Xiaoxue Wu 0001 |
J. Syst. Softw. | 5 |
| 2025 | From Data to Knowledge: Mining Linux Vulnerability Characteristics and Evolution With Knowledge GraphsabstractABSTRACT An operating system is the essence of software, serving as the foundation for the operation of various application software. The security of the operating system is crucial for national informatization construction. Data indicate that many cybersecurity incidents result from exploiting security vulnerabilities in the operating system. Linux is currently the most widely used open‐source operating system, with thousands of Common Vulnerabilities and Exposures (CVEs) related to Linux systems reported each year. Therefore, research and prevention of vulnerabilities in the Linux system are particularly important. To gain a better understanding of the characteristics of Linux system vulnerabilities, this paper leverages knowledge in the field of software security to analyze nearly 10,000 historical vulnerability data in two core systems of Linux: Linux Kernel and Debian Linux. The study explores the evolutionary patterns of vulnerability characteristics. Specific research contents include the following: (1) data collection and cleaning of vulnerability data in Linux Kernel and Debian Linux systems; (2) cross‐statistical analysis of structured data features in vulnerability reports; (3) unstructured data characteristics mining in vulnerability reports based on domain knowledge; (4) analysis of the evolution of vulnerability characteristics. This paper provides empirical lessons and guidance for Linux system vulnerabilities to assist practitioners and researchers in better preventing and detecting vulnerabilities in Linux and Linux‐based systems. Shiyu Weng, Xiaoxue Wu 0001, Wenjing Shan, Xiaobing Sun 0001 |
J. Softw. Evol. Process. | 2 |
| 2025 | NG_MDERANK: A software vulnerability feature knowledge extraction method based on N-gram similarityabstractAbstract As software grows in size and complexity, software vulnerabilities are increasing, leading to a range of serious insecurity issues. Open‐source software vulnerability reports and documentation can provide researchers with great convenience for analysis and detection. However, the quality of different data sources varies, the data are duplicated and lack of correlation, which often requires a lot of manual management and analysis. In order to solve the problems of scattered and heterogeneous data and lack of correlation in traditional vulnerability repositories, this paper proposes a software vulnerability feature knowledge extraction method that combines the N‐gram model and mask similarity. The method generates mask text data based on the extraction of N‐gram candidate keywords and extracts vulnerability feature knowledge by calculating the similarity of mask text. This method analyzes the samples efficiently and stably in the environment of large sample size and complex samples and can obtain high‐value semi‐structured data. Then, the final node, relationship, and attribute information are obtained by secondary knowledge cleaning and extraction of the extracted semi‐structured data results. And based on the extraction results, the corresponding software vulnerability domain knowledge graph is constructed to deeply explore the semantic information features and entity relationships of vulnerabilities, which can help to efficiently study software security problems and solve vulnerability problems. The effectiveness and superiority of the proposed method is verified by comparing it with several traditional keyword extraction algorithms on Common Weakness Enumeration (CWE) and Common Vulnerabilities and Exposures (CVE) vulnerability data. Xiaoxue Wu 0001, Shiyu Weng, Wei Zheng 0006, Xiang Chen 0005, Xiaobing Sun 0001 |
J. Softw. Evol. Process. | 1 |
| 2025 | HgtJIT: Just-in-Time Vulnerability Detection Based on Heterogeneous Graph TransformerabstractVulnerability detection plays a crucial role in the software development lifecycle. Commit-level vulnerability detection aims to detect whether the changed code contributed to potential vulnerabilities by the developer when submitting the code, which is also referred to as Just-In-Time (JIT) vulnerability detection. Previous JIT vulnerability detection approaches relied on code metrics and textual features, which were unable to effectively characterize vulnerability-contributing commits (VCCs). Recently, CodeJIT (a code-centric learning-based approach) has been proposed to detect vulnerability at the commit-level. However, CodeJIT still has its limitations: imprecise feature representation, static code embedding, and underutilized heterogeneous information. In this paper, we propose HgtJIT, a JIT vulnerability detection approach based on a Heterogeneous Graph Transformer (HGT) in order to address several limitations of the state-of-the-art CodeJIT approach. We propose diffPDG to represent code changes and use the CCT5 model (the latest feature encoder pre-trained on a large-scale code change corpus) to embed graph nodes to generate the most meaningful vector representations. In addition, we employ HGT to adequately utilize heterogeneous information of the graph to learn vulnerability features. Extensive experiments have shown that HgtJIT is the best-performing model, with F1 and AUC improvement of 14.6%-37.5% and 12.2%-53.7% compared to the baseline model Xiaobing Sun 0001, Mingxuan Zhou, Sicong Cao, Xiaoxue Wu 0001, Lili Bo, Di Wu 0050, Bin Li 0006, Yang Xiang 0001 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2024 | Coca: Improving and Explaining Graph Neural Network-Based Vulnerability Detection SystemsabstractRecently, Graph Neural Network (GNN)-based vulnerability detection systems have achieved remarkable success. However, the lack of explainability poses a critical challenge to deploy black-box models in security-related domains. For this reason, several approaches have been proposed to explain the decision logic of the detection model by providing a set of crucial statements positively contributing to its predictions. Unfortunately, due to the weakly-robust detection models and suboptimal explanation strategy, they have the danger of revealing spurious correlations and redundancy issue. Sicong Cao, Xiaobing Sun 0001, Xiaoxue Wu 0001, David Lo 0001, Lili Bo, Bin Li 0006, Wei Liu 0010 |
ICSE | 3 |
| 2024 | 1+1>2: Integrating Deep Code Behaviors with Metadata Features for Malicious PyPI Package DetectionabstractPyPI, the official package registry for Python, has seen a surge in the number of malicious package uploads in recent years. Prior studies have demonstrated the effectiveness of learning-based solutions in malicious package detection. However, manually-crafted expert rules are expensive and struggle to keep pace with the rapidly evolving malicious behaviors, while deep features automatically extracted from code are still inaccurate in certain cases. To mitigate these issues, in this paper, we propose Ea4mp, a novel approach which integrates deep code behaviors with metadata features to detect malicious PyPI packages. Specifically, Ea4mp extracts code behavior sequences from all script files and fine-tunes a BERT model to learn deep semantic features of malicious code. In addition, we realize the value of metadata information and construct an ensemble classifier to combine the strengths of deep code behavior features and metadata features for more effective detection. We evaluated Ea4mp against three state-of-the-art baselines on a newly constructed dataset. The experimental results show that Ea4mp improves precision by 6.9%-24.6% and recall by 10.5%-18.4%. With Ea4mp, we successfully identified 119 previously unknown malicious packages from a pool of 46,573 newly-uploaded packages over a three-week period, and 82 out of them have been removed by the PyPI official. Xiaobing Sun 0001, Xingan Gao, Sicong Cao, Lili Bo, Xiaoxue Wu 0001, Kaifeng Huang 0001 |
ASE | 5 |
| 2024 | ChatBR: Automated assessment and improvement of bug report quality using ChatGPTabstractBug reports, containing crucial information such as the Observed Behavior (OB), the Expected Behavior (EB), and the Steps to Reproduce (S2R), can help developers localize and fix bugs efficiently. However, due to the increasing complexity of some bugs and the limited experience of some reporters, large numbers of bug reports miss this crucial information. Although machine learning (ML)-based and information retrieval (IR)-based approaches are proposed to detect and supplement the missing information in bug reports, the performance of these approaches depends heavily on the size and quality of bug report datasets. Lili Bo, Wangjie Ji, Xiaobing Sun 0001, Ting Zhang 0011, Xiaoxue Wu 0001, Ying Wei 0012 |
ASE | 5 |
| 2024 | Snopy: Bridging Sample Denoising with Causal Graph Learning for Effective Vulnerability DetectionabstractDeep Learning (DL) has emerged as a promising means for vulnerability detection due to its ability to automatically derive features from vulnerable code. Unfortunately, current solutions struggle to focus on vulnerability-related parts of vulnerable functions, and tend to exploit spurious correlations for prediction, thus undermining their effectiveness in practice. In this paper, we propose Snopy, a novel DL-based approach, which bridges sample denoising with causal graph learning to capture real vulnerability patterns from vulnerable samples with numerous noise for effective detection. Specifically, Snopy adopts a change-based sample denoising approach to automatically weed out vulnerability-irrelevant code elements in the vulnerable functions without sacrificing the label accuracy. Then, Snopy constructs a novel Causality-Aware Graph Attention Network (CA-GAT) with Feature Caching Scheme (FCS) to learn causal vulnerability features while maintaining efficiency. Experiments on the three public benchmark datasets show that Snopy outperforms the state-of-the-art baselines by an average of 27.22%, 85.89%, and 75.50% in terms of F1-score, respectively. Sicong Cao, Xiaobing Sun 0001, Xiaoxue Wu 0001, David Lo 0001, Lili Bo, Bin Li 0006, Xiaolei Liu 0001, Xingwei Lin, Wei Liu 0010 |
ASE | 3 |
| 2024 | Duplicate Bug Report detection using Named Entity Recognition
Wei Zheng 0006, Xiaoxue Wu 0001, Jingyuan Cheng |
Knowl. Based Syst. | 3 |
| 2024 | ISTA+: Test case generation and optimization for intelligent systems based on coverage analysis
Xiaoxue Wu 0001, Yizeng Gu, Lidan Lin, Wei Zheng 0006, Xiang Chen 0005 |
Sci. Comput. Program. | 1 |
| 2024 | Software bug localization based on optimized and ensembled deep learning modelsabstractAbstract An automated task for finding the essential buggy files among software projects with the help of a given bug report is termed bug localization. The conventional approaches suffer from the challenges of performing lexical matching. Particularly, the terms utilized for describing the bugs in the bug reports are observed to be irrelevant to the terms used in the source code files. To resolve these problems, we propose an optimized and ensemble deep learning model for software bug localization. These features are reduced by the principle component analysis (PCA). Then, they are selected by the weighted convolutional neural network (CNN) model with the support of the Modified Scatter Probability‐based Coyote Optimization Algorithm (MSP‐COA). Finally, the optimal features are subjected to the ensemble deep neural network and long short‐term memory (DNN‐LSTM), with parameter tuning by the MSP‐COA. Experimental results show that the proposed approach can achieve higher bug localization accuracy than individual models. Lili Bo, Xiaobing Sun 0001, Xiaoxue Wu 0001, Aakash Ali, Ying Wei 0012 |
J. Softw. Evol. Process. | 4 |
| 2024 | Hierarchy-Aware Representation Learning for Industrial IoT Vulnerability ClassificationabstractAs with anything connected to the internet, industrial Internet of Things (IIoT) devices are also subject to severe cybersecurity threats because an adversary could exploit vulnerabilities in their internal software to perform malicious attacks. Despite the promising results of deep learning-based approaches, most solutions can only detect the presence of a vulnerability but fail to pinpoint its corresponding type. Recently, TreeVul formalizes the task as a hierarchical multilabel classification problem to predict complete coarse-to-fine vulnerability type hierarchy. Yet, the TreeVul approach is still inaccurate and neglects samples labeled at coarse categories. In this article, we proposeHierVul, a novel hierarchy-aware representation learning approach for IIoT vulnerability classification. Specifically, to make full use of vulnerable samples labeled at any granularity,HierVulconstructs hierarchy-specific extractors as well as classifiers to disentangle level-wise vulnerability features from the code representation learning network backbone, and maximizes their marginal probability in the probability space constrained by the Common Weakness Enumeration tree hierarchy. Furthermore, considering that the distinction between two vulnerability types at the same level of abstraction becomes smaller and smaller as the refinement of classification granularity,HierVulleverages residual connections to add parent-level coarser-grained features to child-level finer-grained features to transfer hierarchical knowledge across levels. The experimental results show thatHierVulachieves 15.25%, 45.16%, and 14.52% relative improvement over TreeVul on Weight F1, Macro F1, and PF, respectively, indicating the effectiveness ofHierVulin the practical scenario. Sicong Cao, Xiaobing Sun 0001, Xiaoxue Wu 0001, Wei Liu 0010, Bin Li 0006 |
IEEE Trans. Ind. Informatics | 4 |
| 2024 | Learning to Detect Memory-related VulnerabilitiesabstractMemory-related vulnerabilities can result in performance degradation or even program crashes, constituting severe threats to the security of modern software. Despite the promising results of deep learning (DL)-based vulnerability detectors, there exist three main limitations: (1) rich contextual program semantics related to vulnerabilities have not yet been fully modeled; (2) multi-granularity vulnerability features in hierarchical code structure are still hard to be captured; and (3) heterogeneous flow information is not well utilized. To address these limitations, in this article, we propose a novel DL-based approach, called MVD+ , to detect memory-related vulnerabilities at the statement-level. Specifically, it conducts both intraprocedural and interprocedural analysis to model vulnerability features, and adopts a hierarchical representation learning strategy, which performs syntax-aware neural embedding within statements and captures structured context information across statements based on a novel Flow-Sensitive Graph Neural Networks, to learn both syntactic and semantic features of vulnerable code. To demonstrate the performance, we conducted extensive experiments against eight state-of-the-art DL-based approaches as well as five well-known static analyzers on our constructed dataset with 6,879 vulnerabilities in 12 popular C/C++ applications. The experimental results confirmed that MVD+ can significantly outperform current state-of-the-art baselines and make a great trade-off between effectiveness and efficiency. Sicong Cao, Xiaobing Sun 0001, Lili Bo, Rongxin Wu, Bin Li 0006, Xiaoxue Wu 0001, Chuanqi Tao, Tao Zhang 0001, Wei Liu 0010 |
ACM Trans. Softw. Eng. Methodol. | 6 |
| 2024 | An Empirical Study on Correlations Between Deep Neural Network Fairness and Neuron Coverage CriteriaabstractRecently, with the widespread use of deep neural networks (DNNs) in high-stakes decision-making systems (such as fraud detection and prison sentencing), concerns have arisen about the fairness of DNNs in terms of the potential negative impact they may have on individuals and society. Therefore, fairness testing has become an important research topic in DNN testing. At the same time, the neural network coverage criteria (such as criteria based on neuronal activation) is considered as an adequacy test for DNN white-box testing. It is implicitly assumed that improving the coverage can enhance the quality of test suites. Nevertheless, the correlation between DNN fairness (a test property) and coverage criteria (a test method) has not been adequately explored. To address this issue, we conducted a systematic empirical study on seven coverage criteria, six fairness metrics, three fairness testing techniques, and five bias mitigation methods on five DNN models and nine fairness datasets to assess the correlation between coverage criteria and DNN fairness. Our study achieved the following findings: 1) with the increase in the size of the test suite, some of the coverage and fairness metrics changed significantly, as the size of the test suite increased; 2) the statistical correlation between coverage criteria and DNN fairness is limited; and 3) after bias mitigation for improving the fairness of DNN, the change pattern in coverage criteria is different; 4) Models debiased by different bias mitigation methods have a lower correlation between coverage and fairness compared to the original models. Our findings cast doubt on the validity of coverage criteria concerning DNN fairness (i.e., increasing the coverage may even have a negative impact on the fairness of DNNs). Therefore, we warn DNN testers against blindly pursuing higher coverage of coverage criteria at the cost of test properties of DNNs (such as fairness). Wei Zheng 0006, Lidan Lin, Xiaoxue Wu 0001, Xiang Chen 0005 |
IEEE Trans. Software Eng. | 3 |
| 2023 | An Intelligent Duplicate Bug Report Detection Method Based on Technical Term ExtractionabstractAs the bug description data generated during the software maintenance cycle, bug reports are usually hastily written by different users, resulting in many redundant and duplicate bug reports (DBRs). Once the DBRs are repeatedly assigned to developers, it will inevitably lead to a serious waste of human resources, especially for large-scale open-source projects. Recently, many experts and scholars have devoted themselves to researching the detection of DBRs and put forward a series of detection methods for DBRs. However, there is still much room for improvement in the performance of DBR prediction. Therefore, this paper proposes a new method for detecting DBR based on technical term extraction, CTEDB (Combination of Term Extraction and DeBERTaV3) for short. This method first extracts technical terms from the text information of bug reports based on Word2Vec and TextRank algorithms. Then it calculates the semantic similarity of technical terms between different bug reports by combining Word2Vec and SBERT models. Finally, it completes the DBR detection task by combining the DeBERTaV3 model. The experimental results show that CTEDB has achieved good results in detecting DBR, and has obviously improved the accuracy, F1-score, recall and precision compared with the baseline approaches. Xiaoxue Wu 0001, Wenjing Shan, Wei Zheng 0006, Xiaobing Sun 0001 |
AST | 1 |
| 2023 | Improving Java Deserialization Gadget Chain Mining via Overriding-Guided Object GenerationabstractJava (de)serialization is prone to causing security-critical vulnerabilities that attackers can invoke existing methods (gadgets) on the application's classpath to construct a gadget chain to perform malicious behaviors. Several techniques have been proposed to statically identify suspicious gadget chains and dynamically generate injection objects for fuzzing. However, due to their incomplete support for dynamic program features (e.g., Java runtime polymorphism) and ineffective injection object generation for fuzzing, the existing techniques are still far from satisfactory. In this paper, we first performed an empirical study to investigate the characteristics of Java deserialization vulnerabilities based on our manually collected 86 publicly known gadget chains. The empirical results show that 1) Java deserialization gadgets are usually exploited by abusing runtime polymorphism, which enables attackers to reuse serializable overridden methods; and 2) attackers usually invoke exploitable overridden methods (gadgets) via dynamic binding to generate injection objects for gadget chain construction. Based on our empirical findings, we propose a novel gadget chain mining approach, GCMiner, which captures both explicit and implicit method calls to identify more gadget chains, and adopts an overriding-guided object generation approach to generate valid injection objects for fuzzing. The evaluation results show that GCMiner significantly outperforms the state-of-the-art techniques, and discovers 56 unique gadget chains that cannot be identified by the baseline approaches. Sicong Cao, Xiaobing Sun 0001, Xiaoxue Wu 0001, Lili Bo, Bin Li 0006, Rongxin Wu, Wei Liu 0010, Biao He 0002, Yu Ouyang |
ICSE | 3 |
| 2023 | ODDFuzz: Discovering Java Deserialization Vulnerabilities via Structure-Aware Directed Greybox FuzzingabstractJava deserialization vulnerability is a severe threat in practice. Researchers have proposed static analysis solutions to locate candidate vulnerabilities and fuzzing solutions to generate proof-of-concept (PoC) serialized objects to trigger them. However, existing solutions have limited effectiveness and efficiency.In this paper, we propose a novel hybrid solution ODDFuzz to efficiently discover Java deserialization vulnerabilities. First, ODDFuzz performs lightweight static taint analysis to identify candidate gadget chains that may cause deserialization vulnerabilities. In this step, ODDFuzz tries to locate all candidates and avoid false negatives. Then, ODDFuzz performs directed greybox fuzzing (DGF) to explore those candidates and generate PoC testcases to mitigate false positives. Specifically, ODDFuzz applies a structure-aware seed generation method to guarantee the validity of the testcases, and adopts a novel hybrid feedback and a step-forward strategy to guide the directed fuzzing.We implemented a prototype of ODDFuzz and evaluated it on the popular Java deserialization repository ysoserial. Results show that, ODDFuzz could discover 16 out of 34 known gadget chains, while two state-of-the-art baselines only identify three of them. In addition, we evaluated ODDFuzz on real-world applications including Oracle WebLogic Server, Apache Dubbo, Sonatype Nexus, and protostuff, and found six previously unreported exploitable gadget chains with five CVEs assigned. Sicong Cao, Biao He 0002, Xiaobing Sun 0001, Yu Ouyang, Chao Zhang 0008, Xiaoxue Wu 0001, Ting Su 0001, Lili Bo, Bin Li 0006, Chuanlei Ma, Tao Wei 0002 |
SP | 6 |
| 2023 | Automated software bug localization enabled by meta-heuristic-based convolutional neural network and improved deep neural network
Lili Bo, Xiaobing Sun 0001, Xiaoxue Wu 0001, Saifullah Memon, Saima Siraj, Ann Suwaree Ashton |
Expert Syst. Appl. | 4 |
| 2023 | VulLoc: vulnerability localization based on inducing commits and fixing commits
Lili Bo, Xiaobing Sun 0001, Xiaoxue Wu 0001, Bin Li 0006 |
Frontiers Comput. Sci. | 4 |
| 2023 | An Abstract Syntax Tree based static fuzzing mutation for vulnerability evolution analysis
Wei Zheng 0006, Peiran Deng, Kui Gui, Xiaoxue Wu 0001 |
Inf. Softw. Technol. | 4 |
| 2023 | Automatic software vulnerability assessment by extracting vulnerability elements
Xiaobing Sun 0001, Zhenlei Ye, Lili Bo, Xiaoxue Wu 0001, Ying Wei 0012, Tao Zhang 0001, Bin Li 0006 |
J. Syst. Softw. | 4 |
| 2023 | RNNtcs: A test case selection method for Recurrent Neural Networks
Xiaoxue Wu 0001, Jinjin Shen, Wei Zheng 0006, Lidan Lin, Yulei Sui, Abubakar Omari Abdallah Semasaba |
Knowl. Based Syst. | 1 |
| 2023 | An empirical evaluation of deep learning-based source code vulnerability detection: Representation versus modelsabstractAbstract Vulnerabilities in the source code of the software are critical issues in the realm of software engineering. Coping with vulnerabilities in software source code is becoming more challenging due to several aspects such as complexity and volume. Deep learning has gained popularity throughout the years as a means of addressing such issues. This paper proposes an evaluation of vulnerability detection performance on source code representations and evaluates how machine learning (ML) strategies can improve them. The structure of our experiment consists of three deep neural networks (DNNs) in conjunction with five different source code representations: abstract syntax trees (ASTs), code gadgets (CGs), semantics‐based vulnerability candidates (SeVCs), lexed code representations (LCRs), and composite code representations (CCRs). Experimental results show that employing different ML strategies in conjunction with the base model structure influences the performance results to a varying degree. However, ML‐based techniques suffer from poor performance on class imbalance handling and dimensionality reduction when used in conjunction with source code representations. Abubakar Omari Abdallah Semasaba, Wei Zheng 0006, Xiaoxue Wu 0001, Samuel Akwasi Agyemang |
J. Softw. Evol. Process. | 3 |
| 2022 | A Comprehensive Analysis of NVD Concurrency VulnerabilitiesabstractConcurrency vulnerabilities caused by synchronization problems will occur in the execution of multi-threaded programs, and the emergence of concurrency vulnerabilities often cause great threats to the system. Once the concurrency vulnerabilities are exploited, the system will suffer various attacks, seriously affecting its availability, confidentiality and security. In this paper, we extract 839 concurrency vulnerabilities from Common Vulnerabilities and Exposures (CVE), and conduct a comprehensive analysis of the trend, classifications, causes, severity, and impact. Finally, we obtained some findings: 1) From 1999 to 2021, the number of concurrency vulnerabilities disclosures show an overall upward trend. 2) In the distribution of concurrency vulnerability, race condition accounts for the largest proportion. 3) The overall severity of concurrency vulnerabilities is medium risk. 4) The number of concurrency vulnerabilities that can be exploited for local access and network access is almost equal, and nearly half of the concurrency vulnerabilities (377/839) can be accessed remotely. 5) The access complexity of 571 concurrency vulnerabilities is medium, and the number of concurrency vulnerabilities with high or low access complexity is almost equal. The results obtained through the empirical study can provide more support and guidance for research in the field of concurrency vulnerabilities. Lili Bo, Xing Meng, Xiaobing Sun 0001, Jingli Xia, Xiaoxue Wu 0001 |
QRS | 5 |
| 2022 | SPVF: security property assisted vulnerability fixing via attention-based models
Lili Bo, Xiaoxue Wu 0001, Xiaobing Sun 0001, Tao Zhang 0001, Bin Li 0006, Jiale Zhang 0001, Sicong Cao |
Empir. Softw. Eng. | 3 |
| 2022 | An approach of method-level bug localizationabstractAbstract Bug localization is an important field in software engineering research. The traditional bug localization approaches based on information retrieval separate words through lexical analysis. In this way, the comments of the source code are ignored or treated as plain text, which will lose some semantic information. In this paper, MBL_SHL, an automatic Method‐level Bug Localization approach, which utilises code Summarization, Historical fixed bugs and code Length, is presented. Based on the code summarization technology, this approach first supplements the comment for uncommented code, and then calculates the Word2vec vector and Term Frequency–Inverse Document Frequency vector for the bug report, methods and comments, respectively. After that the authors calculate separately the similarity between the bug report and each method, the bug report and each comment. The code length information and historical fix information are also considered as a weight and a part of the score, respectively, to calculate the final score of each method. Finally, the scores are sorted to determine the list of methods that may need to be modified when fixing the software bugs. We built a method‐granular bug localization dataset, which contains five open‐source projects. The experimental results show that the proposed approach significantly outperforms the existing approaches on the method level. Zhen Ni, Lili Bo, Bin Li 0006, Tianhao Chen, Xiaobing Sun 0001, Xiaoxue Wu 0001 |
IET Softw. | 6 |
| 2022 | Domain knowledge-based security bug reports prediction
Wei Zheng 0006, Jingyuan Cheng, Xiaoxue Wu 0001, Ruiyang Sun, Xiaobing Sun 0001 |
Knowl. Based Syst. | 3 |
| 2022 | ReCDroid+: Automated End-to-End Crash Reproduction from Bug Reports for Android AppsabstractThe large demand of mobile devices creates significant concerns about the quality of mobile applications (apps). Developers heavily rely on bug reports in issue tracking systems to reproduce failures (e.g., crashes). However, the process of crash reproduction is often manually done by developers, making the resolution of bugs inefficient, especially given that bug reports are often written in natural language. To improve the productivity of developers in resolving bug reports, in this paper, we introduce a novel approach, called ReCDroid+, that can automatically reproduce crashes from bug reports for Android apps. ReCDroid+ uses a combination of natural language processing (NLP) , deep learning, and dynamic GUI exploration to synthesize event sequences with the goal of reproducing the reported crash. We have evaluated ReCDroid+ on 66 original bug reports from 37 Android apps. The results show that ReCDroid+ successfully reproduced 42 crashes (63.6% success rate) directly from the textual description of the manually reproduced bug reports. A user study involving 12 participants demonstrates that ReCDroid+ can improve the productivity of developers when resolving crash bug reports. Yu Zhao 0010, Ting Su 0001, Yang Liu 0003, Wei Zheng 0006, Xiaoxue Wu 0001, Ramakanth Kavuluru, William G. J. Halfond, Tingting Yu 0001 |
ACM Trans. Softw. Eng. Methodol. | 5 |
| 2022 | Data Quality Matters: A Case Study on Data Label Correctness for Security Bug Report PredictionabstractIn the research of mining software repositories, we need to label a large amount of data to construct a predictive model. The correctness of the labels will affect the performance of a model substantially. However, limited studies have been performed to investigate the impact of mislabeled instances on a predictive model. To bridge the gap, in this article, we perform a case study on the security bug report (SBR) prediction. We found five publicly available datasets for SBR prediction contains many mislabeled instances, which lead to the poor performance of SBR prediction models of recent studies (e.g., the work of Peterset al.and Shuet al.). Furthermore, it might mislead the research direction of SBR prediction. In this article, we first improve the label correctness of these five datasets by manually analyzing each bug report, and we find 749 SBRs, which are originally mislabeled as Non-SBRs (NSBRs). We then evaluate the impacts of datasets label correctness by comparing the performance of the classification models on both the noisy (i.e., before our correction) and the clean (i.e., after our correction) datasets. The results show that the cleaned datasets result in improvement in the performance of classification models. The performance of the approaches proposed by Peterset al.and Shuet al.on the clean datasets is much better than on the noisy datasets. Furthermore, with the clean datasets, the simple text classification models could significantly outperform the security keywords-matrix-based approaches applied by Peterset al.and Shuet al. Xiaoxue Wu 0001, Wei Zheng 0006, Xin Xia 0001, David Lo 0001 |
IEEE Trans. Software Eng. | 1 |
| 2021 | Automatically Identifying Bug Reports with Tactical Vulnerabilities by Deep Feature LearningabstractIdentifying and fixing bug reports with tactical vul-nerabilities in a timely and accurate manner is essential to ensure the security of the software architecture. Manually identifying the bug reports with tactical vulnerabilities is labor-intensive and challenging. This paper presents Itactivul, an approach to automatically identify bug reports with tactical vulnerabilities and recommend their tactical categories to guide the fix. Unlike the existing security bug report prediction approach, we are the first attempt to use deep learning to mine discriminative tactical text features only from the vulnerability descriptions of the National Vulnerability Database (NVD) and apply them to identify bug reports with tactical vulnerabilities. We evaluate Itactivul on three bug reports datasets gathered from three large-scale open-source projects, including Chromium, PHP, and Thunderbird. The experimental results show that Itactivul outperforms baselines by an average of 8.88 %, 13.58 %, and 6.61 % in the F1-score of three datasets, respectively. To improve the explainability of the features mined by Itactivul, we manually analyze the high-weight phrases extracted by using attention backtracking. The results show that Itactivul can mine key and potential tactical vulnerabilities text features. Wei Zheng 0006, Manqing Zhang, Yuanfang Cai, Xiang Chen 0005, Xiaoxue Wu 0001, Abubakar Omari Abdallah Semasaba |
ISSRE | 6 |
| 2021 | GrasP: Graph-to-Sequence Learning for Automated Program RepairabstractMany deep learning models, for example, neural machine translation (NMT) models, have been developed for Automated Program Repair (APR). Due to the advantages of NMT model's strong generalization ability and less manual in-tervention, NMT-based methods perform well in APR. However, previous NMT-based APR approaches regard a code snippet as a sequence of tokens, which ignores the inherent structure of code. In this paper, we propose a novel end-to-end approach with Graph-to-Sequence learning, GrasP, to generate patches for buggy methods. To better represent the buggy method, we use a graph based on abstract syntax tree (AST) to represent the source code. In order to learn complex graph representation, we introduce the attention-based encoder-decoder model for graph-to-sequence learning. The empirical evaluation on the popular benchmark Defects4J shows that GrasP can generate compilable patches for 75 bugs, of which 34 patches are correct. Ben Tang, Bin Li 0006, Lili Bo, Xiaoxue Wu 0001, Sicong Cao, Xiaobing Sun 0001 |
QRS | 4 |
| 2021 | Representation vs. Model: What Matters Most for Source Code Vulnerability DetectionabstractVulnerabilities in the source code of software are critical issues in the realm of software engineering. Coping with vulnerabilities in software source code is becoming more challenging due to several aspects of complexity and volume. Deep learning has gained popularity throughout the years as a means of addressing such issues. In this paper, we propose an evaluation of vulnerability detection performance on source code representations and evaluate how Machine Learning (ML) strategies can improve them. The structure of our experiment consists of 3 Deep Neural Networks (DNNs) in conjunction with five different source code representations; Abstract Syntax Trees (ASTs), Code Gadgets (CGs), Semantics-based Vulnerability Candidates (SeVCs), Lexed Code Representations (LCRs), and Composite Code Representations (CCRs). Experimental results show that employing different ML strategies in conjunction with the base model structure influences the performance results to a varying degree. However, ML-based techniques suffer from poor performance on class imbalance handling when used in conjunction with source code representations for software vulnerability detection. Wei Zheng 0006, Abubakar Omari Abdallah Semasaba, Xiaoxue Wu 0001, Samuel Akwasi Agyemang |
SANER | 3 |
| 2021 | A survey of Intel SGX and its applications
Wei Zheng 0006, Xiaoxue Wu 0001, Chen Feng 0005, Yulei Sui, Xiapu Luo, Yajin Zhou |
Frontiers Comput. Sci. | 3 |
| 2021 | Improving high-impact bug report prediction with combination of interactive machine learning and active learning
Xiaoxue Wu 0001, Wei Zheng 0006, Xiang Chen 0005, Yu Zhao 0010, Tingting Yu 0001 |
Inf. Softw. Technol. | 1 |
| 2021 | A Comparative Study of Class Rebalancing Methods for Security Bug Report ClassificationabstractIdentifying security bug reports (SBRs) accurately from a bug repository can reduce a software product’s security risk. However, the class imbalance problem exists for SBR prediction since the number of SBRs is often limited, and this issue has not been thoroughly investigated in previous studies. In our study, we choose six real-world projects of different sizes with over 120 000 bug reports in total as our empirical subjects. We first analyze the impact of the class imbalance issue on SBR prediction and confirm its negative impact on prediction performance. Then we perform a comparative study of six state-of-the-art class rebalancing methods combined with five popular classification algorithms for SBR prediction. By comparing with the baseline method Farsec, using the class rebalancing methods can improve the performance in 78% of cases in the worst case. Moreover, the combination of the Rose and random forest classification algorithm can construct the model with the best performance, which increases the performance by 267% in the best case and 75% on average in terms ofF1-score. Finally, we summarize eight main findings based on our empirical studies’ results, which can provide guidelines for choosing appropriate class rebalancing methods and classifiers for SBR prediction in practice. Wei Zheng 0006, Yuxing Xun, Xiaoxue Wu 0001, Xiang Chen 0005, Yulei Sui |
IEEE Trans. Reliab. | 3 |
| 2020 | Literature survey of deep learning-based vulnerability analysis on source codeabstractVulnerabilities in software source code are one of the critical issues in the realm of software code auditing. Due to their high impact, several approaches have been studied in the past few years to mitigate the damages from such vulnerabilities. Among the approaches, deep learning has gained popularity throughout the years to address such issues. In this literature survey, the authors provide an extensive review of the many works in the field software vulnerability analysis that utilise deep learning-based techniques. The reviewed works are systemised according to their objectives (i.e. the type of vulnerability analysis aspect), the area of focus (i.e. the focus area of the analysis), what information about source code is used (i.e. the features), and what deep learning techniques they employ (i.e. what algorithm is used to process the input and produce the output). They also study the limitations of the papers and topical trends concerning vulnerability analysis. Abubakar Omari Abdallah Semasaba, Wei Zheng 0006, Xiaoxue Wu 0001, Samuel Akwasi Agyemang |
IET Softw. | 3 |
| 2020 | CVE-assisted large-scale security bug report dataset construction method
Xiaoxue Wu 0001, Wei Zheng 0006, Xiang Chen 0005 |
J. Syst. Softw. | 1 |
| 2020 | The impact factors on the performance of machine learning-based vulnerability detection: A comparative study
Wei Zheng 0006, Jialiang Gao, Xiaoxue Wu 0001, Yuxing Xun, Xiang Chen 0005 |
J. Syst. Softw. | 3 |
| 2020 | Invalid bug reports complicate the software aging situation
Xiaoxue Wu 0001, Wei Zheng 0006, Minchao Pu |
Softw. Qual. J. | 1 |
| 2019 | Towards understanding bugs in an open source cloud management stack: An empirical study of OpenStack software bugs
Wei Zheng 0006, Chen Feng 0005, Tingting Yu 0001, Xibing Yang, Xiaoxue Wu 0001 |
J. Syst. Softw. | 5 |
| 2018 | MS-guided many-objective evolutionary optimisation for test suite minimisationabstractTest suite minimisation is a process that seeks to identify and then eliminate the obsolete orredundant test cases from the test suite. It is a trade‐off between cost andother value criteria and is appropriate to be described as a many‐objectiveoptimisation problem. This study introduces a mutation score (MS)‐guidedmany‐objective optimisation approach, which prioritises the fault detectionability of test cases and takes MS, cost and three standard code coveragecriteria as objectives for the test suite minimisation process. They use sixclassical evolutionary many‐objective optimisation algorithms to identifyefficient test suite, and select three small programs from the Software‐ArtefactInfrastructure Repository (SIR) and two larger program space and gzip forexperimental evaluation as well as statistical analysis. The experiment resultsof the three small programs show non‐dominated sorting genetic algorithm II(NSGA‐II) with tuning was the most effective approach. However, MOEA/D‐PBI andMOEA/D‐WS outperform NSGA‐II in the cases of two large programs. On the otherhand, the test cost of the optimal test suite obtained by their proposedMS‐guided many‐objective optimisation approach is much lower than the onewithout it in most situation for both small programs and large programs. Wei Zheng 0006, Xiaoxue Wu 0001, Shichao Cao |
IET Softw. | 2 |