VLDB 2026 Research / reviewers in the wild / expert
Zhou Xu 0003
dblp:00/1568-3
· DBLP profile ↗
58ranked-venue papers
15as first author
31since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 45 · 13 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Erratum to "Effective Prediction of Bug-Fixing Priority via Weighted Graph Convolutional Networks"abstractIn [1], the affiliation of the primary authors should be as follows: Sen Fang, Youshuai Tan, Tao Zhang 0001, Zhou Xu 0003, Hui Liu 0003 |
IEEE Trans. Reliab. | 4 |
| 2023 | Extended Abstract of Graph4Web: A Relation-Aware Graph Attention Network for Web Service ClassificationabstractSoftware reuse, as a means to develop new software products with similar functions by virtue of existing software components, has become a popular way during the software development process. In particular, as the service-oriented architecture became popular, web services turned into an indispensable part in modem software development Web services provide a basic composition with high cohesion and loose coupling to support responses among heterogeneous software components, which is the valuable resources for software reuse. The popular web service repositories, such as Programmable Web, contain a mass of web services for beginners and developers to choose from. Nevertheless, the large number of web services also makes it difficult to select the suitable services. Thus, the key to reuse software components lies in how to find appropriate web services from repositories to meet developers requirements in specific application scenarios. Kunsong Zhao, Jin Liu 0016, Zhou Xu 0003, Xiao Liu 0004, Lei Xue 0001, Zhiwen Xie, Xin Wang 0114 |
SANER | 3 |
| 2023 | An empirical study of the impact of log parsers on the performance of log-based anomaly detection
Meng Yan 0001, Zhou Xu 0003, Xin Xia 0001, Xiaohong Zhang 0002, Dan Yang 0001 |
Empir. Softw. Eng. | 3 |
| 2023 | The impact of class imbalance techniques on crashing fault residence prediction models
Kunsong Zhao, Zhou Xu 0003, Meng Yan 0001, Tao Zhang 0001, Lei Xue 0001, Ming Fan 0002, Jacky W. Keung |
Empir. Softw. Eng. | 2 |
| 2023 | MetaFL: Metamorphic fault localisation using weakly supervised deep learningabstractAbstract Deep‐Learning‐based Fault Localisation (DLFL) leverages deep neural networks to learn the relationship between statement behaviour and program failures, showing promising results. However, since DLFL uses program failures as labels to conduct supervised learning, a labelled dataset is a requisite of applying DLFL. A failure is detected by comparing program output with a test oracle which is the standard answer for the given input. The problem is, test oracles are often difficult, or even impossible to acquire in real life, and that has severely restricted the application of DLFL since we have only unlabelled datasets in most cases. Thus, MetaFL: Metamorphic Fault Localisation Using Weakly Supervised Deep Learning is proposed, to provide a weakly supervised learning solution for DLFL. Instead of using test oracles, MetaFL uses metamorphic relations to prescribe expected behaviour of a program, and defines labels of metamorphic testing groups by verifying integrity in each group of test cases. Hence, a coarse‐grained labelled dataset can be built from the originally unlabelled one, with which DLFL can work now, utilising a weakly supervised learning paradigm. The experiments show that MetaFL yields a performance comparable to plain DLFL under ideal condition (i.e. the labels of datasets are available). MetaFL successfully extends the methodology of DLFL from supervised learning to weakly supervised learning, and a fully labelled dataset is no longer mandatory for applying DLFL. Lingfeng Fu, Yan Lei 0005, Meng Yan 0001, Zhou Xu 0003, Xiaohong Zhang 0002 |
IET Softw. | 5 |
| 2023 | Detecting multi-type self-admitted technical debt with generative adversarial network-based neural networks
Jiaojiao Yu 0001, Zhou Xu 0003, Xiao Liu 0004, Jin Liu 0016, Zhiwen Xie, Kunsong Zhao |
Inf. Softw. Technol. | 2 |
| 2023 | CLG-Trans: Contrastive learning for code summarization via graph attention-based transformer
Jianwei Zeng, Tao Zhang 0001, Zhou Xu 0003 |
Sci. Comput. Program. | 4 |
| 2023 | Java Code Clone Detection by Exploiting Semantic and Syntax Information From Intermediate Code-Based GraphabstractCode clone detection plays a critical role in the field of software engineering. To achieve this goal, developers are required to have rich development experience for finding the “functional” clone code. However, this is unfriendly to novice developers. Although many approaches were proposed to automatically detect code clones, the results are not satisfactory. A major reason is that it is difficult to extract syntax and semantic information from the source code. To resolve this problem, in this article, we develop a novel graph representation approach based on intermediate code to detect the functional code clones. This graph representation is built based on intermediate code compiled from the source code. By using it, we can easily utilize graph embedding techniques to extract syntactic and semantic features from abstract syntax tree, control flow graph, and DFG generated from intermediate code. After that, we use the Softmax classifier to detect functional code clone pairs. We evaluate the performance of the proposed graph representation approach based on intermediate code for the code clone detection task on the BigCloneBench dataset. In order to improve performance, the embedded representation of intermediate code is initialized based on pretrained vectors learned from the collected LLVM IR dataset in advance. The experimental results show that our proposed intermediate code-based graph approach performs better than existing functional code clone detection approaches. Especially for the type-4 code clone detection, our approach outperforms the baseline approaches by an average of 33.49% in the term ofF1 score. Dawei Yuan, Sen Fang, Tao Zhang 0001, Zhou Xu 0003, Xiapu Luo |
IEEE Trans. Reliab. | 4 |
| 2023 | Jointly learning invocations and descriptions for context-aware mashup tagging with graph attention network
Xin Wang 0114, Xiao Liu 0004, Hao Wu 0010, Jin Liu 0016, Zhou Xu 0003 |
World Wide Web (WWW) | 6 |
| 2022 | A Naming Pattern Based Approach for Method Name RecommendationabstractMethod names in software projects are significant for developers to understand the method functionality. Existing state-of-the-art automated approaches tend to explore tokens composing method names from method contexts. However, the method name is not a simple combination of tokens, as it is structured and contains many repetitive naming patterns (e.g. “get __”, “create __”). Through a large-scale empirical analysis on 15M methods from 14K real software projects developed with Java codes, we found repetitive naming patterns in method names. In addition, the names of two function-similar methods usually have the same naming pattern. Based on our empirical study, we propose a naming pattern-based approach for method name recommendation, named Nam-Pat. Specifically, for a target method, NamPat first retrieve the most similar method from the training data by estimating their body code similarity. Then, the name of the most similar method is used as the pattern guider to provide the naming pattern, and NamPat combines it with the context information of the target method to perform method name recommendation. To verify the effectiveness of the proposed approach, we conducted experiments on 17M methods from a widely used Java dataset. Experimental results show that compared with Code2vec, Code2seq, MNire, and Cognac, NamPat improves the state-of-the-art approaches in precision (5.8%-27.1%), recall (11.1%-60.1 %), and F-score (8.5 %-43.9%), which proves the effectiveness of our proposed approach. Meng Yan 0001, Zhou Xu 0003, Zhongyang Deng |
ISSRE | 4 |
| 2022 | One step further: evaluating interpreters using metamorphic testingabstractThe black-box nature of the Deep Neural Network (DNN) makes it difficult for people to understand why it makes a specific decision, which restricts its applications in critical tasks. Recently, many interpreters (interpretation methods) are proposed to improve the transparency of DNNs by providing relevant features in the form of a saliency map. However, different interpreters might provide different interpretation results for the same classification case, which motivates us to conduct the robustness evaluation of interpreters. Ming Fan 0002, Jiali Wei, Wuxia Jin, Zhou Xu 0003, Wenying Wei, Ting Liu 0002 |
ISSTA | 4 |
| 2022 | Fine-grained Co-Attentive Representation Learning for Semantic Code SearchabstractCode search aims to find code snippets from large-scale code repositories based on the developer's query intent. A significant challenge for code search is the semantic gap between programming language and natural language. Recent works have indicated that deep learning (DL) techniques can perform well by automatically learning the relationships between query and code. Among these DL-based approaches, the state-of-the-art model is TabCS, a two-stage attention-based model for code search. However, TabCS still has two limitations: semantic loss and semantic confusion. TabCS breaks the structural information of code into token-level words of abstract syntax tree (AST), which loses the sequential semantics between words in programming statements, and it uses a co-attention mechanism to build the semantic correlation of code-query after fusing all features, which may confuse the correlations between individual code features and query. In this paper, we propose a code search model named FcarCS (Fine-grained Co-Attentive Representation Learning Model for Semantic Code Search). FcarCS extracts code textual features (i.e., method name, API sequence, and tokens) and structural features that introduce a statement-level code structure. Unlike TabCS, FcarCS splits AST into a series of subtrees corresponding to code statements and treats each subtree as a whole to preserve sequential semantics between words in code statements. FcarCS constructs a new fine-grained co-attention mechanism to learn interdependent representations for each code feature and query, respectively, instead of performing one co-attention process for the fused code features like TabCS. Generally, this mechanism leverages row/column-wise CNN to enable our model to focus on the strongly correlated local information between code feature and Query. We train and evaluate FcarCS on an open Java dataset with 475k and 10k code/query pairs, respectively. Experimental results show that FcarCS achieves an MRR of 0.613, outperforming three state-of-the-art models DeepCS, UNIF, and TabCS, by 117.38%, 16.76%, and 12.68%, respectively. We also performed a user study for each model with 50 real-world queries, and the results show that FcarCS returned code snippets that are more relevant than the baseline models. Zhongyang Deng, Chao Liu 0014, Meng Yan 0001, Zhou Xu 0003, Yan Lei 0005 |
SANER | 5 |
| 2022 | BCL-FL: A Data Augmentation Approach with Between-Class Learning for Fault LocalizationabstractAutomated fault localization (FL) techniques collect runtime information as input data and then analyze input data to identify the relationship between program statements and failures. They usually take advantages of the statistics of the input data to develop a suspiciousness evaluation methodology (e.g., spectrum-based formulas and deep neural network models) by exploring the underlying correlation rooted in the input data. Thus, the quality of input data is critical for FL. In the actual process of development, developers seek to generate adequate test cases for testing the function or the robustness of a subject program. However, regarding a fault, most test cases are passed test cases and a very few ones are failed test cases since a very small portion of inputs in input domain will lead to a program failure. It means that FL usually faces a problem of imbalanced data, and this problem has been proven to pose an adverse effect on FL effectiveness. To address this problem, we propose BCL-FL: a data augmentation approach based on between-class learning, which produces new synthesized failed test samples by mixing two classes of real test cases (i.e., a passed test case and a failed one) with a random ratio. Specifically, BCL-FL uses the characteristics of real failed test cases to design a data synthesis formula suitable for failed test samples, which can make the synthesized failed test samples closer to real test cases. Since the synthesized data is different from real data, we ingeniously assign a continuous value between 0 and 1 to label the synthesized sample according to the mixing ratio of original labels. We take the synthesized failed test samples and the original test cases as the balanced input data for FL techniques to address the imbalanced data problem. To evaluate the effectiveness of BCL-FL, we conduct large-scale experiments on 287 faulty versions of eight large-sized programs (from ManyBugs and Defects4J) using six state-of-the-art FL approaches. The experimental results show that BCL-FL significantly improves the effectiveness of existing FL techniques, e.g., BCL-FL improves the CNN-FL approach in Top-1, Top-5, and Top-10 by 150%, 136.36%, and 193.1%, respectively. Yan Lei 0005, Huan Xie 0002, Sheng Huang 0001, Meng Yan 0001, Zhou Xu 0003 |
SANER | 6 |
| 2022 | An unsupervised cross project model for crashing fault residence identificationabstractAbstract It is a critical quality assurance activity to effectively detect the root cause of faults causing the software crashes (i.e. crashing faults). Previous studies extracted features to characterise crash instances and built models to identify whether the residences of crashing faults locate inside the stack traces. These models all belong to supervised learning methods which require labelled crash data to be involved. In this study, the introduction of an unsupervised model, called T ransfer S pectral C lustering ( TSC ), for the task of crashing fault residence identification under the unlabelled data scenario is proposed. Unlike traditional unsupervised methods which are applied to individual project data, TSC transfers the knowledge of auxiliary unlabelled data from the source project to assist the clustering task on the unlabelled data from the target project. TSC is an unsupervised transfer learning method, and simultaneously considers the data manifold information of the individual project and feature manifold information across projects to facilitate the clustering effect. Extensive experiments are conducted on a benchmark dataset containing seven software projects. Five indicators were chosen for performance evaluation. The results show that TSC achieves better performance than four clustering based unsupervised methods, and competitive performance compared with eight supervised cross‐project methods. Xiao Liu 0004, Zhou Xu 0003, Dan Yang 0001, Meng Yan 0001, Weihan Zhang, Haohan Zhao, Lei Xue 0001, Ming Fan 0002 |
IET Softw. | 2 |
| 2022 | A compositional model for effort-aware Just-In-Time defect prediction on android appsabstractAbstract Android apps have played important roles in daily life and work. To meet the new requirements from users, the apps encounter frequent updates, which involves a large quantity of code commits. Previous studies proposed to apply Just‐in‐Time (JIT) defect prediction for apps to timely identify whether the new code commits can introduce defects into apps, aiming to assure their quality. In general, high‐quality features are benefits for improving the classification performance. In addition, the number of defective commit instances is much fewer than that of clean ones, that is the defect data is class imbalanced. In this study, a novel compositional model, called KPIDL, is proposed to conduct the JIT defect prediction task for Android apps. More specifically, KPIDL first exploits a feature learning technique to preprocess original data for obtaining better feature representation, and then introduces a state‐of‐the‐art cost‐sensitive cross‐entropy loss function into the deep neural network to alleviate the class imbalance issue by considering the prior probability of the two types of classes. The experiments were conducted on a benchmark defect data consisting of 15 Android apps. The experimental results show that the proposed KPIDL model performs significantly better than 25 comparative methods in terms of two effort‐aware performance indicators in most cases. Kunsong Zhao, Zhou Xu 0003, Meng Yan 0001, Lei Xue 0001, Wei Li 0121, Gemma Catolino |
IET Softw. | 2 |
| 2022 | PRHAN: Automated Pull Request Description Generation Based on Hybrid Attention Network
Sen Fang, Tao Zhang 0001, Youshuai Tan, Zhou Xu 0003, Zhi-Xin Yuan, Ling-Ze Meng |
J. Syst. Softw. | 4 |
| 2022 | Exploiting gated graph neural network for detecting and explaining self-admitted technical debts
Jiaojiao Yu 0001, Kunsong Zhao, Jin Liu 0016, Xiao Liu 0004, Zhou Xu 0003, Xin Wang 0114 |
J. Syst. Softw. | 5 |
| 2022 | Graph4Web: A relation-aware graph attention network for web service classification
Kunsong Zhao, Jin Liu 0016, Zhou Xu 0003, Xiao Liu 0004, Lei Xue 0001, Zhiwen Xie, Xin Wang 0114 |
J. Syst. Softw. | 3 |
| 2022 | DeepDT: Generative Adversarial Network for High-Resolution Climate PredictionabstractClimate prediction is susceptible to a variety of meteorological factors, and downscaling technology is used for high-resolution climate prediction. This technology can generate small-scale regional climate prediction from large-scale climate output information. Inspired by the concept of image super resolution, we propose to apply the convolutional neural network (CNN) to downscaling technology. However, some unpleasant artifacts always appear in the final climate images generated by existing CNN-based models. To further eliminate these unpleasant artifacts, we present a new training strategy for the generative adversarial network, termed DeepDT. The key idea of our DeepDT is to train a generator and a discriminator separately. More specifically, we apply the residual-in-residual dense block as the basic frame structure to fully extract the features of the input. Additionally, we innovatively use a CNN model to fuse multiple climate elements to generate trainable climate images, and build a high-quality climate data set. Finally, we evaluate the DeepDT using the proposed climate data sets, and the experiments indicate that DeepDT performs best compared to most CNN-based models in climate prediction. Jin Liu 0016, Qiuming Kuang, Zhou Xu 0003, Chenkai Shen |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | Feature-FL: Feature-Based Fault LocalizationabstractFault localization aims at developing an effective methodology identifying suspicious statements potentially responsible for program failures. The spectrum-based fault localization is the widely used methodology by analyzing the statistical coincidences viewed from the spectrum to evaluate the suspiciousness of each statement of being faulty. However, just analyzing statistical coincidences in the coverage information perspective and without combining diverse amount of information may restrict fault localization effectiveness. Thus, this article proposes feature-based fault localization (Feature-FL): A family fault localization methodology of feature-based metrics by combining the feature diversity from the view of program features into suspiciousness evaluation. Specifically,Feature-FLdefines a concept of branching execution probability to abstract program behaviors as the values of features. Then,Feature-FLuses feature selection (i.e., a family of feature-based metrics) to evaluate the relevance of each feature with program failures. Finally,Feature-FLassociates each feature with its corresponding statement, and uses the relevance as the suspiciousness to locate suspicious statements. We present six feature-based metrics forFeature-FL, and conduct an extensive study to evaluate the effectiveness ofFeature-FLand its potential over the state-of-the-art spectrum-based formulas. Our results provide insight into the potential among different feature-based metrics and also showFeature-FLsignificantly outperforms the state-of-the-art spectrum-based formulas, e.g., an averagesavingof at least 30% over spectrum-based formulas in case of real faults. Yan Lei 0005, Huan Xie 0002, Tao Zhang 0001, Meng Yan 0001, Zhou Xu 0003, Chengnian Sun |
IEEE Trans. Reliab. | 5 |
| 2022 | Effort-Aware Just-in-Time Bug Prediction for Mobile Apps Via Cross-Triplet Deep Feature EmbeddingabstractJust-in-time (JIT) bug prediction is an effective quality assurance activity that identifies whether a code commit will introduce bugs into the mobile app, aiming to provide prompt feedback to practitioners for priority review. Since collecting sufficient labeled bug data is not always feasible for some mobile apps, one possible approach is to leverage cross-app models. In this work, we propose a new cross-triplet deep feature embedding method, called CDFE, for cross-app JIT bug prediction task. The CDFE method incorporates a state-of-the-art cross-triplet loss function into a deep neural network to learn high-level feature representation for the cross-app data. This loss function adapts to the cross-app feature learning task and aims to learn a new feature space to shorten the distance of commit instances with the same label and enlarge the distance of commit instances with different labels. In addition, this loss function assigns higher weights to losses caused by cross-app instance pairs than that by intra-app instance pairs, aiming to narrow the discrepancy of cross-app bug data. We evaluate our CDFE method on a benchmark bug dataset from 19 mobile apps with two effort-aware indicators. The experimental results on 342 cross-app pairs show that our proposed CDFE method performs better than 14 baseline methods. Zhou Xu 0003, Kunsong Zhao, Tao Zhang 0001, Chunlei Fu, Meng Yan 0001, Zhiwen Xie, Xiaohong Zhang 0002, Gemma Catolino |
IEEE Trans. Reliab. | 1 |
| 2022 | XDebloat: Towards Automated Feature-Oriented App DebloatingabstractExisting programming practices for building Android apps mainly follow the “one-size-fits-all” strategy to include lots of functions and adapt to most types of devices. However, this strategy can result in software bloat and many serious issues, such as slow download speed, and large attack surfaces. Existing solutions cannot effectively debloat an app as they either lack flexibility or require human efforts. This work proposes a novel feature-oriented debloating approach and builds a prototype, namedXDebloat, to automate this process in a flexible manner. First, We propose three feature location approaches to mine features in an app. XDebloat supports feature location approaches at a fine granularity. It also makes the feature location results editable. Second, XDebloat considers several Android-oriented issues (i.e., callbacks) to perform a more precise analysis. Third, XDebloat supports two major debloating strategies: pruning-based debloating and module-based debloating. We evaluate XDebloat with 200 open-source and 1,000 commercial apps. The results show that XDebloat can successfully remove components from apps or transform apps into on-demand modules within 10 minutes. For thepruning-baseddebloating strategy, on average, XDebloat can remove 32.1% code from an app. For themodule-baseddebloating strategy, XDebloat can help developers build instant apps or app bundles automatically. Yutian Tang, Hao Zhou 0043, Xiapu Luo, Ting Chen 0002, Haoyu Wang 0001, Zhou Xu 0003, Yan Cai 0001 |
IEEE Trans. Software Eng. | 6 |
| 2021 | Contextual-Semantic-Aware Linkable Knowledge Prediction in Stack Overflow via Self-AttentionabstractIn Stack Overflow, a question and its answers are defined as a knowledge unit. These knowledge units can be linked together for different purposes, which typically subdivided into four classes: Duplicate, Directly linkable, Indirectly linkable, and Isolated. Developers usually use these linkable knowledge units to search for more targeted information. Prior studies have found that deep learning or SVM technique can effectively predict the class of linkable knowledge units. However, they focus on short-distance semantic relationship but fail to capture global information (semantic relationship between a word and all the words in the same knowledge unit) and ignore joint semantics (semantic relationship between a word with all the words in different knowledge units). To address the issues, we propose a Self-Attention-based contextual semantic aware Linkable Knowledge prediction model (SALKU). SALKU leverages self-attention to pay attention to all the words in a knowledge unit and fully capture the global information needed for each word, then utilizes a variant of self-attention to extract joint semantics between two knowledge units. Experiment results on an existing dataset show that SALKU out-performs the state-of-the-art approaches CNN, Tuning SVM, and Soft-cos SVM in terms of three metrics, respectively. Additionally, SALKU is faster than the three baseline approaches. Zhaolin Luo, Zhou Xu 0003, Meng Yan 0001, Yan Lei 0005, Can Li 0015 |
ISSRE | 3 |
| 2021 | Predicting Crash Fault Residence via Simplified Deep Forest Based on A Reduced Feature SetabstractThe software inevitably encounters the crash, which will take developers a large amount of effort to find the fault causing the crash (short for crashing fault). Developing automatic methods to identify the residence of the crashing fault is a crucial activity for software quality assurance. Researchers have proposed methods to predict whether the crashing fault resides in the stack trace based on the features collected from the stack trace and faulty code, aiming at saving the debugging effort for developers. However, previous work usually neglected the feature preprocessing operation towards the crash data and only used traditional classification models. In this paper, we propose a novel crashing fault residence prediction framework, called ConDF, which consists of a consistency based feature subset selection method and a state-of-the-art deep forest model. More specifically, first, the feature selection method is used to obtain an optimal feature subset and reduce the feature dimension by reserving the representative features. Then, a simplified deep forest model is employed to build the classification model on the reduced feature set. The experiments on seven open source software projects show that our ConDF method performs significantly better than 17 baseline methods on three performance indicators. Kunsong Zhao, Jin Liu 0016, Zhou Xu 0003, Li Li 0029, Meng Yan 0001, Jiaojiao Yu 0001 |
ICPC | 3 |
| 2021 | DG-Trans: Automatic Code Summarization via Dynamic Graph Attention-based TransformerabstractAutomatic code summarization is an important topic in the software engineering field, which aims to automatically generate the description for the source code. Based on Graph Neural Networks (GNN), most existing methods apply them to Abstract Syntax Tree (AST) to achieve code summarization. However, these methods face two major challenges: 1) they can only capture limited structural information of the source code; 2) they did not effectively solve Out-Of-Vocabulary (OOV) problems by reducing vocabulary size. In order to resolve these problems, in this paper, we propose a novel code summarization model named Dynamic Graph attention-based Transformer (DG-Trans for short), which effectively captures abundant information of the code subword sequence and utilizes the fusion of dynamic graph attention mechanism and Transformer. Extensive experiments show that DG-Trans is able to outperform state-of-the-art models (such as Ast-Attendgru, Transformer, and CodeGNN) by averagely increasing 8.39% and 8.86% on BLEU scores and ROUGUE-L, respectively. Jianwei Zeng, Tao Zhang 0001, Zhou Xu 0003 |
QRS | 3 |
| 2021 | Two-Stage Attention-Based Model for Code Search with Textual and Structural FeaturesabstractSearching and reusing existing code from a large scale codebase can largely improve developers’ programming efficiency. To support code reuse, early code search models leverage information retrieval (IR) techniques to index a large-scale code corpus and return relevant code according to developers’ search query. However, IR-based models fail to capture the semantics in code and query. To tackle this issue, developers applied deep learning (DL) techniques to code search models. However, these models either are too complex to determine an effective method efficiently or learning for semantic correlation between code and query inadequately.To bridge the semantic gap between code and query effectively and efficiently, we propose a code search model TabCS (Two-stage Attention-Based model for Code Search) in this study. TabCS extracts code and query information from the code textual features (i.e., method name, API sequence, and tokens), the code structural feature (i.e., abstract syntax tree), and the query feature (i.e., tokens). TabCS performs a two-stage attention net-work structure. The first stage leverages attention mechanisms to extract semantics from code and query considering their semantic gap. The second stage leverages a co-attention mechanism to capture their semantic correlation and learn better code/query representation. We evaluate the performance of TabCS on two existing large-scale datasets with 485k and 542k code snippets, respectively. Experimental results show that TabCS achieves an MRR of 0.57 on Hu et al.’s dataset, outperforming three state-of-the-art models CARLCS-CNN, DeepCS, and UNIF by 18%, 70%, 12%, respectively. Meanwhile, TabCS gains an MRR of 0.54 on Husain et al.’s, outperforming CARLCS-CNN, DeepCS, and UNIF by 32%, 76%, 29%, respectively. Huanhuan Yang, Chao Liu 0014, Jianhang Shuai, Meng Yan 0001, Yan Lei 0005, Zhou Xu 0003 |
SANER | 7 |
| 2021 | Feature selection and embedding based cross project framework for identifying crashing fault residence
Zhou Xu 0003, Tao Zhang 0001, Jacky W. Keung, Meng Yan 0001, Xiapu Luo, Xiaohong Zhang 0002, Yutian Tang |
Inf. Softw. Technol. | 1 |
| 2021 | A comprehensive investigation of the impact of feature selection techniques on crashing fault residence prediction models
Kunsong Zhao, Zhou Xu 0003, Meng Yan 0001, Tao Zhang 0001, Dan Yang 0001, Wei Li 0121 |
Inf. Softw. Technol. | 2 |
| 2021 | A comprehensive comparative study of clustering-based unsupervised defect prediction models
Zhou Xu 0003, Li Li 0029, Meng Yan 0001, Jin Liu 0016, Xiapu Luo, John C. Grundy, Xiaohong Zhang 0002 |
J. Syst. Softw. | 1 |
| 2021 | Effective Prediction of Bug-Fixing Priority via Weighted Graph Convolutional NetworksabstractWith the increasing number of software bugs, bug fixing plays an important role in software development and maintenance. To improve the efficiency of bug resolution, developers utilize bug reports to resolve given bugs. Especially, bug triagers usually depend on bugs' descriptions to suggest priority levels for reported bugs. However, manual priority assignment is a time-consuming and cumbersome task. To resolve this problem, recent studies have proposed many approaches to automatically predict the priority levels for the reported bugs. Unfortunately, these approaches still face two challenges that include words' nonconsecutive semantics in bug reports and the imbalanced data. In this article, we propose a novel approach that graph convolutional networks (GCN) based on weighted loss function to perform the priority prediction for bug reports. For the first challenge, we build a heterogeneous text graph for bug reports and apply GCN to extract words' semantics in bug reports. For the second challenge, we construct a weighted loss function in the training phase. We conduct the priority prediction on four open-source projects, including Mozilla, Eclipse, Netbeans, and GNU compiler collection. Experimental results show that our method outperforms two baseline approaches in terms of the F-measure by weighted average of 13.22%. Sen Fang, Youshuai Tan, Tao Zhang 0001, Zhou Xu 0003, Hui Liu 0003 |
IEEE Trans. Reliab. | 4 |
| 2021 | Simplified Deep Forest Model Based Just-in-Time Defect Prediction for Android Mobile AppsabstractThe popularity of mobile devices has led to an explosive growth in the number of mobile apps in which Android mobile apps are the mainstream. Android mobile apps usually undergo frequent update due to new requirements proposed by users. Just-in-time (JIT) defect prediction is appropriate for this scenario for quality assurance because it can provide timely feedback by determining whether a new code commit will introduce defects into the apps. As defect-prediction performance usually relies on the quality of the data representation and the used classification model, in this work, we propose a model, called Simplified Deep Forest (SDF), to conduct JIT defect prediction for Android mobile apps. SDF modifies a state-of-the-art deep forest model by removing the multigrained scanning operation that is designed for data with a high-dimensional feature space. It uses a cascade structure with ensemble forests for representation learning and classification. We conduct experiments on 10 Android mobile apps and experimental results show that SDF performs significantly better than comparative methods in terms of 3 performance indicators. Kunsong Zhao, Zhou Xu 0003, Tao Zhang 0001, Yutian Tang, Meng Yan 0001 |
IEEE Trans. Reliab. | 2 |
| 2020 | Blocking Bug Prediction Based on XGBoost with Enhanced FeaturesabstractWith a growing number of software projects, software quality is increasingly crucial. Researchers and engineers in the software engineering field often pay much attention to bug management tasks, such as bug localization, bug triage, and duplicate bug detection. However, there are few researchers to study blocking bug prediction. Blocking bugs prevent other bugs from being fixed and usually need more time to be fixed. Thus, developers need to identify blocking bugs and reduce the impact of blocking bugs. The previous studies utilized supervised algorithms to implement this task. However, they did not consider the dependencies among individual classifiers so that they cannot get the perfect accuracy for blocking bug prediction. In this paper, we propose a new framework XGBlocker that includes two stages. In the first stage, XGBlocker collects more features from bug reports to build an enhanced dataset. In the second stage, XGBlocker exploits XGBoost technique to construct an effective model to perform the prediction task. We conduct experiments on four projects with three evaluation metrics. The experimental results show that our method XGBlocker achieves promising performance compared with baseline methods in most cases. In detail, XGBlocker achieves F1-score, ER@20%, and AUC of up to 0.808, 0.944, and 0.975, respectively. On average across the four projects, XGBlocker improves F1-score, ER@20%, and AUC over the state-of-the-art method ELBlocker by 17.27%, 12.67%, and 4.85%, respectively. Xiaoyun Cheng, Naming Liu, Zhou Xu 0003, Tao Zhang 0001 |
COMPSAC | 4 |
| 2020 | Improving Log-Based Anomaly Detection with Component-Aware AnalysisabstractLogs are universally available in software systems for troubleshooting. They record system run-time states and messages of system activities. Log analysis is an effective way to diagnosis system exceptions, but it will take a long time for engineers to locate anomalies accurately through logs. Many automatic approaches have been proposed for log-based anomaly detection. However, most of the prior approaches did not consider the corresponding system component of a log message. Such component records the log location, which can help detect the location-sequence-related anomalies. In this paper, we propose LogC, a new Log -based anomaly detection approach with Component-aware analysis. LogC contains two phases: (i) turning log messages into log template sequences and component sequences, (ii) feeding such two sequences to train a combined LSTM model for detecting anomalous logs. LogC only needs normal log sequences to train the combined model. We evaluate LogC on two open-source log datasets: HDFS and ThunderBird. Experimental results show that LogC overall outperforms three baselines (i.e., PCA, IM, and DeepLog) in terms of three metrics (precision, recall, and F-measure). Kun Yin, Meng Yan 0001, Zhou Xu 0003, Dan Yang 0001, Xiaohong Zhang 0002 |
ICSME | 4 |
| 2020 | Deep Learning Based Valid Bug Reports Determination and ExplanationabstractBug reports are widely used by developers to fix bugs. Due to the lack of experience, reporters may submit numerous invalid bug reports. Manually determining valid bug reports is a laborious task. Automatically identifying valid bug reports can save time and effort for bug analysis. In this paper, we propose a deep learning-based approach to determine and explain valid bug reports using only textual information i.e., summaries and descriptions of bug reports. Convolutional neural network (CNN) is applied to capture their contextual and semantic features. Moreover, by analyzing the spatial structure of CNN, we backtrack the trained CNN model to get phrases that can explain valid bug reports determination. After inspecting the phrases manually, we summarize some valid bug report patterns. We evaluate our approach on five large-scale open-source projects containing a total of 540491 bug reports. On average, across the five projects, our approach achieves 0.85, 0.80, 0.69 and improves the state-of-the-art approach by 8.97%, 9.59%, 9.52% in terms of AUC, F1-score for valid bug reports, and F1-score for invalid bug reports, respectively. From the summarized patterns, we can find that determining valid bug reports is mainly due to three categories of patterns: Attachment, Environment, and Reproduce. Yuanrui Fan, Zhou Xu 0003, Meng Yan 0001, Yan Lei 0005 |
ISSRE | 4 |
| 2020 | STAN: Towards Describing Bytecodes of Smart ContractabstractMore than eight million smart contracts have been deployed into Ethereum, which is the most popular blockchain that supports smart contract. However, less than 1% of deployed smart contracts are open-source, and it is difficult for users to understand the functionality and internal mechanism of those closed-source contracts. Although a few decompilers for smart contracts have been recently proposed, it is still not easy for users to grasp the semantic information of the contract, not to mention the potential misleading due to decompilation errors. In this paper, we propose the first system named Stan to generate descriptions for the bytecodes of smart contracts to help users comprehend them. In particular, for each interface in a smart contract, Stan can generate four categories of descriptions, including functionality description, usage description, behavior description, and payment description, by leveraging symbolic execution and NLP (Natural Language Processing) techniques. Extensive experiments show that Stan can generate adequate, accurate and readable descriptions for contract's bytecodes, which have practical value for users. Xiaoqi Li 0001, Ting Chen 0002, Xiapu Luo, Tao Zhang 0001, Le Yu 0002, Zhou Xu 0003 |
QRS | 6 |
| 2020 | Simplified Deep Forest Model based Just-In-Time Defect Prediction for Android Mobile AppsabstractThe popularity of mobile devices has led to an explosive growth in the number of mobile apps in which Android mobile apps are the mainstream. Android mobile apps usually undergo frequent update due to new requirements proposed by users. Just-In-Time (JIT) defect prediction is appropriate for this scenario for quality assurance because it can provide timely feedback by determining whether a new code commit will introduce defects into the apps. As defect prediction performance usually relies on the quality of the data representation and the used classification model, in this work, we modify a state-of-the-art model, called Simplified Deep Forest (SDF) to conduct JIT defect prediction for Android mobile apps. This method uses a cascade structure with ensemble forests for representation learning and classification. We conduct experiments on 10 Android mobile apps and experimental results show that SDF performs significantly better than comparative methods in terms of three performance indicators. Kunsong Zhao, Zhou Xu 0003, Tao Zhang 0001, Yutian Tang |
QRS | 2 |
| 2020 | All your app links are belong to us: understanding the threats of instant apps based attacksabstractAndroid deep link is a URL that takes users to a specific page of a mobile app, enabling seamless user experience from a webpage to an app. Android app link, a new type of deep link introduced in Android 6.0, is claimed to offer more benefits, such as supporting instant apps and providing more secure verification to protect against hijacking attacks that previous deep links can not. However, we find that the app link is not as secure as claimed, because the verification process can be bypassed by exploiting instant apps. Yutian Tang, Yulei Sui, Haoyu Wang 0001, Xiapu Luo, Hao Zhou 0043, Zhou Xu 0003 |
ESEC/SIGSOFT FSE | 6 |
| 2020 | Group sparse additive machine with average top-k loss
Peipei Yuan, Xinge You, Hong Chen 0004, Qinmu Peng, Zhou Xu 0003, Xiaoyuan Jing, Zhenyu He 0001 |
Neurocomputing | 6 |
| 2020 | Imbalanced metric learning for crashing fault residence prediction
Zhou Xu 0003, Kunsong Zhao, Meng Yan 0001, Peipei Yuan, Yan Lei 0005, Xiaohong Zhang 0002 |
J. Syst. Softw. | 1 |
| 2020 | Bug severity prediction using question-and-answer pairs from Stack OverflowabstractNowadays, bugs have been common in most software systems. For large-scale software projects, developers usually conduct software maintenance tasks by utilizing software artifacts (e.g., bug reports). The severity of bug reports describes the impact of the bugs and determines how quickly it needs to be fixed. Bug triagers often pay close attention to some features such as severity to determine the importance of bug reports and assign them to the correct developers. However, a large number of bug reports submitted every day increase the workload of developers who have to spend more time on fixing bugs. In this paper, we collect question-and-answer pairs from Stack Overflow and use logical regression to predict the severity of bug reports. In detail, we extract all the posts related to bug repositories from Stack Overflow and combine them with bug reports to obtain enhanced versions of bug reports. We achieve severity prediction on three popular open source projects (e,g., Mozilla, Ecplise, and GCC) with Naïve Bayesian, k-Nearest Neighbor algorithm (KNN), and Long Short-Term Memory (LSTM). The results of our experiments show that our model is more accurate than the previous studies for predicting the severity. Our approach improves by 23.03%, 21.86%, and 20.59% of the average F-measure for Mozilla, Eclipse, and GCC by comparing with the Naïve Bayesian based approach which performs the best among all baseline approaches. Youshuai Tan, Sijie Xu, Zhaowei Wang 0003, Tao Zhang 0001, Zhou Xu 0003, Xiapu Luo |
J. Syst. Softw. | 5 |
| 2020 | Noisy-as-Clean: Learning Self-Supervised Denoising From Corrupted ImageabstractSupervised deep networks have achieved promising performance on image denoising, by learning image priors and noise statistics on plenty pairs of noisy and clean images. Unsupervised denoising networks are trained with only noisy images. However, for an unseen corrupted image, both supervised and unsupervised networks ignore either its particular image prior, the noise statistics, or both. That is, the networks learned from external images inherently suffer from a domain gap problem: the image priors and noise statistics are very different between the training and test images. This problem becomes more clear when dealing with the signal dependent realistic noise. To circumvent this problem, in this work, we propose a novel "Noisy-As-Clean" (NAC) strategy of training self-supervised denoising networks. Specifically, the corrupted test image is directly taken as the "clean" target, while the inputs are synthetic images consisted of this corrupted image and a second yet similar corruption. A simple but useful observation on our NAC is: as long as the noise is weak, it is feasible to learn a self-supervised network only with the corrupted image, approximating the optimal parameters of a supervised network learned with pairs of noisy and clean images. Experiments on synthetic and realistic noise removal demonstrate that, the DnCNN and ResNet networks trained with our self-supervised NAC strategy achieve comparable or better performance than the original ones and previous supervised/unsupervised/self-supervised networks. The code is publicly available at https://github.com/csjunxu/Noisy-As-Clean. Jun Xu 0019, Ming-Ming Cheng, Li Liu 0004, Fan Zhu 0001, Zhou Xu 0003, Ling Shao 0001 |
IEEE Trans. Image Process. | 6 |
| 2020 | Improving Ranking-Oriented Defect Prediction Using a Cost-Sensitive Ranking SVMabstractContext: Ranking-oriented defect prediction (RODP) ranks software modules to allocate limited testing resources to each module according to the predicted number of defects. Most RODP methods overlook that ranking a module with more defects incorrectly makes it difficult to successfully find all of the defects in the module due to fewer testing resources being allocated to the module, which results in much higher costs than incorrectly ranking the modules with fewer defects, and the numbers of defects in software modules are highly imbalanced in defective software datasets. Cost-sensitive learning is an effective technique in handling the cost issue and data imbalance problem for software defect prediction. However, the effectiveness of cost-sensitive learning has not been investigated in RODP models. Aims: In this article, we propose a cost-sensitive ranking support vector machine (SVM) (CSRankSVM) algorithm to improve the performance of RODP models. Method: CSRankSVM modifies the loss function of the ranking SVM algorithm by adding two penalty parameters to address both the cost issue and the data imbalance problem. Additionally, the loss function of the CSRankSVM is optimized using a genetic algorithm. Results: The experimental results for 11 project datasets with 41 releases show that CSRankSVM achieves 1.12%-15.68% higher average fault percentile average (FPA) values than the five existing RODP methods (i.e., decision tree regression, linear regression, Bayesian ridge regression, ranking SVM, and learning-to-rank (LTR)) and 1.08%-15.74% higher average FPA values than the four data imbalance learning methods (i.e., random undersampling and a synthetic minority oversampling technique; two data resampling methods; RankBoost, an ensemble learning method; IRSVM, a CSRankSVM method for information retrieval). Conclusion: CSRankSVM is capable of handling the cost issue and data imbalance problem in RODP methods and achieves better performance. Therefore, CSRankSVM is recommended as an effective method for RODP. Xiao Yu 0008, Jin Liu 0016, Jacky W. Keung, Qing Li 0001, Kwabena Ebo Bennin, Zhou Xu 0003, Xiaohui Cui |
IEEE Trans. Reliab. | 6 |
| 2019 | Identifying Crashing Fault Residence Based on Cross Project ModelabstractAnalyzing the crash reports recorded upon software crashes is a critical activity for software quality assurance. Predicting whether or not the fault causing the crash (crashing fault for short) resides in the stack traces of crash reports can speed-up the program debugging process and determine the priority of the debugging efforts. Previous work mostly collected label information from bug-fixing logs, and extracted crash features from stack traces and source code to train classification models for the Identification of Crashing Fault Residence (ICFR) of newly-submitted crashes. However, labeled data are not always fully available in real applications. Hence the classifier training is not always feasible. In this work, we make the first attempt to develop a cross project ICFR model to address the data scarcity problem. This is achieved by transferring the knowledge from external projects to the current project via utilizing a state-of-the-art Balanced Distribution Adaptation (BDA) based transfer learning method. BDA not only combines both marginal distribution and conditional distribution across projects but also assigns adaptive weights to the two distributions for better adjusting specific cross project pair. The experiments on 7 software projects show that BDA is superior to 9 baseline methods in terms of 6 indicators overall. Zhou Xu 0003, Tao Zhang 0001, Yutian Tang, Jin Liu 0016, Xiapu Luo, Jacky W. Keung, Xiaohui Cui |
ISSRE | 1 |
| 2019 | Demystifying Application Performance Management Libraries for AndroidabstractSince the performance issues of apps can influence users' experience, developers leverage application performance management (APM) tools to locate the potential performance bottleneck of their apps. Unfortunately, most developers do not understand how APMs monitor their apps during the runtime and whether these APMs have any limitations. In this paper, we demystify APMs by inspecting 25 widely-used APMs that target on Android apps. We first report how these APMs implement 8 key functions as well as their limitations. Then, we conduct a large-scale empirical study on 500,000 Android apps from Google Play to explore the usage of APMs. This study has some interesting observations about existing APMs for Android, including 1) some APMs still use deprecated permissions and approaches so that they may not always work properly; 2) some app developers use APMs to collect users' privacy information. Yutian Tang, Xian Zhan, Hao Zhou 0043, Xiapu Luo, Zhou Xu 0003, Yajin Zhou, Qiben Yan 0001 |
ASE | 5 |
| 2019 | MVSE: Effort-Aware Heterogeneous Defect Prediction via Multiple-View Spectral EmbeddingabstractCross-Project Defect Prediction (CPDP) predicts defects in a target project using the defect information of the external project. Existing CPDP methods assume that the data of two projects share identical features. When cross-project data contain heterogeneous features, traditional CPDP methods become ineffective. In this paper, we propose a novel approach called Multiple-View Spectral Embedding (MVSE) to address the heterogeneous CPDP issue. MVSE treats the cross-project data as two different views and exploits the spectral embedding method to map the heterogeneous feature sets into a consistent space where the two mapped feature sets have maximal similarity. To evaluate MVSE in the realistic setting, we employ an effort-aware performance indicator that considers the cost of inspection in the context of heterogeneous CPDP scenario. We have conducted extensive experiments to compare MVSE with two state-of-the-art heterogeneous CPDP methods and within-project setting. The experiments on 94 cross project pairs show that MVSE achieves promising results. Zhou Xu 0003, Sizhe Ye, Tao Zhang 0001, Zhen Xia, Shuai Pang, Yong Wang 0020, Yutian Tang |
QRS | 1 |
| 2019 | An Empirical Study of Learning to Rank Techniques for Effort-Aware Defect PredictionabstractEffort-Aware Defect Prediction (EADP) ranks software modules based on the possibility of these modules being defective, their predicted number of defects, or defect density by using learning to rank algorithms. Prior empirical studies compared a few learning to rank algorithms considering small number of datasets, evaluating with inappropriate or one type of performance measure, and non-robust statistical test techniques. To address these concerns and investigate the impact of learning to rank algorithms on the performance of EADP models, we examine the practical effects of 23 learning to rank algorithms on 41 available defect datasets from the PROMISE repository using a module-based effort-aware performance measure (FPA) and a source lines of code (SLOC) based effort-aware performance measure (Norm(Popt). In addition, we compare the performance of these algorithms when they are trained on a more relevant feature subset selected by the Information Gain feature selection method. In terms of FPA and Norm(Popt), statistically significant differences are observed among these algorithms with BRR (Bayesian Ridge Regression) performing best in terms of FPA, and BRR and LTR (Learning-to-Rank) performing best in terms of Norm (Popt). When these algorithms are trained on a more relevant feature subset selected by Information Gain, LTR and BRR still perform best with significant differences in terms of FPA and Norm(Popt). Therefore, we recommend BRR and LTR for building the EADP model in order to find more defects by inspecting a certain number of modules or lines of codes. Xiao Yu 0008, Kwabena Ebo Bennin, Jin Liu 0016, Jacky W. Keung, Xiaofei Yin, Zhou Xu 0003 |
SANER | 6 |
| 2019 | Labelling issue reports in mobile appsabstractMillions of mobile apps have been released to the market. Developers need to maintain these apps so that they can continue to benefit end users, who usually submit issue reports to describe the bugs, the feature requests, and other changes appearing in apps. The labels (e.g. bug, feature request) are important resources to indicate which issue reports should be resolved first or next. According to the investigation, 35.6% of issue reports in top‐17 popular mobile apps are not labelled. Developers have to spend additional time to manually verify each unlabelled issue report so that they can decide to resolve the most important issues. In order to help developers to reduce the workload, in this study, the authors propose a novel approach to automatically tag the unlabelled issue reports. This approach not only computes the similarity between each unlabelled issue report and user reviews related to bugs and features but also calculates the textual similarity scores between each unlabelled issue report and labelled ones. As a result, among all textual similarity measures, this approach using cosine similarity with MCG shows the best performance. Moreover, this approach performs better than the method proposed in the authors' previous study. Tao Zhang 0001, Haoming Li 0005, Zhou Xu 0003, Rubing Huang, Yiran Shen 0001 |
IET Softw. | 3 |
| 2019 | Watermelon Ripeness Detection via Extreme Learning Machine with Kernel Principal Component Analysis Based on Acoustic SignalsabstractMany investigations have proved that the acoustics method is intuitive and effective for determining watermelon ripeness. The objective of this work is to drive a new robust acoustics classification scheme KPCA-ELM, which is based on the kernel principal component analysis (KPCA) and extreme learning machine (ELM). Acoustic signals are sampled by a microphone from unripe, ripe and over-ripe watermelon samples, which are randomly divided into two sample sets for training and testing. A set of basic signals is first obtained via KPCA of the training sample. Thus, any given signal can be represented as a linear combination of basis signals, and the coefficients of linear combination are extracted as the features of a signal. Corresponding to the unripe, ripe and over-ripe watermelons, a three-class ELM identification model is constructed based on the training data. The scheme presented in this paper is tested with the testing sample and an accuracy of 92% is achieved. To further evaluate the scheme performance, a comparison of ELM and SVM is conducted in terms of the classification results. The results reveal that the proposed scheme can classify faster than SVM, while ELM is better than SVM in accuracy. Xiaoyan Deng, Zhou Xu 0003, Peipei Yuan |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2019 | Software defect prediction based on kernel PCA and weighted extreme learning machine
Zhou Xu 0003, Jin Liu 0016, Xiapu Luo, Zijiang Yang 0006, Peipei Yuan, Yutian Tang, Tao Zhang 0001 |
Inf. Softw. Technol. | 1 |
| 2019 | Cross Project Defect Prediction via Balanced Distribution Adaptation Based Transfer Learning
Zhou Xu 0003, Shuai Pang, Tao Zhang 0001, Xiapu Luo, Jin Liu 0016, Yutian Tang, Xiao Yu 0008, Lei Xue 0001 |
J. Comput. Sci. Technol. | 1 |
| 2019 | TSTSS: A two-stage training subset selection framework for cross version defect predictionabstractCross Version Defect Prediction (CVDP) is a practical scenario by training the classification model on the historical data of the prior version and then predicting the defect labels of modules in the current version. Unfortunately, the differences of data distribution across versions may hinder the effectiveness of the trained CVDP model. Thus, it is not trivial to select a suitable training subset from the prior version to promote the CVDP performance. In this paper, we propose a novel method, called Two-Stage Training Subset Selection (TSTSS), to address this challenging issue. In the first stage, TSTSS utilizes a sparse modeling representative selection method to select an initial module subset from the prior version which can well reconstruct the data of the prior version. In the second stage, TSTSS leverages a dissimilarity-based sparse subset selection method to further refine the selected module subset, which enables the selected modules to well represent the modules of the current version. Finally, we use a novel weighted extreme learning machine classifier to construct the CVDP model. We evaluate the CVDP performance of TSTSS on 50 cross-version pairs using 6 indicators. The experiments show that TSTSS can efficiently improve the CVDP performance compared with 11 baseline methods. Zhou Xu 0003, Shuai Li 0014, Xiapu Luo, Jin Liu 0016, Tao Zhang 0001, Yutian Tang, Jun Xu 0019, Peipei Yuan, Jacky W. Keung |
J. Syst. Softw. | 1 |
| 2019 | LDFR: Learning deep feature representation for software defect prediction
Zhou Xu 0003, Shuai Li 0014, Jun Xu 0019, Jin Liu 0016, Xiapu Luo, Tao Zhang 0001, Jacky W. Keung, Yutian Tang |
J. Syst. Softw. | 1 |
| 2018 | Cross version defect prediction with representative data via sparse subset selectionabstractSoftware defect prediction aims at detecting the defect-prone software modules by mining historical development data from software repositories. If such modules are identified at the early stage of the development, it can save large amounts of resources. Cross Version Defect Prediction (CVDP) is a practical scenario by training the classification model on the historical data of the prior version and then predicting the defect labels of modules of the current version. However, software development is a constantly-evolving process which leads to the data distribution differences across versions within the same project. The distribution differences will degrade the performance of the classification model. In this paper, we approach this issue by leveraging a state-of-the-art Dissimilarity-based Sparse Subset Selection (DS3) method. This method selects a representative module subset from the prior version based on the pairwise dissimilarities between the modules of two versions and assigns each module of the current version to one of the representative modules. These selected modules can well represent the modules of the current version, thus mitigating the distribution differences. We evaluate the effectiveness of DS3 for CVDP performance on total 40 cross-version pairs from 56 versions of 15 projects with three traditional and two effort-aware indicators. The extensive experiments show that DS3 outperforms three baseline methods, especially in terms of two effort-aware indicators. Zhou Xu 0003, Shuai Li 0014, Yutian Tang, Xiapu Luo, Tao Zhang 0001, Jin Liu 0016, Jun Xu 0019 |
ICPC | 1 |
| 2018 | Cross-version defect prediction via hybrid active learning with kernel principal component analysisabstractAs defects in software modules may cause product failure and financial loss, it is critical to utilize defect prediction methods to effectively identify the potentially defective modules for a thorough inspection, especially in the early stage of software development lifecycle. For an upcoming version of a software project, it is practical to employ the historical labeled defect data of the prior versions within the same project to conduct defect prediction on the current version, i.e., Cross-Version Defect Prediction (CVDP). However, software development is a dynamic evolution process that may cause the data distribution (such as defect characteristics) to vary across versions. Furthermore, the raw features usually may not well reveal the intrinsic structure information behind the data. Therefore, it is challenging to perform effective CVDP. In this paper, we propose a two-phase CVDP framework that combines Hybrid Active Learning and Kernel PCA (HALKP) to address these two issues. In the first stage, HALKP uses a hybrid active learning method to select some informative and representative unlabeled modules from the current version for querying their labels, then merges them into the labeled modules of the prior version to form an enhanced training set. In the second stage, HALKP employs a non-linear mapping method, kernel PCA, to extract representative features by embedding the original data of two versions into a high-dimension space. We evaluate the HALKP framework on 31 versions of 10 projects with three prevalent performance indicators. The experimental results indicate that HALKP achieves encouraging results with average F-measure, g-mean and Balance of 0.480, 0.592 and 0.580, respectively and significantly outperforms nearly all baseline methods. Zhou Xu 0003, Jin Liu 0016, Xiapu Luo, Tao Zhang 0001 |
SANER | 1 |
| 2017 | An Empirical Study on the Equivalence and Stability of Feature Selection for Noisy Software Defect DataabstractSoftware Defect Data (SDD) are used to build defect prediction models for software quality assurance.Existing work employs feature selection to eliminate irrelevant features in the data to improve prediction performance.Previous studies have shown that different feature selection methods do not always yield similar prediction performance on SDD, which indicates that these methods are not equivalent.Also, previous studies have shown that SDD usually contains noise that may interfere the process of feature selection.In this work, we empirically investigate and measure the equivalence of different feature selection methods for SDD.Further, we intend to analyze the stability of the methods for noisy SDD.We perform statistical analyses on eight projects from NASA dataset with eight feature selection methods.For the equivalence analysis, we introduce Principal Component Analysis (PCA) and overlap index to qualitatively and quantitatively analyze the equivalence of these methods respectively.For the stability analysis, we apply consistency index to measure the stability of these methods.Experimental results indicate that different feature selection methods are indeed not equivalent to each other, and Correlation and Fisher Score methods achieve better stability. Zhou Xu 0003, Jin Liu 0016, Zhen Xia, Peipei Yuan |
SEKE | 1 |
| 2016 | The Impact of Feature Selection on Defect Prediction Performance: An Empirical ComparisonabstractSoftware defect prediction aims to determine whether a software module is defect-prone by constructing prediction models. The performance of such models is susceptible to the high dimensionality of the datasets that may include irrelevant and redundant features. Feature selection is applied to alleviate this issue. Because many feature selection methods have been proposed, there is an imperative need to analyze and compare these methods. Prior empirical studies may have potential controversies and limitations, such as the contradictory results, usage of private datasets and inappropriate statistical test techniques. This observation leads us to conduct a careful empirical study to reinforce the confidence of the experimental conclusions by considering several potential source of bias, such as the noise in the dataset and the dataset types. In this paper, we investigate the impact of 32 feature selection methods on the defect prediction performance over two versions of the NASA dataset (i.e., the noisy and clean NASA datasets) and one open source AEEEM dataset. We use a state-of-the-art double Scott-Knott test technique to analyze these methods. Experimental results show that the effectiveness of these feature selection methods on defect prediction performance varies significantly over all the datasets. Zhou Xu 0003, Jin Liu 0016, Zijiang Yang 0006, Gege An, Xiangyang Jia |
ISSRE | 1 |
| 2016 | Developer Recommendation with Awareness of Accuracy and CostabstractAs the scale and complexity of software products increase, software maintenance on bug resolution has become a challenging work.In the process of software implementation, developers often use bug reports, source code and change history to help solve bugs.However, hundreds of bug reports are being submitted every day.It is time-consuming and effortless for developers to review all the bug reports.To facilitate the assignment of bug reports, existing developer recommendation systems typically recommend the developer who has the fullest potential.However, bug reports are highly varied; time that the developers may spend fixing them is also important.To address the problem of developer recommendation, we propose a developer recommendation system with awareness of accuracy and cost (DRAC).This recommendation system is based on modern portfolio theory by striking a balance between accuracy and cost (time).We evaluate our approach with experiments on data collected from Bugzilla 1 . Jin Liu 0016, Yiqiuzi Tian, Liang Hong 0001, Xu Chen 0017, Zhou Xu 0003 |
SEKE | 5 |
| 2016 | MICHAC: Defect Prediction via Feature Selection Based on Maximal Information Coefficient with Hierarchical Agglomerative ClusteringabstractDefect prediction aims to estimate software reliability via learning from historical defect data. A defect prediction method identifies whether a software module is defect-prone or not according to metrics that are mined from software projects. These metric values, also known as features, may involve irrelevance and redundancy, which will hurt the performance of defect prediction methods. Existing work employs feature selection to preprocess defect data to filter out useless features. In this paper, we propose a novel feature selection framework, MICHAC, short for defect prediction via Maximal Information Coefficient with Hierarchical Agglomerative Clustering. MICHAC consists of two major stages. First, MICHAC employs maximal information coefficient to rank candidate features to filter out irrelevant ones, second, MICHAC groups features with hierarchical agglomerative clustering and selects one feature from each resulted group to remove redundant features. We evaluate our proposed method on 11 widelystudied NASA projects and four open-source AEEEM projects using three different classifiers with four performance metrics (precision, recall, F-measure, and AUC). Comparison with five existing methods demonstrates that MICHAC is effective in selecting features in defect prediction. Zhou Xu 0003, Jifeng Xuan, Jin Liu 0016, Xiaohui Cui |
SANER | 1 |