Chun Ying Zhou

dblp:313/5982 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
3since 2021 · last 2022
0009-0008-6510-6638ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2022 Data Selection for Cross-Project Defect Prediction with Local and Global Features of Source Code
abstract
An open challenge for cross-project defect prediction (CPDP) is how to select the most appropriate training data for target project to build quality predictor.To our knowledge, existing methods are mostly dominated by traditional hand-crafted features, which do not fully encode the global structure between codes nor the semantics of code tokens.This work is to propose an improved method which is capable of automatically learning features for representing source code, and uses these feataures for training data selection.First, we propose a framework ALGoF to automatically learn the local semantic and global structural features of code files.Then, we analyze the feasibility of the learned features for data selection.Besides, we also validate the effectiveness of ALGoF by comparing with the traditional method.The experiments have been conducted on six defect datasets available at the PROMISE repository.The results show that ALGoF method helps to guide the training data selection for CPDP, and achieves a 48.31% improvement rate of F-measure.Meanwhile, our method has statistically significant advantages over the traditional method, especially when using both the local semantic and global structural features as the representation of code files.The maximum improvement of F-measure can reach 42.6%.
Chun Ying Zhou
SEKE3
2022 An Exploratory Study of Bug Prioritization and Severity Prediction based on Source Code Features
abstract
Software systems generate a large number of bugs during their lifecycles.Managing and assigning these bug reports is a challenging task.Building prediction models for the priority or severity levels of bugs through bug reports can help developers prioritize highly urgent bugs.Traditional prediction models are based on the textual description information in bug reports.However, most of the description is little or no.According to the bug report, developers need to fix the corresponding source code files.If the corresponding source code file is a core module in a software system, the report is likely to have high-level assignment rights.Therefore, in this paper, we investigate the effect of using the source code file feature sets on classification performance.In addition, we evaluate the effect of different sampling methods on the data, namely SMOTE, RUS, SMOTEEN, Adaboost, and GAN.Extensive experiments were conducted on five open-source projects.The experimental results show that the source code file feature sets do not perform as well as the textual description features in bug reports.Besides, over-sampling methods do not alleviate the data imbalance problem in the case of insufficient data, while GAN performs best in the case of sufficient data.
Chun Ying Zhou
SEKE1
2021 GCN2defect : Graph Convolutional Networks for SMOTETomek-based Software Defect Prediction
abstract
With the introduction of network metrics into the field of software defect prediction, the dependency network of software modules is widely used. The network embedding models aim to represent nodes as low-dimensional vectors, thereby preserving the topological structure of the network. However, in software engineering, traditional network embedding models do not concern deep learning strategies, while recently, graph neural networks (GNNs) have been proved to be an effective deep learning framework for learning graph data. As a variant of GNN, graph convolution neural network (GCN) has achieved appealing results in node classification and link prediction. Inspired by the performance of GCN, we propose GCN2defect, which extends GCN to automatically learn to encode the software dependency network and ultimately improve software defect prediction. Specifically, we firstly construct a program's Class Dependency Network, and then use node2vec for embedded learning to obtain the structural features of the network automatically. After that, we combine the learned structural features with traditional software code features to initialize the attributes of nodes in the Class Dependency Network. Next, we feed the dependency network to GCN to get much deeper representation of the class. Meanwhile, to enhance the accuracy of prediction, we also employ the SMOTETomek sampling to solve the problem of data imbalance. Finally, we evaluate the proposed method on eight open-source programs and demonstrate that, on average, GCN2defect improves the state-of-the-art approach by 6.84% ~ 23.85% in terms of the F-measure.
Chun Ying Zhou, Shengkai Lv
ISSRE2