Qiaoling Cao

dblp:346/5342 · DBLP profile ↗
← Back
6ranked-venue papers
0as first author
6since 2021 · last 2026
0009-0003-9245-2716ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 4 · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Software defect prediction based on graph code semantics
Hongwei Tao, Zhenhao Geng, Xiaoxu Niu, Qiaoling Cao
Expert Syst. Appl.5
2025 Prediction of incompatible bug numbers between versions of java open-source software based on deep fusion features
Xiaoxu Niu, Hongwei Tao, Qiaoling Cao, Jianxun Wang 0010, Zhenhao Geng
Expert Syst. Appl.3
2025 Software aging oriented trustworthiness measurement based on weighted Boltzmann entropy
Hongwei Tao, Han Liu 0012, Xiaoxu Niu, Licheng Ding, Yixiang Chen 0001, Qiaoling Cao
Inf. Softw. Technol.6
2024 Software Defect Prediction Method Based on Clustering Ensemble Learning
abstract
The technique of software defect prediction aims to assess and predict potential defects in software projects and has made significant progress in recent years within software development. In previous studies, this technique largely relied on supervised learning methods, requiring a substantial amount of labeled historical defect data to train the models. However, obtaining these labeled data often demands significant time and resources. In contrast, software defect prediction based on unsupervised learning does not depend on known labeled data, eliminating the need for large‐scale data labeling, thereby saving considerable time and resources while providing a more flexible solution for ensuring software quality. This paper conducts software defect prediction using unsupervised learning methods on data from 16 projects across two public datasets (PROMISE and NASA). During the feature selection step, a chi‐squared sparse feature selection method is proposed. This feature selection strategy combines chi‐squared tests with sparse principal component analysis (SPCA). Specifically, the chi‐squared test is first used to filter out the most statistically significant features, and then the SPCA is applied to reduce the dimensionality of these significant features. In the clustering step, the dot product matrix and Pearson correlation coefficient (PCC) matrix are used to construct weighted adjacency matrices, and a clustering overlap method is proposed. This method integrates spectral clustering, Newman clustering, fluid clustering, and Clauset–Newman–Moore (CNM) clustering through ensemble learning. Experimental results indicate that, in the absence of labeled data, using the chi‐squared sparse method for feature selection demonstrates superior performance, and the proposed clustering overlap method outperforms or is comparable to the effectiveness of the four baseline clustering methods.
Hongwei Tao, Qiaoling Cao, Xiaoxu Niu, Zhenhao Geng, Songtao Shang
IET Softw.2
2024 Cross-Project Defect Prediction Using Transfer Learning with Long Short-Term Memory Networks
abstract
With the increasing number of software projects, within‐project defect prediction (WPDP) has already been unable to meet the demand, and cross‐project defect prediction (CPDP) is playing an increasingly significant role in the area of software engineering. The classic CPDP methods mainly concentrated on applying metric features to predict defects. However, these approaches failed to consider the rich semantic information, which usually contains the relationship between software defects and context. Since traditional methods are unable to exploit this characteristic, their performance is often unsatisfactory. In this paper, a transfer long short‐term memory (TLSTM) network model is first proposed. Transfer semantic features are extracted by adding a transfer learning algorithm to the long short‐term memory (LSTM) network. Then, the traditional metric features and semantic features are combined for CPDP. First, the abstract syntax trees (AST) are generated based on the source codes. Second, the AST node contents are converted into integer vectors as inputs to the TLSTM model. Then, the semantic features of the program can be extracted by TLSTM. On the other hand, transferable metric features are extracted by transfer component analysis (TCA). Finally, the semantic features and metric features are combined and input into the logical regression (LR) classifier for training. The presented TLSTM model performs better on the f ‐measure indicator than other machine and deep learning models, according to the outcomes of several open‐source projects of the PROMISE repository. The TLSTM model built with a single feature achieves 0.7% and 2.1% improvement on Log4j‐1.2 and Xalan‐2.7, respectively. When using combined features to train the prediction model, we call this model a transfer long short‐term memory for defect prediction (DPTLSTM). DPTLSTM achieves a 2.9% and 5% improvement on Synapse‐1.2 and Xerces‐1.4.4, respectively. Both prove the superiority of the proposed model on the CPDP task. This is because LSTM capture long‐term dependencies in sequence data and extract features that contain source code structure and context information. It can be concluded that: (1) the TLSTM model has the advantage of preserving information, which can better retain the semantic features related to software defects; (2) compared with the CPDP model trained with traditional metric features, the performance of the model can validly enhance by combining semantic features and metric features.
Hongwei Tao, Lianyou Fu, Qiaoling Cao, Xiaoxu Niu, Songtao Shang, Yang Xian
IET Softw.3
2024 A comparative study of software defect binomial classification prediction models based on machine learning
Hongwei Tao, Xiaoxu Niu, Lianyou Fu, Qiaoling Cao, Songtao Shang, Yang Xian
Softw. Qual. J.5