VLDB 2026 Research / reviewers in the wild / expert
Quanyi Zou
dblp:277/1149
· DBLP profile ↗
10ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0001-6543-2224ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Counterfactual Contrastive Explanations for Software Defect Prediction: Toward Better Model Understanding and AccuracyabstractSoftware defect prediction (SDP) aims to identify potentially defective modules early in the development phase, thereby enhancing testing and improving the overall quality of software. Deep learning has advanced SDP by improving accuracy. However, its lack of interpretability remains a critical limitation. Although some existing methods offer explanations for SDP models, they typically focus on assigning feature importance within model decisions. However, these explanations often remain superficial, failing to clarify the specific roles that features play in the decision-making process. To address these challenges, this article presents counterfactual contrastive explanations for SDP (CCE-SDP), a novel framework that generates counterfactual contrastive explanations to enhance model transparency. Leveraging genetic algorithms, CCE-SDP constructs optimal counterfactual examples to reveal how specific changes in software metrics influence predictions, offering fine-grained, instance-level insights. The core of the CCE-SDP method is to offer counterfactual explanations that minimize the number of altered features while maximizing the counterfactual trust score. In addition, by incorporating synthetic counterfactual samples with defect labels into the training process, the ability to handle imbalanced data is also enhanced. Experiments on 36 benchmark software projects showed that the CCE-SDP model demonstrated to users how specific feature changes impact predictive outcomes, thereby allowing for a better understanding of the model’s decision-making. Furthermore, by embedding counterfactual analysis within the model training, we successfully boosted the predictive performance of the SDP models. Quanyi Zou, Zhanyu Yang, Xuan-Rui Qiu, Jia-Hong Yu, Yue-Yue Shi |
IEEE Trans. Reliab. | 1 |
| 2024 | Ensemble Kernel-Mapping-Based Ranking Support Vector Machine for Software Defect PredictionabstractRank-oriented software defect prediction (ROSDP) aims to establish a model to predict the testing priority of software modules according to defect severity for the reasonable allocation of limited testing resources. Some ROSDP methods construct the prediction model by a linear model with respect to software features. However, in software repositories, the linear condition between testing priority and software feature is not satisfied, and the ranking performance of the linear prediction model is limited. Thus, in order to relax the limitation of the linear prediction model and improve the ranking performance, ensemble kernel-mapping-based ranking support vector machine (EKMRSVM) is developed based on the theories of ranking SVM, which builds a nonlinear ranking function approximated by the kernel-mapping-based method. Furthermore, the sequential minimal optimization algorithm is developed to derive the ideal parameters of the nonlinear ranking function, and ensemble learning is introduced to reduce time costs and guarantee ranking performance. Experimental results on 20 open source datasets indicate that introducing the kernel mapping method in EKMRSVM is very effective in performance improvement, and ensemble learning makes the proposed ranking algorithm very competitive in terms of time costs. Thus, based on the comparative results of some baseline methods, EKMRSVM with the appropriate kernel function can achieve better ranking performance. Zhanyu Yang, Lu Lu 0011, Quanyi Zou |
IEEE Trans. Reliab. | 3 |
| 2023 | A Software Defect Prediction Method based on Multi-type Features and Feature SelectionabstractNumerous software defect prediction methods utilize semantic information and software metrics as code features, neglecting the structural knowledge inherent in the source code.Other studies improve feature completeness by simply combining different types of defect indicators, which causes information redundancy.To address these challenges, this paper proposes a novel software defect prediction method that incorporates multitype features and performs feature selection.Firstly, semantic and structural features are extracted by Text Convolutional Neural Network (TextCNN) and Graph Isomorphism Network (GIN) from Abstract Syntax Tree (AST) and Program Dependency Graph (PDG), respectively, which are combined with software metrics to build a multi-type feature set.Then, Recursive Feature Elimination with Cross-Validation (RFECV) integrating a novel feature importance measure is utilized to remove redundant features and generate a feature subset.Finally, a prediction model for classification is established based on the feature subset.The experiments validated the effectiveness of multi-type features and the improved RFECV.Overall our proposed method outperforms state-of-the-art techniques on nine Java open-source projects. Lu Lu 0011, Quanyi Zou, Zhanyu Yang |
SEKE | 4 |
| 2023 | Software Defect Prediction via Positional Hierarchical Attention Network (S)abstractSoftware Defect Prediction (SDP) aims to identify defect-prone modules in advance to ensure software quality.In SDP research based on deep learning, the mainstream approach is to extract deep semantic features from an Abstract Syntax Tree (AST).Theoretically, the AST as a bi-dimensional structure encloses information at the node level, fragment level, and entire tree level.However, most existing research serializes the whole AST without considering the expression at different granularities.To address this limitation, we introduce a positional hierarchical attention network (PHAN) that acquires semantic features by simultaneously considering contexts between nodes and paths.Specifically, our model incorporates attention mechanisms to capture information of varying importance at separate hierarchies, and relative position representations to distinguish the contributions of different paths.Experimental results demonstrate that PHAN significantly outperforms existing baseline methods. Xinyan Yi, Lu Lu 0011, Quanyi Zou, Zhanyu Yang |
SEKE | 4 |
| 2022 | Two-Stage AST Encoding for Software Defect PredictionabstractSoftware defect prediction (SDP) can find potential containing defect modules, which assists software developers in allocating limited test resources more efficiently.Because traditional software features fail to capture the semantics of source code, various studies have turned to extracting deep learning features.Existing related approaches often parse the program source code into Abstract Syntax Trees (ASTs) for further processing.However, most of these approaches ignore AST nodes' hierarchical and position-sensitive structure.To overcome the aforementioned issues, a two-stage AST encoding (TSE) method is proposed in this paper for software defect prediction.Experiments on eight Java open-source projects showed that our proposed SDP method outperforms several traditional methods and state-of-the-art deep learning methods in terms of F-measure and MCC. Yanwu Zhou, Lu Lu 0011, Quanyi Zou, Cuixu Li |
SEKE | 3 |
| 2021 | Multi-source Cross Project Defect Prediction with Joint Wasserstein Distance and Ensemble LearningabstractCross-Project Defect Prediction (CPDP) refers to transferring knowledge from source software projects to a target software project. Previous research has shown that the impacts of knowledge transferred from different source projects differ on the target task. Therefore, one of the fundamental challenges in CPDP is how to measure the amount of knowledge transferred from each source project to the target task. This article proposed a novel CPDP method called Multi-source defect prediction with Joint Wasserstein Distance and Ensemble Learning (MJWDEL) to learn transferred weights for evaluating the importance of each source project to the target task. In particular, first of all, applying the TCA technique and Logistic Regression (LR) train a sub-model for each source project and the target project. Moreover, the article designs joint Wassertein distance to understand the source-target relationship and then uses this as a basis to compute the transferred weights of different sub-models. After that, the transferred weights can be used to reweight these sub-models to determine their importance in knowledge transfer to the target task. We conducted experiments on 19 software projects from PROMISE, NASA and AEEEM datasets. Compared with several state-of-the-art CPDP methods, the proposed method substantially improves CPDP performance in terms of four evaluation indicators (i.e., F-measure, Balance, G-measure and MMC). Quanyi Zou, Lu Lu 0011, Zhanyu Yang |
ISSRE | 1 |
| 2021 | Correlation feature and instance weights transfer learning for cross project software defect predictionabstractAbstract Due to the differentiation between training and testing data in the feature space, cross‐project defect prediction (CPDP) remains unaddressed within the field of traditional machine learning. Recently, transfer learning has become a research hot‐spot for building classifiers in the target domain using the data from the related source domains. To implement better CPDP models, recent studies focus on either feature transferring or instance transferring to weaken the impact of irrelevant cross‐project data. Instead, this work proposes a dual weighting mechanism to aid the learning process, considering both feature transferring and instance transferring. In our method, a local data gravitation between source and target domains determines instance weight, while features that are highly correlated with the learning task, uncorrelated with other features and minimizing the difference between the domains are rewarded with a higher feature weight. Experiments on 25 real‐world datasets indicate that the proposed approach outperforms the existing CPDP methods in most cases. By assigning weights based on the different contribution of features and instances to the predictor, the proposed approach is able to build a better CPDP model and demonstrates substantial improvements over the state‐of‐the‐art CPDP models. Quanyi Zou, Lu Lu 0011, Shaojian Qiu, Xiaowei Gu 0002 |
IET Softw. | 1 |
| 2021 | Joint feature representation learning and progressive distribution matching for cross-project defect prediction
Quanyi Zou, Lu Lu 0011, Zhanyu Yang, Xiaowei Gu 0002, Shaojian Qiu |
Inf. Softw. Technol. | 1 |
| 2020 | Sentiment key frame extraction in user-generated micro-videos via low-rank and sparse representation
Xiaowei Gu 0002, Lu Lu 0011, Shaojian Qiu, Quanyi Zou, Zhanyu Yang |
Neurocomputing | 4 |
| 2018 | The memory degradation based online sequential extreme learning machine
Quanyi Zou, Xiaojun Wang 0004, Changjun Zhou, Qiang Zhang 0008 |
Neurocomputing | 1 |