EDBT 2026 Demo / reviewers in the wild / expert
Yi Zhu 0008
dblp:67/4972-8
· DBLP profile ↗
22ranked-venue papers
4as first author
20since 2021 · last 2026
0000-0002-0996-0142ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 18 · 4 first-author · 16 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dynamic Feature Selection Based on Model Interpretation for Just-in-Time Software Defect PredictionabstractJust-In-Time Software Defect Prediction (JIT-SDP) aims to predict defects for each code change submitted by developers. Compared with traditional techniques, it offers advantages such as fine granularity, immediacy and traceability. However, most existing JIT-SDP models primarily focus on improving prediction performance, while research on model interpretability remains limited, especially with little attention given to time-series factor. Therefore, this study investigates the interpretability of JIT-SDP models in time-series scenario. It employs model interpretation techniques to analyze feature importance, explores the effectiveness of interpretability-based feature selection methods and proposes a SHAP-based Dynamic Feature Selection (SDFS) method that leverages model explanations in accordance with the temporal evolution of JIT-SDP. The experimental results demonstrate that certain features consistently show high importance across datasets, interpretability-based feature selection methods can enhance prediction performance and the proposed SDFS method effectively improves the performance of defect prediction models. Qiao Yu 0001, Yi Zhu 0008 |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2026 | RTCD4ADS: Runtime traffic rule conflict detection for autonomous driving system
Yi Zhu 0008, Junge Huang, Qiang Zhi, Yuxiao Zheng |
J. Syst. Softw. | 1 |
| 2025 | Chinese Relation Extraction Based on Global-Local Perception and Contrastive Learning
Chengyuan Cao, Yi Zhu 0008 |
WISA | 3 |
| 2025 | CAFEM: Context-Aware Feature Enhancement with Multi-Stage Sampling for Statement-Level Defect PredictionabstractSoftware defect prediction is one of the key technologies to ensure software quality. The current software quality assurance engineering needs more fine-grained defect prediction methods. In this paper, we propose a deep learning-based statement-level software defect prediction method, which integrates lexical analysis features and expert measurement features to accurately locate code defects. In terms of model architecture, a context-aware adaptive feature enhancement module (CAFEM) is proposed, which enhances feature representation capability through self-focusing mechanism and adaptive gating structure. To solve the problem of class imbalance in software defect prediction, an adaptive combination loss function based on binary crossentropy,Focal loss and F1 score was designed. SMOTE oversampling and TomekLinks undersampling were used to optimize the sample distribution. Experiments on a largescale benchmark dataset containing 119,989 code samples showed that our model establishes fine performance with an accuracy of 0.7315, a precision of 0.6364, a recall rate of 0.8941, and an F1 score of 0.7433, reaching the highest values on several metrics compared to existing methods. Yi Zhu 0008, Guosheng Hao |
QRS | 2 |
| 2025 | Software Defect Prediction Method Based on Multi-Feature FusionabstractSoftware defect prediction plays a vital role in software development. During the process of software updates and iterations, it helps developers anticipate potential defects in advance, thereby reducing unnecessary consumption of human, material, and time resources, while optimizing the allocation efficiency of Software Quality Assurance (SQA) resources. Given the widespread use of generative artificial intelligence in code development, effective software defect prediction has become increasingly important. Most previous studies primarily constructed defect prediction models using traditional metric-based features. However, in recent years, research focusing on semantic feature-based defect prediction has gained increasing attention, with researchers leveraging deep learning to automatically extract deep semantic information from source code. Nevertheless, existing approaches often rely on a single type of source code representation, overlooking the advantages and potential contributions of diverse features. To address this issue, this paper proposes a software defect prediction method based on multi-feature fusion, which incorporates multiple types of code representations to capture semantic information from different perspectives. The proposed model, DP-TACT, utilizes a multi-scale Convolutional Neural Network (Multiscale CNN) and Bidirectional Long Short-Term Memory (BiLSTM) network to process Abstract Syntax Tree (AST) and Full-token features. It also employs a Multi-head SelfAttention mechanism to capture critical information. Additionally, a Graph Convolutional Network (GCN) is used to process Control Flow Graph (CFG) features. Finally, traditional features are integrated to validate the effectiveness of the feature fusion strategy. The model is evaluated on the PROMISE dataset, and experimental results demonstrate that the proposed approach outperforms baseline models across multiple metrics. Yi Zhu 0008, Qiao Yu 0001 |
QRS | 2 |
| 2025 | Using composite attribute similarity multi-graph convolutional network for recommendation
Weichao He, Yi Zhu 0008, Yuheng Su, Guosheng Hao |
Appl. Intell. | 2 |
| 2025 | Class Imbalance-oriented Online Feature Selection Method for Just-in-time Software Defect PredictionabstractJust-in-time software defect prediction (JIT-SDP) is a defect prediction technique that targets changes in software code, offering significant advantages in quickly identifying potential defects and improving development efficiency. However, most existing methods assume that the importance of features remains stable over time, overlooking the dynamic changes in feature distributions and the evolution of class imbalance in real-world development environments. This limitation eventually degrades the predictive performance. To address this issue, this paper proposes an Imbalance-oriented Online Feature Selection (IOFS) method, which dynamically adjusts the feature importance and uncertainty parameters to adapt in real time to concept drift and class imbalance in data streams, thereby enhancing model performance and generalization. The experimental validation on 14 open-source project datasets demonstrates that IOFS significantly improves the values of [Formula: see text]Mean on 11 datasets and effectively reduces the average of the absolute differences between recalls for each time step, exhibiting robustness to dynamic feature changes and sensitivity to development-phase feature differences. This study provides an effective solution for online JIT-SDP. Qiao Yu 0001, Yi Zhu 0008 |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2025 | Behavioral decision-making and safety verification approaches for autonomous driving system in extreme scenarios
Yi Zhu 0008, Junge Huang, Qiang Zhi |
J. Syst. Softw. | 2 |
| 2025 | Spatio-Clock Synchronous Constraint Systems Specification and Verification to Ensure Autonomous Driving SafetyabstractABSTRACT Context Ensuring safety in autonomous driving systems is a major challenge, particularly in highly dynamic and complex environments. Traditional verification techniques struggle to capture intertwined spatial and temporal safety requirements in real‐time autonomous behaviors. Objective This study aims to provide a formal specification and verification framework that ensures safety‐critical spatio‐temporal interactions in autonomous driving, supporting rigorous analysis and preventing unsafe behaviors. Method We propose a spatio‐clock synchronous constraint framework, which includes: (1) a formal definition of spatio‐clock constraint trajectories; (2) the development of a specification language (SCSL) integrating RCC and CCSL; (3) the design of Spatio‐Clock Synchronous Automata (SCSA) for modeling driving behaviors; (4) formalization of safety guards and safe transitions; and (5) a verification process based on model checking. Result The framework is validated through a highway autonomous overtaking scenario. The case study demonstrates our approach's effectiveness to formally verify collision‐free guarantees under complex spatial‐temporal interactions and controller decision‐making logic. Conclusion The proposed framework enables precise and modular safety specification and verification for autonomous systems. It has practical value in supporting safety assurance during system design, especially in safety‐critical driving tasks under dynamic conditions. Jinyong Wang, Deyan Yang, Yi Zhu 0008 |
Softw. Pract. Exp. | 5 |
| 2025 | HPDA: An enhanced GNN-based software vulnerability detection approach by hybrid-scale perception and data augmentation
Shengyi Cheng, Qiao Yu 0001, Yi Zhu 0008, Zirui Huang |
Softw. Qual. J. | 3 |
| 2024 | Classification Method of Ethereum Smart Contracts Based on Statistical Model CheckingabstractThe integration of blockchain and smart contracts facilitates efficient and secure data exchange and value transfer. Nevertheless, the reliability of smart contracts has emerged as a critical concern. The vulnerabilities in contracts are intricately linked to their categories, emphasizing the significance of smart contract classification for enhancing code, user, and system security. Formal verification methods offer a robust means to validate the accuracy of contract classification and mitigate vulnerabilities. However, conventional machine learning approaches often lack precision and overlook the impact of account transaction behavior on classification during contract execution. This study introduces an Ethereum smart contract classification methodology based on statistical model detection. Through an examination of five smart contract types and ensuring the logical exclusivity of each, we delineate and formalize the internal logic of each contract type. We establish the contract automata network, devise conversion rules from contract source code to the automata network, and verify class properties using the statistical model detection tool UPPAAL-SMC. Lastly, we showcase the efficacy of our proposed methodology through a practical contract case. Miaoer Li, Yi Zhu 0008, Chan Yin |
QRS | 2 |
| 2024 | An Empirical Study of the Impact of Class Overlap on the Performance and Interpretability of Cross-Version Defect PredictionabstractThe class overlap problem refers to instances from different categories heavily overlapping in the feature space. This issue is one of the challenges in improving the performance of software defect prediction (SDP). Currently, the studies on the impact of class overlap on SDP mainly focused on within-project defect prediction and cross-project defect prediction. Moreover, the existing class overlap instances cleaning methods are not suitable for cross-version defect prediction. In this paper, we propose a class overlap instances cleaning method based on the Ratio of K-nearest neighbors with the Same Label (RKSL). This method removes instances with the abnormal neighbor ratio in the training set. Based on the RKSL method, we investigate the impact of class overlap on the performance and interpretability of the cross-version defect prediction model. The experiment results show that class overlap can affect the performance of cross-version defect prediction models significantly. The RKSL method can handle the class overlap problem in defect datasets, but it may impact the interpretability of models. Through the analysis of feature changes, we consider that class overlap instances cleaning can assist models in identifying more important features. Qiao Yu 0001, Yi Zhu 0008, Shengyi Cheng |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2024 | An Empirical Study on Model-Agnostic Techniques for Source Code-Based Defect PredictionabstractInterpretation is important for adopting software defect prediction in practice. Model-agnostic techniques such as Local Interpretable Model-agnostic Explanation (LIME) can help practitioners understand the factors which contribute to the prediction. They are effective and useful for models constructed on tabular data with traditional features. However, when they are applied on source code-based models, they cannot differentiate the contribution of code tokens in different locations for deep learning-based models with Bag-of-Word features. Besides, only using limited features as explanation may result in information loss about actual riskiness. Such limitations may lead to inaccurate explanation for source code-based models, and make model-agnostic techniques not useful and helpful as expected. Thus, we apply a perturbation-based approach Randomized Input Sampling Explanation (RISE) for source code-based defect prediction. Besides, to fill the gap that there lacks a systematical evaluation on model-agnostic techniques on source code-based defect models, we also conduct an extensive case study on the model-agnostic techniques on both token frequency-based and deep learning-based models. We find that (1) model-agnostic techniques are effective to identify the most important code tokens for an individual prediction and predict defective lines based on the importance scores, (2) using limited features (code tokens) for explanation may result in information loss about actual riskiness, and (3) RISE is more effective than others as it can generate more accurate explanation, achieve better cost-effectiveness for line-level prediction, and result in less information loss about actual riskiness. Based on such findings, we suggest that model-agnostic techniques can be a supplement to file-level source code-based defect models, while such explanations should be used with caution as actual risky tokens may be ignored. Also, compared with LIME, we would recommend RISE for a more effective explanation. Yi Zhu 0008, Yuxiang Gao, Qiao Yu 0001 |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2024 | Evolutionary measures and their correlations with the performance of cross-version defect prediction for object-oriented projectsabstractAbstract Cross‐version defect prediction (CVDP) for evolutionary projects has attracted much attention from researchers in recent years. For multiple versions of an object‐oriented project, the degree of evolution (e.g., the degree of class change) between successive versions can reflect the differences between versions, which could affect the performance of CVDP. Therefore, how to measure the degree of evolution between successive versions and explore the correlations with the performance of CVDP are very important for software defect prediction. Based on the successive versions of evolutionary projects, this paper proposes six evolutionary measures from three aspects of class change, metric change, and label change, including the Ratio of New Classes (RNC), the Ratio of Deleted Classes (RDC), the Average Ratio of Metric Change (ARMC), the Ratio of Label Changed Classes (RLCC), the Ratio of Unchanged Classes (RUC), and the Ratio of Interference Classes (RIC). An empirical study was conducted on 40 versions of 11 object‐oriented projects from the PROMISE repository. Precision, Recall, F‐measure, and AUC were used as the performance indicators. Three correlation approaches (Pearson, Spearman, and Kendall) are applied to show the correlations between evolutionary measures and the performance of CVDP. The statistical results show that RNC, RDC, and RUC show no correlation with four performance indicators. ARMC shows weak or medium positive correlations with Recall and F‐measure. RLCC and RIC show very strong or strong negative correlations with Recall and F‐measure. The results indicate that the correlations between the proposed evolutionary measures and the performance of CVDP are different, which can guide the training set selection of CVDP. Qiao Yu 0001, Yi Zhu 0008, Shujuan Jiang, Junyan Qian |
J. Softw. Evol. Process. | 2 |
| 2024 | A Cascade Domain Clustering Algorithm for Multiview DSM Fusion From Urban Satellite ImagesabstractMulti-view digital surface model (DSM) fusion has emerged as an important technique for three-dimensional (3D) reconstruction of multi-view satellite images. However, existing multi-view DSM fusion approaches are prone to the problems of blurred elevation divisions at the object edges, salt and pepper noises on smooth surfaces, and severe loss of surface details in weakly textured regions. In this paper, we present a cascade domain clustering (CDC) algorithm for fusing multi-view DSMs, which is realized by the combination of salient domain clustering and model domain clustering. Initially, the salient domain clustering is employed to demarcate prominent objects and identify regional edges using two-dimensional (2D) spectral and 3D elevation information. Subsequently, to further segment the intricate surface structures, particularly for objects with low texture attributes, we implement the model domain clustering to iteratively aggregate 3D points and fit geometric models corresponding to the aggregated clusters. Finally, the multi-view DSMs are fused iteratively through the weighted least squares (WLS) method, with model clusters serving as the fundamental units, under the constraints of the geometric models. Sufficient experiments show that the proposed CDC algorithm surpasses other popular multi-view DSM fusion algorithms in terms of completeness and accuracy, achieving 91.51% completeness and 0.93m RMSE, representing a 79.89% improvement in completeness and an 86.40% reduction in RMSE compared to the popular stereo 3D reconstruction pipeline. Yubo Men, Yi Zhu 0008, Chaoguang Men |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | The Change of Code Metrics for Predicting the Label Change on Evolutionary Projects: An Empirical Study
Qiao Yu 0001, Shujuan Jiang, Yi Zhu 0008, Yuanpeng Jiang |
WISA | 3 |
| 2022 | Evolutionary Measures for Object-oriented Projects and Impact on the Performance of Cross-version Defect PredictionabstractCross-version defect prediction (CVDP) has attracted more attention of researchers in recent years. For an evolutionary project, multiple versions will be produced during the process of software evolution. However, for multiple versions of an object-oriented project, the evolution degree (e.g. class change degree) between neighboring versions could affect the performance of CVDP. Therefore, how to measure the evolution degree of neighboring versions and explore the impact on the performance of CVDP are very important. Based on the neighboring versions of evolutionary projects, this paper proposed six evolutionary measures from three aspects of class change, metric change, and label change, including ratio of new classes (RNC), ratio of deleted classes (RDC), average ratio of metric change (ARMC), ratio of label changed classes (RLCC), ratio of unchanged classes (RUC), and ratio of interference classes (RIC). Spearman's rank correlation coefficient was applied to show the correlations between evolutionary measures and the performance of CVDP. An empirical study was conducted on 40 versions of 11 projects from the PROMISE repository. The performance of CVDP was evaluated with F-measure and AUC. The statistical results show that RNC, RDC, and RUC show no correlation with F-measure and AUC. ARMC shows a medium positive correlation with F-measure. RLCC and RIC show very strong or strong negative correlations with F-measure. The results indicate that the correlations between the proposed evolutionary measures and the performance of CVDP are different, which can guide the training set selection of CVDP. Qiao Yu 0001, Yi Zhu 0008, Shujuan Jiang, Junyan Qian |
Internetware | 2 |
| 2022 | Evaluating the effectiveness of local explanation methods on source code-based defect prediction modelsabstractInterpretation has been considered as one of key factors for applying defect prediction in practice. As one way for interpretation, local explanation methods has been widely used for certain predictions on datasets of traditional features. There are also attempts to use local explanation methods on source code-based defect prediction models, but unfortunately, it will get poor results. Since it is unclear how effective those local explanation methods are, we evaluate such methods with automatic metrics which focus on local faithfulness and explanation precision. Based on the results of experiments, we find that the effectiveness of local explanation methods depends on the adopted defect prediction models. They are effective on token frequency-based models, while they may not be effective enough to explain all predictions of deep learning-based models. Besides, we also find that the hyperparameter of local explanation methods should be carefully optimized to get more precise and meaningful explanation. Yuxiang Gao, Yi Zhu 0008, Qiao Yu 0001 |
MSR | 2 |
| 2022 | Statistical Model Checking for Stochastic and Hybrid Autonomous Driving Based on Spatio-Clock ConstraintsabstractAutonomous driving vehicles are a kind of typical cyber-physical systems integrating complex interactions between hardware and software components such as collaborative computation, distributed communication, and spatio-clock synchronous control with surrounding traffic environment. They can percept the environment, communicate with surroundings, and react fast enough to control independently. The purpose of autonomous driving emergence is to improve driving safety, reduce environmental pollution, and ease the traffic congestion. However, new features with surrounding open and dynamic environment make systems design and verification becoming more and more complex than ever, such as stochastic communication delay, hardware spontaneous failure distribution, and natively hybrid behaviors described by ordinary differential equations. Spatial and time collision avoidance remains crucial obstacles on the path to becoming ubiquitous and dependable. In this paper, we adopt statistical model checking (SMC) to enlighten possible hazards affected by stochastic and hybrid features in the design phase of autonomous driving systems. In order to provide safety and accountability, we first propose a dedicated multi-lane spatio-clock stochastic specification language (MLSCL) to describe safety invariants and guards in domain-specific autonomous driving systems. Then, we present the semantic mapping rules between MLSCL and UPPAAL SMC models, and design the spatio-clock stochastic and hybrid automata based on MLSCL in order to model inherently stochastic and hybrid behaviors. Finally, we present an illustrative lane-change case study to verify spatio-clock stochastic and hybrid-related properties adopting SMC, and demonstrate the effectiveness of our proposed approach. Jinyong Wang, Yi Zhu 0008, Guohua Shen |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2022 | Dealing with imbalanced data for interpretable defect predictionabstractContext Interpretation has been considered as a key factor to apply defect prediction in practice. As interpretation from rule-based interpretable models can provide insights about past defects with high quality, many prior studies attempt to construct interpretable models for both accurate prediction and comprehensible interpretation. However, class imbalance is usually ignored, which may bring huge negative impact on interpretation. Objective In this paper, we are going to investigate resampling techniques, a popular solution to deal with imbalanced data , on interpretation for interpretable models. We also investigate the feasibility to construct interpretable defect prediction models directly on original data. Further, we are going to propose a rule-based interpretable model which can deal with imbalanced data directly. Method We conduct an empirical study on 47 publicly available datasets to investigate the impact of resampling techniques on rule-based interpretable models and the feasibility to construct such models directly on original data. We also improve gain function and tolerate lower confidence based on rule induction algorithms to deal with imbalanced data. Results We find that (1) resampling techniques impact on interpretable models heavily from both feature importance and model complexity, (2) it is not feasible to construct meaningful interpretable models on original but imbalanced data due to low coverage of defects and poor performance, and (3) our proposed approach is effective to deal with imbalanced data compared with other rule-based models. Conclusion Imbalanced data heavily impacts on the interpretable defect prediction models. Resampling techniques tend to shift the learned concept, while constructing rule-based interpretable models on original data may also be infeasible. Thus, it is necessary to construct rule-based models which can deal with imbalanced data well in further studies. Yuxiang Gao, Yi Zhu 0008 |
Inf. Softw. Technol. | 2 |
| 2017 | Modeling and verification of Web services composition based on model transformationabstractSummary With the rapid development of Cloud computing, social computing, and Web of Things, an increasing number of requirements of complexity and reliability for modeling Web services composition have emerged too. As more reliable methods are needed to model and verify current complex Web services composition, this paper proposes a method to model and verify Web services composition based on model transformation. First, a modeling and verifying framework based on model transformation is established. Then, Communicating Sequential Process (CSP) is defined according to the features of Web services composition and the corresponding model checking tool Failure Divergence Refinement (FDR) is introduced. The transformation approaches between Business Process Execution Language (BPEL) and CSP are later defined in detail. Lastly, the effect of this method is evaluated by modeling and verifying the Web services composition of a Online Shopping System. The results of the experiments show that this method can greatly increase the reliability of Web services composition. Copyright © 2016 John Wiley & Sons, Ltd. Yi Zhu 0008, Hang Zhou 0002 |
Softw. Pract. Exp. | 1 |
| 2016 | Multi-Resource Modeling of Real-Time Software Based on Resource Timed Process AlgebraabstractWith the process of non-functional properties research on real-time systems, multi-resource estimation and analysis of real-time systems become a hotspot. Process algebra is a formal method that is fit for analyzing the functional properties of real-time systems, but it cannot analyze the multi-resource properties. Resource timed process algebra (RTPA) proposed in this paper can handle it efficiently by extending multi-resource semantics on timed communicating sequential process (TCSP). The multi-resource consumption of real-time instruction is mapped into the resource vector of RTPA, and then the multi-resource consumption of real-time systems can be modeled and optimized by using RTPA; the resource optimal checking algorithms are proposed to check the multi-resource satisfiability of instructions and calculate the minimum resource and accumulated resource of real-time systems. This formal method improves the accuracy and efficiency of multi-resource calculation, and the calculation results can be further used to quantitatively analyze and optimize the multi-resource consumption of real-time systems. Yi Zhu 0008, Guangquan Zhang 0002, Hang Zhou 0002, Fangxiong Xiao |
Int. J. Softw. Eng. Knowl. Eng. | 1 |