Xiaohui Wan

dblp:258/5160 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
7since 2021 · last 2026
0000-0001-6498-8570ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 7 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 ARFT-Transformer: Modeling metric dependencies for cross-project aging-related bug prediction
Shuning Ge, Fangyun Qin, Xiaohui Wan, Yang Liu 0287, Qian Dai, Zheng Zheng 0001
J. Syst. Softw.3
2025 Towards Fault-Tolerant Deep Reinforcement Learning Systems: A Framework Based on N-Version Programming
abstract
As the application of Deep Reinforcement Learning (DRL) systems gradually expands, improving their reliability becomes an important issue. Currently, improving the reliability of DRL systems primarily relies on testing techniques aimed at generating test cases to uncover as many bugs as possible. However, these techniques cannot guarantee the resolution of various bugs that may arise during the execution phase. Therefore, it is necessary to employ fault-tolerant techniques during the design phase as a means to enhance the reliability of DRL systems. This study focuses on developing a fault-tolerant framework for continuous DRL systems, based on the principles of N-Version Programming (NVP). Specifically, it investigates the critical design elements of DRL systems and proposes three types of independent factors, including DNN architecture, hyperparameter and training algorithm, to construct single version models. Subsequently, different mechanisms are explored to combine these models into the framework. Our empirical study conducted on three classical DRL tasks demonstrates that our proposed framework can achieve impressive fault tolerance and validates the feasibility of our proposed independent factors. To the best of our knowledge, this study first explores the fault-tolerant techniques in DRL systems, laying the groundwork for subsequent research on fault-tolerant frameworks in this field.
Yuhuan Fan, Xiaohui Wan, Zheng Zheng 0001
QRS3
2024 DRLFailureMonitor: A Dynamic Failure Monitoring Approach for Deep Reinforcement Learning System
abstract
Over the past decade, deep reinforcement learning (DRL) has seen increasing adoption in addressing various sequential decision-making tasks, such as autonomous driving and robotic control, demonstrating superior performance. However, as its application continues to broaden, the reliability of these DRL systems encounters significant challenges, particularly within safety-critical domains where any system failure could lead to catastrophic consequences. Currently, the assurance of reliability in DRL systems relies on testing techniques, which consist of offline solutions that uncover and address potential defects before deployment. In contrast to existing studies, this paper addresses the issue of online monitoring of DRL systems to detect and alert potential failures in advance, thereby facilitating the transition from automated to manual decision-making when necessary. Specifically, we model online monitoring of DRL systems as a multivariate time series classification problem and propose a novel failure monitoring approach, which is named DRLFailureMonitor. This method employs a temporal dynamic graph neural network to capture hidden spatiotemporal dependencies in the sequences of trajectories planned by the DRL systems. Extensive experiments across five benchmark DRL environments demonstrate that DRLFailureMonitor achieves an average failure detection accuracy of 98.2% and a recall of 100%. In addition, the proposed method offers a failure detection lead time ranging from 11 to 79 steps, indicating that it can detect failures in DRL systems well before the system actually fails and causes any losses. Consequently, this method holds significant importance for enhancing human-machine collaboration and improving the reliability of DRL systems in safety-critical areas.
Xiaohui Wan, Zheng Zheng 0001
ISSRE2
2024 Coverage-guided fuzzing for deep reinforcement learning systems
abstract
While the past decade has witnessed a growing demand for employing deep reinforcement learning (DRL) in various domains to solve real-world problems, the reliability of DRL systems has become more of a concern. In particular, DRL agents are often trained on data from a potentially biased distribution over environmental settings, causing the trained agents to fail in certain cases despite high average-case performance. Hence, it is necessary and urgent to adequately test DRL agents to ensure the reliability of practical DRL systems. However, due to the fundamental difference in the programming paradigm and the development process , traditional software testing methodology cannot be applied directly to DRL systems. Given that, we introduce a novel testing framework for DRL systems, aiming to generate diverse test cases that can drive a DRL system to fail. Specifically, we design, implement and evaluate DRLFuzz, which is a coverage-guided fuzzing (CGF) framework for systematically testing DRL systems. Experimental results demonstrate that DRLFuzz can efficiently discover diverse failures in different DRL systems for various benchmark tasks. Compared with a random search baseline, DRLFuzz can generate 60% more failed cases in general. Additionally, the diversity of failed cases generated by DRLFuzz is increased by 4 . 6 % ∼ 14 . 1 % in terms of mean pairwise distance (MPD). Furthermore, our experiments also indicate that the failed cases generated by DRLFuzz can be utilized to fine-tune the DRL agent to eliminate the failures resulting from inadequate exploration during training and thus improve the reliability of DRL systems.
Xiaohui Wan, Tiancheng Li 0005, Weibin Lin, Zheng Zheng 0001
J. Syst. Softw.1
2024 Data Complexity: A New Perspective for Analyzing the Difficulty of Defect Prediction Tasks
abstract
Defect prediction is crucial for software quality assurance and has been extensively researched over recent decades. However, prior studies rarely focus on data complexity in defect prediction tasks, and even less on understanding the difficulties of these tasks from the perspective of data complexity. In this article, we conduct an empirical study to estimate the hardness of over 33,000 instances, employing a set of measures to characterize the inherent difficulty of instances and the characteristics of defect datasets. Our findings indicate that: (1) instance hardness in both classes displays a right-skewed distribution, with the defective class exhibiting a more scattered distribution; (2) class overlap is the primary factor influencing instance hardness and can be characterized through feature, structural, and instance-level overlap; (3) no universal preprocessing technique is applicable to all datasets, and it may not consistently reduce data complexity, fortunately, dataset complexity measures can help identify suitable techniques for specific datasets; (4) integrating data complexity information into the learning process can enhance an algorithm’s learning capacity. In summary, this empirical study highlights the crucial role of data complexity in defect prediction tasks, and provides a novel perspective for advancing research in defect prediction techniques.
Xiaohui Wan, Zheng Zheng 0001, Fangyun Qin, Xuhui Lu
ACM Trans. Softw. Eng. Methodol.1
2024 Adjusted Trust Score: A Novel Approach for Estimating the Trustworthiness of Software Defect Prediction Models
abstract
Software defect prediction (SDP) techniques play a crucial role in identifying defective code regions and improving testing efficiency. Over recent decades, a plethora of SDP approaches has emerged, with machine learning (ML) models being the most widely employed. Despite their superior predictive performance, their black-box nature and uncertainties make it challenging for developers to trust their predictions. To address this issue, we propose a novel trustworthiness score, the adjusted trust score (ATS), which helps determine when to rely on classifier predictions. Furthermore, we employ ATS to develop a reject option for SDP models. Comprehensive experiments on 32 benchmark datasets and six prevalent ML classifiers reveal that high (low) ATS values successfully yield high precision in identifying correct (or incorrect) predictions. ATS also demonstrates superiority over its counterparts, as evidenced by the Wilcoxon signed-rank test. Furthermore, a comparative analysis of prediction performance, with and without a reject option, confirms the feasibility of designing a reject option for SDP models utilizing ATS. Our work highlights that ATS can assist developers in better comprehending the strengths and weaknesses of SDP models. Therefore, it is an essential component for guaranteeing trust from developers and deserves further investigation.
Xiaohui Wan, Zheng Zheng 0001, Fangyun Qin, Xuhui Lu, Kun Qiu 0001
IEEE Trans. Reliab.1
2022 SPE$^{2}$: Self-Paced Ensemble of Ensembles for Software Defect Prediction
abstract
Software defect prediction aims to predict defect-prone code regions automatically before defects are discovered. Accurate prediction helps software practitioners to prioritize their testing efforts. In recent decades, dozens of approaches have been put forward and acquired good results in this field. However, in practical scenarios, many projects have limited labeled instances; more than that, most of these labeled instances are nondefective. The lack of training data and class imbalance problem together bring serious challenges to software defect prediction tasks. So far, few of prevailing approaches can well handle these two difficulties simultaneously. One important reason is that they do not pay adequate attention to several key instances, which are difficult to classify in a small imbalanced dataset. This article introduces the concept of “instance hardness” to integrate various difficulties of imbalance classification tasks. Based on it, a novel imbalance learning framework named self-paced ensemble of ensembles (SPE$^{2}$) is proposed to perform software defect prediction. SPE$^{2}$aims to generate a strong ensemble of ensembles by self-paced harmonizing instance hardness via undersampling. Finally, SPE$^{2}$is extensively compared with eight imbalance learning approaches on ten open-source defect datasets. Experiments indicate that SPE$^{2}$improves the performance and achieves better and more significant F-measure values than its existing counterparts, based on Brunner’s statistical significance test and Cliff’s effect sizes.
Xiaohui Wan, Zheng Zheng 0001, Yang Liu 0287
IEEE Trans. Reliab.1
2020 An empirical study of factors affecting cross-project aging-related bug prediction with TLAP
Fangyun Qin, Xiaohui Wan, Beibei Yin
Softw. Qual. J.2
2019 Supervised Representation Learning Approach for Cross-Project Aging-Related Bug Prediction
abstract
Software aging, which is caused by Aging-Related Bugs (ARBs), tends to occur in long-running systems and may lead to performance degradation and increasing failure rate during software execution. ARB prediction can help developers discover and remove ARBs, thus alleviating the impact of software aging. However, ARB-prone files occupy a small percentage of all the analyzed files. It is usually difficult to gather sufficient ARB data within a project. To overcome the limited availability of training data, several researchers have recently developed cross-project models for ARB prediction. A key point for cross-project models is to learn a good representation for instances in different projects. Nevertheless, most of the previous approaches neither consider the reconstruction property of new representation nor encode source samples' label information in learning representation. To address these shortcomings, we propose a Supervised Representation Learning Approach (SRLA), which is based on double encoding-layer autoencoder, to perform cross-project ARB prediction. Moreover, we present a transfer cross-validation framework to select the hyper-parameters of cross-project models. Experiments on three large open-source projects demonstrate the effectiveness and superiority of our approach compared with the state-of-the-art approach TLAP.
Xiaohui Wan, Zheng Zheng 0001, Fangyun Qin, Kishor S. Trivedi
ISSRE1