Zheng Zheng 0001

dblp:35/2837-1 · DBLP profile ↗
← Back
75ranked-venue papers
12as first author
37since 2021 · last 2026
0000-0001-7922-9067ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 44 · 3 first-author · 25 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 13 · 6 first-author · 4 since 2021Security and privacy · 9 · 2 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2026 FMMDP: failure monitoring approach for DNN-based Markov decision process
Weibin Lin, Chao Jing, Zheng Zheng 0001
Empir. Softw. Eng.5
2026 ARFT-Transformer: Modeling metric dependencies for cross-project aging-related bug prediction
Shuning Ge, Fangyun Qin, Xiaohui Wan, Yang Liu 0287, Qian Dai, Zheng Zheng 0001
J. Syst. Softw.6
2026 LGMT: Logic-Grounded Metamorphic Testing for evaluating the reasoning reliability of LLMs
Zenghui Zhou, Xiaoke Fang, Weibin Lin, Zheng Zheng 0001
Knowl. Based Syst.6
2026 PCART: Automated Repair of Python API Parameter Compatibility Issues
abstract
In modern software development, Python third-party libraries play a critical role, especially in fields like deep learning and scientific computing. However, API parameters in these libraries often change during evolution, leading to compatibility issues for client applications reliant on specific versions. Python’s flexible parameter-passing mechanism further complicates this, as different passing methods can result in different API compatibility. Currently, no tool can automatically detect and repair Python API parameter compatibility issues. To fill this gap, we introduce PCART, the first solution to fully automate the process of API extraction, code instrumentation, API mapping establishment, compatibility assessment, repair, and validation. PCART handles various types of Python API parameter compatibility issues, including parameter addition, removal, renaming, reordering, and the conversion of positional to keyword parameters. To evaluate PCART, we construct PCBENCH, a large-scale benchmark comprising 47,478 test cases mutated from 844 parameter-changed APIs across 33 popular Python libraries. Evaluation results demonstrate that PCART is both effective and efficient, significantly outperforming existing tools (MLCatchUp and Relancer) and the large language model ChatGPT (GPT-4o), achieving an F1-score of 96.51% in detecting API parameter compatibility issues and a repair precision of 91.97%. Further evaluation on 30 real-world Python projects from GitHub confirms PCART’s practicality. We believe PCART can significantly reduce the time programmers spend maintaining Python API updates and advance the automation of Python API compatibility issue repair.
Shuai Zhang 0046, Guanping Xiao, Jun Wang 0167, Huashan Lei, Gangqiang He, Yepang Liu 0001, Zheng Zheng 0001
IEEE Trans. Software Eng.7
2025 Towards Fault-Tolerant Deep Reinforcement Learning Systems: A Framework Based on N-Version Programming
abstract
As the application of Deep Reinforcement Learning (DRL) systems gradually expands, improving their reliability becomes an important issue. Currently, improving the reliability of DRL systems primarily relies on testing techniques aimed at generating test cases to uncover as many bugs as possible. However, these techniques cannot guarantee the resolution of various bugs that may arise during the execution phase. Therefore, it is necessary to employ fault-tolerant techniques during the design phase as a means to enhance the reliability of DRL systems. This study focuses on developing a fault-tolerant framework for continuous DRL systems, based on the principles of N-Version Programming (NVP). Specifically, it investigates the critical design elements of DRL systems and proposes three types of independent factors, including DNN architecture, hyperparameter and training algorithm, to construct single version models. Subsequently, different mechanisms are explored to combine these models into the framework. Our empirical study conducted on three classical DRL tasks demonstrates that our proposed framework can achieve impressive fault tolerance and validates the feasibility of our proposed independent factors. To the best of our knowledge, this study first explores the fault-tolerant techniques in DRL systems, laying the groundwork for subsequent research on fault-tolerant frameworks in this field.
Yuhuan Fan, Xiaohui Wan, Zheng Zheng 0001
QRS4
2025 Domain generalization for zero-calibration brain-computer interfaces with knowledge distillation-based phase invariant feature extraction
Zilin Liang, Zheng Zheng 0001, Weihai Chen, Xinzhi Ma, Zhongcai Pei, Xiantao Sun
Eng. Appl. Artif. Intell.2
2025 KPAMA: A Kubernetes based tool for Mitigating ML system Aging
Xuhui Lu, Xiaoting Du, Zheng Zheng 0001
J. Syst. Softw.5
2025 Evaluating the effectiveness of neuron coverage metrics: a metamorphic-testing approach
Zenghui Zhou, Pak-Lok Poon, Tsong Yueh Chen, Kun Qiu 0001, Zheng Zheng 0001
Softw. Qual. J.6
2025 Multigranularity Coverage Criteria for Deep Learning Libraries
abstract
ABSTRACT Deep learning (DL) systems are becoming increasingly widely used in safety domains such as self‐driving cars and unmanned aerial vehicles, which arouse natural concerns about their trustworthiness. Underlying DL libraries used in the construction and execution of DL models are involved in the testing processes of DL systems. Therefore, bugs in DL libraries can inevitably cause unexpected behaviours in DL systems. The internal structures of DL libraries are described as APIs with different functionalities, and DL libraries offer model developers access to DL techniques with various API parameter settings. The above characteristics of DL libraries reveal that existing DL coverage criteria are not designed specifically for DL libraries, and traditional software coverage criteria do not apply to DL libraries either. The paper introduces the first set of coverage criteria specifically designed for the systematic measurement of DL libraries across various granularities. APIs, as the fundamental components of DL libraries, are used to define coverage criteria gauging testing adequacy by thoroughly considering their invocation, implementation, parameter quantities and parameter attributes. Furthermore, some properties depicting relations between coverage criteria are investigated. Experiments on the effectiveness of the proposed coverage criteria and comparative analysis are conducted by interval estimate and hypothesis testing techniques for APIs in two well‐known DL libraries. The experimental results demonstrate that the proposed coverage criteria are effective in measuring the test adequacy of DL libraries, and they can be used for the quantitative analysis of test model quality in DL libraries.
Zheng Zheng 0001, Beibei Yin, Zhiyu Xi
Softw. Test. Verification Reliab.2
2025 Metamorphic Relation Generation: State of the Art and Research Directions
abstract
Metamorphic testing has become one mainstream technique to address the notorious oracle problem in software testing, thanks to its great successes in revealing real-life bugs in a wide variety of software systems. Metamorphic relations, the core component of metamorphic testing, have continuously attracted research interests from both academia and industry. In the last decade, a rapidly increasing number of studies have been conducted to systematically generate metamorphic relations from various sources and for different application domains. In this article, based on the systematic review on the state of the art for metamorphic relations’ generation, we summarize and highlight visions for further advancing the theory and techniques for identifying and constructing metamorphic relations and discuss promising research directions in related areas.
Rui Li 0013, Huai Liu, Pak-Lok Poon, Dave Towey, Chang-Ai Sun, Zheng Zheng 0001, Zhiquan Zhou 0001, Tsong Yueh Chen
ACM Trans. Softw. Eng. Methodol.6
2025 DRLMutation: A Comprehensive Framework for Mutation Testing in Deep Reinforcement Learning Systems
abstract
Deep reinforcement learning systems have been increasingly applied in various domains. Testing them, however, remains a major open research problem. Mutation testing is a popular test suite evaluation technique that analyzes the extent to which test suites detect injected faults. It has been widely researched in both traditional software and the field of deep learning. However, due to the fundamental differences between deep reinforcement learning systems and traditional software, as well as deep learning systems, in aspects such as environment interaction, network decision-making, and data efficiency, previous mutation testing techniques cannot be directly applied to deep reinforcement learning systems. In this article, we proposed a comprehensive mutation testing framework specifically designed for deep reinforcement learning systems, DRLMutation , to further fill this gap. We first considered the characteristics of deep reinforcement learning, and based on both the training process and the model of trained agent, examined combinations from three dimensions: objects, operation methods, and injection methods. This approach led to a more comprehensive design methodology for deep reinforcement learning mutation operators. After filtering, we identified a total of 107 applicable deep reinforcement learning mutation operators. Then, in the realm of evaluation, we formulated a set of metrics tailored to assess test suites. Finally, we validated the stealthiness and effectiveness of the proposed mutation operators in the Cart Pole , Mountain Car Continuous , Lunar Lander , Breakout , and CARLA environments. We show inspiring findings that the majority of these designed deep reinforcement learning mutation operators potentially undermine the decision-making capabilities of the agent without affecting normal training. The varying degrees of disruption achieved by these mutation operators can be used to assess the quality of different test suites.
Zheng Zheng 0001, Xiaoting Du, Haoyu Wang 0001
ACM Trans. Softw. Eng. Methodol.2
2025 Identifying the Failure-Revealing Test Cases in Metamorphic Testing: A Statistical Approach
abstract
Metamorphic testing, thanks to its high failure-detection effectiveness especially in the absence of test oracle, has been widely applied in both the traditional context of software testing and other relevant fields such as fault localization and program repair. Its core element is a set of metamorphic relations, which are the necessary properties of the target algorithm in the form of the relationships among multiple inputs and corresponding expected outputs. When a relation is violated by the outputs of a group of test cases, namely metamorphic group of test cases, that are constructed based on the relation, a failure is said to be revealed. Traditionally, the primary task of software testing is to reveal failures. Therefore, from the perspective of software testing, it may not need to know which test case(s) in the metamorphic group cause the violation and thus the failure. However, such information is definitely helpful for other software engineering activities, such as software debugging. The current literature of metamorphic testing lacks a systematic mechanism of identifying the actual failure-revealing test cases, which hinders its applicability and effectiveness in other relevant fields. In this article, we propose a new technique for the FAILure-revealing Test case Identification in Metamorphic testing, namely FAILTIM. The approach is based on a novel application of statistical methods. More specifically, we leverage and adapt the basic ideas of spectrum-based techniques, which are originally used in fault localization, and propose the utilization of a set of risk formulas to estimate the suspiciousness of each individual test case in metamorphic groups. Failure-revealing test cases are then suggested according to their suspiciousness. A series of experiments have been conducted to evaluate the effectiveness and efficiency of FAILTIM using 9 subject programs and 30 risk formulas. The experimental results showed that the new approach can achieve a high accuracy in identifying the actual failure-revealing test cases in metamorphic testing. Consequently, our study will help boost the applicability and performance of metamorphic testing beyond testing to other software engineering areas. The present work also unfolds a number of research directions for further advancing the theory of metamorphic testing and more broadly, software testing.
Zheng Zheng 0001, Dai-Xu Ren, Huai Liu, Tsong Yueh Chen, Tiancheng Li 0005
ACM Trans. Softw. Eng. Methodol.1
2025 Functional and Defect Study in Deep Learning Libraries: A Complex Network Perspective
abstract
Deep learning libraries have emerged as critical software systems for a diverse range of deep learning applications. However, the stability and reliability of these libraries are increasingly challenged by their rapidly expanding scale. In this article, we utilize the function call graph as a model to represent deep learning libraries and conduct an empirical study using an innovative approach based on complex network theory. This method facilitates a thorough exploration of the topological characteristics and functionalities of deep learning libraries, revealing their scale-free and small-world properties. Leveraging the characteristics, we utilizek-core decomposition to pinpoint critical functions within the libraries, and conduct a comprehensive analysis to discern the characteristics of their functionalities. Furthermore, we have compiled a comprehensive dataset comprising 12 774 defective functions within these libraries. This dataset enables us to analyze and compare the distribution and trend of defects across the investigated deep learning libraries, while examining the patterns of defect propagation. Our research presents 14 significant findings, offering insights for researchers in software reliability and testing.
Xuhui Lu, Zheng Zheng 0001, Fangyun Qin, Xiangyue Ma
IEEE Trans. Reliab.2
2024 DRLFailureMonitor: A Dynamic Failure Monitoring Approach for Deep Reinforcement Learning System
abstract
Over the past decade, deep reinforcement learning (DRL) has seen increasing adoption in addressing various sequential decision-making tasks, such as autonomous driving and robotic control, demonstrating superior performance. However, as its application continues to broaden, the reliability of these DRL systems encounters significant challenges, particularly within safety-critical domains where any system failure could lead to catastrophic consequences. Currently, the assurance of reliability in DRL systems relies on testing techniques, which consist of offline solutions that uncover and address potential defects before deployment. In contrast to existing studies, this paper addresses the issue of online monitoring of DRL systems to detect and alert potential failures in advance, thereby facilitating the transition from automated to manual decision-making when necessary. Specifically, we model online monitoring of DRL systems as a multivariate time series classification problem and propose a novel failure monitoring approach, which is named DRLFailureMonitor. This method employs a temporal dynamic graph neural network to capture hidden spatiotemporal dependencies in the sequences of trajectories planned by the DRL systems. Extensive experiments across five benchmark DRL environments demonstrate that DRLFailureMonitor achieves an average failure detection accuracy of 98.2% and a recall of 100%. In addition, the proposed method offers a failure detection lead time ranging from 11 to 79 steps, indicating that it can detect failures in DRL systems well before the system actually fails and causes any losses. Consequently, this method holds significant importance for enhancing human-machine collaboration and improving the reliability of DRL systems in safety-critical areas.
Xiaohui Wan, Zheng Zheng 0001
ISSRE4
2024 Cross-project concurrency bug prediction using domain-adversarial neural network
Fangyun Qin, Zheng Zheng 0001, Yulei Sui, Siqian Gong, Zhi-Ping Shi 0002, Kishor S. Trivedi
J. Syst. Softw.2
2024 Multi-granularity coverage criteria for deep reinforcement learning systems
Beibei Yin, Zheng Zheng 0001
J. Syst. Softw.3
2024 Coverage-guided fuzzing for deep reinforcement learning systems
abstract
While the past decade has witnessed a growing demand for employing deep reinforcement learning (DRL) in various domains to solve real-world problems, the reliability of DRL systems has become more of a concern. In particular, DRL agents are often trained on data from a potentially biased distribution over environmental settings, causing the trained agents to fail in certain cases despite high average-case performance. Hence, it is necessary and urgent to adequately test DRL agents to ensure the reliability of practical DRL systems. However, due to the fundamental difference in the programming paradigm and the development process , traditional software testing methodology cannot be applied directly to DRL systems. Given that, we introduce a novel testing framework for DRL systems, aiming to generate diverse test cases that can drive a DRL system to fail. Specifically, we design, implement and evaluate DRLFuzz, which is a coverage-guided fuzzing (CGF) framework for systematically testing DRL systems. Experimental results demonstrate that DRLFuzz can efficiently discover diverse failures in different DRL systems for various benchmark tasks. Compared with a random search baseline, DRLFuzz can generate 60% more failed cases in general. Additionally, the diversity of failed cases generated by DRLFuzz is increased by 4 . 6 % ∼ 14 . 1 % in terms of mean pairwise distance (MPD). Furthermore, our experiments also indicate that the failed cases generated by DRLFuzz can be utilized to fine-tune the DRL agent to eliminate the failures resulting from inadequate exploration during training and thus improve the reliability of DRL systems.
Xiaohui Wan, Tiancheng Li 0005, Weibin Lin, Zheng Zheng 0001
J. Syst. Softw.5
2024 Deep semi-supervised learning for recovering traceability links between issues and commits
Jianfei Zhu, Guanping Xiao, Zheng Zheng 0001, Yulei Sui
J. Syst. Softw.3
2024 How About Bug-Triggering Paths? - Understanding and Characterizing Learning-Based Vulnerability Detectors
abstract
Machine learning and its promising branch deep learning have proven to be effective in a wide range of application domains. Recently, several efforts have shown success in applying deep learning techniques for automatic vulnerability discovery, as alternatives to traditional static bug detection. In principle, these learning-based approaches are built on top of classification models using supervised learning. Depending on the different granularities to detect vulnerabilities, these approaches rely on learning models which are typically trained with well-labeled source code to predict whether a program method, a program slice, or a particular code line contains a vulnerability or not. The effectiveness of these models is normally evaluated against conventional metrics including precision, recall and F1 score. In this paper, we show that despite yielding promising numbers, the above evaluation strategy can be insufficient and even misleading when evaluating the effectiveness of current learning-based approaches. This is because the underlying learning models only produce the classification results or report individual/isolated program statements, but are unable to pinpoint bug-triggering paths, which is an effective way for bug fixing and the main aim of static bug detection. Our key insight is that a program method or statement can only be stated as vulnerable in the context of a bug-triggering path. In this work, we systematically study the gap between recent learning-based approaches and conventional static bug detectors in terms of fine-grained metrics called BTP metrics using bug-triggering paths. We then characterize and compare the quality of the prediction results of existing learning-based detectors under different granularities. Finally, our comprehensive empirical study reveals several key issues and challenges in developing classification models to pinpoint bug-triggering paths and calls for more advanced learning-based bug detection techniques.
Xiao Cheng 0002, Xu Nie, Ningke Li, Haoyu Wang 0001, Zheng Zheng 0001, Yulei Sui
IEEE Trans. Dependable Secur. Comput.5
2024 Data Complexity: A New Perspective for Analyzing the Difficulty of Defect Prediction Tasks
abstract
Defect prediction is crucial for software quality assurance and has been extensively researched over recent decades. However, prior studies rarely focus on data complexity in defect prediction tasks, and even less on understanding the difficulties of these tasks from the perspective of data complexity. In this article, we conduct an empirical study to estimate the hardness of over 33,000 instances, employing a set of measures to characterize the inherent difficulty of instances and the characteristics of defect datasets. Our findings indicate that: (1) instance hardness in both classes displays a right-skewed distribution, with the defective class exhibiting a more scattered distribution; (2) class overlap is the primary factor influencing instance hardness and can be characterized through feature, structural, and instance-level overlap; (3) no universal preprocessing technique is applicable to all datasets, and it may not consistently reduce data complexity, fortunately, dataset complexity measures can help identify suitable techniques for specific datasets; (4) integrating data complexity information into the learning process can enhance an algorithm’s learning capacity. In summary, this empirical study highlights the crucial role of data complexity in defect prediction tasks, and provides a novel perspective for advancing research in defect prediction techniques.
Xiaohui Wan, Zheng Zheng 0001, Fangyun Qin, Xuhui Lu
ACM Trans. Softw. Eng. Methodol.2
2024 MET-MAPF: A Metamorphic Testing Approach for Multi-Agent Path Finding Algorithms
abstract
The Multi-Agent Path Finding (MAPF) problem, i.e., the scheduling of multiple agents to reach their destinations, has been widely investigated. Testing MAPF systems is challenging, due to the complexity and variety of scenarios and the agents’ distribution and interaction. Moreover, MAPF testing suffers from the oracle problem, i.e., it is not always clear whether a test shows a failure or not. Indeed, only considering whether the agents reach their destinations without collision is not sufficient. Other properties related to the ‘quality’ of the generated paths should be assessed, e.g., an agent should not follow an unnecessarily long path. To tackle this issue, this article proposes MET-MAPF, a Metamorphic Testing approach for MAPF systems. We identified 10 Metamorphic Relations (MRs) that a MAPF system should guarantee, designed over the environment in which agents operate, the behaviour of the single agents and the interactions among agents. Starting from the different MRs, MET-MAPF automatically generates test cases addressing them, so possibly exposing different types of failures. Experimental results show that MET-MAPF can indeed find MR violations not exposed by approaches that only consider the completion of the mission as test oracle. Moreover, experiments show that different MRs expose different types of violations.
Xiao-Yi Zhang 0005, Yang Liu 0287, Paolo Arcaini, Mingyue Jiang, Zheng Zheng 0001
ACM Trans. Softw. Eng. Methodol.5
2024 Adjusted Trust Score: A Novel Approach for Estimating the Trustworthiness of Software Defect Prediction Models
abstract
Software defect prediction (SDP) techniques play a crucial role in identifying defective code regions and improving testing efficiency. Over recent decades, a plethora of SDP approaches has emerged, with machine learning (ML) models being the most widely employed. Despite their superior predictive performance, their black-box nature and uncertainties make it challenging for developers to trust their predictions. To address this issue, we propose a novel trustworthiness score, the adjusted trust score (ATS), which helps determine when to rely on classifier predictions. Furthermore, we employ ATS to develop a reject option for SDP models. Comprehensive experiments on 32 benchmark datasets and six prevalent ML classifiers reveal that high (low) ATS values successfully yield high precision in identifying correct (or incorrect) predictions. ATS also demonstrates superiority over its counterparts, as evidenced by the Wilcoxon signed-rank test. Furthermore, a comparative analysis of prediction performance, with and without a reject option, confirms the feasibility of designing a reject option for SDP models utilizing ATS. Our work highlights that ATS can assist developers in better comprehending the strengths and weaknesses of SDP models. Therefore, it is an essential component for guaranteeing trust from developers and deserves further investigation.
Xiaohui Wan, Zheng Zheng 0001, Fangyun Qin, Xuhui Lu, Kun Qiu 0001
IEEE Trans. Reliab.2
2023 An Empirical Study to Identify Software Aging Indicators for Android OS
abstract
Android mobile devices have been suffering from performance degradation and increased failure rates during long-term operation, known as software aging. With the major changes in performance optimization and resource management in Android, it is spotted that the aging behavior of Android devices in the 2020s differs significantly from previous studies in resource utilization and performance metrics, which makes some classic metrics difficult to measure aging well, and new metrics are required to better describe the new phenomenon. Thus, we propose thread- and interface-level metrics to portray aging at a finer granularity and conduct an empirical study to reidentify classic and new software aging metrics in Android. Analysis confirms that software aging in Android is less reflected in global resources metrics but in more fine-grained ones, so thread- and interface-level metrics combined with specific classic resource and process-level metrics are helpful as indicators of software aging. These metrics have been confirmed and deployed for aging monitoring by our mobile phone manufacturer collaborators. A new experimental method customized for metric studies has also been adopted in this paper, significantly reducing data costs and interference in measurements.
Yulei Chen, Yuge Nie, Beibei Yin, Zheng Zheng 0001, Huayao Wu
QRS4
2023 Application of metamorphic testing on UAV path planning software
Lvyuan Wu, Zhiyu Xi, Zheng Zheng 0001
J. Syst. Softw.3
2023 Editorial: Software Reliability and Dependability Engineering
abstract
As software plays an increasingly important role in our lives, it is essential to maintain its reliability, and generally dependability. Software bugs can cause huge financial losses and dangerous accidents; the safety risks from software are underscored these days to even the non-technical public by the emergence of autonomous software-based systems. Thus, it is important to explore principled approaches to reduce the harm from defects in software, preferably by removing them as early as possible, but also by fault tolerance and by predicting their effects so as to inform mitigation actions.
Zheng Zheng 0001, Lorenzo Strigini, Nuno Antunes, Kishor S. Trivedi
IEEE Trans. Dependable Secur. Comput.1
2022 Taxonomy of Aging-related Bugs in Deep Learning Libraries
abstract
Deep learning libraries are the cornerstone of deep learning systems, and millions of deep learning applications are built on top of deep learning libraries. Due to long-term continuous running, many numerical operations and heavy dependence on resources, deep learning libraries are prone to the effects of software aging. Aging in deep learning libraries can threaten the reliability of deep learning systems and make training and application of deep learning more time-consuming and expensive, causing users to lose confidence in it. In this work, we manually screened 138 bug reports containing aging-related bugs from a total of 13,694 bug reports in four popular deep learning libraries (i.e., TensorFlow, MXNET, PaddlePaddle and MindSpore). We analyzed the information in these 138 bug reports to answer three questions: What categories of aging-related bugs exist in deep learning libraries? What is the distribution of different categories of aging-related bugs in deep learning libraries? Which deep learning phases are most susceptible to software aging? Finally, we conducted a fine-grained taxonomy of aging-related bugs, including four levels and seventeen categories, and obtained eight important findings with corresponding practical implications.
Xiaoting Du, Yanming Miao, Zheng Zheng 0001
ISSRE7
2022 Enhancing Traceability Link Recovery with Unlabeled Data
abstract
Traceability link recovery (TLR) is an important software engineering task for developing trustworthy and reliable software systems. Recently proposed deep learning (DL) models have shown their effectiveness compared to traditional information retrieval-based methods. DL often heavily relies on sufficient labeled data to train the model. However, manually labeling traceability links is time-consuming, labor-intensive, and requires specific knowledge from domain experts. As a result, typically only a small portion of labeled data is accompanied by a large amount of unlabeled data in real-world projects. Our hypothesis is that artifacts are semantically similar if they have the same linked artifact(s). This paper presents TRACEFUN, a new approach to enhance traceability link recovery with unlabeled data. TRACEFUN first measures the similarities between unlabeled and labeled artifacts using two similarity prediction methods (i.e., vector space model and contrastive learning). Then, based on the similarities, newly labeled links are generated between the unlabeled artifacts and the linked objects of the labeled artifacts. Generated links are further used for TLR model training. We have evaluated TRACEFUN on three GitHub projects with two state-of-the-art DL models (i.e., Trace BERT and TraceNN). The results show that TRACEFUN is effective in terms of a maximum improvement of F1-score up to 21% and 1,088%, respectively for Trace BERT and TraceNN.
Jianfei Zhu, Guanping Xiao, Zheng Zheng 0001, Yulei Sui
ISSRE3
2022 A residual convolutional neural network based approach for real-time path planning
Yang Liu 0287, Zheng Zheng 0001, Fangyun Qin, Xiao-Yi Zhang 0005, Haonan Yao
Knowl. Based Syst.2
2022 DeepSIM: Deep Semantic Information-Based Automatic Mandelbug Classification
abstract
Understanding and predicting types of bugs are of practical importance for developers to improve the testing efficiency and take appropriate steps to address bugs in software releases. However, due to the complex conditions under which faults manifest and the complexity of the classification rules, the automatic classification of Mandelbugs is a difficult task. In this article, we present a deep semantic information-based Mandelbug classification method that combines a semantic model with a deep learning classifier and makes use of both labeled and unlabeled bug reports. By training the bug report semantic model on millions of bug reports, each word in the text of a bug report is represented as a word embedding that preserves the semantic relationship among the words. Then, a convolutional neural network model is designed to capture the high-level features of bug reports to obtain a more accurate classification. Moreover, the effects of the semantic model size and domain on the classification results are investigated, and the quality of word embeddings is evaluated by analyzing several important parameters.
Xiaoting Du, Zheng Zheng 0001, Guanping Xiao, Zenghui Zhou, Kishor S. Trivedi
IEEE Trans. Reliab.2
2022 SPE$^{2}$: Self-Paced Ensemble of Ensembles for Software Defect Prediction
abstract
Software defect prediction aims to predict defect-prone code regions automatically before defects are discovered. Accurate prediction helps software practitioners to prioritize their testing efforts. In recent decades, dozens of approaches have been put forward and acquired good results in this field. However, in practical scenarios, many projects have limited labeled instances; more than that, most of these labeled instances are nondefective. The lack of training data and class imbalance problem together bring serious challenges to software defect prediction tasks. So far, few of prevailing approaches can well handle these two difficulties simultaneously. One important reason is that they do not pay adequate attention to several key instances, which are difficult to classify in a small imbalanced dataset. This article introduces the concept of “instance hardness” to integrate various difficulties of imbalance classification tasks. Based on it, a novel imbalance learning framework named self-paced ensemble of ensembles (SPE$^{2}$) is proposed to perform software defect prediction. SPE$^{2}$aims to generate a strong ensemble of ensembles by self-paced harmonizing instance hardness via undersampling. Finally, SPE$^{2}$is extensively compared with eight imbalance learning approaches on ten open-source defect datasets. Experiments indicate that SPE$^{2}$improves the performance and achieves better and more significant F-measure values than its existing counterparts, based on Brunner’s statistical significance test and Cliff’s effect sizes.
Xiaohui Wan, Zheng Zheng 0001, Yang Liu 0287
IEEE Trans. Reliab.2
2022 A Phase-Type Expansion Approach for the Performability of Composite Web Services
abstract
The purpose of web service composition is to satisfy complex requirements by constructing a composite service through loosely coupling web services over the Internet. Services available for composition may experience changes in performance and reliability at any time. Therefore, it is desirable to estimate performance and reliability of the composite service based on the data of atomic services such that it is checked up before service deployment whether the service-level agreement is satisfied. In this article, a phase-type expansion approach to establish an expanded homogeneous continuous time Markov chain of composite web services is proposed. By explicitly including failure states and restart states into the model, both performance and reliability can be computed for the composite service via phase-type fitting method based on the observations on execution times of atomic services. Experimental results based on real-world web services are provided to demonstrate the efficacy of the proposed approach.
Zheng Zheng 0001, Yanjie Liu, Zhiyu Xi
IEEE Trans. Reliab.1
2022 Theoretical and Empirical Analyses of the Effectiveness of Metamorphic Relation Composition
abstract
Metamorphic Relations (MRs) play a key role in determining the fault detection capability of Metamorphic Testing (MT). As human judgement is required for MR identification, systematic MR generation has long been an important research area in MT. Additionally, due to the extra program executions required for follow-up test cases, some concerns have been raised about MT cost-effectiveness. Consequently, the reduction in testing costs associated with MT has become another important issue to be addressed. MR composition can address both of these problems. This technique can automatically generate new MRs by composing existing ones, thereby reducing the number of follow-up test cases. Despite this advantage, previous studies on MR composition have empirically shown that some composite MRs have lower fault detection capability than their corresponding component MRs. To investigate this issue, we performed theoretical and empirical analyses to identify what characteristics component MRs should possess so that their corresponding composite MR has at least the same fault detection capability as the component MRs do. We have also derived a convenient, but effective guideline so that the fault detection capability of MT will most likely not be reduced after composition.
Kun Qiu 0001, Zheng Zheng 0001, Tsong Yueh Chen, Pak-Lok Poon
IEEE Trans. Software Eng.2
2021 Manifold Trial Selection to Reduce Negative Transfer in Motor Imagery-based Brain-Computer Interface
abstract
A major challenge in electroencephalogram (EEG) signal classification is that the EEG signals recorded from different subjects are drawn from different distributions. When the unlabeled EEG data of the new subject arrive, called target domain, classifying them with a classifier trained on prerecorded EEG data of other subjects, called source domain, will greatly decrease the classification accuracy. Being able to use the classifiers trained on data of source domain to accurately classify the data of target domain could reduce the time of the calibration phase in the actual application of the brain-computer interface. This study considers an offline cross-subject classification scenario. We propose a novel manifold trial selection method, which reduces the distribution distance between the source and target domains by manifold transformation and domain adaptation. The proposed method provides a trial selection strategy to suppress negative transfer by removing some abnormal samples. The proposed method is applied to the motor imagery-based brain–computer interface and compared with several existing algorithms. Experimental results show that the proposed method outperforms the state-of-the-art methods.
Zilin Liang, Zheng Zheng 0001, Weihai Chen, Jianbin Zhang, Jianer Chen, Zuobing Chen
IROS2
2021 An Empirical Study on Common Bugs in Deep Learning Compilers
abstract
The highly diversified deep learning (DL) frame-works and target hardware architectures bring big challenges for DL model deployment for industrial production. Up to the present, continuous efforts have been made to develop DL compilers with multiple state-of-the-arts available, e.g., TVM, Glow, nGraph, PlaidML, and Tensor Comprehensions (TC). Unlike traditional compilers, DL compilers take a DL model built by DL frameworks as input and generate optimized code as the output for a particular target device. Similar to other software, DL compilers are also error-prone. Buggy DL compilers can generate incorrect code and result in unexpected model behaviors. To better understand the current status and common bug characteristics of DL compilers, we performed a large-scale empirical study of five popular DL compilers covering TVM, Glow, nGraph, PlaidML, and TC, collecting a total of 2,717 actual bug reports submitted by users and developers. We made large manual efforts to investigate these bug reports and classified them based on their root causes, during which five root causes were identified, including environment, compatibility, memory, document, and semantic. After labeling the types of bugs, we further examined the important consequences of each type of bug and analyzed the correlation between bug types and impacts. Besides, we studied the time required to fix different types of bugs in DL compilers. Seven important findings are eventually obtained, with practical implications provided for both DL compiler developers and users.
Xiaoting Du, Zheng Zheng 0001, Lei Ma 0003, Jianjun Zhao 0001
ISSRE2
2021 Nondeterministic Impact of CPU Multithreading on Training Deep Learning Systems
abstract
With the wide deployment of deep learning (DL) systems, research in reliable and robust DL is not an option but a priority, especially for safety-critical applications. Unfortunately, DL systems are usually nondeterministic. Due to software-level (e.g., randomness) and hardware-level (e.g., GPUs or CPUs) factors, multiple training runs can generate inconsistent models and yield different evaluation results, even with identical settings and training data on the same implementation framework and hardware platform. Existing studies focus on analyzing software-level nondeterminism factors and the nondeterminism introduced by GPUs. However, the nondeterminism impact of CPU multi-threading on training DL systems has rarely been studied. To fill this knowledge gap, we present the first work of studying the variance and robustness of DL systems impacted by CPU multithreading. Our major contributions are fourfold: 1) An experimental framework based on VirtualBox for analyzing the impact of CPU multithreading on training DL systems; 2) Six findings obtained from our experiments and examination on GitHub DL projects; 3) Five implications to DL researchers and practitioners according to our findings; 4) Released the research data (https://github.com/DeterministicDeepLearning).
Guanping Xiao, Zheng Zheng 0001, Yulei Sui
ISSRE3
2021 An Empirical Study on Test Case Prioritization Metrics for Deep Neural Networks
abstract
Deep Neural Networks (DNNs) have been widely applied in safety and security domains. DNN testing is necessary to detect the incorrect behaviors of DNNs and guarantee the reliability of DNNs. Labeling test cases is costly that causes DNN testing a serious efficiency problem, which can be alleviated by just labeling test cases with higher priority rather than labeling them in a messy order. Therefore, test case prioritization for DNNs is extensively studied. This paper studies 11 test case prioritization metrics from the ratio of fault detection, accuracy, and correlation perspectives. We classify them into four categories: surprise adequacy, confidence dispersion, mutation uncertainty, and mutation rate. We perform an empirical study of the metrics on two benchmark datasets and DNN models. Our experimental results demonstrate the metrics based on confidence dispersion outperform others regarding effectiveness and efficiency. Meanwhile, we investigate two impact factors of metrics, including test suite size and mutation.
Beibei Yin, Zheng Zheng 0001, Tiancheng Li 0005
QRS3
2021 Availability Analysis of Systems Deploying Sequences of Environmental-Diversity-Based Recovery Methods
abstract
Mandelbug-caused software failures are significant threats to system availability, especially in the context of mission-critical and safety-critical systems. However, there is still no systematic method for keeping the software free from Mandelbugs before release. To guarantee the availability of systems suffering from Mandelbugs, environmental-diversity-based fault tolerance techniques have been proposed to recover from the failures caused by them. In this article, we develop and study an analytic model to assess the availability of systems that utilize a sequence of environmental-diversity-based recovery methods. Improving over previous relevant studies, the availability formula we obtain in this article works for any number of recovery methods the system is equipped with; it is also independent on both the nature of those recovery methods and the order of their utilization. In addition, we consider the problem of how to arrange the set of available recovery methods to achieve the largest system availability. Based on the results of our analysis, we develop an open-source tool, called OPENS, which assists in the calculation of the optimal system availability. We validate the effectiveness of the proposed modeling approach in two ways, namely by comparing our results with those obtained for specific systems considered in relevant studies and by conducting numerical analyses for more general scenarios of its application.
Kun Qiu 0001, Zheng Zheng 0001, Kishor S. Trivedi, Ivan Mura
IEEE Trans. Reliab.2
2020 An Empirical Study of Code Deobfuscations on Detecting Obfuscated Android Piggybacked Apps
abstract
Android piggybacked malware (i.e., apps that piggyback malicious code) are becoming ubiquitous in app stores. Malware writers often use obfuscation techniques to obfuscate piggybacked apps to evade detection by Android malware detectors. Previous studies in this field have focused on the impact of code obfuscations on the detection of piggybacked malware, but the impact of code deobfuscation on detecting obfuscated piggybacked apps has rarely been studied. Knowing about the impact of code deobfuscation can provide useful insights into obfuscated piggybacked apps and therefore the design of resilient Android malware detectors. In this paper we conduct an empirical study of code deobfuscations on detecting obfuscated Android piggybacked apps, focusing on three types of malware detectors: commercial anti-malware products, machine learning-based detectors, and similarity-based detectors. We observe that code deobfuscations can impact differently depending on the malware detectors. For example, some deobfuscation strategies can improve the precision of detecting obfuscated piggybacked apps. Also we observe that the examined deobfuscation tools (Simplify and Deguard) have a different impact on obfuscated piggybacked apps after deobfuscations.
Guanping Xiao, Zheng Zheng 0001, Tianqing Zhu, Ivor W. Tsang, Yulei Sui
APSEC3
2020 Exploring the Characteristics of Spectra Distribution and Their Impacts on Fault Localization
abstract
Spectrum-Based Fault Localization (SBFL) follows the basic intuitions that the faulty parts are more likely to be covered by failure-revealing test cases and less likely to be covered by passed test cases. However, due to the diversity of programs and faults, many other characteristics (related to program structure, test suites, and type of faulty components) will influence the practical application of SBFL. For example, a statement can be covered by numerous failure-revealing test cases, and also covered by numerous passed test cases. To get more indicators about the faulty components towards a better application of SBFL, we extend the scope of spectrum-based knowledge from the basic intuitions to the Characteristics of Spectra Distribution (CSDs for short). That is, we explore the relationships between different types of statements and their spectra. Firstly, we introduce the concepts of Failure-Independent, Failure-Related, and Failure-Exclusionary to describe the relationships between different types of statements and their executions. Then, we propose two probabilistic models, with and without the noise of fault interference, respectively, to identify various CSDs for each type of statements. As the analysis results, we introduce a visualization technique to generalize the identified CSDs and provide an overall picture of spectra distribution and its dynamics. Finally, based on our analysis and also the observation of the program spectra of current benchmarks, we design a technique to filter the potential non-faulty statements to improve the accuracy of SBFL.
Xiao-Yi Zhang 0005, Zheng Zheng 0001
EASE2
2020 Guest editorial: special issue on modeling and mitigation techniques for software aging
Zheng Zheng 0001, Kishor S. Trivedi
Softw. Qual. J.1
2020 Markov Regenerative Models of WebServers for Their User-Perceived Availability and Bottlenecks
abstract
The Internet world is moving toward a scenario where users and applications have very diverse service expectation, making the current best-effort model inadequate and limiting. To be able to design high-availability service systems, it is essential to consider not only the actual failure and recovery behavior of the service infrastructure, but also the behavioral aspects of its user and their subjective perceptions and reactions in the wake of failure events. In this paper, we propose to use Markov regenerative process (MRGP) models to study the availability of Internet-based services perceived by a Web user on two different online service scenarios: (1) single-user-single-host and (2) single-user-multiple-host. The MRGP models capture the interactions between the service facility and the user. We also detect its parameter bottlenecks by applying the formal sensitivity analysis technique. The trends of the users' perceived unavailability are analyzed with the changed different parameter values, and the necessity of the sophisticated MRGP modeling is evidenced by the comparisons with the corresponding continuous time Markov chain (CTMC) models, which show that the popular convenient CTMC models tend to overestimate user-perceived service unavailability. Finally, controlled experiments are carried out on a real Web service to demonstrate the proposed approach.
Zheng Zheng 0001, Kishor S. Trivedi, Kun Qiu 0001
IEEE Trans. Dependable Secur. Comput.1
2020 Familial Clustering for Weakly-Labeled Android Malware Using Hybrid Representation Learning
abstract
Labeling malware or malware clustering is important for identifying new security threats, triaging and building reference datasets. The state-of-the-art Android malware clustering approaches rely heavily on the raw labels from commercial AntiVirus (AV) vendors, which causes misclustering for a substantial number ofweakly-labeled malwaredue to the inconsistent, incomplete and overly generic labels reported by these closed-source AV engines, whose capabilities vary greatly and whose internal mechanisms are opaque (i.e., intermediate detection results are unavailable for clustering). The raw labels are thus often used as the only important source of information for clustering. To address the limitations of the existing approaches, this paper presents Andre, a new ANDroid Hybrid REpresentation Learning approach to clustering weakly-labeled Android malware by preserving heterogeneous information from multiple sources (including the results of static code analysis, the meta-information of an app, and the raw-labels of the AV vendors) to jointly learn a hybrid representation for accurate clustering. The learned representation is then fed into our outlier-aware clustering to partition the weakly-labeled malware into known and unknown families. The malware whose malicious behaviours are close to those of the existing families on the network, are further classified using a three-layer Deep Neural Network (DNN). The unknown malware are clustered using a standard density-based clustering algorithm. We have evaluated our approach using 5,416 ground-truth malware from Drebin and 9,000 malware from VirusShare (uploaded between Mar. 2017 and Feb. 2018), consisting of 3324 weakly-labeled malware. The evaluation shows that Andre effectively clusters weakly-labeled malware which cannot be clustered by the state-of-the-art approaches, while achieving comparable accuracy with those approaches for clustering ground-truth samples.
Yulei Sui, Shirui Pan, Zheng Zheng 0001, Baodi Ning, Ivor W. Tsang, Wanlei Zhou 0001
IEEE Trans. Inf. Forensics Secur.4
2020 Stress Testing With Influencing Factors to Accelerate Data Race Software Failures
abstract
Software failures caused by data race bugs have always been major concerns in parallel and distributed systems, despite significant efforts spent in software testing. Due to their nondeterministic and hard-to-reproduce features, when evaluating systems' operational reliability, a rather long period of experimental execution time is expected to be spent on observing failures caused by data race conditions. To address this problem, in this paper, we make two contributions. First, this paper proposes stress testing with influencing factors, in which the system runs under certain workloads for a long time with controlled stress conditions to accelerate the occurrence of data race failures. Second, it explores and formulates mathematical relationship models between data races' statistical characteristics of time to failure (TTF) or mean TTF (MTTF) and the influencing factors. Such relationship models are used for TTF/MTTF extrapolation under different operational conditions and are essential to reduce systems' reliability evaluation time. The proposed method is empirically evaluated on six applications suffering from failures caused by real-world data race bugs. Through analysis of the experimental results, we obtain several important findings: First, the reduction in the manifestation time to data race failures achieved by controlling the influencing factors is statistically significant. Second, Power model is the best-fitting model of the relationship between the MTTF and the influencing factors. Third, Power Weibull distribution is the best-fitting probability distribution between the TTF and the influencing factors. Finally, the TTF/MTTF can be accurately estimated with the approach proposed in this paper.
Kun Qiu 0001, Zheng Zheng 0001, Kishor S. Trivedi, Beibei Yin
IEEE Trans. Reliab.2
2020 An Empirical Study of Regression Bug Chains in Linux
abstract
Regression bugs are a type of bugs that cause a feature of software that worked correctly but stop working after a certain software commit. This paper presents a systematic study of regression bug chains, an important but unexplored phenomenon of regression bugs. Our paper is based on the observation that a commit c1, which fixes a regression bug b1, may accidentally introduce another regression bug b2. Likewise, commit c2 repairing b2 may cause another regression bug b3, resulting in a bug chain, i.e., b1 → c1 → b2 → c2 → b3. We have conducted a large-scale study by collecting 1579 regression bugs and 2630 commits from 57 Linux versions (from 2.6.12 to 4.9). The relationships between regression bugs and commits are modeled as a directed bipartite network. Our major contributions and findings are fourfold: 1) a novel concept of regression bug chains and their formulation; 2) compared to an isolated regression bug, a bug on a regression bug chain is much more difficult to repair, costing 2.4× more fixing time, involving 1.3× more developers and 2.8× more comments; 3) 85.8% of bugs on the chains in Linux reside in Drivers, ACPI, Platform Specific/Hardware, and Power Management; and 4) 83% of the chains affect only a single Linux subsystem, while 68% of the chains propagate across Linux versions.
Guanping Xiao, Zheng Zheng 0001, Bo Jiang 0001, Yulei Sui
IEEE Trans. Reliab.2
2019 Supervised Representation Learning Approach for Cross-Project Aging-Related Bug Prediction
abstract
Software aging, which is caused by Aging-Related Bugs (ARBs), tends to occur in long-running systems and may lead to performance degradation and increasing failure rate during software execution. ARB prediction can help developers discover and remove ARBs, thus alleviating the impact of software aging. However, ARB-prone files occupy a small percentage of all the analyzed files. It is usually difficult to gather sufficient ARB data within a project. To overcome the limited availability of training data, several researchers have recently developed cross-project models for ARB prediction. A key point for cross-project models is to learn a good representation for instances in different projects. Nevertheless, most of the previous approaches neither consider the reconstruction property of new representation nor encode source samples' label information in learning representation. To address these shortcomings, we propose a Supervised Representation Learning Approach (SRLA), which is based on double encoding-layer autoencoder, to perform cross-project ARB prediction. Moreover, we present a transfer cross-validation framework to select the hyper-parameters of cross-project models. Experiments on three large open-source projects demonstrate the effectiveness and superiority of our approach compared with the state-of-the-art approach TLAP.
Xiaohui Wan, Zheng Zheng 0001, Fangyun Qin, Kishor S. Trivedi
ISSRE2
2019 A Visualization Analytical Framework for Software Fault Localization Metrics
abstract
The core of Spectra-Based Fault Localization (SBFL) is suspiciousness metric, expressed as a formula to calculate the fault proneness for each program component. Current analysis works on metrics mainly focus on the comparison of their performances based on algebraic reasoning. However, due to the high complexity of real-life programs, there are still challenges in the practical application of SBFL. This paper emphasizes a further exploration of the mechanism of SBFL metrics. We propose a visualization-based framework for metric analyses, in which metrics are interpreted by curves in the identified spectra space, and their performance can be illustrated by geometric properties. Based on the framework, we design a basic approach for metric analysis following the procedures: visualizing representative SBFL instances → generalizing geometric knowledge → obtaining useful guidance. Due to the advantages of visualization, we can get explainable and essential knowledge about SBFL. In particular, we make a comparative analysis among typical metrics and, compared with algebraic reasoning, obtain not only the comparison results but also the explanation about why a metric can outperform others as well as new theoretical findings such as the optimality of continuous maximal metrics. Finally, we make an extended discussion about the possible way to study the influence of fault interferences on SBFL, which indicates the extensibility of our framework.
Xiao-Yi Zhang 0005, Zheng Zheng 0001
PRDC2
2019 Testing Graph Searching Based Path Planning Algorithms by Metamorphic Testing
abstract
Path planning algorithms play critical roles in the systems of robots and unmanned aerial vehicles (UAVs). However, it is always difficult to verify the correctness of the implementations for such algorithms because the "planning oracles", the expected planning results, are usually hard to be obtained for complicate planning tasks. To improve software reliability, in this paper, we present a testing technique for verifying the implementations of graph searching based path planning algorithms deployed on robots and UAVs. Our approach is based on the technique of Metamorphic Testing, which has been shown considerable effectiveness in alleviating the absence of Oracle problems. According to the characteristics of graph searching based path planning problem, we present a framework to systematically design metamorphic relations. Based on the framework, six categories of metamorphic relations are proposed. We conduct the empirical analysis on 21 implements of three different path planning algorithms applied in a released business software project. The experimental results show that our approach can effectively detect dormant faults.
Zheng Zheng 0001, Beibei Yin, Kun Qiu 0001, Yang Liu 0287
PRDC2
2019 Robustness of spectrum-based fault localisation in environments with labelling perturbations
Beibei Yin, Zheng Zheng 0001, Xiao-Yi Zhang 0005, Shunkun Yang
J. Syst. Softw.3
2019 Two-Level Rejuvenation for Android Smartphones and Its Optimization
abstract
The Android operating system (OS) is a sophisticated man-made system and is the dominant OS in the current smartphone market. Due to the accumulation of errors in the system internal state and the incremental consumption of resources, such as the Dalvik heap memory of software applications and the physical memory, software aging is observed frequently and recognized as a chronic problem of Android smartphones. To mitigate this problem, we propose a two-level software rejuvenation, with the two levels referring to software applications and the OS, in this paper. Based on this strategy, a Markov regenerative process model is constructed to evaluate the steady-state availability and to optimize the time required to trigger rejuvenation for Android smartphones. The parameters of the model, such as the degradation rate and failure rate of software applications and the Android OS, are obtained via our testing platform. Experiments on two real Android applications show that the availability of an Android smartphone increases by 10.81% and 10.18% for the two subjects in our experiments, respectively. An empirical study comparing our two-level strategy with one-level strategies (single application-level and system-level rejuvenation) further verifies the effectiveness of our approach.
Zheng Zheng 0001, Yunyu Fang, Fangyun Qin, Kishor S. Trivedi, Kai-Yuan Cai
IEEE Trans. Reliab.2
2019 Studying Aging-Related Bug Prediction Using Cross-Project Models
abstract
In long running systems, software tends to encounter performance degradation and increasing failure rate during execution. This phenomenon has been named software aging, which is caused by aging-related bugs (ARBs). Testing resource allocation can be optimized by identifying ARB-prone modules with ARB prediction. However, due to the low presence and reproducing difficulty of ARBs, it is usually hard to collect sufficient training data to carry out within-project ARB prediction. In this paper, we propose an approach named transfer learning based aging-related bug prediction (TLAP) to perform cross-project ARB prediction. TLAP first takes advantage of transfer learning to reduce distribution difference between training and testing project. Then, class imbalance learning is conducted to mitigate the severe class imbalance between ARB-prone and ARB-free modules. Finally, machine learning methods are used to handle bug prediction tasks. The effectiveness of this approach is validated and evaluated by nine groups of experiments on real software systems. Major conclusions from the experiments include the following: first, TLAP improves cross-project ARB prediction on average compared with traditional machine learning methods; second, utilizing information from multiple-projects can further improve the prediction performance on average. In the best case, it outperforms within-project prediction; third, the number of ARB-prone files and distribution similarity can influence TLAP performance.
Fangyun Qin, Zheng Zheng 0001, Kishor S. Trivedi
IEEE Trans. Reliab.2
2019 An Empirical Study of Fault Triggers in the Linux Operating System: An Evolutionary Perspective
abstract
This paper presents an empirical study of 5741 bug reports for the Linux kernel from an evolutionary perspective, with the aim of obtaining a deep understanding of bug characteristics in the Linux operating system. Bug classification is performed based on the fault triggering conditions, followed by an analysis of the proportions and evolution of the bug types as well as comparisons among versions, products, and repair locations. In addition, an analysis of regression bugs and the relationship between the types of bugs and the time needed to fix them are presented. Moreover, a procedure for the analysis of bug type characteristics based on complex network metrics is proposed, and four network metrics, i.e., degree, clustering coefficient, betweenness, and closeness, are utilized to further investigate the relationship between bug types and software metrics. In this paper, 22 interesting findings based on the empirical results are revealed, and guidance based on these findings is provided for developers and users.
Guanping Xiao, Zheng Zheng 0001, Beibei Yin, Kishor S. Trivedi, Xiaoting Du, Kai-Yuan Cai
IEEE Trans. Reliab.2
2018 Parallel construction of interprocedural memory SSA form
Yulei Sui, Zheng Zheng 0001, Jingling Xue
J. Syst. Softw.3
2018 Exploring the usefulness of unlabelled test cases in software fault localization
Xiao-Yi Zhang 0005, Zheng Zheng 0001, Kai-Yuan Cai
J. Syst. Softw.2
2018 Survey on computational-intelligence-based UAV path planning
Yijing Zhao, Zheng Zheng 0001, Yang Liu 0287
Knowl. Based Syst.2
2018 A Fortification Model for Decentralized Supply Systems and Its Solution Algorithms
abstract
Service disruptions due to deliberate sabotage are serious threats to supply systems. To alleviate the loss of accessibility caused by such disruptions, identifying the system vulnerabilities that would be worth strengthening is a critical problem in critical infrastructure protection. Today's supply systems tend to be organized in a decentralized manner, with different components belonging to different entities, keeping much information private. Therefore, a protection plan must balance its benefits among these entities for universal agreement to be reached. This paper addresses the issue of decentralized supply chain fortification by proposing the R-Interdiction Median problem with Fortification for Decentralized supply systems (D-RIMF). In the D-RIMF, each demand node is private and is a client of a certain facility; each facility evaluates its potential worst-case reduction in accessibility, measured as the increase in service provision costs considering only its own clients, and the objective is to minimize the largest evaluation values. To model the D-RIMF, we introduce a bilevel multiagent framework, in which all facilities and the defender are considered as independent agents. To solve the D-RIMF, both heuristic and optimal algorithms are designed to satisfy different requirements. Finally, the usefulness of the D-RIMF and the performances of the proposed algorithms are observed through simulations performed on typical datasets.
Xiao-Yi Zhang 0005, Zheng Zheng 0001, Kai-Yuan Cai
IEEE Trans. Reliab.2
2017 Understanding the Impacts of Influencing Factors on Time to a DataRace Software Failure
abstract
Datarace is a common problem on shared-memory parallel computers, including multicores. Due to its dependence on the thread scheduling scheme of its execution environment, the time to a datarace failure is usually very long. How to accelerate the occurrence of a datarace failure and further estimate the mean time to failure (MTTF) is an important topic to be studied. In this paper, the influencing factors for failures triggered by datarace bugs are explored and their influences on the time to datarace failure including the relationship with the MTTF are empirically studied. Experiments are conducted on real datarace suffering programs to verify the factors and their influences. Empirical results show that the influencing factors do have influences on the time to datarace failure of the subjects. They can be used to accelerate the occurrence of datarace failures and accurately estimate the MTTF.
Kun Qiu 0001, Zheng Zheng 0001, Kishor S. Trivedi, Beibei Yin
ISSRE2
2017 Experience Report: Fault Triggers in Linux Operating System: from Evolution Perspective
abstract
Linux operating system is a complex system that is prone to suffer failures during usage, and increases difficulties of fixing bugs. Different testing strategies and fault mitigation methods can be developed and applied based on different types of bugs, which leads to the necessity to have a deep understanding of the nature of bugs in Linux. In this paper, an empirical study is carried out on 5741 bug reports of Linux kernel from an evolution perspective. A bug classification is conducted based on fault triggering conditions, followed by the analysis of the evolution of bug type proportions over versions and time, together with their comparisons across versions, products and regression bugs. Moreover, the relationship between bug type proportions and clustering coefficient, as well as the relation between bug types and time to fix are presented. This paper reveals 13 interesting findings based on the empirical results and further provides guidance for developers and users based on these findings.
Guanping Xiao, Zheng Zheng 0001, Beibei Yin, Kishor S. Trivedi, Xiaoting Du, Kai-Yuan Cai
ISSRE2
2017 A Rejuvenation Strategy of Two-Granularity Software Based on Adaptive Control
abstract
In the process of continuous operation in a software system, a series of phenomena could lead to performance degradation of the system, namely software aging. The loss caused by software aging can be reduced through proper rejuvenation strategies, the key to which is to determine the rejuvenation thresholds. Essence of some traditional methods is to set predetermined thresholds based on empirical data. However, in some systems where the memory is shared between operating system and application software (two-granularity software system), as the memory consumption is closely related to system performance and changes constantly, using empirical thresholds may cause system outage or waste of resources. In this paper, an adaptive strategy is adopted to optimize the thresholds. Instead of fixed thresholds, the method regularly regulates the thresholds by taking feedback information in the running process into account. Especially, critical equations are constructed to calculate the thresholds by maximizing the system availability. Simulation results show that the proposed method achieves higher availability and more stable performance than that based on empirical thresholds.
Yunyu Fang, Beibei Yin, Gao-Rong Ning, Zheng Zheng 0001, Kai-Yuan Cai
PRDC4
2017 An Empirical Investigation of Fault Triggers in Android Operating System
abstract
The growing popularity and complexity of Android operating system makes it prone to suffer failures during usage, which increases difficulties of fixing bugs. Different strategies and mitigation methods can be developed and applied based on different types of bugs, which gives rise to the necessity to have a deep understanding of the nature of bugs in this system. In this paper, an empirical study is taken on 513 bug reports from Android operating system. A bug classification is conducted according to fault triggering conditions, followed by the analysis of bug types and bug attributes. Moreover, the comparison of bug types between Android and Linux is carried out. This paper reveals ten interesting findings based on the empirical results from these three aspects and further provides guidance for developers and users based on these findings.
Fangyun Qin, Zheng Zheng 0001, Kishor S. Trivedi
PRDC2
2017 A theoretical analysis on cloning the failed test cases to improve spectrum-based fault localization
Lanfei Yan, Zhenyu Zhang 0004, Jian Zhang 0001, Wing Kwong Chan, Zheng Zheng 0001
J. Syst. Softw.6
2017 Semi-Markov Models of Composite Web Services for their Performance, Reliability and Bottlenecks
abstract
When combining several services into a composite service, it is non-trivial to determine, prior to service deployment, performance and reliability values of the composite service. Moreover, once the service is deployed, it is often the case that during operation it fails to meet its service-level agreement (SLA) and one needs to detect what has gone wrong (i.e., performance/reliability bottlenecks). To study these issues, we develop a Semi-Markov Process (SMP) formulation of composite services with failures and restarts. By explicitly including failure states into the SMP representation of a service, we can compute both its performance and reliability using a single SMP. We can also detect its performance and reliability bottlenecks by applying the formal sensitivity analysis technique. We demonstrate our approach by choosing a representative example that is validated using experiments on real Web services.
Zheng Zheng 0001, Kishor S. Trivedi, Kun Qiu 0001, Ruofan Xia
IEEE Trans. Serv. Comput.1
2016 Exploring the Instability of Spectra Based Fault Localization Performance
abstract
Spectra Based Fault Localization (SBFL) is a technique to improve the efficiency of software fault localization. The performance of SBFL largely depends on the input information provided by an executed test suite. Due to the randomness existing in the testing process, the output of SBFL may not be stable. In practice, testers do not have the chance to run the whole testing process many times. They are not sure whether the actually obtained SBFL output has a large deviation from the ideal output (i.e. the SBFL output obtained under the assumption that the amount of testing resources is unlimited). Thus, concerning the application of SBFL in real cases, such instability of its performance (SBFL instability for short) is a challenge. In this paper, the SBFL instability is discussed and its characteristics are further explored. Specifically, we define SBFL instability as a stochastic quantity and introduce the measure of StabilityLevel to quantify it. Then, based on the definition and measurement, we conduct experimental studies to demonstrate that SBFL instability is indeed a prevalent phenomenon and also a serious problem. Besides, two factors which influence the intensity of SBFL instability, i.e. the test suite size and risk evaluation formula, are observed and analyzed.
Yuanchi Guo, Xiao-Yi Zhang 0005, Zheng Zheng 0001
COMPSAC3
2016 The more obstacle information sharing, the more effective real-time path planning?
Zheng Zheng 0001, Yang Liu 0287, Xiao-Yi Zhang 0005
Knowl. Based Syst.1
2015 Using Partition Information to Prioritize Test Cases for Fault Localization
abstract
Fault Localization Prioritization (FLP) aims at reordering existing test cases so that the location of detected faulty components can be identified earlier, using certain fault localization techniques. Although some researchers have proposed adaptive prioritization strategies with white-box code coverage information, such information may not always be available. In this paper, we address the FLP problem using black-box information derived from partitioning the input domain. Based on the well-known technique of Spectra-Based Fault Localization (SBFL), three test case prioritization strategies are designed following some basic SBFL heuristics. The implementation of these proposed strategies relies only on the partition information, and does not require any test case execution history. Experiments show that our strategies, when compared with pure random selection, result in a faster localization of faulty statements, reducing the number of test case executions required. Here, we analyze the characteristics and merits of the three proposed strategies.
Xiao-Yi Zhang 0005, Dave Towey, Tsong Yueh Chen, Zheng Zheng 0001, Kai-Yuan Cai
COMPSAC4
2015 Cross-Project Aging Related Bug Prediction
abstract
In a long running system, software tends to encounter performance degradation and increasing failure rate during execution, which is called software aging. The bugs contributing to the phenomenon of software aging are defined as Aging Related Bugs (ARBs). Lots of manpower and economic costs will be saved if ARBs can be found in the testing phase. However, due to the low presence probability and reproducing difficulty of ARBs, it is usually hard to predict ARBs within a project. In this paper, we study whether and how ARBs can be located through cross-project prediction. We propose a transfer learning based aging related bug prediction approach (TLAP), which takes advantage of transfer learning to reduce the distribution difference between training sets and testing sets while preserving their data variance. Furthermore, in order to mitigate the severe class imbalance, class imbalance learning is conducted on the transferred latent space. Finally, we employ machine learning methods to handle the bug prediction tasks. The effectiveness of our approach is validated and evaluated by experiments on two real software systems. It indicates that after the processing of TLAP, the performance of ARB bug prediction can be dramatically improved.
Fangyun Qin, Zheng Zheng 0001, Chenggang Bai, Zhenyu Zhang 0004
QRS2
2015 A new solution algorithm for solving rule-sets based bilevel decision problems
abstract
Summary Bilevel decision addresses compromises between two interacting decision entities within a given hierarchical complex system under distributed environments. Bilevel programming typically solves bilevel decision problems. However, formulation of objectives and constraints in mathematical functions is required, which are difficult, and sometimes impossible, in real‐world situations because of various uncertainties. Our study develops a rule‐set based bilevel decision approach, which models a bilevel decision problem by creating, transforming and reducing related rule sets. This study develops a new rule‐sets based solution algorithm to obtain an optimal solution from the bilevel decision problem described by rule sets. A case study and a set of experiments illustrate both functions and the effectiveness of the developed algorithm in solving a bilevel decision problem. Copyright © 2012 John Wiley & Sons, Ltd.
Jie Lu 0001, Zheng Zheng 0001, Guangquan Zhang 0001, Qing He 0003, Zhongzhi Shi
Concurr. Comput. Pract. Exp.2
2014 An experimental study on firewall performance: Dive into the bottleneck for firewall effectiveness
abstract
Performance is an important indicator of firewalls effectiveness, which represents capability of firewalls handling network requests. ModSecurity and iptables, two representative firewalls of packet filtering and application firewall, are studied experimentally in this paper. Firstly, we develop the experiments to test the capacity of these two kinds of firewalls. Secondly, we locate the bottlenecks for system resources such as CPU and memory usage that affect the firewalls performance by analyzing the collecting data from firewalls experiments. Finally, with the same settings, we compare the performance of the two kinds of firewalls by varying the parameters such as request rate, packet length, and maximum concurrent connections.
Cheng-Hong Wang, Donghong Zhang, Hualin Lu, Jing Zhao 0016, Zhenyu Zhang 0004, Zheng Zheng 0001
IAS6
2014 A critical chains based distributed multi-project scheduling approach
Zheng Zheng 0001, Ze Guo, Yueni Zhu, Xiao-Yi Zhang 0005
Neurocomputing1
2013 Bi-level programming based real-time path planning for unmanned aerial vehicles
Zheng Zheng 0001, Kai-Yuan Cai
Knowl. Based Syst.2
2012 Factorising the Multiple Fault Localization Problem: Adapting Single-Fault Localizer to Multi-fault Programs
abstract
Software failures are not rare and fault localizations always an important but laborious activity. Since there is no guarantee that no more than one fault exists in a faulty program, the approach to locate all the faults is necessary. Spectrum-based fault localization techniques collect dynamic program spectra as well as test results of program runs, and estimate the extent of program elements being related to fault(s). A popular solution into generate a ranked list of suspicious candidates, which are checked in order, stopping whenever a fault is found. Such single fault localizers locate one fault in one checking round, terminate, and wait to be triggered by the regression testing to validate the fixing of the located fault. In this paper, we study the manifestation of multiple faults in a program and propose an effective mechanism to indicate their presence. When a fault is reached during the checking round, we use it to interpret the failures observed, and update the indicator to judge whether there remain other faults in the program. Our indicator serves as a stopping criterion of checking the ranked list of suspicious candidates. Our work factories the multiple fault localization problem into developing single-fault localizers and adapting them to multi-fault programs. It both improves the fault localization efficiencies of single-fault localizers, and avoids the ineffective efforts of thoroughly abandoning the many single-fault localizers to develop multi-fault localizers.
Zheng Zheng 0001, Yunqian Zhang, Zhenyu Zhang 0004, Yunzhi Xue
APSEC2
2011 An Algorithm for Solving Rule Sets-Based Bilevel Decision Problems
abstract
Bilevel decision addresses the problem in which two levels of decision makers each tries to optimize their individual objectives under certain constraints, and to act and react in an uncooperative and sequential manner. Given the difficulty of formulating a bilevel decision problem by mathematical functions, a rule sets–based bilevel decision (RSBLD) model was proposed. This article presents an algorithm to solve a RSBLD problem. A case‐based example is given to illustrate the functions of the proposed algorithm. Finally, a set of experiments is analyzed to further show the functions and the effectiveness of the proposed algorithm.
Guangquan Zhang 0001, Zheng Zheng 0001, Jie Lu 0001, Qing He 0003
Comput. Intell.2
2010 Robustness of fuzzy operators in environments with random perturbations
Zheng Zheng 0001, Kai-Yuan Cai
Soft Comput.1
2009 Propagation of Random Perturbations under Fuzzy Algebraic Operators
Zheng Zheng 0001, Shanjie Wu, Kai-Yuan Cai
KSEM1
2009 Rule sets based bilevel decision model and algorithm
Zheng Zheng 0001, Jie Lu 0001, Guangquan Zhang 0001, Qing He 0003
Expert Syst. Appl.1
2004 Rough Set Based Image Texture Recognition Algorithm
Zheng Zheng 0001, Hong Hu 0001, Zhongzhi Shi
KES1