VLDB 2026 Research / reviewers in the wild / expert
Beibei Yin
dblp:171/2320 · also Bei-Bei Yin
· DBLP profile ↗
34ranked-venue papers
2as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 29 · 2 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-authorSecurity and privacy · 2Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Dynamic Test Oracle for Quantum Programs With Separable Output StatesabstractAs quantum software engineering advances, testing techniques are required to assess the quality of quantum programs (QPs). In the test process, the test oracle is vital for determining whether the test result indicates a success or a failure. Most related works directly measure the output states and acquire the corresponding test results by comparing the output distribution with the expected one. While attention has been paid to the capability of fault detection, the guarantee for the correctness of the produced test results remains limited. Unlike classical programs (CPs), the output quantum states of QPs should be transformed into probabilistic classical outcomes through quantum measurement. This additional operation of measurement could cause a test oracle to yield the wrong test results. Especially for high-dimensional output spaces, numerous measurement outcomes are required to capture the distribution characteristics, threatening the effectiveness and cost-efficiency of test oracles. Hence, this paper proposes a novel specified test oracle DOSS employing a dynamic scheme to integrate a quantum algorithm (i.e., swap test) with the direct measurement mode. This innovative approach enables the validation of individual outputs rather than their distribution during the testing phase. Considering acceptable cost, DOSS decomposes the fully or partially separable output states to lower the dimensionality and simplify the quantum circuit for testing. Empirical studies demonstrate that DOSS generally gives more correct test results than baselines, and maintains reasonable cost on an ideal simulator. Besides, DOSS’s effectiveness with quantum noise involved is validated via three noisy simulators. Yuechen Li 0001, Kai-Yuan Cai, Beibei Yin |
IEEE Trans. Software Eng. | 3 |
| 2025 | QuAInth: A Code Comment Approach for Application-Oriented Quantum Programs via N-Version LLMsabstractQuantum computing has recently experienced rapid advancements and promised transformative applications across many fields. To promote the real-world applications of quantum computing, application-oriented quantum programs (AQPs) are designed to explore quantum hardware in performing substantial computational tasks and promote practical use cases for quantum computing. The complexity and scalability of AQPs, along with their dependence on advanced quantum algorithms, make them particularly challenging to understand and maintain. Clear comments offer an effective means of elucidating core logic and filling the knowledge gap for developers unfamiliar with quantum mechanics. Research on code comment for QPs remains scarce, highlighting the need for further investigation into effective comment methods for these complex programs. Given the potential advantages of large language models (LLMs), including their contextual understanding and language generation capabilities, LLMs can significantly reduce the cost of manual comment. Thus, this paper proposes a framework called QuAInth, which utilizes LLMs to annotate AQPs. This framework begins by preprocessing AQPs and segmenting them based on functional signatures. It then employs prompts (i.e., textual instructions carefully structured and given to LLMs) of varying granularity to guide comment generation. Aside from two existing text-based metrics, QuAInth newly adopts a quantum-specific metric that considers 8 indicators to evaluate the domainrelated correctness and clarity of the generated comments. Finally, with the understanding that an individual LLM may produce wrong outputs, QuAInth proposes a vote-enhanced fusion scheme inspired by N-version programming, in which distinct comments output from multiple LLMs are fused into a more reliable and comprehensive comment. Empirical studies are conducted with 4 AQPs written by Qiskit and 3 prevailing open-source LLMs (i.e., Qwen, DeepSeek-Coder, and Llama). The empirical results demonstrate the effectiveness of QuAInth, showing that the fused comments outperform those generated by individual models in the vast majority of cases. Yuechen Li 0001, Jinlong Wen, Kai-Yuan Cai, Beibei Yin |
QRS | 5 |
| 2025 | Markov model based coverage testing of deep learning software systems
Beibei Yin, Jing-Ao Shi |
Inf. Softw. Technol. | 2 |
| 2025 | Multigranularity Coverage Criteria for Deep Learning LibrariesabstractABSTRACT Deep learning (DL) systems are becoming increasingly widely used in safety domains such as self‐driving cars and unmanned aerial vehicles, which arouse natural concerns about their trustworthiness. Underlying DL libraries used in the construction and execution of DL models are involved in the testing processes of DL systems. Therefore, bugs in DL libraries can inevitably cause unexpected behaviours in DL systems. The internal structures of DL libraries are described as APIs with different functionalities, and DL libraries offer model developers access to DL techniques with various API parameter settings. The above characteristics of DL libraries reveal that existing DL coverage criteria are not designed specifically for DL libraries, and traditional software coverage criteria do not apply to DL libraries either. The paper introduces the first set of coverage criteria specifically designed for the systematic measurement of DL libraries across various granularities. APIs, as the fundamental components of DL libraries, are used to define coverage criteria gauging testing adequacy by thoroughly considering their invocation, implementation, parameter quantities and parameter attributes. Furthermore, some properties depicting relations between coverage criteria are investigated. Experiments on the effectiveness of the proposed coverage criteria and comparative analysis are conducted by interval estimate and hypothesis testing techniques for APIs in two well‐known DL libraries. The experimental results demonstrate that the proposed coverage criteria are effective in measuring the test adequacy of DL libraries, and they can be used for the quantitative analysis of test model quality in DL libraries. Zheng Zheng 0001, Beibei Yin, Zhiyu Xi |
Softw. Test. Verification Reliab. | 3 |
| 2025 | Preparation and Utilization of Mixed States for Testing Quantum ProgramsabstractDue to the growing demand for high-quality quantum programs (QPs), unit testing is employed to check the behavior of QPs. As for quantum inputs of testing, most studies limit test inputs to pure states, whereas mixed states representing probabilistic mixtures of pure states are almost excluded from the test process. Besides, when achieving the input domain coverage, lots of pure-state test cases (PSTCs) with pure states as inputs should be employed, leading to high time costs for testing. To handle that, this article explores using mixed states as test inputs for better utilization of quantum information. From the perspective of input domain coverage, applying mixed-state test cases (MSTCs) replacing PSTCs can simplify the test suite and accordingly promote test efficiency. Owing to the mixture of multiple pure states, a single MSTC is more likely to detect a fault than a PSTC, thereby enhancing test effectiveness. This article then proposes a unit testing framework, including generation and execution of MSTCs. Also, this article presents two guidelines and two parameterized quantum circuits to prepare desired mixed states. Empirical studies evaluate the performance of MSTCs and the experimental results demonstrate that MSCTs generally consume less time and detect more faults than PSTCs. Yuechen Li 0001, Kai-Yuan Cai, Beibei Yin |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2024 | A Strategy of Dynamic Random Testing with Hybrid Distance Metrics for Quantum ProgramsabstractQuantum Computing (QC) leverages quantum mechanics to manipulate quantum information, holding greater potential than classical computing. To fully exploit QC’s potential, it is crucial to ensure the reliability and quality of quantum programs. Research on quantum program testing is still at its early stage, in which some distinctive features of quantum programs, e.g., superposition and entanglement, may be overlooked, and the fault detection capability and testing effectiveness are rather limited. Besides, the input space of quantum programs may exponentially grow when the number of qubits increases, posing great challenges to testing quantum programs. It is imperative to develop a proper testing strategy to effectively select the potential failure-causing test cases and detect faults faster. In this paper, test cases with both basis states and superposition ones are considered and generated to cover more input space. A hybrid distance measurement method based on quantum fidelity and Hamming distance is presented for measuring the similarity among quantum test cases. Furthermore, a Dynamic Random Testing strategy based on Hybrid distance metrics (DRT-H) for quantum programs is proposed, which combines the hybrid distance metrics and the feedback mechanism of the classical Dynamic Random Testing (DRT) strategy to adjust the testing profile and guide the test case selection. Experimental studies demonstrate that the proposed DRT-H strategy outperforms the baseline testing strategies in most cases. Linzhi Huang, Hanyu Pei, Yuechen Li 0001, Beibei Yin, Kai-Yuan Cai |
QRS | 4 |
| 2024 | A method of multidimensional software aging prediction based on ensemble learning: A case of Android OS
Yuge Nie, Yulei Chen, Yujia Jiang, Huayao Wu, Beibei Yin, Kai-Yuan Cai |
Inf. Softw. Technol. | 5 |
| 2024 | Multi-granularity coverage criteria for deep reinforcement learning systems
Beibei Yin, Zheng Zheng 0001 |
J. Syst. Softw. | 2 |
| 2024 | Automatic Repair of Quantum Programs via Unitary OperationabstractWith the continuous advancement of quantum computing (QC), the demand for high-quality quantum programs (QPs) is growing. To avoid program failure, in software engineering, the technology of automatic program repair (APR) employs appropriate patches to remove potential bugs without the intervention of a human. However, the method tailored for repairing defective QPs is still absent. This article proposes, to the best of our knowledge, a new APR method named UnitAR that can repair QPs via unitary operation automatically. Based on the characteristics of superposition and entanglement in QC, the article constructs an algebraic model and adopts a generate-and-validate approach for the repair procedure. Furthermore, the article presents two schemes that can respectively promote the efficiency of generating patches and guarantee the effectiveness of applying patches. For the purpose of evaluating the proposed method, the article selects 29 mutated versions as well as five real-world buggy programs as the objects and introduces two traditional APR approaches GenProg and TBar as baselines. According to the experiments, UnitAR can fix 23 buggy programs, and this method demonstrates the highest efficiency and effectiveness among three APR approaches. Besides, the experimental results further manifest the crucial roles of two constituents involved in the framework of UnitAR . Yuechen Li 0001, Hanyu Pei, Linzhi Huang, Beibei Yin, Kai-Yuan Cai |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2023 | An Empirical Study to Identify Software Aging Indicators for Android OSabstractAndroid mobile devices have been suffering from performance degradation and increased failure rates during long-term operation, known as software aging. With the major changes in performance optimization and resource management in Android, it is spotted that the aging behavior of Android devices in the 2020s differs significantly from previous studies in resource utilization and performance metrics, which makes some classic metrics difficult to measure aging well, and new metrics are required to better describe the new phenomenon. Thus, we propose thread- and interface-level metrics to portray aging at a finer granularity and conduct an empirical study to reidentify classic and new software aging metrics in Android. Analysis confirms that software aging in Android is less reflected in global resources metrics but in more fine-grained ones, so thread- and interface-level metrics combined with specific classic resource and process-level metrics are helpful as indicators of software aging. These metrics have been confirmed and deployed for aging monitoring by our mobile phone manufacturer collaborators. A new experimental method customized for metric studies has also been adopted in this paper, significantly reducing data costs and interference in measurements. Yulei Chen, Yuge Nie, Beibei Yin, Zheng Zheng 0001, Huayao Wu |
QRS | 3 |
| 2023 | A dynamic random testing strategy in the context of cloud computing
Hanyu Pei, Beibei Yin, Linzhi Huang, Kai-Yuan Cai |
Softw. Qual. J. | 2 |
| 2022 | A Distance-Based Dynamic Random Testing Strategy for Natural Language Processing DNN ModelsabstractDeep neural networks (DNNs) have achieved tremendous development while they may encounter with incorrect behaviors and result in economic losses. Identifying the most represented data become critical for revealing incorrect behaviours and improving the quality DNN-driven systems. Various testing strategies for DNNs have been proposed. However, DNN testing is still at early stage and existing strategies might not sufficiently effective. Dynamic random testing (DRT) strategy uses the feedback mechanism to guide the test case selection, which has been proved to be effective in fault detection. However, its efficacy for Natural Language Processing (NLP) DNN models has not been thoroughly studied. In this paper, a Distance-based DRT with prioritization (D-DRT-P) is proposed, which combines the priority information and distance information into DRT to guide the selection of test cases and testing profile adjustment. Empirical studies demonstrate that D-DRT-P can improve the fault detecting effectiveness than other test prioritization strategies in most cases. Yuechen Li 0001, Hanyu Pei, Linzhi Huang, Beibei Yin |
QRS | 4 |
| 2021 | An Empirical Study on Test Case Prioritization Metrics for Deep Neural NetworksabstractDeep Neural Networks (DNNs) have been widely applied in safety and security domains. DNN testing is necessary to detect the incorrect behaviors of DNNs and guarantee the reliability of DNNs. Labeling test cases is costly that causes DNN testing a serious efficiency problem, which can be alleviated by just labeling test cases with higher priority rather than labeling them in a messy order. Therefore, test case prioritization for DNNs is extensively studied. This paper studies 11 test case prioritization metrics from the ratio of fault detection, accuracy, and correlation perspectives. We classify them into four categories: surprise adequacy, confidence dispersion, mutation uncertainty, and mutation rate. We perform an empirical study of the metrics on two benchmark datasets and DNN models. Our experimental results demonstrate the metrics based on confidence dispersion outperform others regarding effectiveness and efficiency. Meanwhile, we investigate two impact factors of metrics, including test suite size and mutation. Beibei Yin, Zheng Zheng 0001, Tiancheng Li 0005 |
QRS | 2 |
| 2021 | Dynamic random testing with test case clustering and distance-based parameter adjustment
Hanyu Pei, Beibei Yin, Min Xie 0001, Kai-Yuan Cai |
Inf. Softw. Technol. | 2 |
| 2020 | Cross-project bug type prediction based on transfer learning
Xiaoting Du, Zenghui Zhou, Beibei Yin, Guanping Xiao |
Softw. Qual. J. | 3 |
| 2020 | An empirical study of factors affecting cross-project aging-related bug prediction with TLAP
Fangyun Qin, Xiaohui Wan, Beibei Yin |
Softw. Qual. J. | 3 |
| 2020 | Stress Testing With Influencing Factors to Accelerate Data Race Software FailuresabstractSoftware failures caused by data race bugs have always been major concerns in parallel and distributed systems, despite significant efforts spent in software testing. Due to their nondeterministic and hard-to-reproduce features, when evaluating systems' operational reliability, a rather long period of experimental execution time is expected to be spent on observing failures caused by data race conditions. To address this problem, in this paper, we make two contributions. First, this paper proposes stress testing with influencing factors, in which the system runs under certain workloads for a long time with controlled stress conditions to accelerate the occurrence of data race failures. Second, it explores and formulates mathematical relationship models between data races' statistical characteristics of time to failure (TTF) or mean TTF (MTTF) and the influencing factors. Such relationship models are used for TTF/MTTF extrapolation under different operational conditions and are essential to reduce systems' reliability evaluation time. The proposed method is empirically evaluated on six applications suffering from failures caused by real-world data race bugs. Through analysis of the experimental results, we obtain several important findings: First, the reduction in the manifestation time to data race failures achieved by controlling the influencing factors is statistically significant. Second, Power model is the best-fitting model of the relationship between the MTTF and the influencing factors. Third, Power Weibull distribution is the best-fitting probability distribution between the TTF and the influencing factors. Finally, the TTF/MTTF can be accurately estimated with the approach proposed in this paper. Kun Qiu 0001, Zheng Zheng 0001, Kishor S. Trivedi, Beibei Yin |
IEEE Trans. Reliab. | 4 |
| 2019 | Testing Graph Searching Based Path Planning Algorithms by Metamorphic TestingabstractPath planning algorithms play critical roles in the systems of robots and unmanned aerial vehicles (UAVs). However, it is always difficult to verify the correctness of the implementations for such algorithms because the "planning oracles", the expected planning results, are usually hard to be obtained for complicate planning tasks. To improve software reliability, in this paper, we present a testing technique for verifying the implementations of graph searching based path planning algorithms deployed on robots and UAVs. Our approach is based on the technique of Metamorphic Testing, which has been shown considerable effectiveness in alleviating the absence of Oracle problems. According to the characteristics of graph searching based path planning problem, we present a framework to systematically design metamorphic relations. Based on the framework, six categories of metamorphic relations are proposed. We conduct the empirical analysis on 21 implements of three different path planning algorithms applied in a released business software project. The experimental results show that our approach can effectively detect dormant faults. Zheng Zheng 0001, Beibei Yin, Kun Qiu 0001, Yang Liu 0287 |
PRDC | 3 |
| 2019 | A Distance-Based Dynamic Random Testing with Test Case ClusteringabstractOne goal of software testing strategies is to detect faults faster. Dynamic Random Testing (DRT) strategy uses the testing results to guide the selection of test cases, which have shown to be effective in the fault detection process. However, the effectiveness of DRT still can be improved. In this paper, a distance-based DRT (D-DRT) strategy is proposed. The vectorized test cases are partitioned with k-means clustering method to obtain better classification, and the distance information are used to guide the test case selection, then the test cases that are close to failure-causing test cases are more likely to be selected, thus the testing process can be optimized. In the case study, the performance of D-DRT and other testing strategies are compared. The experiment results show that the proposed D-DRT strategy has better fault detection effectiveness than the others without significant increase in computational cost. Hanyu Pei, Beibei Yin, Kai-Yuan Cai, Min Xie 0001 |
QRS | 2 |
| 2019 | Robustness of spectrum-based fault localisation in environments with labelling perturbations
Beibei Yin, Zheng Zheng 0001, Xiao-Yi Zhang 0005, Shunkun Yang |
J. Syst. Softw. | 2 |
| 2019 | Dynamic Random Testing: Technique and Experimental EvaluationabstractA particularly good software testing strategy is to achieve the underlying testing goal while solving the problems of tradeoffs between testing effectiveness and efficiency. To improve the fault detection effectiveness of software testing, the principle of feedback control theory was adopted, which motivated the proposal of dynamic random testing (DRT). The main idea behind DRT is using the testing results to guide the test case selection to increase the selection probabilities of the subdomains with higher fault detection rates. Previous works show that DRT strategy can achieve better effectiveness than random testing strategy and random partition testing strategy, and has significantly lower computational costs than adaptive testing strategy. However, the essential factors that affect the performance of DRT, i.e., adjusting parameters, initial profile, and test case classification have not been thoroughly investigated. Besides, some experimental assumptions are inconsistent with real scenarios. Therefore, this paper gives a series of investigations on DRT with a set of practical subject programs. More specifically, the effectiveness and efficiency of DRT are presented, and the extended experiments on DRT with relevant factors are conducted. The results indicate that the effectiveness of DRT is robust to different initial profiles and affected noticeably by the adjusting parameter settings and test case classification methods. Hanyu Pei, Kai-Yuan Cai, Beibei Yin, Aditya P. Mathur, Min Xie 0001 |
IEEE Trans. Reliab. | 3 |
| 2019 | An Empirical Study of Fault Triggers in the Linux Operating System: An Evolutionary PerspectiveabstractThis paper presents an empirical study of 5741 bug reports for the Linux kernel from an evolutionary perspective, with the aim of obtaining a deep understanding of bug characteristics in the Linux operating system. Bug classification is performed based on the fault triggering conditions, followed by an analysis of the proportions and evolution of the bug types as well as comparisons among versions, products, and repair locations. In addition, an analysis of regression bugs and the relationship between the types of bugs and the time needed to fix them are presented. Moreover, a procedure for the analysis of bug type characteristics based on complex network metrics is proposed, and four network metrics, i.e., degree, clustering coefficient, betweenness, and closeness, are utilized to further investigate the relationship between bug types and software metrics. In this paper, 22 interesting findings based on the empirical results are revealed, and guidance based on these findings is provided for developers and users. Guanping Xiao, Zheng Zheng 0001, Beibei Yin, Kishor S. Trivedi, Xiaoting Du, Kai-Yuan Cai |
IEEE Trans. Reliab. | 3 |
| 2018 | Test Case Prioritization for GUI Regression Testing Based on Centrality MeasuresabstractRegression testing has been widely used in GUI software testing. For the reason of economy, the prioritization of test cases is particularly important. However, few studies discussed test case prioritization (TCP) for GUI software. Based on GUI software features, a two-layer model is proposed to assist the test case prioritization in this paper, in which, the outer layer is an event handler tree (EHT), and the inner layer is a function call graph (FCG). Compared with the conventional methods, more source code information is used based on the two-layer model for prioritization. What is more, from a global perspective, centrality measure, a complex network viewpoint is used to highlight the importance of modified functions for specific version TCP. The experiment proved the effectiveness of this model and this method. Yijie Ren, Beibei Yin |
COMPSAC (2) | 2 |
| 2017 | Adaptive Resource Allocation of Multiple Servers for Service-Based Systems in Cloud ComputingabstractDue to the advantages of cloud computing, it has been adopted as deployment platform of SBS (Service-based Systems). It also provides an elastic "pay-as-you-go" mode, which creates new resource allocation challenge that satisfying the QoS (Quality of Service) requirements with least resource allocation. There has been much interest in using feedback control to make resource allocation, but these works focus on a single control that does not take the interactions between servers share and compete for the same resource pool. In this paper, we present an adaptive resource allocation approach for SBS in the cloud environment using MIMO (Multi-Input and Multi-Output) control to allocate resource to multiple servers according to multiple workloads. The experimental results show that our approach can ensure the QoS with least resource allocation and increase the resource utilization. Siqian Gong, Beibei Yin, Kai-Yuan Cai |
COMPSAC (2) | 2 |
| 2017 | Understanding the Impacts of Influencing Factors on Time to a DataRace Software FailureabstractDatarace is a common problem on shared-memory parallel computers, including multicores. Due to its dependence on the thread scheduling scheme of its execution environment, the time to a datarace failure is usually very long. How to accelerate the occurrence of a datarace failure and further estimate the mean time to failure (MTTF) is an important topic to be studied. In this paper, the influencing factors for failures triggered by datarace bugs are explored and their influences on the time to datarace failure including the relationship with the MTTF are empirically studied. Experiments are conducted on real datarace suffering programs to verify the factors and their influences. Empirical results show that the influencing factors do have influences on the time to datarace failure of the subjects. They can be used to accelerate the occurrence of datarace failures and accurately estimate the MTTF. Kun Qiu 0001, Zheng Zheng 0001, Kishor S. Trivedi, Beibei Yin |
ISSRE | 4 |
| 2017 | Experience Report: Fault Triggers in Linux Operating System: from Evolution PerspectiveabstractLinux operating system is a complex system that is prone to suffer failures during usage, and increases difficulties of fixing bugs. Different testing strategies and fault mitigation methods can be developed and applied based on different types of bugs, which leads to the necessity to have a deep understanding of the nature of bugs in Linux. In this paper, an empirical study is carried out on 5741 bug reports of Linux kernel from an evolution perspective. A bug classification is conducted based on fault triggering conditions, followed by the analysis of the evolution of bug type proportions over versions and time, together with their comparisons across versions, products and regression bugs. Moreover, the relationship between bug type proportions and clustering coefficient, as well as the relation between bug types and time to fix are presented. This paper reveals 13 interesting findings based on the empirical results and further provides guidance for developers and users based on these findings. Guanping Xiao, Zheng Zheng 0001, Beibei Yin, Kishor S. Trivedi, Xiaoting Du, Kai-Yuan Cai |
ISSRE | 3 |
| 2017 | A Rejuvenation Strategy of Two-Granularity Software Based on Adaptive ControlabstractIn the process of continuous operation in a software system, a series of phenomena could lead to performance degradation of the system, namely software aging. The loss caused by software aging can be reduced through proper rejuvenation strategies, the key to which is to determine the rejuvenation thresholds. Essence of some traditional methods is to set predetermined thresholds based on empirical data. However, in some systems where the memory is shared between operating system and application software (two-granularity software system), as the memory consumption is closely related to system performance and changes constantly, using empirical thresholds may cause system outage or waste of resources. In this paper, an adaptive strategy is adopted to optimize the thresholds. Instead of fixed thresholds, the method regularly regulates the thresholds by taking feedback information in the running process into account. Especially, critical equations are constructed to calculate the thresholds by maximizing the system availability. Simulation results show that the proposed method achieves higher availability and more stable performance than that based on empirical thresholds. Yunyu Fang, Beibei Yin, Gao-Rong Ning, Zheng Zheng 0001, Kai-Yuan Cai |
PRDC | 2 |
| 2016 | Optimization of Two-Granularity Software Rejuvenation Policy Based on the Markov Regenerative ProcessabstractSoftware rejuvenation is a proactive software control technique that is used to improve a computing system performance when it suffers from software aging. In this paper, a two-granularity inspection-based software rejuvenation policy, which works as a closed-loop control technique, is proposed. This policy mitigates the negative impact of two-level software aging. The two levels considered are the user-level applications and the operating system. A Markov regenerative process model is constructed based on the system condition. We obtain the degradation rate of the application software and operating system from fault injection experiments. The diagnostic accuracy of the adopted monitor and analysis system, which is applied to inspect the application software and operating system, is considered as we provide the optimal rejuvenation strategies. Finally, the availability and the overall loss probability with their corresponding optimal inspection time intervals are obtained numerically based on the parameter values estimated from the experiments. Experimental results show that two-granularity software rejuvenation is much more effective than traditional single-level software rejuvenation. In our experimental study, when two-granularity software rejuvenation is used, the unavailability and the overall loss probability of the system were reduced by 17.9% and 2.65%, respectively, in comparison with the single-level rejuvenation. Gao-Rong Ning, Jing Zhao 0016, Yunlong Lou, Javier Alonso 0001, Rivalino Matias, Kishor S. Trivedi, Beibei Yin, Kai-Yuan Cai |
IEEE Trans. Reliab. | 7 |
| 2014 | Estimating confidence interval of software reliability with adaptive testing strategy
Junpeng Lv, Beibei Yin, Kai-Yuan Cai |
J. Syst. Softw. | 2 |
| 2014 | On the Asymptotic Behavior of Adaptive Testing Strategy for Software Reliability AssessmentabstractIn software reliability assessment, one problem of interest is how to minimize the variance of reliability estimator, which is often considered as an optimization goal. The basic idea is that an estimator with lower variance makes the estimates more predictable and accurate. Adaptive Testing (AT) is an online testing strategy, which can be adopted to minimize the variance of software reliability estimator. In order to reduce the computational overhead of decision-making, the implemented AT strategy in practice deviates from its theoretical design that guarantees AT's local optimality. This work aims to investigate the asymptotic behavior of AT to improve its global performance without losing the local optimality. To this end, a new AT strategy named Adaptive Testing with Gradient Descent method (AT-GD) is proposed. Theoretical analysis indicates that AT-GD, a locally optimal testing strategy, converges to the globally optimal solution as the assessment process proceeds. Simulation and experiments are set up to validate AT-GD's effectiveness and efficiency. Besides, sensitivity analysis of AT-GD is also conducted in this study. Junpeng Lv, Beibei Yin, Kai-Yuan Cai |
IEEE Trans. Software Eng. | 2 |
| 2013 | On the Gain of Measuring Test Case PrioritizationabstractTest case prioritization (TCP) techniques aim to schedule the order of regression test suite to maximize some properties, such as early fault detection. In order to measure the abilities of different TCP techniques for early fault detection, a metric named average percentage of faults detected (APFD) is widely adopted. In this paper, we analyze the metric APFD and explore the gain of measuring TCP techniques from a control theory viewpoint. Based on that, we propose a generalized metric for TCP. This new metric focuses on the gain of defining early fault detection and measuring TCP techniques for various needs in different evaluation scenarios. By adopting this new metric, not only flexibility can be guaranteed, but also explicit physical significance for the metric will be provided before evaluation. Junpeng Lv, Beibei Yin, Kai-Yuan Cai |
COMPSAC | 2 |
| 2009 | Software execution processes as an evolving complex network
Kai-Yuan Cai, Beibei Yin |
Inf. Sci. | 2 |
| 2007 | A Data Mining Approach for Software State DefinitionabstractA software system can be modeled by a transition system. In the existing approaches for software modeling, such as FSM, EFSM, and TM etc., states often have specific physics semantics which often represent variables, processes, or modules, etc. In this paper, a data mining approach is introduced into software modeling to do state definition in a different way. The approach is used to extract interesting relationships among program methods and a weighted hypergraph is constructed based on the mining results. Then the hypergraph is partitioned into k clusters which are used to define states in the transition system, using a hypergraph partitioning algorithm. States derived in this way have many particular properties. Some experiments about this approach are also presented in this paper. Beibei Yin, Chenggang Bai, Kai-Yuan Cai |
COMPSAC (1) | 1 |
| 2006 | A Case Study for Invalidating the Markovian Property of GUI Software Structural ProfileabstractSoftware reliability is one of the key factors for determining the quality of software, and a number of software reliability models have been developed for measuring it, among which a group of Markov-based models prevails. In fact, the Markovian assumption can also widely observed in performance evaluation of software system. In this paper, an experimental approach is used to check if GUI software structural profile, which is defined as transfer of control among software modules or states, possesses Markovian property, utilizing x2hypothesis test. The result of the case study indicates that the structural profile of GUI software fails to fit a Markov process Beibei Yin, Chenggang Bai, Kai-Yuan Cai |
COMPSAC (1) | 1 |