Zhenyu Zhang 0004

dblp:01/1844-4 · DBLP profile ↗
← Back
37ranked-venue papers
8as first author
3since 2021 · last 2023
0000-0002-8280-8462ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 29 · 8 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 2 first-authorDatabases, data management, data science and information retrieval · 2Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1
YearPublicationVenuePosition
2023 A study on the impact of pre-trained model on Just-In-Time defect prediction
abstract
Previous researchers conducting Just-In-Time (JIT) defect prediction tasks have primarily focused on the performance of individual pre-trained models, without exploring the relationship between different pre-trained models as backbones. In this study, we build six models: RoBERTaJIT, CodeBERTJIT, BARTJIT, PLBARTJIT, GPT2JIT, and CodeGPTJIT, each with a distinct pre-trained model as its backbone. We systematically explore the differences and connections between these models. Specifically, we investigate the performance of the models when using Commit code and Commit message as inputs, as well as the relationship between training efficiency and model distribution among these six models. Additionally, we conduct an ablation experiment to explore the sensitivity of each model to inputs. Furthermore, we investigate how the models perform in zero-shot and few-shot scenarios. Our findings indicate that each model based on different backbones shows improvements, and when the backbone’s pre-training model is similar, the training resources that need to be consumed are closer. We also observe that Commit code plays a significant role in defect detection, and different pre-trained models demonstrate better defect detection ability with a balanced dataset under few-shot scenarios. These results provide new insights for optimizing JIT defect prediction tasks using pre-trained models and highlight the factors that require more attention when constructing such models. Additionally, CodeGPTJIT and GPT2JIT achieved better performance than DeepJIT and CC2Vec on the two datasets respectively under 2000 training samples. These findings emphasize the effectiveness of transformer-based pre-trained models in JIT defect prediction tasks, especially in scenarios with limited training data.
Yuxiang Guo 0004, Xiaopeng Gao, Zhenyu Zhang 0004, Wing Kwong Chan, Bo Jiang 0001
QRS3
2023 Optimizing Continuous Integration by Dynamic Test Proportion Selection
abstract
Continuous integration is widely used in modern software engineering. However, it is an expensive practice. The proposed approaches focus on either intra- or inter-build cost reduction. Test case prioritization and selection (TCPS) techniques, typical intra-build techniques, are designed to save the high cost of CI by identifying failed test cases for failed builds. However, existing works are inadequate to distinguish characteristics of builds, but to apply an identical test selection proportion to different builds. Build-prediction techniques, typical inter-build techniques, are designed to save CI cost at the build level. If a build is deemed likely to pass based on the prediction of machine learning models, the whole test suite is skipped for it. Apparently, build in such a manner may miss some realistic failed test cases, if a machine learning model provides a false negative result. In this paper, we propose a dynamic test proportion selection technique DTPS, which incorporates intra- and inter-build cost reduction techniques. DTPS uses build features to construct machine learning models to predict the probability of a specific build failure and transform the probability into the necessary test proportion, with respect to a selected test case prioritization technique. Based on the output of machine learning model, it thus selects a prioritized test suite and a variable proportion of test cases with respect to a build. We constructed a large-scale dataset with approximately 115,000 builds, and conducted a controlled experiment using the dataset. The experiment shows that DTPS outperforms existing techniques significantly. It detects 19.9% to 32.5% more failed test cases, compared with state-of-the-art techniques evaluated in the experiment. At the same time, DTPS performs better than all three existing peer techniques on approximately 47% of projects. Moreover, the experiment also shows that our failure prediction model has an improvement of 0.15 in Area Under Curve (AUC), compared to prior machine learning models.
Binyi Cui, Zhenyu Zhang 0004
SANER3
2022 WasmFuzzer: A Fuzzer for WasAssembly Virtual Machines
abstract
WebAssembly is a fast, safe, and portable low-level language suitable for diverse application scenarios.And The WebAssembly virtual machines are widely used by Web browsers or Blockchain platforms as execution engine.When there is a bug in the implementation of the Wasm virtual machine, the execution of WebAssembly may lead to errors or vulnerability in the application.Due to the grammar checks by WASM VMs, fuzzing at the binary level is ineffective to expose the bugs because most inputs cannot reach the deep logic within the WASM VM.In this work, we propose WasmFuzzer, a bytecode level fuzzing tool for WASM VMs.WasmFuzzer proposes to generate initial seeds for Fuzzing at the Wasm bytecode level and it also designs a systematic set of mutation operators for Wasm bytecode.Furthermore, WasmFuzzer proposes an adaptive mutation strategy to search for the best mutation operators for different fuzzing targets.Our evaluation on 3 real-life Wasm VMs shows that WasmFuzzer can significantly outperform AFL in terms of both code coverage and unique crash.
Bo Jiang 0001, Zichao Li 0008, Yuhe Huang, Zhenyu Zhang 0004, Wing Kwong Chan
SEKE4
2020 CUDAsmith: A Fuzzer for CUDA Compilers
abstract
CUDA is a parallel computing platform and programming model for the graphics processing unit (GPU) of NVIDIA. With CUDA programming, general purpose computing on GPU (GPGPU) is possible. However, the correctness of CUDA programs relies on the correctness of CUDA compilers, which is difficult to test due to its complexity. In this work, we propose CUDAsmith, a fuzzing framework for CUDA compilers. Our tool can randomly generate deterministic and valid CUDA kernel code with several different strategies. Moreover, it adopts random differential testing and EMI testing techniques to solve the test oracle problems of CUDA compiler testing. In particular, we lift live code injection to CUDA compiler testing to help generate EMI variants. Our fuzzing experiments with both the NVCC compiler and the Clang compiler for CUDA have detected thousands of failures, some of which have been confirmed by compiler developers. Finally, the cost-effectiveness of CUDAsmith is also thoroughly evaluated in our fuzzing experiment.
Bo Jiang 0001, Wing Kwong Chan, T. H. Tse, Yongfeng Yin, Zhenyu Zhang 0004
COMPSAC7
2020 MCFL: Improving Fault Localization by Differentiating Missing Code and Other Faults
abstract
Software testing is a popular practice to evaluate the software quality, and debugging is one of the most time-consuming tasks. In the last decades, spectrum-based fault localization (SBFL) techniques have been extensively studied and empirically shown effective in locating faults in a program. However, recent researches demonstrated that the accuracy of an SBFL technique may decrease when it is applied to a program containing code-omission faults. In this paper, we present a novel approach - MCFL. It models the behavior of code omission, embeds code-omission probes into programs to identify potential locations of missing code, captures spectra of program execution, and evaluates the suspiciousness of program entities being related to faults. Different from existing SBFL techniques, MCFL synthesizes a ranked list consisting of both suspicious statements and suspicious code-omission sites, which reflect the probability of a normal statement being faulty and the probability of missing code at specific positions in the program, respectively. We conducted a controlled experiment to compare the fault-localization accuracy of MCFL with those of four popular SBFL techniques. Six real-world projects from the dataset Defects4J are used as the experiment subjects. The experiment result showed that (i) MCFL outperforms the experimented SBFL techniques on most subjects, and on average has a 17.47% improvement; (ii) For more than 60% of the faults, MCFL successfully tells whether they are due to code omission.
Zhenyu Zhang 0004, Bo Jiang 0001
COMPSAC3
2020 PEACEPACT: Prioritizing Examples to Accelerate Perturbation-Based Adversary Generation for DNN Classification Testing
abstract
Deep neural networks (DNNs) have been widely used in classification tasks. Studies have shown that DNNs may be fooled by artificial examples known as adversaries. A common technique for testing the robustness of a classification is to apply perturbations (such as random noise) to existing examples and try many of them iteratively, but it is very tedious and time-consuming. In this paper, we propose a technique to select adversaries more effectively. We study the vulnerability of examples by exploiting their class distinguishability. In this way, we can evaluate the probability of generating adversaries from each example, and prioritize all the examples accordingly. We have conducted an empirical study using a classic DNN model on four common datasets. The results reveal that the vulnerability of examples has a strong relationship with distinguishability. The effectiveness of our technique is demonstrated through 98.90 to 99.68% improvements in the F-measure.
Jun Yan 0009, Jian Zhang 0001, Zhenyu Zhang 0004, T. H. Tse
QRS5
2020 Improving Fault-Localization Accuracy by Referencing Debugging History to Alleviate Structure Bias in Code Suspiciousness
abstract
Spectrum-based fault localization (SBFL) techniques can automatically localize software faults. They employ the program spectrum, such as code coverage profile with test verdicts, to rank the program entities based on their code suspiciousness. In the past decades, researchers have proposed many approaches to optimize these techniques; however, the program structure, which can influence their performance, is not taken into consideration in developing and improving these techniques. In this article, we identify and analyze the effect of the program structure on the application of SBFL techniques. We observe that some specific program structures may introduce structure bias to code suspiciousness and negatively influence the output of SBFL techniques. To mitigate these effects and improve the performance of fault localization, we propose Delta4Ts, a structure-aware technique. Delta4Ts references debugging history to alleviate the impact of structure bias in the calculation of code suspiciousness. It reasons from the observable suspicious value towards the desired suspicious value and the impact of structure bias. To evaluate Delta4Ts under practical constraints, we conduct a controlled experiment using nine widely-studied SBFL formulae on 12 C programs and 6 Java programs. The experiment results show that Delta4Ts can significantly improve the accuracy of the studied SBFL formulae by an average of 34.8% on 12 C programs and 30.6% on 6 Java programs, and improve more on subject programs associated with more history versions or having larger code sizes.
Yang Feng 0003, Zhenyu Zhang 0004, Wing Kwong Chan, Jian Zhang 0001, Yuming Zhou
IEEE Trans. Reliab.4
2019 ConRS: A Requests Scheduling Framework for Increasing Concurrency Degree of Server Programs
abstract
Server programs always store a great deal of data and respond to plenty of requests in a short time. Most server programs are highly concurrent. Testing concurrent programs is difficult and costly because of their non-determinism. Researches have shown that increasing the degree of concurrency is more effective for testing process. This paper uses synchronization-pair and data race to quantify the degree of concurrency. The more synchronization-pairs and data races in a trace are, the higher concurrency degree is. A requests scheduling framework ConRS is proposed in this paper to reschedule requests in an existing test case and make the program "more concurrent". Compared with server stress testing tools, ConRS includes more categories of requests. Compared with coverage-guided bug detectors and data race detectors, ConRS are lower overhead and easier to understand. The experiments on MySQL database server show the effectiveness of ConRS. The synchronization-pairs and data races of test cases have been increased by at least 10% and 30% respectively after applying ConRS.
Biyun Zhu, Ruijie Meng, Zhenyu Zhang 0004, Wing Kwong Chan
COMPSAC (1)3
2019 A Systematic Study on Factors Impacting GUI Traversal-Based Test Case Generation Techniques for Android Applications
abstract
Many test case generation algorithms have been proposed to test Android apps through their graphical user interfaces. However, no systematic study on the impact of the core design elements in these algorithms on effectiveness and efficiency has been reported. This paper presents the first controlled experiment to examine three key design factors, each of which is popularly used in GUI traversal-based test case generation techniques. These three major factors are definition of GUI state equivalence, state search strategy, and waiting time strategy in between input events. The empirical results on 33 Android apps with real faults revealed interesting results. First, different choices of GUI state equivalence led to significant difference on failure detection rate and extent of code coverage. Second, searching the GUI state hierarchy randomly is as effective as searching it systematically. Last but not the least, the choices on when to fire the next input event to the app under test is immaterial so long as the length of the test session is practically long enough such as 1 h. We also found two new GUI state equivalence definitions that are statistically as effective as the existing best strategy for GUI state equivalence.
Bo Jiang 0001, Yaoyue Zhang, Wing Kwong Chan, Zhenyu Zhang 0004
IEEE Trans. Reliab.4
2019 ART4SQLi: The ART of SQL Injection Vulnerability Discovery
abstract
SQL injection (SQLi) is one of the chief threats to the security of database-driven Web applications. It can cause serious security issues such as authentication bypassing, privacy leakage, and arbitrary code execution. Dynamic testing techniques are used in SQLi vulnerability discovery, which de-facto approach is to maintain a collection of elaborately designed user inputs (aka. attack payloads) and based on it to compose malicious SQL queries to Web applications. Such techniques are effective to reveal SQLi threats before an application is released, thus reducing the cost of manual analysis, monitoring or postdeployment of other defensive mechanisms. However, because of the diversity of SQLi attacks and the difficulty of SQLi discovery, the process to execute payloads can be costly, time-consuming, and even risky. In this paper, we approach from a test case prioritization perspective to give a more effective SQLi discovery proposal, which is based on adaptive random testing with the aim to successfully trigger an SQLi within limited attempts. To evaluate our method, we conduct an experiment using three extensively adopted open source vulnerable benchmarks. The experiment results indicate that our method ART4SQLi can effectively improve the conventional random testing approach on three common benchmarks by more than 26% in reducing the number of SQLi attempts before accomplishing a successful injection.
Donghong Zhang, Cheng-Hong Wang, Jing Zhao 0016, Zhenyu Zhang 0004
IEEE Trans. Reliab.5
2018 ReTestDroid: Towards Safer Regression Test Selection for Android Application
abstract
Mobile applications are widely used in our daily life and Android is the most popular open source mobile operating system. Because mobile applications update frequently, it is important developers to perform regression testing to ensure their quality. Modeling the control flow of an android application based on the activity lifecycle model only is imprecise for regression testing. Because many Android applications use asynchronous tasks, fragments, and native code frequently, which must be considered during change impact analysis. Otherwise, regression test selection techniques may miss some failure-revealing test cases, compromising the safety of these techniques. In this work, we propose a novel approach to model asynchronous task invocations, fragment-based activity lifecycle, and native code within the control flow graph of an Android application. Furthermore, we designed a regression test selection tool ReTestDroid based on our graph model. Our experiments on five real-life Android applications showed that our approach could enable much safer regression test selection while significantly saving regression-testing time.
Bo Jiang 0001, Yongfei Zhang, Zhenyu Zhang 0004, Wing Kwong Chan
COMPSAC (1)4
2018 The Impact of Lightweight Disassembler on Malware Detection: An Empirical Study
abstract
Malicious software poses serious threats to our lives, and the activity to detect malware is becoming more and more important. An effective approach is to train a classifier using known software samples and malware samples, and recognize malware from new software. To do that, a recent popular trend is to use OpCode, which is extracted from executable modules, as an expression of software entities to drive machine learning. However, we found that the effectiveness of such a framework highly suffers from having insufficient samples, which is caused by the low success rate of disassembly due to the intrinsic complexity of the problem. In this paper, we propose to increase the success rate of disassembly by allowing inaccurate disassembling, with the attempt to increase the number of successful disassembled samples to improve OpCode-driven malware detection. We built a lightweight disassembler D-light based on the linear swap disassembly method to avoid known issues with the recursive descent manner of IDA Pro. We carried out experiment to evaluate the performance, effectiveness, and other design factors of adopting D-light and IDA Pro as disassemblers for malware detection. The empirical study shows the D-light is both more efficient and more effective than IDA Pro in supporting malware detection.
Donghong Zhang, Zhenyu Zhang 0004, Bo Jiang 0001, T. H. Tse
COMPSAC (1)2
2017 Which Factor Impacts GUI Traversal-Based Test Case Generation Technique Most? A Controlled Experiment on Android Applications
abstract
There are many research works on automated GUI traversal-based test case generation techniques for Android application. However, the effect of different factors used in a GUI traversal algorithm has not been systematically explored. In this work, we report a controlled experiment on 33 real-world applications to expose their real failures to systematically study three major factors that are commonly observed in testing tools for this class of applications. They include the notion of GUI state equivalence, the state search (or exploration) strategy, and the amount of time to wait between two input events. Our experimental results clearly show that different notions of GUI state equivalences have significantly different effects on failure detection rate and code coverage, randomized search is comparable to systematic search, and different choices of waiting time strategies do not make significant differences in terms of testing effectiveness. We also report other interesting results in this paper.
Bo Jiang 0001, Yaoyue Zhang, Wing Kwong Chan, Zhenyu Zhang 0004
QRS4
2017 A theoretical analysis on cloning the failed test cases to improve spectrum-based fault localization
Lanfei Yan, Zhenyu Zhang 0004, Jian Zhang 0001, Wing Kwong Chan, Zheng Zheng 0001
J. Syst. Softw.3
2017 Accuracy Graphs of Spectrum-Based Fault Localization Formulas
abstract
The effectiveness of spectrum-based fault localization techniques primarily relies on the accuracy of their fault localization formulas. Theoretical studies prove the relative accuracy orders of selected formulas under certain assumptions, forming a graph of their theoretical accuracy relations. However, it is unclear whether in such a graph the relative positions of these formulas may change when some assumptions are relaxed. On the other hand, empirical studies can measure the actual accuracy of any formula in controlled settings that more closely approximate practical scenarios but in less general contexts. In this paper, we propose an empirical framework of accuracy graphs and their construction that reveal the relative accuracy of formulas. Our work not only evaluates the association between certain assumptions and the theoretical relations among formulas, but also expands our knowledge to reveal new potential accuracy relationships of other formulas which have not been discovered by theoretical analysis. Using our proposed framework, we identified a list of formula pairs in which a formula is consistently statistically more accurate than or similar in accuracy to another, enlightening directions for further theoretical analysis.
Chung Man Tang, Wing Kwong Chan, Yuen-Tak Yu, Zhenyu Zhang 0004
IEEE Trans. Reliab.4
2016 Facilitating Monkey Test by Detecting Operable Regions in Rendered GUI of Mobile Game Apps
abstract
Graphical User Interface (GUI) is a component of many software applications. Many mobile game applications in particular have to provide excellent user experiences using graphical engines to render GUI screens. On a rendered GUI screen such as a treasury map, no GUI widget is embodied in it and the operable GUI regions, each of which is a region that triggers actions when certain events acting on these regions, may only be implicitly determinable. Traditional testing tools like monkey test do not effectively generate effective event sequences over such operable GUI regions. Our insight is that operable regions in a rendered GUI screen of many mobile game applications are given with visible hints to catch user attentions. In this paper, we propose Smart Monkey, which uses the fundamental features of a screen, including color, intensity, and texture, as visual signals to detect operable GUI region candidates, and iteratively identifies and confirms the real operable GUI regions by launching GUI events to the region. We have implemented Smart Monkey as a testing tool for Android apps and conducted case studies on real-world applications to compare it with a peer technique. The empirical results show that it effective in identifying such operable regions and thus able to generate functional event sequences more efficiently.
Chenglong Sun, Zhenyu Zhang 0004, Bo Jiang 0001, Wing Kwong Chan
QRS2
2015 Cross-Project Aging Related Bug Prediction
abstract
In a long running system, software tends to encounter performance degradation and increasing failure rate during execution, which is called software aging. The bugs contributing to the phenomenon of software aging are defined as Aging Related Bugs (ARBs). Lots of manpower and economic costs will be saved if ARBs can be found in the testing phase. However, due to the low presence probability and reproducing difficulty of ARBs, it is usually hard to predict ARBs within a project. In this paper, we study whether and how ARBs can be located through cross-project prediction. We propose a transfer learning based aging related bug prediction approach (TLAP), which takes advantage of transfer learning to reduce the distribution difference between training sets and testing sets while preserving their data variance. Furthermore, in order to mitigate the severe class imbalance, class imbalance learning is conducted on the transferred latent space. Finally, we employ machine learning methods to handle the bug prediction tasks. The effectiveness of our approach is validated and evaluated by experiments on two real software systems. It indicates that after the processing of TLAP, the performance of ARB bug prediction can be dramatically improved.
Fangyun Qin, Zheng Zheng 0001, Chenggang Bai, Zhenyu Zhang 0004
QRS5
2015 A Subsumption Hierarchy of Test Case Prioritization for Composite Services
abstract
Many composite workflow services utilize non-imperative XML technologies such as WSDL, XPath, XML schema, and XML messages. Regression testing should assure the services against regression faults that appear in both the workflows and these artifacts. In this paper, we propose a refinement-oriented level-exploration strategy and a multilevel coverage model that captures progressively the coverage of different types of artifacts by the test cases. We show that by using them, the test case prioritization techniques initialized on top of existing greedy-based test case prioritization strategy form a subsumption hierarchy such that a technique can produce more test suite permutations than a technique that subsumes it. Our experimental study of a model instance shows that a technique generally achieves a higher fault detection rate than a subsumed technique, which validates that the proposed hierarchy and model have the potential to improve the cost-effectiveness of test case prioritization techniques.
Lijun Mei, Yan Cai 0001, Changjiang Jia, Bo Jiang 0001, Wing Kwong Chan, Zhenyu Zhang 0004, T. H. Tse
IEEE Trans. Serv. Comput.6
2015 Are Slice-Based Cohesion Metrics Actually Useful in Effort-Aware Post-Release Fault-Proneness Prediction? An Empirical Study
abstract
Background. Slice-based cohesion metrics leverage program slices with respect to the output variables of a module to quantify the strength of functional relatedness of the elements within the module. Although slice-based cohesion metrics have been proposed for many years, few empirical studies have been conducted to examine their actual usefulness in predicting fault-proneness. Objective. We aim to provide an in-depth understanding of the ability of slice-based cohesion metrics in effort-aware post-release fault-proneness prediction, i.e. their effectiveness in helping practitioners find post-release faults when taking into account the effort needed to test or inspect the code. Method. We use the most commonly used code and process metrics, including size, structural complexity, Halstead's software science, and code churn metrics, as the baseline metrics. First, we employ principal component analysis to analyze the relationships between slice-based cohesion metrics and the baseline metrics. Then, we use univariate prediction models to investigate the correlations between slice-based cohesion metrics and post-release fault-proneness. Finally, we build multivariate prediction models to examine the effectiveness of slice-based cohesion metrics in effort-aware post-release fault-proneness prediction when used alone or used together with the baseline code and process metrics. Results. Based on open-source software systems, our results show that: 1) slice-based cohesion metrics are not redundant with respect to the baseline code and process metrics; 2) most slice-based cohesion metrics are significantly negatively related to post-release fault-proneness; 3) slice-based cohesion metrics in general do not outperform the baseline metrics when predicting post-release fault-proneness; and 4) when used with the baseline metrics together, however, slice-based cohesion metrics can produce a statistically significant and practically important improvement of the effectiveness in effort-aware post-release fault-proneness prediction. Conclusion. Slice-based cohesion metrics are complementary to the most commonly used code and process metrics and are of practical value in the context of effort-aware post-release fault-proneness prediction.
Yibiao Yang, Yuming Zhou, Hongmin Lu, Lin Chen 0015, Zhenyu Chen 0001, Baowen Xu, Hareton K. N. Leung, Zhenyu Zhang 0004
IEEE Trans. Software Eng.8
2014 An experimental study on firewall performance: Dive into the bottleneck for firewall effectiveness
abstract
Performance is an important indicator of firewalls effectiveness, which represents capability of firewalls handling network requests. ModSecurity and iptables, two representative firewalls of packet filtering and application firewall, are studied experimentally in this paper. Firstly, we develop the experiments to test the capacity of these two kinds of firewalls. Secondly, we locate the bottlenecks for system resources such as CPU and memory usage that affect the firewalls performance by analyzing the collecting data from firewalls experiments. Finally, with the same settings, we compare the performance of the two kinds of firewalls by varying the parameters such as request rate, packet length, and maximum concurrent connections.
Cheng-Hong Wang, Donghong Zhang, Hualin Lu, Jing Zhao 0016, Zhenyu Zhang 0004, Zheng Zheng 0001
IAS5
2013 On the adoption of MC/DC and control-flow adequacy for a tight integration of program testing and statistical fault localization
Bo Jiang 0001, Ke Zhai 0002, Wing Kwong Chan, T. H. Tse, Zhenyu Zhang 0004
Inf. Softw. Technol.5
2013 A general noise-reduction framework for fault localization of Java programs
Zhenyu Zhang 0004, Wing Kwong Chan, T. H. Tse, Shanping Li
Inf. Softw. Technol.2
2012 Factorising the Multiple Fault Localization Problem: Adapting Single-Fault Localizer to Multi-fault Programs
abstract
Software failures are not rare and fault localizations always an important but laborious activity. Since there is no guarantee that no more than one fault exists in a faulty program, the approach to locate all the faults is necessary. Spectrum-based fault localization techniques collect dynamic program spectra as well as test results of program runs, and estimate the extent of program elements being related to fault(s). A popular solution into generate a ranked list of suspicious candidates, which are checked in order, stopping whenever a fault is found. Such single fault localizers locate one fault in one checking round, terminate, and wait to be triggered by the regression testing to validate the fixing of the located fault. In this paper, we study the manifestation of multiple faults in a program and propose an effective mechanism to indicate their presence. When a fault is reached during the checking round, we use it to interpret the failures observed, and update the indicator to judge whether there remain other faults in the program. Our indicator serves as a stopping criterion of checking the ranked list of suspicious candidates. Our work factories the multiple fault localization problem into developing single-fault localizers and adapting them to multi-fault programs. It both improves the fault localization efficiencies of single-fault localizers, and avoids the ineffective efforts of thoroughly abandoning the many single-fault localizers to develop multi-fault localizers.
Zheng Zheng 0001, Yunqian Zhang, Zhenyu Zhang 0004, Yunzhi Xue
APSEC4
2012 How well does test case prioritization integrate with statistical fault localization?
Bo Jiang 0001, Zhenyu Zhang 0004, Wing Kwong Chan, T. H. Tse, Tsong Yueh Chen
Inf. Softw. Technol.2
2011 Precise Propagation of Fault-Failure Correlations in Program Flow Graphs
abstract
Statistical fault localization techniques find suspicious faulty program entities in programs by comparing passed and failed executions. Existing studies show that such techniques can be promising in locating program faults. However, coincidental correctness and execution crashes may make program entities indistinguishable in the execution spectra under study, or cause inaccurate counting, thus severely affecting the precision of existing fault localization techniques. In this paper, we propose a Block Rank technique, which calculates, contrasts, and propagates the mean edge profiles between passed and failed executions to alleviate the impact of coincidental correctness. To address the issue of execution crashes, Block Rank identifies suspicious basic blocks by modeling how each basic block contributes to failures by apportioning their fault relevance to surrounding basic blocks in terms of the rate of successful transition observed from passed and failed executions. Block Rank is empirically shown to be more effective than nine representative techniques on four real-life medium-sized programs.
Zhenyu Zhang 0004, Wing Kwong Chan, T. H. Tse, Bo Jiang 0001
COMPSAC1
2011 Non-parametric statistical fault localization
Zhenyu Zhang 0004, Wing Kwong Chan, T. H. Tse, Yuen-Tak Yu, Peifeng Hu
J. Syst. Softw.1
2010 Testing in Parallel - A Need for Practical Regression Testing
Zhenyu Zhang 0004, Zijian Tong, Xiaopeng Gao
ICSOFT (2)1
2010 Fault localization through evaluation sequences
Zhenyu Zhang 0004, Bo Jiang 0001, Wing Kwong Chan, T. H. Tse
J. Syst. Softw.1
2009 Modeling and testing of cloud applications
abstract
What is a cloud application precisely? In this paper, we formulate a computing cloud as a kind of graph, a computing resource such as services or intellectual property access rights as an attribute of a graph node, and the use of a resource as a predicate on an edge of the graph. We also propose to model cloud computation semantically as a set of paths in a subgraph of the cloud such that every edge contains a predicate that is evaluated to be true. Finally, we present algorithms to compose cloud computations and a family of model-based testing criteria to support the testing of cloud applications.
Wing Kwong Chan, Lijun Mei, Zhenyu Zhang 0004
APSCC3
2009 Taming coincidental correctness: Coverage refinement with context patterns to improve fault localization
abstract
Recent techniques for fault localization leverage code coverage to address the high cost problem of debugging. These techniques exploit the correlations between program failures and the coverage of program entities as the clue in locating faults. Experimental evidence shows that the effectiveness of these techniques can be affected adversely by coincidental correctness, which occurs when a fault is executed but no failure is detected. In this paper, we propose an approach to address this problem. We refine code coverage of test runs using control- and data-flow patterns prescribed by different fault types. We conjecture that this extra information, which we call context patterns, can strengthen the correlations between program failures and the coverage of faulty program entities, making it easier for fault localization techniques to locate the faults. To evaluate the proposed approach, we have conducted a mutation analysis on three real world programs and cross-validated the results with real faults. The experimental results consistently show that coverage refinement is effective in easing the coincidental correctness problem in fault localization techniques.
Shing-Chi Cheung, Wing Kwong Chan, Zhenyu Zhang 0004
ICSE4
2009 Adaptive Random Test Case Prioritization
abstract
Regression testing assures changed programs against unintended amendments. Rearranging the execution order of test cases is a key idea to improve their effectiveness. Paradoxically, many test case prioritization techniques resolve tie cases using the random selection approach, and yet random ordering of test cases has been considered as ineffective. Existing unit testing research unveils that adaptive random testing (ART) is a promising candidate that may replace random testing (RT). In this paper, we not only propose a new family of coverage-based ART techniques, but also show empirically that they are statistically superior to the RT-based technique in detecting faults. Furthermore, one of the ART prioritization techniques is consistently comparable to some of the best coverage-based prioritization techniques (namely, the "additional" techniques) and yet involves much less time cost.
Bo Jiang 0001, Zhenyu Zhang 0004, Wing Kwong Chan, T. H. Tse
ASE2
2009 Capturing propagation of infected program states
abstract
Coverage-based fault-localization techniques find the fault-related positions in programs by comparing the execution statistics of passed executions and failed executions. They assess the fault suspiciousness of individual program entities and rank the statements in descending order of their suspiciousness scores to help identify faults in programs. However, many such techniques focus on assessing the suspiciousness of individual program entities but ignore the propagation of infected program states among them. In this paper, we use edge profiles to represent passed executions and failed executions, contrast them to model how each basic block contributes to failures by abstractly propagating infected program states to its adjacent basic blocks through control flow edges. We assess the suspiciousness of the infected program states propagated through each edge, associate basic blocks with edges via such propagation of infected program states, calculate suspiciousness scores for each basic block, and finally synthesize a ranked list of statements to facilitate the identification of program faults. We conduct a controlled experiment to compare the effectiveness of existing representative techniques with ours using standard bench-marks. The results are promising.
Zhenyu Zhang 0004, Wing Kwong Chan, T. H. Tse, Bo Jiang 0001
ESEC/SIGSOFT FSE1
2009 Where to adapt dynamic service compositions
abstract
Peer services depend on one another to accomplish their tasks, and their structures may evolve. A service composition may be designed to replace its member services whenever the quality of the composite service fails to meet certain quality-of-service (QoS) requirements. Finding services and service invocation endpoints having the greatest impact on the quality are important to guide subsequent service adaptations. This paper proposes a technique that samples the QoS of composite services and continually analyzes them to identify artifacts for service adaptation. The preliminary results show that our technique has the potential to effectively find such artifacts in services.
Bo Jiang 0001, Wing Kwong Chan, Zhenyu Zhang 0004, T. H. Tse
WWW3
2009 Test case prioritization for regression testing of service-oriented business applications
abstract
Regression testing assures the quality of modified service-oriented business applications against unintended changes. However, a typical regression test suite is large in size. Earlier execution of those test cases that may detect failures is attractive. Many existing prioritization techniques order test cases according to their respective coverage of program statements in a previous version of the application. On the other hand, industrial service-oriented business applications are typically written in orchestration languages such as WS-BPEL and integrated with workflow steps and web services via XPath and WSDL. Faults in these artifacts may cause the application to extract wrong data from messages, leading to failures in service compositions. Surprisingly, current regression testing research hardly considers these artifacts. We propose a multilevel coverage model to capture the business process, XPath, and WSDL from the perspective of regression testing. We develop a family of test case prioritization techniques atop the model. Empirical results show that our techniques can achieve significantly higher rates of fault detection than existing techniques.
Lijun Mei, Zhenyu Zhang 0004, Wing Kwong Chan, T. H. Tse
WWW2
2009 Is non-parametric hypothesis testing model robust for statistical fault localization?
Zhenyu Zhang 0004, Wing Kwong Chan, T. H. Tse, Peifeng Hu
Inf. Softw. Technol.1
2009 Resource prioritization of code optimization techniques for program synthesis of wireless sensor network applications
Zhenyu Zhang 0004, Wing Kwong Chan, T. H. Tse, Heng Lu 0001, Lijun Mei
J. Syst. Softw.1
2008 Debugging through Evaluation Sequences: A Controlled Experimental Study
abstract
Predicate-based statistical fault-localization techniques locate fault-relevant predicates in a program by contrasting the statistics of the values of individual predicates between successful and failure-causing runs. While short-circuit evaluations are common in program execution, treating predicates as atomic units ignores this fact, masking out various types of important statistics. On the contrary, are such statistics useful for debugging? In this paper, we investigate experimentally the impact of the use of short-circuit evaluation information on fault localization. The results show that, by doing so, it significantly improves predicate-based statistical fault-localization techniques.
Zhenyu Zhang 0004, Bo Jiang 0001, Wing Kwong Chan, T. H. Tse
COMPSAC1