Elaine J. Weyuker

dblp:55/1672 · DBLP profile ↗
← Back
79ranked-venue papers
29as first author
2since 2021 · last 2025
0000-0002-1660-199XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 68 · 23 first-author · 2 since 2021Theory of computation · 5 · 3 first-authorSystems, architecture and hardware · 4 · 2 first-authorSecurity and privacy · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-author
YearPublicationVenuePosition
2025 Impact of "Evaluating Software Complexity Measures"
abstract
We investigate the impact that the paper “Evaluating Software Complexity Measures” has had over the thirty-six year period since it was first published in 1988. We look both at how the frequency of citation increased, and the ways that the research has been used. We describe how citations evolved from the initial obvious intent of assessing newly proposed complexity measures to assessing qualities other than complexity, to using them to assess complexity in entities other than software code, or restricted types of programming paradigms, to the adoption of our approach or properties in fields distinct from computing.
Elaine J. Weyuker
IEEE Trans. Software Eng.1
2025 Impact of "An Applicable Family of Data Flow Testing Criteria"
abstract
We investigate the impact that the paper “An applicable family of data flow testing criteria” has had since it was published in 1988. We contrast its impact with that of an earlier paper, “Selecting software test data using data flow information” which introduced the family of data flow testing criteria and a new formal way of comparing testing criteria. We conclude that while the 1988 paper undoubtedly had a significant impact on the software testing research and practitioner communities, the earlier “Selecting software test data using data flow information” paper had significantly greater impact.
Elaine J. Weyuker
IEEE Trans. Software Eng.1
2020 Intermittently failing tests in the embedded systems domain
abstract
Software testing is sometimes plagued with intermittently failing tests and finding the root causes of such failing tests is often difficult. This problem has been widely studied at the unit testing level for open source software, but there has been far less investigation at the system test level, particularly the testing of industrial embedded systems. This paper describes our investigation of the root causes of intermittently failing tests in the embedded systems domain, with the goal of better understanding, explaining and categorizing the underlying faults. The subject of our investigation is a currently-running industrial embedded system, along with the system level testing that was performed. We devised and used a novel metric for classifying test cases as intermittent. From more than a half million test verdicts, we identified intermittently and consistently failing tests, and identified their root causes using multiple sources. We found that about 1-3% of all test cases were intermittently failing. From analysis of the case study results and related work, we identified nine factors associated with test case intermittence. We found that a fix for a consistently failing test typically removed a larger number of failures detected by other tests than a fix for an intermittent test. We also found that more effort was usually needed to identify fixes for intermittent tests than for consistent tests. An overlap between root causes leading to intermittent and consistent tests was identified. Many root causes of intermittence are the same in industrial embedded systems and open source software. However, when comparing unit testing to system level testing, especially for embedded systems, we observed that the test environment itself is often the cause of intermittence.
Per Erik Strandberg, Thomas J. Ostrand, Elaine J. Weyuker, Wasif Afzal, Daniel Sundmark
ISSTA3
2018 Automated test mapping and coverage for network topologies
abstract
Communication devices such as routers and switches play a critical role in the reliable functioning of embedded system networks. Dozens of such devices may be part of an embedded system network, and they need to be tested in conjunction with various computational elements on actual hardware, in many different configurations that are representative of actual operating networks. An individual physical network topology can be used as the basis for a test system that can execute many test cases, by identifying the part of the physical network topology that corresponds to the configuration required by each individual test case. Given a set of available test systems and a large number of test cases, the problem is to determine for each test case, which of the test systems are suitable for executing the test case, and to provide the mapping that associates the test case elements (the logical network topology) with the appropriate elements of the test system (the physical network topology).
Per Erik Strandberg, Thomas J. Ostrand, Elaine J. Weyuker, Daniel Sundmark, Wasif Afzal
ISSTA3
2017 Software Systems Engineering programmes a capability approach
Carl E. Landwehr, Jochen Ludewig, Robert Meersman, David Lorge Parnas, Peretz Shoval, Yair Wand, David M. Weiss 0001, Elaine J. Weyuker
J. Syst. Softw.8
2016 Experience Report: Automated System Level Regression Test Prioritization Using Multiple Factors
abstract
We propose a new method of determining an effective ordering of regression test cases, and describe its implementation as an automated tool called SuiteBuilder developed by Westermo Research and Development AB. The tool generates an efficient order to run the cases in an existing test suite by using expected or observed test duration and combining priorities of multiple factors associated with test cases, including previous fault detection success, interval since last executed, and modifications to the code tested. The method and tool were developed to address problems in the traditional process of regression testing, such as lack of time to run a complete regression suite, failure to detect bugs in time, and tests that are repeatedly omitted. The tool has been integrated into the existing nightly test framework for Westermo software that runs on large-scale data communication systems. In experimental evaluation of the tool, we found significant improvement in regression testing results. The re-ordered test suites finish within the available time, the majority of fault-detecting test cases are located in the first third of the suite, no important test case is omitted, and the necessity for manual work on the suites is greatly reduced.
Per Erik Strandberg, Daniel Sundmark, Wasif Afzal, Thomas J. Ostrand, Elaine J. Weyuker
ISSRE5
2016 Automated test generation using model checking: an industrial evaluation
Eduard Paul Enoiu, Adnan Causevic, Thomas J. Ostrand, Elaine J. Weyuker, Daniel Sundmark, Paul Pettersson
Int. J. Softw. Tools Technol. Transf.4
2013 The limited impact of individual developer data on software defect prediction
Robert M. Bell, Thomas J. Ostrand, Elaine J. Weyuker
Empir. Softw. Eng.3
2012 On the use of calling structure information to improve fault prediction
Yonghee Shin, Robert M. Bell, Thomas J. Ostrand, Elaine J. Weyuker
Empir. Softw. Eng.4
2011 Empirical Software Engineering Research - The Good, The Bad, The Ugly
abstract
The Software Engineering Research community has slowly recognized that empirical studies are an important way of validating ideas and increasingly our community has stopped accepting the sufficiency of arguing that a smart person has come up with the idea and therefore it must be good. This has led to a flood of Software Engineering papers that contain at least some form of empirical study. However, not all empirical studies are created equal, and many may not even provide any useful information or value. We survey the gradual shift from essentially no empirical studies, to a small number of ones of questionable value, and look at what we need to do to insure that our empirical studies really contribute to the state of knowledge in the field. Thus we have the good, the bad, and the ugly. What are we as a community doing correctly? What are we doing less well than we should be because we either don't have the necessary artifacts or because the time and resources required to do "the good" is perceived to be too great? And where are we missing the boat entirely in terms of not addressing critical questions and often not even recognizing that these questions are central even if we don't know the answers. We look to see whether we can find some commonality in the projects that have really made the transition from research to widespread practice to see whether we can identify some common themes.
Elaine J. Weyuker
ESEM1
2010 We're Finding Most of the Bugs, but What are We Missing?
abstract
We compare two types of model that have been used to predict software fault-proneness in the next release of a software system. Classification models make a binary prediction that a software entity such as a file or module is likely to be either faulty or not faulty in the next release. Ranking models order the entities according to their predicted number of faults. They are generally used to establish a priority for more intensive testing of the entities that occur early in the ranking. We investigate ways of assessing both classification models and ranking models, and the extent to which metrics appropriate for one type of model are also appropriate for the other. Previous work has shown that ranking models are capable of identifying relatively small sets of files that contain 75-95% of the faults detected in the next release of large legacy systems. In our studies of the rankings produced by these models, the faults not contained in the predicted most fault prone files are nearly always distributed across many of the remaining files; i.e., a single file that is in the lower portion of the ranking virtually never contains a large number of faults.
Elaine J. Weyuker, Robert M. Bell, Thomas J. Ostrand
ICST1
2010 Software fault prediction tool
abstract
We have developed an interactive tool that predicts fault likelihood for the individual files of successive releases of large, long-lived, multi-developer software systems. Predictions are the result of a two-stage process: first, the extraction of current and historical properties of the system, and second, application of a negative binomial regression model to the extracted data. The prediction model is presented to the user as a GUI-based tool that requires minimal input from the user, and delivers its output as an ordered list of the system's files together with an expected percent of faults each file will have in the release about to undergo system test. The predictions can be used to prioritize testing efforts, to plan code or design reviews, to allocate human and computer resources, and to decide if files should be rewritten.
Thomas J. Ostrand, Elaine J. Weyuker
ISSTA2
2010 Comparing the effectiveness of several modeling methods for fault prediction
Elaine J. Weyuker, Thomas J. Ostrand, Robert M. Bell
Empir. Softw. Eng.1
2010 Methods and opportunities for rejuvenation in aging distributed software systems
Alberto Avritzer, Robert G. Cole, Elaine J. Weyuker
J. Syst. Softw.3
2010 Editors' introduction
Elaine J. Weyuker, C. Murray Woodside
Perform. Evaluation1
2009 Does calling structure information improve the accuracy of fault prediction?
abstract
Previous studies have shown that software code attributes, such as lines of source code, and history information, such as the number of code changes and the number of faults in prior releases of software, are useful for predicting where faults will occur. In this study of an industrial software system, we investigate the effectiveness of adding information about calling structure to fault prediction models. The addition of calling structure information to a model based solely on non-calling structure code attributes provided noticeable improvement in prediction accuracy, but only marginally improved the best model based on history and non-calling structure code attributes. The best model based on history and non-calling structure code attributes outperformed the best model based on calling and non-calling structure code attributes.
Yonghee Shin, Robert M. Bell, Thomas J. Ostrand, Elaine J. Weyuker
MSR4
2008 What Can Fault Prediction Do for YOU?
Elaine J. Weyuker, Thomas J. Ostrand
TAP1
2008 Do too many cooks spoil the broth? Using the number of developers to enhance defect prediction models
Elaine J. Weyuker, Thomas J. Ostrand, Robert M. Bell
Empir. Softw. Eng.1
2007 Automating algorithms for the identification of fault-prone files
abstract
This research investigates ways of predicting which files would be most likely to contain large numbers of faults in the next release of a large industrial software system. Previous work involved making predictions using several different models ranging from a simple, fully-automatable model (the LOC model) to several different variants of a negative binomial regression model that were customized for the particular software system under study. Not surprisingly, the custom models invariably predicted faults more accurately than the simple model. However, development of customized models requires substantial time and analytic effort, as well as statistical expertise. We now introduce new, more sophisticated models that yield more accurate predictions than the earlier LOC model, but which nonetheless can be fully automated. We also extend our earlier research by presenting another large-scale empirical study of the value of these prediction models, using a new industrial software system over a nine year period.
Thomas J. Ostrand, Elaine J. Weyuker, Robert M. Bell
ISSTA2
2007 Software engineering research: from cradle to grave
abstract
Although this is a talk about the design of predictive models to determine where faults are likely to be in the next release of a large software system, the primary focus of the talk is the process that was followed when doing this type of software engineering research. We follow the project from problem inception (cradle) to productization (grave), describing each of the intermediate stages to try to give a picture of why such research takes so long, and also why it is necessary to perform each of the steps.
Elaine J. Weyuker
ESEC/SIGSOFT FSE1
2007 Ensuring system performance for cluster and single server systems
Alberto Avritzer, Andre B. Bondi, Elaine J. Weyuker
J. Syst. Softw.3
2006 Performance Assurance via Software Rejuvenation: Monitoring, Statistics and Algorithms
abstract
We present three algorithms for detecting the need for software rejuvenation by monitoring the changing values of a customer-affecting performance metric, such as response time. Applying these algorithms can improve the values of this customer-affecting metric by triggering rejuvenation before performance degradation becomes severe. The algorithms differ in the way they gather and use sample values to arrive at a rejuvenation decision. Their effectiveness is evaluated for different sets of control parameters, including sample size, using simulation. The results show that applying the algorithms with suitable choices of control parameters can significantly improve system performance as measured by the response time
Alberto Avritzer, Andre B. Bondi, Michael Grottke, Kishor S. Trivedi, Elaine J. Weyuker
DSN5
2006 Experience Developing Software Using a Globally Distributed Workforce
abstract
Industrial experience assessing the stability of a large mission-critical software project is reported. We observed that the project incurred significant additional delays in resolving the types of problems usually uncovered when assessing mission-critical software stability. We present plausible hypotheses about the possible causes of these additional delays
Alberto Avritzer, Thomas J. Ostrand, Elaine J. Weyuker
ICGSE3
2006 Looking for bugs in all the right places
abstract
We continue investigating the use of a negative binomial regression model to predict which files in a large industrial software system are most likely to contain many faults in the next release. A new empirical study is described whose subject is an automated voice response system. Not only is this system's functionality substantially different from that of the earlier systems we studied (an inventory system and a service provisioning system), it also uses a significantly different software development process. Instead of having regularly scheduled releases as both of the earlier systems did, this system has what are referred to as "continuous releases." We explore the use of three versions of the negative binomial regression model, as well as a simple lines-of-code based model, to make predictions for this system and discuss the differences observed from the earlier studies. Despite the different development process, the best version of the prediction model was able to identify, over the lifetime of the project, 20% of the system's files that contained, on average, nearly three quarters of the faults that were detected in the system's next releases.
Robert M. Bell, Thomas J. Ostrand, Elaine J. Weyuker
ISSTA3
2005 A Different View of Fault Prediction
abstract
We investigated a different mode of using the prediction model to identify the files associated with a fixed percentage of the faults. The tester could ask the tool to identify which files are likely to contain the bulks of faults, with the tester selecting any desired percentage of faults. Again the tool would return a list ordered in decreasing order of the predicted numbers of faults in the files the model expects to be most problematic. If the number of files identified is too large, the tester could reselect a smaller percentage of faults. This would make the number of files requiring particular scrutiny manageable. We expect both modes to be valuable to professional software testers and developers.
Thomas J. Ostrand, Elaine J. Weyuker, Robert M. Bell, Rachel C. W. Ostrand
COMPSAC (2)2
2005 Predicting the Location and Number of Faults in Large Software Systems
abstract
Advance knowledge of which files in the next release of a large software system are most likely to contain the largest numbers of faults can be a very valuable asset. To accomplish this, a negative binomial regression model has been developed and used to predict the expected number of faults in each file of the next release of a system. The predictions are based on the code of the file in the current release, and fault and modification history of the file from previous releases. The model has been applied to two large industrial systems, one with a history of 17 consecutive quarterly releases over 4 years, and the other with nine releases over 2 years. The predictions were quite accurate: for each release of the two systems, the 20 percent of the files with the highest predicted number of faults contained between 71 percent and 92 percent of the faults that were actually detected, with the overall average being 83 percent. The same model was also used to predict which files of the first system were likely to have the highest fault densities (faults per KLOC). In this case, the 20 percent of the files with the highest predicted fault densities contained an average of 62 percent of the system's detected faults. However, the identified files contained a much smaller percentage of the code mass than the files selected to maximize the numbers of faults. The model was also used to make predictions from a much smaller input set that only contained fault data from integration testing and later. The prediction was again very accurate, identifying files that contained from 71 percent to 93 percent of the faults, with the average being 84 percent. Finally, a highly simplified version of the predictor selected files containing, on average, 73 percent and 74 percent of the faults for the two systems.
Thomas J. Ostrand, Elaine J. Weyuker, Robert M. Bell
IEEE Trans. Software Eng.2
2004 Where the bugs are
abstract
The ability to predict which files in a large software system are most likely to contain the largest numbers of faults in the next release can be a very valuable asset. To accomplish this, a negative binomial regression model using information from previous releases has been developed and used to predict the numbers of faults for a large industrial inventory system. The files of each release were sorted in descending order based on the predicted number of faults and then the first 20% of the files were selected. This was done for each of fifteen consecutive releases, representing more than four years of field usage. The predictions were extremely accurate, correctly selecting files that contained between 71% and 92% of the faults, with the overall average being 83%. In addition, the same model was used on data for the same system's releases, but with all fault data prior to integration testing removed. The prediction was again very accurate, ranging from 71% to 93%, with the average being 84%. Predictions were made for a second system, and again the first 20% of files accounted for 83% of the identified faults. Finally, a highly simplified predictor was considered which correctly predicted 73% and 74% of the faults for the two systems.
Thomas J. Ostrand, Elaine J. Weyuker, Robert M. Bell
ISSTA2
2004 How to judge testing progress
Elaine J. Weyuker
Inf. Softw. Technol.1
2004 An AGENDA for testing relational database applications
abstract
Abstract Database systems play an important role in nearly every modern organization, yet relatively little research effort has focused on how to test them. This paper discusses issues arising in testing database systems, presents an approach to testing database applications, and describes AGENDA, a set of tools to facilitate the use of this approach. In testing such applications, the state of the database before and after the user's operation plays an important role, along with the user's input and the system output. A framework for testing database applications is introduced. A complete tool set, based on this framework, has been prototyped. The components of this system are a parsing tool that gathers relevant information from the database schema and application, a tool that populates the database with meaningful data that satisfy database constraints, a tool that generates test cases for the application, a tool that checks the resulting database state after operations are performed by a database application, and a tool that assists the tester in checking the database application's output. The design and implementation of each component of the system are discussed. The prototype described here is limited to applications consisting of a single SQL query. Copyright © 2004 John Wiley & Sons, Ltd.
David Chays, Yuetang Deng, Phyllis G. Frankl, Saikat Dan, Filippos I. Vokolos, Elaine J. Weyuker
Softw. Test. Verification Reliab.6
2004 The Role of Modeling in the Performance Testing of E-Commerce Applications
abstract
An e-commerce scalability case study is presented in which both traditional performance testing and performance modeling were used to help tune the application for high performance. This involved the creation of a system simulation model as well as the development of an approach for test case generation and execution. We describe our experience using a simulation model to help diagnose production system problems, and discuss ways that the effectiveness of performance testing efforts was improved by its use.
Alberto Avritzer, Elaine J. Weyuker
IEEE Trans. Software Eng.2
2002 The distirubtion of faults in a large industrial software system
abstract
A case study is presented using thirteen releases of a large industrial inventory tracking system. Several types of questions are addressed in this study. The first involved examining how faults are distributed over the different files. This included making a distinction between the release during which they were discovered, the lifecycle stage at which they were first detected, and the severity of the fault. The second category of questions we considered involved studying how the size of modules affected their fault density. This included looking at questions like whether or not files with high fault densities at early stages of the lifecycle also had high fault densities during later stages. A third type of question we considered was whether files that contained large numbers of faults during early stages of development, also had large numbers of faults during later stages, and whether faultiness persisted from release to release. Finally, we examined whether newly written files were more fault-prone than ones that were written for earlier releases of the product. The ultimate goal of this study is to help identify characteristics of files that can be used as predictors of fault-proneness, thereby helping organizations determine how best to use their testing resources.
Thomas J. Ostrand, Elaine J. Weyuker
ISSTA2
2001 Difficulties Measuring Software Risk in an Industrial Environment
abstract
Software risk is intended to reflect loss due to software failure. This has traditionally been computed by taking the product of two things: a probability of occurrence and the cost associated with failures. Applying these definitions in practice, however, may be much harder than it at first appears. There are two types of problems that affect the applicability and usefulness of such a computation: that the user has to know detailed information that is not normally available, and that most risk definitions do not use relevant information that is available, including information derived from testing. A definition of risk is introduced that will be usable in industrial settings. We also explore ways of incorporating information about how the software has been tested, the degree to which the software has been tested, and the observed results.
Elaine J. Weyuker
DSN1
2001 Transitioning from Academia to Industrial Research
Elaine J. Weyuker
J. Syst. Softw.1
2001 Empirical Studies of a Prediction Model for Regression Test Selection
abstract
Regression testing is an important activity that can account for a large proportion of the cost of software maintenance. One approach to reducing the cost of regression testing is to employ a selective regression testing technique that: chooses a subset of a test suite that was used to test the software before the modifications; then uses this subset to test the modified software. Selective regression testing techniques reduce the cost of regression testing if the cost of selecting the subset from the test suite together with the cost of running the selected subset of test cases is less than the cost of rerunning the entire test suite. Rosenblum and Weyuker (1997) proposed coverage-based predictors for use in predicting the effectiveness of regression test selection strategies. Using the regression testing cost model of Leung and White (1989; 1990), Rosenblum and Weyuker demonstrated the applicability of these predictors by performing a case study involving 31 versions of the KornShell. To further investigate the applicability of the Rosenblum-Weyuker (RW) predictor, additional empirical studies have been performed. The RW predictor was applied to a number of subjects, using two different selective regression testing tools, Deja vu and TestTube. These studies support two conclusions. First, they show that there is some variability in the success with which the predictors work and second, they suggest that these results can be improved by incorporating information about the distribution of modifications. It is shown how the RW prediction model can be improved to provide such an accounting.
Mary Jean Harrold, David S. Rosenblum, Gregg Rothermel, Elaine J. Weyuker
IEEE Trans. Software Eng.4
2000 Issues in Interoperability and Performance Verification in a Multi-ORB Telecommunications Environment
abstract
The rapid changes in the telecommunications environment demand that new services are developed and deployed quickly, and that legacy systems are integrated seamlessly. While emerging technologies and standards such as CORBA allow rapid solution integration via shared reusable components, new issues arise for interoperability and performance assurance in highly heterogeneous environments. In this paper, we first identify these issues in a multi-vendor and multi-ORB telecommunications environment. We then propose a systematic approach for interoperability verification using a reference ORB and industry-standard test suites. We extend our interoperability testing framework to performance verification, where the scope of the tests is detailed and the design strategy is analyzed.
Cheng J. Lin, Alberto Avritzer, Elaine J. Weyuker, Sai-Lai Lo
DSN3
2000 Testing software to detect and reduce risk
Phyllis G. Frankl, Elaine J. Weyuker
J. Syst. Softw.2
2000 Experience with Performance Testing of Software Systems: Issues, an Approach, and Case Study
abstract
An approach to software performance testing is discussed. A case study describing the experience of using this approach for testing the performance of a system used as a gateway in a large industrial client/server transaction processing application is presented.
Elaine J. Weyuker, Filippos I. Vokolos
IEEE Trans. Software Eng.1
1999 Metrics to Assess the Likelihood of Project Success Based on Architecture Reviews
Alberto Avritzer, Elaine J. Weyuker
Empir. Softw. Eng.2
1999 Evaluation techniques for improving the quality of very large software systems in a cost-effective way
Elaine J. Weyuker
J. Syst. Softw.1
1997 Re-estimation of Software Reliability After Maintenance
Andy Podgurski, Elaine J. Weyuker
ICSE2
1997 Enforcing quality of service of distributed objects
abstract
Three algorithms designed to enforce different quality of service criteria are presented, as well as empirical assessments of the algorithms for three large industrial telecommunications systems. These assessments are made in terms of the simulated performance of each system on average loads selected from operational distributions collected during beta release and field use. In addition, synthetic heavy loads designed to cause the overall CPU utilization rates to exceed 90% of capacity were run. The algorithms build on previously defined load testing algorithms, and use parameters and operational distributions computed for that purpose. This makes the quality of service enforcement algorithms particularly efficient. The primary bases for the assessment of the algorithms were the overall deviation of the response time from the average, and the fraction of service requests that were throttled from clients under varying conditions.
Alberto Avritzer, Elaine J. Weyuker
ISSRE2
1997 Monitoring Smoothly Degrading Systems for Increased Dependability
Alberto Avritzer, Elaine J. Weyuker
Empir. Softw. Eng.2
1997 Lessons Learned from a Regression Testing Case Study
David S. Rosenblum, Elaine J. Weyuker
Empir. Softw. Eng.2
1997 Comments on "Toward a Framework for Software Measurement Validation"
abstract
A view of software measurement that disagrees with the model presented by Kitchenham, Pfleeger, and Fenton (1995), is given. Whereas Kitchenham et al. argue that properties used to define measures should not constrain the scale type of measures, the authors contend that that is an inappropriate restriction. In addition, a misinterpretation of Weyuker's (1988) properties is noted.
Sandro Morasca, Lionel C. Briand, Victor R. Basili, Elaine J. Weyuker, Marvin V. Zelkowitz
IEEE Trans. Software Eng.4
1997 Using Coverage Information to Predict the Cost-Effectiveness of Regression Testing Strategies
abstract
Abstract—Selective regression testing strategies attempt to choose an appropriate subset of test cases from among a previously run test suite for a software system, based on information about the changes made to the system to create new versions. Although there has been a significant amount of research in recent years on the design of such strategies, there has been very little investigation of their cost-effectiveness. This paper presents some computationally efficient predictors of the cost-effectiveness of the two main classes of selective regression testing approaches. These predictors are computed from data about the coverage relationship between the system under test and its test suite. The paper then describes case studies in which these predictors were used to predict the cost-effectiveness of applying two different regression testing strategies to two software systems. In one case study, the TESTTUBE method selected an average of 88.1 percent of the available test cases in each version, while the predictor predicted that 87.3 percent of the test cases would be selected on average. Index Terms—Cost estimation, empirical study, regression testing, software analysis, test coverage.
David S. Rosenblum, Elaine J. Weyuker
IEEE Trans. Software Eng.2
1996 Predicting the Cost-Effectiveness of Regression Testing Strategies
abstract
Selective regression testing strategies aim at choosing an appropriate subset of test cases from among a previously run test suite for a software system, based on information about the changes made to the system to create new versions. Although there has been a significant amount of research in recent years on the design of such strategies, there has been significantly less investigation of their cost-effectiveness. In this paper some computationally efficient predictors of the cost-effectiveness of the two main classes of selective regression testing approaches are presented. A case study is described in which these predictors are used to assess the appropriateness of using a particular regression testing strategy to test multiple versions of a widely-used software system.
David S. Rosenblum, Elaine J. Weyuker
SIGSOFT FSE2
1996 Deriving Workloads for Performance Testing
abstract
An approach is presented to compare the performance of an existing production platform and a proposed replacement architecture. The traditional approach to such a comparison is to develop software for the proposed platform, build the new architecture, and collect performance measurements on both the existing system in production and the new system in the development environment. In this paper we propose a new way to design an application-independent workload for doing such a performance evaluation. We demonstrate the applicability of our approach by describing our experience using it to help an industrial organization determine whether or not a proposed architecture would be adequate to meet their organization's performance requirements.
Alberto Avritzer, Elaine J. Weyuker
Softw. Pract. Exp.2
1996 Using Failure Cost Information for Testing and Reliability Assessment
abstract
A technique for incorporating failure cost information into algorithms designed to automatically generate software-load-testing suites is presented. A previously introduced reliability measure is also modified to incorporate this cost information. examples are presented to show the usefulness of including cost information when testing or assessing software.
Elaine J. Weyuker
ACM Trans. Softw. Eng. Methodol.1
1995 Using the Consequence of Failures for Testing and Reliability Assessment
abstract
article Free Access Share on Using the consequence of failures for testing and reliability assessment Author: Elaine J. Weyuker AT&T Bell Laboratories, 600 Mountain Avenue, Murray Hill, NJ AT&T Bell Laboratories, 600 Mountain Avenue, Murray Hill, NJView Profile Authors Info & Claims ACM SIGSOFT Software Engineering NotesVolume 20Issue 4Oct. 1995 pp 81–91https://doi.org/10.1145/222132.222143Published:01 October 1995Publication History 8citation429DownloadsMetricsTotal Citations8Total Downloads429Last 12 Months14Last 6 weeks1 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Elaine J. Weyuker
SIGSOFT FSE1
1995 The Automatic Generation of Load Test Suites and the Assessment of the Resulting Software
abstract
Three automatic test case generation algorithms intended to test the resource allocation mechanisms of telecommunications software systems are introduced. Although these techniques were specifically designed for testing telecommunications software, they can be used to generate test cases for any software system that is modelable by a Markov chain provided operational profile data can either be collected or estimated. These algorithms have been used successfully to perform load testing for several real industrial software systems. Experience generating test suites for five such systems is presented. Early experience with the algorithms indicate that they are highly effective at detecting subtle faults that would have been likely to be missed if load testing had been done in the more traditional way, using hand-crafted test cases. A domain-based reliability measure is applied to systems after the load testing algorithms have been used to generate test data. Data are presented for the same five industrial telecommunications systems in order to track the reliability as a function of the degree of system degradation experienced.>
Alberto Avritzer, Elaine J. Weyuker
IEEE Trans. Software Eng.2
1995 Correction: "The Automatic Generation of Load Test Suites and the Assessment of the Resulting Software"
Alberto Avritzer, Elaine J. Weyuker
IEEE Trans. Software Eng.2
1995 Reply to "Some Critical Remarks on a Hierarchy of Fault-Detecting Abilities of Test Methods"
Phyllis G. Frankl, Elaine J. Weyuker
IEEE Trans. Software Eng.2
1994 Estimating the software reliability of smoothly degrading systems
abstract
Presents the application of a domain-based reliability measure to systems that can be represented by Markov chains. A load testing algorithm is presented, and the measure is applied to assess the reliability of these systems after they have been tested. Data are presented for three industrial telecommunications systems that had been tested using the load testing algorithm, tracking the reliability as a function of the degree of system degradation experienced.>
Alberto Avritzer, Elaine J. Weyuker
ISSRE2
1994 Generating Test Suites for Software Load Testing
Alberto Avritzer, Elaine J. Weyuker
ISSTA2
1994 A Simplified Domain-Testing Strategy
abstract
article Free Access Share on A simplified domain-testing strategy Authors: Bingchiang Jeng Sun Yat-Sen University Sun Yat-Sen UniversityView Profile , Elaine J. Weyuker New York University New York UniversityView Profile Authors Info & Claims ACM Transactions on Software Engineering and MethodologyVolume 3Issue 3July 1994 pp 254–270https://doi.org/10.1145/196092.193171Published:01 July 1994Publication History 49citation1,628DownloadsMetricsTotal Citations49Total Downloads1,628Last 12 Months60Last 6 weeks6 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Bingchiang Jeng, Elaine J. Weyuker
ACM Trans. Softw. Eng. Methodol.2
1994 Automatically Generating Test Data from a Boolean Specification
abstract
This paper presents a family of strategies for automatically generating test data for any implementation intended to satisfy a given specification that is a Boolean formula. The fault detection effectiveness of these strategies is investigated both analytically and empirically, and the costs, assessed in terms of test set size, are compared.>
Elaine J. Weyuker, Tarak Goradia
IEEE Trans. Software Eng.1
1993 An Analytical Comparison of the Fault-Detecting Ability of Data Flow Testing Techniques
Phyllis G. Frankl, Elaine J. Weyuker
ICSE2
1993 A Formal Analysis of the Fault-Detecting Ability of Testing Methods
abstract
Several relationships between software testing criteria, each induced by a relation between the corresponding multisets of subdomains, are examined. The authors discuss whether for each relation R and each pair of criteria, C/sub 1/ and C/sub 2/, R(C/sub 1/, C/sub 2/) guarantees that C/sub 1/ is better at detecting faults than C/sub 2/ according to various probabilistic measures of fault-detecting ability. It is shown that the fact that C/sub 1/ subsumes C/sub 2/ does not guarantee that C/sub 1/ is better at detecting faults. Relations that strengthen the subsumption relation and that have more bearing on fault-detecting ability are introduced.>
Phyllis G. Frankl, Elaine J. Weyuker
IEEE Trans. Software Eng.2
1993 Provable Improvements on Branch Testing
abstract
This paper compares the fault-detecting ability of several software test data adequacy criteria. It has previously been shown that if C/sub 1/ properly covers C/sub 2/, then C/sub 1/ is guaranteed to be better at detecting faults than C/sub 2/, in the following sense: a test suite selected by independent random selection of one test case from each subdomain induced by C/sub 1/ is at least as likely to detect a fault as a test suite similarly selected using C/sub 2/. In contrast, if C/sub 1/ subsumes but does not properly cover C/sub 2/, this is not necessarily the case. These results are used to compare a number of criteria, including several that have been proposed as stronger alternatives to branch testing. We compare the relative fault-detecting ability of data flow testing, mutation testing, and the condition-coverage techniques, to branch testing, showing that most of the criteria examined are guaranteed to be better than branch testing according to two probabilistic measures. We also show that there are criteria that can sometimes be poorer at detecting faults than substantially less expensive criteria.>
Phyllis G. Frankl, Elaine J. Weyuker
IEEE Trans. Software Eng.2
1993 More Experience with Data Flow Testing
abstract
Experience is provided about the cost and effectiveness of the Rapps-Weyuker data flow testing criteria. This experience is based on studies using a suite of well-known numerical programs, and supplements an earlier study (Weyuker 1990) using different types of programs. The conclusions drawn in the earlier study involving cost are confirmed in this study. New observations about tester variability and cost assessment, as well as fault detection, are also provided.>
Elaine J. Weyuker
IEEE Trans. Software Eng.1
1991 Analyzing Partition Testing Strategies
abstract
Partition testing strategies, which divide a program's input domain into subsets with the tester selecting one or more elements from each subdomain, are analyzed. The conditions that affect the efficiency of partition testing are investigated, and comparisons of the fault detection capabilities of partition testing and random testing are made. The effects of subdomain modifications on partition testing's ability to detect faults are studied.>
Elaine J. Weyuker, Bingchiang Jeng
IEEE Trans. Software Eng.1
1990 The Cost of Data Flow Testing: An Empirical Study
abstract
A family of test data adequacy criteria employing data-flow information was previously proposed, and a theoretical complexity analysis was performed. The author describes an empirical study to determine the actual cost of using these criteria. The aim is to establish the practical usefulness of these criteria in testing software and provide a basis for predicting the amount of testing needed for a given program. The first goal of the study is to confirm the belief that the family of software testing criteria considered is practical to use. An attempt is made to show that even as the program size increases, the amount of testing, expressed in terms of the number of test cases sufficient to satisfy a given criterion, remains modest. Several ways of evaluating this hypothesis are explored. The second goal is to provide the prospective user of these criteria with a way of predicting the number of test cases that will be needed to satisfy a given criterion for a given program. This provides testers with a basis for selecting the most comprehensive criterion that they can expect to satisfy. Several plausible bases for such a prediction are considered.>
Elaine J. Weyuker
IEEE Trans. Software Eng.1
1989 In Defense of Coverage Criteria
abstract
No abstract available.
Elaine J. Weyuker
ICSE1
1988 Metric Space- Based Test-Data Adequacy Criteria
abstract
Since software testing cannot ordinarily be expected to provide conclusive evidence that a program is correct, software engineers have had to be satisfied with the vague notion of a set of test data being adequate for a given program. In this paper a theoretical model is provided for the notion of adequacy. Adequacy criteria are seen as serving to distinguish a given program from a certain class of programs. In particular, notions of distance between programs are studied, and adequacy of a test set is taken to mean that the set successfully distinguishes the program being tested from all programs that are sufficiently near to it, and differ in input-output behaviour from the given program. Certain points, called critical, are identified which must occur in every adequate test set. Finally, lower bounds are obtained on the size of test sets which are minimally adequate, in the sense that they have no adequate proper subsets.
Martin D. Davis, Elaine J. Weyuker
Comput. J.2
1988 An Applicable Family of Data Flow Testing Criteria
abstract
The authors extend the definitions of the previously introduced family of data flow testing criteria to apply to programs written in a large subset of Pascal. They then define a family of adequacy criteria called feasible data flow testing criteria, which are derived from the data-flow testing criteria. The feasible data flow testing criteria circumvent the problem of nonapplicability of the data flow testing criteria by requiring the test data to exercise only those definition-use associations which are executable. It is shown that there are significant differences between the relationships among the data flow testing criteria and the relationships among the feasible data flow testing criteria. The authors discuss a generalized notion of the executability of a path through a program unit. A script of a testing session using their data flow testing tool, ASSET, is included.>
Phyllis G. Frankl, Elaine J. Weyuker
IEEE Trans. Software Eng.2
1988 An Extended Domain-Bases Model of Software Reliability
abstract
A definition of software reliability is proposed in which reliability is treated as a generalization of the probability of correctness of the software in question. A tolerance function is introduced as a method of characterizing an acceptable level of correctness. This in turn is used, together with the probability function defining the operational input distribution, as a parameter of the definition of reliability. It is shown that the definition can be used to provide many natural models of reliability by varying the tolerance function and that it may be reasonably approximated using well-chosen test sets. It is also shown that there is an inherent limitation to the measurement of reliability using finite test sets.>
Stewart N. Weiss, Elaine J. Weyuker
IEEE Trans. Software Eng.2
1988 Evaluating Software Complexity Measures
abstract
A set of properties of syntactic software complexity measures is proposed to serve as a basis for the evaluation of such measures. Four known complexity measures are evaluated and compared using these criteria. This formalized evaluation clarifies the strengths and weaknesses of the examined complexity measures, which include the statement count, cyclomatic number, effort measure, and data flow complexity measures. None of these measures possesses all nine properties, and several are found to fail to possess particularly fundamental properties; this failure calls into question their usefulness in measuring synthetic complexity.>
Elaine J. Weyuker
IEEE Trans. Software Eng.1
1986 Axiomatizing Software Test Data Adequacy
abstract
A test data adequacy criterion is a set of rules used to determine whether or not sufficient testing has been performed. A general axiomatic theory of test data adequacy is developed, and five previously proposed adequacy criteria are examined to see which of the axioms are satisfied. It is shown that the axioms are consistent, but that only two of the criteria satisfy all of the axioms.
Elaine J. Weyuker
IEEE Trans. Software Eng.1
1985 Selecting Software Test Data Using Data Flow Information
abstract
This paper defines a family of program test data selection criteria derived from data flow analysis techniques similar to those used in compiler optimization. It is argued that currently used path selection criteria, which examine only the control flow of a program, are inadequate quate. Our procedure associates with each point in a program at which a variable is defined, those points at which the value is used. Several test data selection criteria, differing in the type and number of these associations, are defined and compared.
Sandra Rapps, Elaine J. Weyuker
IEEE Trans. Software Eng.2
1984 The Complexity of Data Flow Criteria for Test Data Selection
Elaine J. Weyuker
Inf. Process. Lett.1
1984 Collecting and categorizing software error data in an industrial environment
Thomas J. Ostrand, Elaine J. Weyuker
J. Syst. Softw.2
1983 A Formal Notion of Program-Based Test Data Adequacy
Martin D. Davis, Elaine J. Weyuker
Inf. Control.2
1983 Assessing Test Data Adequacy through Program Inference
abstract
Despite the almost universal reliance on testing as the means of locating software errors and its long history of use, few criteria have been proposed for deciding when software has been thoroughly tested.As a basis for the development of usable notions of test data adequacy, an abstract definition is proposed and examined, and approximations to this definition are considered.
Elaine J. Weyuker
ACM Trans. Program. Lang. Syst.1
1982 Data Flow Analysis Techniques for Test Data Selection
Sandra Rapps, Elaine J. Weyuker
ICSE2
1982 On Testing Non-Testable Programs
abstract
A frequently invoked assumption in program testing is that there is an oracle (i.e. the tester or an external mechanism can accurately decide whether or not the output produced by a program is correct). A program is non-testable if either an oracle does not exist or the tester must expend some extraordinary amount of time to determine whether or not the output is correct. The reasonableness of the oracle assumption is examined and the conclusion is reached that in many cases this is not a realistic assumption. The consequences of assuming the availability of an oracle are examined and alternatives investigated.
Elaine J. Weyuker
Comput. J.1
1981 Parsing Regular Grammars with Finite Lookahead
Thomas J. Ostrand, Marvin C. Paull, Elaine J. Weyuker
Acta Informatica3
1980 Theories of Program Testing and the the Application of Revealing Subdomains
abstract
The theory of test data selection proposed by Goodenough and Gerhart is examined. In order to extend and refine this theory, the concepts of a revealing test criterion and a revealing subdomain are proposed. These notions are then used to provide a basis for constructing program tests.
Elaine J. Weyuker, Thomas J. Ostrand
IEEE Trans. Software Eng.1
1979 Modifications of the Program Scheme Model
Elaine J. Weyuker
J. Comput. Syst. Sci.1
1979 Translatability and Decidability Questions for Restricted Classes of Program Schemas
abstract
Two new classes of schemas are introduced: the reachable schemas and the semifree schemas. A schema is reachable if every statement in the schema is executed under some interpretation. A schema is semifree if every test in the schema is necessary in the sense that each exit of the test is taken under some interpretation. It is shown that most of the standard decision problems are unsolvable for schemas in these two classes, and that there can be no algorithm which effectively translates an arbitrary schema into an equivalent reachable or semifree schema, even though such equivalent schemas always exist. These classes are also compared to the free and liberal schemas, and interclass translatability questions are investigated. It is demonstrated that every reachable schema can be effectively translated into a semifree schema, even though it is not decidable whether a reachable schema is semifree.
Elaine J. Weyuker
SIAM J. Comput.1