James H. Andrews

dblp:18/6535 · also Jamie Andrews · DBLP profile ↗
← Back
30ranked-venue papers
21as first author
0since 2021 · last 2015
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 23 · 15 first-authorHuman-computer interaction and ubiquitous computing · 3 · 3 first-authorTheory of computation · 3 · 3 first-authorComputer networks · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
16 papers
Software testing · 86% Debugging and program repair · 7% Software maintenance and evolution · 3%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%

Topics — the 25 heaviest of 27, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Software testing
random testing
0.242008
Random Test Run Length and Effectiveness · ASE 2008
Nighthawk: a two-level genetic-random unit test data generator · ASE 2007
Case Study of Coverage-Checked Random Data Structure Testing · ASE 2004
Software testing
test generation
0.222015
Dynamically Testing GUIs Using Ant Colony Optimization (T) · ASE 2015
Adding Value to Formal Test Oracles · ASE 2002
Software testing › test adequacy
coverage criteria
0.222013
Comparing multi-point stride coverage and dataflow coverage · ICSE 2013
Using Mutation Analysis for Assessing and Comparing Testing Coverage Criteria · IEEE Trans. Software Eng. 2006
Software testing
GUI testing
0.212015
Dynamically Testing GUIs Using Ant Colony Optimization (T) · ASE 2015
Software testing
mutation testing
0.232008
Sufficient mutation operators for measuring test effectiveness · ICSE 2008
Using Mutation Analysis for Assessing and Comparing Testing Coverage Criteria · IEEE Trans. Software Eng. 2006
Is mutation an appropriate tool for testing experiments? · ICSE 2005
Software testing › mutation testing
mutation operators
0.122008
Sufficient mutation operators for measuring test effectiveness · ICSE 2008
Using Mutation Analysis for Assessing and Comparing Testing Coverage Criteria · IEEE Trans. Software Eng. 2006
Software testing
search-based software testing
0.112011
Genetic Algorithms for Randomized Unit Testing · IEEE Trans. Software Eng. 2011
Software testing › test adequacy › coverage criteria › structural coverage criteria
dataflow coverage
0.122013
Using Mutation Analysis for Assessing and Comparing Testing Coverage Criteria · IEEE Trans. Software Eng. 2006
Comparing multi-point stride coverage and dataflow coverage · ICSE 2013
Debugging and program repair
fault localization
0.112009
Evaluating the Accuracy of Fault Localization Techniques · ASE 2009
Debugging and program repair › fault localization
fault localization evaluation
0.112009
Evaluating the Accuracy of Fault Localization Techniques · ASE 2009
Software testing › test suite evaluation
test suite effectiveness
0.112009
The influence of size and coverage on test suite effectiveness · ISSTA 2009
Software testing › regression testing
test suite reduction
0.112009
Evaluating the Accuracy of Fault Localization Techniques · ASE 2009
Software maintenance and evolution
log analysis
0.132003
General Test Result Checking with Log File Analysis · IEEE Trans. Software Eng. 2003
Broad-spectrum studies of log file analysis · ICSE 2000
Testing using Log File Analysis: Tools, Methods, and Issues · ASE 1998
Software testing › mutation testing
mutation score
0.112008
Sufficient mutation operators for measuring test effectiveness · ICSE 2008
Software testing
result verification
0.122003
General Test Result Checking with Log File Analysis · IEEE Trans. Software Eng. 2003
Broad-spectrum studies of log file analysis · ICSE 2000
Software testing
test oracle
0.122002
Adding Value to Formal Test Oracles · ASE 2002
Testing using Log File Analysis: Tools, Methods, and Issues · ASE 1998
Software testing
unit testing
0.132003
General Test Result Checking with Log File Analysis · IEEE Trans. Software Eng. 2003
Broad-spectrum studies of log file analysis · ICSE 2000
Testing using Log File Analysis: Tools, Methods, and Issues · ASE 1998
Program analysis
dynamic analysis
0.112005
Third international workshop on dynamic analysis(WODA 2005) · ICSE 2005
Software testing › mutation testing
fault injection
0.112005
Is mutation an appropriate tool for testing experiments? · ICSE 2005
Software testing
test input generation
0.022007
Nighthawk: a two-level genetic-random unit test data generator · ASE 2007
Case Study of Coverage-Checked Random Data Structure Testing · ASE 2004
Data mining › predictive modeling
classification
0.012009
Evaluating the Accuracy of Fault Localization Techniques · ASE 2009
Data mining › predictive modeling › classification
class imbalance
0.012009
Evaluating the Accuracy of Fault Localization Techniques · ASE 2009
Software testing › test adequacy › coverage criteria
structural coverage criteria
0.012009
The influence of size and coverage on test suite effectiveness · ISSTA 2009
Software testing
failure detection
0.012008
Random Test Run Length and Effectiveness · ASE 2008
Software testing
system testing
0.012000
Broad-spectrum studies of log file analysis · ICSE 2000

Methods — techniques the papers use, named apart from their topics

q-learning · 0.2ant colony optimization · 0.2genetic algorithm · 0.2data mining classifiers · 0.2cost-sensitive learning · 0.2program instrumentation · 0.2statistical analysis · 0.1feature subset selection · 0.1empirical study · 0.1random testing · 0.0
YearPublicationVenuePosition
2015 Dynamically Testing GUIs Using Ant Colony Optimization (T)
abstract
In this paper we introduce a dynamic GUI test generator that incorporates ant colony optimization. We created two ant systems for generating tests. Our first ant system implements the normal ant colony optimization algorithm in order to traverse the GUI and find good event sequences. Our second ant system, called AntQ, implements the antq algorithm that incorporates Q-Learning, which is a behavioral reinforcement learning technique. Both systems use the same fitness function in order to determine good paths through the GUI. Our fitness function looks at the amount of change in the GUI state that each event causes. Events that have a larger impact on the GUI state will be favored in future tests. We compared our two ant systems to random selection. We ran experiments on six subject applications and report on the code coverage and fault finding abilities of all three algorithms.
Santo Carino, James H. Andrews
ASE2
2014 BlackHorse: creating smart test cases from brittle recorded tests
Santo Carino, James H. Andrews, Sheldon Goulding, Pradeepan Arunthavarajah, Jakub Hertyk
Softw. Qual. J.2
2013 Killer App: A Eurogame about software quality
abstract
Games can be effective teaching tools, if the games simulate faithfully some aspect of the material being taught. Eurogames are a recent genre of games that allow for effective simulation of tradeoffs and competing strategies for achieving the same end. We have developed Killer App, a Eurogame which is focused on software development and on software quality tradeoffs. Here we report on the design of Killer App, our early experience in playtesting it with both programmers and non-programmers, and our experience of using it in a senior-year software quality assurance course.
James H. Andrews
CSEE&T1
2013 Comparing multi-point stride coverage and dataflow coverage
abstract
We introduce a family of coverage criteria, called Multi-Point Stride Coverage (MPSC). MPSC generalizes branch coverage to coverage of tuples of branches taken from the execution sequence of a program. We investigate its potential as a replacement for dataflow coverage, such as def-use coverage. We find that programs can be instrumented for MPSC easily, that the instrumentation usually incurs less overhead than that for def-use coverage, and that MPSC is comparable in usefulness to def-use in predicting test suite effectiveness. We also find that the space required to collect MPSC can be predicted from the number of branches in the program.
Mohammad Mahdi Hassan, James H. Andrews
ICSE2
2012 Generating String Test Data for Code Coverage
abstract
String data has traditionally been difficult for test data generation tools to generate. Of particular concern are strings that conform to a given grammar, since coverage of grammar productions does not guarantee high code coverage or fault-finding ability. We address this problem by deriving Java classes from the grammatical categories in the grammar, and then using a collection of deterministic and met heuristic techniques to generate strings from them. We compare the code coverage of these techniques and the standard conformance test suites for the grammars. We conclude that the various techniques have complementary strengths, and that they can be usefully used in combination.
Michael Beyene, James H. Andrews
ICST2
2011 Genetic Algorithms for Randomized Unit Testing
abstract
Randomized testing is an effective method for testing software units. The thoroughness of randomized unit testing varies widely according to the settings of certain parameters, such as the relative frequencies with which methods are called. In this paper, we describe Nighthawk, a system which uses a genetic algorithm (GA) to find parameters for randomized unit testing that optimize test coverage. Designing GAs is somewhat of a black art. We therefore use a feature subset selection (FSS) tool to assess the size and content of the representations within the GA. Using that tool, we can reduce the size of the representation substantially while still achieving most of the coverage found using the full representation. Our reduced GA achieves almost the same results as the full system, but in only 10 percent of the time. These results suggest that FSS could significantly optimize metaheuristic search-based software engineering tools.
James H. Andrews, Tim Menzies, Felix Chun Hang Li
IEEE Trans. Software Eng.1
2009 Johar: a framework for developing accessible applications
abstract
We describe the Johar framework, which supports the development of applications that are accessible to users with a wide range of abilities. A user of Johar applications chooses an "interface interpreter" that best suits them, and then uses it to interact with all Johar applications.
James H. Andrews, Fatima Hussain
ASSETS1
2009 The influence of size and coverage on test suite effectiveness
abstract
We study the relationship between three properties of test suites: size, structural coverage, and fault-finding effectiveness. In particular, we study the question of whether achieving high coverage leads directly to greater effectiveness, or only indirectly through forcing a test suite to be larger. Our experiments indicate that coverage is sometimes correlated with effectiveness when size is controlled for, and that using both size and coverage yields a more accurate prediction of effectiveness than size alone. This in turn suggests that both size and coverage are important to test suite effectiveness. Our experiments also indicate that no linear relationship exists among the three variables of size, coverage and effectiveness, but that a nonlinear relationship does exist.
Akbar Siami Namin, James H. Andrews
ISSTA2
2009 Evaluating the Accuracy of Fault Localization Techniques
abstract
We investigate claims and assumptions made in several recent papers about fault localization (FL) techniques. Most of these claims have to do with evaluating FL accuracy. Our investigation centers on a new subject program having properties useful for FL experiments. We find that Tarantula (Jones et al.) works well on the program, and we show weak support for the assertion that coverage-based test suites help Tarantula to localize faults. Baudry et al. used automatically-generated mutants to evaluate the accuracy of an FL technique that generates many distinct scores for program locations. We find no evidence to suggest that the use of mutants for this purpose is invalid. However, we find evidence that the standard method for evaluating FL accuracy is unfairly biased toward techniques that generate many distinct scores, and we propose a fairer method of accuracy evaluation. Finally, Denmat et al. suggest that data mining techniques may apply to FL. We investigate this suggestion with the data mining tool Weka, using standard techniques for evaluating the accuracy of data mining classifiers. We find that standard classifiers suffer from the class imbalance problem. However, we find that adding cost information improves accuracy.
Shaimaa Ali, James H. Andrews, Tamilselvi Dhandapani, Wantao Wang
ASE2
2008 Sufficient mutation operators for measuring test effectiveness
abstract
Mutants are automatically-generated, possibly faulty variants of programs. The mutation adequacy ratio of a test suite is the ratio of non-equivalent mutants it is able to identify to the total number of non-equivalent mutants. This ratio can be used as a measure of test effectiveness. However, it can be expensive to calculate, due to the large number of different mutation operators that have been proposed for generating the mutants.
Akbar Siami Namin, James H. Andrews, Duncan J. Murdoch
ICSE2
2008 Random Test Run Length and Effectiveness
abstract
A poorly understood but important factor in random testing is the selection of a maximum length for test runs. Given a limited time for testing, it is seldom clear whether executing a small number of long runs or a large number of short runs maximizes utility. It is generally expected that longer runs are more likely to expose failures - which is certainly true with respect to runs shorter than the shortest failing trace. However, longer runs produce longer failing traces, requiring more effort from humans in debugging or more resources for automated minimization. In testing with feedback, increasing ranges for parameters may also cause the probability of failure to decrease in longer runs. We show that the choice of test length dramatically impacts the effectiveness of random testing, and that the patterns observed in simple models and predicted by analysis are useful in understanding effects observed in a large scale case study of a JPL flight software system.
James H. Andrews, Alex Groce, Melissa Weston, Ru-Gang Xu
ASE1
2008 A Useful Bounded Resource Functional Language
Michael J. Burrell, James H. Andrews, Mark Daley
SOFSEM2
2007 Nighthawk: a two-level genetic-random unit test data generator
abstract
Randomized testing has been shown to be an effective method fortesting software units. However, the thoroughness of randomized unit testing varies widely according to the settings of certain parameters, such as the relative frequencies with which methods are called. In this paper, we describe a system which uses agenetic algorithm to find parameters for randomized unit testing that optimize test coverage. We compare our coverage results to previous work, and report on case studies and experiments on system options
James H. Andrews, Felix Chun Hang Li, Tim Menzies
ASE1
2007 An untyped higher order logic with Y combinator
abstract
Abstract We define a higher order logic which has only a notion of sort rather than a notion of type, and which permits all terms of the untyped lambda calculus and allows the use of the Y combinator in writing recursive predicates. The consistency of the logic is maintained by a distinction between use and mention, as in Gilmore's logics. We give a consistent model theory, a proof system which is sound with respect to the model theory, and a cut-elimination proof for the proof system. We also give examples showing what formulas can and cannot be used in the logic.
James H. Andrews
J. Symb. Log.1
2006 Using Mutation Analysis for Assessing and Comparing Testing Coverage Criteria
abstract
The empirical assessment of test techniques plays an important role in software testing research. One common practice is to seed faults in subject software, either manually or by using a program that generates all possible mutants based on a set of mutation operators. The latter allows the systematic, repeatable seeding of large numbers of faults, thus facilitating the statistical analysis of fault detection effectiveness of test suites; however, we do not know whether empirical results obtained this way lead to valid, representative conclusions. Focusing on four common control and data flow criteria (block, decision, C-use, and P-use), this paper investigates this important issue based on a middle size industrial program with a comprehensive pool of test cases and known faults. Based on the data available thus far, the results are very consistent across the investigated criteria as they show that the use of mutation operators is yielding trustworthy results: generated mutants can be used to predict the detection effectiveness of real faults. Applying such a mutation analysis, we then investigate the relative cost and effectiveness of the above-mentioned criteria by revisiting fundamental questions regarding the relationships between fault detection, test suite size, and control/data flow coverage. Although such questions have been partially investigated in previous studies, we can use a large number of mutants, which helps decrease the impact of random variation in our analysis and allows us to use a different analysis approach. Our results are then; compared with published studies, plausible reasons for the differences are provided, and the research leads us to suggest a way to tune the mutation analysis process to possible differences in fault detection probabilities in a specific environment
James H. Andrews, Lionel C. Briand, Yvan Labiche, Akbar Siami Namin
IEEE Trans. Software Eng.1
2005 Is mutation an appropriate tool for testing experiments?
abstract
The empirical assessment of test techniques plays an important role in software testing research. One common practice is to instrument faults, either manually or by using mutation operators. The latter allows the systematic, repeatable seeding of large numbers of faults; however, we do not know whether empirical results obtained this way lead to valid, representative conclusions. This paper investigates this important question based on a number of programs with comprehensive pools of test cases and known faults. It is concluded that, based on the data available thus far, the use of mutation operators is yielding trustworthy results (generated mutants are similar to real faults). Mutants appear however to be different from hand-seeded faults that seem to be harder to detect than real faults.
James H. Andrews, Lionel C. Briand, Yvan Labiche
ICSE1
2005 Third international workshop on dynamic analysis(WODA 2005)
abstract
Dynamic analysis techniques reason over program executions and show promise in aiding the development of robust and reliable large-scale systems. It has become increasingly clear that limitations of static analysis can be overcome by integrating static and dynamic analyses, and that the performance and value of dynamic analysis can be improved by static analysis. Hence, a key focus of the workshop will be on hybrid analyses that involve both static and dynamic components.
James H. Andrews, Lori L. Pollock
ICSE1
2005 Minimization of Randomized Unit Test Cases
abstract
We describe a framework for randomized unit testing, and give empirical evidence that generating unit test cases randomly and then minimizing the failing test cases results in significant benefits. Randomized generation of unit test cases (sequences of method calls) has been shown to allow high coverage and to be highly effective. However, failing test cases, if found, are often very long sequences of method calls. We show that Zeller and Hildebrandt's test case minimization algorithm significantly reduces the length of these sequences. We study the resulting benefits qualitatively and quantitatively, via a case study on found open-source data structures and an experiment on lab-built data structures
James H. Andrews
ISSRE2
2004 Case Study of Coverage-Checked Random Data Structure Testing
James H. Andrews
ASE1
2003 The witness properties and the semantics of the Prolog cut
abstract
The semantics of the Prolog ‘cut’ construct is explored in the context of some desirable properties of logic programming systems, referred to as the witness properties. The witness properties concern the operational consistency of responses to queries. A generalization of Prolog with negation as failure and cut is described, and shown not to have the witness properties. A restriction of the system is then described, which preserves the choice and first-solution behaviour of cut but allows the system to have the witness properties. The notion of cut in the restricted system is more restricted than the Prolog hard cut, but retains the useful first-solution behaviour of hard cut, not retained by other proposed cuts such as the ‘soft cut’. It is argued that the restricted system achieves a good compromise between the power and utility of the Prolog cut and the need for internal consistency in logic programming systems. The restricted system is given an abstract semantics, which depends on the witness properties; this semantics suggests that the restricted system has a deeper connection to logic than simply permitting some computations which are logical. Parts of this paper appeared previously in a different form in the Proceedings of the 1995 International Logic Programming Symposium (Andrews, 1995).
James H. Andrews
Theory Pract. Log. Program.1
2003 General Test Result Checking with Log File Analysis
abstract
We describe and apply a lightweight formal method for checking test results. The method assumes that the software under test writes a text log file; this log file is then analyzed by a program to see if it reveals failures. We suggest a state-machine-based formalism for specifying the log file analyzer programs and describe a language and implementation based on that formalism. We report on empirical studies of the application of log file analysis to random testing of units. We describe the results of experiments done to compare the performance and effectiveness of random unit testing with coverage checking and log file analysis to other unit testing procedures. The experiments suggest that writing a formal log file analyzer and using random testing is competitive with other formal and informal methods for unit testing.
James H. Andrews
IEEE Trans. Software Eng.1
2002 Adding Value to Formal Test Oracles
abstract
Test oracles are programs which check the output of test cases run on other programs. We describe techniques which add value to formally-defined test oracles in three ways: (a) by measuring functional coverage of test suites, (b) by giving automated support to the process of validating the oracles, and (c) by automating the generation of test cases from the oracles. The techniques involve the use of coverage measures and AI-based search algorithms. We describe the application of these techniques in the verification and validation of a complex piece of real-world software.
James H. Andrews, Vicky D. Liu
ASE1
2000 Experience Report: A Software Maintenance Project Course
abstract
A report is made on an experience of teaching a senior-year course on software maintenance, centred around a maintenance project. The main triumphs and pitfalls are recounted, and recommendations are made on project selection and general course conduct.
James H. Andrews, Hanan Lutfiyya
CSEE&T1
2000 Broad-spectrum studies of log file analysis
abstract
This paper reports on research into applying the technique of log file analysis for checking test results to a broad range of testing and other tasks. The studies undertaken included applying log file analysis to both unit- and system-level testing and to requirements of both safety-critical and non-critical systems, and the use of log file analysis in combination with other testing methods. The paper also reports on the technique of using log file analyzers to simulate the software under test, both in order to validate the analyzers and to clarify requirements. It also discusses practical issues to do with the completeness of the approach, and includes comparisons to other recently-published approaches to log file analysis.
James H. Andrews
ICSE1
1998 Testing using Log File Analysis: Tools, Methods, and Issues
abstract
Large software systems often keep log files of events. Such log files can be analyzed to check whether a run of a program reveals faults in the system. We discuss how such log files can be used in software testing. We present a framework for automatically analyzing log files, and describe a language for specifying analyzer programs and an implementation of that language. The language permits compositional, compact specifications of software, which act as test oracles; we discuss the use and efficacy of these oracles for unit- and system-level testing in various settings. We explore methodological issues such as efficiency and logging policies, and the scope and limitations of the framework. We conclude that testing using log file analysis constitutes a useful methodology for software verification, somewhere between current testing practice and formal verification methodologies.
James H. Andrews
ASE1
1997 Using a Formal Description Technique to Model Aspects of a Global Air Traffic Telecommunications Network
James H. Andrews, Nancy A. Day, Jeffrey J. Joyce
FORTE1
1997 A Logical Semantics for Depth-First Prolog with Ground Negation
James H. Andrews
Theor. Comput. Sci.1
1995 Foundational Issues in Implementing Constraint Logic Programming Systems
James H. Andrews
Sci. Comput. Program.1
1994 Foundational Issues in Implementing Constraint Logic Programming Systems
James H. Andrews
ESOP1
1989 Proof-Theoretic Characterisations of Logic Programs
James H. Andrews
MFCS1