A. Jefferson Offutt

dblp:o/AJeffersonOffutt · also Jeff Offutt · DBLP profile ↗
← Back
175ranked-venue papers
86as first author
9since 2021 · last 2025
0000-0002-8657-2557ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 160 · 81 first-author · 9 since 2021Systems, architecture and hardware · 6 · 4 first-authorHuman-computer interaction and ubiquitous computing · 6 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 3 · 1 first-authorSecurity and privacy · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 A systematic review of fault tolerance techniques for smart city applications
Kathiani Elisa de Souza, Fabiano Cutigi Ferrari, Valter Vieira de Camargo, Márcio Ribeiro 0001, A. Jefferson Offutt
J. Syst. Softw.5
2025 Retrospective on: Constraint-Based Automatic Test Data Generation
abstract
This article is a retrospective reflection on our paper Constraint-Based Automatic Test Data Generation (DeMillo and Offutt, 1991), which was selected by the editorial board of the IEEE Transactions on Software Engineering as one of the most influential papers of the second decade of the journal. Published in 1991, this paper introduced a novel approach to automated test case generation using symbolic execution and constraint-solving techniques, a fundamental shift in how to approach software testing. The retrospective explores the impact of the paper on the development of mutation testing, the creation of the Mothra toolset, the evolution of test automation practices, and the technical innovations that made the paper a cornerstone of modern testing research. Beyond its technical contributions, the article has become known for its role in fostering a culture of rigorous systems-oriented research that continues to influence both academia and industry. This reflection highlights the enduring relevance of Constraint-Based Test Case Generation and its role as a foundational work in the field of software engineering.
A. Jefferson Offutt, Richard A. DeMillo
IEEE Trans. Software Eng.1
2024 Automating GUI-based Test Oracles for Mobile Apps
abstract
In automated testing, test oracles are used to determine whether software behaves correctly on individual tests by comparing expected behavior with actual behavior, revealing incorrect behavior. Automatically creating test oracles is a challenging task, especially in domains where software behavior is difficult to model. Mobile apps are one such domain, primarily due to their event-driven, GUI-based nature, coupled with significant ecosystem fragmentation. This paper takes a step toward automating the construction of GUI-based test oracles for mobile apps, first by characterizing common behaviors associated with failures into a behavioral taxonomy, and second by using this taxonomy to create automated oracles. Our taxonomy identifies and categorizes common GUI element behaviors, expected app responses, and failures from 124 reproducible bug reports, which allow us to better understand oracle characteristics. We use the taxonomy to create app-independent oracles and report on their generalizability by analyzing an additional dataset of 603 bug reports. We also use this taxonomy to define an app-independent process for creating automated test oracles, which leverages computer vision and natural language processing, and apply our process to automate five types of app-independent oracles. We perform a case study to assess the effectiveness of our automated oracles by exposing them to 15 real-world failures. The oracles reveal 11 of the 15 failures and report only one false positive. Additionally, we combine our oracles with a recent automated test input generation tool for Android, revealing two bugs with a low false positive rate. Our results can help developers create stronger automated tests that can reveal more problems in mobile apps and help researchers who can use the understanding from the taxonomy to make further advances in test automation.
Kesina Baral, Jack Johnson, Junayed Mahmud, Sabiha Salma, Mattia Fazzini, Julia Rubin, A. Jefferson Offutt, Kevin Moran
MSR7
2023 Test Automation: From Slow & Weak to Fast, Flaky, & Blind to Smart & Effective
abstract
All technical fields add more automation over time as we replace human labor with innovative technologies. Automation comes with many advantages: It creates research opportunities, offers savings in practice, and reduces errors. Automation also comes with disruptive costs. Processes must change to accommodate the automation, and human laborers must adapt by learning new knowledge and skills. Automation also evolves over time as advances inspire more new ideas for automation. This presentation will reflect on automation through history and on years of experience inventing ways to automate software testing. The talk will review achievements in test automation, discuss challenges in cutting edge domains such as games and AI, and present open problems for future research and for practical applications.
A. Jefferson Offutt
ICST1
2023 On subsumption relationships in data flow testing
abstract
Summary Data flow testing creates test requirements as definition‐use (DU) associations, where adefinitionis a program location that assigns a value to a variable and auseis a location where that value is accessed. Data flow testing is expensive, largely because of the number of test requirements. Luckily, many DU‐associations are redundant in the sense that if one test requirement (e.g. node, edge and DU‐association) is covered, other DU‐associations are guaranteed to also be covered. This relationship is calledsubsumption. Thus, testers can save resources by only covering DU‐associations that are not subsumed by other testing requirements. Although this has the potential to significantly decrease the cost of data flow testing, there are roadblocks to its application. Finding data flow subsumptions correctly and efficiently has been an elusive goal; the savings provided by data flow subsumptions and the cost to find them need to be assessed; and the fault detection ability of a reduced set of DU‐associations and the advantages of data flow testing over node and edge coverage need to be verified. This paper presents novel solutions to these problems. We present algorithms that correctly find data flow subsumptions and are asymptotically less costly than previous algorithms. We present empirical data that show that data flow subsumption is effective at reducing the number of DU‐associations to be tested and can be found at scale. Furthermore, we found that using reduced DU‐associations decreased the fault detection ability by less than 2%, and data flow testing adds testing value beyond node and edge coverage.
Marcos Lordello Chaim, Kesina Baral, A. Jefferson Offutt, Mario Concilio, Roberto Paulo Andrioli de Araujo
Softw. Test. Verification Reliab.3
2023 On transforming model-based tests into code: A systematic literature review
abstract
Summary Model‐based test design is increasingly being applied in practice and studied in research. Model‐based testing (MBT) exploits abstract models of the software behaviour to generate abstract tests, which are then transformed into concrete tests ready to run on the code. Given that abstract tests are designed to cover models but are run on code (after transformation), the effectiveness of MBT is dependent on whether model coverage also ensures coverage of key functional code. In this article, we investigate how MBT approaches generate tests from model specifications and how the coverage of tests designed strictly based on the model translates to code coverage. We used snowballing to conduct a systematic literature review. We started with three primary studies, which we refer to as the initial seeds. At the end of our search iterations, we analysed 30 studies that helped answer our research questions. More specifically, this article characterizes how test sets generated at the model level are mapped and applied to the source code level, discusses how tests are generated from the model specifications, analyses how the test coverage of models relates to the test coverage of the code when the same test set is executed and identifies the technologies and software development tasks that are on focus in the selected studies. Finally, we identify common characteristics and limitations that impact the research and practice of MBT:(i) some studies did not fully describe how tools transform abstract tests into concrete tests,(ii) some studies overlooked the computational cost of model‐based approaches and (iii) some studies found evidence that bears out a robust correlation between decision coverage at the model level and branch coverage at the code level. We also noted that most primary studies omitted essential details about the experiments.
Fabiano Cutigi Ferrari, Vinicius H. S. Durelli, Sten F. Andler, A. Jefferson Offutt, Mehrdad Saadatmand, Nils Müllner
Softw. Test. Verification Reliab.4
2021 Graph Representation for Data Flow Coverage
abstract
Data flow testing helps testers design effective tests by requiring the tests to execute sequences of statements from definitions of variables to one or more subsequent uses. These def-use associations are derived from graphs that model software behavior. A "flow graph" that only includes paths that cover defuse associations, and not other control flows, has been defined elsewhere. Although these flow graphs have several advantages over previous graphs, as computed, they omit some valid paths, which are needed to use the graphs to discover subsumption relationships and generate test data. These omissions lead to errors in the results. This paper extends previous solutions by presenting a graph that represents all paths that cover def-use associations. The paper presents empirical data showing that this graph can be generated at reasonable cost and efficiently applied for data flow subsumption discovery.
Mario Concilio, Roberto Paulo Andrioli de Araujo, Marcos Lordello Chaim, A. Jefferson Offutt
COMPSAC4
2021 Self determination: A comprehensive strategy for making automated tests more effective and efficient
abstract
A significant change in software development over the last decade has been the growth of test automation. Most software organizations automate as many tests as possible, which not only saves time and money, but also increases reproducibility and reduces errors during testing. However, as software evolves over time, so must the test suites. For each software change, each test falls into one of four categories: (1) it needs to rerun as is, (2) it does not need to rerun, (3) it needs to change and rerun, (4) it should be deleted. This test management is currently done by hand, leading to shortcuts such as always running all tests (wasteful and expensive), deleting valuable tests that should be fixed, and not deleting unneeded tests. Over time, the test suite becomes larger and more expensive to run while also becoming steadily less effective. This project introduces a novel solution to this problem by giving individual tests the ability to self- manage through self-awareness and self-determination. Each test will encode its purpose (its test requirement), can discover what changed in the software, and then decide whether to run, not run, be changed, or self-delete. We are developing techniques and algorithms to compare syntactically two versions of the same program (previous and new) to identify differences. Tests can then check to see whether their purpose is affected by the change, and decide what to do. We have developed preliminary test framework infrastructure to be used with tests that satisfy edge coverage, based on the control flow graph. We have carried out empirical studies on open-source software to evaluate the accuracy of tests' decisions and the cost of execution. Results are encouraging, indicating strong accuracy and reasonable cost.
Kesina Baral, A. Jefferson Offutt, Fiza Mulla
ICST2
2021 Efficiently Finding Data Flow Subsumptions
abstract
Data flow testing creates test requirements as definition-use (DU) associations, where a definition is a program location that assigns a value to a variable and a use is a location where that value is accessed. Data flow testing is expensive, largely because of the number of test requirements. Luckily, many DU-associations are redundant in the sense that if one test requirement (e.g., node, edge, DU-association) is covered, other DU-associations are guaranteed to also be covered. This relationship is called subsumption. Thus, testers can save resources by only covering DU-associations that are not subsumed by other testing requirements. Although this has the potential to significantly decrease the cost of data flow testing, finding subsumption among DU-associations is quite difficult. Previous solutions are costly and contain subtle flaws that sometimes lead to incorrect results. We model the data flow testing subsumption as a data flow analysis framework, allowing us to use efficient algorithms that quickly discover data flow subsumption relationships. Experimental data suggest that the framework and algorithm can reduce the cost of data flow testing and will work at scale.
Marcos Lordello Chaim, Kesina Baral, A. Jefferson Offutt, Mario Concilio, Roberto Paulo Andrioli de Araujo
ICST3
2020 An Empirical Analysis of Blind Tests
abstract
Modern software engineers automate as many tests as possible. Test automation allows tests to be run hundreds or thousands of times: hourly, daily, and sometimes continuously. This saves time and money, ensures reproducibility, and ultimately leads to software that is better and cheaper. Automated tests must include code to check that the output of the program on the test matches expected behavior. This code is called the test oracle and is typically implemented in assertions that flag the test as passing if the assertion evaluates to true and failing if not. Since automated tests require programming, many problems can occur. Some lead to false positives, where incorrect behavior is marked as correct, and others to false negatives, where correct behavior is marked as incorrect. This paper identifies and studies a common problem where test assertions are written incorrectly, leading to incorrect behavior that is not recognized. We call these tests blind because the test does not see the incorrect behavior. Blind tests cause false positives, essentially wasting the tests. This paper presents results from several human-based studies to assess the frequency of blind tests with different software and different populations of users. In our studies, the percent of blind tests ranged from a low of 39% to a high of 95%.
Kesina Baral, A. Jefferson Offutt
ICST2
2019 Exoneration-based fault localization for SQL predicates
Yun Guo, Nan Li 0008, A. Jefferson Offutt, Amihai Motro
J. Syst. Softw.3
2019 A systematic literature review of techniques and metrics to reduce the cost of mutation testing
Alessandro Viola Pizzoleto, Fabiano Cutigi Ferrari, A. Jefferson Offutt, Leonardo Fernandes, Márcio Ribeiro 0001
J. Syst. Softw.3
2019 Testing concurrent user behavior of synchronous web applications with Petri nets
A. Jefferson Offutt, Sunitha Thummala
Softw. Syst. Model.1
2019 I love journal papers and you should too
abstract
I love
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2019 Farewell and thanks for all the reviews
abstract
Farewell and
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2018 Dazed Droids: A Longitudinal Study of Android Inter-App Vulnerabilities
abstract
Android devices are an integral part of modern life from phone to media boxes to smart home appliances and cameras. With 38.9% of market share, Android is now the most used operating system not just in terms of mobile devices but considering all OSes. As applications' complexity and features increased, Android relied more heavily on code and data sharing among apps for faster response times and richer user experience. To achieve that, Android apps reuse functionality and data by means of inter-app message passing where each app defines the messages it expects to receive. In this paper, we analyze the proliferation of exploitable inter-app communication vulnerabilities using a rich corpus of 1) a representative sample of 32 Android devices, 2) 59 official Google Android versions, and 3) the top 18,583 apps from 2016 to 2017. This corpus covers $91$ Android builds from version 4.4 to present. To the best of our knowledge, ours is the first longitudinal study looking into the propagation of vulnerabilities across AOSP builds, between AOSP and a diverse set of devices, and across app versions over a period of 13 months. To identify inter-app vulnerabilities, we developed Daze as a swift and fully-automated framework for extracting app components and fuzzing all app interfaces. Daze needs only about three hours for full-device analysis or two minutes per app on average. We identified 14,413 vulnerabilities and quantified their exposure time and the number of versions affected. Our findings revealed that $51.7%$ of Android devices and $49%$ of the top $300$ apps on Google Play contained at least one critical inter-app vulnerability. We found that about $15%$ of fixed vulnerabilities lived for more than $100$ days before being patched, more than $20%$ of unpatched vulnerabilities have existed for at least $180$ days, and $45%$ of unpatched vulnerabilities persisted through the latest two to four consecutive app versions in our dataset.
Ryan Johnson 0002, Mohamed Elsabagh, Angelos Stavrou, A. Jefferson Offutt
AsiaCCS4
2018 Using Mutant Stubbornness to Create Minimal and Prioritized Test Sets
abstract
In testing, engineers want to run the most useful tests early (prioritization). When tests are run hundreds or thousands of times, minimizing a test set can result in significant savings (minimization). This paper proposes a new analysis technique to address both the minimal test set and the test case prioritization problems. This paper precisely defines the concept of mutant stubbornness, which is the basis for our analysis technique. We empirically compare our technique with other test case minimization and prioritization techniques in terms of the size of the minimized test sets and how quickly mutants are killed. We used seven C language subjects from the Siemens Repository, specifically the test sets and the killing matrices from a previous study. We used 30 different orders for each set and ran every technique 100 times over each set. Results show that our analysis technique performed significantly better than prior techniques for creating minimal test sets and was able to establish new bounds for all cases. Also, our analysis technique killed mutants as fast or faster than prior techniques. These results indicate that our mutant stubbornness technique constructs test sets that are both minimal in size, and prioritized effectively, as well or better than other techniques.
Loreto Gonzalez-Hernandez, Birgitta Lindström, A. Jefferson Offutt, Sten F. Andler, Pasqualina Potena, Markus Bohlin
QRS3
2018 Automatically Repairing SQL Faults
abstract
SQL is the standard database language, yet SQL statements can be complex and expensive to debug by hand. Automatic program repair techniques have the potential to reduce cost significantly. A previous attempt to repair SQL faults automatically used a decision tree (DT) algorithm that succeeded in some cases, but also generated many patches that passed the automated tests but that were not acceptable to the engineers. This paper proposes a novel fault localization and repair technique to repair faulty SQL statements. It targets faults in two common SQL constructs, JOIN and WHERE. It identifies the fault location and type precisely, and then creates a patch to fix the fault. We implemented this technique in a tool, and evaluated it on five medium to large-scale databases using 825 faulty queries with various complexity and faulty types. Experimental results showed that this technique can identify and repair JOIN faults when the DT approach is infeasible, and repair WHERE faults at about the same rate as the DT approach. Moreover, patches generated by our approach are more acceptable to engineers, and the tool is much faster.
Yun Guo, Nan Li 0008, A. Jefferson Offutt, Amihai Motro
QRS3
2018 Reducing the Cost of Android Mutation Testing
abstract
Due to the high market share of Android mobile devices, Android apps dominate the global market in terms of users, developers, and app releases.However, the quality of Android apps is a significant problem.Previously, we developed a mutation analysis-based approach to testing Android apps and showed it to be very effective.However, the computational cost of Android mutation testing is very high, possibly limiting its practical use.This paper presents a cost-reduction approach based on identifying redundancy among mutation operators used in Android mutation analysis.Excluding them can reduce cost without affecting the test quality.We consider a mutation operator to be redundant if tests designed to kill other types of mutants can also kill all or most of the mutants of this operator.We conducted an empirical study with selected open source Android apps.The results of our study show that three operators are redundant and can be excluded from Android mutation analysis.We also suggest updating one operator's implementation to stop generating trivial mutants.Additionally, we identity subsumption relationships among operators so that the operators subsumed by others can be skipped in Android mutation analysis.
Lin Deng 0001, A. Jefferson Offutt
SEKE2
2018 Experimental Evaluation of Redundancy in Android Mutation Testing
abstract
Because of the widespread usage of Android devices, the Android ecosystem has the highest numbers of users, developers, and app downloads. Researchers find that many Android apps are not sufficiently tested, which may lead to crashes, incorrect behaviors, and security vulnerabilities. Mutation testing is a syntax-based software testing technique that is very effective at designing high-quality tests and evaluating pre-existing tests. Our prior research designed and implemented Android mutation testing technique, and then used experiments to assess its strength. However, the high computational cost of Android mutation testing possibly limits its industrial application. This paper presents an experimental evaluation that investigates redundant mutation operators in Android mutation analysis. While maintaining the test quality, our goal is to reduce the cost by excluding redundant mutation operators or improving their design and implementation. In our evaluation, we first generate mutants and design mutation-adequate tests for each mutation operator. Then, we compute redundancy scores for each pair of mutation operators. Our evaluation results indicate that three operators (AODU, AOIU, and LOI) are redundant in Android mutation analysis. Other three operators (FOB, TVD, and ORL) are very hard to kill. One operator (MDL) needs improvement in its design to eliminate trivial mutants. We also identity subsumption relationships among operators (BWS subsumes BWD, ODL subsumes CDL, COD, and VDL).
Lin Deng 0001, A. Jefferson Offutt
Int. J. Softw. Eng. Knowl. Eng.2
2018 An experimental comparison of edge, edge-pair, and prime path criteria
Vinicius H. S. Durelli, Márcio Eduardo Delamaro, A. Jefferson Offutt
Sci. Comput. Program.3
2018 Editorial: Do we need to teach ethics to PhD students?
abstract
investigates the common problem of mobile apps that fail when we turn our devices sideways.(Recommended by Marcio Delamaro.)"CoopREP: Cooperative Record and Replay of Concurrency Bugs," by Nuno Machado, Paolo Romano, and Luís Rodrigues, presents a system that addresses the difficult problem of replicating concurrent faults, which often appear to be random.
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2018 Editorial: Self-plagiarism is not a thing
abstract
This issue contains two interesting papers that address deep and difficult problems in testing. Heterogeneous fault prediction with cost-sensitive domain adaptation, by Li, Jing, and Zhu, proposes a new approach to predicting faults that incorporates cost. (Recommended by Hyunsook Do.) Improving lazy abstraction for SCR specifications through constraint relaxation, by Degiovanni, Ponzio, Aguirre, and Frias, presents a technique that makes it possible to apply model-checking to requirements specifications that cover large input domains and that include nondeterminism. (Recommended by Shaoying Liu.) I've been ranting on ethics and plagiarism for awhile, both here and elsewhere 1, 2. I have been plagiarised 3, 4, I have been accused of plagiarising (although I did not!), and I deal with plagiarism cases regularly as reviewer and editor. I've also been asked to speak about ethics in publishing to PhD students and young faculty at my university. This is a short rant in the other direction—there is no such thing as self plagiarism! “Taking someone else's work or ideas and passing them off as one's own” — Oxford Dictionary. “To use the words or ideas of another person as if they were your own words or ideas” — Merriam-Webster Dictionary. “an act or instance of using or closely imitating the language and thoughts of another author without authorization and the representation of that author's work as one's own, as by not crediting the original author.” — Dictionary.com The definition of plagiarism makes it quite simple … copying from ourselves is not plagiarism. We are not using “someone else's” work, we are using our own. There could be a copyright problem, depending on who owns the copyright for the paper. But that's very different from plagiarism and not usually something reviewers should be concerned with. There is not much grey area on this topic, but I did find one. I will avoid using names, and just use initials. ‘A' & ‘O' wrote a paper together. O then reused a related work paragraph in another paper with another co-author, ‘L.’ Later, L reused the same paragraph in a draft of a single-author paper. Despite a long talk, A and O could not decide if that secondary reuse amounted to “transitive plagiarism.” Clearly, L had not written the paragraph originally, but by reusing it in a paper with L, O had in some sense given credit to L for the paragraph. In the end, we decided to avoid risk and had L rewrite the questionable paragraph. Back to the original question, we are certainly allowed to reuse our own words. So let's all say this one last time together: “self-plagiarism.” Now, please, never say it again. It's not a thing. Now let's get back to the real issue: is the research sound?
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2018 Proper references is a matter of scholarship, ethics, and courtesy
abstract
Proper references is a matter of
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2018 Why don't we publish more TDD research papers?
abstract
Why don't we publish more TDD research papers?please find a way to infuse TDD into your courses.Modern introductory programming courses now universally include test automation (using JUnit or one of its variants).If you teach one of those courses, have them go through at least one TDD exercise.It's a great exercise in a lab with small teams, where they can go through one test cycle in 5 or 10 minutes.If you teach a general software engineering class, make sure
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2018 What is the value of the peer-reviewing system?
abstract
This issue contains two excellent papers that present novel techniques to improve reliability models and to decrease test suite execution time.
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2018 What is a facade journal?
abstract
presents a new method that uses network theory to estimate reliability of component-based software.
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2018 How can we recognize facade journals?
abstract
How can we recognize
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2018 Why do people publish in facade journals?
abstract
This paper contains two excellent papers on testing non-traditional aspects of software. MobSTer: A model-based security testing framework for web applications, by Michele Peroli, Federico De Meo, Luca Viganò, and Davide Guardini, uses model checking to evaluate security of web applications. (Recommended by Alex Pretschner.) An automated functional testing approach for virtual reality applications, by Alinne C. Corrêa Souza, Fátima L. S. Nunes, and Márcio E. Delamaro, presents a method to test virtual reality applications, a growing area for which testing is quite different from traditional software. (Recommended by Lori Pollock.) I'm happy to say that neither of these papers is published in a facade journal. This editorial completes a series about what I call facade journals. I started by arguing that peer reviews are essential to ensuring quality of published papers 1. I next introduced the term “facade journals,” which pretend to publish quality science but do not 2. Others have called these “predatory journals” 3 4 5, but as I point out here, many journals that publish papers with little or no scientific quality are not always predatory. Most recently, I gave suggestions for how to recognize facade journals 6. The classic book Ender's Game 7 taught me that understanding the competition helps us win. This echoes advice from the great Chinese general, Sun Tzu: “To know your enemy, you must become your enemy.” Before the scientific community can respond to facade journals, we must first understand them. Why do they exist? Why do people publish in them? Let me walk through it. It is hard to be good at research. It takes years of study, and more years of apprenticing to an established scientist. Success requires wisdom to find worthwhile research problems, creativity to invent new solutions, objectivity to evaluate the ideas, communication skills to disseminate the research, resilience to bounce back from failure and criticism, and integrity to avoid shortcuts. Research also takes hard work—many hours in laboratories, libraries, the field, and in front of computers. In many fields (including software engineering), we can make more money with less effort in non-research jobs. With all this, shortcuts are tempting. I use the word “cheat” in a broad sense—stealing, plagiarism, lying, etc. In my experience, a small percentage of people will never cheat, a small percentage will usually cheat, but most are in the middle and will sometimes cheat. Some will cheat if they know they won't get caught, if the consequences of getting caught are low, or if the benefit is high. But many, maybe most, people will cheat if we think the rules are unfairly stacked against us. How many of us drive faster than the speed limit? The speed limits seem ridiculously low, we probably won't get caught, and if we do it's usually a small fine. Paying taxes seems unfair, so many people exaggerate on their tax forms. 30 years of teaching have taught me that if assignments are reasonable, grading is fair, and the professor is skilled and supportive, very few students will cheat. But if they think the system is unfair, many students will cheat. Maybe even most. Publishing in facade journals is a type of cheating. Authors claim credit for scientific publications, when in reality the paper does not advance human knowledge. But why do they publish these papers? Not because they are completely unethical or sociopaths. Most authors publish in facade journals because they think the system is rigged against them! Many universities require professors to publish in international journals and conferences. Positive motivators include raises, bonuses, lower teaching loads, promotion, and other perqs. Negative motivators include reductions in pay, higher teaching loads, and termination. Some universities provide a lot of support, but some don't. Many universities, especially in developing countries, cannot hire professors with strong research training; instead they hire professors who had little or no effective academic training. If your advisor didn't know how to do international quality research, how can you learn? If your dissertation was not competitive, how can you start doing great research after graduation? If a young scientist at a small university in a poor country could not get adequate research training, but is then put under enormous pressure to publish in international journals and conferences, there is no doubt the system is rigged against him! The only path to success is to cheat. Like many other things, cheating has been disrupted by the web 8. First, the journals can compare submitted papers with thousands of published papers. Not only can conferences check automatically, they are also very expensive: travel, visas, and high registration fees bias the field against professors from poor universities. These avenues for cheating may be more limited than in the past, but the web also offers a solution. Instead of the 1990s method of putting ink on paper and mailing thick stacks of paper, journals just put bits on a server. Paper was expensive, but bits? They are cheap. Thus facade journals. If scientists have no hope of following the rules in a system that is rigged against them, facade journals may be their only path to success. So I suggest that facade journals serve a useful purpose—they give some professors a chance. Ender's Game 7 had a quote that I find very compelling at this point: “In the moment when I truly understand my enemy, understand him well enough to defeat him, then in that very moment I also love him.” It's easy for rich people to take a high moral ground and insist that cheaters be harshly published. But how many of us would steal food to feed our starving children? If no matter how hard you work and how well you follow the rules, you can't win because the rules are unfair, would you take a shortcut? If you can't publish in STVR or ICST, but your promotion committee, your chair, and your dean will accept facade publications as real, would you? I would like to claim I wouldn't cheat, but the privilege I was granted by being born in a rich country with a strong education system means I have never been in that position. I'm reluctant to claim that moral superiority without being put to the test. So yes, facade journals serve a purpose. I don't like the purpose, but like any drug, they will not go away as long as there is a demand. What can successful scientists with integrity do? We have a professional responsibility to educate our students about which journals (and conferences) are “real” and which are fake. Not in an elitist, qualitative, manner of “conference X is better than Y because X rejects more papers,” but in the quantitative manner of “this journal publishes papers that advance human knowledge, but that journal does not.” So I've come full circle to the point in my first essay on this topic [1]: the value of a journal is in the rigorous but fair reviewing. The peer review system has survived for centuries, and I'm confident it will survive the web and facade journals.
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2017 A Novel Self-Paced Model for Teaching Programming
abstract
The Self-Paced Learning Increases Retention and Capacity (SPARC) project is responding to the well-documented surge in CS enrollment by creating a self-paced learning environment that blends online learning, automated assessment, collaborative practice, and peer-supported learning. SPARC delivers educational material online, encourages students to practice programming in groups, frees them to learn material at their own pace, and allows them to demonstrate proficiency at any time. This model contrasts with traditional course offerings, which impose a single schedule of due dates and exams for all students. SPARC allows students to complete courses faster or slower at a pace tailored to the individual, thereby allowing universities to teach more students with the same or fewer resources. This paper describes the goals and elements of the SPARC model as applied to CS1. We present results so far and discuss the future of the project.
A. Jefferson Offutt, Paul Ammann, Kinga Dobolyi, Chris Kauffmann, Jaime Lester, Upsorn Praphamontripong, Huzefa Rangwala, Sanjeev Setia, Pearl Y. Wang, Liz White
L@S1
2017 Is Mutation Analysis Effective at Testing Android Apps?
abstract
Not only is Android the most widely used mobile operating system, more apps have been released and downloaded for Android than for any other OS. However, quality is an ongoing problem, with many apps being released with faults, sometimes serious faults. Because the structure of mobile app software differs from other types of software, testing is difficult and traditional methods do not work. Thus we need different approaches to test mobile apps. In this paper, we identify challenges in testing Android apps, and categorize common faults according to fault studies. Then, we present a way to apply mutation testing to Android apps. Additionally, this paper presents results from two empirical studies on fault detection effectiveness using open-source Android applications: one for Android mutation testing, and another for four existing Android testing techniques. The studies use naturally occurring faults as well as crowdsourced faults introduced by experienced Android developers. Our results indicate that Android mutation testing is effective at detecting faults.
Lin Deng 0001, A. Jefferson Offutt, David Samudio
QRS2
2017 Mutation operators for testing Android apps
Lin Deng 0001, A. Jefferson Offutt, Paul Ammann, Nariman Mirzaei
Inf. Softw. Technol.2
2017 Using mutation to design tests for aspect-oriented models
Birgitta Lindström, A. Jefferson Offutt, Daniel Sundmark, Sten F. Andler, Paul Pettersson
Inf. Softw. Technol.2
2017 Editorial: Accepting shortened papers hurts science
abstract
by Jesús M Almendros-Jiménez and Antonio Becerra-Terón, presents a tool that automatically generates tests as XML strings for XQuery programs.
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2017 Is paper reviewing a transaction, a service, or an opportunity?
abstract
This issue contains two papers. Assessment of C++ Object-Oriented Mutation Operators: A Selective Mutation Approach, by Pedro Delgado-Pérez, Sergio Segura, and Inmaculada Medina-Bulo, presents results from an empirical study of redundancy and test quality of class-level mutation operators, finding that choosing operators that ranked higher in both measures reduced the number of mutants without reducing test effectiveness. (Recommended by Paul Ammann.) An Automated Framework to Support Testing for Process-Level Race Conditions, by Tingting Yu, Witty Srisa-an, and Gregg Rothermel, presents dynamic analysis algorithms to detect race conditions at the process level. The paper also presents results from a test framework that found race conditions on 24 applications, with reasonable overhead. (Recommended by Sudipto Ghosh.) We all know the value of the peer review system. Reviews help filter good research from bad. They also improve the research and the presentation of the papers. Without high-quality peer reviewing, readers would need to sift through thousands of uninteresting papers to find two or three that inform them of interesting new ideas and important results. That's the benefit to readers, but why do respected scientists review papers? After all, a good review of a complicated paper requires hours of work, and journals do not pay for that time. At a recent meeting of Wiley editors, I heard a discussion of peer reviewing as if it was purely a transactional activity. That is, I submit a paper, and the editor finds three scientists to review my paper. Therefore, I owe that journal three reviews of other papers. In fact, one editor claimed that he had declined (desk-rejected) a paper because the author had refused more than a dozen invitations to review. I won't call that wrong. But in my opinion, viewing reviewing purely as a tit-for-tat trade, a transaction, misses some very important points. Looking at reviewing as a service turns the work into something much nobler. If we succeed as scientists, then reviewing is a way to give back to the community. It's a way to improve the field to help others succeed, in the same way that the field helped us succeed. But this view still misses an essential benefit of reviewing. Reviewing is an opportunity to learn about important and interesting results early. From a purely self-interest point of view, reviewing has helped my career enormously. Even better, reviewing has taught me more about writing papers than I learned as a student. Reviewing papers taught me about the process and mechanics of carrying out research, how to frame results, and how to present results for public consumption. To paraphrase an old saying: “Most people don't recognize opportunity when it is disguised as hard work.” Journal editors have universally noticed that it is becoming harder to convince scientists to review papers than in the past. We don't know why. I have heard that some advisers tell their former students to never review, or only review two or three papers per year. That's just bad advice. Sure, reviewing takes time. And that time could be spent writing papers or grant proposals. But especially for young scientists, time spent reviewing papers is an investment with a huge payoff. Reviewing helps you write papers faster. Reviewing helps you write first versions better, reducing revisions and rejections. Most importantly, reviewing papers is an opportunity for personal growth; it makes you a better scientist who is capable of having a larger impact on the field. If that's not your goal, then maybe it should be.
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2017 Editorial: Beware of predatory journals
abstract
Debugging-
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2017 Color figures considered harmful
abstract
This issue contains two papers about test automation. Impediments for Software Test Automation: A Systematic Literature Review, by Wiklund, Eldh, Sundmark, and Lundqvist, provides answers for what makes it difficult to apply test automation through a well thought out tour of relevant literature. (Recommended by Bogden Korel.) GUICop: Approach and Toolset for Specification-based GUI Testing, by Hammoud, Zaraket, and Masri, presents a solution to the test oracle problem for GUIs. Small changes in the appearance of a GUi can fool automated tests into thinking the new screen is incorrect. This work addresses the problem by using a user-defined GUI specification. (Recommended by Jane Hayes.) I want to give a very short and direct suggestion to help my less experienced colleagues. I just finished reviewing 15 conference submissions, and nine have figures that are either difficult or impossible to read. Five use color to differentiate elements in the graph. Six use color that prints as a very pale gray in B&W, and therefore cannot be read. Five shrunk a figure to fit into a two-column format, making the words in the figure unreadable. Multiple reviewers are calling the authors out for these mistakes, and at least three papers will be rejected because we cannot evaluate the work. Unreadable figures used to be a rare “rookie mistake” that indicated the advisor or senior co-author did not read the paper. But this mistake is becoming more common, even by relatively senior people who should know better. I also suggest learning to use two-column figures. It's a small option in latex figures and tables, and a straightforward clicking and dragging operation in Word. A few weeks ago I spotted a figure in the paper I was working on that I could not read. So I asked my colleague, who said “it doesn't matter, readers don't need to read the figure anyway.” My answer should be obvious in retrospect: “let's take it out.” The reason only three of the papers mentioned above will be rejected because of poor figures is because most of the unreadable figures don't matter. They don't need to be in the paper. This issue is more important for conference papers then for journal papers, because journals can ask for major revisions. And yes it is becoming more common to ask for figures and table to be made legible in revision. If we can't read the figures, we can't fully evaluate the research. I love good research. Good ideas excite me and solutions that work satisfy me. I really hate to see good research rejected because of bad presentation. Sadly, it happens all the time.
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2017 Test Oracle Strategies for Model-Based Testing
abstract
Testers use model-based testing to design abstract tests from models of the system's behavior. Testers instantiate the abstract tests into concrete tests with test input values and test oracles that check the results. Given the same test inputs, more elaborate test oracles have the potential to reveal more failures, but may also be more costly. This research investigates the ability for test oracles to reveal failures. We define ten new test oracle strategies that vary in amount and frequency of program state checked. We empirically compared them with two baseline test oracle strategies. The paper presents several main findings. (1) Test oracles must check more than runtime exceptions because checking exceptions alone is not effective at revealing failures. (2) Test oracles do not need to check the entire output state because checking partial states reveals nearly as many failures as checking entire states. (3) Test oracles do not need to check program states multiple times because checking states less frequently is as effective as checking states more frequently. In general, when state machine diagrams are used to generate tests, checking state invariants is a reasonably effective low cost approach to creating test oracles.
Nan Li 0008, A. Jefferson Offutt
IEEE Trans. Software Eng.2
2016 Analyzing the validity of selective mutation with dominator mutants
abstract
Various forms of selective mutation testing have long been accepted as valid approximations to full mutation testing. This paper presents counterevidence to traditional selective mutation. The recent development of dominator mutants and minimal mutation analysis lets us analyze selective mutation without the noise introduced by the redundancy inherent in traditional mutation. We then exhaustively evaluate all small sets of mutation operators for the Proteum mutation system and determine dominator mutation scores and required work for each of these sets on an empirical test bed. The results show that all possible selective mutation approaches have poor dominator mutation scores on at least some of these programs. This suggests that to achieve high performance with respect to full mutation analysis, selective approaches will have to become more sophisticated, possibly by choosing mutants based on the specifics of the artifact under test, that is, specialized selective mutation.
Bob Kurtz, Paul Ammann, A. Jefferson Offutt, Márcio Eduardo Delamaro, Mariet Kurtz, Nida Gökçe
SIGSOFT FSE3
2016 What to expect of predicates: An empirical analysis of predicates in real world programs
Vinicius H. S. Durelli, A. Jefferson Offutt, Nan Li 0008, Márcio Eduardo Delamaro, Zengshu Shi, Xinge Ai
J. Syst. Softw.2
2016 How to revise a research paper
abstract
This issue contains three outstanding papers, two that contain strong theory and show promise for immediate practical application, and another that can inform a new generation of researchers. The first, A Lightweight Framework for Dynamic GUI Data Verification Based on Scripts, by Mateo, Ruiz and Pérez, presents a way to integrate verification into a GUI during execution. The runtime verifier reads verification rules from files created by the engineers and checks the state of the GUI for violations while running (recommended by Peter Mueller). The second, Model-Based Security Testing: A Taxonomy and Systematic Classification, by Felderer, Zech, Breu, Büchler and Pretschner, surveys and summarizes 119 papers on model-based security testing. This paper should become the first entry port for anybody doing research in the area (recommended by Bogdan Korel). The third paper, Generating Effective Test Cases Based on Satisfiability Modulo Theory Solvers for Service-Oriented Workflow Applications, by Wang, Xing, Yang, Song and Zhang, address the very technically difficult problem of testing service-oriented applications developed with WS-BPEL. Many execution paths in WS-BPEL applications are infeasible. This paper addresses the problem and shows how to generate tests based on finding test paths from embedded constraints (recommended by Bogdan Korel). A well-crafted process to revising a journal submission is crucial for eventual acceptance. Although the initial reaction to the reviews may be negative, it is very important to be proactive and positive. Researchers, even world-renowned, will always be criticized, fairly or not. We must be able to respond to criticism in positive ways. ‘The reviewers were blind and close-minded’, a common complaint, may be valid—however, expressing that does not help achieve the goal of publishing a paper. Authors cannot make reviewers or editors smarter. This is yet another situation where we must strive to change the things we can and accept the things we cannot. In this editorial, I walk through the process that I have used to revise journal papers for two and a half decades. I start my revision process with three initial steps. First, I look at the decision. If it is an ‘accept’, ‘minor revision’ or ‘major revision’, I celebrate. I view a major revision as an ‘accept after lots of work’. I put off reading the reviews until later that day or the next. Even a decision of minor revision may contain things that are bothersome. A reaction of ‘how could the reviewer be so blind?’ is common. Several days later, I return to the reviews for a deep, detailed analysis of what they said. Being proactive is essential. If the reviewers misunderstood, how can the author change the writing so that reviewers will understand the second time? If the reviewers were not satisfied, can the work be better motivated? If the reviewers did not believe the work truly solved the problem, can the problem be restated? Like software, no paper is ever perfect. Like testers, the reviewers' job is to help the authors improve the paper. Recently a co-author and I got reviews asking for a major revision. The revisions asked us to throw the previous empirical study out and start again. (As an editor, I would define that as a reject, but that is another story 1.) The reviews were strange—as if they read the wrong paper. They reflected neither the paper's goals nor its results. Three reviewers completely misunderstood the paper! We finally found a key review comment that helped us realize that we had buried our important goal inside a subsection in the experimental design …in a formula! Our title, abstract, introduction and research questions all sent the reviewers in the wrong direction. That is an extreme case, but true. And it illustrates the main point of response letters. Take responsibility! After all, authors want a paper accepted, but reviewers do not care. They simply want to write a competent review with minimal effort. And they should not care. If they did, they would be conflicted. The first analysis of the reviews should identify all substantive comments. Some reviews have them neatly organized and numbered; others have long paragraphs that make several comments. As an author, your goal is to make sure you read, understand and address every comment, so organization is essential. My process is to print the reviews and start by numbering every individual comment (‘R1-1’, ‘R1-2’,…‘R2-1’,…). Some are related, so for example, if Reviewers 1 and 3 make the same comment, I write ‘R1-4, see R3-7’ and ‘R3-7, see R1-4’. In my first pass, I try not to consider changes. I just want to understand the comments. My next pass is usually performed in a meeting with my co-authors. This is a terrific opportunity to teach students some of the subtle details of high-quality research. We connect each comment to a location in the paper. Some reviewers make that easy by identifying the page and line number explicitly, others do not. We write our own comment number on a printout of the paper, and if it is not there already, we write the page number on the review. Next, my co-authors and I consider responses. With a different colored pen, we make a short note on the paper and the review about the planned change. The note on the review is usually very brief; for example ‘fix’, ‘reword’ or ‘ignore’. All reviews have a few comments where we were just not sure what to do. We cannot ignore them, but we can postpone. We simply write a question mark to help us remember to come back later. Next, we assign jobs to the co-authors. The lead author usually takes responsibility for the simple jobs, and the more complicated rewriting goes to the expert on that particular part of the paper. The most experienced writer usually works on the abstract and motivation. Just as with the initial version, whoever has the best grammar and writing skills should make the final pass. Next, we modify the paper and develop the response letter simultaneously. I will discuss the response letter in detail in a subsequent editorial; for now, just assume that it lists every comment and has a direct response for each. As we change the paper, we check off the review on the paper copies (with a different color) and draft the responses. Sometimes drafting the responses makes it easier to change the paper, sometimes the change makes the response easier to write, and sometimes it doesn't matter which is performed first. After each co-author takes his or her turn, we are usually left with a few troublesome comments. Sometimes the response is simply ‘no’, sometimes the response is ‘good idea, but we can't do that,' and sometimes authors have to go back to the laboratory. Sometimes making the other changes makes the most difficult comments easier to address. It certainly helps when editors made it clear which comments were mandatory. When the editor's note says ‘Mandatory changes are X, Y, and Z’, then not addressing them makes it very likely the paper will not be accepted. If the author chooses not to make a mandatory change, it is imperative that the response convinces the editor that the change should not be made. Ignoring a mandatory change will certainly lead to a rejection. Arguing against a mandatory change is still risky, so be prepared if the editor and reviewers are not convinced. Revising papers is not easy but is essential. Viewing it as collaboration between the authors, the reviewers, and the editor can make this process easier and more effective. In a very real sense, all parties have similar goals: To publish good papers in the journal. Criticism is painful, but it is also essential. My advice is to ignore unfair criticism, pity those who write ignorant criticism, and learn from justified criticism. If we have to avoid criticism, Aristotle had some excellent advice: Criticism is something we can avoid easily by saying nothing, doing nothing, and being nothing. I would like to acknowledge Mary Jean Harrold for help with this editorial. We developed this process together in the first years of our careers. Naturally, this process was influenced heavily by our PhD advisors, Mary Lou Soffa and Richard DeMillo, as well as dozens of collaborators, most notably Paul Ammann.
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2016 How to write an effective "Response to Reviewers" letter
abstract
This issue presents three new inventions in software testing. A major research topic in software engineering is that of fault localization, that is, finding faulty code in failing software. Probabilistic reasoning in diagnosing causes of program failures, by Junjie Xu, Rong Chen, and Zhenjun Du, invents a new graph to model possible faults probabilistically. (Recommended by Atif Memon.) The second paper, Behaviour abstraction adequacy criteria for API call protocol testing, by Hernan Czemerinski, Victor Braberman, and Sebastian Uchitel, invents new criteria to test whether APIs are used correctly. Specifically, the criteria measure the extent to which software that uses the APIs conforms to the expected protocol such as whether the methods are called in valid orders. (Recommended by Hasan Ural.) The third paper in this issue addresses the difficult oracle problem. Predicting metamorphic relations for testing scientific software: A machine learning approach using graph kernels, by Upulee Kanewala, James M. Bieman, and Asa Ben-Hur, invents a technique to use machine learning to predict metamorphic relations, which are used to create oracles in software for which correct behavior is unknown. (Recommended by TY Chen.)
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2016 Editorial: STVR policy on extending conference papers to journal submissions
abstract
Li Peng and Chin-Yu Huang, presents two approaches to improve the effectiveness of fault tolerant analysis.(Recommended by Min Xie.) Exhaustive test sets for algebraic specifications, by
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2016 Editorial: Changes to STVR's Editorial Board
abstract
This issue presents three strong results related to change. The first, A General Modeling and Analysis Framework for Software Fault Detection and Correction Process, by Liu, Li, Wang, and Hu, presents a Markov model-based approach to evaluate software reliability. The approach deals directly with change by emphasizing the influence of data late in development and after deployment. (Recommended by Rob Hierons.) The second paper, Seeding Strategies in Search-Based Unit Test Generation, by Rojas, Fraser, and Arcuri, uses a large collection of open-source programs to demonstrate that changing the initial seed in a search-based test data generation strategy can have a large effect on the quality of the resulting tests. (Recommended by Giuliano Antoniol.) Finally, Prioritizing Test Cases for Early Detection of Refactoring Faults, by Alves, Machado, Massoni, and Kim, presents a strategy for prioritizing regression tests after changing software to support refactoring. (Recommended by Mark Harman.) Like many things, research thrives best with diversity. Diversity of ideas, diversity of problems, diversity of people, and diversity in the leadership. Leadership in academic fields like software testing must be distributed, and the more widely it is distributed, the better off we all are. Perhaps obviously, research also requires change. After all, that is the purpose of research—to change the way things are done, to change the way we think, to change what we know, and to change our beliefs, philosophies, and habits. When the International Conference on Software Testing, Verification, and Validation was formed a decade ago, the charter included strict rules to continuously change the leadership. The steering committee has term limits: three-year terms and a maximum of two consecutive terms. The technical program committee also has term limits: a maximum of three years in a row. The intent is to bring new people in and avoid stagnation. With the same goals, STVR periodically rotates valued editorial board members off and brings new members in. Thus, we are saying au revoir (but not good-bye!) to Sund-Deok Cha, Byoungju Choi, Wolfgang Grieskampf, Robyn Lutz, Jose Maldonado, Atif Memon, and Allen Nikora. They took on the serious responsibility of finding reviewers for papers and culling through their comments to develop thoughtful and informed recommendations on the papers, a job that requires dedication, energy, professionalism, and honor. We are also welcoming Benoit Baudry, Marcio Delamaro, Hyunsook Do, Robert Feldt, Gordon Fraser, Jeff Gray, Natalia Juristo, Moonzoo Kim, Phil McMinn, Mike Papadakis, Mauro Pezze, Lori Pollock, Tao Xie, and Andreas Zeller. Rob Hierons and I, and hopefully all members of the software testing community, are grateful for their commitment.
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2016 Editorial: How to extend a conference paper to a journal paper
abstract
presents a novel semi-supervised learning technique to predict where faults could be in software.(Recommended by Giuliano Antoniol.)Past-Free[ze] reachability analysis: Reaching further with DAG-directed exhaustive state-space analysis, by Ciprian Teodorov, Luka Le Roux, Zoé Drey, and Philippe Dhaussy, presents a new algorithm to perform reachability analysis in model checkers that significantly reduces the state-space explosion problem.(Recommended by Ronald Olsson.)In a previous editorial, I set out STVR's policy on extending conference papers to journal submissions [1].The major requirements are that the journal version must have at least 30% new material, the journal version must include a citation to the conference paper, and the journal version must discuss the conference paper and summarize the new material.Here I suggest some practical guidelines for how to do the extension.The extension should start with four broad changes:
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2016 Editorial: The Downward Death Spiral Review Process
abstract
This issue contains a real-time verification paper.A simplification of a real-time verification problem, by Suman Roy, Janardan Misra, and Indranil Saha, presents a technique to simplify the real-time verification problem.The paper describes a reduction from an infinite-sized state problem to a finite-sized state problem that can be solved with model checking.(Recommended by Alan Hartman).Although the paper took a long time from initial submission to publication, the paper was in revision, not review, for most of that time.The peer review process has been one of the hallmarks of science for decades, if not centuries.It is far from perfect and often frustrating.However, it is essential.I view paper reviewing as having 3 primary purposes:
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2015 A Scalable Big Data Test Framework
abstract
This paper identifies three problems when testing software that uses Hadoop-based big data techniques. First, processing big data takes a long time. Second, big data is transferred and transformed among many services. Do we need to validate the data at every transition point? Third, how should we validate the transferred and transformed data? We are developing a novel big data test framework to address these problems. The test framework generates a small and representative data set from an original large data set using input space partition testing. Using this data set for development and testing would not hinder the continuous integration and delivery when using agile processes. The test framework also accesses and validates data at various transition points when data is transferred and transformed.
Nan Li 0008, Anthony Escalona, Yun Guo, A. Jefferson Offutt
ICST4
2015 Editorial: Plagiarism Is For Losers
abstract
This issue presents two useful applications of formal modeling. The first paper, Specification guidelines to avoid the state space explosion problem, by Groote, Kouters, and Osaiweran, give a technique for security testing of firewalls. They formally model firewalls, and then combine testing with a proof technique to verify the correctness of the implemented firewall. (Recommended by Alan Hartman.) The second paper, Formal firewall conformance testing: An application of test and proof techniques, by Brucker, Brügger, and Wolff, is based on years of experience building state-based models. The authors have found that some models make testing and verification much easier than others, and give guidance for designing useful models. (Recommended by Ronald Olsson.) This editorial explores different forms of plagiarism, discusses why people plagiarize, and offers strategies for avoiding unintentional plagiarism. The amount of plagiarism detected by STVR's editorial board has increased significantly in recent years. This increase might be due to multiple factors. Like most journals, STVR now uses an automated plagiarism detector that searches tens of thousands of papers for similarities. The increase might also be partly due to the general globalization of SWE research 1. Still, another factor is that STVR receives more submissions than in the past. Regardless of why, we now desk-reject almost a dozen submissions to STVR every year. This editorial discusses plagiarism to help potential authors understand what plagiarism is and how to avoid it. I will address authorship rules in more details in a future editorial. Next, I want to discuss why people plagiarize, then suggest ways to avoid plagiarizing. I discuss intentional and unintentional plagiarizing separately. This list, like the last, is certainly not complete, but is intended to be a representative sample of common reasons. Intentional plagiarizing is the more troubling morally, and multiple reasons may come into play. People sometimes intentionally plagiarize out of desperation. They are required to publish and, for whatever reason, do not have the ability to write publishable papers. Others plagiarize because they simply lack ethics. They have very little sense of right and wrong, or perhaps are outright sociopaths. Another possible reason is simply poor judgment—they believe they would not be caught. Although resources on the Internet can make it easier to find material to plagiarize, that same access makes it easier to detect plagiarism. Sadly, some people simply follow their PhD advisor's lead, thinking plagiarizing is normal behavior. Perhaps the saddest reason is when people simply cannot write, so they copy text from those who can. This probably explains a lot of type 4 plagiarism. Additionally, some lose faith in the review and publication system, lose respect for the process, and plagiarize simply to satisfy the publishing requirements. Of course, people also plagiarize unintentionally. Perhaps the most common reason is by not understanding plagiarism. Some PhD students do not understand plagiarism when they start the research work, so it becomes the advisor's responsibility to teach them. My university has recently started a seminar series for PhD students on ethical issues related to research, which includes specific discussions on plagiarism and authorship. Some authors unintentionally plagiarize out of forgetfulness. They read something, forgot they read it, and later thought they invented it. This is why we must keep good notes! Another reason is working with the wrong co-authors; their co-authors plagiarized, and they did not notice. Sometimes people are simply ignorant; they do not know how to properly quote, and thus indicate that something was theirs when in fact they meant to give credit. Others plan poorly, are late for a deadline, and take a shortcut by copying text from another source. Finally, some try to paraphrase someone else's words, and believe that changing a few words in a paragraph puts the paragraph in ‘their own words.’ Regardless of the reason, it is important to point out that journal editors usually cannot know why a submitter plagiarizes. From a journal's point of view, plagiarism is almost always considered to be knowing, willful, and intentional. We have ‘one strike and you’re out' policies. If someone is caught plagiarizing once, we will not allow that person to publish in the journal again. Scientists' most important asset is their reputations, and being caught plagiarizing is often the end of a research career. It is considered a firing offense throughout the world. In US universities, it is one of the most common reasons why tenure is revoked. Plagiarism can also be a time bomb that goes off years or decades later. Recently, several German politicians lost their jobs after plagiarism was discovered in their PhD dissertations 4. In addition to avoiding plagiarism yourself, most scientists agree that we have a moral obligation to report plagiarism when we observe it. I would certainly be very disappointed if someone I respected knew about plagiarism in a paper submitted to STVR and decided not to tell me. Not informing a journal editor of a case of plagiarism is tantamount to condoning plagiarism, and some even consider not reporting plagiarism to be another form of plagiarism. When in doubt, include a reference to where you got your idea. I am pretty sure no paper has ever been rejected from STVR for ‘too many references.’ Thanks to Paul Ammann, Lionel Briand, Mark Harman, and Rob Hierons for providing helpful comments on an early draft of this editorial.
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2015 Editorial: Who Is An Author?
abstract
This issue contains three papers that invent new ideas and evaluated them empirically.The first paper, Directed test suite augmentation: An empirical investigation, by Xu, Kim, Kim, Cohen, and Rothermel, empirically investigates the effectiveness of strategies for augmenting test sets after changes to software.They compared two different test generation algorithms in two separate studies.(Recommended by Paul Ammann.)The second paper, Reducing execution profiles: Techniques and benefits, by Farjo, Assi, and Masri, presents results of analysis of execution profiles.They invented six ways to reduce the size of execution profiles and empirically measured the effect on the quality of analysis after reducing the profiles.(Recommended by T.Y.Chen.)The third paper, Automated metamorphic testing of variability analysis tools, by Segura, Durán, Sánchez, Le Berre, Lonca, and Ruiz-Cortés, invent a technique for solving the oracle problem when testing tools that analyze the variability of software.(Recommended by T.H. Tse.) Combined, these three papers have a whopping 14 co-authors, which leads in perfectly to the subject of this editorial: determining authorship.
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2015 How the web resuscitated evolutionary design
abstract
How the web resuscitated evolutionary designThis issue contains two exciting papers about test automation.The first, Killing strategies for modelbased mutation testing, by Aichernig, Brandl, Jöbstl, Krenn, Schlick, and Tiran, presents techniques and algorithms for automatically generating tests from UML state machines (recommended by Mark Harman).The second, Assessing and generating test sets in terms of behavioural adequacy, by Fraser and Walkinshaw, turns the notion of criteria on its head, by defining test criteria in terms of the outputs instead of the inputs or source (recommended by Hong Zhu).Both inventions can improve test automation, and thus enhance our ability to have evolutionary design.One of my favorite oldies, The Design of Everyday Things, discusses evolutionary design.It caused me to consider what this concept means to software design, development, and testing.I want to start with cost.All technological artifacts, hardware and software, come with costs.Not being an economist or systems engineer, I may leave some out, but at least four types of costs help us understand a major trend in software:
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2014 Establishing Theoretical Minimal Sets of Mutants
abstract
Mutation analysis generates tests that distinguish variations, or mutants, of an artifact from the original. Mutation analysis is widely considered to be a powerful approach to testing, and hence is often used to evaluate other test criteria in terms of mutation score, which is the fraction of mutants that are killed by a test set. But mutation analysis is also known to provide large numbers of redundant mutants, and these mutants can inflate the mutation score. While mutation approaches broadly characterized as reduced mutation try to eliminate redundant mutants, the literature lacks a theoretical result that articulates just how many mutants are needed in any given situation. Hence, there is, at present, no way to characterize the contribution of, for example, a particular approach to reduced mutation with respect to any theoretical minimal set of mutants. This paper's contribution is to provide such a theoretical foundation for mutant set minimization. The central theoretical result of the paper shows how to minimize efficiently mutant sets with respect to a set of test cases. We evaluate our method with a widely-used benchmark.
Paul Ammann, Márcio Eduardo Delamaro, A. Jefferson Offutt
ICST3
2014 Experimental Evaluation of SDL and One-Op Mutation for C
abstract
Mutation analysis modifies a program by applying syntactic rules, called mutation operators, systematically to create many versions of the program (mutants) that differ in small ways. Testers then design tests to cause the mutants to behave differently from the original program. Mutation testing is widely considered to result in very effective tests, however, it is also quite costly. Cost comes from the many mutants that are created, the number of tests that are needed to kill the mutants, and the difficulty of deciding whether mutants behave equivalently to the original program. One-op mutation theorizes that cost can be reduced by using a single, very powerful, mutation operator that leads to tests that are almost as effective as if all operators are used. Previous research proposed the statement deletion operator (SDL) and found promising results. This paper investigates the use of SDL-mutation in a new context, the language C, and poses additional empirical questions, including whether other operators can be used. We carried out a controlled experiment in which cost and effectiveness of each individual C mutation operator were collected for 39 different subject programs. Experimental data are used to define a cost-effectiveness metric to choose the best single operator for one-op mutation.
Márcio Eduardo Delamaro, Lin Deng 0001, Vinicius H. S. Durelli, Nan Li 0008, A. Jefferson Offutt
ICST5
2014 Designing Deletion Mutation Operators
abstract
As a test criterion, mutation analysis is known for yielding very effective tests. It is also known for creating many test requirements, each of which is represented by a "mutant" that must be "killed." In recent years, researchers have found that these test requirements have a lot of duplication, in that many test requirements yield the same tests. Put another way, hundreds of mutants can usually be killed by only a few dozen tests. If we could reduce this duplication without reducing mutation's effectiveness, mutation testing could become more cost-effective. One avenue of this research has been to use only one type of mutant, the statement deletion mutation operator. Researchers have found that statement deletion mutation has relatively few mutants, but yields tests that are almost as effective as using all mutants, with the significant benefit that fewer equivalent mutants are generated. This paper extends this idea by asking a simple question: if deleting statements is a cost-effective way to design tests, will deleting other program elements also be effective? This paper presents results from mutation operators that delete variables, operators, and constants, finding that indeed, this is an efficient and effective approach.
Márcio Eduardo Delamaro, A. Jefferson Offutt, Paul Ammann
ICST2
2014 An Empirical Analysis of Test Oracle Strategies for Model-Based Testing
abstract
Model-based testing is a technique to design abstract tests from models that partially describe the system's behaviour. Abstract tests are transformed into concrete tests, which include test input values, expected outputs, and test oracles. Although test oracles require significant investment and are crucial to the success of the testing, we have few empirical results about how to write them. With the same test inputs, test oracles that check more of the program state have the potential to reveal more failures, but may also cost more to design and create. This research defines six new test oracle strategies that check different parts of the program state different numbers of times. The experiment compared the six test oracle strategies with two baseline test oracle strategies. The null test oracle strategy just checks whether the program crashes and the state invariant test oracle strategy checks the state invariants in the model. The paper presents five main findings. (1) Testers should check more of the program state than just runtime exceptions. (2) Test oracle strategies that check more program states do not always reveal more failures than strategies that check fewer states. (3) Test oracle strategies that check program states multiple times are slightly more effective than strategies that check the same states just once. (4) Edge-pair coverage did not detect more failures than edge coverage with the same test oracle strategy. (5) If state machine diagrams are used to generate tests, checking state invariants is a reasonably effective low cost approach. In summary, the state invariant test oracle strategy is recommended for testers who do not have enough time. Otherwise, testers should check state invariants, outputs, and parameter objects.
Nan Li 0008, A. Jefferson Offutt
ICST2
2014 An Evaluation of the Effectiveness of the Atomic Section Model
Sunitha Thummala, A. Jefferson Offutt
MoDELS2
2014 An industrial study of applying input space partitioning to test financial calculation engines
A. Jefferson Offutt, Chandra Alluri
Empir. Softw. Eng.1
2014 A case study on bypass testing of web applications
A. Jefferson Offutt, Vasileios Papadimitriou, Upsorn Praphamontripong
Empir. Softw. Eng.1
2014 Globalization - references and citations
abstract
This issue has three exciting papers that show how test tools can help in areas ranging from modelbased testing to embedded software toWeb application software. The first, Tool Support for the Test Template Framework, by Cristiá, Albertengo, Frydman, Plüss, and Rodríguez Monetti, describes a new tool to support model-based testing with Z specifications. Their tool, Fastest, is open-source and available online. (Recommended by Paul Strooper.) The second, Model Checking Trampoline OS: A Case Study on Safety Analysis for Automotive Software, by Choi, presents a study of the use of model checking to check safety properties in automotive operating systems. The author was able to find evidence of safety problems in the Trampoline operating system. (Recommended by Jeff Offutt.) The third, Design and Industrial Evaluation of a Tool Supporting Semi-Automated Website Testing, by Mahmud, Cypher, Haber and Lau, presents experience from using a test automation tool, CoTester, in practical situations. They found that the tool was useful for both professional testers and non-professional testers, but most useful for the non-professional testers. (Recommended by Per Runeson.) I wrote about The Globalization of Software Engineering in a previous editorial 1 and followed up with a discussion on language skills to support globalization 1. Another difference I have noticed is in how the scientific community uses citations and references, and more interestingly, which citations they use. Many of these differences are personal and individual, but some seem cultural. I will first discuss citations in general with some thoughts I share with my PhD students, then talk about some cultural differences from my experience. The first principle is that references must help readers understand the paper. Of course, this is so broad that it is not much help, but it is an important starting point. References have an important role in research papers. They need to explain what the paper is based on (context), they indicate what the authors know about the subject and they summarise what the readers should know to understand this paper. Reviewers also use references as a proxy for the measure of care the authors take with their research. Reviewers also expect certain rules to be followed. The most important is ethical: never reference something you have not read. A secondary citation, where we write something like (Parnas [4] ‘as cited by Burdell [15]’), indicates that you read Burdell's paper, and he referenced Parnas' paper. This should only be used when absolutely necessary if the original reference is unavailable. It is also important to list all authors in a reference list; ‘et al.’ is okay in the text, but if you leave off the name in the references of a person who reviews your paper, it will not help your chances of being accepted. Another expectation is that you write the authors' names as they appeared in the published paper. So if someone changes his or her name, you should not update the old papers. The final note is about grammar. The citations are parenthetical elements, not nouns. That is ‘as said in [52]’ is grammatically wrong and should be written as ‘as said by Liskov [52].’ This last one may be the most common mistake, made even by established scientists, probably because it is a convenient shortcut. These ideas are basic, and generally taught in high school and early college writing classes. So our PhD students should already know these rules, and if not, should certainly absorb them before they ‘leave the nest’ and start independent research. A major question about references is how many? The correct answer is, of course, to use exactly the references the paper needs and no more. But defining what ‘references the paper needs’ is clearly subjective. My general philosophy is that it is better to over-reference than under-reference. We are more likely to confuse readers with too much information than with too little. And as a famous scientist once told me, even famous scientists like to see their names in print. We do not get paid much for publishing papers, so referencing papers is a way to thank those who helped get you to the point where you can write this current paper. Generally, I have noticed that North Americans tend to include more references than scientists in other parts of the world, whereas scientists from Asia tend to include fewer. I do not know why. Of greater concern is that some scientists tend to cite many papers from other scientists in their region, but not as many from other regions. So Americans often cite a lot of papers by American authors but not so many from Europe, French scientists sometimes cite lots of French papers but not many from Asia, and so on. In fact, I recently reviewed a paper by French authors where 17 of 22 references were by themselves or other French authors, and omitted key papers by non-French authors. This is probably extreme, but in my experience, fairly common. Generally, the lesson is that we need to make sure we are not being parochial in our study of the literature. Globalization means that quality research is done all over and good papers can come from anywhere.
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2014 Globalization - standards for research quality
abstract
This issue features three interesting papers, all of which offer real solutions to real problems, and demonstrate their success on industrial software. The first, ‘A practical model-based statistical approach for generating functional test cases: application in the automotive industry’, by Awedikian and Yannou, presents a new way to generate tests from models. The results include a tool that selects test inputs (the test generation problem), predicts the expected results (the oracle problem), and suggests when testing can stop (the stopping problem). They have demonstrated their approach on automotive software. (Recommended by Hong Zhu.) The second, ‘A novel approach to software quality risk management’, by Bubevski, offers an advance in managing the risk of software. Bubevski's technique uses Six Sigma and Monte Carlo simulation and has been successfully used on industrial software. (Recommended by Min Xie.) The third, ‘Automatic test case generation from simulink/stateflow models using model checking’, by Mohalik, Gadkari, Yeolekar, Shashidhar, and Ramesh, uses model checking to solve the problem of test data generation based on models. This technique has also successfully been used on industrial automative software. (Recommended by Peter Mueller.) I wrote about The Globalization of Software Engineering in a previous editorial 1 and followed up with a discussion of language skills to support globalization 2 and then uses of references and citations 3. Another difficult difference that is affected by globalization is the expected standards for research quality. Scientists usually learn about research in graduate school. The process has its roots in the middle ages and is based on the ancient apprenticeship model 4. After finishing our classes, we spend years as an ‘apprentice’ to a ‘master’, the PhD advisor. This advisor is responsible for teaching us the dozens of skills, strategies, and tactics required for a successful research career, including standards for the quality of the research. We also learn about research from other professors, by reading papers and reasoning how the research was conducted, but our advisor has primary responsibility. In most cases, our advisors learned from their advisors, they from their advisors, and so on, sometimes back centuries. This is why we are so interested in our genealogies 5. I have friends who can trace their ‘academic roots’ back to luminaries such as Dijkstra, Poisson, Bernoulli, and Euler. The model of research apprenticeships has a rich tradition in countries that have a long history of scientific research. Historically, many of these countries are in Europe and North America. However, part of globalization is that other countries, without a long history of scientific research, are trying to kick-start this process. When knowledge and skills tend to be handed down by word of mouth, this is, not surprisingly, difficult. Thus, the globalization of research results in a large divergence in the standards for research quality. Who teaches new students? Who teaches the teacher? I see the affects of this divergence at STVR. We have a policy of desk-rejecting papers that are out of scope for the journal or that are low quality. We get papers that have no research results, for example, that explain an existing process or concept with an example. We get papers that have insufficient results for a major journal or whose results are not sufficiently original. And we get papers whose writing is so poor that the reviewers would not be able to understand the paper well enough to fairly assess the results. For these papers, we send polite, regretful rejection letters that are as kind as possible. It is clear that many of these authors are bright enough, are hard working enough, and have sufficient technical strengths to carry out high quality research projects. Unfortunately, they simply have not been adequately prepared. This rarely happens with papers from authors in Europe or North America, which have long traditions of research. Unfortunately, most are from the Indian sub-continent or China. These countries (among others) are aggressively trying to improve their economies, education, and research credentials. Professors are encouraged to submit as many papers as possible and are often funded quite generously. In fact, I teach my own students, the majority of whom are from countries without a long research tradition, all of these things. It seems an important goal, then, is for countries without long research traditions to somehow absorb the institutional knowledge of how to perform and disseminate research from countries that do have these long traditions. How? It is clear to me that pressuring young scientists without adequate training to publish does not work. That method is frustrating to reviewers, editors, and conference program chairs and must be frustrating for the scientists themselves. In fact, plagiarism is all too often the result of this kind of pressure. One of my favorite techniques is Brazil's ‘sandwich’ program, in which PhD students are sent abroad in the middle of their studies to work with a research group in their topic. I have had the pleasure of hosting such students, and it is invariably productive and enjoyable. Paid sabbaticals are also effective. Hiring young faculty who received their PhDs from a country with a strong research tradition can also be effective, although sometimes it is difficult to lure the best and the brightest back. A difficult choice that those of us from countries with a long research tradition must make is whether to take a competitive or cooperative stance with this aspect of globalization. That is, do we view other countries who try to improve their research as competitors, or do we cooperate by helping? I do not see research as a zero-sum game. If another scientist publishes a result before I finish the work (as Lionel Briand has done more than once), it is an opportunity to use those results and go further. It is self-evident that we will not run out of problems in software engineering in our lifetimes, so more help is better.
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2014 Globalization-ethics and plagiarism
abstract
Globalization-
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2014 Globalization - logical flow, motivation, and assumptions
abstract
This issue presents three fascinating papers on ensuring the reliability and correct behavior of software. The first, Analysis and testing of black-box component-based systems by inferring partial models, by Shahbaz and Groz, tackles the problem of integration testing of component-based software. When components have no specifications, models, or source, testers can only infer proper behavior by trial and error. This paper uses a model learning approach to derive finite state machines that describe observed behavior of the software component. (Recommended by Rob Hierons.) The second, Sound and mechanised compositional verification of input-output conformance, by Sampaio, Noguira, Mota, and Isobe, uses process algebra to verify conformance of software with the expected behavior. This test theory was applied to test mobile applications. (Recommended by Alexander Pretschner.) The third, Towards the prioritization of system test cases, by Srikanth, Banerjee, Williams, and Osborne, focuses on the problem of test case prioritization. The approach assumes requirements-based tests, assigns a prioritization value to each requirement, and then prioritizes tests that were designed for requirements with a higher priority. (Recommended by Jeff Offutt.) I wrote about The Globalization of Software Engineering in a previous editorial 1, followed up with a discussion of language skills to support globalization 2, uses of references and citations 3, standards for research quality 4, and cheating and plagiarism 5. Another difficult difference that is affected by globalization is presenting research with logical flow, clear motivation, precision, and without cultural-based assumptions. While it is probably obvious that successful research publications must be based on sound research, hard work, and original ideas, it may be less obvious that presentation is just as important. The ability to clearly present research results can be developed through education, practice, and helpful feedback. This editorial attempts to point out a few issues that are influenced by culture, with the hope of helping authors, teachers, and reviewers to understand and improve. The most obvious cultural aspects of presentation, of course, is language-improving language skills improve the ability to present research results clearly. But telling a coherent story is even more important. Educational systems emphasize different topics, and some spend much more time teaching writing skills than others. Our cultural context also influences our writing. We sometimes make assumptions that are standard in our own culture, but may be different in others. This editorial explores some of these issues from a global perspective. Perhaps the most important aspect of presenting research is to have a logical flow of ideas. Each section must logically flow to the next, each paragraph must logically flow to the next, and each sentence must logically flow from the previous. If not, readers will be confused and not understand the research. Logical flow reflects a structured way of thinking that is influenced by culture, mother language, and scientific training. Outlining allows me to see the logical flow at an abstract level, without being distracted by the details of grammar and sentence structure. I outline sections, paragraphs in each section, and the sentences in each paragraph. I look for ‘data flow’ anomalies in the outline. Finally, when a paragraph or section does not look right but I'm not sure why, I ‘reverse engineer’ the text into an outline, refactor the outline to create a better logical flow, and apply the new outline to the text. Another issue that is heavily influenced by culture is motivation. Motivation essentially answers ‘why’-why the problem is relevant, why the solution technique was chosen, and why the specific validation technique used was chosen. Traditionally, egalitarian cultures have a strong built-in mechanism to develop the skills to present motivation. That is how people have their ideas accepted and used. Authoritarian cultures, on the other hand, can afford to de-emphasize motivation and expect people to do what they are told because an authority says so. Although we could spend hours and pages complaining about English, English has at least one strong advantage. Its rich vocabulary allows us to be wonderfully specific and precise in your writing. Instead of, for example, writing ‘We had a high quality test suite’, we can write ‘Our test suite had 95% branch coverage’. This is more specific as well as quicker to read and understand. This issue of specificity has an unusual cultural aspect, because in some cultures, it is rude to be extremely precise. Without judging any particular culture, or implying that vagueness is bad in general, it is important that scientific papers be as specific as possible. The last topic to mention is about assumptions. We all make certain assumptions about our work and communications. Many of these assumptions are unconscious, often based on a cultural context that we are only peripherally aware of. For example, I might say ‘She hit a home run with that result!’, implicitly assuming that the listener understands the metaphor. This is clearly a culturally contextual assumption, because baseball is only played in a few countries. Assumptions that are based in a specific cultural context can be very confusing in research presentations. Our audience is almost always international, and our assumptions are often unconscious. But it is also quite difficult to recognize our own assumptions. Asking for feedback from people from other cultures will help. Exposing ourselves to different cultures through travel or befriending visitors can help avoid such assumptions. Perhaps the strongest of all is to collaborate with someone from another culture. It always helps to have someone recognize our unconscious mistakes, while making mistakes that we can recognize. Of course, awareness of these issues will not guarantee our papers are accepted, read, or understood. However, these are some of the hardest issues to fix in our writing.
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2014 Editorial: how to get your paper rejected from STVR
abstract
This issue contains three deep and compelling papers on modeling software and generating tests.The first, An improved Pareto distribution for modeling the fault data of open source software, by Luan and Huang, presents a new model, based on the traditional Pareto distribution, that accurately describes the distribution of faults in open-source software.(Recommended by Min Xie.)The second, Extending model checkers for hybrid system verification: The case study of SPIN, by Gallardo and Panizo, studies an unusual type of system, hybrid systems.The authors have developed an extension to model checking that allows engineers to accurately model the behavior of this complex type of software.(Recommended by Paul Ammann.)The third, Search-based testing using constraint-based mutation, by Malburg and Fraser, addresses the key problem of test value generation.They propose and evaluate a hybrid form of test value generation that combines search-based techniques with constraint-based techniques.(Recommended by Mark Harman.)This editorial is based on a talk I gave at the ICST PhD symposium in April 2014.I had fun giving the talk and hope you enjoy reading this summary.First, I want to make it clear that I am highly qualified to give advice on getting papers rejected.I have well over 100 rejections in my time and may well be the most rejected software testing researcher of all time.As examples, let me share some quotes from reviewers: 'As usual, Offutt got it wrong.'-TSE1993.(The paper was accepted on the second revision with two accepts and one reject vote, and currently has over 150 citations.)'A study like this should have been published in about 1980.'-TAV1989.'The presentation needs considerable improvement.'-TAV1989.(This was the complete review; no details were given.)'We are sorry to say your paper has been REJECTED.'-Letterfrom editor, and yes, that word was capitalized and in bold face.'Better than average American academic paper, below the standard of papers written by European (non-English) academics.'-FTCS1990.(This comment, by the way, was about the writing as opposed to the research results.)In this editorial, I assume that your goal is to get your paper rejected.My first concrete piece of advice is to be courteous to the reviewers.Reviewing is hard work so you should try to make it easy for the reviewers to reject your paper.Here is how.The most effective strategy is the only one on this list I have not used-plagiarize!This not only gets the current paper rejected, but future papers.Merriam-Webster [1] defines plagiarism as follows:'To use the words or ideas of another person as if they were your own words or ideas.'To make this easier, I have collected a few specific types of plagiarism: Complete copying of an entire paper Copying key results Copying unpublished workCopying auxiliary text such as related work or background Copying figures Improper quoting
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2013 Workshop on revisions to SE 2004
abstract
We shall conduct a half-day workshop on needed revisions to Software Engineering 2004: Curriculum Guidelines for Undergraduate Degree Programs in Software Engineering (SE 2004). A brief overview of the current guidelines and their revision status will be presented. Workshop attendees will share their experience using the current guidelines and suggest needed changes. We will provide a summary report from the workshop to other CSEE&T attendees at a Birds Of a Feather meeting later during the conference.
Mark A. Ardis, David Budgen, Gregory W. Hislop, A. Jefferson Offutt, Mark J. Sebern, Willem Visser
CSEE&T4
2013 Evasive bots masquerading as human beings on the web
abstract
Web bots such as crawlers are widely used to automate various online tasks over the Internet. In addition to the conventional approach of human interactive proofs such as CAPTCHAs, a more recent approach of human observational proofs (HOP) has been developed to automatically distinguish web bots from human users. Its design rationale is that web bots behave intrinsically differently from human beings, allowing them to be detected. This paper escalates the battle against web bots by exploring the limits of current HOP-based bot detection systems. We develop an evasive web bot system based on human behavioral patterns. Then we prototype a general web bot framework and a set of flexible de-classifier plugins, primarily based on application-level event evasion. We further abstract and define a set of benchmarks for measuring our system's evasion performance on contemporary web applications, including social network sites. Our results show that the proposed evasive system can effectively mimic human behaviors and evade detectors by achieving high similarities between human users and evasive bots.
A. Jefferson Offutt, Feng Mao, Aaron Koehl, Haining Wang 0001
DSN2
2013 Town hall discussion of SE 2004 revisions (panel)
abstract
This panel will engage participants in a discussion of recent changes in software engineering practice that should be reflected in curriculum guidelines for undergraduate software engineering programs. Current progress in revising the guidelines will be presented, including suggestions to update coverage of agile methods, security and service-oriented computing.
Mark A. Ardis, David Budgen, Gregory W. Hislop, A. Jefferson Offutt, Mark J. Sebern, Willem Visser
ICSE4
2013 Empirical Evaluation of the Statement Deletion Mutation Operator
abstract
Mutation analysis is widely considered to be an exceptionally effective criterion for designing tests. It is also widely considered to be expensive in terms of the number of test requirements and in the amount of execution needed to create a good test suite. This paper posits that simply deleting statements, implemented with the statement deletion (SDL) mutation operators in Mothra, is enough to get very good tests. A version of the SDL operator for Java was designed and implemented inside the muJava mutation system. The SDL operator was applied to 40 separate Java classes, tests were designed to kill the non-equivalent SDL mutants, and then run against all mutants.
Lin Deng 0001, A. Jefferson Offutt, Nan Li 0008
ICST2
2013 Transformation Rules for Platform Independent Testing: An Empirical Study
abstract
Most Model-Driven Development projects focus on model-level functional testing. However, our recent study found an average of 67% additional logic-based test requirements from the code compared to the design model. The fact that full coverage at the design model level does not guarantee full coverage at the code level indicates that there are semantic behaviors in the model that model-based tests might miss, e.g., conditional behaviors that are not explicitly expressed as predicates and therefore not tested by logic-based coverage criteria. Avionics standards require that the structure of safety critical software is covered according to logic-based coverage criteria, including MCDC for the highest safety level. However, the standards also require that each test must be derived from the requirements. This combination makes designing tests hard, time consuming and expensive to design. This paper defines a new model that uses transformation rules to help testers define tests at the platform independent model level. The transformation rules have been applied to six large avionic applications. The results show that the new model reduced the difference between model and code with respect to the number of additional test requirements from an average of 67% to 0% in most cases and less than 1% for all applications.
Anders Eriksson, Birgitta Lindström, A. Jefferson Offutt
ICST3
2013 Is bytecode instrumentation as good as source code instrumentation: An empirical study with industrial tools (Experience Report)
abstract
Branch coverage (BC) is a widely used test criterion that is supported by many tools. Although textbooks and the research literature agree on a standard definition for BC tools measure BC in different ways. The general strategy is to “instrument” the program by adding statements that count how many times each branch is taken. But the details for how this is done can influence the measurement for whether a set of tests have satisfied BC. For example, the standard definition is based on program source, yet some tools instrument the bytecode to reduce computation cost. A crucial question for the validity of these tools is whether bytecode instrumentation gives results that are the same as, or at least comparable to, source code instrumentation. An answer to this question will help testers decide which tool to use. This research looked at 31 code coverage tools, finding four that support branch coverage. We chose one tool that instruments the bytecode and two that instrument the source. We acquired tests for 105 methods to discover how these three tools measure branch coverage. We then compared coverage on 64 methods, finding that the bytecode instrumentation method reports the same coverage on 49 and lower coverage on 11. We also found that each tool defined branch coverage differently, and what is called branch coverage in the bytecode instrumentation tool actually matches the standard definition for clause coverage.
Nan Li 0008, A. Jefferson Offutt, Lin Deng 0001
ISSRE3
2013 Revision of the SE 2004 curriculum model
abstract
Software Engineering 2004: Curriculum Guidelines for Undergraduate Degree Programs in Software Engineering (SE 2004) [1] is one volume in a set of computing curricula adopted and supported by the ACM and the IEEE Computer Society. In order to keep the software engineering guidelines up to date the two professional societies began a review and revision project in early 2011. This special session will present the results of the review, present a first draft of the revision, and provide time for discussion and input from the computing education community.
Gregory W. Hislop, Mark A. Ardis, David Budgen, Mark J. Sebern, A. Jefferson Offutt, Willem Visser
SIGCSE5
2013 Improving logic-based testing
Gary Kaminski, Paul Ammann, A. Jefferson Offutt
J. Syst. Softw.3
2013 Mutation at the multi-class and system levels
Pedro Reales Mateo, Macario Polo, A. Jefferson Offutt
Sci. Comput. Program.3
2013 What I have learned from usability
abstract
This issue has three fascinating papers. The first, A new method for testing timed systems, by Bonifácio and Moura, presents a discretization method for timed input/output automata that allows grid automata to be constructed in a more compact manner (recommended by Alan Hartmann). The second, It really does matter how you normalize the branch distance in search based software testing, by Arcuri, analyzes different normalizing functions for branch distances in search algorithms and presents a new and improved normalizing function (recommended by Mark Harman). The third, Ranking of software engineering metrics by fuzzy based matrix methodology, by Garg, Sharma, Nagpal, Garg, Garg, Kumar and Sandhya, presents a new framework for ranking software engineering metrics (recommended by Min Xie). I have recently been working on a project about usable security. An interesting thing I have observed is the many similarities between usability and testing. To start with, they both address emergent properties; that is, they address aspects of the software that do not exist until the entire system is completed. Of course, we can (and should!) test software piecewise, but some faults are not visible until the system is entirely integrated, and we certainly cannot measure reliability until the entire system is present. Usability is even more emergent—there is really no sign of usability until we get the user interface of the system in place. Many engineers and educators view usability and testing in the same way: as being unimportant to the real work of building the backend software. Or worse, usability and testing may get in the way of the real work. I think that the most important way they are similar is in their historical progression in relation to software success. In 2000, both testing and usability were relatively unimportant to the success of many software products. In 2013, however, both are now crucial to the success of most software products. In simple terms, without usability, customers will not buy software. Similarly, without the increased reliability that comes with good testing, customers will not buy software. Finally, there is one other way in which usability and testing are the same. Most practicing software engineers studied computer science in college and were taught little or nothing about usability or testing. How many undergraduate computer science programs have a course in usability? How many undergraduate computer science programs have a course in software testing? Does your program?
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2013 The globalization of software engineering
abstract
This issue has three intriguing papers. The first, Using Concepts of Content-based Image Retrieval to Implement Graphical Testing Oracles, by Delamaro, de Lourdes dos Santos Nunes and de Oliveira, presents a new way to automate the oracle function for programmes that produce images. (Recommended by Atif Memon.) The second, A Measurement-based Ageing Analysis of the JVM, by Cotroneo, Orlando, Pietrantuono and Russo, presents a practical exploration of the issue of aging software. Software is modeled as ‘aging’ by considering such things as the long-term depletion of resources from the operating system, incremental corruption of data and accumulation of numerical errors. This paper analyses the JVM's aging characteristics. (Recommended by Michael Lyu.) The third, Regression Verification: Proving the Equivalence of Similar Programs, by Godlin and Strichman, looks at techniques for proving that two programs have equivalent behavior. This would allow for verification of program changes without needing to refer to a formal specification. (Recommended by Wolfgang Grieskamp.) One of the most exciting trends I've been fortunate enough to join is globalization. When I was growing up in a small town in eastern Kentucky, the James Bond movies always captivated me. The most exciting thing about the movies was that they let me join another world, and not just New York and Los Angeles, but Europe and Asia. Bond took me on a tour of plush hotels, casinos, expensive restaurants and fast cars in every corner of the world. When I started my career in the late 1980s, most conferences I attended were in North America, and foreign attendees were rare. But the world was changing, and we soon started travelling overseas to conferences and welcoming foreign visitors into the USA. Most of my PhD students and many of my MS students are from overseas. It is now common and normal to collaborate with scientists from all over the world. Globalization brings enormous benefits, but of course, with certain costs (jet lag is perhaps the most mundane). Robert Laughlin, Nobel Prize in Physics (1998) said it well: ‘The global economy imposes a tax on young people in the form of learning English.’ He went on to say that native-English speakers have a different tax; by not learning a new language, they are slower to absorb broader lessons about the emerging global culture. A couple of years later, one of my colleagues exasperatedly said that dealing with cheating in the classroom is part of acculturating our students. The vast majority of classroom cheating is with foreign students, some of whom seem to have trouble accepting that we seriously believe it is wrong. Attitude towards cheating and plagiarism is clearly cultural. I plan to expand on each of these issues in the next few editorials. My initial purpose is to simply understand this interesting phenomenon. But once we understand, perhaps we can make this process easier. We can recognize what new members of our community need to know, identify which things individuals do not know and then help them learn. Even better, we might be able to identify flaws in our emerging global community and then develop plans to improve. (For example, how did we get stuck with English, and is that really what we want?) More later…
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2013 Editorial: globalization - language and dialects
abstract
This issue has two excellent papers. The first, Testing and verification in service-oriented architecture: A survey, by Bozkurt, Harman and Hassoun, is another detailed survey from the Centre for Research in Evolution, Search and Testing at University College London. The authors do a wonderful job summarizing testing research on service-oriented architecture, covering no less than 262 papers. (Recommended by Atif Memon.) The second, Parallel mutation testing, by Mateo and Usaola, presents new ideas that dramatically increase the speed at which we can perform mutation analysis. The paper presents a new tool, Bacterio, and results from three studies on five algorithms to execute mutants in parallel. (Recommended by Byoungju Choi.) Note that because of previous papers co-authored with the authors of these papers, Co-EiC Rob Hierons was not involved with the handling of the survey paper, and Offutt was not involved with the mutation paper. I wrote about The Globalization of Software Engineering in my last editorial 1. One of the most obvious aspects of this trend, of course, is language. For software engineers to interact on a global scale, we must be able to communicate. And currently, the primary international language is English. Why English? In his very long history of civil wars among English-speaking people, Philips 2 put forth a compelling argument for how English became our 21st century international language. In brief, British/American colonists beat French hunters and trappers in the 1750s (called the French/Indian war in the USA and the Seven Years war in Europe), English speakers went on to settle the North American continent in the 1800s, the Allies won World War I and the USA finished World War II with less damage than the other combatants. International air flight then became the first global business where everybody (primarily pilots and air traffic controllers) had to be able to communicate in real time, without translators. then, English-language television and movies spread throughout the world, and finally, the World Wide Web gave everybody the ability to talk with everybody else … if we could share a language. That's an interesting theory, but what really counts is that at this point in history, we use English for global scientific communication. Although the hundreds of exceptions and odd quirks (articles, nouns and pronouns with gender, ‘their / there / they're’, ‘though / through / tough’) often make us grumble ‘why English?’, my linguistic friends tell me English has certain advantages. It has a huge vocabulary, which allows us to be very precise and exact. It is also structured to discourage ambiguity and is not well suited to metaphors. Although this makes poetry hard, English works well for law and science. What does this have to do with the globalization of software engineering?We are building a global community of scholars and practitioners. Most people who join this community have advanced education (at least college degrees and often PhDs), thus most join the community with at least moderate fluency in English. But if our papers have too many language problems, readers either do not understand or find the papers too painful to read. Additionally, speaking is different from writing, and scientists with limited speaking abilities have trouble interacting at conferences. Of course, the tail also wags the dog. As English has gone global, it has changed. British use many words that are rare or unknown in North America. And it's very common to mix English with other languages; we've all heard of ‘Hinglish’, ‘Spanglish’ and ‘Chinglish’. Linguists also hypothesize the emergence of new dialects of English. The most widely known dialects are probably ‘standard British English’ (such as on BBC) and ‘standard American English’ (such as on CNN). Indian English is another, or possibly several others. I grew up speaking Appalachian English (‘How you'ns doin'?’ ‘We hain't bad.’) 3. Newly emerging dialects may include Northern European English (often English words with Germanic sentence constructs) and Chinese English (fewer articles and verbs omitted when they are clear from context). It's hard to imagine what will happen in the next century or two, but it's likely that English will continue to partition into more dialects. Perhaps, we will have many dialects and a semi-standardized international English 4. Regardless of how English changes, new members to the international software engineering community will still need years to improve their language skills. This process is explicitly supported in many places. Most graduate classes in Sweden are taught in English. Spain has an ‘international PhD’, which requires study visits to universities abroad, and Brazil has a ‘sandwich programme’, which sends PhD students abroad for a year. Some South Korean universities require MS students to score well on the Test Of English as a Foreign Language, which was created to determine whether foreign students were ready to study in universities in the USA. This is an important role for governments. As individuals, we also must support this process. First, reviewers and editors must be patient with young scientists who are still struggling to master English. Of course, the papers must be understandable, but we should look beyond the writing and focus on the ideas. We must also be prepared to offer constructive criticism for how to improve. As advisors, we must encourage our students to improve their English as much as possible before leaving the nest. I tell my students to make friends who are not from their home country, and send them to our English Language Institute. Students at universities in non-English language countries should be encouraged to find ways to improve their language and writing skills, and be given as many opportunities as possible to do so. A good friend of mine went from almost incomprehensible to being more articulate than most native English speakers in 5 years. One of my students went from being halting and nervous in English to starting conversations and asking public questions at conferences in 2 years. Their successes demonstrate that language skills can be significantly improved. Please help and encourage your students and junior colleagues to work on their communication, and demand that they reach competence before graduation.
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2013 A tribute to Mary Jean Harrold
abstract
This issue has three research papers.'Incremental testing of finite state machines', by Chaves Pedrosa and Vieira Moura, addresses the scalability problem of designing tests from finite state machines.They use a divide and conquer approach to define combined finite state machines, which allow individual tests to be defined on smaller units and allow test suites to be built incrementally.(Recommended by Byougju Choi.) 'A survey of code-based change impact analysis techniques', by Li, Sun, Leung, and Zhang, surveys 30 papers that empirically analyzed 23 change impact analysis techniques.The paper synthesizes these results into a structure of four research questions and proposes several new research questions.(Recommended by Jane Hayes.)'Combining weak and strong mutation for a noninterpretive Java mutation system', by Kim, Ma, and Kwon, looks at the cost of executing mutants in mutation systems.They propose a new way to combine 'strong' and 'weak' mutation that keeps most of the strength of strong mutation, while achieving much of the cost savings from weak mutation.They adapted the muJava mutation tool to demonstrate their ideas.(Recommended by Rob Hierons.)Note that because of previous papers co-authored with the authors of this paper, Offutt was not involved with its handling.Software engineering lost one of its best last month.And I lost a good friend and role model.I first met Mary Jean Harrold at the 1988 TAV symposium (now the International Symposium on Software Testing and Analysis) and was impressed in every possible way.We finished our PhDs in the same year: I in August at the Georgia Institute of Technology and her in December at the University of Pittsburgh.I joined Clemson University in August 1988 and was excited when she applied for the following year.She accepted our offer, and we spent the next 3 years learning how to be professors together.Her years of teaching high school math helped her start as a terrific teacher, and she continued to improve every year.She was also an ideal mentor.Even as a new assistant professor, she somehow knew exactly how to motivate her students, had the insight to understand what knowledge and skills they lacked, and had the patience and abilities to teach what they needed.She set exacting standards with a kind and respectful demeanor.Most importantly, she earned their loyalty and love.Her students worked harder than anybody else because they wanted to impress her and they knew she was working even harder.I have tried to emulate her advising style for more than 20 years.We co-authored four papers, working with three students.The memories of working on those papers are still bright because we taught each other much about research, problem formulation, writing, and how to respond to reviews.Our PhD advisors had very different styles, and I was able to absorb much of what her advisor, Dr. Mary Lou Soffa, taught her, and I think she absorbed some of what my advisor, Dr. Rich DeMillo, tried to teach me.Those three pre-tenure years at Clemson were incredibly formative and bonding.We also grew up very close geographically.Mary Jean was born and raised in Huntington, West Virginia, and I was raised about 50 miles west, near Morehead, Kentucky.Even 20 years later, she teased me because I thought she came from a big city.I responded by reminding her that her home state was even poorer than mine.Appalachians are few and far between in academia, and we always felt that bond.A sharing of an unusual culture that few understand.I firmly believe that Mary Jean Harrold was the best PhD advisor in all of software engineering.Her deft touch shows; I know immediately when I see one of her students give a talk.She was also a wonderful colleague and great scientist.She focused on some of the deepest and most complicated problems in software analysis, testing, and evolution.She did not just focus on research that works on small problems or in the lab but found solutions that were scalable and usable by real engineers.
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2012 Toward Harnessing High-Level Language Virtual Machines for Further Speeding Up Weak Mutation Testing
abstract
High-level language virtual machines (HLL VMs) are now widely used to implement high-level programming languages. To a certain extent, their widespread adoption is due to the software engineering benefits provided by these managed execution environments, for example, garbage collection (GC) and cross-platform portability. Although HLL VMs are widely used, most research has concentrated on high-end optimizations such as dynamic compilation and advanced GC techniques. Few efforts have focused on introducing features that automate or facilitate certain software engineering activities, including software testing. This paper suggests that HLL VMs provide a reasonable basis for building an integrated software testing environment. As a proof-of-concept, we have augmented a Java virtual machine (JVM) to support weak mutation analysis. Our mutation-aware HLL VM capitalizes on the relationship between a program execution and the underlying managed execution environment, thereby speeding up the execution of the program under test and its associated mutants. To provide some evidence of the performance of our implementation, we conducted an experiment to compare the efficiency of our VM-based implementation with a strong mutation testing tool (muJava). Experimental results show that the VM-based implementation achieves speedups of as much as 89% in some cases.
Vinicius H. S. Durelli, A. Jefferson Offutt, Márcio Eduardo Delamaro
ICST2
2012 Better Algorithms to Minimize the Cost of Test Paths
abstract
Model-based testing creates tests from abstract models of the software. These models are often described as graphs, and test requirements are defined as sub paths in the graphs. As a step toward creating concrete tests, complete (test) paths that include the sub paths through the graph are generated. Each test path is then transformed into a test. If we can generate fewer and shorter test paths, the cost of testing can be reduced. The minimum cost test paths problem is finding the test paths that satisfy all test requirements with the minimum cost. This paper presents new algorithms to solve the problem, and then presents data from an empirical comparison. The algorithms adapt approximation algorithms for the shortest super string problem. The comparison is with an existing tool that uses a brute force approach to extend each sub path to a complete path. One new algorithm is based on the greedy set-covering algorithm and the other is based on finding a matching over a prefix graph. The comparison was performed on open software and showed that both new solutions generate fewer test paths than the brute force approach. The prefix-graph based solution takes much less time than the other two solutions when the number of test requirements is large.
Nan Li 0008, A. Jefferson Offutt
ICST3
2012 Adding Criteria-Based Tests to Test Driven Development
abstract
Test driven development (TDD) is the practice of writing unit tests before writing the source. TDD practitioners typically start with example-based unit tests to verify an understanding of the software's intended functionality and to drive software design decisions. Hence, the typical role of test cases in TDD leans more towards specifying and documenting expected behavior, and less towards detecting faults. Conversely, traditional criteria-based test coverage ignores functionality in favor of tests that thoroughly exercise the software. This paper examines whether it is possible to combine both approaches. Specifically, can additional criteria based tests improve the quality of TDD test suites without disrupting the TDD development process? This paper presents the results of an observational study that generated additional criteria-based tests as part of a TDD exercise. The criterion was mutation analysis and the additional tests were designed to kill mutants not killed by the TDD tests. The additional unit tests found several software faults and other deficiencies in the software. Subsequent interviews with the programmers indicated that they welcomed the additional tests, and that the additional tests did not inhibit their productivity.
William Shelton, Nan Li 0008, Paul Ammann, A. Jefferson Offutt
ICST4
2012 The h-index beats the impact factor
abstract
This issue has two papers with results on real industrial software projects. The first, On the testing of user-configurable software systems using firewalls, by Robinson and White, presents the ‘just-in-time’ testing strategy for user-configurable software and includes results on commercial software. The second, A case study in model-based testing of specifications and implementations, by Miller and Strooper, presents a case study of testing software specifications. Before discussing the h-index, I have the pleasure to make an announcement about the journal. In 2012, STVR will have eight issues rather than the four issues per year we have had for the last 21 years. This will allow us to clear the backlog of papers and support the increasing amount of research in the software testing field. My last editorial 1 discussed the reasons why scientists publish papers and emphasized publishing to influence either the research field or the industry. I also wrote an editorial 2 on the problems with the way journals are often evaluated by publishers and universities, the ‘journal impact factor’ 3. My opinion of this deeply flawed criterion has not changed, but I recently learned about another measure that looks more promising. The index h, defined as the number of papers with citation number higher or equal to h. The h-index and the journal impact factor have an essential difference. The journal impact factor has a 2-year ‘window’, that is, it only counts papers published in the last 2 years. The h-index has no window. It counts all papers published by an individual over a lifetime. The h-index also has some other interesting characteristics. It omits papers that are ignored by other scientists, thus encouraging scientists to publish papers on topics that matter and in places that are read. The h-index also rewards longevity and productivity in numbers of papers, but only the papers that other scientists read and cite. This means that the h-index cannot directly compare scientists who have been working for different lengths of time. A derivative measure might be the h-index divided by the years since the first publication, although I have not seen that used or proposed. Our promotion committee was told that a general rule of thumb is that successful scientists should expect to have an h-index approximately the number of years they have been working, excellent scientists should have an h-index of about 1.5 times, and h-indexes of 2 times the years working are very rare. Of course, for a measure to be successful, we need to calculate it. Luckily, the web makes this easy. And not surprisingly, free calculator tools are available on the web, most notably Google Scholar. The h-index is designed for individuals, not journals, so it cannot directly replace the journal impact factor. But it could certainly be adapted. Appropriate modifications would have to be made to account for the age of the journal and the number of papers published per year. So for me, this is the first measure of research productivity that I can support. One thing is missing, though. The fourth reason to publish from my last editorial was to influence practice … the h-index does not measure this. Can we find a measure that does?
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2012 Non-expert reviews considered helpful
abstract
This issue has three innovative papers. In a belated implementation of a decision made at a meeting of STVR's editorial board, this issue will start naming the reviewing editor who took charge of each paper. So when you read ‘recommended by’ later, that identifies the reviewing editor who found reviewers, evaluated the paper and the reviews and made a recommendation to the co-editors in chief. This is immensely more work than it may seem on the surface, and we are happy to acknowledge the reviewing editors’ hard work. The first paper in this issue, A testing strategy for abstract classes, by Clarke, Power, Babich and King, reports on a landmark advance in object-oriented testing. The authors have invented a way to test abstract classes without having to instantiate the class, thereby giving users more confidence when they create inheritance hierarchies. This is recommended by Atif Memon. The second, Test data regeneration: generating new test data from existing test data, by Yoo and Harman, presents an innovative approach to the always challenging problem of automatic test data generation. Their idea is called regeneration, where existing tests are used as a basis for new tests. This is recommended by Paul Strooper. The third, Fuzzy Bayesian system reliability assessment based on prior two-parameter exponential distribution under different loss functions, by Gholizadeh, Shirazi and Gildeh, presents a new approach to reliability assessment that uses fuzzy parameters, fuzzy random variables and fuzzy prior distributions. This is recommended by Tor Stalhane. I want to talk more about reviewing in this editorial. Most reviews are from experts on the topic of the paper. When trying to assess the value of a research paper, reviewing editors quite naturally tend to look at the top experts on the topic and ask two or three of them to review the paper. However, for a research paper to have significant impact, the paper must be accessible beyond the few experts. Thus, it is also valuable to get opinions from non-experts, scientists who are more representative of the hoped for intended audience. Getting one of three reviews from a non-expert can help assess the broad accessibility of a research paper. Reviewing as a non-expert takes different techniques from reviewing as an expert. Generally speaking, a non-expert should be able to read and understand at least the broad points in a research paper. For example, ideally, everybody with some knowledge of testing should be able to understand all introductions and conclusions in testing papers published in STVR, and most should be able to understand empirical sections. If we cannot follow details of algorithms or techniques, that is okay. So as a non-expert, a reviewer should comment on how well he or she understood the paper, whether there were any obvious flaws and whether it was easy to separate the parts that could be understood from the parts that could not. Non-experts may also be particularly suited to assess how hard it will be for the research ideas to move into practice. A non-expert probably cannot assess the originality and significance, but other reviewers should be able to fulfil that role. So if you are asked to review a paper that is a bit out of your area, you should alert the reviewing editor but also be prepared to write a non-expert review. For a reviewing editor, look at the reviews from non-experts differently and make sure that at least some reviewers have the knowledge to understand all the technical details.
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2012 Status and Awards
abstract
by Zhou, Zhang, Hagenbuchner, Tse, Kuo, and Chen, tests search engines.A major problem that is addressed is the 'oracle function', which for these applications means how can we know whether the search results are 'correct' (recommended by Byoungju Choi).The second, Automated Verification and Testing of User-interactive Undo Features in Database Applications, by Ngo and Tan, addresses the problem of testing the ability for users to undo operations.A correspondence between program statements that raise erroneous effects and program statements that can undo those effects is reported and is used to develop a verification technique (recommended by Shaoying Liu).The third, Testing Aspect-oriented Programs with Finite State Machines, by Xu, El-Ariss, Xu, and Wang, reports that aspect-oriented programs have new kinds of faults, which they call aspect faults.This observation is used to develop a test strategy to detect these kinds of faults (recommended by Sudipto Ghosh).The journal held its third editorial board meeting at the Fifth International Conference on Software Testing, Verification, and Reliability (ICST) in Montreal.It was a good meeting with our publishing editor from Wiley and about a dozen members of the board.This meeting is important to resolve issues, discuss strategic directions for the journal, and plan for future directions.Our most important decision was to institute two yearly awards for the journal.The first will be a Best Paper of the Year award.
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2012 Flipping the testing classroom
abstract
Flipping the testing classroomThis issue has three terrific papers.The first paper, 'A testing-based process for component substitutability', by Flores and Polo, shows us how to test reusable components.The research uses back-to-back testing to evaluate the behaviour of the internal functions of a component, helping the tester decide when a component is stable enough to be reused and when new tests are needed.(Recommended by Sudipto Ghosh.)The second, 'A framework for automatic generation of security controllers', by Martinelli and Matteucci, addresses the problem of guaranteeing the security of complex systems.This is done through a formal process of modelling components of the systems, then aggregating those models into a security model of the entire complex system.(Recommended by Jeff Offutt.)The third, 'A formal framework to test soft and hard deadlines in timed systems', by Merayo, Nunez and Rodriguez, presents a method to test timely properties of software.The tests are designed from real-time specifications of the software.(Recommended by Antoniol Guiliano.)I usually write about research issues here, but I am going to diverge to talk about a recent educational experiment I tried.I first heard about 'the flipped classroom' [1] in a presentation by my university's Center for Teaching Excellence.Instead of listening to the professor talk for an hour (or 2.5 h in our one-day-a-week schedule), then going home to do homework, in a 'flipped' classroom, the student listens to recorded lectures at home, then works problems in the classroom.Knewton's website [2] gives a full description.The flipped model has several advantages: (i) students do not need to be in a crowd to listen to a lecture, but working problems in a group can be very helpful; (ii) it is hard to listen for a complete hour (or 2.5 h), so flipping lets students pause the recordings at any time; (iii) flipping lets students go at different speeds, which is a huge relief to gifted students and a benefit for struggling students; (iv) finally, the in-class sessions let professors focus on what each student needs individually, rather than treat all students the same.Although dubious, I decided to try flipping my classroom for 2 weeks in my graduate software testing class [3].These weeks cover Chapter 3 from the green book [4], Logic-based testing, which in the past has been difficult for some students.This chapter also lends itself well to the kind of problem solving that can be done in a classroom.I started by recording lectures for sections 1, 2, and 3. Educational experts suggest making recorded lectures 10-15 min, so I broke my lectures into distinct pieces.I used Camtasia, which records voice over powerpoint.I told the students that they were required to view the lectures before coming to class.My class prep for the in-person meeting was surprisingly light.I prepared and posted the assignment, and wrote a few problems down on a piece of paper.We worked some problems together, then I told the students to start their homework and call out if they needed help.The results were surprisingly positive.Two students submitted the assignment during class, and three others said they had solved the problems but wanted to recopy their answers.The in-class questions were interesting.Much of my time was spent explaining subtle points of the material to students who missed it in the reading and lecture.Other students needed help with manipulating logic expressions.It is easy to say 'they should have known that before', but this format gave us the opportunity to fill in a hole in their knowledge so they could succeed in this part of the class.My second findings are from an opinion poll.I asked the students three questions on our class discussion board: 'Did you view the lectures before class?' 'Did you feel the class was useful?' and 'Do you think you did better on the homework because of the class session?'A total of 100% of the respondents said yes to all three questions.The follow-up comments were extremely positive.
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2011 Using abstraction and Web applications to teach criteria-based test design
abstract
The need for better software continues to rise, as do expectations. This, in turn, puts more emphasis on finding problems before software is released. Industry is responding by testing more, but many test engineers in industry lack a practical, yet theoretically sound, understanding of testing. Software engineering educators must respond by teaching students to test better. An essential testing skill is designing tests, and an efficient way to design high quality tests is to use an engineering approach: test criteria. To achieve the maximum benefit, criteria should be used during unit (developer) testing, as well as integration and system testing. This paper presents an in-depth teaching experience report on how we successfully teach criteria-based test design using abstraction and publicly accessible web applications. Our teaching materials are freely available online or upon request.
A. Jefferson Offutt, Nan Li 0008, Paul Ammann, Wuzhi Xu
CSEE&T1
2011 Teaching software testing: Experiences, lessons learned and the path forward
abstract
According to a study commissioned by the National Institute of Standards and Technology in 2002, software bugs cost the U.S. economy an estimated $59.5 billion annually, or about 0.6 percent of the nation's gross domestic product (GDP). The same study also found that more than one-third of these costs, or an estimated $22.2 billion, could be eliminated by an improved testing infrastructure. These numbers would be significantly higher if the study were conducted today.
W. Eric Wong, Antonia Bertolino, Vidroha Debroy, Aditya P. Mathur, A. Jefferson Offutt, Mladen A. Vouk
CSEE&T5
2011 A logic mutation approach to selective mutation for programs and queries
Garrett Kent Kaminski, Upsorn Praphamontripong, Paul Ammann, A. Jefferson Offutt
Inf. Softw. Technol.4
2011 A mutation carol: Past, present and future
A. Jefferson Offutt
Inf. Softw. Technol.1
2011 Status of the journal
abstract
This issue has three papers that introduce creative solutions to hard testing problems. The first, On the use of a similarity function for test case selection in the context of model-based testing, by Cartaxo, Machado, and Neto, uses a similarity function to reduce model-based tests that are redundant in terms of their coverage. The second, Fault-driven stress testing of distributed real-time software based on UML models, by Garousi, presents a new way to apply stress testing to evaluate real-time constraints in distributed software. The third, On the selection of software defect estimation techniques, by Cangussu, Haider, Cooper, and Baron, introduces a new methodology to analyze techniques for estimating defects in software. I would like to use this editorial to give a brief status report of the journal, starting with a major announcement. Effective immediately, Rob Hierons of Brunel University will join me as a second Editor-in-Chief. Rob has helped run STVR for years and his willingness to take on additional responsibilities is crucial to our continuing success. I have struggled to keep up with the growth of the journal and Editor-in-Chief is now a two-person job. The STVR editorial board met in March at the International Conference on Software Testing, Verification and Validation. The journal is doing very well. The ISI impact factor for 2009 was 1.6, making STVR one of the highest rated journals in software engineering. Our first-submission response time is now down to around four and a half months, which is good, although not quite at our goal of three months. The EiC occasionally has not been responsive enough. Bringing Rob onboard is a major part of the solution, allowing the load to be shared. I am also changing some of my commitments to free more time for the journal. We are setting up several new e-mail notices that will come automatically from Manuscript Central (MC). These will help the editorial board avoid, identify, and rectify delays sooner. Reviewing editors will get copies of review invitation reminder e-mails and late review reminders. Review editors will also get ‘give up’ messages when MC stops sending invitation reminders to reviewers, which will be signals to invite another reviewer. Reviewing editors will also get a message when reviews are two weeks overdue. Finally, reviewers will get ‘friendly reminder’ e-mails one week before their reviews are due. We also plan to rotate some new scientists onto the editorial board, replacing some who will be rotating off. Announcements will appear here in future issues. Finally, STVR has a growing backlog of accepted papers. While this is symptomatic of our success, it is also a problem for authors waiting for papers to appear in print. The Early View system gives us something to reference, but it is not as good as a paper appearing in print. We hope the publisher will soon increase our yearly page budget, which is, of course, partly an economic decision. I want to thank all members of the editorial board and our reviewers for their hard volunteer work. And of course, thanks to the many authors who continue to submit great papers to STVR.
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2011 Editorial: What is the purpose of publishing?
abstract
What is the
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2011 ICST 2009 Special Issue
abstract
This special issue contains extended versions of four papers from the second IEEE International Conference on Software Testing Verification and Validation (ICST 2009). These four papers were selected based on the reviews from members of the program committee and subsequently subjected to additional rounds of review and revision. We issued 10 invitations to this special issue. Three authors declined to submit, two were not selected after submission, one chose not to revise the paper after being reviewed, and four papers were eventually accepted. The first paper is Reducing Logic Test Set Size While Preserving Fault Detection, by Kaminski and Ammann. This paper introduces a new logic criterion, Minimal-MUMCUT, which has less overlap in terms of faults detected than previous logic criteria, and thus requires fewer tests. It is also provably stronger than the widely used MCDC. The second paper is Improving Penetration Testing through Static and Dynamic Analysis, by Halfond, Choudhary, and Orso. Penetration testing seeks vulnerabilities in software by simulating attacks. Halfond et al. use new software analysis techniques to improve penetration testing. The conference version won the best paper award at ICST 2009. The third paper is An Approach for Testing Pointcut Descriptiors in AspectJ, by Romain Delamare, Baudry, Ghosh, Gupta, and Le Traon. Aspect-oriented programming uses ‘crosscutting concerns’ to increase modularity in software, by encapsulating pieces of software that appear in separate units into their own objects. Delamare et al. have invented testing techniques to find faults in the descriptions of the aspects. The fourth paper is JDAMA: Java Database Application Mutation Analyzer, by Zhou and Frankl. Modern software uses databases much more than in the past, opening up a place for software faults to appear. Database queries are often long, complicated logical expressions, with lots of potential for faults. When the faults are subtle, most queries will return the correct information, and when they do fail, the tables that are returned often look correct. Zhou and Frankl extend a previous mutation-based approach to use analysis and instrumentation of bytecode. All four papers in the ICST 2009 special issue are partially drawn from each of the first authors' PhD Dissertations. The ICST organizers take this as a strong measure of success of the conference. Six years ago, if a student asked ‘where is the best place to publish a paper on testing software?’ the answer was far from clear. The field had several special purpose testing conferences, and many conferences included testing as one topic, but not the main topic. We also had conferences that were intentionally kept very small, even as field was growing. The conference that published the most testing papers was the International Symposium on Software Reliability Engineering, even though its main topic was not testing. It was clear that we had a real need; a need for a large-scale, high quality, broad conference that included all facets of testing, verification, and validation. Thus, a small group (initially Anneliese Andrews, Lionel Briand, and Jeff Offutt) started developing ideas for a new conference in the summer of 2005. This group grew quickly to 30 or 40 people who shared this vision of a new major conference in software testing. A steering committee was elected from that group in 2007 (Andrews, Baudry, Briand, Harman, Hierons, Le Traon, Mathur, Offutt, Williams) and a preliminary charter was approved in 2007, followed by a proposal to the IEEE Computer Society in 2007. The first ICST conference was held in Lillehammer, Norway in 2008, the second in Denver, USA in 2009, the third in Paris, France in 2010, and the fourth in Berlin, Germany in 2011. Because of the obvious overlap, STVR committed to offering a special issue to ICST on a yearly basis. The special issue for ICST 2008 appeared earlier this year, and the special issues for ICST 2010 and 2011 should appear next year. On a final note, we express our gratitude to the many hard working people who helped organize ICST 2009 and this special issue. The steering committee members provided valuable advice and we were assisted by industry chairs Wolfgang Grieskamp, Christian Zapf, and Robert Binder, and general chair Anneliese Andrews. The success of ICST 2009 and the high quality papers in this special issue is due to the efforts of these individuals as well as the Technical Program Committee for ICST 2009 and the additional reviewers for this special issue. We also appreciate the hard work from the ICST 2009 local arrangements chair, Susanne Sherba, the web chair, Orest Pilskalns, and our publicity chairs, Robert Feldt, Jens Krink, Shaoying Liu, Jose Maldonado, and Tao Xie. And last but certainly not least, we thank all the authors who spent valuable time in preparing papers for ICST 2009 and this special issue. 7 July 2011
A. Jefferson Offutt, Per Runeson
Softw. Test. Verification Reliab.1
2010 Scalability issues with using FSMWeb to test web applications
Anneliese Amschler Andrews, A. Jefferson Offutt, Curtis E. Dyreson, Christopher J. Mallery, Kshamta Jerath, Roger T. Alexander
Inf. Softw. Technol.2
2010 Modeling presentation layers of web applications for testing
A. Jefferson Offutt
Softw. Syst. Model.1
2010 Testing coupling relationships in object-oriented programs
abstract
Abstract As we move toward developing object‐oriented (OO) programs, the complexity traditionally found in functions and procedures is moving to the connections among components. Different faults occur when components are integrated to form higher‐level structures that aggregate the behavior and state. Consequently, we need to place more effort on testing the connections among components. Although OO technologies provide abstraction mechanisms for building components that can then be integrated to form applications, it also adds new compositional relations that can contain faults. This paper describes techniques for analyzing and testing the polymorphic relationships that occur in OO software. The techniques adapt traditional data flow coverage criteria to consider definitions and uses among state variables of classes, particularly in the presence of inheritance, dynamic binding, and polymorphic overriding of state variables and methods. The application of these techniques can result in an increased ability to find faults and to create an overall higher quality software. Copyright © 2010 John Wiley & Sons, Ltd.
Roger T. Alexander, A. Jefferson Offutt, Andreas Stefik
Softw. Test. Verification Reliab.2
2010 Recognizing authors: an examination of the consistent programmer hypothesis
abstract
Abstract Software developers have individual styles of programming. This paper empirically examines the validity of theconsistent programmer hypothesis: that a facet or set of facets exist that can be used to recognize the author of a given program based on programming style. The paper further postulates that the programming style means that different test strategies work better for some programmers (or programming styles) than for others. For example, all‐edges adequate tests may detect faults for programs written by Programmer A better than for those written by Programmer B. This has several useful applications: to help detect plagiarism/copyright violation of source code, to help improve the practical application of software testing, and to help pursue specific rogue programmers of malicious code and source code viruses. This paper investigates this concept by experimentally examining whether particular facets of the program can be used to identify programmers and whether testing strategies can be reasonably associated with specific programmers. Copyright © 2009 John Wiley & Sons, Ltd.
Jane Huffman Hayes, A. Jefferson Offutt
Softw. Test. Verification Reliab.2
2010 Editorial: OOP and discrete math
abstract
This issue has three terrific papers. The first, Verification of real-time systems design, by M. Emilia Cambronero, Valentín Valero, and Gregorio Díaz, presents a way to verify aspects of software during design. The second, A systematic representation of path constraints for implicit path enumeration technique, by Tai Hyo Kim, Ho Jung Bang, and Sung Deok Cha, proposes to encode path constraints into ‘flow facts’ to improve the accuracy of implicit path enumeration. The third, Fault localization using a model checker, by Andreas Griesmayer, Stefan Staber, and Roderick Bloem, gives a method for using model checking information to not only identify software failures, but to help find where the corresponding faults may be located. I learned a new word last year. A retronym is a descriptive word that was not needed before a new word came into being. When we started writing object-oriented programs, we needed a new word for what we used to do (procedural programming). Procedural programming had the advantage of following directly from algebra; a C function is essentially an example of applied algebra. We all took algebra in high school, so learning procedural programming in college was straightforward. But OO programming uses concepts that do not appear in algebra! My daughter took Java programming in high school last summer. She was introduced to classes, inheritance and polymorphism as follows: ‘These are new concepts unlike anything you've seen before. Thus they are a bit hard to swallow at first.’ I gave her all kinds of examples from real life (watches and wall clocks inherit from general clocks, rocking chairs and desk chairs inherit from chairs, which inherit from furniture, etc.). But it struck me that we have no mathematical context for this. Yet programming is always very mathematical! I didn't learn OO programming until the end of graduate school in the 1980s. In fact, I didn't see information hiding until graduate school because I had undergraduate majors in math and business programming. As a math student, I took some abstract, way out, ‘useless,’ classes such as logic and set theory, Boolean logic, linear algebra, abstract algebra, and number theory. At the time, I thought it helped me have fun with the Rubiks cube, and little did I know these classes were preparing me to be an OO programmer … When I first learned about data structures and information hiding, they were not ‘new concepts.’ They were simply non-rigorous abstract algebras: sets of values with functions. Stacks and queues are not always complete or consistent, but they are useful! What's the point? The point is that teaching procedural programming is easy and natural because students have a solid grounding in elementary algebra. They are only applying concepts they already know. But teaching them OO is hard, because they have not studied logic and set theory, let alone Boolean and abstract algebras! First year CS students today usually take two semesters of OO programming and two semesters of calculus. In their third semester, they take data structures and multi-variable calculus or differential equations. Finally, in their fourth semester (or later) they squeeze in a discrete math class. But by then it is too late to help them with their elementary OO programming!! It's hard to know how much three semesters of calculus will help a CS student, but I believe an early class or two in discrete math would help them enormously. In fact, they probably need to take discrete math much earlier, before high school programming. They can certainly learn discrete math before the traditional pre-calculus, trigonometry, algebra 2, or even geometry. Perhaps this is why high school students think our field is ‘boring,’ ‘hard,’ ‘useless,’ and ‘only for nerds.’ Perhaps this is why programming seems so hard. Perhaps if we taught discrete math before OO programming, students would start studying computing disciplines again—and we wouldn't need so much testing. 25 January 2010
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2010 Editorial: People are approximators, not perfectors
abstract
People are
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2010 Editorial: Agility must be good for testing
abstract
This issue has three papers that advance our knowledge in criteria-based test design. The first, Improving the coverage criteria of UML state machines using data flow analysis, by Lionel Briand, Yvan Labiche, and Q. Lin, proposes coverage criteria to design tests from UML state machines. The second, Automatic string test data generation for detecting domain errors, by Ruilian Zhao, Michael R. Lyu, and Yinghua Min, presents a new method for automatically generating strings as test values. The third, Full predicate coverage for testing SQL database queries, by Javier Tuya, Maria José Suárez-Cabal, and Claudio de la Riva, adapts test criteria in a novel way, to testing database queries. Unless you've been hiding or never look outside of academia, you must have seen the wave of agile processes that have been sweeping the software industry. Is the term agile process just one more of the hundreds of buzzwords that have made the rounds, only to disappear in obscurity, or does it describe something that will have a lasting positive impact on software development? If you hope for an answer to that question here, I am no fool to make such a rash prediction. I am neither an advocate nor a critic of agile processes, and not knowledgeable enough to be either. My mind is truly open on this subject. Through my students and other industrial contacts, I have talked with numerous people who applied an agile process correctly and thought it helped very much. I have also talked with people who failed miserably with an agile process. Many of them seem to have applied the process incorrectly. Without a long tutorial, agile processes involve some speedups to older processes and also some changes that are intended to make major quality improvements. Most companies I have seen fail left out two things: they did not refactor and they did not document. Both are crucial for quality and the companies leave them out to ‘save time’. But leaving those out means after three or four maintenance cycles, the software is in danger of turning into a huge unmaintainable mess. But these issues are well known and discussed by people who are far more knowledgeable about process than I am. My interest is more personal. Whether agile processes are good for software development or, I like them! I like them for the simple reason that they put more emphasis on testing. Perhaps my favorite new term is Developer Testing. Teachers and researchers have been advocating more ‘unit testing’ for decades, but agile processes emphasize ‘developer testing’, where developers are responsible for testing their own classes before integrating them into the rest of the system. Get it? Developer testing, not unit testing. Okay, I don't see the difference either … but the new term works! In the last five years hundreds of software companies have been pushing their programmers to do a better job of developer testing. And the most widely used testing tool is probably JUnit, a simple unit test automation framework that looks a lot like semester projects I assigned in the early 1990s. When practitioners heard about my 1988 PhD dissertation on automatic (unit) test data generation, they often asked how it worked on ‘real’ programs, but now industry needs all our unit testing knowledge. Another major emphasis of many agile processes is Test-Driven Development (TDD). TDD uses tests as requirements, and programmers go through a cyclical development process where each cycle adds enough software to satisfy one more test. Between developer testing and TDD, software testing assumes a muchmore prominent place in software development. In particular, testing is moved from primarily being a downstream activity (system testing after implementation) to an upstream activity, integrated with requirements and design. Whether this is good for software development is still an open question, but it must be good for software testing! A common problem among companies in my area is that they switch to agile processes, then realize they do not have enough testing expertise. If the programmers know little about testing, how can they do well at unit testing? Very few universities teach testing to undergraduates, so most professionals have little testing knowledge. If the testers know the domain but not much about testing or software, the TDD tests will be incomplete. This is an opportunity for educators and researchers. Another issue is that some companies will develop tests for TDD, then use those same tests as system tests. All the TDD books say not to do that, but the apparent potential for savings is simply too attractive for some companies to resist. Unfortunately, TDD tests will not be complete or to do things like test for subtle interactions among features. So what does this mean for researchers and educators? For many years, I felt like I was trying to teach the proverbial farmer a better way to plow only to hear him say ‘I already don't plow as good as I know how’. Now, industry doesn't know how to test as well as they need, which means they need our ideas and inventions more than ever!
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2010 Editorial: Ethics and Publishing
abstract
This issue has two papers with unusual author lists: Roger T. Alexander, Jeff Offutt and Andreas Stefik for the first; and Jane Huffman Hayes and Jeff Offutt for the second. Both papers were submitted before I became EiC and both were handled by Associate Editor Rob Hierons. Effective and sufficient controls were in place to ensure there was no conflict in their handling. These papers were submitted before I agreed to become Editor-in-Chief of STVR (EiC). Moreover, they are primarily the work of the first authors (Hayes and Alexander); I was just their helper. Even after becoming EiC, the privacy and authentication controls that our online submission system, ManuscriptCentral, blinded me from all aspects of the reviewing process . . . just as with any other author. When I became EiC in 2007, I made a commitment to not submit more papers while serving as EiC. While it's a sacrifice to not publish in the main journal in my research area (and somewhat unfair to my students), I strongly feel that is the ethical thing to do. Even if we could define a process that was completely fair and unbiased, the mere appearance of a conflict would be bad for the journal. As an aside, this was a personal decision, not required by the publisher. Editors of other journals have made a different choice, and I do not criticize those decisions. The broader issue of ethics in publishing concerns me a lot, as I'm sure it concerns every journal editor. And yes, the journal has several problems every year. The most obvious problem, of course, is plagiarism. A recent example is a submission that was more than 80% copied from a paper published in the 1970s. The problem was identified by a reviewer; we immediately rejected the paper, informed the authors, the authors' Department Chair and the authors' Dean. These authors were also placed on a list of individuals who are not welcome to submit future papers to STVR. Another recent case concerned a paper that had substantial overlap with a previous research proposal. Ironically, the author of the proposal (which was not funded) was invited to review the paper. The evidence of word-for-word equivalences was overwhelming. These are clear cut cases of plagiarism, and everybody in the world would recognize this behavior as unethical. A more subtle issue that occasionally comes up is that of authorship. Who is listed as authors on the paper, and in what order, is ordinarily a matter for the authors to decide, and editors seldom question their decisions. But the general question of who belongs as author of a paper and how to order those authors is a question that many of us wrestle with. A few years ago I was involved with a strange case on another journal, where the third author on a paper raised a complaint after a revision was submitted. The third author (who I will call Gamma) was a former student of the first (who I will call Alpha), and claimed that Alpha, in effect, stole the paper. Gamma did all the research and wrote the paper alone, and then Alpha heard about the paper and insisted that Gamma add Alpha to the author list as payback for advising Gamma's PhD studies. The first submission's author list was ‘Gamma and Alpha’, but Alpha insisted on submitting the revision, taking the opportunity to change the author order and add an additional author before Gamma's name! Gamma was able to provide early drafts of the paper, and numerous records documenting the research, whereas Alpha clearly did not understand the research. Eventually the paper was not published, but the big loser, of course, was Alpha, whose reputation was irreparably damaged. Another ethics problem is when someone publishes the same results twice. Most journals welcome the process of expanding a conference paper to a journal paper with substantially more results. However, this area has a lot of gray . . . the common ‘30% rule’ is hard to measure. And is it okay to republish in English a paper that was published in a less widely used language? An STVR reviewer recently accused an author of plagiarizing from his own technical report (there is nothing wrong with that, as I clearly expressed to all involved). I was once accused of plagiarizing a result from a journal paper, when in fact, that result had appeared in my prior technical report and in a rejected conference paper. Worse, an author of the paper where the result appeared was on the conference committee that rejected my previous paper. Another gray area is when an author copies and pastes large parts of the background or literature review . . . that is, the key results are not plagiarized but parts of the paper are. This is most common with authors whose English is really bad. Should this situation be treated the same as the above problems, or do we extend some sympathy to those who struggle with writing in this crazy language with which we are all stuck? Editors must constantly be on watch, as do reviewers and readers. Science can only flourish when we have openness and honesty, and although we will never reach that ideal, we continue to try. It is our responsibility to teach each new generation of scientists the expectations and rules of our community. Finally, I hope you read these two papers. Testing coupling relationships in object-oriented programs, by Alexander and others, offers a technically deep and complicated analysis of problems that can arise when using inheritance and polymorphism, and adapts data flow criteria to design tests to find these problems. If nothing else, the possible coupling sequences in Figure 5 should scare any programmer who uses polymorphism. Recognizing authors: An examination of the consistent programmer hypothesis, by Hayes with some help, presents an analysis technique that can help identify which programmer wrote a particular piece of software. This has potential applications in identifying the authors of viruses and other malware.
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2009 Using Coupling-Based Weights for the Class Integration and Test Order Problem
abstract
During component-based and object-oriented software development, software classes exhibit relationships that complicate integration, including method calls, inheritance and aggregation. Classes are integrated and tested in specific orders, where each class is added and tested one by one to see if it integrates successfully. A difficulty arises when cyclic dependencies exist—the functionality that is used by the first class to be tested must be mimicked by creating ‘stubs’ (sometimes called ‘mock objects’), an expensive and error-prone operation. This problem is generally called the class integration and test order (CITO) problem, and solutions must fully be automated for integration and testing to proceed smoothly and efficiently. This paper describes new techniques and algorithms to solve the CITO problem. New results include improved edge weights to more precisely model the cost of stubbing, and the use of node weights, which allows more information to be used. These weights are derived from quantitative measures of couplings between the integrated and the stubbed classes. Also, a new algorithm for computing the integration and test orders is presented. The technique is compared with an existing approach and found to be cheaper, get the same results when using edge weights exclusively, and yield better results when using node weights.
Aynur Abdurazik, A. Jefferson Offutt
Comput. J.2
2009 Test Sequence Generation For Integration Testing Of Component Software
abstract
Ensuring high object interoperability is a goal of integration testing for object-oriented (OO) software. When messages are sent, objects that receive them should respond as intended. Ensuring this is especially difficult when software uses components that are developed by different vendors, in different languages, and the implementation sources are not all available. A finite state machines model of inter-operating OO classes was presented in a previous paper. The previous paper presented details of the method and empirical results from an automatic tool. This paper presents additional details about the tool itself, including how test sequences are generated, how several difficult problems were solved and the introduction of new capabilities to help automate the transformation of test specifications into executable test cases. Although the test method is not 100% automated, it represents a fresh approach to automated testing. It follows accepted theoretical procedures while operating directly on OO software specifications. This yields a data flow graph and executable test cases that adequately cover the graph according to classical graph coverage criteria. The tool supports specification-based testing and helps to bridge the gap between theory and practice.
Leonard Gallagher, A. Jefferson Offutt
Comput. J.2
2009 TAIC PART 2007 and Mutation 2007 special issue editorial
Mark Harman, Zheng Li 0002, Phil McMinn, A. Jefferson Offutt, John A. Clark
J. Syst. Softw.4
2009 Editorial: Is software testing essential or accidental?
abstract
Is software testing essential or accidental?This issue contains three papers.Automatic instantiation of abstract tests on specific configurations for large critical control systems, by Flammini, Mazzocca, and Orazzo, proposes the interesting idea of creating abstract tests from system requirements and automatically instantiating them into concrete tests.Transition covering tests for systems with queues, by Huo and Petrenko, proposes another technique to automatically generate tests, this time for concurrent systems.Generating input data structures for automated program testing, by Chung and Bieman, studies another aspect of automatic test data generation, describing how statements are connected in terms of constraints that can be solved to yield test inputs.The common theme in these three papers is that they are trying to find automated ways to solve the hardest essential problem in testing software-generating test input values.When I was a college senior, a manager from IBM visited one of my classes.I will never forget his message.He said that computing was in its infancy and we should expect great changes during our careers.He also said that programmers spent 90% of their time on activities that were not directly related to the problem the software was trying to solve.The most time-consuming activity was debugging, followed by things such as keeping decks of punched cards in order, struggling with odd syntax and even odder semantics in poorly designed languages, wrestling with computers that had too little memory, convincing poorly designed operating systems to run our programs the way we wanted, and understanding what customers wanted.In graduate school, I read Brooks' famous paper 'No Silver Bullet: Essence and Accidents of Software Engineering' [1].His paper echoed the talk from the IBM manager and connected those issues with philosophy courses I took in college.If you haven't read it, Brooks philosophized that software developers solve accidental problems, which are 'difficulties that today attend its production but are not inherent,' and essential problems, which are 'difficulties inherent in the nature of software.'The IBM talk and Brooks' paper sparked my interest in software engineering, which I still view in large part as reducing the time software developers spend on accidental problems.I tell my students that we have made great progress; we are probably near the 50/50% level, but still far from a mature level, which I estimate to be about 10/90%.As my interest in software testing grew, I sometimes wondered whether testing is solving an accidental or essential problem.Testing is often treated as accidental in industry, where it is sometimes barely done at all, often done poorly, and seldom done well.Testing is often cut when schedules and budgets overrun, which sounds like testing is inessential.On the other hand, processes such
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2009 Editorial: Running a conference program meeting in the 21st century
abstract
The high cost of travel and the continuing growth of the global nature of research mean that we will continue to have online TPC meetings.Although opportunities for mistakes are plentiful, if conducted well, an online meeting can be a rich and productive process.Most importantly, it allows decisions to be made on the basis of technical merit.
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2009 Editorial: confidentiality and plagiarism
abstract
This issue has three interesting papers. The first, Integrating testing with reliability, by Schneidewind, explores issues at the intersection of testing and reliability, with the goal of discovering which test techniques have the strongest impact on reliability. The second paper, Inclusion, subsumption, JJ-paths, and structured path testing: a redress, by Yates and Malevris, explores the JJ-path (LCSAJ) test coverage criterion, correcting errors in previously published theoretical statements about JJ-paths inclusion relationships with other test coverage criteria. The third paper, Testing with model checkers: a survey, by Fraser, Wotawa, and Ammann, provides a comprehensive survey of how model checkers have been applied to software testing problems. A few years ago I attended a lecture by Professor Robert B. Laughlin, physics Nobel Laureate. He said that ‘globalization imposes a tax on young people—they have to learn English.’ We all know that English is the primary international language, but viewing the process of learning English as a tax to join the global culture is thought provoking. Interestingly, I noticed an ‘inverse tax’ on native English speakers on a recent trip to Shanghai. Buying food at a restaurant where the workers speak English costs quite a bit more. Along with the ‘language tax,’ globalization requires adopting a world view that is compatible with the rest of the global culture. This world view has many different aspects. As scientists and educators, we often see this in terms of plagiarism. A recent submission to STVR had entire sections copied verbatim from previous papers. During investigation, we also found that the same paper had been simultaneously submitted to another journal. Naturally, the paper was rejected, the authors' supervisor notified, and the authors are prohibited from submitting to STVR again. Oddly enough, the authors did not seem to understand they'd done anything wrong! Although I'm not a philosopher, this mismatch of world views seems to be at the heart of morals and ethics. We could define level 1 thinking on plagiarism to be ‘I think plagiarism is okay’ and level 2 to be ‘I will be careful when I plagiarize because I might get in trouble.’ Level 3 could be ‘I think plagiarism is wrong, but will do it when the perceived benefits outweigh the tangible risks of getting caught and the intangible risks of upsetting my own conscience.’ Then level 4 could be ‘I think plagiarism is wrong and will not do it.’ Level 1 thinking is clearly not the common world view of the global society, but levels 3 and 4 are. Plagiarism is more important to us than most groups. Science and research cannot thrive without openness and honesty. By accident or intent, our reputation is essential to our success as scientists, and a serious tangible risk of plagiarism is having our reputations destroyed forever. This would be a career-ending event for most scientists. The journal also recently had a breach of confidentiality—a paper under review was shared with some students, one of whom contacted the author directly. It was an innocent mistake, no damage was done, and appropriate apologies were made, but the incident serves as a reminder of how important confidentiality is. Papers submitted to a journal should only be available to editors and reviewers, and all ideas and results in the papers must be protected until the paper is published. Even with a paper that will eventually be published, the authors could be embarrassed by mistakes that are eliminated from later versions. Confidentiality also requires that reviewer's names be withheld from authors and from each other. This is necessary to ensure that reviewers feel safe to provide honest open opinions. Avoiding plagiarism and protecting confidentiality preserves STVR's reputation and its brand name. Happily, both seem to be in good shape. Our publisher tells me that STVR's Journal Citation Reports' impact factor is above 1.0 for the second year in a row, cementing its place as a top software engineering journal.
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2009 Editorial: Testing my new building
abstract
This issue has two exciting papers. The first, Modelling methods for web application verification and testing: state of the art, by Alalfi, Cordy and Dean, provides a comprehensive survey of models that are used to help us evaluate the quality of web applications. The second, System testing for object-oriented systems with test case prioritization, by Kundu, Sarma, Samanta and Mall, presents a method to generate tests from UML sequence diagrams. I recently moved into a new building, and not surprisingly, it made me think about testing. (Right, everything does.) Although I am happy here, it has its flaws like any building. Lots of people have drawn analogies between software engineering and civil engineering, and I took this opportunity to reflect on how the analogy applies to testing. One of the first things I noticed is that my blinds can only be full open or full closed, not part way. Missing functionality! The first day I noticed a space at the bottom of a stairwell with no access. Dead code! Our small seminar room has a built-in projector and an electric screen. Very nice! But the ceiling lights are either all on or all off, and there is no way to turn off just the lights in front of the screen. Usability flaw! Missing requirement? Connecting the computer to the projector requires running a wire across the carpet. Another usability flaw! The acoustics in the large seminar room downstairs is so bad that speakers must use a microphone. Design flaw, usability problem! Several bathroom stall doors do not latch properly because they were hung incorrectly. A classic integration fault! The elevator lights indicate that the wrong elevator is arriving, but only on one floor. Unit fault! (Since this is the CS floor, the students claim a software fault, but my guess is crossed wiring.) Several rooms abut a large glass wall. Because the carpenters did not know how to connect dry wall to glass, they left 6 inch gaps between the office walls and the glass wall. Architectural mistake! When we moved into our offices, we found that some rooms had thermostats in the middle of the wall where we planned to put our whiteboards, and all of us had electrical outlets behind our furniture. Integration faults! Another fault that turns out to be a feature is that the elevators are very slow. The designers claim this was intentional to encourage the use of stairs (it's a ‘green building,’ after all). Just like in software, ‘it's a feature, not a flaw!’ My favorite is our back door. It has a security control and is meant to be locked at all times. Unfortunately, if we simply release the door, it doesn't latch. Hence, the door is seldom closed! The door is fine, the frame is fine, but the door was hung incorrectly. This is an integration fault that leads to a classic emergent property failure: a security flaw. This analysis is fun, but it also tells us that we have the same problems in building construction as in software construction. They've been putting up buildings for thousands of years and they still don't get it right! It also reminds us that there is no such thing as correctness in any but the tiniest of engineering constructs. We would never consider asking whether a building is ‘correct.’ Or a car, or an airplane. Thus why do we ask if computer programs, which are orders of magnitude more complicated than a building, are correct? As George Box said, ‘All models are wrong, some are useful.’ A more interesting question is who is the tester? Who ‘tests’ buildings? When I had a deck added to the back of my house, the county sent inspectors to check the design, the postholes, the concrete bases, the framing, the final deck and the electrical wiring. The inspectors are testers! And how does somebody get to be an inspector for construction projects? By being very knowledgeable; often one of the best engineers, carpenters, electricians, or plumbers! Software engineering will not be considered a mature field until we match all traditional engineering disciplines that operate in the physical domain. That is, until the best programmers and designers hope to be promoted to testing. 11 November 2009
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2008 Programmers Ain't Mathematicians, and Neither Are Testers
A. Jefferson Offutt
ICFEM1
2008 Testability of Dynamic Real-Time Systems: An Empirical Study of Constrained Execution Environment Implications
abstract
Real-time systems must respond to events in a timely fashion; in hard real-time systems the penalty for a missed deadline is high. It is therefore necessary to design hard real-time systems so that the timing behavior of the tasks can be predicted. Static real-time systems have prior knowledge of the worst-case arrival patterns and resource usage. Therefore, a schedule can be calculated off-line and tasks can be guaranteed to have sufficient resources to complete (resource adequacy). Dynamic real-time systems, on the other hand, do not have such prior knowledge, and therefore must react to events when they occur. They also must adapt to changes in the urgencies of various tasks, and fairly allocate resources among the tasks. A disadvantage of static real-time systems is that a requirement on resource adequacy makes them expensive and often impractical. Dynamic realtime systems, on the other hand, have the disadvantage of being less predictable and therefore difficult to test. Hence, in dynamic systems, timeliness is hard to guarantee and reliability is often low. Using a constrained execution environment, we attempt to increase the testability of such systems. An initial step is to identify factors that affect testability. We present empirical results on how various factors in the execution environment impacts testability of real-time systems. The results show that some of the factors, previously identified as possibly impacting testability, do not have an impact, while others do.
Birgitta Lindström, A. Jefferson Offutt, Sten F. Andler
ICST2
2008 An Industrial Case Study of Bypass Testing on Web Applications
abstract
Web applications are interactive programs that are deployed on the world wide Web. Their execution is usually controlled very heavily by user choices and user data. This makes them vulnerable to abnormal behavior from invalid inputs as well as security attacks. Thus, Web applications invest heavily in validating user inputs according to defined constraints on the values. This work focuses on validation done on the client, which uses two types of technologies; restrictions in HTML form fields and scripts that check values. Unfortunately users have the ability to subvert or skip client-side validation. Bypass testing has been developed to test the behavior of Web applications when client-side validation is skipped. This paper presents results from an industry case study of bypass testing applied to a project from Avaya Research Labs, NPP. The paper presents a process for designing, implementing, automating and developing bypass tests. The theory of bypass testing had to be adapted to the unique characteristics of NPP software, which represented a significant engineering challenge. The 184 tests that were generated resulted in 63 unique failures, providing significant experience and numerous lessons learned. The case study also revealed several difficult problems that need to be addressed in future research.
A. Jefferson Offutt, Qingxiang Wang, Joann J. Ordille
ICST1
2008 Quantitatively measuring object-oriented couplings
A. Jefferson Offutt, Aynur Abdurazik, Stephen R. Schach
Softw. Qual. J.1
2008 Editorial: The journal impact factor
abstract
This issue contains three divergent papers.The first, Modular formal verification of specifications of concurrent systems, by Gradara, Santone, Vaglini, and Villani, proposes a bottom-up approach to verifying modular systems.Properties of components are first verified, then emergent properties of the system as a whole are verified.The approach is applied to a web service.The second paper, Simulated time for host-based testing with TTCN-3, by Blom, Deiss, Kontio, Rennoch, and Sidorova, describes a method to test real-time embedded software.When real-time software is tested in a development environment where the timing characteristics do not match the target environment, this research proposes using simulated time to test the real-time properties.The third paper, IPOG/IPOG-D: Efficient test generation for multi-way combinatorial testing, by Lei, Kacker, Kuhn, Okun, and Lawrence, presents two new strategies for t-way combinatorial testing.The strategies have been implemented in a tool, FireEye, which is available on the first author's website.Software engineering researchers have always developed tools, but most tools have not easily been available to other researchers.This positive trend has the ability to multiply the impact of our research, which brings me to the main subject of this editorial.Many readers of Software Testing, Verification and Reliability may not be familiar with the 'impact factor,' but it is the major way that journals currently are evaluated.It is used almost exclusively worldwide by publishers, universities, research organizations, and government agencies to judge the impact of scientific journals.The factor is used to make hiring decisions, determine promotion and tenure, allocate research funding, and determine whether PhD students should be allowed to graduate.I first heard of this measurement when a visitor told me that she could not graduate without a journal paper, and the list of acceptable journals was determined solely by this specific measure.Thomson Scientific's Journal Citation Reports is the recognized authority for evaluating journals [1].It publishes statistics that are intended to objectively evaluate scientific journals and how they influence the global community of researchers.The statistics they compute are broad and multifaceted.Unfortunately, the research community seems to have narrowed down to one specific measure, the impact factor, to measure journals.The primary advantage of the impact factor is that it is very easy to measure.Unfortunately, it does not capture the true quality of research papers, their effectiveness, or the long-term impact of journals.Journal Citation Reports defines the impact factor as the frequency with which the 'average article' in a journal has been cited in a particular year [2].It is calculated on a year-by-year basis.For year Y , the number of citations to papers published in the journal during years Y -1 and Y -2 is
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2008 Editorial: Science Fiction and Fantasy
abstract
This issue contains two papers.The first, IPOG/IPOG-D: Efficient test generation for multi-way combinatorial testing, by Lei, Kacker, Kuhn, Okun and Lawrence, presents two new strategies for t-way combinatorial testing.The strategies have been implemented in a tool, FireEye, which is available on the first author's website.Software engineering researchers have always developed tools, but most tools have not been easily available to other researchers.This positive trend has the ability to multiply the impact of our research.The second, Reconciling perspectives of software logic testing, by Kaminski, Williams and Ammann, describes a rather large family of test criteria based on covering logical expressions.Logical expressions are essential to software at all levelsrequirements, specifications, design, and implementation.These logic test criteria attempt to cover logical expressions in various ways.The paper presents existing logic expressions in a uniform way, relates them by various criteria, and presents a new test criterion.Although both papers contain substantial theory, they both provide new knowledge on testing techniques that are used in industry.Thus, they are the kinds of papers that STVR is looking for . . .I grew up with a writer of science fiction and fantasy.From my perspective, science fiction books postulated new advances in science or engineering and explored their effects on human society, whereas fantasy books postulated a universe where the rules are slightly different.I enjoyed fantasy but my real love was SF, and I voraciously read all the books in the house.This early exposure to SF fueled my interest in being a scientist.I initially wanted a career in physics, but like thousands of students before me, Electricity and Magnetism convinced me that my future would be much brighter with an advanced degree in computing.I went to grad school with the intent of building programs that could think: self-aware with highorder cognitive abilities.Like hundreds before me, my first AI class disabused me of this notion.Disenchanted, I turned to the myriad of accidental problems in software development, which soon led me into my primary research area of testing.Along the way I learned more about the difference between SF and science.Both groups attempt to summon the future with their own ideas.One difference is that SF often assumes sudden, large-scale changes in science or technology.We love to imagine the lone scientist who uses lightning and dead body parts to summon life, the tinkerer who builds a spaceship out of bicycle parts and broken lawnmowers, the bookworm who builds a time machine in her study, or the engineer who creates a faster than light engine on a mountaintop.Unfortunately, scientists have to contend with the reality of taking very small, incremental steps that over decades lead
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2008 Editorial: Software testing is an elephant
abstract
This issue contains two papers. The first, An analysis technique to increase testability of object-oriented components, by Kansomkeat and Rivepiboon, examines the problem of testability of object-oriented software components when the source is not available but something like bytecode is. OO components have low testability because information hiding obscures the state that needs to be controlled and monitored during testing. This research uses a clever idea of extracting a control and data flow graph, which is then used to increase both controllability and observability, making it easier to detect faults. The second paper, The determination of optimal software release times at different confidence levels with consideration of learning effects, by Ho, Fang and Huang, uses stochastic differential equations to build a software reliability model. This model is validated on data that were published in six previous papers. The results will help project managers decide when to release software to maximize its reliability. We all have probably heard the parable about the blind men and the elephant. Each blind man touches a different part of the elephant and describes what he feels. The men get into an argument, which leads to physical violence, and they resolve the conflict with outside help. In some versions the conflict is never resolved. Last spring I attended a workshop on software system testing and was lucky enough to find a wise person who helped dispel some of my blindness by understanding several types of test activities. Most of my research has focused on using test criteria to help design software tests. Test design is amenable to objective, quantitative assessments of the tests, as well as automatic generation of test values, my first research love. Criteria-based test design requires knowledge of mathematics and of programming—it is an engineering approach. An equally important way to generate tests is from human intuition. Objective test criteria can overlook special situations that smart people with knowledge of the domain, testing and user interfaces will easily see. Both approaches are intellectually stimulating, rewarding and challenging, but they appeal to people with different backgrounds. These two approaches are also complementary and, in most projects, help equally to create high-quality software. For efficient and effective testing, especially as software evolves through hundreds or thousands of versions, we also need to automate our tests. Test automation involves programming that is usually relatively straightforward, involving scripting languages, frameworks such as JUnit or capture/replay tools. Test automation requires little knowledge of theory, of algorithms, or about the domain, but test scripts must often solve tricky problems with controllability. Many companies focus heavily on test execution. If we combine test design with test execution and do not use automation, test execution is very hard. This approach is also inefficient and usually ineffective. This is like expecting programmers to take software requirements and immediately start to code without designing or outlining the algorithms. Especially when not fully automated, proper test execution requires meticulous care and attention to detail. A challenge in testing is deciding whether the output was correct or not. Test evaluation is more difficult than most researchers think and often requires deep domain knowledge and the ability to solve problems of observability. Test evaluation is logical and similar to scientific experimentation, so a background in fields such as law, psychology, philosophy and math can be very helpful. These four different engineering activities—test design, automation, execution and evaluation—are essential to testing. But they do not cover all aspects of testing software. Any process requires management to set policy, organize a team and communicate with external groups. Tests also must be well documented and put into revision control. This requires input from all four prime activities as we need to know why each test was designed and how it relates to other software products. The documentation must be included as a part of automation, leading us to maintenance of our tests. Tests must be reused as software evolves, allowing the cost of creating a test to be amortized over hundreds or thousands of versions of the software. A related, and very difficult problem is trimming the test suite. As its size grows, the suite can become redundant and too large for efficient execution. People in our field sometimes see our own piece of this elephant as the only important piece. A couple of years ago I sat on an NSF panel reviewing software testing proposals. As the assessments progressed, it became apparent that one participant was on a mission to block all criteria-based research. A direct quote was ‘Only fools say the human should not be involved in testing.’ Oddly, nobody ever said humans should not be involved! At least, nobody on the panel said that; neither did any of the proposals. That is, the panelist was raising a false conflict. wastes resources, blinds us to differences, creates unnecessary, unpleasant and wasteful conflicts. Software engineering grew out of this attitude with respect to development years ago. Software development teams have personnel who specialize in requirements, architecture, design, integration and coding, among others. It is high time we applied similar specialization in testing. This will allow our testing activities to be more effective and efficient, as well as helping managers to organize, train, and retain testers. The first step is the hardest. We must open our eyes, see the entire elephant in its full glory, and quit thinking that our piece is the only piece. If we do, we can welcome research on all aspects of testing (such as the first paper in this issue, which is on test automation), teach better classes and expand the research community's ability to help industry with real problems.
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2007 Generating Trace-Sets for Model-based Testing
abstract
Model-checkers are powerful tools that can find individual traces through models to satisfy desired properties. These traces provide solutions to a number of problems. Instead of individual traces, software testing needs sets of traces that satisfy coverage criteria. Finding a trace set in a large model is difficult because model checkers generate single traces and use a lot of memory. Space and time requirements of modelchecking algorithms grow exponentially with respect to the number of variables and parallel automata of the model being analyzed. We present a method that generates a set of traces by iteratively invoking a model checker. The method mitigates the memory consumption problem by dynamically building partitions along the traces. This method was applied to a testability case study, and it generated the complete trace set, while ordinary model-checking could only generate 26%.
Birgitta Lindström, Paul Pettersson, A. Jefferson Offutt
ISSRE3
2007 Service Oriented Architecture Empirical Study
Mohammad Abu-Matar, A. Jefferson Offutt
SEKE2
2007 Editorial: Introduction and plans for the future
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2007 Editorial: Standards for reviewing papers
abstract
Many years ago when I was a student, my advisor walked into
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2007 Editorial: Why should I review papers?
abstract
This issue has three fascinating papers. The first paper, Software component composition: A sub-domain-based testing foundation, by Hamlet, proposes a theory of composing software components that is based on testing. Instead of modeling or specifying software systems directly, this paper suggests the novel approach of deriving abstractions of components from testing and then using these abstractions to model the behavior of the entire system. The second paper, A combinatorial testing strategy for concurrent programs, by Lei, Carver, Kacker and Kung, addresses the hard problem of testing concurrent software. Instead of trying to test all possible synchronization sequences, this paper presents a method and algorithm for choosing a subset of synchronization sequences that will lead to effective fault detection. The third paper, Studying the separability relation between finite state machines, by Spitsyna, El-Fakih and Yevtushenko, explores the relationships between finite state machines. It quantifies the ‘distance’ between nondeterministic FSMs. Owing to the difficulty in soliciting reviews, one of these papers has actually been in STVR's reviewing process for more than two years, which brings me to the main subject of this editorial . . . . One of the most important jobs of editors and associate editors is soliciting reviews. Reviewing takes time that we don't get paid for. The effort counts for almost nothing in promotion and tenure decisions. Not surprising, a perennial problem for most journals is getting reviews in a timely fashion. So I'm not surprised when I hear the question: ‘why should I review papers?’ Luckily, I can offer several good answers. The most obvious reason is that it is a service to the community. Service is good for the soul, and donating our time to our field makes us feel better about ourselves. Another obvious reason is to return the favor. Every time we submit a paper to a journal, four or five people work for us: three reviewers, an editor, and possibly an editor-in-chief. We don't pay them in money, but we can pay for their time by reviewing their papers. A less obvious reason is to learn. Reviewing papers is a good way to learn new results early. Of course it's unethical to use results before they're published, but the reviewers know the results as soon as they're published. Moreover, the best ideas don't just give you ideas upon which to build your research, but they change the way you think. And changing your thinking earlier than others always helps. Another thing we learn is how to write papers. Seeing good (and bad!) examples help us develop and refine our own writing skills. Perhaps the most subtle reason for reviewing is to enhance our reputation. This is especially true for young scientists. Reviewing papers allows us to help senior scientists who ask for the reviews, and we can build a reputation as team player who supports journals. If our reviews are of high quality, the editors and associate editors notice and remember them. Reputation is the most important intangible professional asset that we have. Spending time to build that reputation pays off in unexpected ways. The complement for a senior scientist is that reviewing allows us to help shape the field by giving back our experience. This is a way to teach the authors, many of whom are professionally younger. The time required to review a paper varies tremendously by the subject, paper, and reviewer. Theoretical papers take longer, and if the reviewer has read many of the references the review takes less time. Conscientious reviewers will spend between 1 and 4 Hours for a conference paper and between 4 and 12 hours for a journal paper. As we gain experience, we tend to spend less time. Some reviewers will focus on the big picture issues, whereas some will focus more on details (which takes more time). So we have many good reasons to review papers. Of course, it's always okay to say ‘no’. Although saying no to a review request also influences your reputation, nobody can review every paper offered and it's far better to say ‘no’ at the beginning than to say ‘yes’ and then be very slow. The most important thing is to respond to every request. A fast ‘no’ allows the editor to go to the next person on the list, whereas a slow ‘no’ (or no response!) introduces delay into the process, which hurts the authors and the journal. And returning that review within a reasonable time is crucially important. We've all sent off papers to journals only to have to wait … and wait … and wait. I'm glad to report that, with the help of our wonderful editorial board and reviewers, STVR has made significant progress on timeliness. Time to review has come down from close to a year to about five months. My goal as editor is to reduce that further to three months.
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
2007 Editorial: Reflections on the past, present and future
Martin R. Woodward, A. Jefferson Offutt
Softw. Test. Verification Reliab.2
2006 MuJava: a mutation system for java
abstract
Mutation testing is a valuable experimental research technique that has been used in many studies. It has been experimentally compared with other test criteria, and also used to support experimental comparisons of other test criteria, by using mutants as a method to create faults. In effect, mutation is often used as a ``gold standard'' for experimental evaluations of test methods. Although mutation testing is powerful, it is a complicated and computationally expensive testing method. Therefore, automated tool support is indispensable for conducting mutation testing. This demo presents a publicly available mutation system for Java that supports both method-level mutants and class-level mutants. MuJava can be freely downloaded and installed with relative ease under both Unix and Windows. MuJava is offered as a free service to the community and we hope that it will promote the use of mutation analysis for experimental research in software testing.
Yu-Seung Ma, A. Jefferson Offutt, Yong Rae Kwon
ICSE2
2006 An evaluation of combination strategies for test case selection
Mats Grindal, Birgitta Lindström, A. Jefferson Offutt, Sten F. Andler
Empir. Softw. Eng.3
2006 Input validation analysis and testing
Jane Huffman Hayes, A. Jefferson Offutt
Empir. Softw. Eng.2
2006 Maintainability of the kernels of open-source operating systems: A comparison of Linux with FreeBSD, NetBSD, and OpenBSD
Liguo Yu, Stephen R. Schach, Kai Chen 0010, Gillian Z. Heller, A. Jefferson Offutt
J. Syst. Softw.5
2006 Integration testing of object-oriented components using finite state machines
abstract
Abstract In object‐oriented terms, one of the goals of integration testing is to ensure that messages from objects in one class or component are sent and received in the proper order and have the intended effect on the state of the objects that receive the messages. This research extends an existing single‐class testing technique to integration testing of multiple classes. The single‐class technique models the behaviour of a single class as a finite state machine, transforms the representation into a data flow graph that explicitly identifies the definitions and uses of each state variable of the class, and then applies conventional data flow testing to produce test case specifications that can be used to test the class. This paper extends those ideas to inter‐class testing by developing flow graphs, finding paths between pairs of definitions and uses, detecting some infeasible paths and automatically generating tests for an arbitrary number of classes and components. It introduces flexible representations for message sending and receiving among objects and allows concurrency among any or all classes and components. Data flow graphs are stored in a relational database and database queries are used to gather def‐use information. This approach is conceptually simple, mathematically precise, quite powerful and general enough to be used for traditional data flow analysis. This testing approach relies on finite state machines, database modelling and processing techniques and algorithms for analysis and traversal of directed graphs. The paper presents empirical results of the approach applied to an automotive system. This work was prepared by U.S. Government employees as part of their official duties and is, therefore, a work of the U.S. Government and not subject to copyright. Published in 2006 by John Wiley & Sons, Ltd.
Leonard Gallagher, A. Jefferson Offutt, Anthony Cincotta
Softw. Test. Verification Reliab.2
2006 A Tribute to Martin Woodward
A. Jefferson Offutt, Derek Yates, Robert M. Hierons, Michael A. Hennell, Peter Mitchell
Softw. Test. Verification Reliab.2
2005 Testing Web Services by XML Perturbation
abstract
The eXtensible Markup Language (XML) is widely used to transmit data across the Internet. XML schemas are used to defile the syntax of XML messages. XML-based applications can receive messages from arbitrary applications, as long as they follow the protocol defined by the schema. A receiving application must either validate XML messages, process the data in the XML message without validation, or modify the XML message to ensure that it conforms to the XML schema. A problem for developers is how well the application performs the validation, data processing, and, when necessary, transformation. This paper describes and gives examples of a method to generate tests for XML-based communication by modifying and then instantiating XML schemas. The modified schemas are based on precisely defined schema primitive perturbation operators
Wuzhi Xu, A. Jefferson Offutt, Juan Luo
ISSRE2
2005 Testing Web applications by modeling with FSMs
Anneliese Amschler Andrews, A. Jefferson Offutt, Roger T. Alexander
Softw. Syst. Model.2
2005 Combination testing strategies: a survey
abstract
Combination strategies are test case selection methods that identify test cases by combining values of the different test object input parameters based on some combinatorial strategy. This survey presents 16 different combination strategies, covering more than 40 papers that focus on one or several combination strategies. This collection represents most of the existing work performed on combination strategies. This survey describes the basic algorithms used by the combination strategies. Some properties of combination strategies, including coverage criteria and theoretical bounds on the size of test suites, are also included in this description. This survey paper also includes a subsumption hierarchy that attempts to relate the various coverage criteria associated with the identified combination strategies. Copyright © 2005 John Wiley & Sons, Ltd.
Mats Grindal, A. Jefferson Offutt, Sten F. Andler
Softw. Test. Verification Reliab.2
2005 MuJava: an automated class mutation system
abstract
Several module and class testing techniques have been applied to object-oriented (OO) programs, but researchers have only recently begun developing test criteria that evaluate the use of key OO features such as inheritance, polymorphism, and encapsulation. Mutation testing is a powerful testing technique for generating software tests and evaluating the quality of software. However, the cost of mutation testing has traditionally been so high that it cannot be applied without full automated tool support. This paper presents a method to reduce the execution cost of mutation testing for OO programs by using two key technologies, mutant schemata generation (MSG) and bytecode translation. This method adapts the existing MSG method for mutants that change the program behaviour and uses bytecode translation for mutants that change the program structure. A key advantage is in performance: only two compilations are required and both the compilation and execution time for each is greatly reduced. A mutation tool based on the MSG/bytecode translation method has been built and used to measure the speedup over the separate compilation approach. Experimental results show that the MSG/bytecode translation method is about five times faster than separate compilation. Copyright © 2004 John Wiley & Sons, Ltd.
Yu-Seung Ma, A. Jefferson Offutt, Yong Rae Kwon
Softw. Test. Verification Reliab.2
2004 Mutation-Based Testing Criteria for Timeliness
abstract
Temporal correctness is crucial to the dependability of real-time systems. Few methods exist to test for temporal correctness and most existing methods are ad-hoc. A problem with testing real-time applications is the dependency on the execution time and execution order of individual tasks. Thus, the response times for the tasks may be non-deterministic with respect to inputs. Conventional test coverage criteria ignore task interleaving and tinting and, thus do not help determine which execution orders need to be exercised to test for temporal correctness. This paper presents test criteria based on mutation to test timeliness. We also show how previously proposed methods in specification based testing, can be applied to testing real-time systems
Robert Nilsson, A. Jefferson Offutt, Sten F. Andler
COMPSAC2
2004 Bypass Testing of Web Applications
abstract
Web software applications are increasingly being deployed in sensitive situations. Web applications are used to transmit, accept and store data that is personal, company confidential and sensitive. Input validation testing (IVT) checks user inputs to ensure that they conform to the program's requirements, which is particularly important for software that relies on user inputs, including Web applications. A common technique in Web applications is to perform input validation on the client with scripting languages such as JavaScript. An insidious problem with client-side input validation is that end users can bypass this validation. Bypassing validation can cause failures in the software, and can also break the security on Web applications, leading to unauthorized access to data, system failures, invalid purchases and entry of bogus data. We are developing a strategy called bypass testing to create client-side tests for Web applications that intentionally violate explicit and implicit checks on user inputs. This paper describes the strategy, defines specific rules and adequacy criteria for tests, describes a proof-of-concept automated tool, and presents initial empirical results from applying bypass testing.
A. Jefferson Offutt, Xiaochen Du
ISSRE1
2004 Open-Source Change Logs
Kai Chen 0010, Stephen R. Schach, Liguo Yu, A. Jefferson Offutt, Gillian Z. Heller
Empir. Softw. Eng.4
2004 Categorization of Common Coupling and Its Application to the Maintainability of the Linux Kernel
abstract
Data coupling between modules, especially common coupling, has long been considered a source of concern in software design, but the issue is somewhat more complicated for products that are comprised of kernel modules together with optional nonkernel modules. This paper presents a refined categorization of common coupling based on definitions and uses between kernel and nonkernel modules and applies the categorization to a case study. Common coupling is usually avoided when possible because of the potential for introducing risky dependencies among software modules. The relative risk of these dependencies is strongly related to the specific definition-use relationships. In a previous paper, we presented results from a longitudinal analysis of multiple versions of the open-source operating system Linux. This paper applies the new common coupling categorization to version 2.4.20 of Linux, counting the number of instances of common coupling between each of the 26 kernel modules and all the other nonkernel modules. We also categorize each coupling in terms of the definition-use relationships. Results show that the Linux kernel contains a large number of common couplings of all types, raising a concern about the long-term maintainability of Linux.
Liguo Yu, Stephen R. Schach, Kai Chen 0010, A. Jefferson Offutt
IEEE Trans. Software Eng.4
2003 Coverage Criteria for Logical Expressions
abstract
A large number of coverage criteria to generate tests from logical expressions have been proposed. Although there have been large variations in the terminology, the articulation of the criteria and the original source of the expressions, many of these criteria are fundamentally the same. The most commonly known and widely used criterion is that of modified condition decision coverage (MCDC), but some articulations of MCDC have had some ambiguities. This has led to confusion on the part of testers, students, and tool developers on how best to implement these test criteria. This paper presents a complete comprehensive set of criteria that incorporate all the existing criteria, and eliminates the ambiguities by introducing precise definitions of the various possibilities.
Paul Ammann, A. Jefferson Offutt
ISSRE2
2003 Determining the Distribution of Maintenance Categories: Survey versus Measurement
Stephen R. Schach, Liguo Yu, Gillian Z. Heller, A. Jefferson Offutt
Empir. Softw. Eng.5
2003 Quality Impacts of Clandestine Common Coupling
Stephen R. Schach, David R. Wright 0002, Gillian Z. Heller, A. Jefferson Offutt
Softw. Qual. J.5
2003 Generating test data from state-based specifications
abstract
Abstract Although the majority of software testing in industry is conducted at the system level, most formal research has focused on the unit level. As a result, most system‐level testing techniques are only described informally. This paper presents formal testing criteria for system level testing that are based on formal specifications of the software. Software testing can only be formalized and quantified when a solid basis for test generation can be defined. Formal specifications represent a significant opportunity for testing because they precisely describe what functions the software is supposed to provide in a form that can be automatically manipulated. This paper presents general criteria for generating test inputs from state‐based specifications. The criteria include techniques for generating tests at several levels of abstraction for specifications (transition predicates, transitions, pairs of transitions and sequences of transitions). These techniques provide coverage criteria that are based on the specifications and are made up of several parts, including test prefixes that contain inputs necessary to put the software into the appropriate state for the test values. The test generation process includes several steps for transforming specifications to tests. These criteria have been applied to a case study to compare their ability to detect seeded faults. Copyright © 2003 John Wiley & Sons, Ltd.
A. Jefferson Offutt, Shaoying Liu, Aynur Abdurazik, Paul Ammann
Softw. Test. Verification Reliab.1
2002 Syntactic Fault Patterns in OO Programs
abstract
Although program faults are widely studied, there are many aspects of faults that we still do not understand, particularly about OO software. In addition to the simple fact that one important goal during testing is to cause failures and thereby detect faults, a full understanding of the characteristics of faults is crucial to several research areas. The power that inheritance and polymorphism brings to the expressiveness of programming languages also brings a number of new anomalies and fault types. In prior work we presented a fault model for the appearance and realization of OO faults that are specific to the use of inheritance and polymorphism. Many of these faults cannot appear unless certain syntactic patterns are used. The patterns are based on language constructs, such as overriding methods that directly define inherited state variables and non-inherited methods that call inherited methods. If one of these syntactic patterns is used, then we say the software contains an anomaly and possibly a fault. We describe the syntactic patterns for each OO fault type. These syntactic patterns can potentially be found with an automatic tool. Thus, faults can be uncovered and removed early in development.
Roger T. Alexander, A. Jefferson Offutt, James M. Bieman
ICECCS2
2002 An Empirical Comparison of Modularity of Procedural and Object-oriented Software
abstract
A commonly held belief is that applications written in object-oriented languages are more modular than those written in procedural languages. This paper presents results from an experiment that examines this hypothesis. Open source and industrial program modules written in the procedural languages of Fortran and C were compared with open source program modules written in the object-oriented languages of C++ and Java. The metrics examined in this study were lines of code per module and number of parameters per module. The results of the investigation support the hypothesis. The modules of the object-oriented programs were found to be half the size of those of the procedural programs and the average number of parameters per module for the object-oriented programs was approximately half that of the procedural programs. Thus the object-oriented programs were twice as modular as the procedural programs. An unexpected result was that the C++ programs were found to be no more modular than the C programs.
Lisa K. Ferrett, A. Jefferson Offutt
ICECCS2
2002 Fault Detection Capabilities of Coupling-based OO Testing
abstract
Object-oriented programs cause a shift in focus from software units to the way software classes and components are connected. Thus, we are finding that we need less emphasis on unit testing and more on integration testing. The compositional relationships of inheritance and aggregation, especially when combined with polymorphism, introduce new kinds of integration faults, which can be covered using testing criteria that take the effects of inheritance and polymorphism into account. This paper demonstrates, via a set of experiments, the relative effectiveness of several coupling-based OO testing criteria and branch coverage. OO criteria are all more effective at detecting faults due to the use of inheritance and polymorphism than branch coverage.
Roger T. Alexander, A. Jefferson Offutt, James M. Bieman
ISSRE2
2002 Inter-Class Mutation Operators for Java
abstract
The effectiveness of mutation testing depends heavily on the types of faults that the mutation operators are designed to represent. Therefore, the quality of the mutation operators is key to mutation testing. Mutation testing has traditionally been applied to procedural-based languages, and mutation operators have been developed to support most of their language features. Object-oriented programming languages contain new language features, most notably inheritance, polymorphism, and dynamic binding. Not surprisingly; these language features allow new kinds of faults, some of which are not modeled by traditional mutation operators. Although mutation operators for OO languages have previously been suggested, our work in OO faults indicate that the previous operators are insufficient to test these OO language features, particularly at the class testing level. This paper introduces a new set of class mutation operators for the OO language Java. These operators are based on specific OO faults and can be used to detect faults involving inheritance, polymorphism, and dynamic binding, thus are useful for inter-class testing. An initial Java mutation tool has recently been completed, and a more powerful version is currently under construction.
Yu-Seung Ma, Yong Rae Kwon, A. Jefferson Offutt
ISSRE3
2001 Deriving Tests From Software Architectures
abstract
Software architectures are intended to describe essential high level structural and behavioral characteristics of a system. Architecture Description Languages (ADLs) describe these characteristics in ways that can be analyzed and manipulated algorithmically. This provides a unique opportunity for deriving tests at the system level. The paper defines formal testing criteria based on architecture relations, which are paths that architectural components use to communicate. The criteria have been applied to a specific ADL. Results from a comparative empirical study on industrial software are presented.
Zhenyi Jin, A. Jefferson Offutt
ISSRE2
2001 Generating Test Cases for XML-Based Web Component Interactions Using Mutation Analysis
abstract
Web software systems are built using heterogeneous software components. They interact by passing messages that exchange data and activity state information. Such heterogeneous message transfers can be structured using the eXtensible Markup Language (XML), which allows a flexible common data exchange. Parsers have been developed to check the syntax of component interactions, but there are as yet no techniques for checking the semantic correctness of the interactions. The paper presents a technique for using mutation analysis to test the semantic correctness of XML-based component interactions. The Web software interactions are specified using an Interaction Specification Model (ISM) that consists of document type definitions, messaging specifications, and a set of constraints. Test cases are XML messages that are passed between the Web software components. Classes of interaction-specific mutation operators are introduced and applied to the ISM to generate mutant interactions and test cases.
Suet Chun Lee, A. Jefferson Offutt
ISSRE2
2001 A Fault Model for Subtype Inheritance and Polymorphism
abstract
Although program faults are widely studied, there are many aspects of faults that we still do not understand, particularly about OO software. In addition to the simple fact that one important goal during testing is to cause failures and thereby detect faults, a full understanding of the characteristics of faults is crucial to several research areas. The power that inheritance and polymorphism brings to the expressiveness of programming languages also brings a number of new anomalies and fault types. This paper presents a model for the appearance and realization of OO faults and defines and discusses specific categories of inheritance and polymorphic faults. The model and categories can be used to support empirical investigations of object-oriented testing techniques, to inspire further research into object-oriented testing and analysis, and to help improve design and development of object-oriented software.
A. Jefferson Offutt, Roger T. Alexander, Quansheng Xiao, Chuck Hutchinson
ISSRE1
2000 Evaluation of Three Specification-Based Testing Criteria
abstract
This paper compares three specification-based testing criteria using Mathur and Wong's PROBSUBSUMES measure. The three criteria are specification-mutation coverage, full predicate coverage, and transition-pair coverage. A novel aspect of the work is that each criterion is encoded in a model checker, and the model checker is used first to generate test sets for each criterion and then to evaluate test sets against alternate criteria. Significantly, the use of the model checker for generation of test sets eliminates human bias from this phase of the experiment. The strengths and weaknesses of the criteria are discussed.
Aynur Abdurazik, Paul Ammann, Wei Ding 0003, A. Jefferson Offutt
ICECCS4
2000 An Analysis Tool for Coupling-Based Integration Testing
abstract
This research is part of a project to develop practical, effective, formalizable, automatable techniques for integration testing. Integration testing is an important part of the testing process, but few integration testing techniques have been systematically studied or defined. This paper discusses the design and implementation of an analysis tool for measuring the amount of coverage achieved by a set of test data according to a set of previously defined coupling criteria. This tool can be used to support integration testing of software components. The coupling-based testing technique, which has been described elsewhere, is summarized, and coverage algorithms are discussed. The focus of this paper is on the instrumentation techniques and an analysis tool built for Java programs. It was built in Java using the general Java parser JavaCC and the Java Tree Builder (JTB). We are currently using this tool to gather experimental data on the efficacy and the usefulness of the technique.
A. Jefferson Offutt, Aynur Abdurazik, Roger T. Alexander
ICECCS1
2000 Criteria for Testing Polymorphic Relationships
abstract
The emphasis in object oriented programs is on defining abstractions that have both state and behavior. This emphasis causes a shift in focus from software units to the way software components are connected. Thus, we are finding that we need less emphasis on unit testing and more on integration testing. The compositional relationships of inheritance and aggregation, especially when combined with polymorphism, introduce new kinds of integration faults. The paper presents results from an ongoing research project that has the goal of improving the quality of object oriented software. New testing criteria are introduced that take the effects of inheritance and polymorphism into account. These criteria are based on the new analysis technique of quasi-interprocedural data flow analysis. These testing criteria can improve the quality of object oriented software by ensuring that integration tests are high quality.
Roger T. Alexander, A. Jefferson Offutt
ISSRE2
1999 Criteria for Generating Specification-Based Tests
abstract
This paper presents general criteria for generating test inputs from state-based specifications. Software testing can only be formalized and quantified when a solid basis for test generation can be defined. Formal specifications of complex systems represent a significant opportunity for testing because they precisely describe what functions the software is supposed to provide in a form that can easily be manipulated. These techniques provide coverage criteria that are based on the specifications, and are made up of several parts, including test prefixes that contain inputs necessary to put the software into the appropriate state for the test values. The test generation process includes several steps for transforming specifications to tests. Empirical results from a comparative case study application of these criteria are presented.
A. Jefferson Offutt, Yiwei Xiong, Shaoying Liu
ICECCS1
1999 Increased software reliability through input validation analysis and testing
abstract
The input validation testing (IVT) technique has been developed to address the problem of statically analyzing input command syntax as defined in an English textual interface and requirements specifications and then generating test cases for input validation testing. The technique does not require design or code, so it can be applied early in the life cycle. A proof-of-concept tool has been implemented and validation has been performed. Empirical validation on industrial software shows that the IVT method found more requirements specification defects than senior testers, generated test cases with higher syntactic coverage than senior testers, and found defects that were not found by the test cases of senior testers. Additionally, the tool performed at a much-reduced cost.
Jane Huffman Hayes, A. Jefferson Offutt
ISSRE2
1999 Generating test data from SOFL specifications
A. Jefferson Offutt, Shaoying Liu
J. Syst. Softw.1
1999 The Dynamic Domain Reduction Procedure for Test Data Generation
abstract
Test data generation is one of the most technically challenging steps of testing software, but most commercial systems currently incorporate very little automation for this step. This paper presents results from a project that is trying to find ways to incorporate test data generation into practical test processes. The results include a new procedure for automatically generating test data that incorporates ideas from symbolic evaluation, constraint-based testing, and dynamic test data generation. It takes an initial set of values for each input, and dynamically ‘pushes’ the values through the control-flow graph of the program, modifying the sets of values as branches in the program are taken. The result is usually a set of values for each input parameter that has the property that any choice from the sets will cause the path to be traversed. This procedure uses new analysis techniques, offers improvements over previous research results in constraint-based testing, and combines several steps into one coherent process. The dynamic nature of this procedure yields several benefits. Moving through the control flow graph dynamically allows path constraints to be resolved immediately, which is more efficient both in space and time, and more often successful than constraint-based testing. This new procedure also incorporates an intelligent search technique based on bisection. The dynamic nature of this procedure also allows certain improvements to be made in the handling of arrays, loops, and expressions; language features that are traditionally difficult to handle in test data generation systems. The paper presents the test data generation procedure, examples to explain the working of the procedure, and results from a proof-of-concept implementation. Copyright © 1999 John Wiley & Sons, Ltd.
A. Jefferson Offutt, Zhenyi Jin
Softw. Pract. Exp.1
1998 Coupling-based Criteria for Integration Testing
abstract
Integration testing is an important part of the testing process, but few integration testing techniques have been systematically studied or defined. The goal of this research is to develop practical, effective, formalizable, automatable techniques for testing of connections between components during software integration. This paper presents an integration testing technique that is based on couplings between software components. This technique can be used to support integration testing of software components, and satisfies part of the USA's Federal Aviation Authority's requirements for structural coverage analysis of software. The coupling-based testing technique is described, and the coverage criteria for three types of couplings are defined. Techniques and algorithms for developing coverage analysers to measure the extent to which a test set satisfies the criteria are presented, and results from a comparative case study are presented. © 1998 John Wiley & Sons, Ltd.
Zhenyi Jin, A. Jefferson Offutt
Softw. Test. Verification Reliab.2
1998 SOFL: A Formal Engineering Methodology for Industrial Applications
abstract
Formal methods have yet to achieve wide industrial acceptance for several reasons. They are not well integrated into established industrial software processes, their application requires significant abstraction and mathematical skills, and existing tools do not satisfactorily support the entire formal software development process. We have proposed a language called SOFL (Structured-Object-based-formal Language) and a SOFL methodology for system development that attempts to address these problems using an integration of formal methods, structured methods and object oriented methodology. Construction of a system uses structured methods in requirements analysis and specifications, and an object based methodology during design and implementation stages, with formal methods applied throughout the development in a manner that best suits their capabilities. The paper describes the SOFL methodology, which introduces some substantial changes from current formal methods practice. A comprehensive, practical case study of an actual industrial Residential Suites Management System illustrates how SOFL is used.
Shaoying Liu, A. Jefferson Offutt, Chris Ho-Stuart, Mitsuru Ohba
IEEE Trans. Software Eng.2
1997 Maintaining Knowledge Currency in the 21st Century
abstract
Software engineering is a rapidly changing discipline, and will continue to be so for the foreseeable future. This pace of change brings both problems and opportunities to universities that teach software engineering. Engineers ore no longer satisfied with one or two initial university education experiences, but by necessity are becoming lifetime learners, with frequent trips back to educational providers. This recurring education is needed to update engineers' knowledge with new ideas and concepts, and to update engineers' skills. In this paper, we take the position that universities can and should respond to this situation with a new model for graduate software engineering education, which we call professional currency certificates. These courses should offer the depth of knowledge and university academic credit that traditional academic courses offer, but with the convenience and practical nature of corporate training courses. This hybrid model results in a new kind of course that more closely meets the needs of lifetime learners.
Paul Ammann, A. Jefferson Offutt
CSEE&T2
1997 An Approach to Fault Modeling and Fault Seeding Using the Program Dependence Graph
Mary Jean Harrold, A. Jefferson Offutt, Kanupriya Tewary
J. Syst. Softw.2
1997 Automatically Detecting Equivalent Mutants and Infeasible Paths
abstract
Mutation testing is a technique for testing software units that has great potential for improving the quality of testing, and thereby increasing the ability to assure the high reliability of critical software. It will be shown that recent advances in mutation research have brought a practical mutation testing system closer to reality. One recent advance is a partial solution to the problem of automatically detecting equivalent mutant programs. Equivalent mutants are currently detected by hand, which makes it very expensive and time-consuming. The problem of detecting equivalent mutants is a specific instance of a more general problem, commonly called the feasible path problem, which says that for certain structural testing criteria some of the test requirements are infeasible in the sense that the semantics of the program imply that no test case satisfies the test requirements. Equivalent mutants, unreachable statements in path testing techniques, and infeasible DU-pairs in data flow testing are all instances of the feasible path problem. This paper presents a technique that uses mathematical constraints, originally developed for test data generation, to detect some equivalent mutants and infeasible paths automatically. © 1997 John Wiley & Sons, Ltd.
A. Jefferson Offutt
Softw. Test. Verification Reliab.1
1996 Coupling-based Integration Testing
abstract
Integration testing is an important part of the testing process, but few integration testing techniques have been systematically studied or defined. This paper presents an integration testing technique based on couplings between software components. The coupling-based testing technique is described, and coverage criteria for three types of 12 coupling levels are defined. This technique can be used to support integration testing of software components, and satisfies part of the FAA's requirements for structural coverage analysis of software.
Zhenyi Jin, A. Jefferson Offutt
ICECCS2
1996 Algorithmic Analysis of the Impact of Changes to Object-Oriented Software
abstract
As the software industry has matured, we have shifted our resources from being primarily devoted to developing new software systems to primarily making modifications in evolving software systems. A major problem for developers in an evolutionary environment is that seemingly small changes can ripple throughout the system to have major unintended impacts elsewhere. As a result, software developers need mechanisms to understand how a change to a software system will affect the rest of the system. Although the effects of changes in object-oriented software are restricted, they are also more subtle and more difficult to detect. This paper presents algorithms to analyze the potential impacts of changes to object-oriented software, taking into account encapsulation, inheritance and polymorphism. This technique allows software developers to perform "what if" analyses on the effect of proposed changes, and thereby choose the change that has the least influence on the rest of the system. The analysis also adds valuable information to regression testing, by suggesting what classes and methods need to be re-tested, and to project managers, who can use the results for cost estimation and schedule planning.
A. Jefferson Offutt
ICSM2
1996 A Semantic Model of Program Faults
abstract
Program faults are artifacts that are widely studied, but there are many aspects of faults that we still do not understand. In addition to the simple fact that one important goal during testing is to cause failures and thereby detect faults, a full understanding of the characteristics of faults is crucial to several research areas in testing. These include fault-based testing, testability, mutation testing, and the comparative evaluation of testing strategies. In this workshop paper, we explore the fundamental nature of faults by looking at the differences between a syntactic and semantic characterization of faults. We offer definitions of these characteristics and explore the differentiation. Specifically, we discuss the concept of "size" of program faults --- the measurement of size provides interesting and useful distinctions between the syntactic and semantic characterization of faults. We use the fault size observations to make several predictions about testing and present preliminary data that supports this model. We also use the model to offer explanations about several questions that have intrigued testing researchers.
A. Jefferson Offutt, Jane Huffman Hayes
ISSTA1
1996 An Experimental Evaluation of Data Flow and Mutation Testing
abstract
Two experimental comparisons of data flow and mutation testing are presented. These techniques are widely considered to be effective for unit-level software testing, but can only be analytically compared to a limited extent. We compare the techniques by evaluating the effectiveness of test data developed for each. We develop ten independent sets of test data for a number of programs: five to satisfy the mutation criterion and five to satisfy the all-uses data-flow criterion. These test sets are developed using automated tools, in a manner consistent with the way a test engineer might be expected to generate test data in practice. We use these test sets in two separate experiments. First we measure the effectiveness of the test data that was developed for one technique in terms of the other. Second, we investigate the ability of the test sets to find faults. We place a number of faults into each of our subject programs, and measure the number of faults that are detected by the test sets. Our results indicate that while both techniques are effective, mutation-adequate test sets are closer to satisfying the data flow criterion, and detect more faults.
A. Jefferson Offutt, Kanupriya Tewary, Tong Zhang 0006
Softw. Pract. Exp.1
1996 An Experimental Determination of Sufficient Mutant Operators
abstract
Mutation testing is a technique for unit-testing software that, although powerful, is computationally expensive, The principal expense of mutation is that many variants of the test program, called mutants, must be repeatedly executed.This article quantifies the expense of mutation in terms of the number of mutants that are created, then proposes and evaluates a technique that reduces the number of mutants by an order of magnitude.Selective mutation reduces.the cost of mutation testing by reducing the number of mutants, This article reports experimental results that compare selective mutation testing with standard, or nonselective, mutation testing, and results that quantify the savings achieved by selective mutation testing, The results support the hypothesis that selective mutation is almost as strong as nonselective mutation: in experimental trials selective mutation provides almost the same coverage as nonselective mutation.with a four-fold or more reduction in the number of mutants.
A. Jefferson Offutt, Ammei Lee, Gregg Rothermel, Roland H. Untch, Christian Zapf
ACM Trans. Softw. Eng. Methodol.1
1994 A Practical System for Mutation Testing: Help for the Common Programmer
abstract
Mutation testing is a technique for unit testing software that, although powerful, is computationally expensive. Recent engineering advances have given us techniques and algorithms for significantly reducing the cost of mutation testing. These techniques include a new algorithmic execution technique called schema-based mutation, an approximation technique called weak mutation, a reduction technique called selective mutation, and algorithms for automatic test data generation. This paper outlines a design for a system that will approximate mutation, but in a way that will be accessible to everyday programmers. We envisage a system to which a programmer can submit a program unit, and get back a set of input/output pairs that are guaranteed to form an effective test of the unit by being close to mutation adequate.
A. Jefferson Offutt
ITC1
1994 Using Compiler Optimization Techniques to Detect Equivalent Mutants
abstract
Abstract Mutation analysis is a software testing technique that requires the tester to generate test data that will find specific, well‐defined errors. Mutation testing executes many slightly differing versions, called mutants, of the same program to evaluate the quality of the data used to test the program. Although these mutants are generated and executed efficiently by automated methods, many of the mutants are functionally equivalent to the original program and are not useful for testing. Recognizing and eliminating equivalent mutants is currently done by hand, a time‐consuming and arduous task. This problem is currently a major obstacle to the practical application of mutation testing. This paper presents extensions to previous work in detecting equivalent mutants; specifically, algorithms for determining several classes of equivalent mutants are presented, an implementation of these algorithms is discussed, and results from using this implementation are presented. These algorithms are based on data flow analysis and six compiler optimization techniques. Each of these techniques is described together with how they are used to detect equivalent mutants. The design of the tool and some experimental results using it are also presented.
A. Jefferson Offutt, W. M. Craft
Softw. Test. Verification Reliab.1
1994 An Empirical Evaluation of Weak Mutation
abstract
Mutation testing is a fault-based technique for unit-level software testing. Weak mutation was proposed as a way to reduce the expense of mutation testing. Unfortunately, weak mutation is also expected to provide a weaker test of the software than mutation testing does. This paper presents results from an implementation of weak mutation, which we used to evaluate the effectiveness versus the efficiency of weak mutation. Additionally, we examined several options in an attempt to find the most appropriate way to implement weak mutation. Our results indicate that weak mutation can be applied in a manner that is almost as effective as mutation testing, and with significant computational savings.>
A. Jefferson Offutt, Stephen D. Lee
IEEE Trans. Software Eng.1
1994 Correction to "An Empirical Evaluation of Weak Mutation"
A. Jefferson Offutt, Stephen D. Lee
IEEE Trans. Software Eng.1
1993 An Experimental Evaluation of Selective Mutation
A. Jefferson Offutt, Gregg Rothermel, Christian Zapf
ICSE1
1993 Mutation Analysis Using Mutant Schemata
abstract
Mutation analysis is a powerful technique for assessing and improving the quality of test data used to unit test software. Unfortunately, current automated mutation analysis systems suffer from severe performance problems. This paper presents a new method for performing mutation analysis that uses program schemata to encode all mutants for a program into one metaprogram, which is subsequently compiled and run at speeds substantially higher than achieved by previous interpretive systems. Preliminary performance improvements of over 300% are reported. This method has the additional advantages of being easier to implement than interpretive systems, being simpler to port across a wide range of hardware and software platforms, and using the same compiler and run-time support system that is used during development and/or deployment.
Roland H. Untch, A. Jefferson Offutt, Mary Jean Harrold
ISSTA2
1993 A software metric system for module coupling
A. Jefferson Offutt, Mary Jean Harrold, Priyadarshan Kolte
J. Syst. Softw.1
1993 Experimental Results from an Automatic Test Case Generator
abstract
Constraint-based testing is a novel way of generating test data to detect specific types of common programming faults. The conditions under which faults will be detected are encoded as mathematical systems of constraints in terms of program symbols. A set of tools, collectively called Godzilla, has been implemented that automatically generates constraint systems and solves them to create test cases for use by the Mothra testing system. Experimental results from using Godzilla show that the technique can produce test data that is very close in terms of mutation adequacy to test data that is produced manually, and at substantially reduced cost. Additionally, these experiments have suggested a new procedure for unit testing, where test cases are viewed as throw-away items rather than scarce resources.
Richard A. DeMillo, A. Jefferson Offutt
ACM Trans. Softw. Eng. Methodol.2
1992 Mutation Testing of Software Using MIMD Computer
A. Jefferson Offutt, Roy P. Pargas, Scott V. Fichter, Prashant K. Khambekar
ICPP (2)1
1992 Estimation and Enhancement of Real-Time Software Reliability Through Mutation Analysis
abstract
A simulation-based method for obtaining numerical estimates of the reliability of N-version, real-time software is proposed. An extended stochastic Petri net is used to represent the synchronization structure of N versions of the software, where dependencies among versions are modeled through correlated sampling of module execution times. The distributions of execution times are derived from automatically generated test cases that are based on mutation testing. Since these test cases are designed to reveal software faults, the associated execution times and reliability estimates are likely to be conservative. Experimental results using specifications for NASA's planetary lander control software suggest that mutation-based testing could hold greater potential for enhancing reliability than the desirable but perhaps unachievable goal of independence among N versions. Nevertheless, some support for N-version enhancement of high-quality, mutation-tested code is also offered. Mutation analysis could also be valuable in the design of fault-tolerant software systems.>
Robert Geist, A. Jefferson Offutt, Frederick C. Harris Jr.
IEEE Trans. Computers2
1992 Investigations of the Software Testing Coupling Effect
abstract
Fault-based testing strategies test software by focusing on specific, common types of faults. The coupling effect hypothesizes that test data sets that detect simple types of faults are sensitive enough to detect more complex types of faults. This paper describes empirical investigations into the coupling effect over a specific class of software faults. All of the results from this investigation support the validity of the coupling effect. The major conclusion from this investigation is the fact that by explicitly testing for simple faults, we are also implicitly testing for more complicated faults, giving us confidence that fault-based testing is an effective way to test software.
A. Jefferson Offutt
ACM Trans. Softw. Eng. Methodol.1
1991 Unit Testing Versus Integration Testing
A. Jefferson Offutt
ITC1
1991 A Fortran Language System for Mutation-based Software Testing
abstract
Abstract Mutation analysis is a powerful technique for testing software systems. The Mothra software testing project uses mutation analysis as the basis for an integrated software testing environment. Mutation analysis requires executing many slightly differing versions of the same program to evaluate the quality of the data used to test the program. The current version of Mothra includes a complete language system that translates a program to be tested into intermediate code so that it and its mutated versions can be executed by an interpreter. In this paper, we discuss some of the unique requirements of a language system used in a mutation‐based testing environment. We then describe how these requirements affected the design and implementation of the Fortran 77 version of the Mothra system. We also describe the intermediate language used by Mothra and the features of the language system that are needed for software testing. The appendices contain a full description of the intermediate language and the mutation operators used by Mothra. The design and implementation techniques that were developed for Mothra are applicable for constructing not just software testing systems, but any type of program analysis system or language system for a special‐purpose application. In particular, we discuss decisions made and techniques developed by the Mothra team that can be useful in such applications as debuggers, program measurement tools, software development environments and other types of program analysis systems.
K. N. King, A. Jefferson Offutt
Softw. Pract. Exp.2
1991 Constraint-Based Automatic Test Data Generation
abstract
A novel technique for automatically generating test data is presented. The technique is based on mutation analysis and creates test data that approximate relative adequacy. It is a fault-based technique that uses algebraic constraints to describe test cases designed to find particular types of faults. A set of tools (collectively called Godzilla) that automatically generates constraints and solves them to create test cases for unit and module testing has been implemented. Godzilla has been integrated with the Mothra testing system and has been used as an effective way to generate test data that kill program mutants. The authors present an initial list of constraints and discuss some of the problems that have been solved to develop the complete implementation of the technique.>
Richard A. DeMillo, A. Jefferson Offutt
IEEE Trans. Software Eng.2
1988 Anatomy of a software engineering project
abstract
This paper discusses a complete software development project carried out in a one quarter undergraduate software engineering course. The project was the design and implementation of a complete system by 25 students. They worked in smaller groups on four functionally separate subsystems that were successfully integrated into a complete system. This was accomplished by using five advanced students to manage the groups, real users to criticize each step of the process, and UNIX tools to implement the subsystems. This paper describes the project, presents the methodologies used, and discusses both the positive and negative aspects of this course. It concludes by presenting a set of recommendations based on our experience with this project.
Catherine L. Bullard, Inez Caldwell, James Harrell, Cis Hinkle, A. Jefferson Offutt
SIGCSE5
1987 A Fortran 77 interpreter for mutation analysis
A. Jefferson Offutt, K. N. King
PLDI1