VLDB 2026 Research / reviewers in the wild / expert
Adam A. Porter
dblp:p/AdamAPorter
· DBLP profile ↗
63ranked-venue papers
14as first author
6since 2021 · last 2026
0000-0001-7764-7433ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 59 · 13 first-author · 5 since 2021Systems, architecture and hardware · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Understanding and Improving ML-based Static Analysis Result Classification via Explainable AI
Sai S. Yerramreddy, Mohammad Rafieian, Shiyi Wei, Adam A. Porter |
ICST | 4 |
| 2025 | A Reflection on "Advances in Software Inspections"abstractMichael Fagan’s work on the software inspection process defined a practical approach for improving software quality that was broadly adopted throughout the software industry. This work sparked numerous follow-on efforts aimed at understanding and improving the approach. Finally, the process has flexibly evolved over roughly 50 years, incorporating new languages, tools, techniques, and new understandings of its goals and benefits. Adam A. Porter, Harvey P. Siy, Lawrence G. Votta |
IEEE Trans. Software Eng. | 1 |
| 2023 | An empirical assessment of machine learning approaches for triaging reports of static analysis tools
Sai S. Yerramreddy, Austin Mordahl, Ugur Koc, Shiyi Wei, Jeffrey S. Foster, Marine Carpuat, Adam A. Porter |
Empir. Softw. Eng. | 7 |
| 2021 | Invited: Independent Verification and Validation of Security-Aware EDA Tools and IPabstractSecure silicon requires a seamless integration of new tools, new IP, and design flows to help designers protect integrated circuits from increasingly sophisticated attacks. Independent Validation and Verification (IV&V) of this integrated technology is important to ensure that the tools actually deliver on their security claims when used by independent parties (i.e., people who were not involved in designing the tools). This work discusses the principles and approaches for IV&V of such a complex design environment, including validation of the security strength of the various hardware security techniques, such as combinational and sequential logic locking, Trojan Detection, side-channel mitigation, and blockchain-based asset management. The main challenge in running an IV&V effort is to ensure that the process provides rigorous, methodical and provable evaluation of the claims of not only the component tools and IP, but whether such an integrated environment can produce security-hardened designs by a non-security expert. CCS Concepts • Hardware $\rightarrow$ Very large scale integration design; Methodologies for EDA; • Security and privacy $\rightarrow$ Security in hardware. Benjamin Tan 0001, Siddharth Garg, Ramesh Karri, Yuntao Liu 0001, Michael Zuzak, Abhisek Chakraborty, Ankur Srivastava 0001, Omid Aramoon, Qian Xu 0022, Gang Qu 0001, Adam A. Porter, Jeno Szep, Warren Savage |
DAC | 11 |
| 2021 | SATune: A Study-Driven Auto-Tuning Approach for Configurable Software Verification ToolsabstractMany program verification tools can be customized via run-time configuration options that trade off performance, precision, and soundness. However, in practice, users often run tools under their default configurations, because understanding these tradeoffs requires significant expertise. In this paper, we ask how well a single, default configuration can work in general, and we propose SATune, a novel tool for automatically configuring program verification tools for given target programs. To answer our question, we gathered a dataset that runs four well-known program verification tools against a range of C and Java benchmarks, with results labeled as correct, incorrect, or inconclusive (e.g., timeout). Examining the dataset, we find there is generally no one-size-fits-all best configuration. Moreover, a statistical analysis shows that many individual configuration options do not have simple tradeoffs: they can be better or worse depending on the program.Motivated by these results, we developed SATune, which constructs configurations using a meta-heuristic search. The search is guided by a surrogate fitness function trained on our dataset. We compare the performance of SATune to three baselines: a single configuration with the most correct results in our dataset; the most precise configuration followed by the most correct configuration (if needed); and the most precise configuration followed by random search (also if needed). We find that SATune outperforms these approaches by completing more correct tasks with high precision. In summary, our work shows that good configurations for verification tools are not simple to find, and SATune takes an important step towards automating the process of finding them. Ugur Koc, Austin Mordahl, Shiyi Wei, Jeffrey S. Foster, Adam A. Porter |
ASE | 5 |
| 2021 | Classifying User Requirements from Online Feedback in Small Dataset Environments using Deep LearningabstractAn overwhelming number of users access app repositories like App Store/Google Play and social media platforms like Twitter, where they provide feedback on digital experiences. This vast textual corpus comprising user feedback has the potential to unearth detailed insights regarding the users’ opinions on products and services. Various tools have been proposed that employ natural language processing (NLP) and traditional machine learning (ML) based models as an inexpensive mechanism to identify requirements in user feedback. However, they fall short on their classification accuracy over unseen data due to factors like the cost of generating voluminous de-biased labeled datasets and general inefficiency. Recently, Van Vliet et al. [1] achieved state-of-the-art results extracting and classifying requirements from user reviews through traditional crowdsourcing. Based on their reference classification tasks and outcomes, we successfully developed and validated a deep-learning-backed artificial intelligence pipeline to achieve a state-of-the-art averaged classification accuracy of ∼87% on standard tasks for user feedback analysis. This approach, which comprises a BERT-based sequence classifier, proved effective even in extremely low-volume dataset environments. Additionally, our approach drastically reduces the time and costs of evaluation, and improves on the accuracy measures achieved using traditional ML-/NLP-based techniques. Rohan Reddy Mekala, Asif Irfan, Eduard C. Groen, Adam A. Porter, Mikael Lindvall |
RE | 4 |
| 2019 | SeqScreen: a biocuration platform for robust taxonomic and biological process characterization of nucleic acid sequences of interestabstractRapid advancements in synthetic biology and nucleic acid synthesis, in particular concerns about its intentional or accidental misuse, call for more sophisticated screening tools to identify genes of interest within short sequence fragments. One major gap in predicting genes of concern is the inadequacy of current tools and ontologies to describe the specific biological processes of pathogenic proteins. The objective of this work is to design software that sensitively assigns taxonomic classifications, functional annotations, and biological processes of interest to short nucleotide sequences of unknown origin (50bp-1,000bp). The overarching goal is to perform sensitive characterization of short sequences and highlight specific pathogenic biological processes of interest (BPoIs). The SeqScreen software executes these tasks in analytical workflows with Nextflow and outputs results in a tab-delimited report. Local and global alignments differentiate hits to taxonomically-related sequences from similar but unrelated sequences, and an ensemble approach leverages multiple tools and databases to assign a variety of functional terms to each query sequence. Final biological process assessments are made from the predicted functional annotations, which leverage information in pre-existing databases, as well as new custom biocurations. Machine learning models predict each biological process of interest on large protein databases before incorporation into the SeqScreen framework to streamline computational efficiency, ensure reproducible results, allow for version control, and facilitate the review of the automated predictions by expert biocurators. The SeqScreen source code is available at https://gitlab.com/treangenlab/seqscreen. Dreycey Albin, Pravin Muthu, Gene Godbold, Mikael Lindvall, Madeline Diep, Adam A. Porter, Mihai Pop, Krista Ternus, Todd J. Treangen, Dan Nasko, Ryan A. Leo Elworth, Jacob Lu, Advait Balaji, Christian Diaz, Nidhi Shah, Jeremy D. Selengut, Chris Hulme-Lowe |
BIBM | 6 |
| 2019 | An Empirical Assessment of Machine Learning Approaches for Triaging Reports of a Java Static Analysis ToolabstractDespite their ability to detect critical bugs in software, developers consider high false positive rates to be a key barrier to using static analysis tools in practice. To improve the usability of these tools, researchers have recently begun to apply machine learning techniques to classify and filter false positive analysis reports. Although initial results have been promising, the long-term potential and best practices for this line of research are unclear due to the lack of detailed, large-scale empirical evaluation. To partially address this knowledge gap, we present a comparative empirical study of four machine learning techniques, namely hand-engineered features, bag of words, recurrent neural networks, and graph neural networks, for classifying false positives, using multiple ground-truth program sets. We also introduce and evaluate new data preparation routines for recurrent neural networks and node representations for graph neural networks, and show that these routines can have a substantial positive impact on classification accuracy. Overall, our results suggest that recurrent neural networks (which learn over a program's source code) outperform the other subject techniques, although interesting tradeoffs are present among all techniques. Our observations provide insight into the future research needed to speed the adoption of machine learning approaches in practice. Ugur Koc, Shiyi Wei, Jeffrey S. Foster, Marine Carpuat, Adam A. Porter |
ICST | 5 |
| 2016 | Coordinated Collaborative Testing of Shared Software ComponentsabstractSoftware developers commonly build their software systems by reusing other components developed and maintained by third-party developer groups. As the components evolve over time, new end-user machine configurations that contain new component versions will be added continuously for the potential user base. Therefore developers must test whether their components function correctly in the new configurations to ensure the quality of the overall systems. This would be achievable if developers could provision the configurations in house and conduct regression testing over the configurations. However, this is often very time-consuming and also there can be redundancy in test effort between developers when a common set of components is reused for providing the functionality of the systems. In this paper, we present a coordinated collaborative regression testing process for multiple developer groups. It involves a scheduling method for distributing test effort across the groups at component updates, with the objectives of reducing test redundancy between the groups and also shortening the time window in which compatibility faults are exposed to user community. The process is implemented on Conch, a collaborative test data repository and services we developed in our previous work. Conch has been modified to function as the test process coordinator, as well as the shared repository of test data. Our experiments over the 1.5-year evolution history of eleven components in the Ubuntu developer community show that developers can quickly discover compatibility faults by applying the coordinated process. Moreover, total testing time is comparable to the scenario where the developers conduct regression testing only at updates of their own components. Teng Long 0003, Il-Chul Yoon, Adam A. Porter, Atif M. Memon, Alan Sussman |
ICST | 3 |
| 2016 | iGen: dynamic interaction inference for configurable softwareabstractTo develop, analyze, and evolve today's highly configurable software systems, developers need deep knowledge of a system's configuration options, e.g., how options need to be set to reach certain locations, what configurations to use for testing, etc. Today, acquiring this detailed information requires manual effort that is difficult, expensive, and error prone. In this paper, we propose iGen, a novel, lightweight dynamic analysis technique that automatically discovers a program's interactions---expressive logical formulae that give developers rich and detailed information about how a system's configuration option settings map to particular code coverage. iGen employs an iterative algorithm that runs a system under a small set of configurations, capturing coverage data; processes the coverage data to infer potential interactions; and then generates new configurations to further refine interactions in the next iteration. We evaluated iGen on 29 programs spanning five languages; the breadth of this study would be unachievable using prior interaction inference tools. Our results show that iGen finds precise interactions based on a very small fraction of the number of possible configurations. Moreover, iGen's results confirm several earlier hypotheses about typical interaction distributions and structures. ThanhVu Nguyen, Ugur Koc, Javran Cheng, Jeffrey S. Foster, Adam A. Porter |
SIGSOFT FSE | 5 |
| 2016 | Enhancing Rules For Cloud Resource Provisioning Via Learned Software Performance ModelsabstractIn cloud computing, stakeholders deploy and run their software applications on a sophisticated infrastructure that is owned and managed by third-party providers. The ability of a given cloud infrastructure to effectively re-allocate resources to applications is referred to as elasticity. To enable elasticity, programmers study the behavior of applications and write scripts that guide the cloud to provision resources for these applications. This is an imprecise, laborious, manual and expensive approach that drastically increases the cost of application deployment and maintenance in the cloud. We propose an approach, coined as Provisioning Resources with Experimental SofTware mOdeling (PRESTO), to automatically learn behavioral models of software applications during performance testing in order to recommend programmers how to improve provisioning strategies that guide the cloud to (de)allocate resources to these applications. We applied PRESTO to two software applications and our experiments demonstrate that with PRESTO programmers can create rules for provisioning resources with a high degree of precision when the performance is about to worsen, so that the applications maintain their throughputs at the desired level. Mark Grechanik, Denys Poshyvanyk, Adam A. Porter |
ICPE | 4 |
| 2014 | iTree: Efficiently Discovering High-Coverage Configurations Using Interaction TreesabstractModern software systems are increasingly configurable. While this has many benefits, it also makes some software engineering tasks,such as software testing, much harder. This is because, in theory,unique errors could be hiding in any configuration, and, therefore,every configuration may need to undergo expensive testing. As this is generally infeasible, developers need cost-effective technique for selecting which specific configurations they will test. One popular selection approach is combinatorial interaction testing (CIT), where the developer selects a strength t and then computes a covering array (a set of configurations) in which all t-way combinations of configuration option settings appear at least once. In prior work, we demonstrated several limitations of the CIT approach. In particular, we found that a given system's effective configuration space - the minimal set of configurations needed to achieve a specific goal - could comprise only a tiny subset of the system's full configuration space. We also found that effective configuration space may not be well approximated by t-way covering arrays. Based on these insights we have developed an algorithm called interaction tree discovery (iTree). iTree is an iterative learning algorithm that efficiently searches for a small set of configurations that closely approximates a system's effective configuration space. On each iteration iTree tests the system on a small sample of carefully chosen configurations, monitors the system's behaviors, and then applies machine learning techniques to discover which combinations of option settings are potentially responsible for any newly observed behaviors. This information is used in the next iteration to pick a new sample of configurations that are likely to reveal further new behaviors. In prior work, we presented an initial version of iTree and performed an initial evaluation with promising results. This paper presents an improved iTree algorithm in greater detail. The key improvements are based on our use of composite proto-interactions - a construct that improves iTree's ability to correctly learn key configuration option combinations, which in turn significantly improves iTree's running time, without sacrificing effectiveness. Finally, the paper presents a detailed evaluation of the improved iTree algorithm by comparing the coverage it achieves versus that of covering arrays and randomly generated configuration sets, including a significantly expanded scalability evaluation with the ~ 1M-LOC MySQL. Our results strongly suggest that the improved iTree algorithm is highly scalable and can identify a high-coverage test set of configurations more effectively than existing methods. Charles Song, Adam A. Porter, Jeffrey S. Foster |
IEEE Trans. Software Eng. | 2 |
| 2014 | Reducing Masking Effects in CombinatorialInteraction Testing: A Feedback DrivenAdaptive ApproachabstractThe configuration spaces of modern software systems are too large to test exhaustively. Combinatorial interaction testing (CIT) approaches, such as covering arrays, systematically sample the configuration space and test only the selected configurations. The basic justification for CIT approaches is that they can cost-effectively exercise all system behaviors caused by the settings of t or fewer options. We conjecture, however, that in practice some of these behaviors are not actually tested because of unanticipated masking effects - test case failures that perturb system execution so as to prevent some behaviors from being exercised. While prior research has identified this problem, most solutions require knowing the masking effects a priori. In practice this is impractical, if not impossible. In this work, we reduce the harmful consequences of masking effects. First we define a novel interaction testing criterion, which aims to ensure that each test case has a fair chance to test all valid t-way combinations of option settings. We then introduce a feedback driven adaptive combinatorial testing process (FDA-CIT) to materialize this criterion in practice. At each iteration of FDA-CIT, we detect potential masking effects, heuristically isolate their likely causes (i.e., fault characterization), and then generate new samples that allow previously masked combinations to be tested in configurations that avoid the likely failure causes. The iterations end when the new interaction testing criterion has been satisfied. This paper compares two different fault characterization approaches - an integral part of the proposed approach, and empirically assesses their effectiveness and efficiency in removing masking effects on two widely used open source software systems. It also compares FDA-CIT against error locating arrays, a state of the art approach for detecting and locating failures. Furthermore, the scalability of the proposed approach is evaluated by comparing it with perfect test scenarios, in which all masking effects are known a priori. Our results suggest that masking effects do exist in practice, and that our approach provides a promising and efficient way to work around them, without requiring that masking effects be known a priori. Cemal Yilmaz 0001, Emine Dumlu, Myra B. Cohen, Adam A. Porter |
IEEE Trans. Software Eng. | 4 |
| 2013 | Testing component compatibility in evolving configurations
Il-Chul Yoon, Alan Sussman, Atif M. Memon, Adam A. Porter |
Inf. Softw. Technol. | 4 |
| 2012 | iTree: Efficiently discovering high-coverage configurations using interaction treesabstractSoftware configurability has many benefits, but it also makes programs much harder to test, as in the worst case the program must be tested under every possible configuration. One potential remedy to this problem is combinatorial interaction testing (CIT), in which typically the developer selects a strength t and then computes a covering array containing all t-way configuration option combinations. However, in a prior study we showed that several programs have important high-strength interactions (combinations of a subset of configuration options) that CIT is highly unlikely to generate in practice. In this paper, we propose a new algorithm called interaction tree discovery (iTree) that aims to identify sets of configurations to test that are smaller than those generated by CIT, while also including important high-strength interactions missed by practical applications of CIT. On each iteration of iTree, we first use low-strength CIT to test the program under a set of configurations, and then apply machine learning techniques to discover new interactions that are potentially responsible for any new coverage seen. By repeating this process, iTree builds up a set of configurations likely to contain key high-strength interactions. We evaluated iTree by comparing the coverage it achieves versus covering arrays and randomly generated configuration sets. Our results strongly suggest that iTree can identify high-coverage sets of configurations more effectively than traditional CIT or random sampling. Charles Song, Adam A. Porter, Jeffrey S. Foster |
ICSE | 2 |
| 2012 | Overlap and Synergy in Testing Software Components across Loosely Coupled CommunitiesabstractComponent integration rather than from-scratch programming increasingly defines software development. As a result software developers often play diverse roles, including that of a component provider -- packaging a component for others to use, a component user -- integrating other providers' components into their software, and a component tester -- ensuring that other providers' components work as part of an integrated system. In this paper, we explore the conjecture that we can better utilize testing resources by focusing not just on individual component-based systems, but on groups of systems that form what we refer to as loosely-coupled software development communities, meaning a set of independently-managed systems that use many of the same components. We demonstrate that such communities do in fact exist, and that there are overlaps and synergies in their test efforts. Teng Long 0003, Il-Chul Yoon, Adam A. Porter, Alan Sussman, Atif M. Memon |
ISSRE | 3 |
| 2011 | Smartphones in the curriculum workshop (SMACK 2011)abstractSmartphone sales are expected to outpace desktop/laptop computer sales in 2011. It is critical for software engineers to understand the key issues of building applications for this new platform. The mobile sensing and networking capabilities of smartphones create a unique platform for building cyberphysical and other applications that sense and respond to the environment. Moreover, social networking capabilities of these platforms offer new paradigms for dissemination of knowledge, harvesting of user relationship information, and following current events. This workshop will foster new ideas, approaches, and artifacts that can be used to support the introduction of smartphones into traditional software engineering courses (e.g., a survey course, or senior capstone design course, among others). Jules White, Jeffrey G. Gray, Adam A. Porter |
CSEE&T | 3 |
| 2011 | Feedback driven adaptive combinatorial testingabstractThe configuration spaces of modern software systems are too large to test exhaustively. Combinatorial interaction testing (CIT) approaches, such as covering arrays, systematically sample the configuration space and test only the selected configurations. The basic justification for CIT approaches is that they can cost-effectively exercise all system behaviors caused by the settings of t or fewer options. We conjecture, however, that in practice many such behaviors are not actually tested because of masking effects -- failures that perturb execution so as to prevent some behaviors from being exercised. In this work we present a feedback-driven, adaptive, combinatorial testing approach aimed at detecting and working around masking effects. At each iteration we detect potential masking effects, heuristically isolate their likely causes, and then generate new covering arrays that allow previously masked combinations to be tested in the subsequent iteration. We empirically assess the effectiveness of the proposed approach on two large widely used open source software systems. Our results suggest that masking effects do exist and that our approach provides a promising and efficient way to work around them. Emine Dumlu, Cemal Yilmaz 0001, Myra B. Cohen, Adam A. Porter |
ISSTA | 4 |
| 2010 | Using symbolic evaluation to understand behavior in configurable software systemsabstractMany modern software systems are designed to be highly configurable, which increases flexibility but can make programs hard to test, analyze, and understand. We present an initial empirical study of how configuration options affect program behavior. We conjecture that, at certain levels of abstraction, configuration spaces are far smaller than the worst case, in which every configuration is distinct. We evaluated our conjecture by studying three configurable software systems: vsftpd, ngIRCd, and grep. We used symbolic evaluation to discover how the settings of run-time configuration options affect line, basic block, edge, and condition coverage for our subjects under a given test suite. Our results strongly suggest that for these subject programs, test suites, and configuration options, when abstracted in terms of the four coverage criteria above, configuration spaces are in fact much smaller than combinatorics would suggest and are effectively the composition of many small, self-contained groupings of options. Elnatan Reisner, Charles Song, Kin-Keung Ma, Jeffrey S. Foster, Adam A. Porter |
ICSE (1) | 5 |
| 2010 | Combining hardware and software instrumentation to classify program executionsabstractSeveral research efforts have studied ways to infer properties of software systems from program spectra gathered from the running systems, usually with software-level instrumentation. One specific application of this general approach, which is the focus of this paper, is distinguishing failed executions from successful executions. While existing efforts appear to produce accurate classifications, detailed understanding of their costs and potential cost-benefit tradeoffs is lacking. In this work, we present a hybrid instrumentation approach which uses hardware performance counters to gather program spectra at very low cost. This underlying data is further augmented with data captured by minimal amounts of software-level instrumentation. We also evaluate this hybrid approach by comparing it to other existing approaches. We conclude that these hybrid spectra can reliably distinguish failed executions from successful executions at a fraction of the runtime overhead cost of using software-only spectra. Cemal Yilmaz 0001, Adam A. Porter |
SIGSOFT FSE | 2 |
| 2009 | Prioritizing component compatibility tests via user preferencesabstractMany software systems rely on third-party components during their build process. Because the components are constantly evolving, quality assurance demands that developers perform compatibility testing to ensure that their software systems build correctly over all deployable combinations of component versions, also called configurations. However, large software systems can have many configurations, and compatibility testing is often time and resource constrained. We present a prioritization mechanism that enhances compatibility testing by examining the ldquomost importantrdquo configurations first, while distributing the work over a cluster of computers. We evaluate our new approach on two large scientific middleware systems and examine tradeoffs between the new prioritization approach and a previously developed lowest-cost-configuration-first approach. Il-Chul Yoon, Alan Sussman, Atif M. Memon, Adam A. Porter |
ICSM | 4 |
| 2009 | Incremental covering array failure characterization in large configuration spacesabstractThe increasing complexity of configurable software systems has created a need for more intelligent sampling mechanisms to detect and characterize failure-inducing dependencies between configurations. Prior work - in idealized environments - has shown that test schedules based on a mathematical object, called a covering array, in combination with classification techniques, can meet this need. Applying this approach in practice, however, is tricky because testing time and resource availability are unpredictable, and because failure characteristics can change from release to release. With current approaches developers must set a key covering array parameter (its strength) based on estimated release times and failure characterizations. This will influence the outcome of their results. Sandro Fouché, Myra B. Cohen, Adam A. Porter |
ISSTA | 3 |
| 2008 | Effective and scalable software compatibility testingabstractToday's software systems are typically composed of multiple components, each with different versions. Software compatibility testing is a quality assurance task aimed at ensuring that multi-component based systems build and/or execute correctly across all their versions' combinations, or configurations. Because there are complex and changing interdependencies between components and their versions, and because there are such a large number of configurations, it is generally infeasible to test all potential configurations. Consequently, in practice, compatibility testing examines only a handful of default or popular configurations to detect problems; as a result costly errors can and do escape to the field. Il-Chul Yoon, Alan Sussman, Atif M. Memon, Adam A. Porter |
ISSTA | 4 |
| 2007 | Towards a Distributed Continuous Certification ProcessabstractSoftware scale and complexity are growing by every measure: more hardware and software, more communication links, more interdependency, more lines of code, more storage and data, etc. At the same time business trends are increasingly squeezing development resources. In particular, development processes are straining under severe cost and time-to-market pressures. Global competition and market deregulation are shrinking profit margins and thus limiting budgets for the development and QA of software. In response to these trends, developers have begun to change the way they build and validate software systems by (among other things) moving towards more flexible product designs allowing dynamic reconfiguration. This approach promises to improve cost, quality, and development-time, but creates other problems, especially when used in the context of safety-critical systems. To realize this promise, however, effective certification becomes more important than ever since as static controls are removed or reduced, it becomes even more vital that (1) problems be caught as quickly as possible and (2) systems not be allowed to drift so far from their intended functional and performance requirements that rework costs overwhelm the hoped-for efficiencies. This article can present and discuss some of our recent efforts to address these problems. Adam A. Porter |
IPDPS | 1 |
| 2007 | Direct-dependency-based software compatibility testingabstractSoftware compatibility testing is an important quality assurance task aimed at ensuring that component-based software systems build and/or execute properly across a broad range of user system configurations. Because each configuration can involve multiple components with different versions, and because there are complex and changing interdependencies between components and their versions, it is generally infeasible to test all potential configurations. Therefore, compatibility testing usually means examining only a handful of default or popular configurations to detect problems, and as a result costly errors can and do escape to the field Il-Chul Yoon, Alan Sussman, Atif M. Memon, Adam A. Porter |
ASE | 4 |
| 2007 | Towards incremental adaptive covering arraysabstractThe increasing complexity of configurable software systems creates a need for more intelligent sampling mechanisms to detect and locate failure-inducing dependencies between configurations. Prior work shows that test schedules basedon a mathematical object, called a covering array, can beused to detect and locate failures in combination with a classification tree analysis. This paper addresses limitations ofthe earlier approach. First, the previous work requires developers to choose the covering array's strength, even thought here is no scientific or historical basis for doing so. Second, if a single covering array is insufficient to classify specificfailures, the entire process must be rerun from scratch. To address these issues, our new approach incrementally and adaptively builds covering array schedules. It begins witha low strength, and continually increases this as resources allow, or poor classification results require. At each stage, previous tests are reused. This allows failures due to only one or two configurations settings to be found and classified as early as possible, and also limits duplication of work when multiple covering arrays must be used. Sandro Fouché, Myra B. Cohen, Adam A. Porter |
ESEC/SIGSOFT FSE | 3 |
| 2007 | Techniques for Classifying Executions of Deployed Software to Support Software Engineering TasksabstractThere is an increasing interest in techniques that support analysis and measurement of fielded software systems. These techniques typically deploy numerous instrumented instances of a software system, collect execution data when the instances run in the field, and analyze the remotely collected data to better understand the system's in-the-field behavior. One common need for these techniques is the ability to distinguish execution outcomes (e.g., to collect only data corresponding to some behavior or to determine how often and under which condition a specific behavior occurs). Most current approaches, however, do not perform any kind of classification of remote executions and either focus on easily observable behaviors (e.g., crashes) or assume that outcomes' classifications are externally provided (e.g., by the users). To address the limitations of existing approaches, we have developed three techniques for automatically classifying execution data as belonging to one of several classes. In this paper, we introduce our techniques and apply them to the binary classification of passing and failing behaviors. Our three techniques impose different overheads on program instances and, thus, each is appropriate for different application scenarios. We performed several empirical studies to evaluate and refine our techniques and to investigate the trade-offs among them. Our results show that 1) the first technique can build very accurate models, but requires a complete set of execution data; 2) the second technique produces slightly less accurate models, but needs only a small fraction of the total execution data; and 3) the third technique allows for even further cost reductions by building the models incrementally, but requires some sequential ordering of the software instances' instrumentation. Murali Haran, Alan F. Karr, Michael Last, Alessandro Orso, Adam A. Porter, Ashish P. Sanil, Sandro Fouché |
IEEE Trans. Software Eng. | 5 |
| 2007 | Skoll: A Process and Infrastructure for Distributed Continuous Quality AssuranceabstractSoftware engineers increasingly emphasize agility and flexibility in their designs and development approaches. They increasingly use distributed development teams, rely on component assembly and deployment rather than green field code writing, rapidly evolve the system through incremental development and frequent updating, and use flexible product designs supporting extensive end-user customization. While agility and flexibility have many benefits, they also create an enormous number of potential system configurations built from rapidly changing component implementations. Since today's quality assurance (QA) techniques do not scale to handle highly configurable systems, we are developing and validating novel software QA processes and tools that leverage the extensive computing resources of user and developer communities in a distributed, continuous manner to improve software quality significantly. This paper provides several contributions to the study of distributed, continuous QA (DCQA). First, it shows the structure and functionality of Skoll, which is an environment that defines a generic around-the-world, around-the-clock QA process and several sophisticated tools that support this process. Second, it describes several novel QA processes built using the Skoll environment. Third, it presents two studies using Skoll: one involving user testing of the Mozilla browser and another involving continuous build, integration, and testing of the ACE+TAO communication software package. The results of our studies suggest that the Skoll environment can manage and control distributed continuous QA processes more effectively than conventional QA processes. For example, our DCQA processes rapidly identified problems that had taken the ACE+TAO developers much longer to find and several of which they had not found. Moreover, the automatic analysis of QA results provided developers information that enabled them to quickly find the root causes of problems Adam A. Porter, Cemal Yilmaz 0001, Atif M. Memon, Douglas C. Schmidt, Balachandran Natarajan |
IEEE Trans. Software Eng. | 1 |
| 2007 | Reliable Effects Screening: A Distributed Continuous Quality Assurance Process for Monitoring Performance Degradation in Evolving Software SystemsabstractDevelopers of highly configurable performance-intensive software systems often use in-house performance-oriented "regression testing" to ensure that their modifications do not adversely affect their software's performance across its large configuration space. Unfortunately, time and resource constraints can limit in-house testing to a relatively small number of possible configurations, followed by unreliable extrapolation from these results to the entire configuration space. As a result, many performance bottlenecks escape detection until systems are fielded. In our earlier work, we improved the situation outlined above by developing an initial quality assurance process called "main effects screening". This process 1) executes formally designed experiments to identify an appropriate subset of configurations on which to base the performance-oriented regression testing, 2) executes benchmarks on this subset whenever the software changes, and 3) provides tool support for executing these actions on in-the-field and in-house computing resources. Our initial process had several limitations, however, since it was manually configured (which was tedious and error-prone) and relied on strong and untested assumptions for its accuracy (which made its use unacceptably risky in practice). This paper presents a new quality assurance process called "reliable effects screening" that provides three significant improvements to our earlier work. First, it allows developers to economically verify key assumptions during process execution. Second, it integrates several model-driven engineering tools to make process configuration and execution much easier and less error prone. Third, we evaluate this process via several feasibility studies of three large, widely used performance-intensive software frameworks. Our results indicate that reliable effects screening can detect performance degradation in large-scale systems more reliably and with significantly less resources than conventional techniques Cemal Yilmaz 0001, Adam A. Porter, Arvind S. Krishna, Atif M. Memon, Douglas C. Schmidt, Aniruddha S. Gokhale, Balachandran Natarajan |
IEEE Trans. Software Eng. | 2 |
| 2006 | Covering Arrays for Efficient Fault Characterization in Complex Configuration SpacesabstractMany modern software systems are designed to be highly configurable so they can run on and be optimized for a wide variety of platforms and usage scenarios. Testing such systems is difficult because, in effect, you are testing a multitude of systems, not just one. Moreover, bugs can and do appear in some configurations, but not in others. Our research focuses on a subset of these bugs that are "option-related"-those that manifest with high probability only when specific configuration options take on specific settings. Our goal is not only to detect these bugs, but also to automatically characterize the configuration subspaces (i.e., the options and their settings) in which they manifest. To improve efficiency, our process tests only a sample of the configuration space, which we obtain from mathematical objects called covering arrays. This paper compares two different kinds of covering arrays for this purpose and assesses the effect of sampling strategy on fault characterization accuracy. Our results strongly suggest that sampling via covering arrays allows us to characterize option-related failures nearly as well as if we had tested exhaustively, but at a much lower cost. We also provide guidelines for using our approach in practice. Cemal Yilmaz 0001, Myra B. Cohen, Adam A. Porter |
IEEE Trans. Software Eng. | 3 |
| 2005 | Main effects screening: a distributed continuous quality assurance process for monitoring performance degradation in evolving software systemsabstractDevelopers of highly configurable performance-intensive software systems often use a type of in-house performance-oriented "regression testing" to ensure that their modifications have not adversely affected their software's performance across its large configuration space. Unfortunately, time and resource constraints often limit developers to in-house testing of a small number of configurations and unreliable extrapolation from these results to the entire configuration space, which allows many performance bottlenecks and sources of QoS degradation to escape detection until systems are fielded. To improve performance assessment of evolving systems across large configuration spaces, we have developed a distributed continuous quality assurance (DCQA) process called main effects screening that uses in-the-field resources to execute formally designed experiments to help reduce the configuration space, thereby allowing developers to perform more targeted in-house QA. We have evaluated this process via several feasibility studies on several large, widely-used performance-intensive software systems. Our results indicate that main effects screening can detect key sources of performance degradation in large-scale systems with significantly less effort than conventional techniques. Cemal Yilmaz 0001, Arvind S. Krishna, Atif M. Memon, Adam A. Porter, Douglas C. Schmidt, Aniruddha S. Gokhale, Balachandran Natarajan |
ICSE | 4 |
| 2005 | Applying classification techniques to remotely-collected program execution dataabstractThere is an increasing interest in techniques that support measurement and analysis of fielded software systems. One of the main goals of these techniques is to better understand how software actually behaves in the field. In particular, many of these techniques require a way to distinguish, in the field, failing from passing executions. So far, researchers and practitioners have only partially addressed this problem: they have simply assumed that program failure status is either obvious (i.e., the program crashes) or provided by an external source (e.g., the users). In this paper, we propose a technique for automatically classifying execution data, collected in the field, as coming from either passing or failing program runs. (Failing program runs are executions that terminate with a failure, such as a wrong outcome.) We use statistical learning algorithms to build the classification models. Our approach builds the models by analyzing executions performed in a controlled environment (e.g., test cases run in-house) and then uses the models to predict whether execution data produced by a fielded instance were generated by a passing or failing program execution. We also present results from an initial feasibility study, based on multiple versions of a software subject, in which we investigate several issues vital to the applicability of the technique. Finally, we present some lessons learned regarding the interplay between the reliability of classification models and the amount and type of data collected. Murali Haran, Alan F. Karr, Alessandro Orso, Adam A. Porter, Ashish P. Sanil |
ESEC/SIGSOFT FSE | 4 |
| 2005 | An empirical study of regression test application frequencyabstractRegression testing is an expensive process used to revalidate modified software. Regression test selection (RTS) techniques reduce the cost of regression testing by selecting a subset of a test suite. Many RTS techniques have been proposed, and studies have shown that they produce savings; other studies have shown that their cost-effectiveness varies with characteristics of the workloads to which they are applied. It seems plausible, however, that another factor that impacts RTS techniques involves the process by which they are applied. In particular, issues such as the frequency with which regression testing is performed affect the techniques. Thus, in earlier work an experiment was conducted to assess the effects of test application frequency on the cost-effectiveness of RTS techniques. The results exposed tradeoffs to consider when using these techniques over a series of releases. This work, however, was limited in external validity; in particular, the programs studied were relatively small. Thus, the previous experiment has been replicated on a large, multi-version program. This second experiment confirms the findings of the first. In particular, the cost of using safe RTS techniques was strongly and negatively affected by testing interval. Conversely, the effectiveness of minimization RTS techniques was strongly and positively affected. Copyright © 2005 John Wiley & Sons, Ltd. Jung-Min Kim, Adam A. Porter, Gregg Rothermel |
Softw. Test. Verification Reliab. | 2 |
| 2004 | Skoll: Distributed Continuous Quality AssuranceabstractQuality assurance (QA) tasks, such as testing, profiling, and performance evaluation, have historically been done in-house on developer-generated workloads and regression suites. Since this approach is inadequate for many systems, tools and processes are being developed to improve software quality by increasing user participation in the QA process. A limitation of these approaches is that they focus on isolated mechanisms, not on the coordination and control policies and tools needed to make the global QA process efficient, effective, and scalable. To address these issues, we have initiated the Skoll project, which is developing and validating novel software QA processes and tools that leverage the extensive computing resources of worldwide user communities in a distributed, continuous manner to significantly and rapidly improve software quality. This paper provides several contributions to the study of distributed continuous QA. First, it illustrates the structure and functionality of a generic around-the-world, around-the-clock QA process and describes several sophisticated tools that support this process. Second, it describes several QA scenarios built using these tools and process. Finally, it presents a feasibility study applying these scenarios to a 1MLOC+ software package called ACE+TAO. While much work remains to be done, the study suggests that the Skoll process and tools effectively manage and control distributed, continuous QA processes. Using Skoll we rapidly identified problems that had taken the ACE+TAO developers substantially longer to find and several of which had previously not been found. Moreover, automatic analysis of QA task results often provided developers information that quickly led them to the root cause of the problems. Atif M. Memon, Adam A. Porter, Cemal Yilmaz 0001, Adithya Nagarajan, Douglas C. Schmidt, Balachandran Natarajan |
ICSE | 2 |
| 2004 | Second ICSE Workshop on Remote Analysis and Measurement of Software Systems (RAMSS)
Alessandro Orso, Adam A. Porter |
ICSE | 2 |
| 2004 | Validating Quality of Service for Reusable Software Via Model-Integrated Distributed Continuous Quality Assurance
Arvind S. Krishna, Douglas C. Schmidt, Atif M. Memon, Adam A. Porter, Diego Sevilla Ruiz |
ICSR | 4 |
| 2004 | Covering arrays for efficient fault characterization in complex configuration spacesabstractTesting systems with large configurations spaces that change often is a challenging problem. The cost and complexity of QA explodes because often there isn't just one system, but a multitude of related systems. Bugs may appear in certain configurations, but not in others.The Skoll system and process has been developed to test these types of systems through distributed, continuous quality assurance, leveraging user resources around-the-world, around-the-clock. It has been shown to be effective in automatically characterizing configurations in which failures manifest. The derived information helps developers quickly narrow down the cause of failures which then improves turn around time for fixes. However, this method does not scale well. It requires one to exhaustively test each configuration in the configuration space.In this paper we examine an alternative approach. The idea is to systematically sample the configuration space, test only the selected configurations, and conduct fault characterization on the resulting data. The sampling approach we use is based on calculating a mathematical object called a covering array. We empirically assess the effect of using covering array derived test schedules on the resulting fault characterizations and provide guidelines to practitioners for their use. Cemal Yilmaz 0001, Myra B. Cohen, Adam A. Porter |
ISSTA | 3 |
| 2003 | ICSE Workshop on Remote Analysis and Measurement of Software Systems (RAMSS)
Alessandro Orso, Adam A. Porter |
ICSE | 2 |
| 2002 | A history-based test prioritization technique for regression testing in resource constrained environmentsabstractRegression testing is an expensive and frequently executed maintenance process used to revalidate modified software. To improve it, regression test selection (RTS) techniques strive to lower costs without overly reducing effectiveness by carefully selecting a subset of the test suite. Under certain conditions, some can even guarantee that the selected test cases perform no worse than the original test suite.But this ignores certain software development realities such as resource and time constraints that may prevent using RTS techniques as intended (e.g., regression testing must be done overnight, but RTS selection returns two days worth of tests). In practice, testers work around this by prioritizing the test cases and running only those that fit within existing constraints. Unfortunately this generally violates key RTS assumptions, voiding RTS technique guarantees and making regression testing performance unpredictable.Despite this, existing prioritization techniques are memoryless, implicitly assuming that local choices can ensure adequate long run performance. Instead, we proposed a new technique that bases prioritization on historical execution data. We conducted an experiment to assess its effects on the long run performance of resource constrained regression testing. Our results expose essential tradeoffs that should be considered when using these techniques over a series of software releases. Jung-Min Kim, Adam A. Porter |
ICSE | 2 |
| 2002 | Reducing Inspection Interval in Large-Scale Software DevelopmentabstractWe have found that, when software is developed by multiple, geographically separated teams, the cost-benefit trade-offs of software inspection change. In particular, this situation can significantly lengthen the inspection interval (calendar time needed to complete an inspection). Our research goal was to find a way to reduce the inspection interval without reducing inspection effectiveness. We believed that Internet technology offered some potential solutions, but we were not sure which technology to use nor what effects it would have on effectiveness. To conduct this research, we drew on the results of several empirical studies we had previously performed. These results clarified the role that meetings and individuals play in inspection effectiveness and interval. We conducted further studies showing that manual inspections without meetings were just as effective as manual inspections with them. On the basis of these and other findings and our understanding of Internet technology, we built an economical and effective tool that reduced the interval without reducing effectiveness. This tool, Hypercode, supports meetingless software inspections with geographically distributed reviewers. HyperCode is a platform-independent tool, developed on top of an Internet browser, that integrates seamlessly into the current development process. By seamless, we mean the tool produces a paper flow that is almost identical to the current inspection process. HyperCode's acceptance by its user community has been excellent. Moreover, we estimate that using HyperCode has reduced the inspection interval by 20 to 25 percent. We believe that, had we focused solely on technology (without considering the information our studies had uncovered), we would have created a more complex, but not necessarily more effective tool. We probably would have supported group meetings, restricted each participant's access to review comments, and supported a wider variety of inspection methods. In other words, the principles derived from our empirical studies dramatically and successfully directed our search for a technological solution. Dewayne E. Perry, Adam A. Porter, Michael W. Wade, Lawrence G. Votta, James Perpich |
IEEE Trans. Software Eng. | 2 |
| 2001 | An empirical study of regression test selection techiquesabstractRegression testing is the process of validating modified software to detect whether new errors have been introduced into previously tested code and to provide confidence that modifications are correct. Since regression testing is an expensive process, researchers have proposed regression test selection techniques as a way to reduce some of this expense. These techniques attempt to reduce costs by selecting and running only a subset of the test cases in a program's existing test suite. Although there have been some analytical and empirical evaluations of individual techniques, to our knowledge only one comparative study, focusing on one aspect of two of these techniques, has been reported in the literature. We conducted an experiment to examine the relative costs and benefits of several regression test selection techniques. The experiment examined five techniques for reusing test cases, focusing on their relative ablilities to reduce regression testing effort and uncover faults in modified programs. Our results highlight several differences between the techiques, and expose essential trade-offs that should be considered when choosing a technique for practical application. Todd L. Graves, Mary Jean Harrold, Jung-Min Kim, Adam A. Porter, Gregg Rothermel |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2000 | An empirical study of regression test application frequencyabstractRegression testing is an expensive maintenance process used to revalidate modified software. Regression test selection (RTS) techniques try to lower the cost of regression testing by selecting and running a subset of the existing test cases. Many such techniques have been proposed and initial studies show that they can produce savings. We believe, however, that issues such as the frequency with which testing is done have a strong effect on the behavior of these techniques. Therefore, we conducted an experiment to assess the effects of test application frequency on the costs and benefits of regression test selection techniques. Our results expose essential tradeoffs that should be considered when using these techniques over a series of software releases. Jung-Min Kim, Adam A. Porter, Gregg Rothermel |
ICSE | 2 |
| 1998 | An Empirical Study of Regression Test Selection TechniquesabstractRegression testing is an expensive maintenance process directed at validating modified software. Regression test selection techniques attempt to reduce the cost of regression testing by selecting tests from a program's existing test suite. Many regression test selection techniques have been proposed. Although there have been some analytical and empirical evaluations of individual techniques, to our knowledge only one comparative study, focusing on one aspect of two of these techniques, has been performed. We conducted an experiment to examine the relative costs and benefits of several regression test selection techniques. The experiment examined five techniques for reusing tests, focusing on their relative abilities to reduce regression testing effort and uncover faults in modified programs. Our results highlight several differences between the techniques, and expose essential tradeoffs that should be considered when choosing a technique for practical application. Todd L. Graves, Mary Jean Harrold, Jung-Min Kim, Adam A. Porter, Gregg Rothermel |
ICSE | 4 |
| 1998 | Comparing Detection Methods For Software Requirements Inspections: A Replication Using Professional Subjects
Adam A. Porter, Lawrence G. Votta |
Empir. Softw. Eng. | 1 |
| 1998 | Specification-based testing of reactive software: A case study in technology transfer
Lalita Jategaonkar Jagadeesan, Lawrence G. Votta, Adam A. Porter, Carlos Puchol, J. Christopher Ramming |
J. Syst. Softw. | 3 |
| 1998 | Understanding the Sources of Variation in Software InspectionsabstractIn a previous experiment, we determined how various changes in three structural elements of the software inspection process (team size and the number and sequencing of sessions) altered effectiveness and interval. Our results showed that such changes did not significantly influence the defect detection rate, but that certain combinations of changes dramatically increased the inspection interval. We also observed a large amount of unexplained variance in the data, indicating that other factors must be affecting inspection performance. The nature and extent of these other factors now have to be determined to ensure that they had not biased our earlier results. Also, identifying these other factors might suggest additional ways to improve the efficiency of inspections. Acting on the hypothesis that the “inputs” into the inspection process (reviewers, authors, and code units) were significant sources of variation, we modeled their effects on inspection performance. We found that they were responsible for much more variation in detect detection than was process structure. This leads us to conclude that better defect detection techniques, not better process structures, are the key to improving inspection effectiveness. The combined effects of process inputs and process structure on the inspection interval accounted for only a small percentage of the variance in inspection interval. Therefore, there must be other factors which need to be identified. Adam A. Porter, Harvey P. Siy, Audris Mockus, Lawrence G. Votta |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 1997 | Specification-based Testing of Reactive Software: Tools and Experiments (Experience Report)abstractTesting commercial software is expensive and time consuming.Automated testing methods promise to save a great deal of time and money throughout the software industry.One approach that is well-suited for the reactive systems found in telephone switching systems is specification-based testing.We have built a set of tools to automatically test software applications for violations of safety properties expressed in temporal logic.Our testing system automatically constructs finite state machine oracles corresponding to safety properties, builds test harnesses, and integrates them with the application.The test harness then generates inputs automatically to test the application.We describe a study examining the feasibility of this approach for testing industrial applications.To conduct this study we formally modeled an Automatic Protection Switching system (APS), which is an application common to many telephony systems.We then asked a number of computer science graduate students to develop several versions of the APS and use our tools to test them.We found that the tools are very effective, save significant amounts of human effort (at the expense of machine resources), and are easy to use.We also dis- Lalita Jategaonkar Jagadeesan, Adam A. Porter, Carlos Puchol, J. Christopher Ramming, Lawrence G. Votta |
ICSE | 2 |
| 1997 | Anywhere, Anytime Code Inspections: Using the Web to Remove Inspection Bottlenecks in Large-Scale Software Developmentabstractand the time savings from the reduction in distribution interval of the inspection package (sometimes involving international mailings) have been substantial.These savings together with the seamless integration into the existing environment are the major factors for this acceptance.From our viewpoint as experimentalists, the acceptance came too readily.Therefore we lost our opportunity to explore this tool using a series of controlled experiments to isolate the underlying factors or its effectiveness.Nevertheless, by using historical data we can show that the new process is less expensive in terms of cost and at least as effective in terms of quality (defect detection effectiveness). James Perpich, Dewayne E. Perry, Adam A. Porter, Lawrence G. Votta, Michael W. Wade |
ICSE | 3 |
| 1997 | A Primer on Empirical Studies (Tutorial)abstractNo abstract available. Dewayne E. Perry, Adam A. Porter, Lawrence G. Votta |
ICSE | 2 |
| 1997 | Understanding the Effects of Developer Activities on Inspection IntervalabstractWe have conducted an industrial experiment to assess the cost-benefit tradeoffs of several software inspection processes. Our results to date explain the variation in observed effectiveness very well, but are unable to satisfactorily explain variation in inspection interval. In this article we examine the effect of a new factor - process environment - on inspection interval (calendar time needed to complete the inspection). Our analysis suggests that process environment does indeed influence inspection interval. in particular, we found that non-uniform work priorities, time-varying workloads, and deadlines have significant effects. Moreover, these experiences suggest that regression models are inherently inadequate for interval modeling, and that queueing models may be more effective. (Also cross-referenced as UMIACS-TR-97-19) Adam A. Porter, Harvey P. Siy, Lawrence G. Votta |
ICSE | 1 |
| 1997 | Fundamental Laws and Assumptions of Software Maintenance
Adam A. Porter |
Empir. Softw. Eng. | 1 |
| 1997 | Assessing Software Review Meetings: Results of a Comparative Analysis of Two Experimental StudiesabstractSoftware review is a fundamental tool for software quality assurance. Nevertheless, there are significant controversies as to the most efficient and effective review method. One of the most important questions currently being debated is the utility of meetings. Although almost all industrial review methods are centered around the inspection meeting, recent findings call their value into question. In prior research the authors separately and independently conducted controlled experimental studies to explore this issue. The paper presents new research to understand the broader implications of these two studies. To do this, they designed and carried out a process of "reconciliation" in which they established a common framework for the comparison of the two experimental studies, reanalyzed the experimental data with respect to this common framework, and compared the results. Through this process they found many striking similarities between the results of the two studies, strengthening their individual conclusions. It also revealed interesting differences between the two experiments, suggesting important avenues for future research. Adam A. Porter, Philip M. Johnson |
IEEE Trans. Software Eng. | 1 |
| 1997 | An Experiment ot Assess the Cost-Benefits of Code Inspections in Large Scale Software DevelopmentabstractWe conducted a long term experiment to compare the costs and benefits of several different software inspection methods. These methods were applied by professional developers to a commercial software product they were creating. Because the laboratory for this experiment was a live development effort, we took special care to minimize cost and risk to the project, while maximizing our ability to gather useful data. The article has several goals: (1) to describe the experiment's design and show how we used simulation techniques to optimize it; (2) to present our results and discuss their implications for both software practitioners and researchers; and (3) to discuss several new questions raised by our findings. For each inspection, we randomly assigned three independent variables: (1) the number of reviewers on each inspection team (1, 2, or 4); (2) the number of teams inspecting the code unit (1 or 2); and (3) the requirement that defects be repaired between the first and second team's inspections. The reviewers for each inspection were randomly selected without replacement from a pool of 11 experienced software developers. The dependent variables for each inspection included inspection interval (elapsed time), total effort, and the defect detection rate. Our results showed that these treatments did not significantly influence the defect detection effectiveness, but that certain combinations of changes dramatically increased the inspection interval. Adam A. Porter, Harvey P. Siy, Carol A. Toman, Lawrence G. Votta |
IEEE Trans. Software Eng. | 1 |
| 1995 | Experimental Software Engineering: A Report on the State of the ArtabstractArticle Free Access Share on Experimental software engineering: a report on the state of the art Authors: Lawrence G. Votta Software & Systems Research Lab, AT&T Bell Laboratories, Naperville, Illinois Software & Systems Research Lab, AT&T Bell Laboratories, Naperville, IllinoisView Profile , Adam Porter University of Maryland, Computer Science Department, College Park, Maryland University of Maryland, Computer Science Department, College Park, MarylandView Profile Authors Info & Claims ICSE '95: Proceedings of the 17th international conference on Software engineeringApril 1995Pages 277–279https://doi.org/10.1145/225014.225040Published:23 April 1995Publication History 29citation568DownloadsMetricsTotal Citations29Total Downloads568Last 12 Months90Last 6 weeks3 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF Lawrence G. Votta, Adam A. Porter |
ICSE | 2 |
| 1995 | An Experiment to Assess the Cost-Benefits of Code Inspections in Large Scale Software DevelopmentabstractWe are conducting a long-term experiment (in progress) to compare the costs and benefits of several different software inspection methods.These methods are being applied by beliefs about the most cost-effective ways to conduct inspections and raise some questions about the feasibility of recently proposed methods.professional developers to a commercial software product Adam A. Porter, Harvey P. Siy, Carol A. Toman, Lawrence G. Votta |
SIGSOFT FSE | 1 |
| 1995 | Comparing Detection Methods for Software Requirements Inspections: A Replicated ExperimentabstractSoftware requirements specifications (SRS) are often validated manually. One such process is inspection, in which several reviewers independently analyze all or part of the specification and search for faults. These faults are then collected at a meeting of the reviewers and author(s). Usually, reviewers use Ad Hoc or Checklist methods to uncover faults. These methods force all reviewers to rely on nonsystematic techniques to search for a wide variety of faults. We hypothesize that a Scenario-based method, in which each reviewer uses different, systematic techniques to search for different, specific classes of faults, will have a significantly higher success rate. We evaluated this hypothesis using a 3/spl times/2/sup 4/ partial factorial, randomized experimental design. Forty eight graduate students in computer science participated in the experiment. They were assembled into sixteen, three-person teams. Each team inspected two SRS using some combination of Ad Hoc, Checklist or Scenario methods. For each inspection we performed four measurements: (1) individual fault detection rate, (2) team fault detection rate, (3) percentage of faults first identified at the collection meeting (meeting gain rate), and (4) percentage of faults first identified by an individual, but never reported at the collection meeting (meeting loss rate). The experimental results are that (1) the Scenario method had a higher fault detection rate than either Ad Hoc or Checklist methods, (2) Scenario reviewers were more effective at detecting the faults their scenarios are designed to uncover, and were no less effective at detecting other faults than both Ad Hoc or Checklist reviewers, (3) Checklist reviewers were no more effective than Ad Hoc reviewers, and (4) Collection meetings produced no net improvement in the fault detection rate-meeting gains were offset by meeting losses.> Adam A. Porter, Lawrence G. Votta, Victor R. Basili |
IEEE Trans. Software Eng. | 1 |
| 1994 | An Experiment to Assess Different Defect Detection Methods for Software Requirements Inspections
Adam A. Porter, Lawrence G. Votta |
ICSE | 1 |
| 1993 | Using measurement-driven modeling to provide empirical feedback to software developers
Adam A. Porter |
J. Syst. Softw. | 1 |
| 1992 | An improved classification tree analysis of high cost modules based upon an axiomatic definition of complexityabstractIdentification of high cost modules has been viewed as one mechanism to improve overall system reliability, since such modules tend to produce more than their fair share of problems. A decision tree model has previously been used to identify such modules. In this paper, a previously developed axiomatic model of program complexity is merged with the previously developed decision tree process for an improvement in the ability to identify such modules. This improvement has been tested using data from the NASA Software Engineering Laboratory.> Jeff Tian, Adam A. Porter, Marvin V. Zelkowitz |
ISSRE | 2 |
| 1991 | Metric-Driven Analysis and Feedback Systems for Enabling Empirically Guided Software Development
Richard W. Selby, Adam A. Porter, Douglas C. Schmidt, Jim Berney |
ICSE | 2 |
| 1990 | Evaluating techniques for generating metric-based classification trees
Adam A. Porter, Richard W. Selby |
J. Syst. Softw. | 1 |
| 1989 | Software metric classification trees help guide the maintenance of large-scale systemsabstractThe 80:20 rule states that approximately 20% of a software system is responsible for 80% of its errors. The authors propose an automated method for generating empirically-based models of error-prone software objects. These models are intended to help localize the troublesome 20%. The method uses a recursive algorithm to automatically generate classification trees whose nodes are multivalued functions based on software metrics. The purpose of the classification trees is to identify components that are likely to be error prone or costly, so that developers can focus their resources accordingly. A feasibility study was conducted using 16 NASA projects. On average, the classification trees correctly identified 79.3% of the software modules that had high development effort or faults.> Richard W. Selby, Adam A. Porter |
ICSM | 2 |
| 1988 | Learning from Examples: Generation and Evaluation of Decision Trees for Software Resource AnalysisabstractA general solution method for the automatic generation of decision (or classification) trees is investigated. The approach is to provide insights through in-depth empirical characterization and evaluation of decision trees for one problem domain, specifically, that of software resource data analysis. The purpose of the decision trees is to identify classes of objects (software modules) that had high development effort, i.e. in the uppermost quartile relative to past data. Sixteen software systems ranging from 3000 to 112000 source lines have been selected for analysis from a NASA production environment. The collection and analysis of 74 attributes (or metrics), for over 4700 objects, capture a multitude of information about the objects: development effort, faults, changes, design style, and implementation style. A total of 9600 decision trees are automatically generated and evaluated. The analysis focuses on the characterization and evaluation of decision tree accuracy, complexity, and composition. The decision trees correctly identified 79.3% of the software modules that had high development effort or faults, on the average across all 9600 trees. The decision trees generated from the best parameter combinations correctly identified 88.4% of the modules on the average. Visualization of the results is emphasized, and sample decision trees are included.> Richard W. Selby, Adam A. Porter |
IEEE Trans. Software Eng. | 2 |