VLDB 2026 Research / reviewers in the wild / expert
Brendan Murphy
dblp:24/4066
· DBLP profile ↗
39ranked-venue papers
6as first author
6since 2021 · last 2025
0000-0001-9327-389XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 24 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 3Security and privacy · 3Theory of computation · 3 · 3 first-authorDatabases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scaling Trends for Data Poisoning in LLMsabstractLLMs produce harmful and undesirable behavior when trained on datasets containing even a small fraction of poisoned data. We demonstrate that GPT models remain vulnerable to fine-tuning on poisoned data, even when safeguarded by moderation systems. Given the persistence of data poisoning vulnerabilities in today's most capable models, this paper investigates whether these risks increase with model scaling. We evaluate three threat models—malicious fine-tuning, imperfect data curation, and intentional data contamination—across 24 frontier LLMs ranging from 1.5 to 72 billion parameters. Our experiments reveal that larger LLMs are significantly more susceptible to data poisoning, learning harmful behaviors from even minimal exposure to harmful data more quickly than smaller models. These findings underscore the need for leading AI companies to thoroughly red team fine-tuning APIs before public release and to develop more robust safeguards against data poisoning, particularly as models continue to scale in size and capability. Dillon Bowen, Brendan Murphy, Will Cai, David Khachaturov, Adam Gleave, Kellin Pelrine |
AAAI | 2 |
| 2025 | Jailbreak-Tuning: Models Efficiently Learn Jailbreak SusceptibilityabstractBrendan Murphy, Dillon Bowen, Shahrad Mohammadzadeh, Tom Tseng, Julius Broomfield, Adam Gleave, Kellin Pelrine. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Brendan Murphy, Dillon Bowen, Shahrad Mohammadzadeh, Tom Tseng, Julius Broomfield, Adam Gleave, Kellin Pelrine |
EMNLP | 1 |
| 2025 | On Targeted Manipulation and Deception when Optimizing LLMs for User FeedbackabstractAs LLMs become more widely deployed, there is increasing interest in directly optimizing for feedback from end users (e.g. thumbs up) in addition to feedback from paid annotators. However, training to maximize human feedback creates a perverse incentive structure for the AI to resort to manipulative or deceptive tactics to obtain positive feedback from users who are vulnerable to such strategies. We study this phenomenon by training LLMs with Reinforcement Learning with simulated user feedback in environments of practical LLM usage. In our settings, we find that: 1) Extreme forms of "feedback gaming" such as manipulation and deception are learned reliably; 2) Even if only 2% of users are vulnerable to manipulative strategies, LLMs learn to identify and target them while behaving appropriately with other users, making such behaviors harder to detect; 3) To mitigate this issue, it may seem promising to leverage continued safety training or LLM-as-judges during training to filter problematic outputs. Instead, we found that while such approaches help in some of our settings, they backfire in others, sometimes even leading to subtler manipulative behaviors. We hope our results can serve as a case study which highlights the risks of using gameable feedback sources -- such as user feedback -- as a target for RL. Our code is publicly available. Warning: some of our examples may be upsetting. Marcus Williams, Micah Carroll, Adhyyan Narang, Constantin Weisser, Brendan Murphy, Anca D. Dragan |
ICLR | 5 |
| 2022 | Discovering feature flag interdependencies in Microsoft officeabstractFeature flags are a popular method to control functionality in released code. They enable rapid development and deployment, but can also quickly accumulate technical debt. Complex interactions between feature flags can go unnoticed, especially if interdependent flags are located far apart in the code, and these unknown dependencies could become a source of serious bugs. Testing all possible combinations of feature flags is infeasible in large systems like Microsoft Office, which has about 12000 active flags. The goal of our research is to aid product teams in improving system reliability by providing an approach to automatically discover feature flag interdependencies. We use probabilistic reasoning to infer causal relationships from feature flag query logs. Our approach is language-agnostic, scales easily to large heterogeneous codebases, and is robust against noise such as code drift or imperfect log data. We evaluated our approach on real-world query logs from Microsoft Office and are able to achieve over 90% precision while recalling non-trivial indirect feature flag relationships across different source files. We also investigated re-occurring patterns of relationships and describe applications for targeted testing, determining deployment velocity, error mitigation, and diagnostics. Michael Schröder 0005, Katja Kevic, Daniel Gopstein, Brendan Murphy, Jennifer Beckmann |
ESEC/SIGSOFT FSE | 4 |
| 2021 | Remote Pair ProgrammingabstractPair programming is routinely used in industry and taught in face-to-face programming classes. Research indicates that it improves self-confidence and students' programming, communication and team working skills. We investigate the technology, barriers, and soft-skill benefits for distance-learning students who pair program with a remote partner online. In one study, students watched two tutors pair programming and then performed a pair programming task remotely with a student partner. Students felt significantly more positive with the latter compared to the former. As universities strive to provide a good student experience during a pandemic, these findings highlight the value of active remote pair programming using standard online communication tools. Adeola Adeliyi, Janet M. Hughes, Karen Kear, Bobby Law, Brendan Murphy, Jon Rosewell, Ann Walshe, Michel Wermelinger |
SIGCSE | 5 |
| 2021 | Towards a Theory of Software Developer Job Satisfaction and Perceived ProductivityabstractDeveloper satisfaction and work productivity are important considerations for software companies. Enhanced developer satisfaction may improve the attraction, retention and health of employees, while higher productivity should reduce costs and increase customer satisfaction through faster software improvements. Many researchers and companies assume that perceived productivity and job satisfaction are related and may be used as proxies for one another, but these claims are a current topic of debate. There are also many social and technical factors that may impact satisfaction and productivity, but which factors have the most impact is not clear, especially for specific development contexts. Through our research, we developed a theory articulating a bi-directional relationship between software developer job satisfaction and perceived productivity, and identified what additional social and technical factors, challenges and work context variables influence this relationship. The constructs and relationships in our theory were derived in part from related literature in software engineering and knowledge work, and we validated and extended these concepts through a rigorously designed survey instrument. We instantiate our theory with a large software company, which suggests a number of propositions about the relative impact of various factors and challenges on developer satisfaction and perceived productivity. Our survey instrument and analysis approach can be applied to other development settings, while our findings lead to concrete recommendations for practitioners and researchers. Margaret-Anne D. Storey, Thomas Zimmermann 0001, Christian Bird, Jacek Czerwonka, Brendan Murphy, Eirini Kalliamvakou |
IEEE Trans. Software Eng. | 5 |
| 2020 | Remote Pair ProgrammingabstractComputing students often learn to program individually or in variously-sized groups whilst studying in computing laboratories and face-to-face classes. Previous research indicates that learning via pair programming can lead to students improving the quality of their programming, enhancing their programming skills and increasing their self confidence when programming. Pair programming also is well established as a mechanism that supports peer learning and self-assessment for novice and more experienced students of programming. Observed benefits include increased self-efficacy, sharing of expertise, improved communication and team-working – all enhancing employability. A considerable amount of existing work has examined pair programming benefits as they relate to campus-based students pairing face-to-face in a laboratory class – but how can distance learning students experience such benefits? This paper describes the preliminary results from a pilot study to investigate the benefits to distance learning students of engaging in Remote Pair Programming in their learning. Our investigation goes beyond academic learning to explore community, social and employability benefits, all of which are relevant to national measures of student satisfaction. Janet M. Hughes, Ann Walshe, Bobby Law, Brendan Murphy |
CSEDU (2) | 4 |
| 2020 | Growth in Some Finite Three-Dimensional Matrix GroupsabstractWe study the growth of product sets in some finite three-dimensional matrix groups. In particular, we prove two results about the group of $2\times 2$ upper triangular matrices over arbitrary finite fields: a product set estimate using techniques from multiplicative combinatorics and an energy estimate using incidence geometry. The energy method gives better quantitative results but only applies to small sets. We also prove an energy result for the Heisenberg group. Brendan Murphy, James Wheeler |
SIAM J. Discret. Math. | 1 |
| 2019 | How Do Developers Act on Static Analysis Alerts? An Empirical Study of Coverity UsageabstractStatic analysis tools (SATs) often fall short of developer satisfaction despite their many benefits. An understanding of how developers in the real-world act on the alerts detected by SATs can help improve the utility of these tools and determine future research directions. The goal of this paper is to aid researchers and tool makers in improving the utility of static analysis tools through an empirical study of developer action on the alerts detected by Coverity, a state-of-the-art static analysis tool. In this paper, we analyze five open source projects as case studies (Linux, Firefox, Samba, Kodi, and Ovirt-engine) that have been actively using Coverity over a period of at least five years. We investigate the alert occurrences and developer triage of the alerts from the Coverity database; identify the alerts that were fixed through code changes (i.e. actionable) by mining the commit history of the projects; analyze the time an alert remain in the code base (i.e. lifespan) and the complexity of code changes (i.e. fix complexity) in fixing the alert. We find that 27.4% to 49.5% (median: 36.7%) of the alerts are actionable across projects, a rate higher than previously reported. We also find that the fixes of Coverity alerts are generally low in complexity (2 to 7 lines of code changes in the affected file, median: 4). However, developers still take from 36 to 245 days (median: 96) to fix these alerts. Finally, our data suggest that severity and fix complexity may correlate with an alert's lifespan in some of the projects. Nasif Imtiaz, Brendan Murphy, Laurie A. Williams |
ISSRE | 2 |
| 2018 | Axiomatic Foundations and Algorithms for Deciding Semantic Equivalences of SQL QueriesabstractDeciding the equivalence of SQL queries is a fundamental problem in data management. As prior work has mainly focused on studying the theoretical limitations of the problem, very few implementations for checking such equivalences exist. In this paper, we present a new formalism and implementation for reasoning about the equivalences of SQL queries. Our formalism, U-semiring, extends SQL's semiring semantics with unbounded summation and duplicate elimination. U-semiring is defined using only very few axioms and can thus be easily implemented using proof assistants such as Lean for automated query reasoning. Yet, they are sufficient enough to enable us reason about sophisticated SQL queries that are evaluated over bags and sets, along with various integrity constraints. To evaluate the effectiveness of U-semiring, we have used it to formally verify 68 equivalent queries and rewrite rules from both classical data management research papers and real-world SQL engines, where many of them have never been proven correct before. Shumo Chu, Brendan Murphy, Jared Roesch, Alvin Cheung, Dan Suciu |
Proc. VLDB Endow. | 2 |
| 2017 | Variations on the Sum-Product Problem IIabstractThis paper is a sequel to a paper entitled Variations on the sum-product problem by the same authors [SIAM J. Discrete Math., 29 (2015), pp. 514-540]. In this sequel, we quantitatively improve several of the main results of the first paper as well as generalize a method from it to give a near-optimal bound for a new expander. The main new results are the following bounds, which hold for any finite set $A \subset \mathbb R$: $\exists a \in A$ such that $|A(A+a)| \gtrsim |A|^{\frac{3}{2}+\frac{1}{186}}, |A(A-A)| \gtrsim |A|^{\frac{3}{2}+\frac{1}{34}}, |A(A+A)| \gtrsim |A|^{\frac{3}{2}+\frac{5}{242}}, |\{(a_1+a_2+a_3+a_4)^2+\log a_5 : a_i \in A \}| \gg \frac{|A|^2}{\log |A|}$. Brendan Murphy, Oliver Roche-Newton, Ilya D. Shkredov |
SIAM J. Discret. Math. | 1 |
| 2016 | Switching to Git: The Good, the Bad, and the UglyabstractSince its introduction 10 years ago, GIT has taken the world of version control systems (VCS) by storm. Its success is partly due to creating opportunities for new usage patterns that empower developers to work more efficiently. However, the resulting change in both user behavior and the way GIT stores changes impacts data mining and data analytics procedures [6], [13]. While some of these unique characteristics can be managed by adjusting mining and analytical techniques, others can lead to severe data loss and the inability to audit code changes, e.g. knowing the full history of changes of code related to security and privacy functionality. Thus, switching to GIT comes with challenges to established development process analytics. This paper is based on our experience in attempting to provide continuous process analysis for Microsoft product teams who switching to GIT as their primary VCS. We illustrate how GIT's concepts and usage patterns create a need for changing well-established data analytic processes. The goal of this paper is to raise awareness how certain GIT operations may damage or even destroy information about historical code changes necessary for continuous data development process analytics. To that end, we provide a list of common GIT usage patterns with a description of how these operations impact data mining applications. Finally, we provide examples of how one may counteract the effects of such destructive operations in the future. We further provide a new algorithm to detect integration paths that is specific to distributed version control systems like GIT, which allows us to reconstruct the information that is crucial to most development process analytics. Sascha Just, Kim Herzig, Jacek Czerwonka, Brendan Murphy |
ISSRE | 4 |
| 2015 | The Art of Testing Less without Sacrificing QualityabstractTesting is a key element of software development processes for the management and assessment of product quality. In most development environments, the software engineers are responsible for ensuring the functional correctness of code. However, for large complex software products, there is an additional need to check that changes do not negatively impact other parts of the software and they comply with system constraints such as backward compatibility, performance, security etc. Ensuring these system constraints may require complex verification infrastructure and test procedures. Although such tests are time consuming and expensive and rarely find defects they act as an insurance process to ensure the software is compliant. However, long lasting tests increasingly conflict with strategic aims to shorten release cycles. To decrease production costs and to improve development agility, we created a generic test selection strategy called THEO that accelerates test processes without sacrificing product quality. THEO is based on a cost model, which dynamically skips tests when the expected cost of running the test exceeds the expected cost of removing it. We replayed past development periods of three major Microsoft products resulting in a reduction of 50% of test executions, saving millions of dollars per year, while maintaining product quality. Kim Herzig, Michaela Greiler, Jacek Czerwonka, Brendan Murphy |
ICSE (1) | 4 |
| 2015 | Approximating Attack Surfaces with Stack TracesabstractSecurity testing and reviewing efforts are a necessity for software projects, but are time-consuming and expensive to apply. Identifying vulnerable code supports decision-making during all phases of software development. An approach for identifying vulnerable code is to identify its attack surface, the sum of all paths for untrusted data into and out of a system. Identifying the code that lies on the attack surface requires expertise and significant manual effort. This paper proposes an automated technique to empirically approximate attack surfaces through the analysis of stack traces. We hypothesize that stack traces from user-initiated crashes have several desirable attributes for measuring attack surfaces. The goal of this research is to aid software engineers in prioritizing security efforts by approximating the attack surface of a system via stack trace analysis. In a trial on Windows 8, the attack surface approximation selected 48.4% of the binaries and contained 94.6% of known vulnerabilities. Compared with vulnerability prediction models (VPMs) run on the entire codebase, VPMs run on the attack surface approximation improved recall from .07 to .1 for binaries and from .02 to .05 for source files. Precision remained at .5 for binaries, while improving from .5 to .69 for source files. Christopher Theisen, Kim Herzig, Patrick Morrison, Brendan Murphy, Laurie A. Williams |
ICSE (2) | 4 |
| 2015 | Variations on the Sum-Product ProblemabstractThis paper considers various formulations of the sum-product problem. It is shown that, for a finite set $A\subset{\mathbb{R}}$, $|A(A+A)|\gg{|A|^{\frac{3}{2}+\frac{1}{178}}},$ giving a partial answer to a conjecture of Balog. In a similar spirit, it is established that $|A(A+A+A+A)|\gg{\frac{|A|^2}{\log{|A|}}},$ a bound which is optimal up to constant and logarithmic factors. We also prove several new results concerning sum-product estimates and expanders, for example, showing that $|A(A+a)|\gg{|A|^{3/2}}$ holds for a typical element of $A$. Brendan Murphy, Oliver Roche-Newton, Ilya D. Shkredov |
SIAM J. Discret. Math. | 1 |
| 2013 | Dwelling in Software: Aspects of the Felt-Life of Engineers in Large Software Projects
Richard Harper 0001, Christian Bird, Thomas Zimmermann 0001, Brendan Murphy |
ECSCW | 4 |
| 2013 | Have Agile Techniques been the Silver Bullet for Software Development at Microsoft?abstractBackground. The pressure to release high-quality, valuable software products at an increasingly faster rate is forcing software development organizations to adapt their development practices. Agile techniques began emerging in the mid-1990s in response to this pressure and to increased volatility of customer requirements and technical change. Theoretically, agile techniques seem to be the silver bullet for responding to these pressures on the software industry. Aims. This paper tracks the changing attitudes to agile adoption and techniques, within Microsoft, in one of the largest longitudinal surveys of its kind (2006-2012). Method. We collected the opinions of 1,969 agile and non-agile practitioners in five surveys over a six-year period. Results. The survey results reveal that despite intense market pressure, the growth of agile adoption at Microsoft is slower than would be expected. Additionally, no individual agile practice exhibited strong growth trends. We also found that while development practices of teams may be similar, some perceive and declare themselves to be following an agile methodology while others do not. Both agile and non-agile practitioners agree on the relative benefits and problem areas of agile techniques. Conclusions. We found no clear trends in practice adoption. Non-agile practitioners are less enamored of the benefits and more strongly in agreement with the problem areas. The ability for agile practices to be used by large-scale teams generally concerned all respondents, which may limit its future adoption. Brendan Murphy, Christian Bird, Thomas Zimmermann 0001, Laurie A. Williams, Nachiappan Nagappan, Andrew Begel |
ESEM | 1 |
| 2012 | Characterizing and predicting which bugs get reopenedabstractFixing bugs is an important part of the software development process. An underlying aspect is the effectiveness of fixes: if a fair number of fixed bugs are reopened, it could indicate instability in the software system. To the best of our knowledge there has been on little prior work on understanding the dynamics of bug reopens. Towards that end, in this paper, we characterize when bug reports are reopened by using the Microsoft Windows operating system project as an empirical case study. Our analysis is based on a mixed-methods approach. First, we categorize the primary reasons for reopens based on a survey of 358 Microsoft employees. We then reinforce these results with a large-scale quantitative study of Windows bug reports, focusing on factors related to bug report edits and relationships between people involved in handling the bug. Finally, we build statistical models to describe the impact of various metrics on reopening bugs ranging from the reputation of the opener to how the bug was found. Thomas Zimmermann 0001, Nachiappan Nagappan, Philip J. Guo, Brendan Murphy |
ICSE | 4 |
| 2012 | The difficulties of building generic reliability models for software
Brendan Murphy |
Empir. Softw. Eng. | 1 |
| 2011 | Empirical software engineering at Microsoft ResearchabstractWe describe the activities of the Empirical Software Engi-neering (ESE) group at Microsoft Research. We highlight our research themes and activities using examples from our research on socio technical congruence, bug reporting and triaging, and data-driven software engineering to illustrate our relationship to the CSCW community. We highlight our unique ability to leverage industrial data and developers and the ability to make near term impact on Microsoft via the results of our studies. We also present the collaborations our group has with academic researchers. Christian Bird, Brendan Murphy, Nachiappan Nagappan, Thomas Zimmermann 0001 |
CSCW | 2 |
| 2011 | "Not my bug!" and other reasons for software bug report reassignmentsabstractBug reporting/fixing is an important social part of the soft-ware development process. The bug-fixing process inher-ently has strong inter-personal dynamics at play, especially in how to find the optimal person to handle a bug report. Bug report reassignments, which are a common part of the bug-fixing process, have rarely been studied. Philip J. Guo, Thomas Zimmermann 0001, Nachiappan Nagappan, Brendan Murphy |
CSCW | 4 |
| 2011 | Characterizing the differences between pre- and post- release versions of softwareabstractMany software producers utilize beta programs to predict post-release quality and to ensure that their products meet quality expectations of users. Prior work indicates that software producers need to adjust predictions to account for usage environments and usage scenarios differences between beta populations and post-release populations. However, little is known about how usage characteristics relate to field quality and how usage characteristics differ between beta and post-release. In this study, we examine application crash, application hang, system crash, and usage information from millions of Windows® users to 1) examine the effects of usage characteristics differences on field quality (e.g. which usage characteristics impact quality), 2) examine usage characteristics differences between beta and post-release (e.g. do impactful usage characteristics differ), and 3) report experiences adjusting field quality predictions for Windows. Among the 18 usage characteristics that we examined, the five most important were: the number of application executed, whether the machines was pre-installed by the original equipment manufacturer, two sub-populations (two language/geographic locales), and whether Windows was 64-bit (not 32-bit). We found each of these usage characteristics to differ between beta and post-release, and by adjusting for the differences, accuracy of field quality predictions for Windows improved by ~59%. Paul Luo Li, Ryan Kivett, Zhiyuan Zhan, Sung-eok Jeon, Nachiappan Nagappan, Brendan Murphy, Amy J. Ko |
ICSE | 6 |
| 2011 | The impact of fault models on software robustness evaluationsabstractFollowing the design and in-lab testing of software, the evaluation of its resilience to actual operational perturbations in the field is a key validation need. Software-implemented fault injection (SWIFI) is a widely used approach for evaluating the robustness of software components. Recent research [24, 18] indicates that the selection of the applied fault model has considerable influence on the results of SWIFI-based evaluations, thereby raising the question how to select appropriate fault models (i.e. that provide justified robustness evidence). This paper proposes several metrics for comparatively evaluating fault models's abilities to reveal robustness vulnerabilities. It demonstrates their application in the context of OS device drivers by investigating the influence (and relative utility) of four commonly used fault models, i.e. bit flips (in function parameters and in binaries), data type dependent parameter corruptions, and parameter fuzzing. We assess the efficiency of these models at detecting robustness vulnerabilities during the SWIFI evaluation of a real embedded operating system kernel and discuss application guidelines for our metrics alongside. Stefan Winter 0001, Constantin Sârbu, Neeraj Suri, Brendan Murphy |
ICSE | 4 |
| 2011 | Don't touch my code!: examining the effects of ownership on software qualityabstractOwnership is a key aspect of large-scale software development. We examine the relationship between different ownership measures and software failures in two large software projects: Windows Vista and Windows 7. We find that in all cases, measures of ownership such as the number of low-expertise developers, and the proportion of ownership for the top owner have a relationship with both pre-release faults and post-release failures. We also empirically identify reasons that low-expertise developers make changes to components and show that the removal of low-expertise contributions dramatically decreases the performance of contribution based defect prediction. Finally we provide recommendations for source code change policies and utilization of resources such as code inspections based on our results. Christian Bird, Nachiappan Nagappan, Brendan Murphy, Harald C. Gall, Premkumar T. Devanbu |
SIGSOFT FSE | 3 |
| 2010 | Characterizing and predicting which bugs get fixed: an empirical study of Microsoft WindowsabstractWe performed an empirical study to characterize factors that affect which bugs get fixed in Windows Vista and Windows 7, focusing on factors related to bug report edits and relationships between people involved in handling the bug. We found that bugs reported by people with better reputations were more likely to get fixed, as were bugs handled by people on the same team and working in geographical proximity. We reinforce these quantitative results with survey feedback from 358 Microsoft employees who were involved in Windows bugs. Survey respondents also mentioned additional qualitative influences on bug fixing, such as the importance of seniority and interpersonal skills of the bug reporter. Philip J. Guo, Thomas Zimmermann 0001, Nachiappan Nagappan, Brendan Murphy |
ICSE (1) | 4 |
| 2010 | Change Bursts as Defect PredictorsabstractIn software development, every change induces a risk. What happens if code changes again and again in some period of time? In an empirical study on Windows Vista, we found that the features of such change bursts have the highest predictive power for defect-prone components. With precision and recall values well above 90%, change bursts significantly improve upon earlier predictors such as complexity metrics, code churn, or organizational structure. As they only rely on version history and a controlled change process, change bursts are straight-forward to detect and deploy. Nachiappan Nagappan, Andreas Zeller, Thomas Zimmermann 0001, Kim Herzig, Brendan Murphy |
ISSRE | 5 |
| 2009 | Does distributed development affect software quality? An empirical case study of Windows VistaabstractIt is widely believed that distributed software development is riskier and more challenging than collocated development. Prior literature on distributed development in software engineering and other fields discuss various challenges, including cultural barriers, expertise transfer difficulties, and communication and coordination overhead. We evaluate this conventional belief by examining the overall development of Windows Vista and comparing the post-release failures of components that were developed in a distributed fashion with those that were developed by collocated teams. We found a negligible difference in failures. This difference becomes even less significant when controlling for the number of developers working on a binary. We also examine component characteristics such as code churn, complexity, dependency information, and test code coverage and find very little difference between distributed and collocated components to investigate if less complex components are more distributed. Further, we examine the software process and phenomena that occurred during the Vista development cycle and present ways in which the development process utilized may be insensitive to geography by mitigating the difficulties introduced in prior work in this area. Christian Bird, Nachiappan Nagappan, Premkumar T. Devanbu, Harald C. Gall, Brendan Murphy |
ICSE | 5 |
| 2009 | Putting It All Together: Using Socio-technical Networks to Predict FailuresabstractStudies have shown that social factors in development organizations have a dramatic effect on software quality. Separately, program dependency information has also been used successfully to predict which software components are more fault prone. Interestingly, the influence of these two phenomena have only been studied separately. Intuition and practical experience suggests,however, that task assignment (i.e. who worked on which components and how much) and dependency structure (which components have dependencies on others)together interact to influence the quality of the resulting software. We study the influence of combined socio-technical software networks on the fault-proneness of individual software components within a system. The network properties of a software component in this combined network are able to predict if an entity is failure prone with greater accuracy than prior methods which use dependency or contribution information in isolation. We evaluate our approach in different settings by using it on Windows Vista and across six releases of the Eclipse development environment including using models built from one release to predict failure prone components in the next release. We compare this to previous work. In every case, our method performs as well or better and is able to more accurately identify those software components that have more post-release failures, with precision and recall rates as high as 85%. Christian Bird, Nachiappan Nagappan, Harald C. Gall, Brendan Murphy, Premkumar T. Devanbu |
ISSRE | 4 |
| 2009 | Cross-project defect prediction: a large scale experiment on data vs. domain vs. processabstractPrediction of software defects works well within projects as long as there is a sufficient amount of data available to train any models. However, this is rarely the case for new software projects and for many companies. So far, only a few have studies focused on transferring prediction models from one project to another. In this paper, we study cross-project defect prediction models on a large scale. For 12 real-world applications, we ran 622 cross-project predictions. Our results indicate that cross-project prediction is a serious challenge, i.e., simply using models from projects in the same domain or with the same process does not lead to accurate predictions. To help software engineers choose models wisely, we identified factors that do influence the success of cross-project predictions. We also derived decision trees that can provide early estimates for precision, recall, and accuracy before a prediction is attempted. Thomas Zimmermann 0001, Nachiappan Nagappan, Harald C. Gall, Emanuel Giger, Brendan Murphy |
ESEC/SIGSOFT FSE | 5 |
| 2008 | The influence of organizational structure on software quality: an empirical case studyabstractOften software systems are developed by organizations consisting of many teams of individuals working together. Brooks states in the Mythical Man Month book that product quality is strongly affected by organization structure. Unfortunately there has been little empirical evidence to date to substantiate this assertion. In this paper we present a metric scheme to quantify organizational complexity, in relation to the product development process to identify if the metrics impact failure-proneness. In our case study, the organizational metrics when applied to data from Windows Vista were statistically significant predictors of failure-proneness. The precision and recall measures for identifying failure-prone binaries, using the organizational metrics, was significantly higher than using traditional metrics like churn, complexity, coverage, dependencies, and pre-release bug measures that have been used to date to predict failure-proneness. Our results provide empirical evidence that the organizational metrics are related to, and are effective predictors of failure-proneness. Nachiappan Nagappan, Brendan Murphy, Victor R. Basili |
ICSE | 2 |
| 2008 | DEFECTS 2008: international workshop on defects in large software systemsabstractBugs are everywhere in today's software and because of the huge economic damage they are actively studied. The goal of this one-day workshop is to connect the different research communities with each other and with industry. Premkumar T. Devanbu, Brendan Murphy, Nachiappan Nagappan, Thomas Zimmermann 0001, Valentin Dallmeier |
ISSTA | 2 |
| 2008 | Can developer-module networks predict failures?abstractSoftware teams should follow a well defined goal and keep their work focused. Work fragmentation is bad for efficiency and quality. In this paper we empirically investigate the relationship between the fragmentation of developer contributions and the number of post-release failures. Our approach is to represent developer contributions with a developer-module network that we call contribution network. We use network centrality measures to measure the degree of fragmentation of developer contributions. Fragmentation is determined by the centrality of software modules in the contribution network. Our claim is that central software modules are more likely to be failure-prone than modules located in surrounding areas of the network. We analyze this hypothesis by exploring the network centrality of Microsoft Windows Vista binaries using several network centrality measures as well as linear and logistic regression analysis. In particular, we investigate which centrality measures are significant to predict the probability and number of post-release failures. Results of our experiments show that central modules are more failure-prone than modules located in surrounding areas of the network. Results further confirm that number of authors and number of commits are significant predictors for the probability of post-release failures. For predicting the number of post-release failures the closeness centrality measure is most significant. Martin Pinzger 0001, Nachiappan Nagappan, Brendan Murphy |
SIGSOFT FSE | 3 |
| 2007 | On the Selection of Error Model(s) for OS Robustness EvaluationabstractThe choice of error model used for robustness evaluation of operating systems (OSs) influences the evaluation run time, implementation complexity, as well as the evaluation precision. In order to find an "effective" error model for OS evaluation, this paper systematically compares the relative effectiveness of three prominent error models, namely bit-flips, data type errors and fuzzing errors using fault injection at the interface between device drivers OS. Bit-flips come with higher costs (time) than the other models, but allow for more detailed results. Fuzzing is cheaper to implement but is found to be less precise. A composite error model is presented where the low cost of fuzzing is combined with the higher level of details of bit-flips, resulting in high precision with moderate setup and execution costs. Andréas Johansson, Neeraj Suri, Brendan Murphy |
DSN | 3 |
| 2007 | On the Impact of Injection Triggers for OS Robustness EvaluationabstractTraditionally, in fault injection-based robustness evaluation of software (specifically for operating systems - OS's), faults or errors are injected at specific code locations. This paper studies the sensitivity and accuracy of the robustness evaluation results arising from varying the timing of injecting the faults into the OS. A strategy to guide the triggering of fault injection is proposed, based on the observation that the operational usage profile of a driver shows a high degree of regularity in the calls being made. The concept of call blocks (i.e., a distinct sequence of calls made to the driver) can be used to guide injections into different system states, corresponding to the driver operations carried out. A real-world case study compares the effectiveness of the proposed strategy to traditional location-based approaches, demonstrating that significant and useful insights can be gained by modulating the injection instants. Andréas Johansson, Neeraj Suri, Brendan Murphy |
ISSRE | 3 |
| 2007 | Post-release reliability growth in software productsabstractMost software reliability growth models work under the assumption that reliability of software grows due to the removal of bugs that cause failures. However, another phenomenon has often been observed—the failure rate of a software product following its release decreases with time even if no bugs are corrected. In this article we present a simple model to represent this phenomenon. We introduce the concept of initial transient failure rate of the product and assume that it decays with a factor α per unit time thereby increasing the product reliability with time. When the transient failure rate decays away, the product displays a steady state failure rate. We discuss how the parameters in this model—initial transient failure rate, decay factor, and steady state failure rate—can be determined from the failure and sales data of a product. We also describe how, using the model, we can determine the product stabilization time—a product quality metric that describes how long it takes a product to reach close to its stable failure rate. We provide many examples where this model has been applied to data from released products. Pankaj Jalote, Brendan Murphy, Vibhu Saujanya Sharma |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2006 | Using Historical In-Process and Product Metrics for Early Estimation of Software FailuresabstractThe benefits that a software organization obtains from estimates of product quality are dependent upon how early in the product cycle that these estimates are available. Early estimation of software quality can help organizations make informed decisions about corrective actions. To provide such early estimates we present an empirical case study of two large scale commercial operating systems, Windows XP and Windows Server 2003. In particular, we leverage various historical in-process and product metrics from Windows XP binaries to create statistical predictors to estimate the post-release failures/failure-proneness of Windows Server 2003 binaries. These models estimate the failures and failure-proneness of Windows Server 2003 binaries at statistically significant levels. Our study is unique in showing that historical predictors for a software product line can be useful, even at the very large scale of the Windows operating system Nachiappan Nagappan, Thomas Ball 0001, Brendan Murphy |
ISSRE | 3 |
| 2005 | Dependability Benchmarking of Computing Systems - Panel StatementabstractThe importance of benchmarking is increasing as every aspect of human life is relying on correct operation of computing systems. Although considerable efforts have been made, presently there are no widely accepted dependability benchmarks. The panelists, in the order shown above, address benchmarking of computer hardware, operating systems, applications, and systems. Cristian Constantinescu, Karama Kanoun, Henrique Madeira, Brendan Murphy, Ira Pramanick, Aaron B. Brown |
DSN | 4 |
| 2004 | Reliability Growth in Software ProductsabstractMost of the software reliability growth models work under the assumption that reliability of software grows due to the bugs that cause failures being removed from the software. While correcting bugs will improve reliability, another phenomenon has often been observed - the failure rate of a software product, as observed by the user, improves with time irrespective of whether bugs are corrected or not. Consequently, the reliability of a product, as observed by users, varies, depending on the length of time they have been using the product. One reason for this reliability growth is that as the users gain experience with the product, they learn to use the product correctly and find work-around for failure-causing situations. Another factor that affects this growth is that following the product installation, the user discovers that other actions may be required, like installing new drivers, upgrading other software to a compatible version, etc. to properly configure the new product. In this paper we present a simple model to represent this phenomenon - we assume that the failure rate for a product decays with a factor /spl alpha/ per unit time. Applying this failure rate decay model to the data collected on reported failures and number of units of the product sold, it is possible to determine the initial failure rate, the decay factor, and the steady state failure rate of a product. The paper provides a number of examples where this model has been applied to data captured from released products. Pankaj Jalote, Brendan Murphy |
ISSRE | 2 |
| 2002 | Joint Panel - IPDS and Workshop on Dependability Benchmarking
Ravishankar K. Iyer, Zbigniew T. Kalbarczyk, Philip Koopman, Henrique Madeira, Gunter Heiner, Karama Kanoun, Haim Levendel, Brendan Murphy, Lawrence G. Votta, Don Wilson |
DSN | 8 |