EDBT 2026 Demo / reviewers in the wild / expert
Tracy Hall
dblp:69/3214
· DBLP profile ↗
71ranked-venue papers
11as first author
14since 2021 · last 2026
0000-0002-2728-9014ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 69 · 10 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | How vulnerability explanations help software practitioners confirm and fix code vulnerabilitiesabstractContext: Most current code vulnerability detection tools provide only a binary classification (vulnerable/non-vulnerable) with little to no additional context. This paper explores the impact of providing explanations for vulnerabilities alongside code labelled as vulnerable. Objective: We investigate the influence of explanations on the ability of software practitioners to confirm such labelled code as actually vulnerable (i.e., a true positive vulnerability) and to fix such vulnerable code correctly. Method: We surveyed 99 software practitioners to establish their use of code-vulnerability detection tools and to evaluate the impact of explanations on their behaviour towards code labelled as vulnerable in a series of coding exercises. Participants were presented with four forms of explanation: vulnerable lines , vulnerability type , short-form text , and long-form text . Results: Software practitioners performed better at confirming and fixing code vulnerabilities when presented with any of the four forms of explanation. Although practitioners stated a preference for long-form text explanations, they achieved the highest confirmation and fixing performance with short-form text explanations. Practitioners also indicated willingness to accept modest drops in detection precision and recall if richer explanations were provided, and their preferences for explanation types and performance trade-offs varied according to where a detection tool is used in the software-development pipeline. Conclusions: Vulnerability-detection and prediction tools should provide explanatory output and allow different explanation types tailored to their deployment stage in the development workflow. Few current tools provide any explanations, and none identified in this study provide text-based explanations. Fahad Al Debeyan, Tracy Hall, Lech Madeyski, Emily Winter 0001 |
Inf. Softw. Technol. | 2 |
| 2024 | Semgrep*: Improving the Limited Performance of Static Application Security Testing (SAST) ToolsabstractVulnerabilities in code should be detected and patched quickly to reduce the time in which they can be exploited. There are many automated approaches to assist developers in detecting vulnerabilities, most notably Static Application Security Testing (SAST) tools. However, no single tool detects all vulnerabilities and so relying on any one tool may leave vulnerabilities dormant in code. In this study, we use a manually curated dataset to evaluate four SAST tools on production code with known vulnerabilities. Our results show that the vulnerability detection rates of individual tools range from 11.2% to 26.5%, but combining these four tools can detect 38.8% of vulnerabilities. We investigate why SAST tools are unable to detect 61.2% of vulnerabilities and identify missing vulnerable code patterns from tool rule sets. Based on our findings, we create new rules for Semgrep, a popular configurable SAST tool. Our newly configured Semgrep tool detects 44.7% of vulnerabilities, more than using a combination of tools, and a 181% improvement in Semgrep’s detection rate. Gareth Bennett, Tracy Hall, Emily Winter 0001, Steve Counsell |
EASE | 2 |
| 2024 | Do Developers Use Static Application Security Testing (SAST) Tools Straight Out of the Box? A large-scale Empirical StudyabstractStatic application Security Testing (SAST) tools are an established means of detecting vulnerabilities early in development. Previous studies have reported low detection rates from SAST tools and recommend either combining SAST tools or configuring rule sets to detect more vulnerabilities. However, while previous work suggests that developers rarely combine or configure any of the Automatic Static Analysis Tools (ASATs) they use, it is currently unclear whether SAST tools are used directly “out of the box”. To understand how developers use SAST tools, we performed a large-scale survey involving 1,263 developers. We pre-screened developers to establish their SAST use and found that only 20% (204/1,003) used SAST tools. Of those developers who did use SAST tools, we found a large number did not use multiple tools (59%), did not configure tools (54%) or did neither (40%). Our results suggest that more work is needed to help developers combine and configure tools, since doing so is likely to detect significantly more vulnerabilities. Gareth Bennett, Tracy Hall, Steve Counsell, Emily Winter 0001, Thomas Shippey |
ESEM | 2 |
| 2024 | Different Strokes for Different Folks: A Comparison of Developer and Tester Views on TestingabstractIn this paper, we re-analyse the data from a previous study of Straubinger et al., which asked 284 industrial IT staff about their views on testing. In that study, as well as developers, the dedicated role of tester was included in the data - both roles were treated as the same role. In this paper, we posit that the two roles (i.e., developer and tester) are so very different that we should analyse each role separately; testers will have unique insights into testing, so separating their views and experiences from developers is important. To this end, we analyse six of the same research questions as the original study, using separate developer and tester data. Results showed that for almost every question we re-visited, testers differed in their opinions from developers, whether on the type of testing they did, measures of code quality, effort to write tests and motivation for testing. Steve Counsell, Stephen Swift, Mahir Arzoky, Tracy Hall, Emily Winter 0001, Gareth Bennett, Thomas Shippey |
SEAA | 4 |
| 2024 | The impact of hard and easy negative training data on vulnerability prediction performanceabstractVulnerability prediction models have been shown to perform poorly in the real world. We examine how the composition of negative training data influences vulnerability prediction model performance. Inspired by other disciplines (e.g. image processing), we focus on whether distinguishing between negative training data that is ‘easy’ to recognise from positive data (very different from positive data) and negative training data that is ‘hard’ to recognise from positive data (very similar to positive data) impacts on vulnerability prediction performance. We use a range of popular machine learning algorithms, including deep learning, to build models based on vulnerability patch data curated by Reis and Abreu, as well as the MSR dataset. Our results suggest that models trained on higher ratios of easy negatives perform better, plateauing at 15 easy negatives per positive instance. We also report that different ML algorithms work better based on the negative sample used. Overall, we found that the negative sampling approach used significantly impacts model performance, potentially leading to overly optimistic results. The ratio of ‘easy’ versus ‘hard’ negative training data should be explicitly considered when building vulnerability prediction models for the real world. Fahad Al Debeyan, Lech Madeyski, Tracy Hall, David Bowes |
J. Syst. Softw. | 3 |
| 2023 | The Paradox of Analysing Gender-Based DataabstractIn this short paper, we analyse "gender" perspectives from a survey of three hundred and seventy-eight industry developers on two aspects of IT industry developer practice: bugs and Automatic Program Repair. We also explore questions of how developers view their job satisfaction. Our key motivation was to show whether there was a difference in the way that males and females viewed these three important concepts. From a total of thirteen survey questions analysed, only two showed any statistical difference between the responses of females compared to males. Those differences were found exclusively in the job satisfaction part of the survey. In terms of the way that male or female developers think about technical activities per se and diversity and inclusivity more generally, we therefore have the paradoxical issue of whether gender comparisons have any basis. We all think the same way about technical-oriented activities, so perhaps we need to stop trying to find differences and division. Steve Counsell, Emily Winter 0001, Tracy Hall, Vesna Nowack |
SEAA | 3 |
| 2023 | Fault-insertion and fault-fixing behavioural patterns in Apache Software Foundation ProjectsabstractDevelopers inevitably make human errors while coding. These errors can lead to faults in code, some of which may result in system failures. It is important to reduce the faults inserted by developers as well as fix any that slip through. To investigate the fault insertion and fault fixing activities of developers. We identify developers who insert and fix faults, ask whether code topic ‘experts’ insert fewer faults, and experts fix more faults and whether patterns of insertion and fixing change over time. We perform a time-based analysis of developer activity on twelve Apache projects using Latent Dirichlet Allocation (LDA), Network Analysis and Topic Modelling. We also build three models (using Petri-net, Markov Chain and Hawkes Processes) which describe and simulate developers’ bug-introduction and fixing behaviour. We show that: the majority of the projects we analysed have developers who dominate in the insertion and fixing of faults; Faults are less likely to be inserted by developers with code topic expertise; Different projects have different patterns of fault inserting and fixing over time. We recommend that projects identify the code topic expertise of developers and use expertise information to inform the assignment of project work. Marco Ortu, Giuseppe Destefanis, Tracy Hall, David Bowes |
Inf. Softw. Technol. | 3 |
| 2023 | How do Developers Really Feel About Bug Fixing? Directions for Automatic Program RepairabstractAutomatic program repair (APR) is a rapidly advancing field of software engineering that aims to supplement or replace manual bug fixing with an automated tool. For APR to be successfully adopted in industry, it is vital that APR tools respond to developer needs and preferences. However, very little research has considered developers' general attitudes to APR or developers' current bug fixing practices (the activity APR aims to replace). This paper responds to this gap by reporting on a survey of 386 software developers about their bug finding and fixing practices and experiences, and their instinctive attitudes towards APR. We find that bug finding and fixing is not necessarily as onerous for developers as has often been suggested, being rated as more satisfying than developers' general work. The fact that developers derive satisfaction and benefit from bug fixing indicates that APR adoption is not as simple as APR replacing an unwanted activity. When it comes to potential APR approaches, we find a strong preference for developers being kept in the loop (for example, choosing between different fixes or validating fixes) as opposed to a fully automated process. This suggests that advances in APR should be careful to consider the agency of the developer, as well as what information is presented to developers alongside fixes. It also indicates that there are key barriers related to trust that would need to be overcome for full scale APR adoption, supported by the fact that even those developers who stated that they were positive about APR listed several caveats and concerns. We find very few statistically significant relationships between particular demographic variables (for example, developer experience, age, education) and key attitudinal variables, suggesting that developers' instinctive attitudes towards APR are little influenced by experience level but are held widely across the developer community. Emily Winter 0001, David Bowes, Steve Counsell, Tracy Hall, Saemundur O. Haraldsson, Vesna Nowack, John R. Woodward |
IEEE Trans. Software Eng. | 4 |
| 2023 | Let's Talk With Developers, Not About Developers: A Review of Automatic Program Repair ResearchabstractAutomatic program repair (APR) offers significant potential for automating some coding tasks. Using APR could reduce the high costs historically associated with fixing code faults and deliver significant benefits to software engineering. Adopting APR could also have profound implications for software developers’ daily activities, transforming their work practices. To realise the benefits of APR it is vital that we consider how developers feel about APR and the impact APR may have on developers’ work. Developing APR tools without consideration of the developer is likely to undermine the success of APR deployment. In this paper, we critically review how developers are considered in APR research by analysing how human factors are treated in 260 studies from Monperrus’s Living Review of APR. Over half of the 260 studies in our review were motivated by a problem faced by developers (e.g., the difficulty associated with fixing faults). Despite these human-oriented motivations, fewer than 7% of the 260 studies included a human study. We looked in detail at these human studies and found their quality mixed (for example, one human study was based on input from only one developer). Our results suggest that software developers are often talkedaboutin APR studies, but are rarely talkedwith. A more comprehensive and reliable understanding of developer human factors in relation to APR is needed. Without this understanding, it will be difficult to develop APR tools and techniques which integrate effectively into developers’ workflows. We recommend a future research agenda to advance the study of human factors in APR. Emily Winter 0001, Vesna Nowack, David Bowes, Steve Counsell, Tracy Hall, Saemundur O. Haraldsson, John R. Woodward |
IEEE Trans. Software Eng. | 5 |
| 2022 | An 80-20 Analysis of Buggy and Non-buggy Refactorings in Open-Source CommitsabstractIn this short paper, we explore the Pareto principle, sometimes known as the “80-20” rule as part of the refactoring process. We explore five frequently applied refactorings, namely extract method, extract variable, rename variable, rename method and change variable type from a data set of forty open-source systems and nearly two hundred thousand refactorings. We address two key research questions. Firstly, do 80% of “buggy” refactorings (where a refactoring has induced a bug fix) arise from just 20% of commits and, secondly, does the same rule apply to “non-buggy” refactorings when applied to the same systems? To facilitate our analysis, we used refactoring and bug data from a study by Di Penta et al. Results showed that refactorings inducing bugs were clustered around a more concentrated set of commits than refactorings that did not induce bugs. One refactoring ‘change variable type’ stood out - it almost conformed to an 80-20 rule. The take-away message is, as the saying goes, that too much of a “good” thing [refactoring] could actually be a “bad” thing. Steve Counsell, Vesna Nowack, Tracy Hall, David Bowes, Saemundur O. Haraldsson, Emily Winter 0001, John R. Woodward |
SEAA | 3 |
| 2022 | Towards developer-centered automatic program repair: findings from BloombergabstractThis paper reports on qualitative research into automatic program repair (APR) at Bloomberg. Six focus groups were conducted with a total of seventeen participants (including both developers of the APR tool and developers using the tool) to consider: the development at Bloomberg of a prototype APR tool (Fixie); developers’ early experiences using the tool; and developers’ perspectives on how they would like to interact with the tool in future. APR is developing rapidly and it is important to understand in greater detail developers' experiences using this emerging technology. In this paper, we provide in-depth, qualitative data from an industrial setting. We found that the development of APR at Bloomberg had become increasingly user-centered, emphasising how fixes were presented to developers, as well as particular features, such as customisability. From the focus groups with developers who had used Fixie, we found particular concern with the pragmatic aspects of APR, such as how and when fixes were presented to them. Based on our findings, we make a series of recommendations to inform future APR development, highlighting how APR tools should 'start small', be customisable, and fit with developers' workflows. We also suggest that APR tools should capitalise on the promise of repair bots and draw on advances in explainable AI. Emily Winter 0001, Vesna Nowack, David Bowes, Steve Counsell, Tracy Hall, Saemundur O. Haraldsson, John R. Woodward, Serkan Kirbas, Etienne Windels, Olayori McBello, Abdurahman Atakishiyev, Kevin Kells, Matthew W. Pagano |
ESEC/SIGSOFT FSE | 5 |
| 2022 | How Software Developers Mitigate Their Errors When Developing CodeabstractCode remains largely hand-made by humans and, as such, writing code is prone to error. Many previous studies have focused on the technical reasons for these errors and provided developers with increasingly sophisticated tools. Few studies have looked in detail at why code errors have been made from a human perspective. We use Human Error Theory to frame our exploratory study and use semi-structured interviews to uncover a preliminary understanding of the errors developers make while coding. We look particularly at the Skill-based (SB) errors reported by 27 professional software developers. We found that the complexity of the development environment is one of the most frequently reported reasons for errors. Maintaining concentration and focus on a particular task also underpins many developer errors. We found that developers struggle with effective mitigation strategies for their errors, reporting strategies largely based on improving their own willpower to concentrate better on coding tasks. We discuss how using Reason’s Swiss Cheese model may help reduce errors during software development. This model ensures that layers of tool, process and management mitigation are in place to prevent developer errors from causing system failures. Bhaveet Nagaria, Tracy Hall |
IEEE Trans. Software Eng. | 2 |
| 2021 | Expanding Fix Patterns to Enable Automatic Program RepairabstractAutomatic Program Repair (APR) has been proposed to help developers and reduce the time spent repairing programs. Recent APR tools have applied learned templates (fix patterns) to fix code using knowledge from fixes successfully applied in the past. However, there is still no general agreement on the representation of fix patterns, making their application and comparison with a baseline difficult. As a consequence, it is also difficult to expand fix patterns and further enable APR. We automatically generate fix patterns from similar fixes and compare the generated fix patterns against a state-of-the-art taxonomy. Our automated approach splits fixes into smaller, method-level chunks and calculates their similarity. A threshold-based clustering algorithm groups similar chunks and finds matches with state-of-the-art fix patterns. In our evaluation, we present 33 clusters whose fix patterns were generated from the fixes of 835 Defects4J bugs. Of those 33 clusters, 22 matched a state-of-the-art taxonomy with good agreement. The remaining 11 clusters were thematically analysed and generated new fix patterns that expanded the taxonomy. Our new fix patterns should enable APR researchers and practitioners to expand their tools to fix a greater range of bugs in the future. Vesna Nowack, David Bowes, Steve Counsell, Tracy Hall, Saemundur O. Haraldsson, Emily Winter 0001, John R. Woodward |
ISSRE | 4 |
| 2021 | Using Machine Learning to Recognise Novice and Expert Programmers
Chi Hong Lee, Tracy Hall |
PROFES | 2 |
| 2020 | Which Software Faults Are Tests Not Detecting?abstractContext: Software testing plays an important role in assuring the reliability of systems. Assessing the efficacy of testing remains challenging with few established test effectiveness metrics. Those metrics that have been used (e.g. coverage and mutation analysis) have been criticised for insufficiently differentiating between the faults detected by tests. Objective: We investigate how effective tests are at detecting different types of faults and whether some types of fault evade tests more than others. Our aim is to suggest to developers specific ways in which their tests need to be improved to increase fault detection. Method: We investigate seven fault types and analyse how often each goes undetected in 10 open source systems. We statistically look for any relationship between the test set and faults. Results: Our results suggest that the fault detection rates of unit tests are relatively low, typically finding only about a half of all faults. In addition, conditional boundary and method call removals are less well detected by tests than other fault types. Conclusions: We conclude that the testing of these open source systems needs to be improved across the board. In addition, despite boundary cases being long known to attract faults, tests covering boundaries need particular improvement. Overall, we recommend that developers do not rely only on code coverage and mutation score to measure the effectiveness of their tests. Jean Petric, Tracy Hall, David Bowes |
EASE | 2 |
| 2020 | BugVis: Commit Slicing for Fault VisualisationabstractIn this paper we present BugVis, our tool which allows the visualisation of the lifetime of a code fault. The commit history of the fault from insertion to fix is visualised. Unlike previous similar tools, BugVis visualises only the lines of each commit involved in the fault. The visualisation creates a commit slice throughout the history of the fault which enables comprehension of the evolution of the code involved in the fault. David Bowes, Jean Petric, Tracy Hall |
ICPC | 3 |
| 2019 | Automatically identifying code features for software defect prediction: Using AST N-grams
Thomas Shippey, David Bowes, Tracy Hall |
Inf. Softw. Technol. | 3 |
| 2018 | Code Cleaning for Software Defect Prediction: A Cautionary TaleabstractIn this paper, we describe our experience of developing a new technique to improve defect prediction (code cleaning) which performed very encouragingly on the first two systems on which we evaluated it (both systems had their origins in one company). Code cleaning also worked well on an additional open source system (Eclipse). But our code cleaning technique then performed disappointingly on all 69 subsequent open source systems on which we evaluated it. Without our round two evaluations on these 69 open source systems we would have published misleading prediction results. We discuss the need for performance evaluations to be performed on carefully selected samples of systems if reliable conclusions are to be drawn. Thomas Shippey, David Bowes, Steve Counsell, Tracy Hall |
SEAA | 4 |
| 2018 | The relationship between evolutionary coupling and defects in large industrial software (journal-first abstract)abstractIn this study, we investigate the effect of EC on the defect-proneness of large industrial software systems and explain why the effects vary. Serkan Kirbas, Bora Caglayan, Tracy Hall, Steve Counsell, David Bowes, Alper Sen 0001, Ayse Basar Bener |
SANER | 3 |
| 2018 | Reproducibility and replicability of software defect prediction studies
Zaheed Mahmood, David Bowes, Tracy Hall, Peter C. R. Lane, Jean Petric |
Inf. Softw. Technol. | 3 |
| 2018 | Software defect prediction: do different classifiers find the same defects?abstractDuring the last 10 years, hundreds of different defect prediction models have been published. The performance of the classifiers used in these models is reported to be similar with models rarely performing above the predictive performance ceiling of about 80% recall. We investigate the individual defects that four classifiers predict and analyse the level of prediction uncertainty produced by these classifiers. We perform a sensitivity analysis to compare the performance of Random Forest, Naïve Bayes, RPart and SVM classifiers when predicting defects in NASA, open source and commercial datasets. The defect predictions that each classifier makes is captured in a confusion matrix and the prediction uncertainty of each classifier is compared. Despite similar predictive performance values for these four classifiers, each detects different sets of defects. Some classifiers are more consistent in predicting defects than others. Our results confirm that a unique subset of defects can be detected by specific classifiers. However, while some classifiers are consistent in the predictions they make, other classifiers vary in their predictions. Given our results, we conclude that classifier ensembles with decision-making strategies not based on majority voting are likely to perform best in defect prediction. David Bowes, Tracy Hall, Jean Petric |
Softw. Qual. J. | 2 |
| 2018 | Authors' Reply to "Comments on 'Researcher Bias: The Use of Machine Learning in Software Defect Prediction'"abstractIn 2014 we published a meta-analysis of software defect prediction studies [1] . This suggested that the most important factor in determining results was Research Group, i.e., who conducts the experiment is more important than the classifier algorithms being investigated. A recent re-analysis [2] sought to argue that the effect is less strong than originally claimed since there is a relationship between Research Group and Dataset. In this response we show (i) the re-analysis is based on a small (21 percent) subset of our original data, (ii) using the same re-analysis approach with a larger subset shows that Research Group is more important than type of Classifier and (iii) however the data are analysed there is compelling evidence that who conducts the research has an effect on the results. This means that the problem of researcher bias remains. Addressing it should be seen as a matter of priority amongst those of us who conduct and publish experiments comparing the performance of competing software defect prediction systems. Martin J. Shepperd, Tracy Hall, David Bowes |
IEEE Trans. Software Eng. | 2 |
| 2017 | Introduction to the Special Section from the Empirical Track of the XP2016 conference
Helen Sharp, Tracy Hall, Nat Pryce |
Inf. Softw. Technol. | 2 |
| 2017 | Evolutionary coupling measurement: Making sense of the current chaos
Serkan Kirbas, Tracy Hall, Alper Sen 0001 |
Sci. Comput. Program. | 2 |
| 2017 | The relationship between evolutionary coupling and defects in large industrial softwareabstractAbstract Evolutionary coupling (EC) is defined as the implicit relationship between 2 or more software artifacts that are frequently changed together. Changing software is widely reported to be defect‐prone. In this study, we investigate the effect of EC on the defect proneness of large industrial software systems and explain why the effects vary. We analysed 2 large industrial systems: a legacy financial system and a modern telecommunications system. We collected historical data for 7 years from 5 different software repositories containing 176 thousand files. We applied correlation and regression analysis to explore the relationship between EC and software defects, and we analysed defect types, size, and process metrics to explain different effects of EC on defects through correlation. Our results indicate that there is generally a positive correlation between EC and defects, but the correlation strength varies. Evolutionary coupling is less likely to have a relationship to software defects for parts of the software with fewer files and where fewer developers contributed. Evolutionary coupling measures showed higher correlation with some types of defects (based on root causes) such as code implementation and acceptance criteria. Although EC measures may be useful to explain defects, the explanatory power of such measures depends on defect types, size, and process metrics. Serkan Kirbas, Bora Caglayan, Tracy Hall, Steve Counsell, David Bowes, Alper Sen 0001, Ayse Basar Bener |
J. Softw. Evol. Process. | 3 |
| 2016 | The jinx on the NASA software defect data setsabstractBackground: The NASA datasets have previously been used extensively in studies of software defects. In 2013 Shepperd et al. presented an essential set of rules for removing erroneous data from the NASA datasets making this data more reliable to use. Jean Petric, David Bowes, Tracy Hall, Bruce Christianson, Nathan Baddoo |
EASE | 3 |
| 2016 | Building an Ensemble for Software Defect Prediction Based on Diversity SelectionabstractBackground: Ensemble techniques have gained attention in various scientific fields. Defect prediction researchers have investigated many state-of-the-art ensemble models and concluded that in many cases these outperform standard single classifier techniques. Almost all previous work using ensemble techniques in defect prediction rely on the majority voting scheme for combining prediction outputs, and on the implicit diversity among single classifiers. Aim: Investigate whether defect prediction can be improved using an explicit diversity technique with stacking ensemble, given the fact that different classifiers identify different sets of defects. Method: We used classifiers from four different families and the weighted accuracy diversity (WAD) technique to exploit diversity amongst classifiers. To combine individual predictions, we used the stacking ensemble technique. We used state-of-the-art knowledge in software defect prediction to build our ensemble models, and tested their prediction abilities against 8 publicly available data sets. Conclusion: The results show performance improvement using stacking ensembles compared to other defect prediction models. Diversity amongst classifiers used for building ensembles is essential to achieving these performance improvements. Jean Petric, David Bowes, Tracy Hall, Bruce Christianson, Nathan Baddoo |
ESEM | 3 |
| 2016 | So You Need More Method Level Datasets for Your Software Defect Prediction?: Voilà!abstractContext: Defect prediction research is based on a small number of defect datasets and most are at class not method level. Consequently our knowledge of defects is limited. Identifying defect datasets for prediction is not easy and extracting quality data from identified datasets is even more difficult. Goal: Identify open source Java systems suitable for defect prediction and extract high quality fault data from these datasets. Method: We used the Boa to identify candidate open source systems. We reduce 50,000 potential candidates down to 23 suitable for defect prediction using a selection criteria based on the system's software repository and its defect tracking system. We use an enhanced SZZ algorithm to extract fault information and calculate metrics using JHawk. Result: We have produced 138 fault and metrics datasets for the 23 identified systems. We make these datasets (the ELFF datasets) and our data extraction tools freely available to future researchers. Conclusions: The data we provide enables future studies to proceed with minimal effort. Our datasets significantly increase the pool of systems currently being used in defect analysis studies. Thomas Shippey, Tracy Hall, Steve Counsell, David Bowes |
ESEM | 2 |
| 2016 | Mutation-aware fault predictionabstractWe introduce mutation-aware fault prediction, which leverages additional guidance from metrics constructed in terms of mutants and the test cases that cover and detect them. We report the results of 12 sets of experiments, applying 4 different predictive modelling techniques to 3 large real-world systems (both open and closed source). The results show that our proposal can significantly (p ≤ 0.05) improve fault prediction performance. Moreover, mutation-based metrics lie in the top 5% most frequently relied upon fault predictors in 10 of the 12 sets of experiments, and provide the majority of the top ten fault predictors in 9 of the 12 sets of experiments. David Bowes, Tracy Hall, Mark Harman, Yue Jia 0001, Federica Sarro, Fan Wu 0009 |
ISSTA | 2 |
| 2015 | Editorial for the special section on Empirical Studies in Software Engineering Selected, and extended papers from the Eighteenth International Conference on Evaluation and Assessment in Software Engineering, May 13th-14th 2014, London, UK
Tracy Hall, Steve Counsell, Ingunn Myrtveit |
Inf. Softw. Technol. | 1 |
| 2014 | Filling the Gaps of Development Logs and Bug Issue DataabstractIt has been suggested that the data from bug repositories is not always in sync or complete compared to the logs detailing the actions of developers on source code. Bilyaminu Auwal Romo, Andrea Capiluppi, Tracy Hall |
OpenSym | 3 |
| 2014 | DConfusion: a technique to allow cross study performance evaluation of fault prediction studies
David Bowes, Tracy Hall, David Gray |
Autom. Softw. Eng. | 2 |
| 2014 | Factors that motivate software engineering teams: A four country empirical study
June M. Verner, Muhammad Ali Babar 0001, Narciso Cerpa, Tracy Hall, Sarah Beecham |
J. Syst. Softw. | 4 |
| 2014 | Some Code Smells Have a Significant but Small Effect on FaultsabstractWe investigate the relationship between faults and five of Fowler et al.'s least-studied smells in code: Data Clumps, Switch Statements, Speculative Generality, Message Chains, and Middle Man. We developed a tool to detect these five smells in three open-source systems: Eclipse, ArgoUML, and Apache Commons. We collected fault data from the change and fault repositories of each system. We built Negative Binomial regression models to analyse the relationships between smells and faults and report the McFadden effect size of those relationships. Our results suggest that Switch Statements had no effect on faults in any of the three systems; Message Chains increased faults in two systems; Message Chains which occurred in larger files reduced faults; Data Clumps reduced faults in Apache and Eclipse but increased faults in ArgoUML; Middle Man reduced faults only in ArgoUML, and Speculative Generality reduced faults only in Eclipse. File size alone affects faults in some systems but not in all systems. Where smells did significantly affect faults, the size of that effect was small (always under 10 percent). Our findings suggest that some smells do indicate fault-prone code in some circumstances but that the effect that these smells have on faults is small. Our findings also show that smells have different effects on different systems. We conclude that arbitrary refactoring is unlikely to significantly reduce fault-proneness and in some cases may increase fault-proneness. Tracy Hall, Min Zhang 0008, David Bowes, Yi Sun 0001 |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2014 | Researcher Bias: The Use of Machine Learning in Software Defect PredictionabstractBackground. The ability to predict defect-prone software components would be valuable. Consequently, there have been many empirical studies to evaluate the performance of different techniques endeavouring to accomplish this effectively. However no one technique dominates and so designing a reliable defect prediction model remains problematic. Objective. We seek to make sense of the many conflicting experimental results and understand which factors have the largest effect on predictive performance. Method. We conduct a meta-analysis of all relevant, high quality primary studies of defect prediction to determine what factors influence predictive performance. This is based on 42 primary studies that satisfy our inclusion criteria that collectively report 600 sets of empirical prediction results. By reverse engineering a common response variable we build a random effects ANOVA model to examine the relative contribution of four model building factors (classifier, data set, input metrics and researcher group) to model prediction performance. Results. Surprisingly we find that the choice of classifier has little impact upon performance (1.3 percent) and in contrast the major (31 percent) explanatory factor is the researcher group. It matters more who does the work than what is done. Conclusion. To overcome this high level of researcher bias, defect prediction researchers should (i) conduct blind analysis, (ii) improve reporting protocols and (iii) conduct more intergroup studies in order to alleviate expertise issues. Lastly, research is required to determine whether this bias is prevalent in other applications domains. Martin J. Shepperd, David Bowes, Tracy Hall |
IEEE Trans. Software Eng. | 3 |
| 2012 | A mapping study of software code cloningabstractBackground: Software Code Cloning is widely used by developers to produce code in which they have confidence and which reduces development costs and improves the software quality. However, Fowler and Beck suggest that the maintenance of clones may lead to defects and therefore clones should be re-factored out. Objective: We investigate the purpose of code cloning, the detection techniques developed and the datasets used in software code cloning studies between the years of 2007 and 2011. This is to analyse the current research trends in code cloning to try and find techniques which have been successful in identifying clones used for defect prediction. Method: We used a mapping study to identify 220 software code cloning studies published from January 2007 to December 2011. We use these papers to answer six research questions by analysing their abstracts, titles and reading the papers themselves. Results: The main focus of studies is the technique of software code clone detection. In the past four years the number of studies being accepted at conferences and in journals has risen by 71%. Most datasets are only used once, therefore the performance reported by one paper is not comparable with the performance reported by another study. Conclusion: The techniques used to detect clones seem to be the main focus of studies. However it is difficult to compare the performance of the detection tools reported in different studies because the same dataset is rarely used in more than one paper. There are few benchmark datasets where the clones have been correctly identified. Few studies apply code cloning detection to defect prediction. Thomas Shippey, David Bowes, Bruce Christianson, Tracy Hall |
EASE | 4 |
| 2012 | The State of Machine Learning Methodology in Software Fault PredictionabstractThe aim of this paper is to investigate the quality of methodology in software fault prediction studies using machine learning. Over two hundred studies of fault prediction have been published in the last 10 years. There is evidence to suggest that the quality of methodology used in some of these studies does not allow us to have confidence in the predictions reported by them. We evaluate the machine learning methodology used in 21 fault prediction studies. All of these studies use NASA data sets. We score each study from 1 to 10 in terms of the quality of their machine learning methodology (e.g. whether or not studies report randomising their cross validation folds). Only 10 out of the 21 studies scored 5 or more out of 10. Furthermore 1 study scored only 1 out of 10. When we plot these scores over time there is no evidence that the quality of machine learning methodology is better in recent studies. Our results suggest that there remains much to be done by both researchers and reviewers to improve the quality of machine learning methodology used in software fault prediction. We conclude that the results reported in some studies need to be treated with caution. Tracy Hall, David Bowes |
ICMLA (2) | 1 |
| 2012 | A Systematic Literature Review on Fault Prediction Performance in Software EngineeringabstractBackground: The accurate prediction of where faults are likely to occur in code can help direct test effort, reduce costs, and improve the quality of software. Objective: We investigate how the context of models, the independent variables used, and the modeling techniques applied influence the performance of fault prediction models. Method: We used a systematic literature review to identify 208 fault prediction studies published from January 2000 to December 2010. We synthesize the quantitative and qualitative results of 36 studies which report sufficient contextual and methodological information according to the criteria we develop and apply. Results: The models that perform well tend to be based on simple modeling techniques such as Naive Bayes or Logistic Regression. Combinations of independent variables have been used by models that perform well. Feature selection has been applied to these combinations when models are performing particularly well. Conclusion: The methodology used to build models seems to be influential to predictive performance. Although there are a set of fault prediction studies in which confidence is possible, more studies are needed that use a reliable methodology and which report their context, methodology, and performance comprehensively. Tracy Hall, Sarah Beecham, David Bowes, David Gray, Steve Counsell |
IEEE Trans. Software Eng. | 1 |
| 2011 | Code Bad Smells: a review of current knowledgeabstractAbstract Fowler et al. identified 22 Code Bad Smells to direct the effective refactoring of code. These are increasingly being taken up by software engineers. However, the empirical basis of using Code Bad Smells to direct refactoring and to address ‘trouble’ in code is not clear, i.e., we do not know whether using Code Bad Smells to target code improvement is effective. This paper aims to identify what is currently known about Code Bad Smells. We have performed a systematic literature review of 319 papers published since Fowler et al. identified Code Bad Smells (2000 to June 2009). We analysed in detail 39 of the most relevant papers. Our findings indicate that Duplicated Code receives most research attention, whereas some Code Bad Smells, e.g., Message Chains, receive little. This suggests that our knowledge of some Code Bad Smells remains insufficient. Our findings also show that very few studies report on the impact of using Code Bad Smells, with most studies instead focused on developing tools and methods to automatically detect Code Bad Smells. This indicates an important gap in the current knowledge of Code Bad Smells. Overall this review suggests that there is little evidence currently available to justify using Code Bad Smells. Copyright © 2010 John Wiley & Sons, Ltd. Min Zhang 0008, Tracy Hall, Nathan Baddoo |
J. Softw. Maintenance Res. Pract. | 2 |
| 2010 | Evaluating Three Approaches to Extracting Fault Data from Software Change Repositories
Tracy Hall, David Bowes, Gernot Armin Liebchen, Paul Wernick |
PROFES | 1 |
| 2010 | A Theoretical and Empirical Analysis of Three Slice-Based Metrics for CohesionabstractSound empirical research suggests that we should analyze software metrics from a theoretical and practical perspective. This paper describes the result of an investigation into the respective merits of two cohesion-based metrics for program slicing. The Tightness and Overlap metrics were those originally proposed by Weiser for the procedural paradigm. We compare and contrast these two metrics with a third metric for the OO paradigm first proposed by Counsell et al. based on Hamming Distance and based on a matrix-based notation. We theoretically validated the three metrics using the properties of Kitchenham and then empirically validated the same three metrics; some revealing properties of the metrics were found as a result. In particular, that the OO-based metric was the most stable of the three; module length was not a confounding factor for the Hamming Distance-based metric; it was however for the two slice-based metrics supporting previous work by Meyers and Binkley. The number of module slices however, was found to be an even stronger influence on the values of the two slice-based metrics, whose near perfect correlation with each other suggests that they may be measuring the same software attribute. We calculated and then compared the three metrics using first, a set of manufactured, pre-determined modules as a preliminary analysis and second, approximately nine thousand functions from the modules of multiple versions of the Barcode system, used previously by Meyers and Binkley in their empirical study. The over-arching message of the research is that a combination of theoretical and empirical analysis can help significantly in comparing the viability and indeed choice of a metric or set of metrics. More specifically, although cohesion is a subjective measure, there are certain properties of a metric that are less desirable than others and it is these 'relative' features that distinguish metrics, make their comparison possible and their value more evident. Steve Counsell, Tracy Hall, David Bowes |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2009 | Empirical Support for Two Refactoring Studies Using Commercial C# Software
Matt Gatrell, Steve Counsell, Tracy Hall |
EASE | 3 |
| 2009 | Models of motivation in software engineering
Helen Sharp, Nathan Baddoo, Sarah Beecham, Tracy Hall, Hugh Robinson |
Inf. Softw. Technol. | 4 |
| 2009 | A systematic review of theory use in studies investigating the motivations of software engineersabstractMotivated software engineers make a critical contribution to delivering successful software systems. Understanding the motivations of software engineers and the impact of motivation on software engineering outcomes could significantly affect the industry's ability to deliver good quality software systems. Understanding the motivations of people generally in relation to their work is underpinned by eight classic motivation theories from the social sciences. We would expect these classic motivation theories to play an important role in developing a rigorous understanding of the specific motivations of software engineers. In this article we investigate how this theoretical basis has been exploited in previous studies of software engineering. We analyzed 92 studies of motivation in software engineering that were published in the literature between 1980 and 2006. Our main findings are that many studies of software engineers' motivations are not explicitly underpinned by reference to the classic motivation theories. Furthermore, the findings presented in these studies are often not explicitly interpreted in terms of those theories, despite the fact that in many cases there is a relationship between those findings and the theories. Our conclusion is that although there has been a great deal of previous work looking at motivation in software engineering, the lack of reference to classic theories of motivation means that the current body of work in the area is weakened and our understanding of motivation in software engineering is not as rigorous as it may at first appear. This weakness in the current state of knowledge highlights important areas for future researchers to contribute towards developing a rigorous and usable body of knowledge in motivating software engineers. Tracy Hall, Nathan Baddoo, Sarah Beecham, Hugh Robinson, Helen Sharp |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2008 | Building a Narrative Based Requirements Engineering Mediation Model
Tracy Hall, Trevor Barker |
EuroSPI | 2 |
| 2008 | Improving the Precision of Fowler's Definitions of Bad SmellsabstractCurrent approaches to detecting bad smells in code are mainly based on software metrics. We suggest that these methods lack precision in detecting bad smells, and we propose a code pattern-based approach to detecting bad smells. However before such a pattern-based approach can be implemented, Fowler's original definitions of bad smells need to be made more precise. Currently Fowler's definitions are too informal to implement in a pattern-searching tool. In this paper we use an expert panel to evaluate our enhanced definitions for five of Fowler's bad smells. We use a questionnaire to survey four experts' opinions of our bad smell definitions. Our results show that the experts basically agree with our enhanced definitions of the message chains, middle man and speculative generality bad smells. However, there are strong disagreements on our definitions of the data clumps and switch statements bad smells. We present enhanced definitions on the basis of these expert opinions. Min Zhang 0008, Nathan Baddoo, Paul Wernick, Tracy Hall |
SEW | 4 |
| 2008 | Investigating the Role of Trust in Agile Methods Using a Light Weight Systematic Literature Review
Eisha Hasnain, Tracy Hall |
XP | 2 |
| 2008 | Motivation in Software Engineering: A systematic literature review
Sarah Beecham, Nathan Baddoo, Tracy Hall, Hugh Robinson, Helen Sharp |
Inf. Softw. Technol. | 3 |
| 2007 | Reducing Regression Test Size by ExclusionabstractOperational software is constantly evolving. Regression testing is used to identify the unintended consequences of evolutionary changes. As most changes affect only a small proportion of the system, the challenge is to ensure that the regression test set is both safe (all relevant tests are used) and inclusive (only relevant tests are used). Previous approaches to reducing test sets struggle to find safe and inclusive tests by looking only at the changed code. We use decomposition program slicing to safely reduce the size of regression test sets by identifying those parts of a system that could not have been affected by a change; this information will then direct the selection of regression tests by eliminating tests that are not relevant to the change. The technique properly accounts for additions and deletions of code. We extend and use Rothermel and Harrold's framework for measuring the safety of regression test sets and introduce new safety and precision measures that do not require a priori knowledge of the exact number of modification-revealing tests. We then analytically evaluate and compare our techniques for producing reduced regression test sets. Keith B. Gallagher, Tracy Hall, Sue Black 0001 |
ICSM | 2 |
| 2007 | Exploring motivational differences between software developers and project managersabstractIn this paper, we describe our investigation of the motivational differences between project managers and developers. Motivation has been found to be a central factor in successful software projects. However the motivation of software engineers is generally poorly understood and previous work done in the area is thought to be largely out-of-date. We present data collected from 6 software developers and 4 project managers at a workshop we organized at the XP2006 international conference. Helen Sharp, Tracy Hall, Nathan Baddoo, Sarah Beecham |
ESEC/SIGSOFT FSE | 2 |
| 2007 | Motivating developer performance to improve project outcomes in a high maturity organization
Tracy Hall, Dorota Jagielska, Nathan Baddoo |
Softw. Qual. J. | 1 |
| 2006 | A Preliminary Empirical Investigation of the Use of Evidence Based Software Engineering by Under-graduate Students
Austen Rainer, Tracy Hall, Nathan Baddoo |
EASE | 2 |
| 2006 | Trust in software outsourcing relationships: An empirical investigation of Indian software companies
Nilay V. Oza, Tracy Hall, Austen Rainer, Susan Grey |
Inf. Softw. Technol. | 2 |
| 2005 | Using an expert panel to validate a requirements process improvement model
Sarah Beecham, Tracy Hall, Carol Britton, Michaela Cottee, Austen Rainer |
J. Syst. Softw. | 2 |
| 2005 | Defining a Requirements Process Improvement Model
Sarah Beecham, Tracy Hall, Austen Rainer |
Softw. Qual. J. | 2 |
| 2004 | The Impact of Using Pair Programming on System Evolution: A Simulation-Based StudyabstractWe investigate the impact of pair programming on the long term evolution of software systems. We use system dynamics to build simulation models which predict the trend in system growth with and without pair programming. Initial results suggest that the extra effort needed for two people to code together may generate sufficient benefit to justify pair programming. Paul Wernick, Tracy Hall |
ICSM | 2 |
| 2004 | Can Thomas Kuhn's paradigms help us understand software engineering?abstractRecent articles in EJIS have discussed whether or not Information Systems is a ‘discipline’. In The Structure of Scientific Revolutions, Kuhn states that a scientific discipline can be identified by reference to its underlying belief system, the ‘paradigm’ or ‘disciplinary matrix’, to which all workers in that field must commit. An important element of Kuhn's model is the notion of ‘scientific communities’. We consider here the belief system underlying Software Engineering (SE). We examine the extent to which a belief system analogous to the disciplinary matrix of a Kuhnian science can be identified in SE. Our preliminary fieldwork has comprised an examination of books used by SE students and practitioners, and in-depth interviews with a number of practitioners. The results of this study suggest that the current status of the theory of SE parallels Kuhn's ‘pre-paradigm’ stage of scientific development. At this early stage, theorists and practitioners are divided into schools. These schools are based on differences in the beliefs and models forming their disciplinary matrices. We conclude that the application by analogy of Kuhn's view of scientific activity to SE is justifiable. Our findings can assist both SE theorists and practitioners in improving the understanding of how and why software development projects succeed or fail. Our findings also provide a framework within which to place the beliefs, models and values which underlie SE. Such a framework can contribute to the discussion as to whether the software development-related aspects of Information Systems can be considered to be a discipline, and if so how that discipline is structured. Paul Wernick, Tracy Hall |
Eur. J. Inf. Syst. | 2 |
| 2003 | Software Process Improvement Problems in Twelve Software Companies: An Empirical Analysis
Sarah Beecham, Tracy Hall, Austen Rainer |
Empir. Softw. Eng. | 2 |
| 2003 | De-motivators for software process improvement: an analysis of practitioners' views
Nathan Baddoo, Tracy Hall |
J. Syst. Softw. | 2 |
| 2003 | A quantitative and qualitative analysis of factors affecting software processes
Austen Rainer, Tracy Hall |
J. Syst. Softw. | 2 |
| 2002 | Software Process Improvement Motivators: An Analysis using Multidimensional Scaling
Nathan Baddoo, Tracy Hall |
Empir. Softw. Eng. | 2 |
| 2002 | Motivators of Software Process Improvement: an analysis of practitioners' views
Nathan Baddoo, Tracy Hall |
J. Syst. Softw. | 2 |
| 2002 | Key success factors for implementing software process improvement: a maturity-based analysis
Austen Rainer, Tracy Hall |
J. Syst. Softw. | 2 |
| 2001 | An Empirical Study of Maintenance Issues within Process Improvement Programmes in the Software IndustryabstractAnecdotal evidence from our work with software developers suggests that maintenance is a significant problem for software development companies. A problem that is absorbing increasing amounts of precious development effort. In parallel, software companies are increasingly applying process improvement principles to development problems. In this paper we discuss how maintenance is addressed in process improvement programmes. We look at how well maintenance is addressed by formal process models like CMM. We also present empirical evidence from our study of process improvement in UK software companies. Our main findings are that although developers report that maintenance is indeed a problem, it is not always their most important problem. Furthermore, our findings also suggest that companies are often not well prepared for the maintenance phase of developments and that formal process improvement models do not pay enough attention to maintenance. Tracy Hall, Austen Rainer, Nathan Baddoo, Sarah Beecham |
ICSM | 1 |
| 2001 | Ethical Issues in Software Engineering Research: A Survey of Current Practice
Tracy Hall, Valerie Flynn |
Empir. Softw. Eng. | 1 |
| 2001 | A framework for evaluation and prediction of software process improvement success
David Wilson 0001, Tracy Hall, Nathan Baddoo |
J. Syst. Softw. | 2 |
| 1999 | A Critical Analysis of Current OO Design Metrics
Tobias Mayer 0003, Tracy Hall |
Softw. Qual. J. | 2 |
| 1999 | A Framework for Improving the Requirements Engineering Process Management
Ddembe Williams, Tracy Hall |
Softw. Qual. J. | 2 |
| 1998 | Perceptions of software quality : a pilot study
David Wilson 0001, Tracy Hall |
Softw. Qual. J. | 2 |
| 1997 | Managing Software Quality, by Brian Hambling, McGraw-Hill, 1996 (Book Review)
Tracy Hall |
Softw. Test. Verification Reliab. | 1 |
| 1996 | Software quality programmes: a snapshot of theory versus reality
Tracy Hall, Norman E. Fenton |
Softw. Qual. J. | 1 |