EDBT 2026 Demo / reviewers in the wild / expert
Steve Counsell
dblp:93/6557 · also Stephen J. Counsell
· DBLP profile ↗
104ranked-venue papers
31as first author
16since 2021 · last 2026
0000-0002-2939-8919ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 84 · 23 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 22 · 14 first-author · 6 since 2021Artificial intelligence and machine learning · 12 · 6 first-author · 1 since 2021Databases, data management, data science and information retrieval · 8 · 2 first-authorHuman-computer interaction and ubiquitous computing · 5 · 3 first-authorComputer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An audit of machine learning experiments on software defect predictionabstractMachine learning algorithms are increasingly being proposed to solve the problem of predicting defect-prone software components. In this literature, computational experiments are the primary means of evaluating and comparing learners and the credibility of findings depends critically on their experimental design and reporting. This paper audits recent software defect prediction (SDP) experiments by assessing their experimental design, analysis and reporting practices against widely accepted norms from statistics, machine learning and empirical software engineering. Our aim is to characterise the current state of practice and evaluate the reproducibility of published findings. We undertook an audit of relevant studies published from the SCOPUS database (2019-2023) focusing on their experimental design and analysis choices e.g., the outcome variables such as F-measure and the type of out of sample (OOS) validation regime, e.g., cross-validation, plus the statistical analysis and inference mechanisms. In all, we evaluated nine different study issues. This was complemented by an assessment of reproducibility using the instrument proposed by González-Barahona and Robles. Our search located approximately 1,585 experiments in SDP (2019-2023), a substantial body of work. From this, we randomly sampled 101 ( $$ \approx 6.4\%$$ ) papers, 61 journal and 40 conference papers. Almost 50% are behind ‘paywalls’. We found considerable divergence in research practice. The number of datasets used ranged 1-365, the number of learners or learner variants evaluated from 1-34 and the number of performance metrics from 1 to 9. Approximately 45% of papers made use of formal statistical inference. We detected a total of 427 issues distributed across 101 papers (median=4) with only one paper being entirely issue-free. In terms of reproducibility, experiments ranged from near perfect to lacking almost all required information. We also found two examples of tortured phrases and potential “paper mill” activity. Approaches to designing and reporting computational experiments varied greatly, but almost half the studies provided insufficient information such that reproduction would be challenging. Overall, our audit suggests that as a research community, we have considerable scope for improvement. Fortunately, many improvements should be neither difficult nor costly to achieve. Giuseppe Destefanis, Leila Yousefi, Martin J. Shepperd, Allan Tucker, Stephen Swift, Steve Counsell, Mahir Arzoky |
Empir. Softw. Eng. | 6 |
| 2024 | Semgrep*: Improving the Limited Performance of Static Application Security Testing (SAST) ToolsabstractVulnerabilities in code should be detected and patched quickly to reduce the time in which they can be exploited. There are many automated approaches to assist developers in detecting vulnerabilities, most notably Static Application Security Testing (SAST) tools. However, no single tool detects all vulnerabilities and so relying on any one tool may leave vulnerabilities dormant in code. In this study, we use a manually curated dataset to evaluate four SAST tools on production code with known vulnerabilities. Our results show that the vulnerability detection rates of individual tools range from 11.2% to 26.5%, but combining these four tools can detect 38.8% of vulnerabilities. We investigate why SAST tools are unable to detect 61.2% of vulnerabilities and identify missing vulnerable code patterns from tool rule sets. Based on our findings, we create new rules for Semgrep, a popular configurable SAST tool. Our newly configured Semgrep tool detects 44.7% of vulnerabilities, more than using a combination of tools, and a 181% improvement in Semgrep’s detection rate. Gareth Bennett, Tracy Hall, Emily Winter 0001, Steve Counsell |
EASE | 4 |
| 2024 | Do Developers Use Static Application Security Testing (SAST) Tools Straight Out of the Box? A large-scale Empirical StudyabstractStatic application Security Testing (SAST) tools are an established means of detecting vulnerabilities early in development. Previous studies have reported low detection rates from SAST tools and recommend either combining SAST tools or configuring rule sets to detect more vulnerabilities. However, while previous work suggests that developers rarely combine or configure any of the Automatic Static Analysis Tools (ASATs) they use, it is currently unclear whether SAST tools are used directly “out of the box”. To understand how developers use SAST tools, we performed a large-scale survey involving 1,263 developers. We pre-screened developers to establish their SAST use and found that only 20% (204/1,003) used SAST tools. Of those developers who did use SAST tools, we found a large number did not use multiple tools (59%), did not configure tools (54%) or did neither (40%). Our results suggest that more work is needed to help developers combine and configure tools, since doing so is likely to detect significantly more vulnerabilities. Gareth Bennett, Tracy Hall, Steve Counsell, Emily Winter 0001, Thomas Shippey |
ESEM | 3 |
| 2024 | Different Strokes for Different Folks: A Comparison of Developer and Tester Views on TestingabstractIn this paper, we re-analyse the data from a previous study of Straubinger et al., which asked 284 industrial IT staff about their views on testing. In that study, as well as developers, the dedicated role of tester was included in the data - both roles were treated as the same role. In this paper, we posit that the two roles (i.e., developer and tester) are so very different that we should analyse each role separately; testers will have unique insights into testing, so separating their views and experiences from developers is important. To this end, we analyse six of the same research questions as the original study, using separate developer and tester data. Results showed that for almost every question we re-visited, testers differed in their opinions from developers, whether on the type of testing they did, measures of code quality, effort to write tests and motivation for testing. Steve Counsell, Stephen Swift, Mahir Arzoky, Tracy Hall, Emily Winter 0001, Gareth Bennett, Thomas Shippey |
SEAA | 1 |
| 2023 | An "80-20" Approach to the Study of CouplingabstractIn this paper, we use both afferent (i.e., incoming) and efferent (i.e., outgoing) class coupling to analyse trends in seven open-source and four industrial systems. We determine whether an 80-20 rule (or Pareto’s Law) applies to either type of coupling and whether 20% of classes contain 80% of coupling (whether incoming and outgoing). We also explore the relationship between bugs and both types of coupling. Results showed that an 80-20 rule applied overwhelmingly in the industrial systems (but not in the open-source). For the two types of coupling, the open-source systems also had a stronger, positive relationship with bugs; opposite effects were found in the industrial systems. These results highlight the need for adoption of different and impactful approaches to understanding systems which reflect on developer practice. Steve Counsell, Stephen Swift, Amjed Tahir |
SEAA | 1 |
| 2023 | The Paradox of Analysing Gender-Based DataabstractIn this short paper, we analyse "gender" perspectives from a survey of three hundred and seventy-eight industry developers on two aspects of IT industry developer practice: bugs and Automatic Program Repair. We also explore questions of how developers view their job satisfaction. Our key motivation was to show whether there was a difference in the way that males and females viewed these three important concepts. From a total of thirteen survey questions analysed, only two showed any statistical difference between the responses of females compared to males. Those differences were found exclusively in the job satisfaction part of the survey. In terms of the way that male or female developers think about technical activities per se and diversity and inclusivity more generally, we therefore have the paradoxical issue of whether gender comparisons have any basis. We all think the same way about technical-oriented activities, so perhaps we need to stop trying to find differences and division. Steve Counsell, Emily Winter 0001, Tracy Hall, Vesna Nowack |
SEAA | 1 |
| 2023 | How do Developers Really Feel About Bug Fixing? Directions for Automatic Program RepairabstractAutomatic program repair (APR) is a rapidly advancing field of software engineering that aims to supplement or replace manual bug fixing with an automated tool. For APR to be successfully adopted in industry, it is vital that APR tools respond to developer needs and preferences. However, very little research has considered developers' general attitudes to APR or developers' current bug fixing practices (the activity APR aims to replace). This paper responds to this gap by reporting on a survey of 386 software developers about their bug finding and fixing practices and experiences, and their instinctive attitudes towards APR. We find that bug finding and fixing is not necessarily as onerous for developers as has often been suggested, being rated as more satisfying than developers' general work. The fact that developers derive satisfaction and benefit from bug fixing indicates that APR adoption is not as simple as APR replacing an unwanted activity. When it comes to potential APR approaches, we find a strong preference for developers being kept in the loop (for example, choosing between different fixes or validating fixes) as opposed to a fully automated process. This suggests that advances in APR should be careful to consider the agency of the developer, as well as what information is presented to developers alongside fixes. It also indicates that there are key barriers related to trust that would need to be overcome for full scale APR adoption, supported by the fact that even those developers who stated that they were positive about APR listed several caveats and concerns. We find very few statistically significant relationships between particular demographic variables (for example, developer experience, age, education) and key attitudinal variables, suggesting that developers' instinctive attitudes towards APR are little influenced by experience level but are held widely across the developer community. Emily Winter 0001, David Bowes, Steve Counsell, Tracy Hall, Saemundur O. Haraldsson, Vesna Nowack, John R. Woodward |
IEEE Trans. Software Eng. | 3 |
| 2023 | Let's Talk With Developers, Not About Developers: A Review of Automatic Program Repair ResearchabstractAutomatic program repair (APR) offers significant potential for automating some coding tasks. Using APR could reduce the high costs historically associated with fixing code faults and deliver significant benefits to software engineering. Adopting APR could also have profound implications for software developers’ daily activities, transforming their work practices. To realise the benefits of APR it is vital that we consider how developers feel about APR and the impact APR may have on developers’ work. Developing APR tools without consideration of the developer is likely to undermine the success of APR deployment. In this paper, we critically review how developers are considered in APR research by analysing how human factors are treated in 260 studies from Monperrus’s Living Review of APR. Over half of the 260 studies in our review were motivated by a problem faced by developers (e.g., the difficulty associated with fixing faults). Despite these human-oriented motivations, fewer than 7% of the 260 studies included a human study. We looked in detail at these human studies and found their quality mixed (for example, one human study was based on input from only one developer). Our results suggest that software developers are often talkedaboutin APR studies, but are rarely talkedwith. A more comprehensive and reliable understanding of developer human factors in relation to APR is needed. Without this understanding, it will be difficult to develop APR tools and techniques which integrate effectively into developers’ workflows. We recommend a future research agenda to advance the study of human factors in APR. Emily Winter 0001, Vesna Nowack, David Bowes, Steve Counsell, Tracy Hall, Saemundur O. Haraldsson, John R. Woodward |
IEEE Trans. Software Eng. | 4 |
| 2022 | An 80-20 Analysis of Buggy and Non-buggy Refactorings in Open-Source CommitsabstractIn this short paper, we explore the Pareto principle, sometimes known as the “80-20” rule as part of the refactoring process. We explore five frequently applied refactorings, namely extract method, extract variable, rename variable, rename method and change variable type from a data set of forty open-source systems and nearly two hundred thousand refactorings. We address two key research questions. Firstly, do 80% of “buggy” refactorings (where a refactoring has induced a bug fix) arise from just 20% of commits and, secondly, does the same rule apply to “non-buggy” refactorings when applied to the same systems? To facilitate our analysis, we used refactoring and bug data from a study by Di Penta et al. Results showed that refactorings inducing bugs were clustered around a more concentrated set of commits than refactorings that did not induce bugs. One refactoring ‘change variable type’ stood out - it almost conformed to an 80-20 rule. The take-away message is, as the saying goes, that too much of a “good” thing [refactoring] could actually be a “bad” thing. Steve Counsell, Vesna Nowack, Tracy Hall, David Bowes, Saemundur O. Haraldsson, Emily Winter 0001, John R. Woodward |
SEAA | 1 |
| 2022 | Timing is Everything! A Test and Production Class View of Self-Admitted Technical Debt
Steve Counsell, Stephen Swift |
SEAA | 1 |
| 2022 | Exploring the Explicit Modelling of Bias in Machine Learning Classifiers: A Deep Multi-label ConvNet ApproachabstractThis paper addresses the problem that many machine learning classifiers make decisions based on data that are biased and can therefore result in prejudiced decisions. For example, in education (which this paper focuses on) a student may be rejected from a course based on historical decisions in the data that only exist due to historical biases in society or due to the skewed sampling of the data. Other approaches to dealing with bias in data include resampling methods (to counter imbalanced samples) and dimensionality reduction (to focus only on relevant features to the classification task). In this paper, we explore issues of modelling bias explicitly so that we can identify the types of bias and whether they are accounting for inflated predictive accuracies. In particular, we compare graphical model approaches to building classifiers, that are transparent in how they make decisions, with two forms of Deep Multi-label Convolutional Neural Networks to investigate if models can be built that maximise accuracy and minimise bias. We carry out this comparison on student entry and performance data from a higher educational institution. Mashael Al-Luhaybi, Stephen Swift, Steve Counsell, Allan Tucker |
ICMLA | 3 |
| 2022 | Towards developer-centered automatic program repair: findings from BloombergabstractThis paper reports on qualitative research into automatic program repair (APR) at Bloomberg. Six focus groups were conducted with a total of seventeen participants (including both developers of the APR tool and developers using the tool) to consider: the development at Bloomberg of a prototype APR tool (Fixie); developers’ early experiences using the tool; and developers’ perspectives on how they would like to interact with the tool in future. APR is developing rapidly and it is important to understand in greater detail developers' experiences using this emerging technology. In this paper, we provide in-depth, qualitative data from an industrial setting. We found that the development of APR at Bloomberg had become increasingly user-centered, emphasising how fixes were presented to developers, as well as particular features, such as customisability. From the focus groups with developers who had used Fixie, we found particular concern with the pragmatic aspects of APR, such as how and when fixes were presented to them. Based on our findings, we make a series of recommendations to inform future APR development, highlighting how APR tools should 'start small', be customisable, and fit with developers' workflows. We also suggest that APR tools should capitalise on the promise of repair bots and draw on advances in explainable AI. Emily Winter 0001, Vesna Nowack, David Bowes, Steve Counsell, Tracy Hall, Saemundur O. Haraldsson, John R. Woodward, Serkan Kirbas, Etienne Windels, Olayori McBello, Abdurahman Atakishiyev, Kevin Kells, Matthew W. Pagano |
ESEC/SIGSOFT FSE | 4 |
| 2022 | Code smells detection via modern code review: a study of the OpenStack and Qt communities
Amjed Tahir, Peng Liang 0001, Steve Counsell, Kelly Blincoe, Bing Li 0010, Yajing Luo |
Empir. Softw. Eng. | 4 |
| 2021 | Are 20% of Classes Responsible for 80% of Refactorings?abstractThe 80-20 rule is well-known in the real-world. When applied to bugs, it suggests that 80% of bugs arise in just 20% of classes. One research question that has yet to be explored is whether the same rule applies to refactoring activity. In other words, do 20% of classes account for 80% of refactorings applied to a system? In this short paper, we explore this question using data from seven open-source systems drawn from two previous studies. In each case, we explore whether the 80-20 rule applies and suggest why. Results showed limited evidence of an 80-20 rule; in the two systems where it was evident, the refactoring profile implied firstly, a large-scale movement of class fields and methods and, secondly, the deliberate aim of collapsing the class hierarchy using inheritance-based refactorings. Steve Counsell, Robert M. Hierons, Krishna Patel |
SEAA | 1 |
| 2021 | Expanding Fix Patterns to Enable Automatic Program RepairabstractAutomatic Program Repair (APR) has been proposed to help developers and reduce the time spent repairing programs. Recent APR tools have applied learned templates (fix patterns) to fix code using knowledge from fixes successfully applied in the past. However, there is still no general agreement on the representation of fix patterns, making their application and comparison with a baseline difficult. As a consequence, it is also difficult to expand fix patterns and further enable APR. We automatically generate fix patterns from similar fixes and compare the generated fix patterns against a state-of-the-art taxonomy. Our automated approach splits fixes into smaller, method-level chunks and calculates their similarity. A threshold-based clustering algorithm groups similar chunks and finds matches with state-of-the-art fix patterns. In our evaluation, we present 33 clusters whose fix patterns were generated from the fixes of 835 Defects4J bugs. Of those 33 clusters, 22 matched a state-of-the-art taxonomy with good agreement. The remaining 11 clusters were thematically analysed and generated new fix patterns that expanded the taxonomy. Our new fix patterns should enable APR researchers and practitioners to expand their tools to fix a greater range of bugs in the future. Vesna Nowack, David Bowes, Steve Counsell, Tracy Hall, Saemundur O. Haraldsson, Emily Winter 0001, John R. Woodward |
ISSRE | 3 |
| 2021 | Understanding Code Smell Detection via Code Review: A Study of the OpenStack CommunityabstractCode review plays an important role in software quality control. A typical review process would involve a careful check of a piece of code in an attempt to find defects and other quality issues/violations. One type of issues that may impact the quality of the software is code smells - i.e., bad programming practices that may lead to defects or maintenance issues. Yet, little is known about the extent to which code smells are identified during code reviews. To investigate the concept behind code smells identified in code reviews and what actions reviewers suggest and developers take in response to the identified smells, we conducted an empirical study of code smells in code reviews using the two most active OpenStack projects (Nova and Neutron). We manually checked 19,146 review comments obtained by keywords search and random selection, and got 1,190 smell-related reviews to study the causes of code smells and actions taken against the identified smells. Our analysis found that 1) code smells were not commonly identified in code reviews, 2) smells were usually caused by violation of coding conventions, 3) reviewers usually provided constructive feedback, including fixing (refactoring) recommendations to help developers remove smells, and 4) developers generally followed those recommendations and actioned the changes. Our results suggest that 1) developers should closely follow coding conventions in their projects to avoid introducing code smells, and 2) review-based detection of code smells is perceived to be a trustworthy approach by developers, mainly because reviews are context-sensitive (as reviewers are more aware of the context of the code given that they are part of the project's development team). Amjed Tahir, Peng Liang 0001, Steve Counsell, Yajing Luo |
ICPC | 4 |
| 2020 | Using the Lexicon from Source Code to Determine Application DomainabstractContext: The vast majority of software engineering research is reported independently of the application domain: techniques and tools usage is reported without any domain context. As reported in previous research, this has not always been so: early in the computing era, the research focus was frequently application domain specific (for example, scientific and data processing). Andrea Capiluppi, Nemitari Ajienka, Nour Ali, Mahir Arzoky, Steve Counsell, Giuseppe Destefanis, Alina Dana Miron, Bhaveet Nagaria, Rumyana Neykova, Martin J. Shepperd, Stephen Swift, Allan Tucker |
EASE | 5 |
| 2020 | Themes and Difficulties in Distributed Agile Email Activity: A Qualititative Team-Based StudyabstractIn a previous study by Niinimaki, the main use of emails in a distributed setting was found to be for sending non-urgent, group-wide information. In this paper, we delve deeper into this question and examine interview text from seven industrial development staff in a distributed agile setting to explore the underlying rationale for using email. Exploring communication and co-ordination patterns between teams was the chief motivation for the interviews and study. Results showed that while, in some cases, email was indeed used for team-wide communication (in support of the earlier work) a number of other uses and a set of email 'themes' emerged. We examine these themes and compare them with a set of fourteen communication difficulties associated with GSD listed by Monasor and detailed in a Systematic Literature Review. Preliminary findings from our study suggest further that email reflects a microcosm of many of the listed difficulties, but only for half of the total. The overarching conclusion is that, while email might be perceived as relatively unimportant to GSD activity, it is often symptomatic of larger (and recognized) issues and challenges often underpinned by the personalities involved. Steve Counsell, Sunila Modi, Pamela Abbott |
EASE | 1 |
| 2020 | On Clones and Comments in Production and Test Classes: An Empirical Study
Steve Counsell, Steve Swift, Mahir Arzoky, Giuseppe Destefanis |
PROFES | 1 |
| 2020 | On the Link Between Refactoring Activity and Class Cohesion Through the Prism of Two Cohesion-Based MetricsabstractThe practice of refactoring has evolved over the past thirty years to become standard developer practice; for almost the same amount of time, proposals for measuring object-oriented cohesion have also been suggested. Yet, we still know very little about their inter-relationship empirically, despite the fact that classes exhibiting low cohesion would be strong candidates for refactoring. In this paper, we use a large set of refactorings to understand the characteristics of two cohesion metrics from a refactoring perspective. Firstly, through the well-known LCOM metric of Chidamber and Kemerer and, secondly, the C3 metric proposed more recently by Marcus et al. Our research question is motivated by the premise that different refactorings will be applied to classes with low cohesion compared with those applied to classes with high cohesion. We used three open-source systems as a basis of our analysis and on data from the lower and upper quartiles of metric data. Results showed that the set of refactoring types across both upper and lower quartiles was broadly the same, although very different in actual numbers. The `rename method' refactoring stood out from the rest, being applied over three times as often to classes with low cohesion than to classes with high cohesion. Steve Counsell, Giuseppe Destefanis, Steve Swift, Mahir Arzoky, Davide Taibi 0001 |
QRS | 1 |
| 2020 | A large scale study on how developers discuss code smells and anti-pattern in Stack Exchange sites
Amjed Tahir, Jens Dietrich 0001, Steve Counsell, Sherlock A. Licorish, Aiko Fallas Yamashita |
Inf. Softw. Technol. | 3 |
| 2020 | The effect of multiple developers on structural attributes: A Study based on java software
Andrea Capiluppi, Nemitari Ajienka, Steve Counsell |
J. Syst. Softw. | 3 |
| 2019 | Predicting Academic Performance: A Bootstrapping Approach for Learning Dynamic Bayesian Networks
Mashael Al-Luhaybi, Leila Yousefi, Stephen Swift, Steve Counsell, Allan Tucker |
AIED (1) | 4 |
| 2019 | An Empirical Study of the AGIS Visual Field Metric and Its Seasonal VariationsabstractThe severity of the glaucoma eye disease is usually measured by the Advanced Glaucoma Intervention Studies (AGIS) metric. The metric provides a value between zero and twenty inclusive, where the former represents no evidence of glaucoma and the latter the most advanced of measurements. In a previous study by Montolio et al., the season in which the test was undertaken was shown to affect the value of eye measurements; the lowest sensitivity was found in Summer and the highest sensitivity found in Winter and Spring. In this paper, we partially replicate that study with a different set of data from 2468 patients obtained from Moorfields Eye Hospital, London. We decomposed the data according to the four seasonal dates to determine if extra sensitivity meant that patients' results improved in Winter. Steve Counsell, Stephen Swift, Mahir Arzoky, Giuseppe Destefanis |
CBMS | 1 |
| 2019 | On the Relationship Between Coupling and Refactoring: An Empirical ViewpointabstractBackground: Refactoring has matured over the past twenty years to become part of a developer's toolkit. However, many fundamental research questions still remain largely unexplored. Aim: The goal of this paper is to investigate the highest and lowest quartile of refactoring-based data using two coupling metrics - the Coupling between Objects metric and the more recent Conceptual Coupling between Classes metric to answer this question. Can refactoring trends and patterns be identified based on the level of class coupling? Method: In this paper, we analyze over six thousand refactoring operations drawn from releases of three open-source systems to address one such question. Results: Results showed no meaningful difference in the types of refactoring applied across either lower or upper quartile of coupling for both metrics; refactorings usually associated with coupling removal were actually more numerous in the lower quartile in some cases. A lack of inheritance-related refactorings across all systems was also noted. Conclusions: The emerging message (and a perplexing one) is that developers seem to be largely indifferent to classes with high coupling when it comes to refactoring types - they treat classes with relatively low coupling in almost the same way. Steve Counsell, Mahir Arzoky, Giuseppe Destefanis, Davide Taibi 0001 |
ESEM | 1 |
| 2019 | The Prevalence of Errors in Machine Learning Experiments
Martin J. Shepperd, Ning Li 0022, Mahir Arzoky, Andrea Capiluppi, Steve Counsell, Giuseppe Destefanis, Stephen Swift, Allan Tucker, Leila Yousefi |
IDEAL (1) | 6 |
| 2019 | A comparison and evaluation of variants in the coupling between objects metric
Mike Child, Peter Rosner, Steve Counsell |
J. Syst. Softw. | 3 |
| 2019 | Special issue on evaluation and assessment in software engineering
Emilia Mendes, Nauman Bin Ali, Steve Counsell, Maria Teresa Baldassarre |
J. Syst. Softw. | 3 |
| 2018 | Can you tell me if it smells?: A study on how developers discuss code smells and anti-patterns in Stack OverflowabstractThis paper investigates how developers discuss code smells and anti-patterns over Stack Overflow to understand better their perceptions and understanding of these two concepts. Understanding developers' perceptions of these issues are important in order to inform and align future research efforts and direct tools vendors in the area of code smells and anti-patterns. In addition, such insights could lead the creation of solutions to code smells and anti-patterns that are better fit to the realities developers face in practice. We applied both quantitative and qualitative techniques to analyse discussions containing terms associated with code smells and anti-patterns. Our findings show that developers widely use Stack Overflow to ask for general assessments of code smells or anti-patterns, instead of asking for particular refactoring solutions. An interesting finding is that developers very often ask their peers 'to smell their code' (i.e., ask whether their own code 'smells' or not), and thus, utilize Stack Overflow as an informal, crowd-based code smell/anti-pattern detector. We conjecture that the crowd-based detection approach considers contextual factors, and thus, tends to be more trusted by developers over automated detection tools. We also found that developers often discuss the downsides of implementing specific design patterns, and 'flag' them as potential anti-patterns to be avoided. Conversely, we found discussions on why some anti-patterns previously considered harmful should not be flagged as anti-patterns. Our results suggest that there is a need for: 1) more context-based evaluations of code smells and anti-patterns, and 2) better guidelines for making trade-offs when applying design patterns or eliminating smells/anti-patterns in industry. Amjed Tahir, Aiko Fallas Yamashita, Sherlock A. Licorish, Jens Dietrich 0001, Steve Counsell |
EASE | 5 |
| 2018 | Re-visiting a Test Taxonomy with Refactoring and Defect-fix DataabstractIn a previous empirical study by Bavota et al., multiple releases of three open-source systems reported the extent to which refactorings induced defect-fixes. In a much earlier study, van Deursen and Moonen (vD&M) provided a test taxonomy in which Fowler's seventy-two refactorings were categorized according to the post refactoring test burden of each (i.e., the changes required to unit tests after each refactoring had been undertaken). A refactoring was categorized as 'Type B' if it required no change to the original tests and 'Type E' if significant changes were necessary. In this paper, we investigate nine refactorings spread across vD&M's taxonomy and the corresponding defect-fix data provided by Bavota et al., to explore the relationship between defect-fixes due to refactoring and vD&M's taxonomy. Results showed that, in contrast to our intuition, the most defect-fix prone refactorings were of Types C and D and not, as we thought, of Type E. The 'Extract method' refactoring stood out as particularly 'defect-fix' inducing, suggesting that while it may solve one problem (i.e., in decomposing an excessively long method), it may well introduce other problems and required defect-fixes as a by-product. Steve Counsell, Stephen Swift, Roberto Tonelli, Michele Marchesi, Michael Felderer |
SEAA | 1 |
| 2018 | Code Cleaning for Software Defect Prediction: A Cautionary TaleabstractIn this paper, we describe our experience of developing a new technique to improve defect prediction (code cleaning) which performed very encouragingly on the first two systems on which we evaluated it (both systems had their origins in one company). Code cleaning also worked well on an additional open source system (Eclipse). But our code cleaning technique then performed disappointingly on all 69 subsequent open source systems on which we evaluated it. Without our round two evaluations on these 69 open source systems we would have published misleading prediction results. We discuss the need for performance evaluations to be performed on carefully selected samples of systems if reliable conclusions are to be drawn. Thomas Shippey, David Bowes, Steve Counsell, Tracy Hall |
SEAA | 3 |
| 2018 | An empirical study on the interplay between semantic coupling and co-change of software classesabstractThe evolution of software systems is an inevitable process which has to be managed effectively to enhance software quality. Change impact analysis (CIA) is a technique that identifies impact sets, i.e., the set of classes that require correction as a result of a change made to a class or artefact. These sets can also be considered as ripple effects and typically non-local: changes propagate to different parts of a system. Nemitari Ajienka, Andrea Capiluppi, Steve Counsell |
ICSE | 3 |
| 2018 | Do Developers Really Worry About Refactoring Re-test? An Empirical Study of Open-Source Systems
Steve Counsell, Stephen Swift, Mahir Arzoky, Giuseppe Destefanis |
PROFES | 1 |
| 2018 | The relationship between evolutionary coupling and defects in large industrial software (journal-first abstract)abstractIn this study, we investigate the effect of EC on the defect-proneness of large industrial software systems and explain why the effects vary. Serkan Kirbas, Bora Caglayan, Tracy Hall, Steve Counsell, David Bowes, Alper Sen 0001, Ayse Basar Bener |
SANER | 4 |
| 2018 | An empirical study on the interplay between semantic coupling and co-change of software classesabstractSoftware systems continuously evolve to accommodate new features and interoperability relationships between artifacts point to increasingly relevant software change impacts. During maintenance, developers must ensure that related entities are updated to be consistent with these changes. Studies in the static change impact analysis domain have identified that a combination of source code and lexical information outperforms using each one when adopted independently. However, the extraction of lexical information and the measure of how loosely or closely related two software artifacts are, considering the semantic information embedded in their comments and identifiers has been carried out using somewhat complex information retrieval (IR) techniques. The interplay between software semantic and change relationship strengths has also not been extensively studied. This work aims to fill both gaps by comparing the effectiveness of measuring semantic coupling of OO software classes using (i) simple identifier based techniques and (ii) the word corpora of the entire classes in a software system. Afterwards, we empirically investigate the interplay between semantic and change coupling. The empirical results show that: (1) identifier based methods have more computational efficiency but cannot always be used interchangeably with corpora-based methods of computing semantic coupling of classes and (2) there is no correlation between semantic and change coupling. Furthermore we found that (3) there is a directional relationship between the two, as over 70% of the semantic dependencies are also linked by change coupling but not vice versa. Nemitari Ajienka, Andrea Capiluppi, Steve Counsell |
Empir. Softw. Eng. | 3 |
| 2018 | The role and value of replication in empirical software engineering results
Martin J. Shepperd, Nemitari Ajienka, Steve Counsell |
Inf. Softw. Technol. | 3 |
| 2017 | A Deconstructed Replication of a Time of Test Study Using the AGIS MetricabstractIn medical practice, glaucoma severity is usually measured using the Advanced Glaucoma Intervention Studies (AGIS) metric. In a previous study [2], we replicated the work of Montolio et al., [5] and demonstrated that, for a larger dataset, time of day of test using the AGIS metric did make a difference to the measurement of glaucoma, supporting Montolio et als work. However, in our earlier study, we used the AGIS scores for both eyes combined. In this paper, we use the measurement from just one eye at a time. A dataset of 14389 left eye AGIS scores and the same number for the right eye from 2468 Moorfield Eye Hospital patients was used as the empirical basis. We then re-compared time of test results with those of Montolios study. Results revealed that using the values from just one eye (as opposed to both) may give a distorted picture of the AGIS scores; differences in the same time period were found between the two eyes. This may have implications for choice of sampling data and analysis of glaucoma using the AGIS metric. Steve Counsell, Stephen Swift, Allan Tucker |
CBMS | 1 |
| 2017 | Managing Hidden Dependencies in OO Software: A Study Based on Open Source ProjectsabstractDependency-based software change impact analysis is the domain concerned with estimating the sets of artifacts impacted by a change to a related artifact. Research has shown that analysing the various class dependency types independently will never completely reveal the impact sets. Therefore, dependency types are combined to improve the precision of estimated when compared to impact sets. Software classes can be linked in different ways; for instance semantically, if their meaning is somewhat related or, structurally, if one class depends on the services of other classes. 'Hidden' dependencies arise when two classes, linked structurally, do not share the same semantic namespace or when semantically dependent classes do not share a structural link. With the goal of revealing hidden dependencies during change impact analysis, we empirically investigated the relationship between structural and semantic class dependencies in object-oriented software systems. Results show that (i) semantic and structural links are significantly associated, (ii) the strengths of those links do not play a significant role and, (iii) a significant number of dependencies are hidden. We propose two refactoring techniques to deal with hidden dependencies, based on existing design patterns. We plan to investigate them further to assert whether either has the potential for reducing refactoring and testing effort. Nemitari Ajienka, Andrea Capiluppi, Steve Counsell |
ESEM | 3 |
| 2017 | An experimental search-based approach to cohesion metric evaluationabstractIn spite of several decades of software metrics research and practice, there is little understanding of how software metrics relate to one another, nor is there any established methodology for comparing them. We propose a novel experimental technique, based on search-based refactoring, to ‘animate’ metrics and observe their behaviour in a practical setting. Our aim is to promote metrics to the level of active, opinionated objects that can be compared experimentally to uncover where they conflict, and to understand better the underlying cause of the conflict. Our experimental approaches include semi-random refactoring, refactoring for increased metric agreement/disagreement, refactoring to increase/decrease the gap between a pair of metrics, and targeted hypothesis testing. We apply our approach to five popular cohesion metrics using ten real-world Java systems, involving 330,000 lines of code and the application of over 78,000 refactorings. Our results demonstrate that cohesion metrics disagree with each other in a remarkable 55 % of cases, that Low-level Similarity-based Class Cohesion (LSCC) is the best representative of the set of metrics we investigate while Sensitive Class Cohesion (SCOM) is the least representative, and we discover several hitherto unknown differences between the examined metrics. We also use our approach to investigate the impact of including inheritance in a cohesion metric definition and find that doing so dramatically changes the metric. Mel Ó Cinnéide, Iman Hemati Moghadam, Mark Harman, Steve Counsell, Laurence Tratt |
Empir. Softw. Eng. | 4 |
| 2017 | The relationship between evolutionary coupling and defects in large industrial softwareabstractAbstract Evolutionary coupling (EC) is defined as the implicit relationship between 2 or more software artifacts that are frequently changed together. Changing software is widely reported to be defect‐prone. In this study, we investigate the effect of EC on the defect proneness of large industrial software systems and explain why the effects vary. We analysed 2 large industrial systems: a legacy financial system and a modern telecommunications system. We collected historical data for 7 years from 5 different software repositories containing 176 thousand files. We applied correlation and regression analysis to explore the relationship between EC and software defects, and we analysed defect types, size, and process metrics to explain different effects of EC on defects through correlation. Our results indicate that there is generally a positive correlation between EC and defects, but the correlation strength varies. Evolutionary coupling is less likely to have a relationship to software defects for parts of the software with fewer files and where fewer developers contributed. Evolutionary coupling measures showed higher correlation with some types of defects (based on root causes) such as code implementation and acceptance criteria. Although EC measures may be useful to explain defects, the explanatory power of such measures depends on defect types, size, and process metrics. Serkan Kirbas, Bora Caglayan, Tracy Hall, Steve Counsell, David Bowes, Alper Sen 0001, Ayse Basar Bener |
J. Softw. Evol. Process. | 4 |
| 2016 | An Empirical Study into the Relationship Between Class Features and Test SmellsabstractWhile a substantial body of prior research has investigated the form and nature of production code, comparatively little attention has examined characteristics of test code, and, in particular, test smells in that code. In this paper, we explore the relationship between production code properties (at the class level) and a set of test smells, in five open source systems. Specifically, we examine whether complexity properties of a production class can be used as predictors of the presence of test smells in the associated unit test. Our results, derived from the analysis of 975 production class-unit test pairs, show that the Cyclomatic Complexity (CC) and Weighted Methods per Class (WMC) of production classes are strong indicators of the presence of smells in their associated unit tests. The Lack of Cohesion of Methods in a production class (LCOM) also appears to be a good indicator of the presence of test smells. Perhaps more importantly, all three metrics appear to be good indicators of particular test smells, especially Eager Test and Duplicated Code. The Depth of the Inheritance Tree (DIT), on the other hand, was not found to be significantly related to the incidence of test smells. The results have important implications for large-scale software development, particularly in a context where organizations are increasingly using, adopting or adapting open source code as part of their development strategy and need to ensure that classes and methods are kept as simple as possible. Amjed Tahir, Steve Counsell, Stephen G. MacDonell |
APSEC | 2 |
| 2016 | The AGIS Metric and Time of Test: A Replication StudyabstractVisual Field (VF) tests and corresponding data are commonly used in clinical practices to manage glaucoma. The standard metric used to measure glaucoma severity is the Advanced Glaucoma Intervention Studies (AGIS) metric. We know that time of day when VF tests are applied can influence a patient's AGIS metric value; a previous study showed that this was the case for a data set of 160 patients. In this paper, we replicate that study using data from 2468 patients obtained from Moorfields Eye Hospital. This may provide further evidence and support of this phenomenon in a replication sense. Results did indeed show a tendency for the metric to be lower for early onset patients in the morning; equally, for advanced patients, the effect was less pronounced. We thus found support for the earlier work of Montolio et al. [4] and add to the body of evidence on the AGIS metric. Steve Counsell, Stephen Swift, Allan Tucker |
CBMS | 1 |
| 2016 | So You Need More Method Level Datasets for Your Software Defect Prediction?: Voilà!abstractContext: Defect prediction research is based on a small number of defect datasets and most are at class not method level. Consequently our knowledge of defects is limited. Identifying defect datasets for prediction is not easy and extracting quality data from identified datasets is even more difficult. Goal: Identify open source Java systems suitable for defect prediction and extract high quality fault data from these datasets. Method: We used the Boa to identify candidate open source systems. We reduce 50,000 potential candidates down to 23 suitable for defect prediction using a selection criteria based on the system's software repository and its defect tracking system. We use an enhanced SZZ algorithm to extract fault information and calculate metrics using JHawk. Result: We have produced 138 fault and metrics datasets for the 23 identified systems. We make these datasets (the ELFF datasets) and our data extraction tools freely available to future researchers. Conclusions: The data we provide enables future studies to proceed with minimal effort. Our datasets significantly increase the pool of systems currently being used in defect analysis studies. Thomas Shippey, Tracy Hall, Steve Counsell, David Bowes |
ESEM | 3 |
| 2016 | Comparing Test and Production Code Quality in a Large Commercial Multicore SystemabstractA fundamental goal of software engineering practice is to ensure that code quality is maintained throughout its lifetime. Measuring and maintaining the quality of test code should be as important as measuring production (in-the-field) code. However, test code often seems to be a second class citizen compared to production code in terms of its upkeep and general maintenance. Many of the code features we might expect in test code are either absent or, included when they should not be. In this paper, we investigate four releases of an industrial embedded multi-core system from four perspectives and compare results for test code with corresponding production code. The four perspectives we considered as indicators of code quality. Firstly, we looked at whether test and production code conformed to a set of in-house designated design rules. Secondly, we explored whether test code contained a reasonable proportion of comment to code lines ratio relative to production code. Thirdly, we examined test and production code and the number of assertions in that code. Finally we investigated the relationship between faults and code features. In terms of results, test code did not fare well when compared with production code. An interesting and startling result related to the use of assertions, they were used liberally in test and production code. However, their effect, if triggered, was much larger in production code. Steve Counsell, Giuseppe Destefanis, Xiaohui Liu 0001, Sigrid Eldh, Andreas Ermedahl, Kenneth Andersson |
SEAA | 1 |
| 2016 | Arsonists or Firefighters? Affectiveness in Agile Software DevelopmentabstractIn this paper, we present an analysis of more than 500 K comments from open-source repositories of software systems developed using agile methodologies. Our aim is to empirically determine how developers interact with each other under certain psychological conditions generated by politeness, sentiment and emotion expressed within developers’ comments. Developers involved in an open-source projects do not usually know each other; they mainly communicate through mailing lists, chat, and tools such as issue tracking systems. The way in which they communicate affects the development process and the productivity of the people involved in the project. We evaluated politeness, sentiment and emotions of comments posted by agile developers and studied the communication flow to understand how they interacted in the presence of impolite and negative comments (and vice versa ). Our analysis shows that “firefighters” prevail. When in presence of impolite or negative comments, the probability of the next comment being impolite or negative is 13 % and 25 %, respectively; ANGER however, has a probability of 40 % of being followed by a further ANGER comment. The result could help managers take control the development phases of a system, since social aspects can seriously affect a developer’s productivity. In a distributed agile environment this may have a particular resonance. Marco Ortu, Giuseppe Destefanis, Steve Counsell, Stephen Swift, Roberto Tonelli, Michele Marchesi |
XP | 3 |
| 2015 | 6th International Workshop on Emerging Trends in Software Metrics (WETSoM 2015)abstractWETSoM is a gathering of researchers and practitioners to discuss the progress on software metrics knowledge. Motivations for this workshop include the low impact that software metrics have on current software development and the increased interest in research. The goals of this workshop include critically examining the evidence for the effectiveness of existing metrics and identifying new directions for metrics. Evidence for existing metrics includes how the metrics have been used in practice and studies showing their effectiveness. Identifying new directions includes use of new theories, such as complex network theory, on which to base metrics. Steve Counsell, Corrado Aaron Visaggio, Roberto Tonelli, Ewan D. Tempero |
ICSE (2) | 1 |
| 2015 | Detection of violation causes in reflexion modelsabstractReflexion Modelling is a well-understood technique to detect architectural violations that occur during software architecture erosion. Resolving these violations can be difficult when erosion has reached a critical level and the causes of the violations are interwoven and difficult to understand. This article outlines a novel technique to automatically detect typical causes of violations in reflexion models, based on the definition and detection of typical symptoms for these causes. Preliminary results show that the proposed technique can support software architects' navigation through reflexion models of eroded systems to understand causes of violations and to systematically take actions against them. Sebastian Herold, Michael English, Jim Buckley, Steve Counsell, Mel Ó Cinnéide |
SANER | 4 |
| 2015 | Would you mind fixing this issue? - An Empirical Analysis of Politeness and Attractiveness in Software Developed Using Agile Boards
Marco Ortu, Giuseppe Destefanis, Mohamad Kassab, Steve Counsell, Michele Marchesi, Roberto Tonelli |
XP | 4 |
| 2015 | Editorial for the special section on Empirical Studies in Software Engineering Selected, and extended papers from the Eighteenth International Conference on Evaluation and Assessment in Software Engineering, May 13th-14th 2014, London, UK
Tracy Hall, Steve Counsell, Ingunn Myrtveit |
Inf. Softw. Technol. | 2 |
| 2015 | The effect of refactoring on change and fault-proneness in commercial C# software
Matt Gatrell, Steve Counsell |
Sci. Comput. Program. | 2 |
| 2014 | An Approach to Controlling the Runtime for Search Based Modularisation of Sequential Source Code Check-ins
Mahir Arzoky, Stephen Swift, Steve Counsell, James Cain 0002 |
IDA | 3 |
| 2014 | Comparing Pre-defined Software Engineering Metrics with Free-Text for the Prediction of Code 'Ripples'
Steve Counsell, Allan Tucker, Stephen Swift, Guy Fitzgerald, Jason Peters |
IDA | 1 |
| 2014 | System performance analyses through object-oriented fault and coupling prismsabstractA fundamental aspect of a system's performance over time is the number of faults it generates. The relationship between the software engineering concept of "coupling" (i.e., the degree of inter-connectedness of a system's components) and faults is still a research question attracting attention and a relationship with strong implications for performance; excessive coupling is generally acknowledged to contribute to fault-proneness. In this paper, we explore the relationship between faults and coupling. Two releases from each of three open-source Eclipse projects (six releases in total) were used as an empirical basis and coupling and fault data extracted from those systems. A contrasting coupling profile between fault-free and fault-prone classes was observed and this result was statistically supported. Object-oriented (OO) classes with low values of fan-in (incoming coupling) and fan-out (outgoing coupling) appeared to support fault-free classes, while classes with high fan-out supported relatively fault-prone classes. We also considered size as an influence on fault-proneness. The study thus emphasizes the importance of minimizing coupling where possible (and particularly that of fan-out); failing to control coupling may store up problems for later in a system's life; equally, controlling class size should be a concomitant goal. Alessandro Murgia, Roberto Tonelli, Michele Marchesi, Giulio Concas, Steve Counsell, Stephen Swift |
ICPE | 5 |
| 2014 | Software Metrics in Agile Software: An Empirical Study
Giuseppe Destefanis, Steve Counsell, Giulio Concas, Roberto Tonelli |
XP | 2 |
| 2014 | Improving predictive models of glaucoma severity by incorporating quality indicatorsabstractOBJECTIVE: In this paper we present an evaluation of the role of reliability indicators in glaucoma severity prediction. In particular, we investigate whether it is possible to extract useful information from tests that would be normally discarded because they are considered unreliable. METHODS: We set up a predictive modelling framework to predict glaucoma severity from visual field (VF) tests sensitivities in different reliability scenarios. Three quality indicators were considered in this study: false positives rate, false negatives rate and fixation losses. Glaucoma severity was evaluated by considering a 3-levels version of the Advanced Glaucoma Intervention Study scoring metric. A bootstrapping and class balancing technique was designed to overcome problems related to small sample size and unbalanced classes. As a classification model we selected Naïve Bayes. We also evaluated Bayesian networks to understand the relationships between the different anatomical sectors on the VF map. RESULTS: The methods were tested on a data set of 28,778 VF tests collected at Moorfields Eye Hospital between 1986 and 2010. Applying Friedman test followed by the post hoc Tukey's honestly significant difference test, we observed that the classifiers trained on any kind of test, regardless of its reliability, showed comparable performance with respect to the classifier trained only considering totally reliable tests (p-value>0.01). Moreover, we showed that different quality indicators gave different effects on prediction results. Training classifiers using tests that exceeded the fixation losses threshold did not have a deteriorating impact on classification results (p-value>0.01). On the contrary, using only tests that fail to comply with the constraint on false negatives significantly decreased the accuracy of the results (p-value<0.01). Meaningful patterns related to glaucoma evolution were also extracted. CONCLUSIONS: Results showed that classification modelling is not negatively affected by the inclusion of less reliable tests in the training process. This means that less reliable tests do not subtract useful information from a model trained using only completely reliable data. Future work will be devoted to exploring new quantitative thresholds to ensure high quality testing and low re-test rates. This could assist doctors in tuning patient follow-up and therapeutic plans, possibly slowing down disease progression. Lucia Sacchi, Allan Tucker, Steve Counsell, David F. Garway-Heath, Stephen Swift |
Artif. Intell. Medicine | 3 |
| 2013 | Negotiating Common Ground in Distributed Agile Development: A Case Study PerspectiveabstractDistributed Agile Development is gaining prevalence in the global software engineering field. However establishing and negotiating common ground across geographical, temporal and cultural borders can be a challenging process for distributed team members. This paper reports on early findings of one case study and investigates how common ground or mutually shared understanding takes place within one globally distributed agile team. The paper presents an extended version of the 3C Collaboration model, drawing upon existing literature of raising awareness cues through the use of boundary objects. The research seeks a greater understanding of how common ground is negotiated across boundaries. The case study data was obtained from semi-structured interviews within a financial context. The findings suggest that team members use multifaceted techniques to enhance common ground for better collaborative practices to take place. Sunila Modi, Pamela Abbott, Steve Counsell |
ICGSE | 3 |
| 2013 | 4th international workshop on emerging trends in software metrics (WETSoM 2013)abstractThe International Workshop on Emerging Trends in Software Metrics aims at gathering together researchers and practitioners to discuss the progress of software metrics. The motivation for this workshop is the low impact that software metrics has on current software development. The goals of this workshop includes critically examining the evidence for the effectiveness of existing metrics and identifying new directions for metrics. Evidence for existing metrics includes how the metrics have been used in practice and studies showing their effectiveness. Identifying new directions includes use of new theories, such as complex network theory, on which to base metrics. Steve Counsell, Michele Marchesi, Ewan D. Tempero, Corrado Aaron Visaggio |
ICSE | 1 |
| 2013 | The Modelling of Glaucoma Progression through the Use of Cellular Automata
Stelios Pavlidis, Stephen Swift, Allan Tucker, Steve Counsell |
IDA | 4 |
| 2013 | Testing Real-Time Embedded Systems using Timed Automata based approaches
Mohammad Saeed Abou Trab, Michael J. Brockway, Steve Counsell, Robert M. Hierons |
J. Syst. Softw. | 3 |
| 2013 | Code smells as system-level indicators of maintainability: An empirical study
Aiko Fallas Yamashita, Steve Counsell |
J. Syst. Softw. | 2 |
| 2012 | Applying Knowledge Elicitation to Improve Web Effort Estimation: A Case StudyabstractOBJECTIVE - The objective of this paper is to describe a case study where Bayesian Networks (BNs) were used to construct an expert-based Web effort model. METHOD - We built a single-company BN model solely elicited from expert knowledge, where the domain expert was an experienced Web project manager from a small Web company in Auckland, New Zealand. This model was validated using data from 22 past finished Web projects. RESULTS - The BN model has to date been successfully used to estimate effort for numerous Web projects. CONCLUSIONS - Our results suggest that, at least for the Web Company that participated in this case study, the use of a model that allows the representation of uncertainty, inherent in effort estimation, can outperform expert-based estimates. Another nine companies have also benefited from using Bayesian Networks, with very promising results. Emilia Mendes, Manar Abu Talib, Steve Counsell |
COMPSAC | 3 |
| 2012 | Specification Mutation Analysis for Validating Timed Testing Approaches Based on Timed AutomataabstractTesting real-time systems is a non-trivial validation task, especially after adding time as a new dimension to its complexity. In previous research, we introduced a 'priority-based' approach which tested the logical and timing behaviour of real-time systems modelled formally as UPPAAL Timed Automata (UTA). In this paper, we validate the 'priority-based' approach with a comparison to four well-known timed testing approaches based on a Timed Automata (TA) formalism using Specification Mutation Analysis (SMA). We introduce a set of timed and functional mutation operators based on TA. Three case studies are used to run the mutation analysis and mutants are generated according to the proposed mutation operators. The effectiveness of timed testing approaches are determined and contrasted according to the mutation score; we show that our testing approach achieves high mutation adequacy score when compared with others. Mohammad Saeed Abou Trab, Steve Counsell, Robert M. Hierons |
COMPSAC | 2 |
| 2012 | Experimental assessment of software metrics using automated refactoringabstractA large number of software metrics have been proposed in the literature, but there is little understanding of how these metrics relate to one another. We propose a novel experimental technique, based on search-based refactoring, to assess software metrics and to explore relationships between them. Our goal is not to improve the program being refactored, but to assess the software metrics that guide the auto- mated refactoring through repeated refactoring experiments. Mel Ó Cinnéide, Laurence Tratt, Mark Harman, Steve Counsell, Iman Hemati Moghadam |
ESEM | 4 |
| 2012 | Use of General Purpose GPU Programming to Enhance the Classification of Leukaemia Blast Cells in Blood Smear Images
Stefan Skrobanski, Stelios Pavlidis, Waidah Ismail, Rosline Hassan, Steve Counsell, Stephen Swift |
IDA | 5 |
| 2012 | A meta-analysis of relationships between organizational characteristics and IT innovation adoption in organizations
Mumtaz Abdul Hameed, Steve Counsell, Stephen Swift |
Inf. Manag. | 2 |
| 2012 | A framework for pathologies of message sequence charts
Haitao Dan, Robert M. Hierons, Steve Counsell |
Inf. Softw. Technol. | 3 |
| 2012 | Region-Based RTSJ Memory Management: State of the art
Hamza Hamza, Steve Counsell |
Sci. Comput. Program. | 2 |
| 2012 | A Systematic Literature Review on Fault Prediction Performance in Software EngineeringabstractBackground: The accurate prediction of where faults are likely to occur in code can help direct test effort, reduce costs, and improve the quality of software. Objective: We investigate how the context of models, the independent variables used, and the modeling techniques applied influence the performance of fault prediction models. Method: We used a systematic literature review to identify 208 fault prediction studies published from January 2000 to December 2010. We synthesize the quantitative and qualitative results of 36 studies which report sufficient contextual and methodological information according to the criteria we develop and apply. Results: The models that perform well tend to be based on simple modeling techniques such as Naive Bayes or Logistic Regression. Combinations of independent variables have been used by models that perform well. Feature selection has been applied to these combinations when models are performing particularly well. Conclusion: The methodology used to build models seems to be influential to predictive performance. Although there are a set of fault prediction studies in which confidence is possible, more studies are needed that use a reliable methodology and which report their context, methodology, and performance comprehensively. Tracy Hall, Sarah Beecham, David Bowes, David Gray, Steve Counsell |
IEEE Trans. Software Eng. | 5 |
| 2011 | Design patterns and fault-proneness a study of commercial C# softwareabstractIn this paper, we document a study of design patterns in commercial, proprietary software and determine whether design pattern participants (i.e. the constituent classes of a pattern) had a greater propensity for faults than non-participants. We studied a commercial software system for a 24 month period and identified design pattern participants by inspecting the design documentation and source code; we also extracted fault data for the same period to determine whether those participant classes were more fault-prone than non-participant classes. Results showed that design pattern participant classes were marginally more fault-prone than non-participant classes, The Adaptor, Method and Singleton patterns were found to be the most fault-prone of thirteen patterns explored. However, the primary reason for this fault-proneness was the propensity of design classes to be changed more often than non-design pattern classes. Matt Gatrell, Steve Counsell |
RCIS | 2 |
| 2011 | A Model for Predicting Class Movement in an Inheritance HierarchyabstractIn this paper, we present an empirical study to investigate whether class movement and re-location within inheritance hierarchy can be predicted based on size, coupling and cohesion for four Java open-source systems. Our results showed that class movement may not be predicted based on coupling and cohesion, and while class size was found to be a factor that may help predict class movement, it does not per se predict class movement within an inheritance hierarchy. We found a significantly higher odds ratio for larger classes to be moved within an inheritance hierarchy than that of smaller classes, suggesting that, counter-intuitively, larger classes tend to be more susceptible to movement than smaller classes. We also found that in the four systems, while classes with high coupling, low cohesion and larger size tended to be moved within their respective inheritance hierarchy, classes with high coupling, low cohesion and relatively smaller size tended to be candidate classes for deletion. Finally, while we found that class coupling and size tended to rise as the systems evolved we found no statistical support for class cohesion to decline. Directed towards developers and project managers, the message that the research conveys is that excessive growth in class size is at the root of a class' deterioration in terms of movement; developmental controls should be exercised to avoid such growth. Emal Nasseri, Steve Counsell |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2010 | An Evolutionary Study of Fan-in and Fan-out Metrics in OSS
Asma Mubarak, Steve Counsell, Robert M. Hierons |
RCIS | 2 |
| 2010 | Non-local Choice and Implied ScenariosabstractA number of issues, such as non-local choice and implied scenarios, that arise in Message Sequence Charts (MSCs) have been investigated in the past. However, existing research on these two issues show disagreements regarding how they are related. In this paper, we analyse the relations among existing conditions for non-local choice free and Closure Conditions (CCs) for implied scenarios. On the basis of this, we propose a new definition for non-local choice and a non-local choice free condition derived from CCs of implied scenarios. Compared to existing conditions, we argue that the new condition covers more non-local choices that satisfy the informal idea of non-local choice. We formally show that the existence of non-local choices in an MSC specification results in implied scenarios and the appearance of implied scenarios according to corresponding CCs means there are non-local choices in the specification. Haitao Dan, Robert M. Hierons, Steve Counsell |
SEFM | 3 |
| 2010 | An Empirical Study of Fan-In and Fan-Out in Java OSSabstractCoupling is a well researched topic in the Object-Oriented (OO) research community and its influence on class cohesion is well understood. In this paper, we present an empirical study exploring the effect of method calling on class cohesion using two coupling metrics, namely fan-in and fan-out. Three Java, open-source systems (OSS) were used as a basis of the study. A small number of classes were found to account for the vast majority of fan-in and fan-out. We also found the impact of fan-out on class cohesion to be higher than that of fan-in. Classes containing fan-out tended to have lower cohesion than those containing fan-in. Emal Nasseri, Steve Counsell, Ewan D. Tempero |
SERA | 2 |
| 2010 | A Theoretical and Empirical Analysis of Three Slice-Based Metrics for CohesionabstractSound empirical research suggests that we should analyze software metrics from a theoretical and practical perspective. This paper describes the result of an investigation into the respective merits of two cohesion-based metrics for program slicing. The Tightness and Overlap metrics were those originally proposed by Weiser for the procedural paradigm. We compare and contrast these two metrics with a third metric for the OO paradigm first proposed by Counsell et al. based on Hamming Distance and based on a matrix-based notation. We theoretically validated the three metrics using the properties of Kitchenham and then empirically validated the same three metrics; some revealing properties of the metrics were found as a result. In particular, that the OO-based metric was the most stable of the three; module length was not a confounding factor for the Hamming Distance-based metric; it was however for the two slice-based metrics supporting previous work by Meyers and Binkley. The number of module slices however, was found to be an even stronger influence on the values of the two slice-based metrics, whose near perfect correlation with each other suggests that they may be measuring the same software attribute. We calculated and then compared the three metrics using first, a set of manufactured, pre-determined modules as a preliminary analysis and second, approximately nine thousand functions from the modules of multiple versions of the Barcode system, used previously by Meyers and Binkley in their empirical study. The over-arching message of the research is that a combination of theoretical and empirical analysis can help significantly in comparing the viability and indeed choice of a metric or set of metrics. More specifically, although cohesion is a subjective measure, there are certain properties of a metric that are less desirable than others and it is these 'relative' features that distinguish metrics, make their comparison possible and their value more evident. Steve Counsell, Tracy Hall, David Bowes |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2010 | Class movement and re-location: An empirical study of Java inheritance evolution
Emal Nasseri, Steve Counsell, Martin J. Shepperd |
J. Syst. Softw. | 2 |
| 2009 | Empirical Support for Two Refactoring Studies Using Commercial C# Software
Matt Gatrell, Steve Counsell, Tracy Hall |
EASE | 2 |
| 2009 | Does an 80: 20 rule apply to Java coupling?
Asma Mubarak, Steve Counsell, Robert M. Hierons |
EASE | 2 |
| 2009 | An Application of Intelligent Data Analysis Techniques to a Large Software Engineering Dataset
James Cain 0002, Steve Counsell, Stephen Swift, Allan Tucker |
IDA | 2 |
| 2009 | An Empirical Study of Java System Evolution at the Method LevelabstractExploring the evolution of systems can provide valuable insights into the traits of developers and inform our understanding of system dynamics. While we usually expect an object-oriented system to grow (in classes) as it ages, what are not so obvious are patterns in the evolution of specific class features. In this paper, we explore empirical traits of four Java open-source systems using data extracted by two tools and informed by a previous study of inheritance depth evolution. We analyse evolution at a lower level of granularity given by the `methods' of a class on an incremental (change per version) basis rather than absolute class size per version. Evolution at a finer-grain can identify trends not possible on a class-wide basis; the approach thus represents a `white-box' view of the investigation of evolutionary forces. Our analysis also allowed direct comparison with a set of low-level refactorings extracted by an automated tool in a previous study. Scrutiny of trends in methods was further motivated by the fact that the vast majority of refactorings apply not at the class level but at the method level. Emal Nasseri, Steve Counsell |
SERA | 2 |
| 2008 | Refactoring Steps, Java Refactorings and Empirical EvidenceabstractWhile we can determine the likely testing effort of a single refactoring through simple visual inspection, the inter-relationships between many of the seventy- two refactorings mean that a chain of refactorings and hence a chain of tests may be required for completion of each. In this paper, we establish the properties of, and the inter-relationships between, fourteen of the seventy-two refactorings described in Fowler from a testing chain perspective. We provide an empirical analysis of those refactorings and their associated testing chains. We also inform our understanding of testing effort with recourse to refactoring data from 7 Java OSS. Steve Counsell, Stephen Swift |
COMPSAC | 1 |
| 2008 | Inheritance, 'Warnings' and Potential Refactorings: An Empirical StudyabstractA recent empirical study of seven Java systems showed that approximately 96% of all classes added to those systems over the course of multiple versions were at inheritance level one and two and only 4% at all other lower levels. In this paper, we use code 'warnings' extracted by a tool to explore potential problems and benefits of system evolution according to this inheritance profile. We analyze the type of warning across single and multiple versions of three Java systems for commonalities based on information provided in the warnings. We compare the frequency of warning for classes added at different levels of the inheritance hierarchy and explore the possibilities for refactoring code on that basis. The research illustrates how tools can inform our understanding of potential problems in evolutionary code, allow us to assess the impact those problems may have and present opportunities for rectifying those problems through techniques such as refactoring. Emal Nasseri, Steve Counsell |
ICSEA | 2 |
| 2008 | Is the need to follow chains a possible deterrent to certain refactorings and an inducement to others?abstractA current and difficult challenge in the software engineering arena is assessment of code smells and subsequent re-engineering decisions. The mechanics of seventy-two individual, object-oriented refactorings are specified in the seminal text by Fowler, providing the steps that need to be undertaken to complete each. While it is relatively easy to identify dasiarelatedpsila refactorings, i.e., those that each refactoring itself directly uses as part of those mechanics, what is not so clear is the chain of required refactorings that may emerge due to these indirect (and composite) relationships. In this paper, we investigate the characteristics of fourteen of the seventy-two refactorings, identifying, for each, its related refactorings and the implications this may have for the overall time and effort required to carry out each. We supported our analysis with data from a previous empirical analysis. The key result was that refactorings inducing long chains tended to be utilized less by developers than refactorings with short chains, suggesting that complexity given by long chains may be a real consideration prior to refactoring; empirically, long chains were found to be composed of sets of smaller, inter-related refactorings. On a general note, understanding the composition of refactorings is recognized as an emerging yet under-researched area but has significant implications for the amount of effort that a developer might have to invest in any single dasiachangepsila to a system involving refactoring. Steve Counsell |
RCIS | 1 |
| 2007 | A Meta-analysis Approach to Refactoring and XPabstractThe mechanics of seventy-two different Java refactorings are described fully in Fowler's text. In the same text, Fowler describes seven categories of refactoring, into which each of the seventy-two refactorings can be placed. A current research problem in the refactoring and XP community is assessing the likely time and testing effort for each refactoring, since any single refactoring may use any number of other refactorings as part of its mechanics and, in turn, can be used by many other refactorings. In this paper, we draw on a dependency analysis carried out as part of our research in which we identify the 'Use' and 'Used By' relationships of refactorings in all seven categories. We offer reasons why refactorings in the 'Dealing with Generalisation' category seem to embrace two distinct refactoring sub-categories and how refactorings in the 'Moving Features between Objects' category also exhibit specific characteristics. In a wider sense, our meta-analysis provides a developer with concrete guidelines on which refactorings, due to their explicit dependencies, will prove problematic from an effort and testing perspective. Steve Counsell, Robert M. Hierons, George Loizou |
AICCSA | 1 |
| 2007 | Thread-Based Analysis of Sequence Diagrams
Haitao Dan, Robert M. Hierons, Steve Counsell |
FORTE | 3 |
| 2007 | A Thread-tag Based Semantics for Sequence DiagramsabstractThe sequence diagram is one of the most popular behaviour modelling languages which offers an intuitive and visual way of describing expected behaviour of object-oriented software. Much research work has investigated ways of providing a formal semantics for sequence diagrams. However, these proposed semantics may not properly interpret sequence diagrams when lifelines do not correspond to threads of controls. In this paper, we address this problem and propose a thread-tag based sequence diagram as a solution. A formal, partially ordered multiset based semantics for the thread-tag based sequence diagrams is proposed. Haitao Dan, Robert M. Hierons, Steve Counsell |
SEFM | 3 |
| 2007 | Quality of manual data collection in Java software: an empirical investigation
Steve Counsell, George Loizou, Rajaa Najjar |
Empir. Softw. Eng. | 1 |
| 2006 | The Concerns of Prototypers and Their Mitigating Practices: An Industrial Case-Study
Steve Counsell, Keith Phalp, Emilia Mendes, Stella Geddes |
PROFES | 1 |
| 2006 | The interpretation and utility of three cohesion metrics for object-oriented designabstractThe concept of cohesion in a class has been the subject of various recent empirical studies and has been measured using many different metrics. In the structured programming paradigm, the software engineering community has adopted an informal yet meaningful and understandable definition of cohesion based on the work of Yourdon and Constantine. The object-oriented (OO) paradigm has formalised various cohesion measures, but the argument over the most meaningful of those metrics continues to be debated. Yet achieving highly cohesive software is fundamental to its comprehension and thus its maintainability. In this article we subject two object-oriented cohesion metrics, CAMC and NHD, to a rigorous mathematical analysis in order to better understand and interpret them. This analysis enables us to offer substantial arguments for preferring the NHD metric to CAMC as a measure of cohesion. Furthermore, we provide a complete understanding of the behaviour of these metrics, enabling us to attach a meaning to the values calculated by the CAMC and NHD metrics. In addition, we introduce a variant of the NHD metric and demonstrate that it has several advantages over CAMC and NHD. While it may be true that a generally accepted formal and informal definition of cohesion continues to elude the OO software engineering community, there seems considerable value in being able to compare, contrast, and interpret metrics which attempt to measure the same features of software. Steve Counsell, Stephen Swift, Jason Crampton |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2005 | Towards a Taxonomy of Hypermedia and Web Application Size Metrics
Emilia Mendes, Steve Counsell, Nile Mosley |
ICWE | 2 |
| 2005 | What Formal Models Cannot Show Us: People Issues During the Prototyping Process
Steve Counsell, Keith Phalp, Emilia Mendes, Stella Geddes |
PROFES | 1 |
| 2005 | Investigating Web size metrics for early Web cost estimation
Emilia Mendes, Nile Mosley, Steve Counsell |
J. Syst. Softw. | 3 |
| 2005 | Applications of dynamic proxies in distributed environmentsabstractIn object-oriented programming (OOP), proxies are entities that act as an intermediary between client objects and target objects. Dynamic proxies can be used to construct distributed systems that support the open implementation approach and promote code reuse. The OO paradigm supports code reuse through various ways including inheritance, polymorphism and aggregation. In this paper, we adopt a definition of software reuse restricted to reuse of code components and address the question of constructing distributed systems based on dynamic proxies. Different networking techniques and programming paradigms such as Java's Remote Method Invocation (RMI), the Common Object Request Broker Architecture (CORBA) and Java Servlets are used to implement the distributed client/server architecture. Copyright © 2004 John Wiley & Sons, Ltd. Youssef Hassoun, Roger Johnson, Steve Counsell |
Softw. Pract. Exp. | 3 |
| 2004 | Design Level Hypothesis Testing Through Reverse Engineering Of Object-Oriented SoftwareabstractComprehension of an object-oriented (OO) system, its design and use of OO features such as aggregation, generalisation and other forms of association is a difficult task to undertake without the original design documentation for reference. In this paper, we describe the collection of high-level class metrics from the UML design documentation of five industrial-sized C++ systems. Two of the systems studied were libraries of reusable classes. Three hypotheses were tested between these high-level features and the low-level class features of a number of class methods and attributes in each of the five systems. A further two conjectures were then investigated to determine features of key classes in a system and to investigate any differences between library-based systems and the other systems studied in terms of coupling. Results indicated that, for the three application-based systems, no clear patterns emerged for hypotheses relating to generalisation. There was, however, a clear (positive) statistical significance for all three systems studied between aggregation, other types of association and the number of methods and attributes in a class. Key classes in the three application-based systems tended to contain large numbers of methods, attributes, and associations, significant amounts of aggregation but little inheritance. No consistent, identifiable key features could be found in the two library-based systems; both showed a distinct lack of any form of coupling (including inheritance) other than through the C++ friend facility. Steve Counsell, Peter Newson, Emilia Mendes |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2003 | Applying Intelligent Data Analysis to Coupling Relationships in Object-Oriented Software
Steve Counsell, Xiaohui Liu 0001, Rajaa Najjar, Stephen Swift, Allan Tucker |
IDA | 1 |
| 2003 | A Comparative Study of Cost Estimation Models for Web Hypermedia Applications
Emilia Mendes, Ian D. Watson, Chris Triggs, Nile Mosley, Steve Counsell |
Empir. Softw. Eng. | 5 |
| 2002 | The Application of Case-Based Reasoning to Early Web Project Cost EstimationabstractLiterature shows that over the years numerous techniques for estimating development effort have been suggested, derived from late project measures. However, to the successful management of software projects, estimates are necessary throughout the whole development life cycle. The objective is twofold. First, we describe the application of case-based reasoning (CBR) for estimating Web hypermedia development effort using measures collected at different stages in the development cycle. Second, we compare the prediction accuracy of those measures, obtained using different CBR configurations. Contrary to the expected, late measures did not show statistically significant better predictions than early measures. Emilia Mendes, Nile Mosley, Steve Counsell |
COMPSAC | 3 |
| 2002 | Evolutionary algorithms for grouping high dimensional Email data
Steve Counsell, Xiaohui Liu 0001, Janet McFall, Stephen Swift, Allan Tucker |
Intell. Data Anal. | 1 |
| 2001 | The cognitive flexibility theory0: an approach for teaching Hypermedia EngineeringabstractHypermedia engineering constitutes the employment of an engineering approach to the development of hypermedia applications. Its main teaching objectives are for students to learn what an engineering approach means and how measurement can be applied.This paper presents the application of the Cognitive Flexibility Theory as an instructional theory to teach Hypermedia Engineering principles.Early results have shown that students presented a greater learning variability (suggested by their exam marks) when exposed to the CFT as a teaching practice, compared to conventional methods. Emilia Mendes, Nile Mosley, Steve Counsell |
ITiCSE | 3 |
| 2001 | Coupling Trends in Industrial Prototyping Roles: An Empirical Investigation
Keith Phalp, Steve Counsell |
Softw. Qual. J. | 2 |
| 2000 | Use of friends in C++ software: an empirical investigation
Steve Counsell, Peter Newson |
J. Syst. Softw. | 1 |
| 2000 | Experimental assessment of the effect of inheritance on the maintainability of object-oriented systems
Rachel Harrison, Steve Counsell, Reuben V. Nithi |
J. Syst. Softw. | 2 |
| 1999 | Empirical Studies of Object-Oriented Artifacts, Methods, and Processes: State of the Art and Future Directions
Lionel C. Briand, Erik Arisholm, Steve Counsell, Frank Houdek, Pascale Thévenod-Fosse |
Empir. Softw. Eng. | 3 |
| 1998 | An Investigation into the Applicability and Validity of Object-Oriented Design Metrics
Rachel Harrison, Steve Counsell, Reuben V. Nithi |
Empir. Softw. Eng. | 2 |
| 1998 | An Evaluation of the MOOD Set of Object-Oriented Software MetricsabstractThis paper describes the results of an investigation into a set of metrics for object-oriented design, called the MOOD metrics. The merits of each of the six MOOD metrics is discussed from a measurement theory viewpoint, taking into account the recognized object-oriented features which they were intended to measure: encapsulation, inheritance, coupling, and polymorphism. Empirical data, collected from three different application domains, is then analyzed using the MOOD metrics, to support this theoretical validation. Results show that (with appropriate changes to remove existing problematic discontinuities) the metrics could be used to provide an overall assessment of a software system, which may be helpful to managers of software development projects. However, further empirical studies are needed before these results can be generalized. Richard H. Carver, Steve Counsell, Reuben V. Nithi |
IEEE Trans. Software Eng. | 2 |