Nigel Stanger

dblp:95/5248 · DBLP profile ↗
← Back
20ranked-venue papers
4as first author
11since 2021 · last 2025
0000-0003-3450-7443ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 17 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Comprehensive predictive analytics for collaborators' answers, code quality, and dropout: stack overflow case study
Elijah Zolduoarrati, Sherlock A. Licorish, Nigel Stanger
Empir. Softw. Eng.3
2025 Stack overflow's hidden nuances: How does zip code define user contribution?
Elijah Zolduoarrati, Sherlock A. Licorish, Nigel Stanger
J. Syst. Softw.3
2024 The need for more informative defect prediction: A systematic literature review
abstract
Software defect prediction is crucial for prioritising quality assurance tasks, however, there are still limitations to the use of defect models. For example, the outputs often do not provide the defect type, severity, or the cause of the defect. Current models are also often complex in implementation (they use low transparency classifiers such as random forest or support vector machines) and primarily output binary predic- tions. They lack directly actionable outputs, that is, outputs that provide additional information (e.g., defect severity or defect type) to aid in fixing the defect. One approach is to utilise tools of explainable AI. In order to improve current models and plan the direction for explainability in software defect prediction, we need to understand how explainable current models are. Starting from 861 papers from multiple databases, we inves- tigated a sample of 132 papers in a systematic literature review. We extracted the following information to answer our research questions: (i) information about the outputs (e.g., how informative they were) and ex- plainability methods used, (ii) how explainability and performance is mea- sured and (iii) explainability in future research. Our results were sum- marised by manually labelling the data so that trends could be analysed across selected papers, along with a thematic analysis. We found that 71% of current models used binary outputs, while 68% of models were not yet utilising any explainability techniques. Only 7% of studies considered explainability in their future research sug- gestions. There is still a lack of awareness among researchers for the need for explainability and motivation to invest further research into more explainable and more informative software defect prediction models.
Natalie Grattan, Daniel Alencar da Costa, Nigel Stanger
Inf. Softw. Technol.3
2024 How are decisions made in open source software communities? - Uncovering rationale from python email repositories
abstract
Abstract Group decision‐making (GDM) processes shape the evolution of open source software (OSS) products, thus playing an important role in the governance of open source software communities. While these GDM processes have attracted the attention of researchers, the rationale behind decisions, that is, how decisions are made that enhance the OSS, have not received much attention. This work bridges this gap by extracting these rationales from a large open source repository comprising 1.55 million emails available in Python development archives. This work makes a methodological contribution by presenting a heuristics‐based rationale extraction system called Rationale Miner that employs information retrieval, natural language processing, and heuristics‐based techniques. Using these techniques, it extracts the rationale behind specific decisions (for example, whether a new module was added based on core developer consensus or a benevolent dictator's pronouncement). This work unearths 11 such rationales behind decisions in the Python community and thus makes a knowledge contribution. It also analyzes the prevalence of these rationales across all PEPs and three sub‐types of PEPs: Process, Informational, and Standard Track PEPs. The effectiveness of our contributions has been positively evaluated using quantitative and qualitative approaches (e.g., comparison against baselines for rationale identification showed up to 47% improvement in the most conservative case, and feedback from the Python steering committee showed the accurate identification of rationales respectively). The approach proposed in this work can be used and extended to discover the rationale behind decisions that remain hidden in communication repositories of other OSS projects, which will make the decision‐making (DM) process transparent to stakeholders and encourage decision‐makers to be more accountable.
Pankajeshwara Sharma, Bastin Tony Roy Savarimuthu, Nigel Stanger
J. Softw. Evol. Process.3
2024 Harmonising Contributions: Exploring Diversity in Software Engineering through CQA Mining on Stack Overflow
abstract
The need for collective intelligence in technology means that online Q&A platforms, such as Stack Overflow and Reddit, have become invaluable in building the global knowledge ecosystem. Despite literature demonstrating a prevalence of inclusion and contribution disparities in online communities, studies investigating the underlying reasons behind such fluctuations remain scarce. The current study examines Stack Overflow users’ contribution profiles, both in isolation and relative to various diversity metrics, including GDP and access to electricity. This study also examines whether such profiles propagate to the city and state levels, supplemented by granular data such as per capita income and education, before validating quantitative findings using content analysis. We selected 143 countries and compared the profiles of their respective users to assess implicit diversity-related complications that impact how users contribute. Results show that countries with high GDP, prominent R&D presence, less wealth inequality and sufficient access to infrastructure tend to have more users, regardless of their development status. Similarly, cities and states where technology is more prevalent (e.g., San Francisco and New York) have more users who tend to contribute more often. Qualitative analysis reveals distinct communication styles based on users’ locations. Urban users exhibited assertive, solution-oriented behaviour, actively sharing information. Conversely, rural users engaged through inquiries and discussions, incorporating personal anecdotes, gratitude and conciliatory language. Findings from this study may benefit scholars and practitioners, allowing them to develop sustainable mechanisms to bridge the inclusion and diversity gaps.
Elijah Zolduoarrati, Sherlock A. Licorish, Nigel Stanger
ACM Trans. Softw. Eng. Methodol.3
2023 Studying the characteristics of SQL-related development tasks: An empirical study
abstract
Abstract A key function of a software system is its ability to facilitate the manipulation of data, which is often implemented using a flavour of the Structured Query Language (SQL). To develop the data operations of software (i.e, creating, retrieving, updating, and deleting data), developers are required to excel in writing and combining both SQL and application code. The problem is that writing SQL code in itself is already challenging (e.g., SQL anti-patterns are commonplace) and combining SQL with application code (i.e., for SQL development tasks) is even more demanding. Meanwhile, we have little empirical understanding regarding the characteristics of SQL development tasks. Do SQL development tasks typically need more code changes? Do they typically have a longer time-to-completion? Answers to such questions would prepare the community for the potential challenges associated with such tasks. Our results obtained from 20 Apache projects reveal that SQL development tasks have a significantly longer time-to-completion than SQL-unrelated tasks and require significantly more code changes. Through our qualitative analyses, we observe that SQL development tasks require more spread out changes, effort in reviews and documentation. Our results also corroborate previous research highlighting the prevalence of SQL anti-patterns. The software engineering community should make provision for the peculiarities of SQL coding, in the delivery of safe and secure interactive software.
Daniel Alencar da Costa, Natalie Grattan, Nigel Stanger, Sherlock A. Licorish
Empir. Softw. Eng.3
2023 Secondary studies on human aspects in software engineering: A tertiary study
Elijah Zolduoarrati, Sherlock A. Licorish, Nigel Stanger
J. Syst. Softw.3
2022 Impact of individualism and collectivism cultural profiles on the behaviour of software developers: A study of stack overflow
Elijah Zolduoarrati, Sherlock A. Licorish, Nigel Stanger
J. Syst. Softw.3
2022 Unearthing open source decision-making processes: A case study of python enhancement proposals
abstract
Abstract Good governance practices are pivotal to the success of Open Source Software (OSS) projects. However, the decision‐making processes that are made available to stakeholders are at times incomplete and may remain buried and hidden in large amounts of software repository data. This work bridges this gap by unearthing enacted decision‐making processes available for Python Enhancement Proposals (PEPs) from 1.54 million email messages that embody decisions made during the evolution of the Python language. This work employs a design science approach in operationalizing a framework calledDeMaP minerthat is used to discover hidden processes using information retrieval and information extraction techniques. It also uses process mining techniques to visualize the processes, and comparative structural analysis techniques to compare different decision processes. The work identifies a richer set of decision‐making activities than those reported on the Python website and in prior research work (48 new decision activities, 199 new pathways and 6 new stages). The extracted decision process has been positively evaluated by a prominent member of the Python steering council. The extracted process can be used for process compliance checking and process improvement in OSS communities. Additionally, the DeMaP Miner framework can be extended and customized to suit other OSS projects, such as the OpenJDK project.
Pankajeshwara Sharma, Bastin Tony Roy Savarimuthu, Nigel Stanger
Softw. Pract. Exp.3
2021 Influence of Roles in Decision-Making during OSS Development - A Study of Python
abstract
Governance has been highlighted as a key factor in the success of an Open Source Software (OSS) project. It is generally seen that in a mixed meritocracy and autocracy governance model, the decision-making (DM) responsibility regarding what features are included in the OSS is shared among members from select roles; prominently the project leader. However, less examination has been made whether members from these roles are also prominent in DM discussions and how decisions are made, to show they play an integral role in the success of the project. We believe that to establish their influence, it is necessary to examine not only discussions of proposals in which the project leader makes the decisions, but also those where others make the decisions. Therefore, in this study, we examine the prominence of members performing different roles in: (i) making decisions, (ii) performing certain social roles in DM discussions (e.g., discussion starters), (iii) contributing to the OSS development social network through DM discussions, and (iv) how decisions are made under both scenarios. We examine these aspects in the evolution of the well-known Python project. We carried out a data-driven longitudinal study of their email communication spanning 20 years, comprising about 1.5 million emails. These emails contain decisions for 466 Python Enhancement Proposals (PEPs) that document the language’s evolution. Our findings make the influence of different roles transparent to future (new) members, other stakeholders, and more broadly, to the OSS research community.
Pankajeshwara Sharma, Bastin Tony Roy Savarimuthu, Nigel Stanger
EASE3
2021 Extracting Rationale for Open Source Software Development Decisions - A Study of Python Email Archives
abstract
A sound Decision-Making (DM) process is key to the successful governance of software projects. In many Open Source Software Development (OSSD) communities, DM processes lie buried amongst vast amounts of publicly available data. Hidden within this data lie the rationale for decisions that led to the evolution and maintenance of software products. While there have been some efforts to extract DM processes from publicly available data, the rationale behind 'how' the decisions are made have seldom been explored. Extracting the rationale for these decisions can facilitate transparency (by making them known), and also promote accountability on the part of decision-makers. This work bridges this gap by means of a large-scale study that unearths the rationale behind decisions from Python development email archives comprising about 1.5 million emails. This paper makes two main contributions. First, it makes a knowledge contribution by unearthing and presenting the rationale behind decisions made. Second, it makes a methodological contribution by presenting a heuristics-based rationale extraction system called Rationale Miner that employs multiple heuristics, and follows a data-driven, bottom-up approach to infer the rationale behind specific decisions (e.g., whether a new module is implemented based on core developer consensus or benevolent dictator's pronouncement). Our approach can be applied to extract rationale in other OSSD communities that have similar governance structures.
Pankajeshwara Sharma, Bastin Tony Roy Savarimuthu, Nigel Stanger
ICSE3
2020 Mining Decision-Making Processes in Open Source Software Development: A Study of Python Enhancement Proposals (PEPs) using Email Repositories
abstract
Open source software (OSS) communities are often able to produce high quality software comparable to proprietary software. The success of an open source software development (OSSD) community is often attributed to the underlying governance model, and a key component of these models is the decision-making (DM) process. While there have been studies on the decision-making processes publicized by OSS communities (e.g., through published process diagrams), little has been done to study decision-making processes that can be extracted using a bottom-up, data-driven approach, which can then be used to assess whether the publicized processes conform to the extracted processes. To bridge this gap, we undertook a large-scale data-driven study to understand how decisions are made in an OSSD community, using the case study of Python Enhancement Proposals (PEPs), which embody decisions made during the evolution of the Python language. Our main contributions are:
Pankajeshwara Sharma, Bastin Tony Roy Savarimuthu, Nigel Stanger
EASE3
2018 Linking User Requests, Developer Responses and Code Changes: Android OS Case Study
abstract
Since software systems are designed to satisfy customers' needs, developers have an obligation to address users' requirements and demands logged via issue trackers and other forums. Having to respond to a large number of requests while developing and perfecting systems presents prioritization challenges, however. Android Operating System (OS) developers have largely overcome this obstacle by responding to specific user requests, which may be traced back to actual software code changes, providing lessons for the software engineering community. This study applies text and data mining techniques to investigate the Android community as an ecosystem, exploring how developers responded to issues raised by the community over several versions of the OS. Results show a strong relationship between issues raised by the community and developer responses to these issues. This relationship also extended to actual source code changes made by developers. Furthermore, the findings show a correlation between user requests and developer responses enacted via code changes across specific Android versions and important functionalities. This evidence suggests that developers have invested in the Android platform to guarantee its survival and overall success, largely through addressing user demands. We outline implications for software engineering professionals and software systems success.
Sherlock A. Licorish, Elijah Zolduoarrati, Nigel Stanger
EASE3
2018 Semi-Automated Assessment of SQL Schemas via Database Unit Testing
Nigel Stanger
ICCE1
2017 Boundary Spanners in Open Source Software Development: A Study of Python Email Archives
abstract
In many open source software development communities, a significant proportion of development is undertaken by a relatively small number of individuals, the "core members". The stability and longevity of this group of most active developers are crucial for the success of the project. While there has been prior work on identifying key individuals in open source development, little attention has been devoted to the identification of cross-cutting core individuals (boundary spanners) whose responsibilities span across different functional areas of open source development (e.g., who are involved both in development-centric activities and user-centric activities). To address this gap, we propose an approach to identify the core cross-cutting members and their roles within the community through analyzing email communication repositories. We use Social Network Analysis (SNA) tools to identify the most active core members in different forums (that have different focus such as Python-dev that focuses on language evolution and Python Lists that focus on user support), and their activities over time, thus identifying the core developers and their involvement in different community mailing lists. Based on the involvement of a core developer and the overall social structure of the network of core developers, we also present an approach for identifying a potential replacement for a community administrator that steps down. Using email repositories of six main Python forums as the case study domain, we computed several social network analysis metrics to characterize the core developers and their importance in the Python community.
Pankajeshwara Sharma, Bastin Tony Roy Savarimuthu, Nigel Stanger
APSEC3
2017 Investigating developers' email discussions during decision-making in Python language evolution
abstract
Context: Open Source Software (OSS) developers use mailing lists as their main forum for discussing the evolution of a project. However, the use of mailing lists by developers for decision-making has not received much research attention. Objective: We have explored this issue by studying developers' email discussions around Python Enhancement Proposals (PEPs). Method: Our dataset comprised 42,672 emails from six different mailing lists pertaining to PEP development. We performed multiple forms of analysis on these emails, involving both quantitative measures (e.g., frequency) and deeper analysis of specific PEP discussions (i.e., outlier analysis). Results: Out of three PEP types (Informational, Process and Standard Track), Standard Track PEPs attract a large amount of discussion (both in volume and average number of messages per proposal). Our study also identified specific PEP states and topics that generated a disproportionate amount of discussion. Conclusion: Our outcomes point to several opportunities for improving the management of an OSS team based on the knowledge generated from discussions. We have also identified several interesting avenues for future work such as identifying individuals or groups that present persuasive arguments during decision-making.
Pankajeshwara Sharma, Bastin Tony Roy Savarimuthu, Nigel Stanger, Sherlock A. Licorish, Austen Rainer
EASE3
2013 Context identification of sentences in research articles: Towards developing intelligent tools for the research community
abstract
Abstract Scientific literature is an important medium for disseminating scientific knowledge. However, in recent times, a dramatic increase in research output has resulted in challenges for the research community. An increasing need is felt for tools that exploit the full content of an article and provide insightful services with value beyond quantitative measures such as impact factors and citation counts. However, the intricacies of language and thought, and the unstructured format of research articles present challenges in providing such services. The identification of sentence contexts that encode the role of specific sentences in advancing an article's scientific argument can facilitate in developing intelligent tools for the research community. This paper describes our research work in this direction. First, we investigate the possibility of identifying contexts associated with sentences and propose a scheme of thirteen context type definitions for sentences, based on the generic rhetorical pattern found in scientific articles. We then present the results of our experiments using sequential classifiers – conditional random fields – for achieving automatic context identification. We also describe our Semantic Web application developed for providing citation context based information services for the research community. Finally, we present a comparison and analysis of our results with similar studies and explain the distinct features of our application.
M. A. Angrosh, Stephen Cranefield, Nigel Stanger
Nat. Lang. Eng.3
2000 Translating descriptions of a viewpoint among different representations
abstract
An important part of the systems development process is building models of real-world phenomena. These phenomena are described by many different kinds of information, and this diversity has resulted in a wide variety of modelling representations. Different types of information are better expressed by some representations than others, so it is sensible to use multiple representations to describe a phenomenon. This paper describes an approach to facilitating the use of multiple representations within a single viewpoint by translating descriptions of the viewpoint among different representations. An important issue with such translations is their quality, or how well they map the constructs of one representation to the constructs of another. Two methods are proposed for improving translation quality: heuristics and enrichment, and a preliminary metric for measuring relative translation quality is described.
Nigel Stanger
APSEC1
2000 A Viewpoint-Based Framework for Discussing the Use of Multiple Modelling Representations
Nigel Stanger
ER1
1997 Exploiting the advantages of object oriented programming in the implementation of a database design environment
abstract
We describe the implementation of a database design environment (Swift) that incorporates several novel features. Swift's data modelling approach is derived from viewpoint-oriented methods. Swift is implemented in Java, which allows us to easily construct a client-server based environment. The repository is implemented using PostgreSQL, which allows us to store the actual application code in the database. The combination of Java and PostgreSQL reduces the impedance mismatch between the application and the repository.
Nigel Stanger, Richard Pascoe
APSEC1