VLDB 2026 Research / reviewers in the wild / expert
Gemma Catolino
dblp:200/9066
· DBLP profile ↗
40ranked-venue papers
7as first author
32since 2021 · last 2026
0000-0002-4689-3401ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 36 · 7 first-author · 28 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Tracing Stereotypes in Pre-trained Transformers: From Biased Neurons to Fairer ModelsabstractThe advent of transformer-based language models has reshaped how AI systems process and generate text. In software engineering (SE), these models now support diverse activities, accelerating automation and decision-making. Yet, evidence shows that these models can reproduce or amplify social biases, raising fairness concerns. Recent work on neuron editing has shown that internal activations in pre-trained transformers can be traced and modified to alter model behavior. Building on the concept of knowledge neurons—neurons that encode factual information—we hypothesize the existence of biased neurons that capture stereotypical associations within pre-trained transformers. Gianmario Voria, Moses Openja, Foutse Khomh, Gemma Catolino, Fabio Palomba |
MSR | 4 |
| 2026 | From values to adoption: on the role of individual cultural values on fairness toolkit adoption in software developmentabstractAbstract Fairness is a critical concern in the integration of machine learning and AI—e.g., Generative AI, LLMs, and Agents—into decision-making, yet the adoption of fairness toolkits by software practitioners remains limited. This gap hinders efforts to operationalize fairness, especially when cultural and ethical values are overlooked. Given that fairness is socially constructed, individual cultural values may significantly influence how practitioners perceive and adopt fairness tools. The objective of this study is, therefore, to investigate whether and how individual cultural values influence software practitioners’ intention to adopt and actual use of fairness toolkits. Specifically, we integrate the UTAUT2 model with Hofstede’s cultural dimensions to examine both direct and moderating cultural effects within the adoption process. A survey of 181 software professionals was conducted, and data were analyzed using Partial Least Squares Structural Equation Modeling. Findings show that cultural values—specifically Power Distance, Collectivism, and Long-Term Orientation—not only directly affect adoption intention and behavior but also moderate key relationships in the adoption process. For example, collectivist individuals were more likely to act on their intention to use fairness tools, highlighting the importance of shared team goals. Overall, our findings indicate that fairness toolkit adoption is shaped not only by technology-related perceptions but also by culturally grounded value orientations. These results provide actionable insights for promoting fairness tool adoption through culturally aware strategies in software development environments. Stefano Lambiase, Gianmario Voria, Maria Concetta Schiavone, Gemma Catolino, Fabio Palomba |
Empir. Softw. Eng. | 4 |
| 2026 | Fairness set and forgotten: Mining fairness toolkit usage in open-source machine learning projectsabstractThe development of machine learning (ML) systems in high-stakes domains has amplified concerns about fairness, prompting the creation of fairness toolkits offering metrics and mitigation techniques. Open-source software (OSS) ecosystems, a critical driver of AI innovation, present a unique opportunity to study the practical adoption of these toolkits. This paper aims to empirically characterize the adoption of fairness toolkits in OSS ML projects by investigating for what purposes they are used and how their usage evolves over time. We conducted a mining study on GitHub repositories related to real-world ML projects that integrate fairness toolkits such as AIF360 and Fairlearn . Starting from 1,096 candidate repositories, we applied systematic filtering to identify a final dataset of 20 relevant ML projects (comprising 5,777 total commits). We analyzed toolkit usage by examining invoked APIs and commit histories to uncover patterns of adoption and evolution. Our findings reveal that fairness toolkits are predominantly used for diagnostic purposes, with analytic components integrated early in the project lifecycle and rarely modified thereafter. In contrast, mitigation techniques are infrequently adopted, tend to appear later, and exhibit short, unstable lifespans. Our results show that the adoption of fairness toolkits in OSS ML projects is limited and often restricted to initial diagnostic phases, with active mitigation practices remaining rare. These findings highlight the need for improved support to foster more sustained and effective integration of fairness practices within open-source development. Alfonso Cannavale, Gianmario Voria, Antonio Scognamiglio, Giammaria Giordano, Gemma Catolino, Fabio Palomba |
Inf. Softw. Technol. | 5 |
| 2026 | Fair and square? Evaluating fairness of LLM-generated synthetic datasetsabstractContext: Machine Learning (ML) is driving advancements across various industries, including healthcare, finance, and entertainment, but it also raises significant ethical concerns, particularly regarding fairness. Biases in training data can lead to unfair outcomes, perpetuating or even amplifying existing disparities. Prior research in the Software Engineering (SE) and ML communities has developed numerous bias mitigation techniques, yet two key limitations persist: (1) most approaches intervene at later stages of development, such as after data collection or model training, rather than addressing fairness from the outset; and (2) these methods often mitigate bias without fully eliminating it, since the root issue frequently lies in the data itself. Objective: In this paper, we explore an alternative approach to mitigate unfairness: synthetic data generation , which involves creating artificial datasets that mimic the statistical properties of real-world data. We aim to assess how this approach can contribute to generating data that positively impacts the trade-off between performance and fairness by creating datasets that reduce the influence of real-world biases through synthetic feature generation. Methods: To this end, we conducted an empirical study comparing ML models trained on synthetic datasets generated by large language models to ML models trained on real-world data, evaluating performance and fairness indicators. Results: Our results demonstrate that models trained with synthetic data, particularly those generated using simpler prompts, can achieve competitive performance while enhancing fairness. Conclusion: Our work suggests that synthetic data generation may be a viable approach to addressing fairness requirements in ML systems. Gianmario Voria, Benedetto Scala, Leopoldo Todisco, Carlo Venditto, Giammaria Giordano, Gemma Catolino, Fabio Palomba |
Inf. Softw. Technol. | 6 |
| 2026 | Investigating the Role of Cultural Values in Adopting Large Language Models for Software EngineeringabstractAs a socio-technical activity, software development involves the close interconnection of people and technology. The integration of Large Language Models (LLMs) into this process exemplifies the socio-technical nature of software development. Although LLMs influence the development process, software development remains fundamentally human-centric, necessitating an investigation of the human factors in this adoption. Thus, with this study we explore the factors influencing the adoption of LLMs in software development, focusing on the role of professionals’ cultural values. Guided by the Unified Theory of Acceptance and Use of Technology (UTAUT2) and Hofstede’s cultural dimensions, we hypothesized that cultural values moderate the relationships within the UTAUT2 framework. Using Partial Least Squares-Structural Equation Modelling and data from 188 software engineers, we found that habit and performance expectancy are the primary drivers of LLM adoption, while cultural values do not significantly moderate this process. These findings suggest that, by highlighting how LLMs can boost performance and efficiency, organizations can encourage their use, no matter the cultural differences. Practical steps include offering training programs to demonstrate LLM benefits, creating a supportive environment for regular use, and continuously tracking and sharing performance improvements from using LLMs. Stefano Lambiase, Gemma Catolino, Fabio Palomba, Filomena Ferrucci, Daniel Russo 0002 |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2025 | How Do Communities of ML-Enabled Systems Smell? A Cross-Sectional Study on the Prevalence of Community SmellsabstractEffective software development relies on managing both collaboration and technology, but socio-technical challenges can harm team dynamics and increase technical debt. Although teams working on ML-enabled systems are interdisciplinary, research has largely focused on technical issues, leaving their socio-technical dynamics underexplored. This study aims to address this gap by examining the prevalence, evolution, and interrelations of “community smells”, in open-source ML projects. We conducted an empirical study on 188 repositories from the NICHE dataset using the CADOCS tool to identify and analyze community smells. Our analysis focused on their prevalence, interrelations, and temporal variations. We found that certain smells—such as Prima Donna Effects and Sharing Villainy—are more prevalent and fluctuate over time compared to others like Radio Silence or Organizational Skirmish. These insights might provide valuable support for ML project managers in addressing socio-technical issues and improving team coordination. Giusy Annunziata, Stefano Lambiase, Fabio Palomba, Gemma Catolino, Filomena Ferrucci |
EASE | 4 |
| 2025 | Fairness on a budget, across the board: A cost-effective evaluation of fairness-aware practices across contexts, tasks, and sensitive attributesabstractContext: Machine Learning (ML) is widely used in critical domains like finance, healthcare, and criminal justice, where unfair predictions can lead to harmful outcomes. Although bias mitigation techniques have been developed by the Software Engineering (SE) community, their practical adoption is limited due to complexity and integration issues. As a simpler alternative, fairness-aware practices, namely conventional ML engineering techniques adapted to promote fairness, e.g., MinMax Scaling, which normalizes feature values to prevent attributes linked to sensitive groups from disproportionately influencing predictions, have recently been proposed, yet their actual impact is still unexplored. Objective: Building on our prior work that explored fairness-aware practices in different contexts, this paper extends the investigation through a large-scale empirical study assessing their effectiveness across diverse ML tasks, sensitive attributes, and datasets belonging to specific application domains. Methods: We conduct 5940 experiments, evaluating fairness-aware practices from two perspectives: contextual bias mitigation and cost-effectiveness . Contextual evaluation examines fairness improvements across different ML models, sensitive attributes, and datasets. Cost-effectiveness analysis considers the trade-off between fairness gains and performance costs. Results: Findings reveal that the effectiveness of fairness-aware practices depends on specific contexts’ datasets and configurations, while cost-effectiveness analysis highlights those that best balance ethical gains and efficiency. Conclusion: These insights guide practitioners in choosing fairness-enhancing practices with minimal performance impact, supporting ethical ML development. Alessandra Parziale, Gianmario Voria, Giammaria Giordano, Gemma Catolino, Gregorio Robles, Fabio Palomba |
Inf. Softw. Technol. | 4 |
| 2025 | Fairness-aware practices from developers' perspective: A surveyabstractMachine Learning (ML) technologies have shown great promise in many areas, but when used without proper oversight, they can produce biased results that discriminate against historically underrepresented groups. In recent years, the software engineering research community has contributed to addressing the need for ethical machine learning by proposing a number of fairness-aware practices, e.g., fair data balancing or testing approaches, that may support the management of fairness requirements throughout the software lifecycle. Nonetheless, the actual validity of these practices, in terms of practical application, impact, and effort, from the developers’ perspective has not been investigated yet. This paper addresses this limitation, assessing the developers’ perspective of a set of 28 fairness practices collected from the literature. We perform a survey study involving 155 practitioners who have been working on the development and maintenance of ML-enabled systems, analyzing the answers via statistical and clustering analysis to group fairness-aware practices based on their application frequency, impact on bias mitigation, and effort required for their application. While all the practices are deemed relevant by developers, those applied at the early stages of development appear to be the most impactful. More importantly, the effort required to implement the practices is average and sometimes high, with a subsequent average application. The findings highlight the need for effort-aware automated approaches that ease the application of the available practices, as well as recommendation systems that may suggest when and how to apply fairness-aware practices throughout the software lifecycle. Gianmario Voria, Giulia Sellitto, Carmine Ferrara, Francesco Abate, Andrea De Lucia, Filomena Ferrucci, Gemma Catolino, Fabio Palomba |
Inf. Softw. Technol. | 7 |
| 2025 | Examining the impact of bias mitigation algorithms on the sustainability of ML-enabled systems: A benchmark studyabstractContext: As machine learning (ML) systems become increasingly prevalent across various industries, concerns regarding fairness have intensified. Bias mitigation algorithms—that aim to reduce bias in ML models—serve as solutions to mitigate this issue. However, these techniques can affect more than just social sustainability . They may alter the computational overhead and energy usage of ML systems, affecting their environmental sustainability. Similarly, they can influence businesses’ economic sustainability by shaping resource allocation and consumer trust. Goal: This work aims to provide a benchmark study of the implications of applying bias mitigation algorithms on the sustainability of ML solutions. We first corroborate previous findings by examining their effect on social sustainability metrics. Additionally, we complement existing studies by offering a comprehensive analysis of how bias mitigation affects environmental and economic sustainability, aiming to highlight trade-offs for practitioners designing ML solutions. Method: We evaluate six bias mitigation algorithms by conducting 3,360 experiments across multiple configurations of four ML algorithms and datasets. From these experiments, we compute metrics for social, environmental, and economic sustainability, evaluating them using statistical analysis. Results: Our quantitative findings show that all bias mitigation algorithms affect the three sustainability dimensions differently, indicating that applying these algorithms involves complex trade-offs. Furthermore, we expand our discussion with qualitative insights that arise from our results, also providing implications for both research and practice. Conclusions: Our study emphasizes the need for a deeper investigation into the trade-offs bias mitigation algorithms introduce and how they impact various non-functional requirements of ML systems. Editor’s note: Open Science material was validated by the Journal of Systems and Software Open Science Board . Vincenzo De Martino, Gianmario Voria, Ciro Troiano, Gemma Catolino, Fabio Palomba |
J. Syst. Softw. | 4 |
| 2025 | A Novel, Tool-Supported Catalog of Community Smell SymptomsabstractABSTRACT Software development is a multifaceted endeavor, requiring a profound grasp of both social dynamics and technical intricacies. Poor collaboration often leads to the accumulation of social debt , manifesting as unforeseen project costs due to sub‐optimal team interactions. Community smells have emerged as indicators of these socio‐technical inefficiencies and potential social debt. While previous research has focused on automated detection of community smells through analyzing developer communication patterns, our study offers a complementary approach. We emphasize the critical role of project managers in assessing socio‐technical dynamics and propose a novel, tool‐supported catalog of symptoms. This catalog can be used for manual inspections to identify early signs of community smells at the individual level, allowing managers to address issues before they escalate. Using a mixed‐method design that leveraged an existing literature review and a user survey, we cataloged symptoms related to four community smell types. Additionally, we developed TOAST, a tool that operationalizes this catalog, and assessed its usability and practical usefulness through an experiment involving project managers. The study showed that even participants unfamiliar with the term “community smells” were able to interpret the tool's output, reflect on team dynamics, and recognize problematic behavioral patterns when supported by structured symptom‐based information. The paper concludes by shedding light on the potential impact of our work and its contribution to advancing the detection and analysis of community smells. Antonio Della Porta, Stefano Lambiase, Gemma Catolino, Filomena Ferrucci, Fabio Palomba |
J. Softw. Evol. Process. | 3 |
| 2025 | Uncovering Community Smells in Machine Learning-Enabled Systems: Causes, Effects, and Mitigation StrategiesabstractSuccessful software development hinges on effective communication and collaboration, which are significantly influenced by human and social dynamics. Poor management of these elements can lead to the emergence of ‘community smells’, i.e., negative patterns in socio-technical interactions that gradually accumulate as ‘social debt’. This issue is particularly pertinent in machine learning-enabled systems, where diverse actors such as data engineers and software engineers interact at various levels. The unique collaboration context of these systems presents an ideal setting to investigate community smells and their impact on development communities. This article addresses a gap in the literature by identifying the types, causes, effects, and potential mitigation strategies of community smells in machine learning-enabled systems. Using Partial Least Squares Structural Equation Modeling (PLS-SEM), we developed hypotheses based on existing literature and interviews, and conducted a questionnaire-based study to collect data. Our analysis resulted in the construction and validation of five models that represent the causes, effects, and strategies for five specific community smells. These models can help practitioners identify and address community smells within their organizations, while also providing valuable insights for future research on the socio-technical aspects of machine learning-enabled system communities. Giusy Annunziata, Stefano Lambiase, Damian A. Tamburri, Willem-Jan van den Heuvel, Fabio Palomba, Gemma Catolino, Filomena Ferrucci, Andrea De Lucia |
ACM Trans. Softw. Eng. Methodol. | 6 |
| 2025 | RECOVER: Toward Requirements Generation From Stakeholders' ConversationsabstractStakeholders’ conversations in requirements elicitation meetings hold valuable insights into system and client needs. However, manually extracting requirements is time-consuming, labor-intensive, and prone to errors and biases. While current state-of-the-art methods assist in summarizing stakeholder conversations and classifying requirements based on their nature, there is a noticeable lack of approaches capable of both identifying requirements within these conversations and generating corresponding system requirements. These approaches would assist requirement identification, reducing engineers’ workload, time, and effort. They would also enhance accuracy and consistency in documentation, providing a reliable foundation for further analysis. To address this gap, this paper introduces RECOVER (Requirements EliCitation frOm conVERsations), a novel conversational requirements engineering approach that leverages natural language processing and large language models (LLMs) to support practitioners in automatically extracting system requirements from stakeholder interactions by analyzing individual conversation turns. The approach is evaluated using a mixed-method research design that combines statistical performance analysis with a user study involving requirements engineers, targeting two levels of granularity. First, at the conversation turn level, the evaluation measures RECOVER’s accuracy in identifying requirements-relevant dialogue and the quality of generated requirements in terms of correctness, completeness, and actionability. Second, at the entire conversation level, the evaluation assesses the overall usefulness and effectiveness of RECOVER in synthesizing comprehensive system requirements from full stakeholder discussions. Empirical evaluation of RECOVER shows promising performance, with generated requirements demonstrating satisfactory correctness, completeness, and actionability. The results also highlight the potential of automating requirements elicitation from conversations as an aid that enhances efficiency while maintaining human oversight. Gianmario Voria, Francesco Casillo, Carmine Gravino, Gemma Catolino, Fabio Palomba |
IEEE Trans. Software Eng. | 4 |
| 2024 | Security Risk Assessment on Cloud: A Systematic Mapping StudyabstractCloud computing has become integral to modern organizational operations, offering efficiency and agility. However, security challenges such as data loss and downtime necessitate tailored compliance solutions. Risk assessment is crucial for identifying and mitigating cloud-related threats, yet a standardized approach remains elusive. Our study aims to fill this gap by conducting a systematic mapping study on the prevailing methodologies. Through a meticulous analysis of 21 scholarly papers, we explore various aspects of security risk assessment for the cloud. The results provide valuable insights into delivery models, standards, and validation practices, contributing to a comprehensive understanding of cloud risk assessment. Giusy Annunziata, Alexandra Sheykina, Fabio Palomba, Andrea De Lucia, Gemma Catolino, Filomena Ferrucci |
EASE | 5 |
| 2024 | FRINGE: context-aware FaiRness engineerING in complex software systEmsabstractMachine learning (ML) is essential in modern technology, driving complex data-driven decisions. By 2025, daily data generation will exceed 463 exabytes, increasing ML’s influence and ethical risks of data exploitation and discrimination. The European Union’s Artificial Intelligence Act highlights the need for ethical AI solutions. Fabio Palomba, Andrea Di Sorbo, Davide Di Ruscio, Filomena Ferrucci, Gemma Catolino, Giammaria Giordano, Dario Di Dario, Gianmario Voria, Viviana Pentangelo, Maria Tortorella, Arnaldo Sgueglia, Claudio Di Sipio, Giordano d'Aloisio, Antinisca Di Marco |
ESEM | 5 |
| 2024 | An Empirical Study on the Relation Between Programming Languages and the Emergence of Community SmellsabstractTo provide a measurable representation of social issues in software teams, the research community defined a set of anti-patterns that may lead to the emergence of both social and technical debt, i.e., “community smells”. Researchers have investigated community smells from different perspectives; in particular, they have analyzed how product-related aspects of software development, such as architecture and introducing a new language, could influence community smells. However, how technical project characteristics may be in relation to the emergence of community smells is still unknown. Different from those works, we aim to investigate how adopting specific programming languages might influence the socio-technical alignment and congruence of the development community, possibly inducing their overall ability to communicate and collaborate, leading to the emergence of social anti-patterns, i.e., community smells. We studied the relationship between the most used programming languages and the community smells in 100 open-source projects on G ITHub. Key results of the study show a low statistical correlation for specific community smells like Prima Donna Effects, Solution Defiance, and Organizational Skirmish, highlighting the fact that for some programming languages, its adoption could not be an indicator of the presence or absence of community smells. Giusy Annunziata, Carmine Ferrara, Stefano Lambiase, Fabio Palomba, Gemma Catolino, Filomena Ferrucci, Andrea De Lucia |
SEAA | 5 |
| 2024 | On the adoption and effects of source code reuse on defect proneness and maintenance effortabstractAbstract Software reusability mechanisms, like inheritance and delegation in Object-Oriented programming, are widely recognized as key instruments of software design that reduce the risks of source code being affected by defects, other than to reduce the effort required to maintain and evolve source code. Previous work has traditionally employed source code reuse metrics for prediction purposes, e.g., in the context of defect prediction. However, our research identifies two noticeable limitations of the current literature. First, still little is known about the extent to which developers actually employ code reuse mechanisms over time. Second, it is still unclear how these mechanisms may contribute to explaining defect-proneness and mainten0ance effort during software evolution. We aim at bridging this gap of knowledge, as an improved understanding of these aspects might provide insights into the actual support provided by these mechanisms, e.g., by suggesting whether and how to use them for prediction purposes. We propose an exploratory study, conducted on 12Javaprojects–over 44,900 commits–of theDefects4Jdataset, aiming at (1) assessing how developers use inheritance and delegation during software evolution; and (2) statistically analyzing the impact of inheritance and delegation on fault proneness and maintenance effort. Our results let emerge various usage patterns that describe the way inheritance and delegation vary over time. In addition, we find out that inheritance and delegation are statistically significant factors that influence both source code defect-proneness and maintenance effort. Giammaria Giordano, Gerardo Festa, Gemma Catolino, Fabio Palomba, Filomena Ferrucci, Carmine Gravino |
Empir. Softw. Eng. | 3 |
| 2024 | An Empirical Investigation Into the Influence of Software Communities' Cultural and Geographical Dispersion on ProductivityabstractEstimating and understanding software development productivity represent crucial tasks for researchers and practitioners. Although different works focused on evaluating the impact of human factors on productivity, a few explored the influence of cultural/geographical diversity in software development communities. More particularly, all previous treatise addresses cultural aspects as abstract concepts without providing a quantitative representation. Improved knowledge of these matters might help project managers to assemble more productive teams and tool vendors to design software analytics toolkits that may better estimate productivity. This paper has the goal of enlarging the existing body of knowledge on the factors affecting productivity by focusing on cultural and geographical dispersion of a development community—namely, how diverse a community is in terms of cultural attitudes and geographical collocation of the members who belong to it. To reach this goal, we performed a mixed-method empirical study. First, we built a statistical model relating dispersion metrics with the productivity of 25 open-source communities on Github. Then, we performed a confirmatory survey with 140 practitioners. The key results of our study indicate that cultural and geographical dispersion considerably impact productivity, thus encouraging managers and practitioners to consider such aspects during all the phases of the software development lifecycle. We conclude our paper by elaborating on the main insights from our analyses and instilling implications that may drive further research. Stefano Lambiase, Gemma Catolino, Fabiano Pecorelli, Damian A. Tamburri, Fabio Palomba, Willem-Jan van den Heuvel, Filomena Ferrucci |
J. Syst. Softw. | 2 |
| 2024 | Technical debt in AI-enabled systems: On the prevalence, severity, impact, and management strategies for code and architectureabstractArtificial Intelligence (AI) is pervasive in several application domains and promises to be even more diffused in the next decades. Developing high-quality AI-enabled systems — software systems embedding one or multiple AI components, algorithms, and models — could introduce critical challenges for mitigating specific risks related to the systems’ quality. Such development alone is insufficient to fully address socio-technical consequences and the need for rapid adaptation to evolutionary changes. Recent work proposed the concept of AI technical debt, a potential liability concerned with developing AI-enabled systems whose impact can affect the overall systems’ quality. While the problem of AI technical debt is rapidly gaining the attention of the software engineering research community, scientific knowledge that contributes to understanding and managing the matter is still limited. In this paper, we leverage the expertise of practitioners to offer useful insights to the research community, aiming to enhance researchers’ awareness about the detection and mitigation of AI technical debt. Our ultimate goal is to empower practitioners by providing them with tools and methods. Additionally, our study sheds light on novel aspects that practitioners might not be fully acquainted with, contributing to a deeper understanding of the subject. We develop a survey study featuring 53 AI practitioners, in which we collect information on the practical prevalence, severity, and impact of AI technical debt issues affecting the code and the architecture other than the strategies applied by practitioners to identify and mitigate them. The key findings of the study reveal the multiple impacts that AI technical debt issues may have on the quality of AI-enabled systems (e.g., the high negative impact that Undeclared consumers has on security, whereas Jumbled Model Architecture can induce the code to be hard to maintain) and the little support practitioners have to deal with them, limited to apply manual effort for identification and refactoring. We conclude the article by distilling lessons learned and actionable insights for researchers. Gilberto Recupito, Fabiano Pecorelli, Gemma Catolino, Valentina Lenarduzzi, Davide Taibi 0001, Dario Di Nucci, Fabio Palomba |
J. Syst. Softw. | 3 |
| 2024 | Generative AI in Software Engineering Must Be Human-Centered: The Copenhagen Manifesto
Daniel Russo 0002, Sebastian Baltes, Niels van Berkel, Paris Avgeriou, Fabio Calefato, Beatriz Cabrero-Daniel, Gemma Catolino, Jürgen Cito, Neil A. Ernst, Thomas Fritz 0001, Hideaki Hata, Reid Holmes, Maliheh Izadi, Foutse Khomh, Mikkel Baun Kjærgaard, Grischa Liebel, Alberto Lluch-Lafuente, Stefano Lambiase, Walid Maalej, Gail C. Murphy, Nils Brede Moe, Gabrielle O'Brien, Elda Paja, Mauro Pezzè, John Stouby Persson, Rafael Prikladnicki, Paul Ralph, Martin P. Robillard, Thiago Rocha Silva, Klaas-Jan Stol, Margaret-Anne D. Storey, Viktoria Stray, Paolo Tell, Christoph Treude, Bogdan Vasilescu |
J. Syst. Softw. | 7 |
| 2024 | "Through the looking-glass ..." An empirical study on blob infrastructure blueprints in the Topology and Orchestration Specification for Cloud ApplicationsabstractAbstract Infrastructure‐as‐code (IaC) helps keep up with the demand for fast, reliable, high‐quality services by provisioning and managing infrastructures through configuration files. Those files ensure efficient and repeatable routines for system provisioning, but they might be affected by code smells that negatively affect quality and code maintenance. Research has broadly studied code smells for traditional source code development; however, none explored them in the “Topology and Orchestration Specification for Cloud Applications” (TOSCA), the technology‐agnostic OASIS standard for IaC. In this paper, we investigate a prominent traditional implementation code smell potentially applicable to TOSCA: Large Class, or “Blob Blueprint” in IaC terms. We compare metrics‐based and unsupervised learning‐based detectors on a large dataset of manually validated observations related to Blob Blueprints. We provide insights on code metrics that corroborate previous findings and empirically show that metrics‐based detectors perform highly in detecting Blob Blueprints. We deem our results put forward a new research path toward dealing with this problem, for example, in the scope of fully automated service pipelines. Stefano Dalla Palma, Chiel van Asseldonk, Gemma Catolino, Dario Di Nucci, Fabio Palomba, Damian A. Tamburri |
J. Softw. Evol. Process. | 3 |
| 2024 | Few Images, Many Insights: Illicit Content Detection Using a Limited Number of ImagesabstractThe anonymity and untraceability benefits of the dark web increased its popularity exponentially. The cost of these technical benefits is that such anonymity has created a suitable womb for illicit activity. Hence—in collaboration with cybersecurity practitioners and law-enforcement agencies—the research community provided approaches for recognizing and classifying illicit activities. Most of these approaches exploit textual content from dark web markets, whereas few used images that originated from them. This article investigates alternative techniques for recognizing illegal activities from images. The significant contributions of our work are threefold: (a) We investigate label-agnostic learning techniques like One-Shot and Few-Shot learning that use Siamese Neural Networks. Our approach manages to handle small-scale datasets with promising accuracy. In particular, the Siamese Neural Network approach reaches 90.9% on 5-Shot experiments over a 10-class dataset. (b) This study’s satisfactory findings facilitate the creation of potent tools to assist authorities in identifying illicit content on the Web. Moreover, our proof-of-concept approach demonstrated the ability to recognize illegal images using a limited number of files, reducing the time constraint in collecting illegal images. (c) We provide a complete labeled dataset of 3,570 images from 55 different categories from dark web markets that can be used for future research activities. Giuseppe Cascavilla, Gemma Catolino, Mauro Conti, Dimos Mellios, Damian A. Tamburri |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2023 | When the Few Outweigh the Many: Illicit Content Recognition with Few-Shot LearningabstractOnline Appendix Giuseppe Cascavilla, Gemma Catolino, Mauro Conti, Dimos Mellios, Damian A. Tamburri |
SECRYPT | 2 |
| 2022 | "There and Back Again?" On the Influence of Software Community Dispersion Over ProductivityabstractEstimating and understanding productivity still represents a crucial task for researchers and practitioners. Researchers spent significant effort identifying the factors that influence software developers’ productivity, providing several approaches for analyzing and predicting such a metric. Although different works focused on evaluating the impact of human factors on productivity, little is known about the influence of cultural/geographical diversity in software development communities. Indeed, in previous studies, researchers treated cultural aspects like an abstract concept without providing a quantitative representation. This work provides an empirical assessment of the relationship between cultural and geographical dispersion of a development community—namely, how diverse a community is in terms of cultural attitudes and geographical collocation of the members who belong to it—and its productivity. To reach our aim, we built a statistical model that contained product and socio-technical factors as independent variables to assess the correlation with productivity, i.e., the number of commits performed in a given time. Then, we ran our model considering data of 25 open-source communities on GitHub. Results of our study indicate that cultural and geographical dispersion impact productivity, thus encouraging managers and practitioners to consider such aspects during all the phases of the software development lifecycle. Stefano Lambiase, Gemma Catolino, Fabiano Pecorelli, Damian A. Tamburri, Fabio Palomba, Willem-Jan van den Heuvel, Filomena Ferrucci |
SEAA | 2 |
| 2022 | A Multivocal Literature Review of MLOps Tools and FeaturesabstractDevOps has become increasingly widespread, with companies employing its methods in different fields. In this context, MLOps automates Machine Learning pipelines by applying DevOps practices. Considering the high number of tools available and the high interest of the practitioners to be supported by tools to automate the steps of Machine Learning pipelines, little is known concerning MLOps tools and their functionalities. To this aim, we conducted a Multivocal Literature Review (MLR) to (i) extract tools that allow for and support the creation of MLOps pipelines and (ii) analyze their main characteristics and features to provide a comprehensive overview of their value. Overall, we investigate the functionalities of 13 MLOps Tools. Our results show that most MLOps Tools support the same features but apply different approaches that can bring different advantages, depending on user requirements. Gilberto Recupito, Fabiano Pecorelli, Gemma Catolino, Sergio Moreschini, Dario Di Nucci, Fabio Palomba, Damian A. Tamburri |
SEAA | 3 |
| 2022 | "When the Code becomes a Crime Scene" Towards Dark Web Threat Intelligence with Software Quality MetricsabstractThe increasing growth of illegal online activities in the so-called dark web—that is, the hidden collective of internet sites only accessible by a specialized web browsers—has challenged law enforcement agencies in recent years with sparse research efforts to help. For example, research has been devoted to supporting law enforcement by employing Natural Language Processing (NLP) to detect illegal activities on the dark web and build models for their classification. However, current approaches strongly rely upon the linguistic characteristics used to train the models, e.g., language semantics, which threatens their generalizability. To overcome this limitation, we tackle the problem of predicting illegal and criminal activities—a process defined as threat intelligence—on the dark web from a complementary perspective—that of dark web code maintenance and evolution— and propose a novel approach that uses software quality metrics and dark website appearance parameters instead of linguistic characteristics. We performed a preliminary empirical study on 10.367 web pages and collected more than 40 code metrics and website parameters using sonarqube. Results show an accuracy of up to 82% for predicting the three types of illegal activities (i.e., suspicious, normal, and unknown) and 66% for detecting 26 specific illegal activities, such as drugs or weapons trafficking. We deem our results can influence the current trends in detecting illegal activities on the dark web and put forward a completely novel research avenue toward dealing with this problem from a software maintenance and evolution perspective. Giuseppe Cascavilla, Gemma Catolino, Felipe Ebert, Damian A. Tamburri, Willem-Jan van den Heuvel |
ICSME | 2 |
| 2022 | Community Smell Detection and Refactoring in SLACK: The CADOCS ProjectabstractSoftware engineering is a human-centered activity involving various stakeholders with different backgrounds that have to communicate and collaborate to reach shared objectives. The emergence of conflicts among stakeholders may lead to undesired effects on software maintainability, yet it is often unavoidable in the long run. Community smells, i.e., sub-optimal communication and collaboration practices, have been defined to map recurrent conflicts among developers. While some community smell detection tools have been proposed in the recent past, these can be mainly used for research purposes because of their limited level of usability and user engagement. To facilitate a wider use of community smell-related information by practitioners, we present CADOCS, a client-server conversational agent that builds on top of a previous community smell detection tool proposed by Almarini et al. to (1) make it usable within a well-established communication channel like Slack and (2) augment it by providing initial support to software analytics instruments useful to diagnose and refactor community smells. We describe the features of the tool and the preliminary evaluation conducted to assess and improve robustness and usability. Gianmario Voria, Viviana Pentangelo, Antonio Della Porta, Stefano Lambiase, Gemma Catolino, Fabio Palomba, Filomena Ferrucci |
ICSME | 5 |
| 2022 | Illicit Darkweb Classification via Natural-language Processing: Classifying Illicit Content of Webpages based on Textual InformationabstractThis work aims at expanding previous works done in the context of illegal activities classification, performing three different steps. First, we created a heterogeneous dataset of 113995 onion sites and dark marketplaces. Then, we compared pre-trained transferable models, i.e., ULMFit (Universal Language Model Fine-tuning), Bert (Bidirectional Encoder Representations from Transformers), and RoBERTa (Robustly optimized BERT approach) with a traditional text classification approach like LSTM (Long short-term memory) neural networks. Finally, we developed two illegal activities classification approaches, one for illicit content on the Dark Web and one for identifying the specific types of drugs. Results show that Bert obtained the best approach, classifying the dark web's general content and the types of Drugs with 96.08% and 91.98% of accuracy. Giuseppe Cascavilla, Gemma Catolino, Mirella Sangiovanni |
SECRYPT | 2 |
| 2022 | On the Evolution of Inheritance and Delegation Mechanisms and Their Impact on Code QualityabstractSource code reuse is considered one of the holy grails of modern software development. Indeed, it has been widely demonstrated that this activity decreases software development and maintenance costs while increasing its overall trustwor-thiness. The Object-Oriented (OO) paradigm provides different internal mechanisms to favor code reuse, i.e., specification inheritance, implementation inheritance, and delegation. While previous studies investigated how inheritance relations impact source code quality, there is still a lack of understanding of their evolutionary aspects and, more particular, of how these mechanisms may impact source code quality over time. To bridge this gap of knowledge, this paper proposes an empirical investigation into the evolution of specification inheritance, implementation inheritance, and delegation and their impact on the variability of source code quality attributes. First, we assess how the implementation of those mechanisms varies over 15 releases of three software systems. Second, we devise a statistical approach with the aim of understanding how inheritance and delegation let source code quality—as indicated by the severity of code smells—vary in either positive or negative manner. The key results of the study indicate that inheritance and delegation evolve over time, but not in a statistically significant manner. At the same time, their evolution often leads code smell severity to be reduced, hence possibly contributing to improve code maintainability. Giammaria Giordano, Antonio Fasulo, Gemma Catolino, Fabio Palomba, Filomena Ferrucci, Carmine Gravino |
SANER | 3 |
| 2022 | Gender Diversity and Community Smells: A Double-Replication Study on Brazilian Software TeamsabstractSocial debts in software teams are gaining increasing attention from the research community due to their potential adverse effects on software quality. For instance, community smells are indicators of sub-optimal organizational structures and may well lead to the emergence of social debt. Previous studies analyzed which factors influence the emergence/mitigation of such smells. In particular, studies by Catolino et al. showed how factors related to team composition, particularly gender diversity, correlated to the mitigation of community smells. However, a confirmation survey on 60 practitioners suggested that these results were not aligned with the experts' perceptions. In a separate survey, Catolino et al. collected the most common team refactoring strategies for those community smells. In this work we replicate two studies by those authors, focusing on the Brazilian software teams; culture-specific expectations on the behavior of people of different genders might have affected the perception of the importance of gender diversity and refactoring strategies when mitigating community smells. We translated the survey instrument used by Catolino et al. to Brazilian Portuguese and recruited 184 Brazilian developers. Re-sults did not show significant differences from the original study; indeed, participants perceived gender diversity as less valuable to mitigate community smells than such factors like experience or team size. Additionally, we performed a qualitative analysis of an open question within the questionnaire for the refactoring strategies. Brazilian developers agree with the original studies for most smells, mainly promoting restructuring communities, creating a communication plan and mentoring. We believe these results provide further evidence on the problem and its implications when managing software teams, avoiding technical debt and maintenance issues due to team communication and coordination problems. Camila Sarmento, Tiago Massoni, Alexander Serebrenik, Gemma Catolino, Damian A. Tamburri, Fabio Palomba |
SANER | 4 |
| 2022 | Software testing and Android applications: a large-scale empirical study
Fabiano Pecorelli, Gemma Catolino, Filomena Ferrucci, Andrea De Lucia, Fabio Palomba |
Empir. Softw. Eng. | 2 |
| 2022 | A compositional model for effort-aware Just-In-Time defect prediction on android appsabstractAbstract Android apps have played important roles in daily life and work. To meet the new requirements from users, the apps encounter frequent updates, which involves a large quantity of code commits. Previous studies proposed to apply Just‐in‐Time (JIT) defect prediction for apps to timely identify whether the new code commits can introduce defects into apps, aiming to assure their quality. In general, high‐quality features are benefits for improving the classification performance. In addition, the number of defective commit instances is much fewer than that of clean ones, that is the defect data is class imbalanced. In this study, a novel compositional model, called KPIDL, is proposed to conduct the JIT defect prediction task for Android apps. More specifically, KPIDL first exploits a feature learning technique to preprocess original data for obtaining better feature representation, and then introduces a state‐of‐the‐art cost‐sensitive cross‐entropy loss function into the deep neural network to alleviate the class imbalance issue by considering the prior probability of the two types of classes. The experiments were conducted on a benchmark defect data consisting of 15 Android apps. The experimental results show that the proposed KPIDL model performs significantly better than 25 comparative methods in terms of two effort‐aware performance indicators in most cases. Kunsong Zhao, Zhou Xu 0003, Meng Yan 0001, Lei Xue 0001, Wei Li 0121, Gemma Catolino |
IET Softw. | 6 |
| 2022 | Effort-Aware Just-in-Time Bug Prediction for Mobile Apps Via Cross-Triplet Deep Feature EmbeddingabstractJust-in-time (JIT) bug prediction is an effective quality assurance activity that identifies whether a code commit will introduce bugs into the mobile app, aiming to provide prompt feedback to practitioners for priority review. Since collecting sufficient labeled bug data is not always feasible for some mobile apps, one possible approach is to leverage cross-app models. In this work, we propose a new cross-triplet deep feature embedding method, called CDFE, for cross-app JIT bug prediction task. The CDFE method incorporates a state-of-the-art cross-triplet loss function into a deep neural network to learn high-level feature representation for the cross-app data. This loss function adapts to the cross-app feature learning task and aims to learn a new feature space to shorten the distance of commit instances with the same label and enlarge the distance of commit instances with different labels. In addition, this loss function assigns higher weights to losses caused by cross-app instance pairs than that by intra-app instance pairs, aiming to narrow the discrepancy of cross-app bug data. We evaluate our CDFE method on a benchmark bug dataset from 19 mobile apps with two effort-aware indicators. The experimental results on 342 cross-app pairs show that our proposed CDFE method performs better than 14 baseline methods. Zhou Xu 0003, Kunsong Zhao, Tao Zhang 0001, Chunlei Fu, Meng Yan 0001, Zhiwen Xie, Xiaohong Zhang 0002, Gemma Catolino |
IEEE Trans. Reliab. | 8 |
| 2020 | A blessing in disguise? Assessing the Relationship between Code Smells and SustainabilityabstractCode smells have always been associated with bad code source quality and maintenance problems by the research community. However, a few code smells remain within code for years without influence a project, and some of the developers do not even recognize them. This late-breaking idea paper wants to investigate the role of code smells within the context of software quality and the sustainability of projects. For this reason, we first plan to understand the impact of code smells on sustainability and then perform a fine-grain analysis to analyze their evolution, focusing on the possible benefits of their presence. Gemma Catolino |
ICSME | 1 |
| 2020 | Testing of Mobile Applications in the Wild: A Large-Scale Empirical Study on Android AppsabstractNowadays, mobile applications (a.k.a., apps) are used by over two billion users for every type of need, including social and emergency connectivity. Their pervasiveness in today's world has inspired the software testing research community in devising approaches to allow developers to better test their apps and improve the quality of the tests being developed. In spite of this research effort, we still notice a lack of empirical studies aiming at assessing the actual quality of test cases developed by mobile developers: this perspective could provide evidence-based findings on the current status of testing in the wild as well as on the future research directions in the field. As such, we performed a large-scale empirical study targeting 1,780 open-source Android apps and aiming at assessing (1) the extent to which these apps are actually tested, (2) how well-designed are the available tests, and (3) what is their effectiveness. The key results of our study show that mobile developers still tend not to properly test their apps. Furthermore, we discovered that the test cases of the considered apps have a low (i) design quality, both in terms of test code metrics and test smells, and (ii) effectiveness when considering code coverage as well as assertion density. Fabiano Pecorelli, Gemma Catolino, Filomena Ferrucci, Andrea De Lucia, Fabio Palomba |
ICPC | 2 |
| 2020 | Improving change prediction models with code smell-related information
Gemma Catolino, Fabio Palomba, Francesca Arcelli Fontana, Andrea De Lucia, Andy Zaidman, Filomena Ferrucci |
Empir. Softw. Eng. | 1 |
| 2019 | How the Experience of Development Teams Relates to Assertion Density of Test ClassesabstractThe impact of developers' experience on several development practices has been widely investigated in the past. One of the most promising research fields is software testing, as many researchers found significant correlations between developers' experience and testing effectiveness. In this paper, we aim at further studying this relation, by focusing on how development teams' experience is associated with the assertion density, i.e., the number of assertions per test class KLOC, that has previously been shown as an effective way to decrease fault density. We perform a mixed-methods empirical study. First, we devise a statistical model relating development teams' experience and other control factors to the assertion density of test classes belonging to 12 software projects. This model enables us to investigate whether experience comes out as a statistically significant factor to explain assertion density. Second, we contrast the statistical findings with a survey study conducted with 57 developers, who were asked their opinions on how developer's experience is related to the way they add assertions in test code. Our findings suggest the existence of a relationship: on the one hand, the development team's experience is a statistically significant factor in most of the systems that we have investigated; on the other hand, developers confirm the importance of experience and team composition for the effective testing of production code. Gemma Catolino, Fabio Palomba, Andy Zaidman, Filomena Ferrucci |
ICSME | 1 |
| 2019 | Not all bugs are the same: Understanding, characterizing, and classifying bug types
Gemma Catolino, Fabio Palomba, Andy Zaidman, Filomena Ferrucci |
J. Syst. Softw. | 1 |
| 2019 | An extensive evaluation of ensemble techniques for software change predictionabstractAbstract Predicting the areas of the source code having a higher likelihood to change in the future represents an important activity to allow developers to plan preventive maintenance operations. For this reason, several change prediction models have been proposed. Moreover, research community demonstrated how different classifiers impact on the performance of devised models as well as classifiers tend to perform similarly even though they are able to correctly predict the change proneness of different code elements, possibly indicating the presence of some complementarity among them. In this paper, we deeper investigated whether the use of ensemble approaches, ie, machine learning techniques able to combine multiple classifiers, can improve the performances of change prediction models. Specifically, we built three change prediction models based on different predictors, ie, product‐, process‐ metrics‐, and developer‐related factors, comparing the performances of four ensemble techniques (ie, Boosting, Random Forest, Bagging, and Voting) with those of standard machine learning classifiers (ie, Logistic Regression, Naive Bayes, Simple Logistic, and Multilayer Perceptron). The study was conducted on 33 releases of 10 open‐source systems, and the results showed how ensemble methods and in particular Random Forest provide a significant improvement of more than 10% in terms of F measure. Indeed, the statistical analyses conducted confirm the superiority of this ensemble technique. Moreover, the model built using developer‐related factors performed better than the other models that exploit product and process metrics and achieves an overall median of F measure around 77%. Gemma Catolino, Filomena Ferrucci |
J. Softw. Evol. Process. | 1 |
| 2018 | Enhancing change prediction models using developer-related factors
Gemma Catolino, Fabio Palomba, Andrea De Lucia, Filomena Ferrucci, Andy Zaidman |
J. Syst. Softw. | 1 |
| 2017 | Developer-related factors in change prediction: an empirical assessmentabstractPredicting the areas of the source code having a higher likelihood to change in the future is a crucial activity to allow developers to plan preventive maintenance operations such as refactoring or peer-code reviews. In the past the research community was active in devising change prediction models based on structural metrics extracted from the source code. More recently, Elish et al. showed how evolution metrics can be more efficient for predicting change-prone classes. In this paper, we aim at making a further step ahead by investigating the role of different developer-related factors, which are able to capture the complexity of the development process under different perspectives, in the context of change prediction. We also compared such models with existing change-prediction models based on evolution and code metrics. Our findings reveal the capabilities of developer-based metrics in identifying classes of a software system more likely to be changed in the future. Moreover, we observed interesting complementarities among the experimented prediction models, that may possibly lead to the definition of new combined models exploiting developer-related factors as well as product and evolution metrics. Gemma Catolino, Fabio Palomba, Andrea De Lucia, Filomena Ferrucci, Andy Zaidman |
ICPC | 1 |