Gianmario Voria

dblp:337/2934 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
11since 2021 · last 2026
0009-0002-5394-8148ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 11 · 5 first-author · 11 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Tracing Stereotypes in Pre-trained Transformers: From Biased Neurons to Fairer Models
abstract
The advent of transformer-based language models has reshaped how AI systems process and generate text. In software engineering (SE), these models now support diverse activities, accelerating automation and decision-making. Yet, evidence shows that these models can reproduce or amplify social biases, raising fairness concerns. Recent work on neuron editing has shown that internal activations in pre-trained transformers can be traced and modified to alter model behavior. Building on the concept of knowledge neurons—neurons that encode factual information—we hypothesize the existence of biased neurons that capture stereotypical associations within pre-trained transformers.
Gianmario Voria, Moses Openja, Foutse Khomh, Gemma Catolino, Fabio Palomba
MSR1
2026 From values to adoption: on the role of individual cultural values on fairness toolkit adoption in software development
abstract
Abstract Fairness is a critical concern in the integration of machine learning and AI—e.g., Generative AI, LLMs, and Agents—into decision-making, yet the adoption of fairness toolkits by software practitioners remains limited. This gap hinders efforts to operationalize fairness, especially when cultural and ethical values are overlooked. Given that fairness is socially constructed, individual cultural values may significantly influence how practitioners perceive and adopt fairness tools. The objective of this study is, therefore, to investigate whether and how individual cultural values influence software practitioners’ intention to adopt and actual use of fairness toolkits. Specifically, we integrate the UTAUT2 model with Hofstede’s cultural dimensions to examine both direct and moderating cultural effects within the adoption process. A survey of 181 software professionals was conducted, and data were analyzed using Partial Least Squares Structural Equation Modeling. Findings show that cultural values—specifically Power Distance, Collectivism, and Long-Term Orientation—not only directly affect adoption intention and behavior but also moderate key relationships in the adoption process. For example, collectivist individuals were more likely to act on their intention to use fairness tools, highlighting the importance of shared team goals. Overall, our findings indicate that fairness toolkit adoption is shaped not only by technology-related perceptions but also by culturally grounded value orientations. These results provide actionable insights for promoting fairness tool adoption through culturally aware strategies in software development environments.
Stefano Lambiase, Gianmario Voria, Maria Concetta Schiavone, Gemma Catolino, Fabio Palomba
Empir. Softw. Eng.2
2026 Fairness set and forgotten: Mining fairness toolkit usage in open-source machine learning projects
abstract
The development of machine learning (ML) systems in high-stakes domains has amplified concerns about fairness, prompting the creation of fairness toolkits offering metrics and mitigation techniques. Open-source software (OSS) ecosystems, a critical driver of AI innovation, present a unique opportunity to study the practical adoption of these toolkits. This paper aims to empirically characterize the adoption of fairness toolkits in OSS ML projects by investigating for what purposes they are used and how their usage evolves over time. We conducted a mining study on GitHub repositories related to real-world ML projects that integrate fairness toolkits such as AIF360 and Fairlearn . Starting from 1,096 candidate repositories, we applied systematic filtering to identify a final dataset of 20 relevant ML projects (comprising 5,777 total commits). We analyzed toolkit usage by examining invoked APIs and commit histories to uncover patterns of adoption and evolution. Our findings reveal that fairness toolkits are predominantly used for diagnostic purposes, with analytic components integrated early in the project lifecycle and rarely modified thereafter. In contrast, mitigation techniques are infrequently adopted, tend to appear later, and exhibit short, unstable lifespans. Our results show that the adoption of fairness toolkits in OSS ML projects is limited and often restricted to initial diagnostic phases, with active mitigation practices remaining rare. These findings highlight the need for improved support to foster more sustained and effective integration of fairness practices within open-source development.
Alfonso Cannavale, Gianmario Voria, Antonio Scognamiglio, Giammaria Giordano, Gemma Catolino, Fabio Palomba
Inf. Softw. Technol.2
2026 Fair and square? Evaluating fairness of LLM-generated synthetic datasets
abstract
Context: Machine Learning (ML) is driving advancements across various industries, including healthcare, finance, and entertainment, but it also raises significant ethical concerns, particularly regarding fairness. Biases in training data can lead to unfair outcomes, perpetuating or even amplifying existing disparities. Prior research in the Software Engineering (SE) and ML communities has developed numerous bias mitigation techniques, yet two key limitations persist: (1) most approaches intervene at later stages of development, such as after data collection or model training, rather than addressing fairness from the outset; and (2) these methods often mitigate bias without fully eliminating it, since the root issue frequently lies in the data itself. Objective: In this paper, we explore an alternative approach to mitigate unfairness: synthetic data generation , which involves creating artificial datasets that mimic the statistical properties of real-world data. We aim to assess how this approach can contribute to generating data that positively impacts the trade-off between performance and fairness by creating datasets that reduce the influence of real-world biases through synthetic feature generation. Methods: To this end, we conducted an empirical study comparing ML models trained on synthetic datasets generated by large language models to ML models trained on real-world data, evaluating performance and fairness indicators. Results: Our results demonstrate that models trained with synthetic data, particularly those generated using simpler prompts, can achieve competitive performance while enhancing fairness. Conclusion: Our work suggests that synthetic data generation may be a viable approach to addressing fairness requirements in ML systems.
Gianmario Voria, Benedetto Scala, Leopoldo Todisco, Carlo Venditto, Giammaria Giordano, Gemma Catolino, Fabio Palomba
Inf. Softw. Technol.1
2025 Teaching Software Engineering for Artificial Intelligence: An Experience Report
Fabio Palomba, Gianmario Voria, Alessandra Parziale, Viviana Pentangelo, Antonio Della Porta, Vincenzo De Martino, Gilberto Recupito, Giammaria Giordano
SEAA (3)2
2025 Fairness on a budget, across the board: A cost-effective evaluation of fairness-aware practices across contexts, tasks, and sensitive attributes
abstract
Context: Machine Learning (ML) is widely used in critical domains like finance, healthcare, and criminal justice, where unfair predictions can lead to harmful outcomes. Although bias mitigation techniques have been developed by the Software Engineering (SE) community, their practical adoption is limited due to complexity and integration issues. As a simpler alternative, fairness-aware practices, namely conventional ML engineering techniques adapted to promote fairness, e.g., MinMax Scaling, which normalizes feature values to prevent attributes linked to sensitive groups from disproportionately influencing predictions, have recently been proposed, yet their actual impact is still unexplored. Objective: Building on our prior work that explored fairness-aware practices in different contexts, this paper extends the investigation through a large-scale empirical study assessing their effectiveness across diverse ML tasks, sensitive attributes, and datasets belonging to specific application domains. Methods: We conduct 5940 experiments, evaluating fairness-aware practices from two perspectives: contextual bias mitigation and cost-effectiveness . Contextual evaluation examines fairness improvements across different ML models, sensitive attributes, and datasets. Cost-effectiveness analysis considers the trade-off between fairness gains and performance costs. Results: Findings reveal that the effectiveness of fairness-aware practices depends on specific contexts’ datasets and configurations, while cost-effectiveness analysis highlights those that best balance ethical gains and efficiency. Conclusion: These insights guide practitioners in choosing fairness-enhancing practices with minimal performance impact, supporting ethical ML development.
Alessandra Parziale, Gianmario Voria, Giammaria Giordano, Gemma Catolino, Gregorio Robles, Fabio Palomba
Inf. Softw. Technol.2
2025 Fairness-aware practices from developers' perspective: A survey
abstract
Machine Learning (ML) technologies have shown great promise in many areas, but when used without proper oversight, they can produce biased results that discriminate against historically underrepresented groups. In recent years, the software engineering research community has contributed to addressing the need for ethical machine learning by proposing a number of fairness-aware practices, e.g., fair data balancing or testing approaches, that may support the management of fairness requirements throughout the software lifecycle. Nonetheless, the actual validity of these practices, in terms of practical application, impact, and effort, from the developers’ perspective has not been investigated yet. This paper addresses this limitation, assessing the developers’ perspective of a set of 28 fairness practices collected from the literature. We perform a survey study involving 155 practitioners who have been working on the development and maintenance of ML-enabled systems, analyzing the answers via statistical and clustering analysis to group fairness-aware practices based on their application frequency, impact on bias mitigation, and effort required for their application. While all the practices are deemed relevant by developers, those applied at the early stages of development appear to be the most impactful. More importantly, the effort required to implement the practices is average and sometimes high, with a subsequent average application. The findings highlight the need for effort-aware automated approaches that ease the application of the available practices, as well as recommendation systems that may suggest when and how to apply fairness-aware practices throughout the software lifecycle.
Gianmario Voria, Giulia Sellitto, Carmine Ferrara, Francesco Abate, Andrea De Lucia, Filomena Ferrucci, Gemma Catolino, Fabio Palomba
Inf. Softw. Technol.1
2025 Examining the impact of bias mitigation algorithms on the sustainability of ML-enabled systems: A benchmark study
abstract
Context: As machine learning (ML) systems become increasingly prevalent across various industries, concerns regarding fairness have intensified. Bias mitigation algorithms—that aim to reduce bias in ML models—serve as solutions to mitigate this issue. However, these techniques can affect more than just social sustainability . They may alter the computational overhead and energy usage of ML systems, affecting their environmental sustainability. Similarly, they can influence businesses’ economic sustainability by shaping resource allocation and consumer trust. Goal: This work aims to provide a benchmark study of the implications of applying bias mitigation algorithms on the sustainability of ML solutions. We first corroborate previous findings by examining their effect on social sustainability metrics. Additionally, we complement existing studies by offering a comprehensive analysis of how bias mitigation affects environmental and economic sustainability, aiming to highlight trade-offs for practitioners designing ML solutions. Method: We evaluate six bias mitigation algorithms by conducting 3,360 experiments across multiple configurations of four ML algorithms and datasets. From these experiments, we compute metrics for social, environmental, and economic sustainability, evaluating them using statistical analysis. Results: Our quantitative findings show that all bias mitigation algorithms affect the three sustainability dimensions differently, indicating that applying these algorithms involves complex trade-offs. Furthermore, we expand our discussion with qualitative insights that arise from our results, also providing implications for both research and practice. Conclusions: Our study emphasizes the need for a deeper investigation into the trade-offs bias mitigation algorithms introduce and how they impact various non-functional requirements of ML systems. Editor’s note: Open Science material was validated by the Journal of Systems and Software Open Science Board .
Vincenzo De Martino, Gianmario Voria, Ciro Troiano, Gemma Catolino, Fabio Palomba
J. Syst. Softw.2
2025 RECOVER: Toward Requirements Generation From Stakeholders' Conversations
abstract
Stakeholders’ conversations in requirements elicitation meetings hold valuable insights into system and client needs. However, manually extracting requirements is time-consuming, labor-intensive, and prone to errors and biases. While current state-of-the-art methods assist in summarizing stakeholder conversations and classifying requirements based on their nature, there is a noticeable lack of approaches capable of both identifying requirements within these conversations and generating corresponding system requirements. These approaches would assist requirement identification, reducing engineers’ workload, time, and effort. They would also enhance accuracy and consistency in documentation, providing a reliable foundation for further analysis. To address this gap, this paper introduces RECOVER (Requirements EliCitation frOm conVERsations), a novel conversational requirements engineering approach that leverages natural language processing and large language models (LLMs) to support practitioners in automatically extracting system requirements from stakeholder interactions by analyzing individual conversation turns. The approach is evaluated using a mixed-method research design that combines statistical performance analysis with a user study involving requirements engineers, targeting two levels of granularity. First, at the conversation turn level, the evaluation measures RECOVER’s accuracy in identifying requirements-relevant dialogue and the quality of generated requirements in terms of correctness, completeness, and actionability. Second, at the entire conversation level, the evaluation assesses the overall usefulness and effectiveness of RECOVER in synthesizing comprehensive system requirements from full stakeholder discussions. Empirical evaluation of RECOVER shows promising performance, with generated requirements demonstrating satisfactory correctness, completeness, and actionability. The results also highlight the potential of automating requirements elicitation from conversations as an aid that enhances efficiency while maintaining human oversight.
Gianmario Voria, Francesco Casillo, Carmine Gravino, Gemma Catolino, Fabio Palomba
IEEE Trans. Software Eng.1
2024 FRINGE: context-aware FaiRness engineerING in complex software systEms
abstract
Machine learning (ML) is essential in modern technology, driving complex data-driven decisions. By 2025, daily data generation will exceed 463 exabytes, increasing ML’s influence and ethical risks of data exploitation and discrimination. The European Union’s Artificial Intelligence Act highlights the need for ethical AI solutions.
Fabio Palomba, Andrea Di Sorbo, Davide Di Ruscio, Filomena Ferrucci, Gemma Catolino, Giammaria Giordano, Dario Di Dario, Gianmario Voria, Viviana Pentangelo, Maria Tortorella, Arnaldo Sgueglia, Claudio Di Sipio, Giordano d'Aloisio, Antinisca Di Marco
ESEM8
2022 Community Smell Detection and Refactoring in SLACK: The CADOCS Project
abstract
Software engineering is a human-centered activity involving various stakeholders with different backgrounds that have to communicate and collaborate to reach shared objectives. The emergence of conflicts among stakeholders may lead to undesired effects on software maintainability, yet it is often unavoidable in the long run. Community smells, i.e., sub-optimal communication and collaboration practices, have been defined to map recurrent conflicts among developers. While some community smell detection tools have been proposed in the recent past, these can be mainly used for research purposes because of their limited level of usability and user engagement. To facilitate a wider use of community smell-related information by practitioners, we present CADOCS, a client-server conversational agent that builds on top of a previous community smell detection tool proposed by Almarini et al. to (1) make it usable within a well-established communication channel like Slack and (2) augment it by providing initial support to software analytics instruments useful to diagnose and refactor community smells. We describe the features of the tool and the preliminary evaluation conducted to assess and improve robustness and usability.
Gianmario Voria, Viviana Pentangelo, Antonio Della Porta, Stefano Lambiase, Gemma Catolino, Fabio Palomba, Filomena Ferrucci
ICSME1