EDBT 2026 Demo / reviewers in the wild / expert
Giordano d'Aloisio
dblp:304/9990
· DBLP profile ↗
12ranked-venue papers
8as first author
12since 2021 · last 2026
0000-0001-7388-890XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 11 · 7 first-author · 11 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SafeTune: Search-based Harmfulness Minimisation for Large Language Models
Giordano d'Aloisio, Giusy Annunziata, Zhiwei Fei, Antinisca Di Marco, Federica Sarro |
SSBSE | 1 |
| 2026 | How do generative models draw a software engineer? An empirical study on implicit bias of open-source image generation modelsabstractContext: Generative models are nowadays widely used to generate graphical content used for multiple purposes. However, it has been shown that the images generated by these models could reinforce societal biases already existing in specific contexts. The Software Engineering (SE) community is not immune to gender and ethnicity disparities, which could be amplified by the use of these models. Hence, if used without consciousness, artificially generated images could reinforce these biases in the SE domain. Objective: In this paper, we focus on understanding the implicit bias exposed by general-purpose open-source image generation models towards SE tasks. In addition, we investigate the extent to which it is possible to mitigate the bias by using prompt engineering techniques. Methods: We perform an extensive empirical evaluation of the implicit gender and ethnicity bias exposed by six popular open-source image generation models towards SE tasks. We obtain 20,160 images by feeding each model with three sets of prompts describing different software-related tasks: One set does not include any specification of the person performing the task, one set specifies that the person performing the task is a Software Engineer , and the last set explicitly request a fair representation of different genders and ethnicities. Next, we evaluate the gender and ethnicity disparities in the generated images. Results: The results indicate that all models exhibit a significant bias related to gender and ethnicity in SE tasks. Furthermore, we demonstrate that prompt engineering effectively reduces gender bias in only one of the six models; however, none of the models achieves fair representation with respect to ethnicity. Conclusion: The results of our analysis highlight serious concerns about the adoption of these models to generate content for SE tasks and open the field for future research on bias mitigation in this context. Giordano d'Aloisio, Tosin Fadahunsi, Antinisca Di Marco, Federica Sarro |
Inf. Softw. Technol. | 1 |
| 2026 | How fair are we? From conceptualization to automated assessment of fairness definitionsabstractAbstract Fairness is a critical concept in ethics and social domains, but it is also a challenging property to engineer in software systems. With the increasing use of machine learning in software systems, researchers have been developing techniques to assess the fairness of software systems automatically. Nonetheless, many of these techniques rely upon pre-established fairness definitions, metrics, and criteria, which may fail to encompass the wide-ranging needs and preferences of users and stakeholders. To overcome this limitation, we propose a novel approach, called MODNESS, that enables users to customize and define their fairness concepts using a dedicated modeling environment. Our approach guides the user through the definition of new fairness concepts also in emerging domains, and the specification and composition of metrics for its evaluation through a dedicated domain-specific language. Ultimately, MODNESS generates the source code to implement fair assessment based on these custom definitions. In addition, we elucidate the process we followed to collect and analyze relevant literature on fairness assessment in software engineering (SE). We compare MODNESS with the selected approaches and evaluate how they support the distinguishing features identified by our study. Our findings reveal that i) most of the current approaches do not support user-defined fairness concepts; ii) our approach can cover additional application domains not addressed by currently available tools, e.g., mitigating bias in recommender systems for software engineering and Arduino software component recommendations; iii) MODNESS demonstrates the capability to overcome the limitations of the only two other model-driven engineering-based approaches for fairness assessment. Giordano d'Aloisio, Claudio Di Sipio, Antinisca Di Marco, Davide Di Ruscio |
Softw. Syst. Model. | 1 |
| 2025 | Investigating the Role of LLMs Hyperparameter Tuning and Prompt Engineering to Support Domain Modeling
Vladyslav Bulhakov, Giordano d'Aloisio, Claudio Di Sipio, Antinisca Di Marco, Davide Di Ruscio |
SEAA | 2 |
| 2025 | On the Compression of Language Models for Code: An Empirical Study on CodeBERTabstractLanguage models have proven successful across a wide range of software engineering tasks, but their significant computational costs often hinder their practical adoption. To address this challenge, researchers have begun applying various compression strategies to improve the efficiency of language models for code. These strategies aim to optimize inference latency and memory usage, though often at the cost of reduced model effectiveness. However, there is still a significant gap in understanding how these strategies influence the efficiency and effectiveness of language models for code. Here, we empirically investigate the impact of three well-known compression strategies - knowledge distillation, quantization, and pruning - across three different classes of software engineering tasks: vulnerability detection, code summarization, and code search. Our findings reveal that the impact of these strategies varies greatly depending on the task and the specific compression method employed. Practitioners and researchers can use these insights to make informed decisions when selecting the most appropriate compression strategy, balancing both efficiency and effectiveness based on their specific needs. Giordano d'Aloisio, Luca Traini, Federica Sarro, Antinisca Di Marco |
SANER | 1 |
| 2025 | Towards early detection of algorithmic bias from dataset's bias symptoms: An empirical studyabstractThe rise of AI software has made fairness auditing essential, particularly where biased decisions have serious impacts. This entails identifying sensitive variables and calculating fairness metrics based on predictions from a baseline model. Since model training is computationally intensive, recent research focuses on early bias assessment to detect bias before extensive training starts. This paper presents an empirical study to evaluate how dataset statistics, named bias symptoms , can assist in the early identification of variables that may lead to bias in the system. The aim of this study is to avoid training a machine learning model before assessing - and, in case, mitigating - its bias, thus increasing the sustainability of the development process. We first identify a bias symptoms dataset, employing 24 datasets from diverse application domains commonly used in fairness auditing. Through extensive empirical analysis, we investigate the ability of these bias symptoms to predict variables associated with bias under three fairness definitions. Our results demonstrate that bias symptoms are effective in supporting early predictions of bias-inducing variables under specific fairness definitions. These findings offer valuable insights for practitioners and researchers, encouraging further exploration in developing methods for proactive bias mitigation involving bias symptoms. Giordano d'Aloisio, Claudio Di Sipio, Antinisca Di Marco, Davide Di Ruscio |
Inf. Softw. Technol. | 1 |
| 2024 | FRINGE: context-aware FaiRness engineerING in complex software systEmsabstractMachine learning (ML) is essential in modern technology, driving complex data-driven decisions. By 2025, daily data generation will exceed 463 exabytes, increasing ML’s influence and ethical risks of data exploitation and discrimination. The European Union’s Artificial Intelligence Act highlights the need for ethical AI solutions. Fabio Palomba, Andrea Di Sorbo, Davide Di Ruscio, Filomena Ferrucci, Gemma Catolino, Giammaria Giordano, Dario Di Dario, Gianmario Voria, Viviana Pentangelo, Maria Tortorella, Arnaldo Sgueglia, Claudio Di Sipio, Giordano d'Aloisio, Antinisca Di Marco |
ESEM | 13 |
| 2024 | Exploring LLM-Driven Explanations for Quantum AlgorithmsabstractBackground: Quantum computing is a rapidly growing new programming paradigm that brings significant changes to the design and implementation of algorithms. Understanding quantum algorithms requires knowledge of physics and mathematics, which can be challenging for software developers. Aims: In this work, we provide a first analysis of how LLMs can support developers’ understanding of quantum code. Method: We empirically analyse and compare the quality of explanations provided by three widely adopted LLMs (Gpt3.5, Llama2, and Tinyllama) using two different human-written prompt styles for seven state-of-the-art quantum algorithms. We also analyse how consistent LLM explanations are over multiple rounds and how LLMs can improve existing descriptions of quantum algorithms. Results: Llama2 provides the highest quality explanations from scratch, while Gpt3.5 emerged as the LLM best suited to improve existing explanations. In addition, we show that adding a small amount of context to the prompt significantly improves the quality of explanations. Finally, we observe how explanations are qualitatively and syntactically consistent over multiple rounds. Conclusions: This work highlights promising results, and opens challenges for future research in the field of LLMs for quantum code explanation. Future work includes refining the methods through prompt optimisation and parsing of quantum code explanations, as well as carrying out a systematic assessment of the quality of explanations. Giordano d'Aloisio, Sophie Fortz, Carol Hanna, Daniel Fortunato, Avner Bensoussan, Eñaut Mendiluze, Federica Sarro |
ESEM | 1 |
| 2024 | GreenStableYolo: Optimizing Inference Time and Image Quality of Text-to-Image Generation
Jingzhi Gong, Giordano d'Aloisio, Zishuo Ding, Yulong Ye, William B. Langdon, Federica Sarro |
SSBSE | 3 |
| 2024 | Uncovering gender gap in academia: A comprehensive analysis within the software engineering communityabstractGender gap in education has gained considerable attention in recent years, as it carries profound implications for the academic community. However, while the problem has been tackled from a student perspective, research is still lacking from an academic point of view. In this work, our main objective is to address this unexplored area by shedding light on the intricate dynamics of gender gap within the Software Engineering (SE) community. To this aim, we first review how the problem of gender gap in the SE community and in academia has been addressed by the literature so far. Results show that men in SE build more tightly-knit clusters but less global co-authorship relations than women, but the networks do not exhibit homophily. Concerning academic promotions, the Software Engineering community presents a higher bias in promotions to Associate Professors and a smaller bias in promotions to Full Professors than the overall Informatics community. Andrea D'Angelo, Giordano d'Aloisio, Francesca Marzi, Antinisca Di Marco, Giovanni Stilo |
J. Syst. Softw. | 2 |
| 2023 | Democratizing Quality-Based Machine Learning Development through Extended Feature ModelsabstractAbstract ML systems have become an essential tool for experts of many domains, data scientists and researchers, allowing them to find answers to many complex business questions starting from raw datasets. Nevertheless, the development of ML systems able to satisfy the stakeholders’ needs requires an appropriate amount of knowledge about the ML domain. Over the years, several solutions have been proposed to automate the development of ML systems. However, an approach taking into account the new quality concerns needed by ML systems (like fairness, interpretability, privacy, and others) is still missing. In this paper, we propose a new engineering approach for the quality-based development of ML systems by realizing a workflow formalized as a Software Product Line through Extended Feature Models to generate an ML System satisfying the required quality constraints. The proposed approach leverages an experimental environment that applies all the settings to enhance a given Quality Attribute, and selects the best one. The experimental environment is general and can be used for future quality methods’ evaluations. Finally, we demonstrate the usefulness of our approach in the context of multi-class classification problem and fairness quality attribute. Giordano d'Aloisio, Antinisca Di Marco, Giovanni Stilo |
FASE | 1 |
| 2023 | Debiaser for Multiple Variables to enhance fairness in classification tasksabstractNowadays assuring that search and recommendation systems are fair and do not apply discrimination among any kind of population has become of paramount importance. This is also highlighted by some of the sustainable development goals proposed by the United Nations. Those systems typically rely on machine learning algorithms that solve the classification task. Although the problem of fairness has been widely addressed in binary classification, unfortunately, the fairness of multi-class classification problem needs to be further investigated lacking well-established solutions. For the aforementioned reasons, in this paper, we present the Debiaser for Multiple Variables (DEMV), an approach able to mitigate unbalanced groups bias (i.e., bias caused by an unequal distribution of instances in the population) in both binary and multi-class classification problems with multiple sensitive variables. The proposed method is compared, under several conditions, with a set of well-established baselines using different categories of classifiers. At first we conduct a specific study to understand which is the best generation strategies and their impact on DEMV’s ability to improve fairness. Then, we evaluate our method on a heterogeneous set of datasets and we show how it overcomes the established algorithms of the literature in the multi-class classification setting and in the binary classification setting when more than two sensitive variables are involved. Finally, based on the conducted experiments, we discuss strengths and weaknesses of our method and of the other baselines. Giordano d'Aloisio, Andrea D'Angelo, Antinisca Di Marco, Giovanni Stilo |
Inf. Process. Manag. | 1 |