EDBT 2026 Demo / reviewers in the wild / expert
Enrico Fregnan
dblp:235/0216
· DBLP profile ↗
13ranked-venue papers
7as first author
10since 2021 · last 2025
0000-0002-6897-7665ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 11 · 6 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Leveraging LLMs Towards Assistant-based Support for Industrial Threat ModelsabstractThreat models contribute to strengthen the security of an enterprise system by listing the cybersecurity threats that might affect it as well as possible mitigations to such threats. Given the importance of these models, different tools have been devised to support creating or updating threat models. However, these tools focus mainly on the needs of cybersecurity experts, often overlooking non-expert users.To address this gap, in the first part of this paper we present a chat-based LLM assistant to answer both expert and non-expert queries on threat models. The assistant is based on a microservice architecture and uses Retrieval-Augmented Generation (RAG) to extract relevant information from an industrial threat model.Guaranteeing the reliability of an LLM’s answers in sensitive fields such as cybersecurity is paramount. For this reason, in the second part of this paper, we evaluate the LLM assistant using (i) Bleu and Rouge score metrics, (ii) human evaluation, and (iii) an automatic LLM-as-a-judge approach.The evaluation, conducted on 275 test cases, confirms the accuracy of the answers provided by our LLM-based threat model assistant. Additionally, the results of the human and LLM-as-a-judge evaluations are consistent. This corroborates the effectiveness of automatic LLM-as-a-judge methods in assessing the performance of LLM-based industrial solutions. Enrico Fregnan, Christian Göttel, Balz Maag, Abdallah Dawoud, Georgios Nakas |
ETFA | 1 |
| 2025 | Correction to: A preliminary investigation on using multi-task learning to predict change performance in code reviews
Enrico Fregnan, Fernando Petrulio, Linda Di Geronimo, Alberto Bacchelli |
Empir. Softw. Eng. | 1 |
| 2023 | Graph-based visualization of merge requests for code reviewabstractCode review is a software development practice aimed at assessing code quality, finding defects, and sharing knowledge among developers. Despite its wide adoption, code review is a challenging task for developers, who often struggle to understand the content of a review change-set. Visualization techniques represent a promising approach to support reviewers. In this paper we present a new visualization approach that displays classes and methods in review changes as nodes in a graph. Then, we implemented our graph-based approach in a tool (ReviewVis) and performed a two-step feedback collection phase to assess the developers’ perceptions on the tool’s benefits through (1) an in-company study with nine professional software developers and (2) an online survey with 37 participants. Given the positive results obtained by this first evaluation, we performed a second survey with 31 participants with a specific focus on supporting developers’ understanding of a review change-set. The collected feedback showed that the developers indeed perceive that ReviewVis can help them navigate and understand the changes under review. The results achieved also indicate possible future paths to use software visualization for code review. Data and Materials: https://doi.org/10.5281/zenodo.7047993 Enrico Fregnan, Josua Fröhlich, Davide Spadini, Alberto Bacchelli |
J. Syst. Softw. | 1 |
| 2022 | An Exploratory Study on Regression VulnerabilitiesabstractBackground: Security regressions are vulnerabilities introduced in a previously unaffected software system. They often happen as a result of code changes (e.g., a bug fix) and can have severe effects. Larissa Braz, Enrico Fregnan, Vivek Arora, Alberto Bacchelli |
ESEM | 2 |
| 2022 | First come first served: the impact of file position on code reviewabstractThe most popular code review tools (e.g., Gerrit and GitHub) present the files to review sorted in alphabetical order. Could this choice or, more generally, the relative position in which a file is presented bias the outcome of code reviews? We investigate this hypothesis by triangulating complementary evidence in a two-step study. First, we observe developers’ code review activity. We analyze the review comments pertaining to 219,476 Pull Requests (PRs) from 138 popular Java projects on GitHub. We found files shown earlier in a PR to receive more comments than files shown later, also when controlling for possible confounding factors: e.g., the presence of discussion threads or the lines added in a file. Second, we measure the impact of file position on defect finding in code review. Recruit- ing 106 participants, we conduct an online controlled experiment in which we measure participants’ performance in detecting two unrelated defects seeded into two different files. Participants are assigned to one of two treatments in which the position of the defective files is switched. For one type of defect, participants are not affected by its file’s position; for the other, they have 64% lower odds to identify it when its file is last as opposed to first. Overall, our findings provide evidence that the relative position in which files are presented has an impact on code reviews’ outcome; we discuss these results and implications for tool design and code review. Enrico Fregnan, Larissa Braz, Marco D'Ambros, Gül Çalikli, Alberto Bacchelli |
ESEC/SIGSOFT FSE | 1 |
| 2022 | The evolution of the code during review: an investigation on review changesabstractAbstract Code review is a software engineering practice in which reviewers manually inspect the code written by a fellow developer and propose any change that is deemed necessary or useful. The main goal of code review is to improve the quality of the code under review. Despite the widespread use of code review, only a few studies focused on the investigation of its outcomes, for example, investigating the code changes that happen to the code under review. The goal of this paper is to expand our knowledge on the outcome of code review while re-evaluating results from previous work. To this aim, we analyze changes that happened during the review process, which we define as review changes. Considering three popular open-source software projects, we investigate the types of review changes (based on existing taxonomies) and what triggers them; also, we study which code factors in a code review are most related to the number of review changes. Our results show that the majority of changes relate to evolvability concerns, with a strong prevalence of documentation and structure changes at type-level. Furthermore, differently from past work, we found that the majority of review changes are not triggered by reviewers’ comments. Finally, we find that the number of review changes in a code review is related to the size of the initial patch as well as the new lines of code that it adds. However, other factors, such as lines deleted or the author of the review patchset, do not always show an empirically supported relationship with the number of changes. Enrico Fregnan, Fernando Petrulio, Alberto Bacchelli |
Empir. Softw. Eng. | 1 |
| 2022 | What happens in my code reviews? An investigation on automatically classifying review changesabstractAbstract Code reviewing is a widespread practice used by software engineers to maintain high code quality. To date, the knowledge on the effect of code review on source code is still limited. Some studies have addressed this problem by classifying the types of changes that take place during the review process (a.k.a. review changes), as this strategy can, for example, pinpoint the immediate effect of reviews on code. Nevertheless, this classification (1) is not scalable, as it was conducted manually, and (2) was not assessed in terms of how meaningful the provided information is for practitioners. This paper aims at addressing these limitations: First, we investigate to what extent a machine learning-based technique can automatically classify review changes. Then, we evaluate the relevance of information on review change types and its potential usefulness, by conducting (1) semi-structured interviews with 12 developers and (2) a qualitative study with 17 developers, who are asked to assess reports on the review changes of their project. Key results of the study show that not only it is possible to automatically classify code review changes, but this information is also perceived by practitioners as valuable to improve the code review process. Data and materials: 10.5281/zenodo.5592254 Enrico Fregnan, Fernando Petrulio, Linda Di Geronimo, Alberto Bacchelli |
Empir. Softw. Eng. | 1 |
| 2022 | Do explicit review strategies improve code review performance? Towards understanding the role of cognitive loadabstractAbstract Code review is an important process in software engineering – yet, a very expensive one. Therefore, understanding code review and how to improve reviewers’ performance is paramount. In the study presented in this work, we test whether providing developers with explicit reviewing strategies improves their review effectiveness and efficiency. Moreover, we verify if review guidance lowers developers’ cognitive load. We employ an experimental design where professional developers have to perform three code review tasks. Participants are assigned to one of three treatments: ad hoc reviewing, checklist, and guided checklist. The guided checklist was developed to provide an explicit reviewing strategy to developers. While the checklist is a simple form of signaling (a method to reduce cognitive load), the guided checklist incorporates further methods to lower cognitive demands of the task such as segmenting and weeding. The majority of the participants are novice reviewers with low or no code review experience. Our results indicate that the guided checklist is a more effective aid for a simple review,while the checklist supports reviewers’ efficiency and effectiveness in a complex task. However, we did not identify a strong relationship between the guidance provided and code review performance. The checklist has the potential to lower developers’ cognitive load, but higher cognitive load led to better performance possibly due to the generally low effectiveness and efficiency of the study participants. Data and materials: 10.5281/zenodo.5653341 . Registered report: 10.17605/OSF.IO/5FPTJ . Pavlína Wurzel Gonçalves, Enrico Fregnan, Tobias Baum, Kurt Schneider, Alberto Bacchelli |
Empir. Softw. Eng. | 2 |
| 2021 | Why Don't Developers Detect Improper Input Validation? '; DROP TABLE Papers; -abstractImproper Input Validation (IIV) is a software vulnerability that occurs when a system does not safely handle input data. Even though IIV is easy to detect and fix, it still commonly happens in practice. In this paper, we study to what extent developers can detect IIV and investigate underlying reasons. This knowledge is essential to better understand how to support developers in creating secure software systems. We conduct an online experiment with 146 participants, of which 105 report at least three years of professional software development experience. Our results show that the existence of a visible attack scenario facilitates the detection of IIV vulnerabilities and that a significant portion of developers who did not find the vulnerability initially could identify it when warned about its existence. Yet, a total of 60 participants could not detect the vulnerability even after the warning. Other factors, such as the frequency with which the participants perform code reviews, influence the detection of IIV. Preprint: https://arxiv.org/abs/2102.06251. Data and materials: https://doi.org/10.5281/zenodo.3996696. Larissa Braz, Enrico Fregnan, Gül Çalikli, Alberto Bacchelli |
ICSE | 2 |
| 2021 | ChangeViz: Enhancing the GitHub Pull Request Interface with Method Call InformationabstractCode review is a widely adopted software development practice aimed at finding defects, improving software quality, and transferring knowledge among developers. Performing an effective code review is a challenging task for developers. Two of the main challenges reviewers face are (1) understanding the content of a review change-set and (2) assessing the impact of a change on the codebase. Visualization techniques can be used to increase developers’ understanding of a changeset to review and its context. However, only a few attempts have been made to apply visualization to code review.In this paper, we present a novel approach we devised to support developers in understanding GitHub pull requests. Our approach expands the GitHub interface with two lateral bars to let developers navigate to the definition/uses of the methods in the changeset under review.We evaluated our approach’s interface through (1) interviews with eight developers and (2) a survey with 12 participants. Based on the results of this evaluation, we implemented our approach in a web-based tool, ChangeViz.Pre-print, data and materials, and demo video: https://doi.org/10.5281/zenodo.5175927. Lorenzo Gasparini, Enrico Fregnan, Larissa Braz, Tobias Baum, Alberto Bacchelli |
VISSOFT | 2 |
| 2020 | UI Dark Patterns and Where to Find Them: A Study on Mobile Applications and User PerceptionabstractA Dark Pattern (DP) is an interface maliciously crafted to deceive users into performing actions they did not mean to do. In this work, we analyze Dark Patterns in 240 popular mobile apps and conduct an online experiment with 589 users on how they perceive Dark Patterns in such apps. The results of the analysis show that 95% of the analyzed apps contain one or more forms of Dark Patterns and, on average, popular applications include at least seven different types of deceiving interfaces. The online experiment shows that most users do not recognize Dark Patterns, but can perform better in recognizing malicious designs if informed on the issue. We discuss the impact of our work and what measures could be applied to alleviate the issue. Linda Di Geronimo, Larissa Braz, Enrico Fregnan, Fabio Palomba, Alberto Bacchelli |
CHI | 3 |
| 2020 | Do Explicit Review Strategies Improve Code Review Performance?abstractContext: Code review is a fundamental, yet expensive part of software engineering. Therefore, research on understanding code review and its efficiency and performance is paramount. Pavlína Wurzel Gonçalves, Enrico Fregnan, Tobias Baum, Kurt Schneider, Alberto Bacchelli |
MSR | 2 |
| 2019 | A survey on software coupling relations and tools
Enrico Fregnan, Tobias Baum, Fabio Palomba, Alberto Bacchelli |
Inf. Softw. Technol. | 1 |