VLDB 2026 Research / reviewers in the wild / expert
Marco Castelluccio
dblp:167/4096
· DBLP profile ↗
11ranked-venue papers
3as first author
3since 2021 · last 2025
0000-0002-3285-5121ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 11 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Automated Generation of Issue-Reproducing Tests by Combining LLMs and Search-Based TestingabstractIssue-reproducing tests fail on buggy code and pass once a patch is applied, thus increasing developers’ confidence that the issue has been resolved and will not be re-introduced. However, past research has shown that developers often commit patches without such tests, making the automated generation of issue-reproducing tests an area of interest. We propose BLAST, a tool for automatically generating issue-reproducing tests from issue-patch pairs by combining LLMs and search-based software testing (SBST). For the LLM part, we complement the issue description and the patch by extracting relevant context through Git history analysis, static analysis, and SBST-generated tests. For the SBST part, we adapt SBST for generating issue-reproducing tests; the issue description and the patch are fed into the SBST optimization through an intermediate LLM-generated seed, which we deserialize into SBST-compatible form. BLAST successfully generates issue-reproducing tests for 151/426 (35.4%) of the issues from a curated Python benchmark, outperforming the state-of-the-art (23.5%).Additionally, to measure the real-world impact of BLAST, we built a GitHub bot that runs BLAST whenever a new pull request (PR) linked to an issue is opened, and if BLAST generates an issue-reproducing test, the bot proposes it as a comment in the PR. We deployed the bot in three open-source repositories for three months, gathering data from 32 PRs-issue pairs. BLAST generated an issue-reproducing test in 11 of these cases, which we proposed to the developers. By analyzing the developers’ feedback, we discuss challenges and opportunities for researchers and tool builders.Data and material: https://doi.org/10.5281/zenodo.16949042 Konstantinos Kitsios, Marco Castelluccio, Alberto Bacchelli |
ASE | 2 |
| 2024 | Predicting the Impact of Crashes Across Release ChannelsabstractSoftware maintenance faces a persistent challenge with crash bugs, especially across diverse release channels catering to distinct user bases. Nightly builds, favoured by enthusiasts, often reveal crashes that are cheaper to fix but may differ significantly from those in stable releases. In this paper, we emphasize the need for a data-driven solution to predict the impact of crashes happening on nightly channels once they are released to stable channels. We also list the challenges that need to be considered when approaching this problem. Suhaib Mujahid, Diego Costa 0001, Marco Castelluccio |
MSR | 3 |
| 2022 | Works for Me! Cannot Reproduce - A Large Scale Empirical Study of Non-reproducible Bugs
Mohammad Masudur Rahman 0001, Foutse Khomh, Marco Castelluccio |
Empir. Softw. Eng. | 3 |
| 2020 | Why are Some Bugs Non-Reproducible? : -An Empirical Investigation using Data Fusion-abstractSoftware developers attempt to reproduce software bugs to understand their erroneous behaviours and to fix them. Unfortunately, they often fail to reproduce (or fix) them, which leads to faulty, unreliable software systems. However, to date, only a little research has been done to better understand what makes the software bugs non-reproducible. In this paper, we conduct a multimodal study to better understand the non-reproducibility of software bugs. First, we perform an empirical study using 576 non-reproducible bug reports from two popular software systems (Firefox, Eclipse) and identify 11 key factors that might lead a reported bug to non-reproducibility. Second, we conduct a user study involving 13 professional developers where we investigate how the developers cope with non-reproducible bugs. We found that they either close these bugs or solicit for further information, which involves long deliberations and counter-productive manual searches. Third, we offer several actionable insights on how to avoid non-reproducibility (e.g., false-positive bug report detector) and improve reproducibility of the reported bugs (e.g., sandbox for bug reproduction) by combining our analyses from multiple studies (e.g., empirical study, developer study). Mohammad Masudur Rahman 0001, Foutse Khomh, Marco Castelluccio |
ICSME | 3 |
| 2019 | Understanding flaky tests: the developer's perspectiveabstractFlaky tests are software tests that exhibit a seemingly random outcome (pass or fail) despite exercising unchanged code. In this work, we examine the perceptions of software developers about the nature, relevance, and challenges of flaky tests. Moritz Eck, Fabio Palomba, Marco Castelluccio, Alberto Bacchelli |
ESEC/SIGSOFT FSE | 3 |
| 2019 | An empirical study of DLL injection bugs in the Firefox ecosystem
Marco Castelluccio, Foutse Khomh |
Empir. Softw. Eng. | 2 |
| 2019 | An empirical study of patch uplift in rapid release development pipelines
Marco Castelluccio, Foutse Khomh |
Empir. Softw. Eng. | 1 |
| 2018 | Why Did This Reviewed Code Crash? An Empirical Study of Mozilla FirefoxabstractCode review, i.e., the practice of having other team members critique changes to a software system, is a pillar of modern software quality assurance approaches. Although this activity aims at improving software quality, some high-impact defects, such as crash-related defects, can elude the inspection of reviewers and escape to the field, affecting user satisfaction and increasing maintenance overhead. In this research, we investigate the characteristics of crash-prone code, observing that such code tends to have high complexity and depend on many other classes. In the code review process, developers often spend a long time on and have long discussions about crash-prone code. We manually classify a sample of reviewed crash-prone patches according to their purposes and root causes. We observe that most crash-prone patches aim to improve performance, refactor code, add functionality, or fix previous crashes. Memory and semantic errors are identified as major root causes of the crashes. Our results suggest that software organizations should apply more scrutiny to these types of patches, and provide better support for reviewers to focus their inspection effort by using static analysis tools. Foutse Khomh, Shane McIntosh, Marco Castelluccio |
APSEC | 4 |
| 2018 | What makes a code change easier to review: an empirical investigation on code change reviewabilityabstractPeer code review is a practice widely adopted in software projects to improve the quality of code. In current code review practices, code changes are manually inspected by developers other than the author before these changes are integrated into a project or put into production. We conducted a study to obtain an empirical understanding of what makes a code change easier to review. To this end, we surveyed published academic literature and sources from gray literature (blogs and white papers), we interviewed ten professional developers, and we designed and deployed a reviewability evaluation tool that professional developers used to rate the reviewability of 98 changes. We find that reviewability is defined through several factors, such as the change description, size, and coherent commit history. We provide recommendations for practitioners and researchers. Public preprint [https://doi.org/10.5281/zenodo.1323659]; data and materials [https://doi.org/10.5281/zenodo.1323659]. Achyudh Ram, Anand Ashok Sawant, Marco Castelluccio, Alberto Bacchelli |
ESEC/SIGSOFT FSE | 3 |
| 2017 | Is it Safe to Uplift this Patch?: An Empirical Study on Mozilla FirefoxabstractIn rapid release development processes, patches that fix critical issues, or implement high-value features are often promoted directly from the development channel to a stabilization channel, potentially skipping one or more stabilization channels. This practice is called patch uplift. Patch uplift is risky, because patches that are rushed through the stabilization phase can end up introducing regressions in the code. This paper examines patch uplift operations at Mozilla, with the aim to identify the characteristics of uplifted patches that introduce regressions. Through statistical and manual analyses, we quantitatively and qualitatively investigate the reasons behind patch uplift decisions and the characteristics of uplifted patches that introduced regressions. Additionally, we interviewed three Mozilla release managers to understand organizational factors that affect patch uplift decisions and outcomes. Results show that most patches are uplifted because of a wrong functionality or a crash. Uplifted patches that lead to faults tend to have larger patch size, and most of the faults are due to semantic or memory errors in the patches. Also, release managers are more inclined to accept patch uplift requests that concern certain specific components, and- or that are submitted by certain specific developers. Marco Castelluccio, Foutse Khomh |
ICSME | 1 |
| 2017 | Automatically analyzing groups of crashes for finding correlationsabstractWe devised an algorithm, inspired by contrast-set mining algorithms such as STUCCO, to automatically find statistically significant properties (correlations) in crash groups. Many earlier works focused on improving the clustering of crashes but, to the best of our knowledge, the problem of automatically describing properties of a cluster of crashes is so far unexplored. This means developers currently spend a fair amount of time analyzing the groups themselves, which in turn means that a) they are not spending their time actually developing a fix for the crash; and b) they might miss something in their exploration of the crash data (there is a large number of attributes in crash reports and it is hard and error-prone to manually analyze everything). Our algorithm helps developers and release managers understand crash reports more easily and in an automated way, helping in pinpointing the root cause of the crash. The tool implementing the algorithm has been deployed on Mozilla's crash reporting service. Marco Castelluccio, Carlo Sansone, Luisa Verdoliva, Giovanni Poggi |
ESEC/SIGSOFT FSE | 1 |