VLDB 2026 Research / reviewers in the wild / expert
Mahi Begoug
dblp:360/8476
· DBLP profile ↗
11ranked-venue papers
6as first author
11since 2021 · last 2026
0009-0007-5914-9968ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 11 · 6 first-author · 11 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Single Code Changes: An Empirical Study of Topic-Based Code Review Practices in Gerrit for OpenStack
Moataz Chouchen, Mahi Begoug, Ali Ouni 0001 |
MSR | 2 |
| 2026 | On the Reliability of Agentic AI in Continuous Integration PipelinesabstractAgentic AI systems powered by Large Language Models (LLMs) are increasingly used to autonomously contribute code in modern software development. While prior work has shown that such systems can accelerate development tasks, their reliability and maintenance behavior in real-world Continuous Integration (CI) workflows remain poorly understood. In this study, we analyze 11,771 pull requests (PRs) from GitHub, including 7,619 agentic and 4,152 human-authored PRs, to investigate how agentic code behaves during CI workflows. We examine (1) CI failure rates at the pull-request level, (2) responsibility for introducing and fixing CI failures, and (3) time-to-fix at the commit level using fail–fix mappings. Our results show that human-authored CI fixes exhibit a median time to fix of 71.70 minutes, whereas AI agentic-authored CI fixes resolve failures nearly four times faster, with a median of 17.23 minutes. Our results show that agent-authored fixes resolve CI failures nearly four times faster than human fixes (median 17.23 vs. 71.70 minutes). However, agents introduce most CI failures (79.15%) while performing a smaller share of fixes (60.63%), indicating that human developers remain heavily involved in failure resolution despite faster agent responses. Moataz Chouchen, Jasem Khelifi, Mahi Begoug, Ali Ouni 0001, Mohammed Sayagh, Mohamed Aymen Saied |
MSR | 3 |
| 2026 | When AI Code Doesn't Stick: An Empirical Study on Reverted Changes Introduced by AI Coding AgentsabstractAgentic AI systems are increasingly integrated into software development workflows, contributing code alongside human developers. However, some AI-authored changes are later reverted, reflecting situations where agent-generated contributions are judged unsuitable after integration. This paper presents a large-scale empirical study of reverted changes introduced by AI coding agents to better understand the causes behind their rejection. We analyze 33,580 agentic pull requests comprising 86,315 commits authored by five major AI coding agents: Claude, Copilot, Cursor, Devin, and OpenAI Codex. Our results show that 2.66% of agentic pull requests contain at least one reverting commit, with substantial variation across agents, ranging from 0.7% for OpenAI Codex to 7.6% for GitHub Copilot indicating notable differences in code generation reliability. Through a manual analysis of 500 reverting commits, we derive a taxonomy comprising eight categories and 25 themes that explain why agent-generated code is reverted. The most common causes are unintended side effects (22.33%), overengineering (22.13%), functional incorrectness (17.71%), and dependency management problems (12.47%). Overall, our findings indicate that AI coding agents struggle primarily with scope management and contextual understanding, rather than purely functional defects. This study provides actionable guidance for practitioners, informs the design of human-AI collaboration workflows, and highlights priority areas for improving agentic code generation systems. Issam Oukhay, Mahi Begoug, Moataz Chouchen, Ali Ouni 0001 |
MSR | 2 |
| 2025 | How Do Infrastructure-as-Code Practitioners Update Their Dependencies? An Empirical Study on Terraform Module UpdatesabstractInfrastructure-as-Code (IaC) enables practitioners to configure and manage software infrastructure through machine-readable code files. Various IaC tools facilitate code reuse and modularity via IaC modules that act as dependencies. These modules are maintained by IaC providers to introduce new features, resolve bugs, or address security vulnerabilities. However, there is a limited understanding of how practitioners update their IaC module dependencies in their software projects, including updates frequency, delays, as well as motivations behind such updates. To fill this gap, this paper aims to understand current update practices in IaC module dependencies, focusing on Terraform (TF), being currently one of the most popular IaC tools. In particular, we investigate (i) the frequency in which IaC practitioners update their module dependencies, (ii) the technical lag phenomena, which represents the time that the infrastructure configurations remain outdated relative to their upstream modules, and (iii) the motivations that drive these updates. To achieve these, we conduct an empirical study on 13,490 TF-related commits from 131 open-source projects. Our results reveal that only 1.2% of the analyzed commits involve updating module dependencies. Furthermore, we observe an increasing technical lag from 2021 until 2024, reaching ten months on average by 2024. Then, we conduct a qualitative study using thematic analysis on code changes involving TF module dependencies updates to investigate practitioners’ motivations behind such updates. We identify that TF practitioners revolve around six main motivations, with IaC Ecosystem Compatibility, Security Vulnerabilities Fixes, and IaC Code Quality Improvement being the three most prevalent motivations. Our findings advocate that TF practitioners need customized IaC tool support for safe module dependency updates while addressing compatibility concerns. Mahi Begoug, Ali Ouni 0001, Moataz Chouchen |
MSR | 1 |
| 2025 | Understanding AWS Provider Dependency Updates in Infrastructure-As-Code: Empirical Study, Taxonomy, and InsightsabstractInfrastructure-as-Code (IaC) automates the configuration of cloud platforms through code. As business needs evolve, IaC files often become complex, containing hundreds of lines and multiple dependencies. These configurations rely on third-party providers to provision system infrastructure. Practitioners regularly update IaC code to align with evolving cloud provider specifications (i.e., AWS, GCP, Azure) and to address security issues or defects. Although prior work highlights the risks of outdated dependencies, it remains unclear whether IaC practitioners consistently update provider dependencies in accordance with official releases. To address this gap, we conduct a mixed-method empirical study focused on the Amazon Web Services (AWS) provider, one of the most widely used providers for provisioning cloud infrastructures. We analyze 23,404 Terraform (TF) related commits from 194 open-source TF projects, focusing on: (i) technical lag, which captures how long AWS provider dependencies remain unchanged in code; (ii) the frequency of dependency updates; (iii) the code review effort involved in updating AWS provider dependencies; and (iv) the motivations behind such updates. Our findings reveal that Terraform developers frequently rely on outdated provider versions, with the technical lag increasing steadily from 2017 to early 2025, reaching a monthly average of approximately 9 months by 2025. Quantitative analysis reveals that only 1.86% of TF-related commits involve updates to AWS provider dependencies, indicating that such updates are not a priority. Moreover, related code reviews are substantial, affecting a median of 7 files across multiple directories. Through thematic analysis, we identify nine key motivations for updating the AWS provider dependencies, with the top three being: Providers Dependency Management, Terraform Compatibility Management, and Security Management. These insights highlight a clear need for better support and tooling to help practitioners manage provider updates more effectively, minimizing disruption while modernizing infrastructure. We recommend adopting automated dependency management tools and improved update workflows to reduce technical lag and lower the cost of staying up to date. Mahi Begoug, Ali Ouni 0001, Jasem Khelifi |
IEEE Trans. Serv. Comput. | 1 |
| 2024 | How Do Infrastructure-as-Code Practitioners Update their Provider Dependencies? An Empirical Study on the AWS Provider
Mahi Begoug, Ali Ouni 0001 |
ICSOC (2) | 1 |
| 2024 | TerraMetrics: An Open Source Tool for Infrastructure-as-Code (IaC) Quality Metrics in TerraformabstractInfrastructure-as-Code (IaC) constitutes a pivotal DevOps methodology, leading edge of software deployment onto cloud platforms. IaC relies on source code files rather than manual configuration to manage the infrastructure of a software system. Terraform, an IaC tool and its declarative configuration language named HCL, has recently garnered considerable attention among IaC practitioners. Like other software artefacts, Terraform files could be affected by misconfigurations, faults, and smells. Therefore, DevOps practitioners might benefit from a quality assurance tool to help them perform quality assurance activities on Terrafrom artefacts. This paper introduces TerraMetrics, an open-source tool designed to characterize the quality of Terraform artefacts by providing a catalogue of 40 quality metrics. TerraMetrics leverages the Terraform Abstract Syntax Tree (AST) to extract the metric list, offering a potentially enduring solution compared to conventional regular expressions. This tool comprises three main components: (i) a parser transforming HCL code into an AST, (ii) visitors that traverse the AST nodes to extract the metrics, and (iii) collectors for storing the collected metrics in JSON format. The TerraMetrics tool is publicly available as an Open Source tool, with a demo video, at: https://github.com/stilab-ets/terametrics. Mahi Begoug, Moataz Chouchen, Ali Ouni 0001 |
ICPC | 1 |
| 2024 | Fine-Grained Just-In-Time Defect Prediction at the Block Level in Infrastructure-as-Code (IaC)abstractInfrastructure-as-Code (IaC) is an emerging software engineering practice that leverages source code to facilitate automated configuration of software systems' infrastructure. IaC files are typically complex, containing hundreds of lines of code and dependencies, making them prone to defects, which can result in breaking online services at scale. To help developers early identify and fix IaC defects, research efforts have introduced IaC defect prediction models at the file level. However, the granularity of the proposed approaches remains coarse-grained, requiring developers to inspect hundreds of lines of code in a file, while only a small fragment of code is defective. To alleviate this issue, we introduce a machine-learning-based approach to predict IaC defects at a fine-grained level, focusing on IaC blocks, i.e., small code units that encapsulate specific behaviours within an IaC file. We trained various machine learning algorithms based on a mixture of code, process, and change-level metrics. We evaluated our approach on 19 open-source projects that use Terraform, a widely used IaC tool. The results indicated that there is no single algorithm that consistently outperforms the others in 19 projects. Overall, among the six algorithms, we observed that the LightGBM model achieved a higher average of 0.21 in terms of MCC and 0.71 in terms of AUC. Models analysis reveals that the developer's experience and the relative number of added lines tend to be the most important features. Additionally, we found that blocks belonging to the most frequent types are more prone to defects. Our defect prediction models have also shown sensitivity to concept drift, indicating that IaC practitioners should regularly retrain their models. Mahi Begoug, Moataz Chouchen, Ali Ouni 0001, Eman Abdullah AlOmar, Mohamed Wiem Mkaouer |
MSR | 1 |
| 2024 | How Do So ware Developers Use ChatGPT? An Exploratory Study on GitHub Pull RequestsabstractNowadays, Large Language Models (LLMs) play a pivotal role in software engineering. Developers can use LLMs to address software development-related tasks such as documentation, code refactoring, debugging, and testing. ChatGPT, released by OpenAI, has become the most prominent LLM. In particular, ChatGPT is a cutting-edge tool for providing recommendations and solutions for developers in their pull requests (PRs). However, little is known about the characteristics of PRs that incorporate ChatGPT compared to those without it and what developers usually use it for. To this end, we quantitatively analyzed 243 PRs that listed at least one ChatGPT prompt against a representative sample of 384 PRs without any ChatGPT prompts. Our findings show that developers use ChatGPT in larger, time-consuming pull requests that are five times slower to be closed than PRs that do not use ChatGPT. Furthermore, we perform a qualitative analysis to build a taxonomy of the topics developers primarily address in their prompts. Our analysis results in a taxonomy comprising 8 topics and 32 sub-topics. Our findings highlight that ChatGPT is often used in review-intensive pull requests. Moreover, our taxonomy enriches our understanding of the developer's current applications of ChatGPT. Moataz Chouchen, Narjes Bessghaier, Mahi Begoug, Ali Ouni 0001, Eman Abdullah AlOmar, Mohamed Wiem Mkaouer |
MSR | 3 |
| 2024 | On the Prevalence, Co-occurrence, and Impact of Infrastructure-as-Code SmellsabstractIn modern software systems, Infrastructure-as-Code (IaC) tools play a pivotal role in automating the management of various infrastructure resources such as networks, databases, and services. This automation is done through code-based specification files, commonly known as IaC files. Similarly to other code files, IaC files can suffer from violations of established implementation and design standards, i.e., IaC smells. Although prior research has studied various aspects of traditional smells in non-IaC artifacts, there is little knowledge of how IaC smells are prevalent, co-occurring, and impacting the change and defect proneness of IaC code. To fill this gap, we conduct an empirical study encompassing 82 Puppet-based open-source projects. Our investigation focused on 12 types of IaC smells in both implementation and design levels. Our findings reveal that IaC smells do not manifest uniformly, as IaC smells that are particularly associated with modularity issues, exhibit high prevalence rates across projects. Additionally, we found that 74% of IaC files are smelly and over 52% of the smelly IaC files have at least two co-occurring IaC smells. Furthermore, our findings highlight that, on average, smelly IaC files are modified nearly 3.8 times, in terms of number of commits, more frequently than non-smelly IaC files. Furthermore, smelly IaC files are found to be 3.1 times more prone to larger code changes, in terms of code churn, than non-smelly IaC files. Additionally, we found that smelly IaC files are 3.3 times more prone to the introduction of defects that are likely to persist in 1.65 more commits before being fixed than non-smelly IaC files. These findings advocate developers to be more aware of IaC smells in their projects and consider their correction. Narjes Bessghaier, Mahi Begoug, Chemseddine Mebarki, Ali Ouni 0001, Mohammed Sayagh, Mohamed Wiem Mkaouer |
SANER | 2 |
| 2023 | What Do Infrastructure-as-Code Practitioners Discuss: An Empirical Study on Stack OverflowabstractBackground. Infrastructure-as-Code (IaC) is an emerging practice to manage cloud infrastructure resources for software systems. Modern software development has evolved to embrace IaC as a best practice for consistently provisioning and managing infrastructure using various tools such as Terraform and Ansible. However, recent studies highlighted that developers still encounter various challenges with IaC tools. Aims. We aim in this paper to understand the different challenges that developers encounter with IaC and analyze the trend of seeking assistance on Q&A platforms in the context of IaC. To this end, we conduct a large-scale empirical study investigating developers' discussions in Stack Overflow. Method. We first collect IaC-relevant tags on Stack Overflow, constituting a dataset that comprises 52,692 questions and 64,078 answers. Then, we group questions into specific topics using the Latent Dirichlet Allocation (LDA) method, which we optimize using a Genetic Algorithm (GA) for parameter's fine-tuning. Finally, to gain better insights, we analyze the identified topics based on different criteria such as popularity and difficulty. Results. Our findings reveal an average yearly increase of 150% in terms of IaC-related questions and 135% in terms of users between 2011 and 2022. Furthermore, we observe that IaC questions revolve around seven main topics: server configuration, policy configuration, networking, deployment pipelines, variable management, templating, and file management. Notably, we found that server configuration and file management are the most popular topics, i.e., the most discussed among IaC developers, while the deployment pipelines and templating topics are the most difficult. Conclusions. Our results shed light on IaC challenges that are often encountered by developers on popular Q&A platforms. These findings reveal important implications for practitioners seeking better support for IaC tools in real-world settings and for researchers to better understand the IaC community needs and further investigate IaC in different aspects. Mahi Begoug, Narjes Bessghaier, Ali Ouni 0001, Eman Abdullah AlOmar, Mohamed Wiem Mkaouer |
ESEM | 1 |