VLDB 2026 Research / reviewers in the wild / expert
Mohammed Sayagh
dblp:171/5853
· DBLP profile ↗
36ranked-venue papers
6as first author
28since 2021 · last 2026
0000-0002-2724-0034ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 35 · 6 first-author · 27 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Understanding the Rejection of Fixes Generated by Agentic Pull Requests - Insights from the AIDev DatasetabstractAI coding agents are increasingly used to generate pull requests (PRs) that propose code fixes in software projects. From a first exploration of the AIDev dataset, we find that 46.41% of the fixes proposed by the agents Copilot, Devin, Cursor, and Claude are rejected. This represents a significant amount of wasted resources that require human reviews, verifications, and running tests and validations for fixes that are merely discarded. Our goal in this paper is to understand the failure modes of AI-agents, an understanding that is crucial for better integrating AI-agents as efficient teammates. In this paper, we conduct a qualitative study on a representative sample of 306 non-merged pull requests created or co-authored by the agents mentioned earlier, followed by a quantitative analysis of the reasons for rejection. Our qualitative findings identify 14 reasons divided into four high-level categories for rejecting AI-agent fixes. We observe that developers can reject fixes due to fixes whose implementation is incorrect (e.g., incomplete, wrong approach), fixes that do not pass the continuous integration (CI) pipelines and fail tests, fixes for which the agent is unable to perform the implementation (e.g., no code generated, sessions lost), and fixes whose priority is low. Our results shed light on the importance of better guiding the model at these levels: (1) proposing hints about the approach to follow for fixing an issue, (2) outlining constraints or limitations regarding the approaches that should not be taken, and (3) instructing the agent on how to validate the implementation through CI pipelines and without introducing a breaking change. Our results suggest the need for good prioritization of tasks so that generated fixes do not lead to wasted human review efforts or wasted agent resources (e.g., tokens, compute, or allowed number of requests). Mahmoud Abujadallah, Ali Arabat, Mohammed Sayagh |
MSR | 3 |
| 2026 | Toward Instructions-as-Code: Understanding the Impact of Instruction Files on Agentic Pull RequestsabstractAI-agents (e.g., GitHub Copilot) collaborate as teammates in different software engineering tasks, including code generation proposed through pull requests (Agentic-PRs). For better agent efficiency, developers create instruction files that guide the AI-agents, including how to navigate the project, locate the right components, run tests, respect best practices, and more. In this paper, we investigate the relationship between the creation of these instructions and the performance of AI-agents in creating better pull requests, which have a higher chance of success (i.e., the merge rate), address more complex tasks (e.g., code churn), and require less effort to be merged (e.g., time to merge). To this end, we analyze 15,549 agentic PRs from 148 projects in the AIDev dataset. Using the three dimensions, we compare each project before and after the creation of the instruction files. We find that specifying instructions for AI-agents does not necessarily lead to better results. With the instruction files, 27.7% of the projects increased their merge rate by at least 20%, while 26.35% decreased it. The same observation is seen with the amount of changes (e.g., code churn, number of modified files) and with the efforts to merge an agentic PR (e.g., merge time and number of comments). From a first exploration, we find that projects that managed to increase their merge rate have substantially longer instruction files, which are also well structured into a higher number of sections and sub-sections. Our results motivate the need for research to assist practitioners in framing the development of instruction files as a software engineering activity (aka, Instructions-as-Code). Ali Arabat, Mohammed Sayagh |
MSR | 2 |
| 2026 | Characterizing Self-Admitted Technical Debt Generated by AI Coding AgentsabstractLarge Language Models (LLMs) are increasingly used through autonomous agents (e.g., Copilot, Cursor, Devin, Claude) to perform complex software development tasks. However, little is known about how these agents introduce and document technical debt through Self-Admitted Technical Debt (SATD) comments. Understanding SATD in AI-generated code is critical, as such comments explicitly reveal acknowledged limitations and deferred fixes that affect long-term maintenance. In this study, we quantitatively and qualitatively analyze 525 SATD comments authored by AI agents using the AIDev dataset. Our results show that AI-generated SATD is slightly more technically detailed than human-authored SATD, yet both often describe problems without clear guidance on resolution. Through thematic analysis, we identify 34 SATD topics grouped into 10 categories, with AI agents predominantly documenting requirement- and design-related debt. While many SATD topics overlap between AI and humans, our taxonomy reveals new debt categories and emphases specific to AI-authored SATD, particularly related to infrastructure, pipelines, dependency management, and requirement interpretation driven by developer prompts. Overall, our findings suggest that AI- and human-authored SATD share common characteristics but differ in expression and focus, highlighting the need for deeper investigation into how agentic systems communicate and manage technical debt. Zaki Brahmi, Ali Ouni 0001, Mohammed Sayagh, Mohamed Aymen Saied |
MSR | 3 |
| 2026 | On the Reliability of Agentic AI in Continuous Integration PipelinesabstractAgentic AI systems powered by Large Language Models (LLMs) are increasingly used to autonomously contribute code in modern software development. While prior work has shown that such systems can accelerate development tasks, their reliability and maintenance behavior in real-world Continuous Integration (CI) workflows remain poorly understood. In this study, we analyze 11,771 pull requests (PRs) from GitHub, including 7,619 agentic and 4,152 human-authored PRs, to investigate how agentic code behaves during CI workflows. We examine (1) CI failure rates at the pull-request level, (2) responsibility for introducing and fixing CI failures, and (3) time-to-fix at the commit level using fail–fix mappings. Our results show that human-authored CI fixes exhibit a median time to fix of 71.70 minutes, whereas AI agentic-authored CI fixes resolve failures nearly four times faster, with a median of 17.23 minutes. Our results show that agent-authored fixes resolve CI failures nearly four times faster than human fixes (median 17.23 vs. 71.70 minutes). However, agents introduce most CI failures (79.15%) while performing a smaller share of fixes (60.63%), indicating that human developers remain heavily involved in failure resolution despite faster agent responses. Moataz Chouchen, Jasem Khelifi, Mahi Begoug, Ali Ouni 0001, Mohammed Sayagh, Mohamed Aymen Saied |
MSR | 5 |
| 2026 | MLStractor: LLM-Powered Search-Based Monolith-to-Microservice Decomposition
Ilyes Kasdallah, Mostafa Anouar Ghorab, Oussama Jebbar, Khaled Sellami, Mohammed Sayagh, Ali Ouni 0001, Mohamed Aymen Saied |
SSBSE | 5 |
| 2026 | A recommendation system for predicting dependencies among software changes insights from an empirical study on OpenStack
Ali Arabat, Mohammed Sayagh, Jameleddine Hassine |
Empir. Softw. Eng. | 2 |
| 2026 | Code-level challenges and opportunities in hardware description languages: Insights from StackOverflow discussions
Kais Belwafi, Mohammed Sayagh, Ali Ouni 0001 |
Integr. | 2 |
| 2026 | A search-based file recommendation approach for infrastructure-as-code evolution
Narjes Bessghaier, Ali Ouni 0001, Mohammed Sayagh, Mohamed Wiem Mkaouer |
J. Syst. Softw. | 3 |
| 2025 | An Empirical Study on Hugging Face Trends, Topics and Challenges on Stack OverflowabstractHugging Face (HF) has emerged as a pivotal platform for the Machine Learning (ML) community, functioning as a central hub where developers collaborate, share models, and exchange datasets. By offering a vast repository of pre-trained models (PTMs), HF has democratized access to advanced ML resources, promoting model reuse and accelerating the development of ML-based systems. Despite its rapid adoption in recent years, there remains a limited understanding of the challenges developers encounter when working with HF in general and PTMs in particular. Understanding these challenges is crucial for guiding future research and developing support strategies for the software engineering community. Consequently, in this study we investigate HF-related Stack Overflow (SO) posts, one of the most popular discussion platforms for developers, to uncover the relevance of the topics, key challenges, and trends in HF-related discussions. This understanding will help future studies and the HF community improve the use of HF by focusing on the challenges developers face according to the prevalence and complexity of each of these challenges. To do so, we apply a topic modeling technique to categorize the topics discussed in SO posts that are related to HF. We then assess the popularity and difficulty of these topics to gain deeper insight into the specific challenges developers encounter. Our findings reveal an average annual growth rate of 31.3% in the number of HF-related questions on SO from 2019 to 2024. Furthermore, we identify eight major topics, with the usage and understanding of large language models (LLMs) being the most popular, while the distributed computing and resource management of PTMs stands out as the most challenging topic for developers. Hatem Feki, Manel Abdellatif, Mohammed Sayagh |
COMPSAC | 3 |
| 2025 | An Empirical Study on Microservices Deployment Trends, Topics and Challenges in Stack OverflowabstractMicroservices architecture is increasingly adopted in modern software projects. Microservices deployment is often managed by tools like Spring Cloud, Consul, and Docker. Although there is existing research on microservices, practical deployment challenges are still under-explored, impacting the efficiency and success of applications. In this paper, we aim to identify and understand the challenges developers encounter with microservices deployment. We analyze trends in help requests on Stack Overflow, one of the most popular Q&A platforms for developers, to identify and categorize these challenges and highlight the most popular and difficult ones. First, we examined 1,214 Stack Overflow posts related to microservices deployment using topic modelling based on the BERTopic method to extract and analyze challenge topics. To obtain a more comprehensive understanding, we also analyzed the identified topics according to their popularity and difficulty. Our results reveal that discussions related to microservices deployment vary over time from 2013 to 2023. We identified nine distinct topics related to microservices deployment challenges, including deployment strategies, data management, composition and discovery, containerization, configuration, and orchestration in Kubernetes, security management, CI/CD pipeline automation, exposure to external clients, and post-deployment monitoring. Results reveal that microservices containerization is the most popular topic that poses numerous challenges to many users, with 2,148 average views and a 3.19 average score. While composition and discovery and post deployment monitoring are the most challenging topics, with 78 % of questions on post deployment monitoring lacking accepted answers, and 28 % of questions about composition and discovery remaining unanswered. This study identifies critical areas in microservices deployment that need further investigation, particularly, difficult and popular ones. Amina Bouaziz, Mohamed Aymen Saied, Mohammed Sayagh, Ali Ouni 0001, Mohamed Wiem Mkaouer |
SANER | 3 |
| 2025 | GHAminer: An Open Source Tool to Extract GitHub Actions Build MetricsabstractGitHub Actions (GHA) has become among the most popular Continuous Integration (CI) platforms in open-source software (OSS) and commercial projects. Collecting such build data remains crucial for practitioners and researchers to allow build performance monitoring, optimization and improvement. However, mining GHA builds to collect build-related data and metrics remains challenging and time-consuming. This paper introduces GHAminer, an open-source tool designed to collect build-related metrics for GitHub Actions. GHAminer covers various aspects of data such as the build-related code changes and tests, the build duration and status (e.g., passed, failed, timeout, etc.), and repository metadata, which would be useful for practitioners and researchers to make data-driven decisions to enhance CI efficiency and quality. The tool has a modular architecture that supports efficient data extraction with minimal API load. Specifically, it consists of a set of modules that are related to repository information collection, build analysis, commit history analysis, and build log parsing. We evaluate the performance of GHAminer on a representative sample of 3,151 OSS projects. Results show that GHAminer is efficient in handling projects of various sizes with relatively stable performance to collect build data for larger projects. GHAminer is publicly available with a demo video at: https:lIgithub.com/stilab-ets/GHAminer Jasem Khelifi, Yacine Benzina, Moataz Chouchen, Ali Ouni 0001, Mohammed Sayagh, Salah Bouktif |
SANER | 5 |
| 2025 | Towards understanding code review practices for infrastructure-as-code: An empirical study on OpenStack projects
Narjes Bessghaier, Ali Ouni 0001, Mohammed Sayagh, Moataz Chouchen, Mohamed Wiem Mkaouer |
Empir. Softw. Eng. | 3 |
| 2024 | How Much Logs Does My Source Code File Need? Learning to Predict the Density of LogsabstractSoftware logging is the practice of recording different events that occur within a software system, which are useful for several analysis activities. However, striking the right balance between logging and system overhead is challenging. Prior work has conducted various machine learning-based solutions to suggest where to insert logging statements. But most importantly, before answering the question “where to log?’’, practitioners first need to determine whether a file needs logging at the first place. To do so, we conduct in this paper an empirical study to characterize the log density (i.e., ratio of log lines over the total lines of code) in seven open-source software projects. Then, we propose a deep learning based approach to predict the log density based on syntactic and semantic features of the source code. We find that the percentage of files with at least one log line ranges from 5% to 33% across the studied projects. Additionally, the median log density in the files with at least one log line ranges from 0.95% to 1.85% across the seven projects and can go up to 18%. Our findings resonate with the hypothesis that not all source code files require logging. On the other hand, our log density models achieve an average accuracy of 84%. Whereas our cross-project log density prediction results show a promising performance with an average accuracy of 72%, which represents over 86% (ratio of cross/within) of the corresponding within-project predictions using syntactic features. Our results show that we can accurately predict whether a file needs logging and such predictions may be generalized across projects. Mohamed Amine Batoun, Mohammed Sayagh, Ali Ouni 0001 |
EASE | 2 |
| 2024 | On the Impact of Draft Pull Requests on Accelerating FeedbackabstractThe pull request (PR) mechanism provides a structured way for developers to present their proposed modifications, engage in code review, and address any concerns before the changes are incorporated into the main codebase. As the adoption of PRs has grown over time, a variant known as draft PRs has gained traction, offering developers a novel way to share ideas and get feedback and collaboration on code changes during their development. Despite the benefits that draft PRs offer, there is still a lack of research on their efficiency in reaching the goal for which they are conceived, which is getting early feedback. This research paper aims to fill this gap by exploring how the draft mechanism is used, the impact of the draft mechanism on getting feedback and on the integration of new changes, and the different factors contributing to the responsiveness of comments. We observe that the draft is used by practitioners for complex changes and for different purposes, including the discussion of features to implement, experiments to conduct, and new versions to release. However, the goal of the draft mechanism is not fully reached as practitioners do not receive feedback (i.e., issue or review comments) on their drafts. For that, we leverage explanatory machine learning models to understand the differences between drafts that receive comments from these that did not before becoming ready to review. Our models show a median AUC performance of 0.67 to 0.77 and 0.66 to 0.75 for the issue comments and review comments models, respectively. The interpretation of our models shows that the actions that developers take as events, the description of the draft, the engagement of the author of the draft, and the engagement and responsiveness of the main reviewers play a crucial role in encouraging the reception of comments for draft pull requests. Firas Harbaoui, Mohammed Sayagh, Rabe Abdalkareem |
ICSME | 2 |
| 2024 | On the Prevalence, Co-occurrence, and Impact of Infrastructure-as-Code SmellsabstractIn modern software systems, Infrastructure-as-Code (IaC) tools play a pivotal role in automating the management of various infrastructure resources such as networks, databases, and services. This automation is done through code-based specification files, commonly known as IaC files. Similarly to other code files, IaC files can suffer from violations of established implementation and design standards, i.e., IaC smells. Although prior research has studied various aspects of traditional smells in non-IaC artifacts, there is little knowledge of how IaC smells are prevalent, co-occurring, and impacting the change and defect proneness of IaC code. To fill this gap, we conduct an empirical study encompassing 82 Puppet-based open-source projects. Our investigation focused on 12 types of IaC smells in both implementation and design levels. Our findings reveal that IaC smells do not manifest uniformly, as IaC smells that are particularly associated with modularity issues, exhibit high prevalence rates across projects. Additionally, we found that 74% of IaC files are smelly and over 52% of the smelly IaC files have at least two co-occurring IaC smells. Furthermore, our findings highlight that, on average, smelly IaC files are modified nearly 3.8 times, in terms of number of commits, more frequently than non-smelly IaC files. Furthermore, smelly IaC files are found to be 3.1 times more prone to larger code changes, in terms of code churn, than non-smelly IaC files. Additionally, we found that smelly IaC files are 3.3 times more prone to the introduction of defects that are likely to persist in 1.65 more commits before being fixed than non-smelly IaC files. These findings advocate developers to be more aware of IaC smells in their projects and consider their correction. Narjes Bessghaier, Mahi Begoug, Chemseddine Mebarki, Ali Ouni 0001, Mohammed Sayagh, Mohamed Wiem Mkaouer |
SANER | 5 |
| 2024 | An empirical study on cross-component dependent changes: A case study on the components of OpenStack
Ali Arabat, Mohammed Sayagh |
Empir. Softw. Eng. | 2 |
| 2024 | A literature review and existing challenges on software logging practices
Mohamed Amine Batoun, Mohammed Sayagh, Roozbeh Aghili, Ali Ouni 0001, Heng Li 0007 |
Empir. Softw. Eng. | 2 |
| 2024 | The impact of concept drift and data leakage on log level prediction models
Youssef Esseddiq Ouatiti, Mohammed Sayagh, Noureddine Kerzazi, Bram Adams, Ahmed E. Hassan |
Empir. Softw. Eng. | 2 |
| 2024 | What Constitutes the Deployment and Runtime Configuration System? An Empirical Study on OpenStack ProjectsabstractModern software systems are designed to be deployed in different configured environments (e.g., permissions, virtual resources, network connections) and adapted at runtime to different situations (e.g., memory limits, enabling/disabling features, database credentials). Such a configuration during the deployment and runtime of a software system is implemented via a set of configuration files, which together constitute what we refer to as a “configuration system.” Recent research efforts investigated the evolution and maintenance of configuration files. However, they merely focused on a limited part of the configuration system (e.g., specific infrastructure configuration files or Dockerfiles), and their results do not generalize to the whole configuration system. To cope with such a limitation, we aim to better capture and understand what files constitute a configuration system. To do so, we leverage an open card sort technique to qualitatively study 1,756 configuration files from OpenStack, a large and widely studied open source software ecosystem. Our investigation reveals the existence of nine types of configuration files, which cover the creation of the infrastructure on top of which OpenStack will be deployed, along with other types of configuration files used to customize OpenStack after its deployment. These configuration files are interconnected while being used at different deployment stages. For instance, we observe specific configuration files used during the deployment stage to create other configuration files that are used in the runtime stage. We also observe that identifying and classifying these types of files is not straightforward, as five out of the nine types can be written in similar programming languages (e.g., Python and Bash) as regular source code files. We also found that the same file extensions (e.g., Yaml ) can be used for different configuration types, making it difficult to identify and classify configuration files. Thus, we first leverage a machine learning model to identify configuration from non-configuration files, which achieved a median area under the curve (AUC) of 0.91, a median Brier score of 0.12, a median precision of 0.86, and a median recall of 0.83. Thereafter, we leverage a multi-class classification model to classify configuration files based on the nine configuration types. Our multi-class classification model achieved a median weighted AUC of 0.92, a median Brier score of 0.04, a median weighted precision of 0.84, and a median weighted recall of 0.82. Our analysis also shows that with only 100 labeled configuration and non-configuration files, our model reached a median AUC higher than 0.69. Furthermore, our configuration model requires a minimum of 100 configuration files to reach a median weighted AUC higher than 0.75. Narjes Bessghaier, Mohammed Sayagh, Ali Ouni 0001, Mohamed Wiem Mkaouer |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2023 | IoPV: On Inconsistent Option Performance VariationsabstractMaintaining a good performance of a software system is a primordial task when evolving a software system. The performance regression issues are among the dominant problems that large software systems face. In addition, these large systems tend to be highly configurable, which allows users to change the behaviour of these systems by simply altering the values of certain configuration options. However, such flexibility comes with a cost. Such software systems suffer throughout their evolution from what we refer to as “Inconsistent Option Performance Variation” (IoPV ). An IoPV indicates, for a given commit, that the performance regression or improvement of different values of the same configuration option is inconsistent compared to the prior commit. For instance, a new change might not suffer from any performance regression under the default configuration (i.e., when all the options are set to their default values), while altering one option’s value manifests a regression, which we refer to as a hidden regression as it is not manifested under the default configuration. Similarly, when developers improve the performance of their systems, performance regression might be manifested under a subset of the existing configurations. Unfortunately, such hidden regressions are harmful as they can go unseen to the production environment. In this paper, we first quantify how prevalent (in)consistent performance regression or improvement is among the values of an option. In particular, we study over 803 Hadoop and 502 Cassandra commits, for which we execute a total of 4,902 and 4,197 tests, respectively, amounting to 12,536 machine hours of testing. We observe that IoPV is a common problem that is difficult to manually predict. 69% and 93% of the Hadoop and Cassandra commits have at least one configuration that hides a performance regression. Worse, most of the commits have different options or tests leading to IoPV and hiding performance regressions. Therefore, we propose a prediction model that identifies whether a given combination of commit, test, and option (CTO) manifests an IoPV. Our evaluation for different models shows that random forest is the best performing classifier, with a median AUC of 0.91 and 0.82 for Hadoop and Cassandra, respectively. Our paper defines and provides scientific evidence about the IoPV problem and its prevalence, which can be explored by future work. In addition, we provide an initial machine learning model for predicting IoPV. Jinfu Chen 0002, Zishuo Ding, Yiming Tang 0002, Mohammed Sayagh, Heng Li 0007, Bram Adams, Weiyi Shang |
ESEC/SIGSOFT FSE | 4 |
| 2023 | An Empirical Study on GitHub Pull Requests' ReactionsabstractThe pull request mechanism is commonly used to propose source code modifications and get feedback from the community before merging them into a software repository. On GitHub, practitioners can provide feedback on a pull request by either commenting on the pull request or simply reacting to it using a set of pre-defined GitHub reactions, i.e., “Thumbs-up”, “Laugh”, “Hooray”, “Heart”, “Rocket”, “Thumbs-down”, “Confused”, and “Eyes”. While a large number of prior studies investigated how to improve different software engineering activities (e.g., code review and integration) by investigating the feedback on pull requests, they focused only on pull requests’ comments as a source of feedback. However, the GitHub reactions, according to our preliminary study, contain feedback that is not manifested within the comments of pull requests. In fact, our preliminary analysis of six popular projects shows that a median of 100% of the practitioners who reacted to a pull request did not leave any comment suggesting that reactions can be a unique source of feedback to further improve the code review and integration process. To help future studies better leverage reactions as a feedback mechanism, we conduct an empirical study to understand the usage of GitHub reactions and understand their promises and limitations. We investigate in this article how reactions are used, when and who use them on what types of pull requests, and for what purposes. Our study considers a quantitative analysis on a set of 380 k reactions on 63 k pull requests of six popular open-source projects on GitHub and three qualitative analyses on a total number of 989 reactions from the same six projects. We find that the most common used GitHub reactions are the positive ones (i.e., “Thumbs-up”, “Hooray”, “Heart”, “Rocket”, and “Laugh”). We observe that reactors use positive reactions to express positive attitude (e.g., approval, appreciation, and excitement) on the proposed changes in pull requests. A median of just 1.95% of the used reactions are negative ones, which are used by reactors who disagree with the proposed changes for six reasons, such as feature modifications that might have more downsides than upsides or the use of the wrong approach to address certain problems. Most (a median of 78.40%) reactions on a pull request come before the closing of the corresponding pull requests. Interestingly, we observe that non-contributors (i.e., outsiders who potentially are the “end-users” of the software) are also active on reacting to pull requests. On top of that, we observe that core contributors, peripheral contributors, casual contributors and outsiders have different behaviors when reacting to pull requests. For instance, most core contributors react in the early stages of a pull request, while peripheral contributors, casual contributors and outsiders react around the closing time or, in some cases, after a pull request is merged. Contributors tend to react to the pull request’s source code, while outsiders are more concerned about the impact of the pull request on the end-user experience. Our findings shed light on common patterns of GitHub reactions usage on pull requests and provide taxonomies about the intention of reactors, which can inspire future studies better leverage pull requests’ reactions. Mohamed Amine Batoun, Ka Lai Yung, Yuan Tian 0008, Mohammed Sayagh |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2023 | The Co-evolution of the WordPress Platform and Its PluginsabstractOne can extend the features of a software system by installing a set of additional components called plugins. WordPress, as a typical example of such plugin-based software ecosystems, is used by millions of websites and has a large number (i.e., 54,777) of available plugins. These plugin-based software ecosystems are different from traditional ecosystems (e.g., NPM dependencies) in the sense that there is high coupling between a platform and its plugins compared to traditional ecosystems for which components might not necessarily depend on each other (e.g., NPM libraries do not depend on a specific version of NPM or a specific version of a client software system). The high coupling between a plugin and its platform and other plugins causes incompatibility issues that occur during the co-evolution of a plugin and its platform as well as other plugins. In fact, incompatibility issues represent a major challenge when upgrading WordPress or its plugins. According to our study of the top 500 most-released WordPress plugins, we observe that incompatibility issues represent the third major cause for bad releases, which are rapidly (within the next 24 hours) fixed via urgent releases. Thirty-two percent of these incompatibilities are between a plugin and WordPress while 19% are between peer plugins. In this article, we study how plugins co-evolve with the underlying platform as well as other plugins, in an effort to understand the practices that are related support such co-evolution and reduce incompatibility issues. In particular, we investigate how plugins support the latest available versions of WordPress, as well as how plugins are related to each other, and how they co-evolve. We observe that a plugin’s support of new versions of WordPress with a large amount of code change is risky, as the releases that declare such support have a higher chance to be followed by an urgent release compared to ordinary releases. Although plugins support the latest WordPress version, plugin developers omit important changes such as deleting the use of removed WordPress APIs, which are removed a median of 873 days after the APIs have been removed from the source code of WordPress. Plugins introduce new releases that are made according to a median of five other plugins, which we refer to as peer-triggered releases. A median of 20% of the peer-triggered releases are urgent releases that fix problems in their previous releases. The most common goal of peer-triggered releases is the fixing of incompatibility issues that a plugin detects as late as after a median of 36 days since the last release of another plugin. Our work sheds light on the co-evolution of WordPress plugins with their platform as well as peer plugins in an effort to uncover the practices of plugin evolution, so WordPress can accordingly design approaches to avoid incompatibility issues. Jiahuei Lin, Mohammed Sayagh, Ahmed E. Hassan |
ACM Trans. Softw. Eng. Methodol. | 2 |
| 2023 | An Empirical Study on Log Level Prediction for Multi-Component SystemsabstractLogging statements are used to trace the execution of a software system. Practitioners leverage different logging information (e.g., the content of a log message) to decide for each logging statement an appropriate log level, which is leveraged to adjust the verbosity of logs so that only important log messages are traced. Deciding for the log level can be done differently from one to another component of a multi-component system, such as OpenStack and its 28 components. For example, a component might aim for increasing the verbosity of its log messages, while another component for the same multi-component system might aim at decreasing such a verbosity. Such different logging strategies can exist since each component can be developed and maintained by a different team. While a prior work leveraged an ordinal regression model to recommend the appropriate log level for a new logging statement, their evaluation did not consider the particularities that each component can have within a multi-component system. For instance, their model might not perform well at each component level of a multi-component system. The same model’s interpretability can mislead the developers of each component that has its unique logging strategy. In this paper, we quantify the impact of the particularities of each component of a multi-component system on the performance and interpretability of the log level prediction model of prior work. We observe that the performance of the log level prediction models that are trained at the whole project level (aka., global models) have lower performances (AUC) on 72% to 100% of the components of our five evaluated multi-component systems, compared to the same models when evaluated on the whole multi-component system. We observe that the models that are trained at the component level (aka., local models) statistically outperform the global model on 33% to 77% of the components of our evaluated multi-component systems. Furthermore, we observe that the rankings of the most important features that are obtained from the global models are statistically different from the feature importance rankings of 50% to 87% of the local models of our evaluated multi-component systems. Finally, we observe that 60% and 35% of the Spring and OpenStack components do not have enough data points to train their own local models (aka., data lacking components). Leveraging a peer-local model for such type of components is more promising than using the global model. Youssef Esseddiq Ouatiti, Mohammed Sayagh, Noureddine Kerzazi, Ahmed E. Hassan |
IEEE Trans. Software Eng. | 2 |
| 2022 | Improving State-of-the-Art Compression Techniques for Log Management ToolsabstractLog data records important runtime information about the running of a software system for different purposes including performance assurance, capacity planning, and anomaly detection. Log management tools such as ELK Stack and Splunk are widely adopted to manage and leverage log data in order to assist DevOps in real-time log analytics and decision making. To enable fast queries and to save storage space, such tools split log data into small blocks (e.g., 16KB), then index and compress each block separately. Previous log compression studies focus on improving the compression of either large-sized log files or log streams, without considering improving the compression of small log blocks (the actual compression need by modern log management tools). The evaluation of four state-of-the-art compression approaches (e.g.,Logzip, a variation ofLogzipby pre-extracting log templates namedLogzip-E,LogArchiveandCowic) indicates that these approaches do not perform well on small log blocks. In fact, the compressed blocks that are preprocessed usingLogzip,Logzip-E,LogArchiveorCowicare even larger (on median 1.3 times, 1.5 times, 0.2 times or 6.6 times) than the compressed blocks without any preprocessing. Hence, we propose an approach namedLogBlockto preprocess small log blocks before compressing them with a general compressor such asgzip,deflateandlz4, which are widely adopted by log management tools.LogBlockreduces the repetitiveness of logs by preprocessing the log headers and rearranging the log content leading to an improved compression ratio for a log file. Our evaluation on 16 log files shows that, for 16KB to 128KB block sizes, the compressed blocks byLogBlockare on median 5 to 21 percent smaller than the same compressed blocks without preprocessing (outperforming the state-of-the-art compression approaches).LogBlockachieves both a higher compression ratio (a median of 1.7 to 8.4 times, 1.9 to 10.0 times, 1.3 to 1.9 times and 6.2 to 11.4 times) and a faster compression speed (a median of 30.8 to 49.7 times, 42.6 to 53.8 times, 4.5 to 6.0 times and 2.5 to 4.0 times) thanLogzip,Logzip-E,LogArchiveandCowic.LogBlockcan help improve the storage efficiency of log management tools. Kundi Yao, Mohammed Sayagh, Weiyi Shang, Ahmed E. Hassan |
IEEE Trans. Software Eng. | 2 |
| 2021 | A study of how Docker Compose is used to compose multi-component systems
Md Hasan Ibrahim, Mohammed Sayagh, Ahmed E. Hassan |
Empir. Softw. Eng. | 2 |
| 2021 | An Empirical Study of the Impact of Data Splitting Decisions on the Performance of AIOps SolutionsabstractAIOps (Artificial Intelligence for IT Operations) leverages machine learning models to help practitioners handle the massive data produced during the operations of large-scale systems. However, due to the nature of the operation data, AIOps modeling faces several data splitting-related challenges, such as imbalanced data, data leakage, and concept drift. In this work, we study the data leakage and concept drift challenges in the context of AIOps and evaluate the impact of different modeling decisions on such challenges. Specifically, we perform a case study on two commonly studied AIOps applications: (1) predicting job failures based on trace data from a large-scale cluster environment and (2) predicting disk failures based on disk monitoring data from a large-scale cloud storage environment. First, we observe that the data leakage issue exists in AIOps solutions. Using a time-based splitting of training and validation datasets can significantly reduce such data leakage, making it more appropriate than using a random splitting in the AIOps context. Second, we show that AIOps solutions suffer from concept drift. Periodically updating AIOps models can help mitigate the impact of such concept drift, while the performance benefit and the modeling cost of increasing the update frequency depend largely on the application data and the used models. Our findings encourage future studies and practices on developing AIOps solutions to pay attention to their data-splitting decisions to handle the data leakage and concept drift challenges. Yingzhe Lyu, Heng Li 0007, Mohammed Sayagh, Zhen Ming (Jack) Jiang, Ahmed E. Hassan |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2021 | A Qualitative Study of the Benefits and Costs of Logging From Developers' PerspectivesabstractSoftware developers insert logging statements in their source code to collect important runtime information of software systems. In practice, logging appropriately is a challenge for developers. Prior studies aimed to improve logging by proactively inserting logging statements in certain code snippets or by learningwhere to logfrom existing logging code. However, there exists no work that systematically studies developers’ logging considerations, i.e., the benefits and costs of logging from developers’ perspectives. Without understanding developers’ logging considerations, automated approaches for logging decisions are based primarily on researchers’ intuition which may not be convincing to developers. In order to fill the gap between developers’ logging considerations and researchers’ intuition, we performed a qualitative study that combines a survey of 66 developers and a case study of 223 logging-related issue reports. The findings of our qualitative study draw a comprehensive picture of the benefits and costs of logging from developers’ perspectives. We observe that developers consider a wide range of logging benefits and costs, while most of the uncovered benefits and costs have never been observed nor discussed in prior work. We also observe that developers usead hocstrategies to balance the benefits and costs of logging. Developers need to be fully aware of the benefits and costs of logging, in order to better benefit from logging (e.g., leveraging logging to enable users to solve problems by themselves) and avoid unnecessary negative impact (e.g., exposing users’ sensitive information). Future research needs to consider such a wide range of logging benefits and costs when developing automated logging strategies. Our findings also inspire opportunities for researchers and logging library providers to help developers balance the benefits and costs of logging, for example, to support different log levels for different parts of a logging statement, or to help developers estimate and reduce the negative impact of logging statements. Heng Li 0007, Weiyi Shang, Bram Adams, Mohammed Sayagh, Ahmed E. Hassan |
IEEE Trans. Software Eng. | 4 |
| 2021 | ConfigMiner: Identifying the Appropriate Configuration Options for Config-Related User Questions by Mining Online ForumsabstractWhile the behavior of a software system can be easily changed by modifying the values of a couple of configuration options, finding one out of hundreds or thousands of available options is, unfortunately, a challenging task. Therefore, users often spend a considerable amount of time asking and searching around for the appropriate configuration options in online forums such as StackOverflow. In this paper, we propose ConfigMiner, an approach to automatically identify the appropriate option(s) to config-related user questions by mining already-answered config-related questions in online forums. Our evaluation on 2,062 config-related user questions for seven software systems shows that ConfigMiner can identify the appropriate option(s) for a median of 83 percent (up to 91 percent) of user questions within the top-20 recommended options, improving over state-of-the-art approaches by a median of 130 percent. Besides, ConfigMiner reports the relevant options at a median rank of 4, compared to a median of 16-20.5 as reported by the state-of-the-art approaches. Mohammed Sayagh, Ahmed E. Hassan |
IEEE Trans. Software Eng. | 1 |
| 2020 | Too many images on DockerHub! How different are images for the same system?
Md Hasan Ibrahim, Mohammed Sayagh, Ahmed E. Hassan |
Empir. Softw. Eng. | 2 |
| 2020 | An empirical study of the characteristics of popular Minecraft mods
Gopi Krishnan Rajbahadur, Dayi Lin, Mohammed Sayagh, Cor-Paul Bezemer, Ahmed E. Hassan |
Empir. Softw. Eng. | 4 |
| 2020 | What should your run-time configuration framework do to help developers?
Mohammed Sayagh, Noureddine Kerzazi, Fábio Petrillo, Khalil Bennani, Bram Adams |
Empir. Softw. Eng. | 1 |
| 2020 | Software Configuration Engineering in Practice Interviews, Survey, and Systematic Literature ReviewabstractModern software applications are adapted to different situations (e.g., memory limits, enabling/disabling features, database credentials) by changing the values of configuration options, without any source code modifications. According to several studies, this flexibility is expensive as configuration failures represent one of the most common types of software failures. They are also hard to debug and resolve as they require a lot of effort to detect which options are misconfigured among a large number of configuration options and values, while comprehension of the code also is hampered by sprinkling conditional checks of the values of configuration options. Although researchers have proposed various approaches to help debug or prevent configuration failures, especially from the end users' perspective, this paper takes a step back to understand the process required by practitioners to engineer the run-time configuration options in their source code, the challenges they experience as well as best practices that they have or could adopt. By interviewing 14 software engineering experts, followed by a large survey on 229 Java software engineers, we identified 9 major activities related to configuration engineering, 22 challenges faced by developers, and 24 expert recommendations to improve software configuration quality. We complemented this study by a systematic literature review to enrich the experts' recommendations, and to identify possible solutions discussed and evaluated by the research community for the developers' problems and challenges. We find that developers face a variety of challenges for all nine configuration engineering activities, starting from the creation of options, which generally is not planned beforehand and increases the complexity of a software system, to the non-trivial comprehension and debugging of configurations, and ending with the risky maintenance of configuration options, since developers avoid touching and changing configuration options in a mature system. We also find that researchers thus far focus primarily on testing and debugging configuration failures, leaving a large range of opportunities for future work. Mohammed Sayagh, Noureddine Kerzazi, Bram Adams, Fábio Petrillo |
IEEE Trans. Software Eng. | 1 |
| 2017 | On cross-stack configuration errorsabstractToday's web applications are deployed on powerful software stacks such as MEAN (JavaScript) or LAMP (PHP), which consist of multiple layers such as an operating system, web server, database, execution engine and application framework, each of which provide resources to the layer just above it. These powerful software stacks unfortunately are plagued by so-called cross-stack configuration errors (CsCEs), where a higher layer in the stack suddenly starts to behave incorrectly or even crash due to incorrect configuration choices in lower layers. Due to differences in programming languages and lack of explicit links between configuration options of different layers, sysadmins and developers have a hard time identifying the cause of a CsCE, which is why this paper (1) performs a qualitative analysis of 1,082 configuration errors to understand the impact, effort and complexity of dealing with CsCEs, then (2) proposes a modular approach that plugs existing source code analysis (slicing) techniques, in order to recommend the culprit configuration option. Empirical evaluation of this approach on 36 real CsCEs of the top 3 LAMP stack layers shows that our approach reports the misconfigured option with an average rank of 2.18 for 32 of the CsCEs, and takes only few minutes, making it practically useful. Mohammed Sayagh, Noureddine Kerzazi, Bram Adams |
ICSE | 1 |
| 2017 | Does the Choice of Configuration Framework Matter for Developers? Empirical Study on 11 Java Configuration FrameworksabstractConfiguration frameworks are routinely used in software systems to change application behavior without recompilation. Selecting a suitable configuration framework among the vast variety of existing choices is a crucial decision for developers, as it can impact project reliability and its maintenance profile. In this paper, we analyze almost 2,000 Java projects on GitHub to investigate the features and properties of 11 major Java configuration frameworks. We analyze the popularity of the frameworks and try to identify links between the maintenance effort involved with the usage of these frameworks and the frameworks' properties. More basic frameworks turn out to be the most popular, but in half of the cases are complemented by more complex frameworks. Furthermore, younger, more active frameworks with more detailed documentation, support for hierarchical configuration models and/or more data formats seem to require more maintenance by client developers. Mohammed Sayagh, Zhen Dong 0004, Artur Andrzejak 0001, Bram Adams |
SCAM | 1 |
| 2016 | Energy profiles of Java collections classesabstractWe created detailed profiles of the energy consumed by common operations done on Java List, Map, and Set abstractions. The results show that the alternative data types for these abstractions differ significantly in terms of energy consumption depending on the operations. For example, an ArrayList consumes less energy than a LinkedList if items are inserted at the middle or at the end, but consumes more energy than a LinkedList if items are inserted at the start of the list. To explain the results, we explored the memory usage and the bytecode executed during an operation. Expensive computation tasks in the analyzed bytecode traces appeared to have an energy impact, but memory usage did not contribute. We evaluated our profiles by using them to selectively replace Collections types used in six applications and libraries. We found that choosing the wrong Collections type, as indicated by our profiles, can cost even 300% more energy than the most efficient choice. Our work shows that the usage context of a data structure and our measured energy profiles can be used to decide between alternative Collections implementations. Samir Hasan, Zachary King, Munawar Hafiz, Mohammed Sayagh, Bram Adams, Abram Hindle |
ICSE | 4 |
| 2015 | Multi-layer software configuration: Empirical study on wordpressabstractSoftware can be adapted to different situations and platforms by changing its configuration. However, incorrect configurations can lead to configuration errors that are hard to resolve or understand, especially in the case of multi-layer architectures, where configuration options in each layer might contradict each other or be hard to trace to each other. Hence, this paper performs an empirical study on the occurrence of multi-layer configuration options across Wordpress (WP) plugins, WP, and the PHP engine. Our analyses show that WP and its plugins use on average 76 configuration options, a number that increases across time. We also find that each plugin uses on average 1.49% to 9.49% of all WP database options, and 1.38% to 15.18% of all WP configurable constants. 85.16% of all WP database options, 78.88% of all WP configurable constants, and 52 PHP configuration options are used by at least two plugins at the same time. Finally, we show how the latter options have a larger potential for questions and confusion amongst users. Mohammed Sayagh, Bram Adams |
SCAM | 1 |