Alexandre Decan

dblp:84/7943 · DBLP profile ↗
← Back
31ranked-venue papers
11as first author
21since 2021 · last 2026
0000-0002-5824-5823ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 30 · 11 first-author · 21 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2026 An Empirical Analysis of Code Clones in GitHub Actions Workflows
abstract
peer reviewed
Guillaume Cardoen, Tom Mens, Alexandre Decan
SANER3
2026 An empirical study of the evolution of GitHub actions workflows
Pooya Rostami Mazrae, Alexandre Decan, Tom Mens, Mairieli Santos Wessel
J. Syst. Softw.2
2025 A Dataset of Contributor Activities in the NumFocus Open-Source Community
abstract
Large open-source software (OSS) communities are composed of multiple interrelated projects, hosting numerous repositories involving thousands of interacting contributors. Socio-technical studies about a community’s collaboration dynamics can benefit from historical data logs of the detailed activities performed by the projects’ contributors.This paper provides an automated mapping of raw public events in GitHub repositories to structured activities that more accurately capture the intent of contributors. It also contributes a large dataset containing three years of activities of the 180K+ contributors of NUMFOCUS, a large OSS community supporting scientific research and data science. The dataset covers 58 projects, including 2.2M+ activities across 2,851 GitHub repositories. This dataset allows advanced studies of the NUMFOCUS community collaboration dynamics, and the activity mapping process enables the possibility to create and use similar datasets for other OSS communities.
Youness Hourri, Alexandre Decan, Tom Mens
MSR2
2025 A bot identification model and tool based on GitHub activity sequences
Natarajan Chidambaram, Alexandre Decan, Tom Mens
J. Syst. Softw.2
2024 A dataset of GitHub Actions workflow histories
abstract
GitHub Actions is the de facto workflow automation tool for GitHub repositories. Its popularity has increased dramatically over the recent years, opening up opportunities for empirical studies related to its usage. To enable such studies, we implemented gigawork, an open source tool for extracting the commit histories of changes to work-flow files in GitHub repositories. Using this tool we collected and publicly released a dataset of 160K+ commit histories of workflow files in 32K+ public GitHub repositories, covering 1.5M+ workflow file versions. In order to facilitate its use by other researchers, the dataset includes relevant metadata related to workflow file changes in each commit. gigawork is publicly released on PyPi. Its associated dataset can be found on Zenodo (DOI: 10.5281/zenodo.10259013).
Guillaume Cardoen, Tom Mens, Alexandre Decan
MSR3
2024 RABBIT: A tool for identifying bot accounts based on their recent GitHub event history
abstract
Collaborative software development through GitHub repositories frequently relies on bot accounts to automate repetitive and error-prone tasks. This highlights the need to have accurate and efficient bot identification tools. Several such tools have been proposed in the past, but they tend to rely on a substantial amount of historical data, or they limit themselves to a reduced subset of activity types, making them difficult to use at large scale. To overcome these limitations, we developed RABBIT, an open source command-line tool that queries the GitHub Events API to retrieve the recent events of a given GitHub account and predicts whether the account is a human or a bot. RABBIT is based on an XGBoost classification model that relies on six features related to account activities and achieves high performance, with an AUC, F1 score, precision and recall of 0.92. Compared to the state-of-the-art in bot identification, RABBIT exhibits a similar performance in terms of precision, recall and F1 score, while being more than an order of magnitude faster and requiring considerably less data. This makes RABBIT usable on a large scale, capable of processing several thousand accounts per hour efficiently.
Natarajan Chidambaram, Tom Mens, Alexandre Decan
MSR3
2024 Quantifying Security Issues in Reusable JavaScript Actions in GitHub Workflows
abstract
GitHub's integrated automated workflow mechanism called GitHub Actions promotes the use of Actions as reusable building blocks in workflows. The majority of those Actions are developed in JavaScript and depend on packages distributed through the npm package manager. Those packages can suffer from security vulnerabilities, potentially affecting the Actions that rely on them. Using a dataset of 8,107 JavaScript Actions, we analysed to which extent dependencies on npm packages expose these Actions to vulnerabilities. We observed that JavaScript Actions tend to rely on dozens of npm packages, and that the vast majority of them depend on npm package releases with known vulnerabilities. Most of these vulnerabilities are caused by indirect dependencies, making it difficult for Actions maintainers to analyse their exposure to security vulnerabilities. Moreover, indirect dependencies are more likely to suffer from vulnerabilities of higher severity. We also studied to which extent security weaknesses occur in the source code of JavaScript Actions. To do so, we used CodeQL to detect security weaknesses, revealing that more than 54% of the studied JavaScript Actions contain at least one security weakness, and a small subset of these weaknesses recur frequently in their code. This justifies the need for further studies and more advanced tool support for addressing security issues in the GitHub Actions ecosystem.
Hassan Onsori Delicheh, Alexandre Decan, Tom Mens
MSR2
2024 gawd: A Differencing Tool for GitHub Actions Workflows
abstract
The GitHub social coding platform introduced GitHub Actions as a way to automate different aspects of collaborative software development through the use of workflow files. It is the most popular CI/CD and workflow automation tool for GitHub. To maintain workflow code over time, it is useful to rely on differencing tools to identify the changes made during successive commits. Unfortunately, existing code differencing tools are not able to correctly identify changes made to workflow files. We therefore implemented gawd, a syntactic differencing tool for GitHub Actions workflows. The tool is capable of reporting the addition, deletion, modification and move of syntactic components in workflow files, taking into account the specific syntax of workflows. gawd has been evaluated on manually classified sets of workflow changes taken from existing commits in 40 different GitHub repositories, and was able to successfully identify these changes. gawd is publicly released as an open source Python tool distributed on PyPI.
Pooya Rostami Mazrae, Alexandre Decan, Tom Mens
MSR2
2023 A Dataset of Bot and Human Activities in GitHub
abstract
Software repositories hosted on GitHub frequently use development bots to automate repetitive, effort intensive and error-prone tasks. To understand and study how these bots are used, state-of-the-art bot identification tools have been developed to detect bots based on their comments in commits, issues and pull requests. Given that bots can be involved in many other activity types, there is a need to consider more activities that they are carrying out in the software repositories they are involved in. We therefore propose a curated dataset of such activities carried out by bots and humans involved in GitHub repositories. The dataset was constructed by identifying 24 high-level activity types that could be extracted from 15 lower-level event types that were queried from GitHub’s event stream API for all considered bots and humans. The proposed dataset contains around 834K activities performed by 385 bots and 616 humans involved in GitHub repositories, during an observation period ranging from 25 November 2022 to 9 March 2023. By analysing the activity patterns of bots and humans, this dataset could lead to better bot identification tools and empirical studies on how bots play a role in collaborative software development.
Natarajan Chidambaram, Alexandre Decan, Tom Mens
MSR2
2023 On the usage, co-usage and migration of CI/CD tools: A qualitative analysis
Pooya Rostami Mazrae, Tom Mens, Mehdi Golzadeh, Alexandre Decan
Empir. Softw. Eng.4
2023 On the outdatedness of workflows in the GitHub Actions ecosystem
Alexandre Decan, Tom Mens, Hassan Onsori Delicheh
J. Syst. Softw.1
2022 On the Use of GitHub Actions in Software Development Repositories
abstract
GitHub Actions was introduced in 2019 and constitutes an integrated alternative to CI/CD services for GitHub repositories. The deep integration with GitHub allows repositories to easily automate software development workflows. This paper empirically studies the use of GitHub Actions on a dataset comprising 68K repositories on GitHub, of which 43.9% are using GitHub Actions workflows. We analyse which workflows are automated and identify the most frequent automation practices. We show that reuse of actions is a common practice, even if this reuse is concentrated in a limited number of actions. We study which actions are most frequently used and how workflows refer to them. Furthermore, we discuss the related security and versioning aspects. As such, we provide an overview of the use of GitHub Actions, constituting a necessary first step towards a better understanding of this emerging ecosystem and its implications on collaborative software development in the GitHub social coding platform.
Alexandre Decan, Tom Mens, Pooya Rostami Mazrae, Mehdi Golzadeh
ICSME1
2022 PaReco: patched clones and missed patches among the divergent variants of a software family
abstract
Re-using whole repositories as a starting point for new projects is often done by maintaining a variant fork parallel to the original. However, the common artifacts between both are not always kept up to date. As a result, patches are not optimally integrated across the two repositories, which may lead to sub-optimal maintenance between the variant and the original project. A bug existing in both repositories can be patched in one but not the other (we see this as a missed opportunity) or it can be manually patched in both probably by different developers (we see this as effort duplication). In this paper we present a tool (named PaReCo) which relies on clone detection to mine cases of missed opportunity and effort duplication from a pool of patches. We analyzed 364 (source to target) variant pairs with 8,323 patches resulting in a curated dataset containing 1,116 cases of effort duplication and 1,008 cases of missed opportunities. We achieve a precision of 91%, recall of 80%, accuracy of 88%, and F1-score of 85%. Furthermore, we investigated the time interval between patches and found out that, on average, missed patches in the target variants have been introduced in the source variants 52 weeks earlier. Consequently, PaReCo can be used to manage variability in “time” by automatically identifying interesting patches in later project releases to be backported to supported earlier releases.
Poedjadevie Ramkisoen, John Businge, Brent van Bladel, Alexandre Decan, Serge Demeyer, Coen De Roover, Foutse Khomh
ESEC/SIGSOFT FSE4
2022 Variant Forks - Motivations and Impediments
abstract
Social coding platforms centred around git provide explicit facilities to share code between projects: forks, pull requests, cherry-picking to name but a few. Variant forks are an interesting phenomenon in that respect, as they permit for different projects to peacefully co-exist, yet explicitly acknowledge the common ancestry. Several researchers analysed forking practices on open source platforms and observed that variant forks get created frequently. However, little is known on the motivations for launching such a variant fork. Is it mainly technical (e.g., diverging features), governance (e.g., diverging interests), legal (e.g., diverging licences), or do other factors come into play? We report the results of an exploratory qualitative analysis on the motivations behind creating and maintaining variant forks. We surveyed 105 maintainers of different active open source variant projects hosted on GitHub. Our study extends previous findings, identifying a number of fine-grained common motivations for launching a variant fork and listing concrete impediments for maintaining the co-existing projects.
John Businge, Ahmed Zerouali, Alexandre Decan, Tom Mens, Serge Demeyer, Coen De Roover
SANER3
2022 On the rise and fall of CI services in GitHub
abstract
Continuous integration (CI) services are used in collaborative open source projects to automate parts of the development workflow. Such services have been in widespread use for over a decade, with new CIs being introduced over the years, sometimes overtaking other CIs in popularity. We conducted a longitudinal empirical study over a period of nine years, aiming to better understand this rapidly evolving CI landscape. By analysing the development history of 91,810 GitHub repositories of active npm packages having used at least one CI service, we quantitatively studied the evolution of seven popular CIs, specifically focusing on their co-usage and migration in the considered repositories. We provide statistical evidence of the rise of GitHub Actions, that has become the dominant CI service in less than 18 months time. This coincides with the fall of Travis that has seen an important decrease in usage, likely due to a combination of policy changes and migrations to GitHub Actions.
Mehdi Golzadeh, Alexandre Decan, Tom Mens
SANER2
2022 On the impact of security vulnerabilities in the npm and RubyGems dependency networks
Ahmed Zerouali, Tom Mens, Alexandre Decan, Coen De Roover
Empir. Softw. Eng.3
2022 Back to the Past - Analysing Backporting Practices in Package Dependency Networks
abstract
The practice of backporting aims to bring the benefits of a bug or vulnerability fix from a higher to a lower release of a software package. When such a package adheres to semantic versioning, backports can be recognised as new releases in a lower major train. This is particularly useful in case a substantial number of software packages continues to depend on that lower major train. In this article, we study the backporting practices in four popular package distributions, namelyCargo,npm,PackagistandRubyGems. We observe that many dependent packages could benefit from backports provided by their dependencies. In particular, we find that a majority of security vulnerabilities affect more than one major train but are only fixed in the highest one, letting thousands of dependent packages exposed to the vulnerability. Despite that, we find that backporting updates is quite infrequent, and mostly practised by long-lived and more active packages for a variety of reasons.
Alexandre Decan, Tom Mens, Ahmed Zerouali, Coen De Roover
IEEE Trans. Software Eng.1
2021 A multi-dimensional analysis of technical lag in Debian-based Docker images
Ahmed Zerouali, Tom Mens, Alexandre Decan, Jesús M. González-Barahona, Gregorio Robles
Empir. Softw. Eng.3
2021 A ground-truth dataset and classification model for detecting bots in GitHub issue and PR comments
Mehdi Golzadeh, Alexandre Decan, Damien Legay, Tom Mens
J. Syst. Softw.2
2021 Lost in zero space - An empirical comparison of 0.y.z releases in software package distributions
Alexandre Decan, Tom Mens
Sci. Comput. Program.1
2021 What Do Package Dependencies Tell Us About Semantic Versioning?
abstract
The semantic versioning (semver) policy is commonly accepted by open source package management systems to inform whether new releases of software packages introduce possibly backward incompatible changes. Maintainers depending on such packages can use this information to avoid or reduce the risk of breaking changes in their own packages by specifying version constraints on their dependencies. Depending on the amount of control a package maintainer desires to have over her package dependencies, these constraints can range from very permissive to very restrictive. This article empirically comparessemvercompliance of four software packaging ecosystems (Cargo, npm, Packagist and Rubygems), and studies how this compliance evolves over time. We explore to what extent ecosystem-specific characteristics or policies influence the degree of compliance. We also propose an evaluation based on the “wisdom of the crowds” principle to help package maintainers decide which type of version constraints they should impose on their dependencies.
Alexandre Decan, Tom Mens
IEEE Trans. Software Eng.1
2020 On Package Freshness in Linux Distributions
abstract
The open-source Linux operating system is available through a wide variety of distributions, each containing a collection of installable software packages. It can be important to keep these packages as fresh as possible to benefit from new features, bug fixes and security patches. However, not all distributions place the same emphasis on package freshness. We conducted a survey in the first half of 2020 with 170 Linux users to gauge their perception of package freshness in the distributions they employ, the value they place on package freshness and the reasons why they do so, and the methods they use to update packages. The results of this survey reveal that, for the aforementioned reasons, keeping packages up to date is an important concern to Linux users and that they install and update packages through their distribution’s official repositories whenever possible, but often resort to third-party repositories and package managers for proprietary software and programming language libraries. Some distributions are perceived to be much quicker in deploying package updates than others. These results are useful to assess the expectations and requirements of Linux users in terms of package freshness and guide them in choosing a fitting distribution.
Damien Legay, Alexandre Decan, Tom Mens
ICSME2
2020 GAP: Forecasting commit activity in git projects
Alexandre Decan, Eleni Constantinou, Tom Mens, Henrique Rocha
J. Syst. Softw.1
2019 An empirical comparison of dependency network evolution in seven software packaging ecosystems
abstract
Nearly every popular programming language comes with one or more package managers. The software packages distributed by such package managers form large software ecosystems. These packaging ecosystems contain a large number of package releases that are updated regularly and that have many dependencies to other package releases. While packaging ecosystems are extremely useful for their respective communities of developers, they face challenges related to their scale, complexity, and rate of evolution. Typical problems are backward incompatible package updates, and the risk of (transitively) depending on packages that have become obsolete or inactive. This manuscript uses the libraries.io dataset to carry out a quantitative empirical analysis of the similarities and differences between the evolution of package dependency networks for seven packaging ecosystems of varying sizes and ages: Cargo for Rust , CPAN for Perl , CRAN for R , npm for JavaScript , NuGet for the .NET platform, Packagist for PHP , and RubyGems for Ruby . We propose novel metrics to capture the growth, changeability, reusability and fragility of these dependency networks, and use these metrics to analyze and compare their evolution. We observe that the dependency networks tend to grow over time, both in size and in number of package updates, while a minority of packages are responsible for most of the package updates. The majority of packages depend on other packages, but only a small proportion of packages accounts for most of the reverse dependencies. We observe a high proportion of “fragile” packages due to a high and increasing number of transitive dependencies. These findings are instrumental for assessing the quality of a package dependency network, and improving it through dependency management tools and imposed policies.
Alexandre Decan, Tom Mens, Philippe Grosjean
Empir. Softw. Eng.1
2019 A formal framework for measuring technical lag in component repositories - and its application to npm
abstract
Abstract Reusable Open Source Software (OSS) components for major programming languages are available in package repositories. Developers rely on package management tools to automate deployments, specifying which package releases satisfy the needs of their applications. However, these specifications may lead to deploying package releases that are outdated, or otherwise undesirable, because they do not include bug fixes, security fixes, or new functionality. In contrast, automatically updating to a more recent release may introduce incompatibility issues. To capture this delicate balance, we formalise a generic model of technical lag, a concept that quantifies to which extent a deployed collection of components is outdated, with respect to the ideal deployment. We operationalise this model for the npm package manager. We empirically analyze the history of package update practices and technical lag for more than 500K packages with about 4M package releases over a seven‐year period. We consider both development and runtime dependencies, and study both direct and transitive dependencies. We also analyze the technical lag of external GitHub applications depending on npm packages. We report our findings, suggesting the need for more awareness of, and integrated tool support for, controlling technical lag in software libraries.
Ahmed Zerouali, Tom Mens, Jesús M. González-Barahona, Alexandre Decan, Eleni Constantinou, Gregorio Robles
J. Softw. Evol. Process.4
2019 A method for testing and validating executable statechart models
Tom Mens, Alexandre Decan, Nikolaos I. Spanoudakis
Softw. Syst. Model.2
2018 On the Evolution of Technical Lag in the npm Package Dependency Network
abstract
Software packages developed and distributed through package managers extensively depend on other packages. These dependencies are regularly updated, for example to add new features, resolve bugs or fix security issues. In order to take full advantage of the benefits of this type of reuse, developers should keep their dependencies up to date by relying on the latest releases. In practice, however, this is not always possible, and packages lag behind with respect to the latest version of their dependencies. This phenomenon is described as technical lag in the literature. In this paper, we perform an empirical study of technical lag in the npm dependency network by investigating its evolution for over 1.4M releases of 120K packages and 8M dependencies between these releases. We explore how technical lag increases over time, taking into account the release type and the use of package dependency constraints. We also discuss how technical lag can be reduced by relying on the semantic versioning policy.
Alexandre Decan, Tom Mens, Eleni Constantinou
ICSME1
2018 On the impact of security vulnerabilities in the npm package dependency network
abstract
Security vulnerabilities are among the most pressing problems in open source software package libraries. It may take a long time to discover and fix vulnerabilities in packages. In addition, vulnerabilities may propagate to dependent packages, making them vulnerable too. This paper presents an empirical study of nearly 400 security reports over a 6-year period in the npm dependency network containing over 610k JavaScript packages. Taking into account the severity of vulnerabilities, we analyse how and when these vulnerabilities are discovered and fixed, and to which extent they affect other packages in the packaging ecosystem in presence of dependency constraints. We report our findings and provide guidelines for package maintainers and tool developers to improve the process of dealing with security issues.
Alexandre Decan, Tom Mens, Eleni Constantinou
MSR1
2017 An empirical comparison of dependency issues in OSS packaging ecosystems
abstract
Nearly every popular programming language comes with one or more open source software packaging ecosystem(s), containing a large collection of interdependent software packages developed in that programming language. Such packaging ecosystems are extremely useful for their respective software development community. We present an empirical analysis of how the dependency graphs of three large packaging ecosystems (npm, CRAN and RubyGems) evolve over time. We study how the existing package dependencies impact the resilience of the three ecosystems over time and to which extent these ecosystems suffer from issues related to package dependency updates. We analyse specific solutions that each ecosystem has put into place and argue that none of these solutions is perfect, motivating the need for better tools to deal with package dependency update problems.
Alexandre Decan, Tom Mens, Maëlick Claes
SANER1
2016 When GitHub Meets CRAN: An Analysis of Inter-Repository Package Dependency Problems
abstract
When developing software packages in a software ecosystem, an important and well-known challenge is how to deal with dependencies to other packages. In presence of multiple package repositories, dependency management tends to become even more problematic. For the R ecosystem of statistical computing, dependency management is currently insufficient to deal with multiple package versions and inter-repository package dependencies. We explore how the use of GitHub influences the R ecosystem, both for the distribution of R packages and for inter-repository package dependency management. We also discuss how these problems could be addressed.
Alexandre Decan, Tom Mens, Maëlick Claes, Philippe Grosjean
SANER1
2009 On First-Order Query Rewriting for Incomplete Database Histories
abstract
Multiwords are defined as words in which single symbols can be replaced by nonempty sets of symbols. Such a set of symbols captures uncertainty about the exact symbol. Words are obtained from multiwords by selecting a single symbol from every set. A pattern is certain in a multiword W if it occurs in every word that can be obtained from W. For a given pattern, we are interested in finding a logic formula that recognizes the multiwords in which that pattern is certain. This problem can be seen as a special case of consistent query answering (CQA). We show how our results can be applied in CQA on database histories under primary key constraints.
Véronique Bruyère, Alexandre Decan, Jef Wijsen
TIME2