VLDB 2026 Research / reviewers in the wild / expert
Adem Ait
dblp:322/8422 · also Adem Ait Fonollà
· DBLP profile ↗
8ranked-venue papers
7as first author
8since 2021 · last 2025
0000-0002-5334-9041ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 8 · 7 first-author · 8 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Towards Modeling Human-Agentic Collaborative Workflows: A BPMN Extension
Adem Ait, Javier Luis Cánovas Izquierdo, Jordi Cabot |
SEAA | 1 |
| 2025 | Towards Automated Governance: A DSL for Human-Agent Collaboration in Software ProjectsabstractThe stakeholders involved in software development are becoming increasingly diverse, with both human contributors from varied backgrounds and AI-powered agents collaborating together in the process. This situation presents unique governance challenges, particularly in Open-Source Software (OSS) projects, where explicit policies are often lacking or unclear. This paper presents the vision and foundational concepts for a novel Domain-Specific Language (DSL) designed to define and enforce rich governance policies in systems involving diverse stakeholders, including agents. This DSL offers a pathway towards more robust, adaptable, and ultimately automated governance, paving the way for more effective collaboration in software projects, especially OSS ones. Adem Ait, Gwendal Jouneaux, Javier Luis Cánovas Izquierdo, Jordi Cabot |
ASE | 1 |
| 2025 | On the suitability of hugging face hub for empirical studies
Adem Ait, Javier Luis Cánovas Izquierdo, Jordi Cabot |
Empir. Softw. Eng. | 1 |
| 2024 | On the Creation of Representative Samples of Software RepositoriesabstractSoftware repositories is one of the sources of data in Empirical Software Engineering, primarily in the Mining Software Repositories field, aimed at extracting knowledge from the dynamics and practice of software projects. With the emergence of social coding platforms such as GitHub, researchers have now access to millions of software repositories to use as source data for their studies. With this massive amount of data, sampling techniques are needed to create more manageable datasets. The creation of these datasets is a crucial step, and researchers have to carefully select the repositories to create representative samples according to a set of variables of interest. However, current sampling methods are often based on random selection or rely on variables which may not be related to the research study (e.g., popularity or activity). In this paper, we present a methodology for creating representative samples of software repositories, where such representativeness is properly aligned with both the characteristics of the population of repositories and the requirements of the empirical study. We illustrate our approach with use cases based on Hugging Face repositories. June Gorostidi, Adem Ait, Jordi Cabot, Javier Luis Cánovas Izquierdo |
ESEM | 2 |
| 2024 | HFCommunity: An extraction process and relational database to analyze Hugging Face Hub dataabstractSocial coding platforms such as GitHub or GitLab have become the de facto standard for developing Open-Source Software (OSS) projects. With the emergence of Machine Learning (ML), platforms specifically designed for hosting and developing ML-based projects have appeared, being Hugging Face Hub (HFH) one of the most popular ones. HFH aims at sharing datasets, pre-trained ML models and the applications built with them. With over 400 K repositories, and growing fast, HFH is becoming a promising source of empirical data on all aspects of ML project development. However, apart from the API provided by the platform, there are no easy-to-use solutions to collect the data, nor prepackaged datasets to explore the different facets of HFH. We present HFCommunity, an extraction process for HFH data and a relational database to facilitate an empirical analysis on the growing number of ML projects. Adem Ait, Javier Luis Cánovas Izquierdo, Jordi Cabot |
Sci. Comput. Program. | 1 |
| 2023 | A Tool for the Definition and Deployment of Platform-Independent Bots on Open Source ProjectsabstractThe development of Open Source Software (OSS) projects is a collaborative process that heavily relies on active contributions by passionate developers. Creating, retaining and nurturing an active community of developers is a challenging task; and finding the appropriate expertise to drive the development process is not always easy. To alleviate this situation, many OSS projects try to use bots to automate some development tasks, thus helping community developers to cope with the daily workload of their projects. However, the techniques and support for developing bots is specific to the code hosting platform where the project is being developed (e.g., GitHub or GitLab). Furthermore, there is no support for orchestrating bots deployed in different platforms nor for building bots that go beyond pure development activities. In this paper, we propose a tool to define and deploy bots for OSS projects, which besides automation tasks they offer a more social facet, improving community interactions. The tool includes a Domain-Specific Language (DSL) which allows defining bots that can be deployed on top of several platforms and that can be triggered by different events (e.g., creation of a new issue or a pull request). We describe the design and the implementation of the tool, and illustrate its use with examples. Adem Ait, Javier Luis Cánovas Izquierdo, Jordi Cabot |
SLE | 1 |
| 2023 | HFCommunity: A Tool to Analyze the Hugging Face Hub CommunityabstractIn recent years, empirical studies on software engineering practices have primarily relied on general-purpose social coding platforms such as GITHUB or GITLAB. With the emergence of Machine Learning (ML), platforms specifically designed for hosting and developing ML-based projects have appeared, being HUGGING FACE HUB one of the most popular ones. HUGGING FACE HUB focuses on facilitating the sharing of datasets, pre-trained ML models and applications built with them (spaces in HUGGING FACE HUB terminology). Besides, the Hub is adding more and more collaborative features, such as issues and pull requests, to facilitate the building of these artifacts within the platform itself. With over 100K repositories, and growing fast, HUGGING FACE HUB is therefore becoming a promising source of data on all aspects of ML projects and the community interactions around them. As such, we believe it is a promising source for all types of empirical studies aimed at analyzing the collaborative development and evolution of ML artifacts. Nevertheless, apart from the API provided by the platform, there are no easy-to-use solutions to collect and explore the different facets of HUGGING FACE HUB data, including the repositories, discussions and code evolution. To overcome this situation, in this paper we present HFCOMMUNITY, a relational database populated with HUGGING FACE HUB data to facilitate empirical analysis on the growing number of ML-related development projects. Adem Ait, Javier Luis Cánovas Izquierdo, Jordi Cabot |
SANER | 1 |
| 2022 | An Empirical Study on the Survival Rate of GitHub ProjectsabstractThe number of Open Source projects hosted in social coding platforms such as GitHub is constantly growing. However, many of these projects are not regularly maintained and some are even abandoned shortly after they were created. In this paper we analyze early project development dynamics in software projects hosted on GitHub, including their survival rate. To this aim, we collected all 1,127 GitHub repositories from four different ecosystems (i.e., NPM packages, R packages, WordPress plugins and Laravel packages) created in 2016. We stored their activity in a time series database and analyzed their activity evolution along their lifespan, from 2016 to now. Our results reveal that the prototypical development process consists of intensive coding-driven active periods followed by long periods of inactivity. More importantly, we have found that a significant number of projects die in the first year of existence with the survival rate decreasing year after year. In fact, the probability of surviving longer than five years is less than 50% though some types of projects have better chances of survival. Adem Ait, Javier Luis Cánovas Izquierdo, Jordi Cabot |
MSR | 1 |