VLDB 2026 Research / reviewers in the wild / expert
Souhaila Serbout
dblp:262/9466
· DBLP profile ↗
20ranked-venue papers
8as first author
18since 2021 · last 2026
0000-0002-8144-2606ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 19 · 8 first-author · 17 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PoolinGH: Fast, Efficient, and Robust GitHub Repository MiningabstractResearchers in Mining (open-source) Software Repositories (MSR) often create datasets that should survive the single paper and support long-term investigation of specific phenomena. Although popular, these studies recurrently deal with similar technical limitations. For instance, public collaborative development platforms, such as GitHub, impose hourly rate limits on their API requests. Furthermore, depending on network and API conditions, queries can fail and disrupt the process. These unexpected events can slow down or even invalidate the mining. Nevertheless, there are ways to minimize the undesirable effects in a reusable way while still complying with such limitations. However, best practices are often (re-)implemented on an ad hoc basis. Whatever works. Maxime André 0001, Marco Raglianti, Souhaila Serbout, Anthony Cleve, Michele Lanza 0001 |
MSR | 3 |
| 2026 | Consistent or Sensitive? Automated Code Revision Tools Against Semantics-Preserving PerturbationsabstractAutomated Code Revision (ACR) tools aim to reduce manual effort by automatically generating code revisions based on reviewer feedback. While ACR tools have shown promising performance on historical data, their real-world utility depends on their ability to handle similar code variants expressing the same issue—a property we define as consistency. However, the probabilistic nature of ACR tools often compromises consistency, which may lead to divergent revisions even for semantically equivalent code variants. Shirin Pirouzkhah, Souhaila Serbout, Alberto Bacchelli |
MSR | 2 |
| 2025 | Mining Security Documentation Practices in OpenAPI DescriptionsabstractSecurity is an integral requirement of any trustworthy software architecture, particularly critical for application programming interfaces (APIs). In this paper, we survey security documentation practices, specifically API security schemes related to authentication and authorization, by mining a large collection of OpenAPI descriptions retrieved from open-source GitHub repositories. Our study focuses on detecting existing security schemes and evaluating their prevalence and positioning within API descriptions. We distinguish whether security schemes are introduced locally (at the path or operation level) or globally (for the entire API). Our analysis highlights scenarios where security schemes are featured in APIs in different proportions over time, thus tracking whether the API documentation tends to include more (or less) security details as the API evolves. Diana Carolina Muñoz Hurtado, Souhaila Serbout, Cesare Pautasso |
ICSA | 2 |
| 2025 | PrioTestCI: Efficient Test Case Prioritization in GitHub Workflows for CI OptimizationabstractContinuous Integration (CI) is a widely adopted practice in software development to automatically verify code changes across diverse environments. However, executing the full test suite on every pull request update can lead to redundant runs, slower feedback loops, and inefficient utilization of CI resources. To address this issue, we introduce PrioTestCI, a prioritization technique within GitHub Actions that focuses on re-executing test cases that have previously failed. If these prioritized tests succeed, the remaining tests proceed; otherwise, the workflow terminates early, saving computation resources and providing early feedback to developers. PrioTestCI utilizes commit-to-commit test result tracking to inform future test runs, thereby reducing unnecessary repetition and accelerating validation cycles. We evaluated our technique on the Pytest project, a real-world open-source project with an extensive test matrix. PrioTestCI resulted in a CI runtime reduction of 1h57m39s compared to the normal workflow, with individual configuration improvements ranging from 63.75% to 91.94% (81.55% on average). Demo video: https://youtu.be/_3CF9LJdv0I?si=XyE_8mBnDxk1lMnD Repository: https://github.com/ShubhamDesai/CI-Optimization Shubham Vasudeo Desai, Shonil Bhide, Souhaila Serbout, Luciano Marchezan, Wesley K. G. Assunção |
ASE | 3 |
| 2025 | Quirx: A Mutation-Based Framework for Evaluating Prompt Robustness in LLM-based SoftwareabstractLarge Language Models (LLMs) increasingly power critical business processes, yet prompt robustness remains under-explored. Small variations—such as synonym changes or instruction reordering—can cause significant output shifts, undermining reliability in domains like customer service and finance. Existing evaluations rely on ad-hoc manual testing, limiting scalability in production environments.We present Quirx, a mutation-based fuzzing framework for systematically evaluating prompt robustness across LLM providers. Quirx applies tri-dimensional mutations (lexical, semantic, structural), executes them against target models, and measures response consistency via multi-level similarity analysis. It produces robustness scores, reveals failure patterns, and supports informed model selection.We evaluate Quirx on four models (GPT-3.5-turbo, GPT-4o-mini, Claude-3.5-Sonnet, Claude-Sonnet-4) across three tasks. Results show sentiment classification is uniformly robust (1.00), summarization is highly provider-sensitive (0.23–0.58) with Claude models 2.5× more robust than OpenAI, and SQL generation is consistently strong (0.80–1.00). Structural mutations cause 50–67% of summarization failures but have minimal effect on other tasks.Demo video: https://youtu.be/Sm3Gk2X2-vk Souhaila Serbout |
ASE | 1 |
| 2025 | EvoScat: Exploring Software Change Dynamics in Large-Scale Historical DatasetsabstractLong lived software projects encompass a large number of artifacts, which undergo many revisions throughout their history. Empirical software engineering researchers studying software evolution gather and collect datasets with millions of events, representing changes introduced to specific artifacts. In this paper, we propose EvoScat, a tool that attempts addressing temporal scalability through the usage of interactive density scatterplot to provide a global overview of large historical datasets mined from open source repositories in a single visualization. EvoScat intents to provide researchers with a mean to produce scalable visualizations that can help them explore and characterize evolutions datasets, as well as comparing the histories of individual artifacts, both in terms of 1) observing how rapidly different artifacts age over multiple-year-long time spans 2) how often metrics associated with each artifacts tend towards an improvement or worsening. The paper shows how the tool can be tailored to specific analysis needs (pace of change comparison, clone detection, freshness assessment) thanks to its support for flexible configuration of history scaling and alignment along the time axis, artifacts sorting and interactive color mapping, enabling the analysis of millions of events obtained by mining the histories of tens of thousands of software artifacts. We include in this paper a gallery showcasing datasets gathering specific artifacts (OpenAPI descriptions, GitHub workflow definitions) across multiple repositories, as well as diving into the history of specific popular open source projects. Souhaila Serbout, Diana Carolina Muñoz Hurtado, Hassan Atwi, Edoardo Riggio, Cesare Pautasso |
VISSOFT | 1 |
| 2025 | APICANVAS: Graphically Designing Web APIsabstractWeb APIs are essential in modern software systems, with OpenAPI as the standard specification for describing RESTful APIs. Despite widespread adoption, OpenAPI specifications are mainly textual documents that become verbose, deeply nested, and difficult to navigate at scale. Current tools like Swagger Editor lack graphical design capabilities, impeding effective prototyping for large teams and non-technical users. To address this gap, we present APICAnvas, a visual design tool enabling users to interact with OpenAPI specifications through both editable YAML and graphical interfaces with bidirectional synchronization. We evaluated the tool with 13 participants across two tasks: creating new APIs and modifying existing ones. Results show 83% successfully created APIs from scratch and 75% correctly updated existing APIs, demonstrating the tool’s effectiveness across diverse experience levels. Griffin Tomaszewski, Souhaila Serbout, Wesley K. G. Assunção |
VL/HCC | 2 |
| 2025 | Understanding security tactics in microservice APIs using annotated software architecture decomposition models - a controlled experimentabstractWhile microservice architectures have become a widespread option for designing distributed applications, designing secure microservice systems remains challenging. Although various security-related guidelines and practices exist, these systems' sheer size, complex communication structures, and polyglot tech stacks make it difficult to manually validate whether adequate security tactics are applied throughout their architecture. To address these challenges, we have devised a novel solution that involves the automatic generation of security-annotated software decomposition models and the utilization of security-based metrics to guide software architectures through the assessment of security tactics employed within microservice systems. To evaluate the effectiveness of our artifacts, we conducted a controlled experiment where we asked 60 students from two universities and ten experts from the industry to identify and assess the security features of two microservice reference systems. During the experiment, we tracked the correctness of their answers and the time they needed to solve the given tasks to measure how well they could understand the security tactics applied in the reference systems. Our results indicate that the supplemental material significantly improved the correctness of the participants' answers without requiring them to consult the documentation more. Most participants also stated in a self-assessment that their understanding of the security tactics used in the systems improved significantly because of the provided material, with the additional diagrams considered very helpful. In contrast, the perception of architectural metrics varied widely. We could also show that novice developers benefited most from the supplementary diagrams. In contrast, senior developers could rely on their experience to compensate for the lack of additional help. Contrary to our expectations, we found no significant correlation between the time spent solving the tasks and the overall correctness score achieved, meaning that participants who took more time to read the documentation did not automatically achieve better results. As far as we know, this empirical study is the first analysis that explores the influence of security annotations in component diagrams to guide software developers when assessing microservice system security. Patric Genfer, Souhaila Serbout, Georg Simhandl, Uwe Zdun, Cesare Pautasso |
Empir. Softw. Eng. | 2 |
| 2024 | How Many Web APIs Evolve Following Semantic Versioning?
Souhaila Serbout, Cesare Pautasso |
ICWE | 1 |
| 2024 | APIstic: A Large Collection of OpenAPI MetricsabstractIn the rapidly evolving landscape of web services, the significance of efficiently designed and well-documented APIs is paramount. In this paper, we present APIstic an API analytics dataset and exploration tool to navigate and segment APIs based on an extensive set of pre-computed metrics extracted from OpenAPI specifications, sourced from GitHub, SwaggerHub, BigQuery and APIs.guru. These pre-computed metrics are categorized into structure, data model, natural language description, and security metrics. The extensive dataset of varied API metrics provides crucial insights into API design and documentation for both researchers and practitioners. Researchers can use APIstic as an empirical resource to extract refined samples, analyze API design trends, best practices, smells, and patterns. For API designers, it serves as a benchmarking tool to assess, compare, and improve API structures, data models, and documentation using metrics to select points of references among 1,275,568 valid OpenAPI specifications. The paper discusses potential use cases of the collected data and presents a descriptive analysis of selected API analytics metrics. Souhaila Serbout, Cesare Pautasso |
MSR | 1 |
| 2024 | How Are Web APIs Versioned in Practice? A Large-Scale Empirical StudyabstractWeb APIs form the cornerstone of modern software ecosystems, facilitating seamless data exchange and service integration. Ensuring the compatibility and longevity of these APIs is paramount. This study delves into the intricate realm of API versioning practices, a crucial mechanism for managing API evolution. Exploring an expanded and diverse dataset of 603 293 APIs specifications created during the 2015–2023 timeframe and gathered from four different sources, we examined the adoption of the following versioning practices: Metadata-based, URL-based, Header-based and Dynamic versioning, with one or more versions in production. API developers use more than 50 different version identifier formats to encode information about the changes introduced with respect to the previous version (i.e., semantic versioning), about when the version was released (i.e., age versioning) and about which phase of the API development lifecycle the version belongs (i.e., stable vs. preview releases). Souhaila Serbout, Cesare Pautasso |
J. Web Eng. | 1 |
| 2023 | An Empirical Study of Web API Versioning Practices
Souhaila Serbout, Cesare Pautasso |
ICWE | 1 |
| 2023 | Interactively Exploring API Changes and Versioning ConsistencyabstractApplication Programming Interfaces (APIs) evolve over time. As they change, they are expected to be versioned based on how changes might affect their clients. In this paper, we present two novel visualizations specifically designed to represent all structural changes and the level of adherence to semantic versioning practices over time. They can also serve for characterizing and comparing the evolution history of different Web APIs. The API Version Clockhelps to visualize the sequence of API changes over time and highlight inconsistencies between major, minor, or patch version changes and the corresponding introduced breaking or non-breaking changes applied to the API. The API Changesoverview aggregates all changes to an OpenAPI (OAS) description, highlighting the unstable vs. the stable elements of the API over its entire history. Both visualizations can be automatically created using the APIcTURE, a command-line and web-based tool that analyzes the histories of git code repositories containing OAS descriptions, extracting the necessary data for generating visualizations and computing metrics related to API evolution and versioning. The visualizations have been successfully applied to classify, compare, and interactively explore the multi-year evolution history of APIs with up to hundreds of individual commits. Video URL: https://youtu.be/WtFm6VvKi20 Souhaila Serbout, Diana Carolina Muñoz Hurtado, Cesare Pautasso |
VISSOFT | 1 |
| 2022 | To Deprecate or to Simply Drop Operations? An Empirical Study on the Evolution of a Large OpenAPI Collection
Fabio Di Lauro, Souhaila Serbout, Cesare Pautasso |
ECSA | 2 |
| 2022 | How Composable is the Web? An Empirical Study on OpenAPI Data model CompatibilityabstractComposing Web APIs is a widely adopted practice by developers to speed up the development process of complex Web applications, mashups, and data processing pipelines. However, since most publicly available APIs are built independently of each other, developers often need to invest their efforts in solving incompatibility issues by writing ad-hoc glue code, adapters and message translation mappings. How likely are Web APIs to be directly composable?The paper presents an empirical study to determine the potential composability of a large collection of 20,587 public Web APIs by verifying their schemas’ compatibility. We define three levels of data model elements compatibility – considering matches between property names and/or data types – which can be determined statically based on API descriptions conforming to the OpenAPI specification. The study research questions address: to which extent are Web APIs compatible; the average number of compatible endpoints within each API; the likelihood of finding two APIs with at least one pair of compatible endpoints.To perform the analysis we developed a compatibility checker tool which can statically determine API schema compatibility on the three levels and find matching pairs of API responses which can be directly forwarded as requests to the same or other APIs. We run the tool on a dataset of 751,390 request and response message schemas extracted from publicly available OpenAPI descriptions.The results indicate a relatively high number of compatible APIs when matching their data models only on the level of their elements’ data type. However, this number gets lower narrowing the scope to only the ones handling data objects having identical properties name. The average likelihood of finding two compatible APIs with both matching property names and data types reaches 21%. Also, the number of compatible endpoints within the same API is very low. Souhaila Serbout, Cesare Pautasso, Uwe Zdun |
ICWS | 1 |
| 2022 | A Large-scale Empirical Assessment of Web API Size EvolutionabstractLike any other type of software, also Web Application Programming Interfaces (APIs) evolve over time. In the case of widely used API, introducing changes is never a trivial task, because of the risk of breaking thousands of clients relying on the API. In this paper we conduct an empirical study over a large collection of OpenAPI descriptions obtained by mining open source repositories. We measure the speed at which Web APIs change and how changes affect their size, simply defined as the number of operations. The dataset of API descriptions was collected over a period of one year and includes APIs with histories spanning across up to 7 years of commits. The main finding is that APIs tend to grow, although some do reduce their size, as shown in the case study examples included in the appendix. Fabio Di Lauro, Souhaila Serbout, Cesare Pautasso |
J. Web Eng. | 2 |
| 2022 | Live process modeling with the BPMN Sketch MinerabstractBPMN Sketch Miner is a modeling environment for generating visual business process models starting from constrained natural language textual input. Its purpose is to support business process modelers who need to rapidly sketch visual BPMN models during interviews and design workshops, where participants should not only provide input but also give feedback on whether the sketched visual model represents accurately what has been described during the discussion. In this article, we present a detailed description of the BPMN Sketch Miner design decisions and list the different control flow patterns supported by the current version of its textual DSL. We also summarize the user study and survey results originally published in MODELS 2020 concerning the tool usability and learnability and present a new performance evaluation regarding the visual model generation pipeline under actual usage conditions. The goal is to determine whether it can support a rapid model editing cycle, with live synchronization between the textual description and the visual model. This study is based on a benchmark including a large number of models (1350 models) exported by users of the tool during the year 2020. The main results indicate that the performance is sufficient for a smooth live modeling user experience and that the end-to-end execution time of the text-to-model-to-visual pipeline grows linearly with the model size, up to the largest models (with 195 lines of textual description) found in the benchmark workload. Ana Ivanchikj, Souhaila Serbout, Cesare Pautasso |
Softw. Syst. Model. | 2 |
| 2021 | Towards Large-Scale Empirical Assessment of Web APIs Evolution
Fabio Di Lauro, Souhaila Serbout, Cesare Pautasso |
ICWE | 2 |
| 2020 | From text to visual BPMN process models: design and evaluationabstractMost existing Business Process Model and Notation (BPMN) editing tools are graphical, and as such based on explicit modeling, requiring good knowledge of the notation and its semantics, as well as the ability to analyze and abstract business requirements and capture them by correctly using the notation. As a consequence, their use can be cumbersome for live modeling during interviews and design workshops, where participants should not only provide input but also give feedback on how it has been represented in a model. To overcome this, in this paper we present the design and evaluation of BPMN Sketch Miner, a tool which combines notes taking in constrained natural language with process mining to automatically produce BPMN diagrams in real-time as interview participants describe them with stories. In this work we discuss the design decisions regarding the trade-off between using mining vs. modelling in order to: 1) support a larger number of BPMN constructs in the textual language; 2) target both BPMN beginners and business analysts, in addition to the process participants themselves. The evaluation of the new version of the tool in terms of how it balances the expressiveness and learnability of its DSL with the usability of the text-to-visual sketching environment shows encouraging results. Namely, while BPMN beginners could model a non-trivial process with the tool in a relatively short time and with good accuracy, business analysts appreciated the usability of the tool and the expressiveness of the language in terms of supported BPMN constructs. Ana Ivanchikj, Souhaila Serbout, Cesare Pautasso |
MoDELS | 2 |
| 2020 | Defining Referential Integrity Constraints in Graph-oriented DatastoresabstractNowadays, the volume of data manipulated by our information systems is growing so rapidly that they cannot be efficiently managed and exploited only by means of standard relational data management systems. Hence the recent emergence of NoSQL datastores as alternative/complementary choices for big data management. While NoSQL datastores are usually designed with high performance and scalability as primary concerns, this often comes at a cost of tolerating (temporary) data inconsistencies. This is the case, in particular, for managing referential integrity in graph-oriented datastores, for which no support currently exists. This paper presents a MDE-based, tool-supported approach to the definition and enforcement of referential integrity constraints (RICs) in graph-oriented NoSQL datastores. This approach relies on a domain-specific language allowing users to specify RICs as well as the way they must be managed. This specification is then exploited to support the automated identification and correction of RICs violations in a graph-oriented datastore. We illustrate the application of our approach, currently implemented for Neo4J, through a small experiment. Thibaud Masson, Romain Ravet, Francisco Javier Bermudez Ruiz, Souhaila Serbout, Diego Sevilla Ruiz, Anthony Cleve |
MODELSWARD | 4 |