VLDB 2026 Research / reviewers in the wild / expert
Amine Barrak
dblp:224/1604
· DBLP profile ↗
13ranked-venue papers
9as first author
12since 2021 · last 2026
0000-0002-0046-2454ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 8 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ML in a Box: Analyzing Containerization Practices in Open Source ML ProjectsabstractContainerization has become increasingly essential in the machine learning (ML) domain, providing reproducibility, portability, and environment consistency. While prior studies have analyzed Dockerfile structures and best practices, none have examined ML projects in depth to reveal how the iterative nature of ML workflows influences container footprint, build performance, and caching behavior. Faten Jebari, Emna Ksontini, Amine Barrak, Wael Kessentini |
MSR | 3 |
| 2026 | A Large-Scale Dataset of MCP Implementations on GitHubabstractThe rapid emergence of the Model Context Protocol (MCP) has introduced a new standard for connecting large language models to external tools and services. Despite its rapid adoption in open-source development, systematic understanding of how MCP is implemented, structured, and maintained remains limited. This study presents the first large-scale, evidence-based dataset of real-world MCP implementation collected directly from GitHub. Using a hybrid pipeline that integrates the GitHub REST and GraphQL APIs with custom Python verification scripts, 3,238 candidate repositories were discovered, filtered, and validated through multi-stage evidence checks. Each verified project was classified by operational role (e.g., client, server, gateway) and exported in a reproducible JSONL schema. A manual review of a representative subset confirmed an overall precision of 83% at a 95% confidence level, and additionally revealed a set of repositories functioning primarily as educational samples, tutorials, or demonstration templates. A targeted exclusion rule was then applied to remove these non-operational repositories, resulting in a final dataset of 2,297 validated MCP projects. The analysis shows that Python and TypeScript dominate MCP development, with hybrid architectures emerging as the most common design pattern. By emphasizing transparent verification strategies, structured evidence tagging, and reproducible data organization, this work establishes a foundational benchmark for studying real-world MCP ecosystems and supports future research on integration, connectivity, and compatibility across the broader developer community. Benny Toeppe, Amine Barrak, Emna Ksontini |
MSR | 2 |
| 2025 | Cost-Performance Analysis: A Comparative Study of CPU-Based Serverless and GPU-Based Training Architectures
Amine Barrak, Fábio Petrillo, Fehmi Jaafar |
PDCAT | 1 |
| 2025 | FaaSGuard: Secure CI/CD for Serverless Applications - An OpenFaaS Case StudyabstractServerless computing significantly alters software development by abstracting infrastructure management and enabling rapid, modular, event-driven deployments. Despite its benefits, the distinct characteristics of serverless functions, such as ephemeral execution and fine-grained scalability, pose unique security challenges, particularly in open-source platforms like OpenFaaS. Existing approaches typically address isolated phases of the DevSecOps lifecycle, lacking an integrated and comprehensive security strategy. To bridge this gap, we propose FaaSGuard, a unified DevSecOps pipeline explicitly designed for open-source serverless environments. FaaSGuard systematically embeds lightweight, fail-closed security checks into every stage of the development lifecycle—planning, coding, building, deployment, and monitoring—effectively addressing threats such as injection attacks, hard-coded secrets, and resource exhaustion. We validate our approach empirically through a case study involving 20 real-world serverless functions from public GitHub repositories. Results indicate that FaaSGuard effectively detects and prevents critical vulnerabilities, demonstrating high precision (95%) and recall (91%) without significant disruption to established CI/CD practices. Amine Barrak, Emna Ksontini, Ridouane Atike, Fehmi Jaafar |
SCAM | 1 |
| 2024 | The Promise of Serverless Computing within Peer-to-Peer Architectures for Distributed ML TrainingabstractMy thesis focuses on the integration of serverless computing with Peer to Peer (P2P) architectures in distributed Machine Learning (ML). This research aims to harness the decentralized, resilient nature of P2P systems, combined with the scalability and automation of serverless platforms. We explore using databases not just for communication but also for in-database model updates and gradient averaging, addressing the challenges of statelessness in serverless environments. Amine Barrak |
AAAI | 1 |
| 2024 | Incorporating Serverless Computing into P2P Networks for ML Training: In-Database Tasks and Their Scalability Implications (Student Abstract)abstractDistributed ML addresses challenges from increasing data and model complexities. Peer to peer (P2P) networks in distributed ML offer scalability and fault tolerance. However, they also encounter challenges related to resource consumption, and communication overhead as the number of participating peers grows. This research introduces a novel architecture that combines serverless computing with P2P networks for distributed training. Serverless computing enhances this model with parallel processing and cost effective scalability, suitable for resource-intensive tasks. Preliminary results show that peers can offload expensive computational tasks to serverless platforms. However, their inherent statelessness necessitates strong communication methods, suggesting a pivotal role for databases. To this end, we have enhanced an in memory database to support ML training tasks. Amine Barrak |
AAAI | 1 |
| 2024 | Securing AWS Lambda: Advanced Strategies and Best PracticesabstractThe emergence of the serverless paradigm, embodied by AWS Lambda functions, has revolutionized the landscape of cloud computing. This model empowers users to offload server management tasks, allowing them to focus their efforts on core business logic while achieving substantial cost savings. However, this transition to serverless exposes significant vulnerabilities, especially in terms of security. This article delves into the specific security challenges associated with AWS Lambda functions, with a focus on major threats such as malicious code injection, sensitive data leaks, DDoS attacks, excessive privileges, vulnerable dependencies, and certificate issues. Our investigation, centered around the AWS Lambda platform, thoroughly analyzes these challenges by identifying underlying mechanisms and inherent risks. We review the state of the art solutions from the literature while examining the strategies adopted by AWS and the industry to enhance security. By implementing these solutions on an AWS server, we concretely illustrate possible protective measures. In this paper, we aims to provide a comprehensive understanding of security issues in the context of Lambda functions, paving the way for recommendations and research directions to bolster the resilience of this essential serverless cloud technology. Amine Barrak, Gildas Fofe, Léo Mackowiak, Emmanuel Kouam, Fehmi Jaafar |
CSCloud | 1 |
| 2023 | Exploring the Impact of Serverless Computing on Peer To Peer Training Machine LearningabstractThe increasing demand for computational power in big data and machine learning has driven the development of distributed training methodologies. Among these, peer-to-peer (P2P) networks provide advantages such as enhanced scalability and fault tolerance. However, they also encounter challenges related to resource consumption, costs, and communication overhead as the number of participating peers grows. In this paper, we introduce a novel architecture that combines serverless computing with P2P networks for distributed training and present a method for efficient parallel gradient computation under resource constraints.Our findings show a significant enhancement in gradient computation time, with up to a 97.34% improvement compared to conventional P2P distributed training methods. As for costs, our examination confirmed that the serverless architecture could incur higher expenses, reaching up to 5.4 times more than instance-based architectures. It is essential to consider that these higher costs are associated with marked improvements in computation time, particularly under resource-constrained scenarios.Despite the cost-time trade-off, the serverless approach still holds promise due to its pay-as-you-go model. Utilizing dynamic resource allocation, it enables faster training times and optimized resource utilization, making it a promising candidate for a wide range of machine learning applications. Amine Barrak, Ranim Trabelsi, Fehmi Jaafar, Fábio Petrillo |
IC2E | 1 |
| 2023 | SPIRT: A Fault-Tolerant and Reliable Peer-to-Peer Serverless ML Training ArchitectureabstractThe advent of serverless computing has ushered in notable advancements in distributed machine learning, particularly within parameter server-based architectures. Yet, the integration of serverless features within peer-to-peer (P2P) distributed networks remains largely uncharted. In this paper, we introduce SPIRT, a fault-tolerant, reliable, scalable and secure serverless P2P ML training architecture. designed to bridge this existing gap. Capitalizing on the inherent robustness and reliability innate to P2P systems, we emphasized Intra-peer scalability for concurrent gradient to mitigate communication overhead from increased peer interactions. SPIRT, employs RedisAI for in-database operations, achieves an 82% reduction in model update times. This architecture showcases resilience against peer failures and adeptly manages the integration of new peers. Furthermore, SPIRT ensures secure communication between peers, enhancing the reliability of distributed machine learning tasks. Even in the face of Byzantine attacks, the system’s robust aggregation algorithms maintain high levels of accuracy. These findings illuminate the promising potential of serverless architectures in P2P distributed machine learning, offering a significant stride towards the development of more efficient, scalable, and resilient applications. Amine Barrak, Mayssa Jaziri, Ranim Trabelsi, Fehmi Jaafar, Fábio Petrillo |
QRS | 1 |
| 2021 | Identification of Compromised IoT Devices: Combined Approach Based on Energy Consumption and Network Traffic AnalysisabstractIn the burgeoning age of digitalization, the Internet of Things presents a core part of the digital ecosystem. Unfortunately, as the deployment of connected devices is increasing tremendously, so are cyber-attacks. The consequences of cyber-attacks could be devastating as they gain access to sensitive data and even damages critical infrastructures. This urges the development and integration of proactive and intelligent security breach detection mechanisms in different levels of the IoT platforms including the devices themselves. Several empirical observations indicated a change in the energy consumption and network behaviour of compromised devices. Thus, we propose in this paper a machine learning based approach to identify compromised IoT devices using their energy consumption footprint and network traffic. We base our study on real data collected from real experiments using different commercially available IoT devices infected with authentic IoT botnets. Our results show that machine learning algorithms can classify correctly attacks reaching 98.40% precision for Mirai, over 99.91% for Ufonet and respectively 97.63% and 99.93% performance. Overall, our exploratory study is one of the very first of its kind to explore the energy consumption combined with network behavior analysis to detect IoT compromised devices and its outcomes will be a starting point for further research on this topic. Fehmi Jaafar, Darine Ameyed, Amine Barrak, Mohamed Cheriet |
QRS | 3 |
| 2021 | On the Co-evolution of ML Pipelines and Source Code - Empirical Study of DVC ProjectsabstractThe growing popularity of machine learning (ML) applications has led to the introduction of software engineering tools such as Data Versioning Control (DVC), MLFlow and Pachyderm that enable versioning ML data, models, pipelines and model evaluation metrics. Since these versioned ML artifacts need to be synchronized not only with each other, but also with the source and test code of the software applications into which the models are integrated, prior findings on co-evolution and coupling between software artifacts might need to be revisited. Hence, in order to understand the degree of coupling between ML-related and other software artifacts, as well as the adoption of ML versioning features, this paper empirically studies the usage of DVC in 391 Github projects, 25 of which in detail. Our results show that more than half of the DVC files in a project are changed at least once every one-tenth of the project's lifetime. Furthermore, we observe a tight coupling between DVC files and other artifacts, with 1/4 pull requests changing source code and 1/2 pull requests changing tests requiring a change to DVC files. As additional evidence of the observed complexity associated with adopting ML-related software engineering tools like DVC, an average of 78% of the studied projects showed a non-constant trend in pipeline complexity. Amine Barrak, Ellis E. Eghan, Bram Adams |
SANER | 1 |
| 2021 | Why do builds fail? - A conceptual replication study
Amine Barrak, Ellis E. Eghan, Bram Adams, Foutse Khomh |
J. Syst. Softw. | 1 |
| 2018 | The State of Practice on Virtual Reality (VR) Applications: An Exploratory Study on Github and Stack OverflowabstractVirtual Reality (VR) is a computer technology that holds the promise of revolutionizing the way we live. The release in 2016 of new-generation headsets from Facebook-owned Oculus and HTC has renewed the interest in that technology. Thousands of VR applications have been developed over the past years, but most software developers lack formal training on this technology. In this paper, we propose descriptive information on the state of practice of VR applications' development to understand the level of maturity of this new technology from the perspective of Software Engineering (SE). To do so, we focused on the analysis of 320 VR open source projects from Github to determine which are the most popular languages and engines used in VR projects, and evaluate the quality of the projects from a software metric perspective. To get further insights on VR development, we also manually analyzed nearly 300 questions from Stack Overflow. Our results show that (1) VR projects on GitHub are currently mostly small to medium projects, and (2) the most popular languages are JavaScript and C#. Unity is the most used game engine during VR development and the most discussed topic on Stack Overflow. Overall, our exploratory study is one of the very first of its kind for VR projects and provides material that is hopefully a starting point for further research on challenges and opportunities for VR software development. Naoures Ghrairi, Segla Kpodjedo, Amine Barrak, Fábio Petrillo, Foutse Khomh |
QRS | 3 |