Loïc Cudennec

dblp:76/6367 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
5since 2021 · last 2026
0000-0002-6476-4574ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 1 first-author · 5 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Evaluating Federated Learning Beyond Simulation: A Deployment-Aware Methodology
Mathis Valli, Cédric Tedeschi, Gabriel Antoniu, Loïc Cudennec, Alexandru Costan
CCGrid4
2025 On the Reproducibility Challenges of Federated Learning: Investigating the Gap Between Simulation, Emulation and Real-World Deployments
abstract
Federated Learning (FL) is an emerging paradigm for decentralized training of Machine Learning models. It has been the subject of a large corpus of research due to its innovative approach to handling sensitive data. A common practice in the FL literature is to run simulations on a single compute node to assess the performance of FL algorithms. While simulation enables fast prototyping and validation of algorithmic concepts, it may face limitations in reproducing the real system's performance in heterogeneous environments such as the Computing Continuum, and particularly on resource-constrained Edge devices. Conversely, emulation on distributed testbeds offers more effective means to accurately reproduce the performance of real-world devices. However, to the best of our knowledge, no prior research has investigated the differences between simulation and emulation in FL experiments. In this paper, we study the complementarity of these approaches and discuss their respective challenges, as a first step towards reproducibility of FL experiments. We illustrate our study with a real-life application used as a baseline: an outdoor air quality forecasting framework with real-world sensors. Our results show that simulation can be used to accurately reproduce model performance metrics, while emulation can effectively reproduce the system performance of real-world experiments. Finally, we present a set of lessons learned on the challenges of FL reproducibility and the selection of experimental infrastructures for FL experiments and applications.
Cèdric Prigent, Kate Keahey, Alexandru Costan, Loïc Cudennec, Gabriel Antoniu
CCGrid4
2024 Efficient Resource-Constrained Federated Learning Clustering with Local Data Compression on the Edge-to-Cloud Continuum
abstract
Federated Learning (FL) has been proposed as a privacy-preserving approach for distributed learning over decen-tralized resources. While it can be a highly efficient tool for large-scale collaborative training of Machine Learning (ML) models, its efficiency may be strongly impacted by a high variability in data distributions among clients. Clustered FL tackles this problem by grouping clients with similar data distributions and training personalized models. Despite increasing model accuracy for federated peers, existing clustering approaches overlook system and infrastructure constraints leading to sustainability problems for resource-constrained devices. This paper introduces a new method for resource-constrained FL clustering. We leverage pre-trained autoencoders to compress client data into low dimensional space and build lightweight em-bedding vectors used to cluster federated clients. A randomized quantization approach specifically secures the client embedding vectors against data reconstruction. Extensive experiments using a multi-GPU testbed with multiple scenarios introducing concept drift between clients demonstrate the generalitity of our approach to personalized FL. By minimizing the overall system overhead and improving the model convergence, our approach reduces model training cost by up to l.44× − 4.32× communication, l.03× − 2.40× training time compared to IFCA and l.0× − 8.60× communication, 0.87× − 0.87× training time compared to LADD to achieve similar accuracy in the different evaluation scenarios. While each of the baselines encounters performance degradation in at least one of the scenarios, our strategy demonstrates top efficiency in all of them.
Cèdric Prigent, Melvin Chelli, Alexandru Costan, Loïc Cudennec, René Schubotz, Gabriel Antoniu
HiPC4
2024 Enabling federated learning across the computing continuum: Systems, challenges and future directions
Cèdric Prigent, Alexandru Costan, Gabriel Antoniu, Loïc Cudennec
Future Gener. Comput. Syst.4
2023 FedGuard: Selective Parameter Aggregation for Poisoning Attack Mitigation in Federated Learning
abstract
Minimizing the attack surface of Federated Learning (FL) systems is a field of active research. FL turns out to be highly vulnerable to various threats coming from the edge of the network. Current approaches rely on robust aggregation, anomaly detection and generative models for defending against poisoning attacks. Yet, they either have limited defensive capabilities due to their underlying design or are impractical to use as they rely on constraining building blocks.We introduce FedGuard, a novel FL framework that utilizes the generative capabilities of Conditional Variational AutoEncoders (CVAE) to effectively defend against poisoning attacks with tuneable overhead in communication and computation. Whilst the idea of hardening a FL system using generative models is not entirely new, FedGuard’s original contribution is in its selective parameter aggregation operator with parameter selection being driven by synthetic validation data sampled from the CVAEs trained locally by each participating party.Experimental evaluations in a 100-client setup demonstrates FedGuard to be more effective than previous approaches against several types of attacks (label and sign flipping, additive noise, same value attacks). FedGuard successfully defends in scenarios with up to 50% malicious peers where other strategies fail. In addition, FedGuard does not require auxiliary datasets or centralized (pre-) training. It provides resilience against poisoning attacks from the very first round of federated training.
Melvin Chelli, Cèdric Prigent, René Schubotz, Alexandru Costan, Gabriel Antoniu, Loïc Cudennec, Philipp Slusallek
CLUSTER6
2020 A combined fast/cycle accurate simulation tool for reconfigurable accelerator evaluation: application to distributed data management
abstract
Parallel computing systems based on reconfigurable accelerators are becoming (1) increasingly heterogeneous, (2) difficult to design and (3) complex to model. Such modeling of a parallel computing system helps to evaluate its performance and to improve its architecture before prototyping. This paper presents a simulation tool aiming to study the integration of reconfigurable accelerators in scalable distributed systems and runtimes, such as S-DSM systems, where S-DSM (software-distributed shared memory) is a paradigm to ease data management among distributed nodes. This tool allows us to simulate the execution of irregular compute kernels accessing distributed data. To deal with the complexity of modeling (3) the complete system we used a hybrid methodology. We integrated the simulation engine into the S-DSM. The distributed data management part is executed on the physical architecture allowing to generate precise and faithful latencies, and the accelerator simulation is cycle accurate. We used general sparse matrix-matrix multiplication (SpGEMM) as a case study. We show that the use of this tool makes it possible to analyze the behavior of an heterogeneous system (1) with rapid prototyping and simulation. The analysis of the results allowed to determine the correct sizing of the architecture (2) to obtain the best performance. The tool allowed to identify the bottleneck of our architecture and confirmed the possibility of hiding data access latencies. Our simulation platform allows to emulate a heterogeneous distributed system by introducing a slowdown between 1.2 and 3.7 times compared to the compute kernel simulation alone.
Erwan Lenormand, Thierry Goubier, Loïc Cudennec, Henri-Pierre Charles
RSP3
2020 Adaptive message passing polling for energy efficiency: Application to software-distributed shared memory over heterogeneous computing resources
abstract
Summary Autonomous vehicles, smart manufacturing, heterogeneous systems, and new high‐performance embedded computing (HPEC) applications can benefit from the reuse of code coming from the high‐performance computing world. However, unlike for HPC, energy efficiency is critical in embedded systems, especially when running on battery power. Code base from HPC mostly relies on the message passing interface (MPI) message passing runtime to deal with distributed systems. MPI has been designed primarily for performance and not for energy efficiency. One drawback is the way messages are received, in an energy‐consuming busy‐wait fashion. In this work, we study a simple approach in which receiving processes are put to sleep instead of constantly polling. We implement this strategy at the user level to be transparent to the MPI runtime and the application. Experiments are conducted with OpenMPI, MPICH, and MPC, using a video processing application and a software‐distributed shared memory system deployed over two heterogeneous platforms, including the Christmann RECS|Box Antares Microserver. Results show significant energy savings. In some particular cases involving process colocation, we also observe better performance using our strategy which can be explained by a better sharing of the computing resource.
Loïc Cudennec
Concurr. Comput. Pract. Exp.1
2016 The M2DC Project: Modular Microserver DataCentre
abstract
The Modular Microserver DataCentre (M2DC) project will investigate, develop and demonstrate a modular, highly-efficient, cost-optimized server architecture composed of heterogeneous microserver computing resources, being able to be tailored to meet requirements from various application domains such as image processing, cloud computing or HPC. M2DC will be built on three main pillars: a flexible server architecture that can be easily customised, maintained and updated, advanced management strategies and system efficiency enhancements (SEE), well-defined interfaces to surrounding software data centre ecosystem.
Mariano Cecowski, Giovanni Agosta, Ariel Oleksiak, Michal Kierzynka, Micha vor dem Berge, Wolfgang Christmann, Stefan Krupop, Mario Porrmann, Jens Hagemeyer, René Griessl, Meysam Peykanu, Lennart Tigges, Sven Rosinger, Daniel Schlitt, Christian Pieper, Carlo Brandolese, William Fornaciari, Gerardo Pelosi, Robert Plestenjak, Justin Cinkelj, Loïc Cudennec, Thierry Goubier, Jean-Marc Philippe, Udo Janssen, Chris Adeniyi-Jones
DSD21
2016 Deploying software functionality to manufacturing resources safely at runtime
abstract
Automated manufacturing systems are becoming increasingly flexible in order to support a growing number of different products and product variations, as well as shortening lot sizes and product life cycles. Many manufacturing resource are already multi-purpose. But they are integrated in an automation infrastructure that may require significant effort to adapt. In this contribution, we present a system architecture for the deployment of software functionality to manufacturing resources safely at runtime. The architecture integrates software model verification to ensure the integrity of new functionality, software-defined networking (SDN) to establish new communication pathways for collaboration, and the close interaction between a flexible runtime execution environment and its safety-critical counterpart.
Julius Pfrommer, Miriam Schleipen, Selma Azaiez, Michael Boc, Loïc Cudennec, Selma Kchir, Xenia Klinge
ETFA5
2008 Building Hierarchical Grid Storage Using the GfarmGlobal File System and the JuxMemGrid Data-Sharing Service
Gabriel Antoniu, Loïc Cudennec, Majd Ghareeb, Osamu Tatebe
Euro-Par2
2007 Performance scalability of the JXTA P2P framework
abstract
Features of the P2P model, such as scalability and volatility tolerance, have motivated its use in distributed systems. Several generic P2P libraries have been proposed for building distributed applications. However, very few experimental evaluations of these frameworks have been conducted, especially at large scales. Such experimental analyses are important, since they can help system designers to optimize P2P protocols and better understand the benefits of the P2P model. This is particularly important when the P2P model is applied to special use cases, such as grid computing. This paper focuses on the scalability of two main protocols proposed by the JXTA P2P platform. First, we provide a detailed description of the underlying mechanisms used by JXTA to manage its overlay and propagate messages over it: the rendezvous protocol. Second, we describe the discovery protocol used to find resources inside a JXTA network. We then report a detailed, large-scale, multi-site experimental evaluation of these protocols, using the nine clusters of the French Grid'5000 testbed.
Gabriel Antoniu, Loïc Cudennec, Mathieu Jan, Mike Duigou
IPDPS2
2006 Extending the Entry Consistency Model to Enable Efficient Visualization for Code-Coupling Grid Applications
abstract
This paper addresses the problem of efficient visualization of shared data within code coupling grid applications. These applications are structured as a set of distributed, autonomous, weakly-coupled codes. We focus on the case where the codes are able to interact using the abstraction of a shared data space. We propose an efficient visualization scheme by adapting the mechanisms used to maintain the data consistency. We introduce a new operation called relaxed read, as an extension to the entry consistency model. This operation can efficiently take place without locking, in parallel with write operations. We discuss the benefits and the constraints of the proposed approach.
Gabriel Antoniu, Loïc Cudennec, Sébastien Monnet
CCGRID2