EDBT 2026 Demo / reviewers in the wild / expert
Olivier Richard
dblp:81/1910
· DBLP profile ↗
26ranked-venue papers
2as first author
3since 2021 · last 2025
0009-0005-8679-2874ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1Computer networks · 1Security and privacy · 1Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Dynamic Resource Management in HPC Systems Using Dynamic Processes with PSetsabstractWith the increasing scale of High-Performance Computing (HPC) systems and a new awareness of the environmental impact of HPC, new strategies are required to improve the efficiency of resource usage on these systems. One such strategy is Dynamic Resource Management (DRM), which allows changing the resources assigned to a job dynamically during its execution. This increased flexibility in resource allocation and job scheduling can lead to improvements in several system efficiency metrics. Despite these benefits, DRM has not yet been established as a ready-to-use technology for production HPC systems. This is caused by the significant changes required in all the layers of the HPC system software stack, which are only achievable with an extensive and holistic co-design process between resource management software and applications. In this work, we demonstrate the applicability of a recently introduced, generic design approach for dynamic resources called Dynamic Processes with PSets (DPP), to enable DRM in realworld systems. To this end, we developed an exemplary, dynamic system software stack implementation following the DPP design principles throughout all layers. Based on this, we assess the applicability and performance of our approach using both synthetic benchmarks and job mixes consisting of several dynamic, real-world applications. On up to$\mathbf{1 0 0}$nodes, we measure moderate overheads for process reconfiguration in applications while significantly improving the system throughput and average job turnaround time compared to static scheduling in crowded system scenarios. Dominik Huber, Keerthi Gaddameedi, Tobias Neckel, Hans-Joachim Bungartz, Martin Schulz 0001, Pierre-François Dutot, Olivier Richard, Martin Schreiber 0001, Sergio Iserte, Antonio J. Peña |
HiPC | 7 |
| 2022 | Painless Transposition of Reproducible Distributed Environments with NixOS ComposeabstractDevelopment of environments for distributed systems is a tedious and time-consuming iterative process. The reproducibility of such environments is a crucial factor for rigorous scientific contributions. We think that being able to smoothly test environments both locally and on a target dis-tributed platform makes development cycles faster and reduces the friction to adopt better experimental practices. To address this issue, this paper introduces the notion of environment transposition and implements it in NixOS Compose, a tool that generates reproducible distributed environments. It enables users to deploy their environments on virtualized (docker, QEMU) or physical (Grid'5000) platforms with the same unique description of the environment. We show that NixOS Compose enables to build reproducible environments without overhead by comparing it to state-of-the-art solutions for the generation of distributed environments (EnOSlib and Kameleon). NixOS Compose actually enables substantial performance improvements on image building time over Kameleon (up to 11x faster for initial builds and up to 19x faster when building a variation of an existing environment). Quentin Guilloteau, Jonathan Bleuzen, Millian Poquet, Olivier Richard |
CLUSTER | 4 |
| 2022 | A Methodology to Scale Containerized HPC Infrastructures in the Cloud
Nicolas Grenèche, Tarek Menouer, Christophe Cérin, Olivier Richard |
Euro-Par | 4 |
| 2018 | Reducing the number of response time service level objective violations by a cloud-HPC convergence schedulerabstractSummary Job scheduling is an old topic in High‐Performance Computing (HPC), and it is more and more studied in data centers. Large data centers are often split into separate partitions for cloud computing and HPC; each partition normally has its specific scheduler. The possibility of migrating jobs from the HPC partition to the cloud one is a topic widely discussed in the literature. However, job migration from cloud to HPC is a much less explored topic. Nevertheless, such migration may be useful in many situations, in particular when the HPC platform has a low resource usage level, and the cloud usage level is high. A large number of jobs that could migrate from the cloud to the HPC partition may be observed in Google data center workloads. Job scheduling using overbooking strategy is seen as the main reason for the high resource usage level in clouds. However, overbooking can lead to a high rate of rescheduling and job dumping, which potentially causes response time violations. This work shows that HPC platforms can host and execute some cloud jobs with low interference in HPC jobs and a low number of response time violations. We introduce the definition of a cloud‐HPCconvergence areaand propose a job scheduling strategy for it, aiming at reducing the number of response time violations of cloud jobs without interfering with HPC jobs execution. Our proposal is formally defined and then evaluated in different execution scenarios, using theSimGridsimulation framework, with workload data from production HPC grid. The experimental results show that often, there is a large number of empty areas in the scheduling plan of HPC platforms, which makes it possible to allocate cloud jobs by backfilling. This is due to the sparse HPC job submission pattern and the low resource usage level in some HPC platforms. One performed simulation scenario considered a set of 11K parallel HPC jobs running on a 2560‐processor platform having an average resource usage level of 38.0%. The proposed convergence scheduler succeeded to inject around 267K cloud jobs in the HPC platform, with a response time violation rate under 0.00094% for such jobs, considering 80 processors in theconvergence areaand no effects on the HPC workload. This experiment considered cloud jobs based on job features of Google public cloud workloads, with a processing time slack factor of 1.25 (which is considered as high priority in the Google cloud SLA—Service Level Agreement). Usually, most cloud jobs show a slack factor higher than 1.25 (most cloud jobs are medium or low priority). The same simulation, repeated with a higher slack factor (4), showed no response time violations. Alessandro Kraemer, Carlos Maziero, Olivier Richard, Denis Trystram |
Concurr. Comput. Pract. Exp. | 3 |
| 2017 | Big data and HPC collocation: Using HPC idle resources for Big Data analyticsabstractExecuting Big Data workloads upon High Performance Computing (HPC) infrastractures has become an attractive way to improve their performances. However, the collocation of HPC and Big Data workloads is not an easy task, mainly because of their core concepts' differences. This paper focuses on the challenges related to the scheduling of both Big Data and HPC workloads on the same computing platform. In classic HPC workloads, the rigidity of jobs tends to create holes in the schedule: we can use those idle resources as a dynamic pool for Big Data workloads. We propose a new idea based on Resource and Job Management System's (RJMS) configuration, that makes HPC and Big Data systems to communicate through a simple prolog/epilog mechanism. It leverages the built-in resilience of Big Data frameworks, while minimizing the disturbance on HPC workloads. We present the first study of this approach, using the production RJMS middleware OAR and Hadoop YARN from the HPC and Big Data ecosystems respectively. Our new technique is evaluated with real experiments upon the Grid5000 platform. Our experiments validate our assumptions and show promising results. The system is capable of running an HPC workload with 70% cluster utilization, with a Big Data workload that fills the schedule holes to reach a full 100% utilization. We observe a penalty on the mean waiting time for HPC jobs of less than 17% and a Big Data effectiveness of more than 68% in average. Michael Mercier, David Glesser, Yiannis Georgiou 0002, Olivier Richard |
IEEE BigData | 4 |
| 2016 | Batsim: A Realistic Language-Independent Resources and Jobs Management Systems Simulator
Pierre-François Dutot, Michael Mercier, Millian Poquet, Olivier Richard |
JSSPP | 4 |
| 2015 | A survey of general-purpose experiment management tools for distributed systems
Tomasz Buchert, Cristian Ruiz, Lucas Nussbaum, Olivier Richard |
Future Gener. Comput. Syst. | 4 |
| 2014 | Platform Calibration for Load Balancing of Large Simulations: TLM CaseabstractThe heterogeneous nature of distributed platforms such as computational Grids is one of the main barriers to effectively deploy tightly-coupled applications. For those applications, one common problem that appears due to the hardware heterogeneity is the load imbalance which slows down the application to the pace of the slower processor. One solution is to distribute the load adequately taking into account hardware capacities. To do so, an estimation of the hardware capacities for running the application has to be obtained. In this paper, we present a static load balancing for iterative tightly-coupled applications based on a profile prediction model. This technique is presented as a successful example of the interaction between experiment management tools and parallel applications. The experiment management tool Expo is used that enabled to: (1) provide a general, lightweight and descriptive way to capture the tuning and deployment of a parallel application in a computing infrastructure, (2) perform the tuning of the application efficiently in terms of human effort and resources needed. This paper reports the costs for carrying out the tuning of a large electromagnetic simulation based on TLM for the platform Grid'5000 and the improvements obtained on the total execution time of the application. Cristian Ruiz, Mihai Alexandru, Olivier Richard, Thierry Monteil 0001, Hervé Aubert |
CCGRID | 3 |
| 2013 | Analysis of the Jobs Resource Utilization on a Production System
Joseph Emeras, Cristian Ruiz, Jean-Marc Vincent, Olivier Richard |
JSSPP | 4 |
| 2011 | Large scale P2P discovery middleware demonstrationabstractSpades aims at offering a solution to deal with distributed, volatile and heterogeneous computing resources. The targeted platform is a large one with potentially huge number of resources. In a seamless way, our proposal includes i) an abstraction of the resources in a computing overlay; ii) a P2P distributed resource/service discovery; iii) an user job scheduling workflow; iv) an auto-stabilizing solution and v) optimized algorithms for Petascale architecture. In order to orchestrate Spades platform, we have designed and implemented Sbam middleware. Eddy Caron, Florent Chuffart, Haiwu He, Anissa Lamani, Philippe Le Brouster, Olivier Richard |
Peer-to-Peer Computing | 6 |
| 2010 | DSL-Lab: A Low-Power Lightweight Platform to Experiment on Domestic Broadband InternetabstractThis article presents the design and building of DSL-Lab, a platform to experiment on distributed computing over broadband domestic Internet. Experimental platforms such as PlanetLab and Grid'5000 are promising methodological approaches to study distributed systems. However, both platforms focus on high-end service and network deployments only available on a restricted part of the Internet, leaving aside the possibility for researchers to experiment in conditions close to what is usually available with domestic connection to the Internet. DSL-Lab is a complementary approach to PlanetLab and Grid'5000 to experiment with distributed computing in an environment closer to how Internet appears, when applications are run on end-user PCs. DSL-Lab is a set of 40 low-power and low-noise nodes, which are hosted by participants, using the participants' xDSL or cable access to the Internet. The objective is to provide a validation and experimentation platform for new protocols, services, simulators and emulators for these systems. In this paper, we report on the software design (security, resources allocation, power management) as well as on the first experiments achieved. Gilles Fedak, Jean-Patrick Gelas, Thomas Hérault, Victor Iniesta, Derrick Kondo, Laurent Lefèvre, Paul Malecot, Lucas Nussbaum, Ala Rezmerita, Olivier Richard |
ISPDC | 10 |
| 2009 | TakTuk, adaptive deployment of remote executionsabstractThis article deals with TakTuk, a middleware that deploys efficiently parallel remote executions on large scale grids (thousands of nodes). This tool is mostly intended for interactive use: distributed machines administration and parallel applications development. Thus, it has to minimize the time required to complete the whole deployment process. Benoit Claudel, Guillaume Huard, Olivier Richard |
HPDC | 3 |
| 2009 | The GREEN-NET framework: Energy efficiency in large scale distributed systemsabstractThe question of energy savings has been a matter of concern since a long time in the mobile distributed systems and battery-constrained systems. However, for large-scale non-mobile distributed systems, which nowadays reach impressive sizes, the energy dimension (electrical consumption) just starts to be taken into account. In this paper, we present the GREEN-NET1framework which is based on 3 main components: an ON/OFF model based on an Energy Aware Resource Infrastructure (EARI), an adapted Resource Management System (OAR) for energy efficiency and a trust delegation component to assume network presence of sleeping nodes. Georges Da Costa, Jean-Patrick Gelas, Yiannis Georgiou 0002, Laurent Lefèvre, Anne-Cécile Orgerie, Jean-Marc Pierson, Olivier Richard, K. Sharma |
IPDPS | 7 |
| 2009 | On Robust Covert Channels Inside DNS
Lucas Nussbaum, Pierre Neyron, Olivier Richard |
SEC | 3 |
| 2008 | Lightweight emulation to study peer-to-peer systemsabstractAbstract The current methods used to test and study peer‐to‐peer systems (namely modeling, simulation or execution on real testbeds) often show limits regarding scalability, realism and accuracy. This paper describes and evaluates P2PLab, our framework to study peer‐to‐peer systems by combining emulation (use of the real studied application within a configured synthetic environment) and virtualization. P2PLab is scalable (it uses a distributed network model) and has good virtualization characteristics (many virtual nodes can be executed on the same physical node by using process‐level virtualization). Experiments with the BitTorrent file‐sharing system complete this article and demonstrate the usefulness of this platform. Copyright © 2007 John Wiley & Sons, Ltd. Lucas Nussbaum, Olivier Richard |
Concurr. Comput. Pract. Exp. | 2 |
| 2007 | Evaluations of the Lightweight Grid CIGRI upon the Grid5000 PlatformabstractA widely used method for large scale experiment execution upon P2P or cluster computing platforms, is the exploitation of idle resources. Specifically in the case of clusters, administrators share their cluster's idle cycles into computational grids, for the execution of the so called bag-of-tasks applications. Fault-tolerance and scheduling are some of the important challenges that have arisen on the specific research field. On this paper, we present a simple, scalable, fault tolerant and user-transparent approach of harnessing idle cluster resources for executing grid bag-of-tasks applications. Our main interest lies on the large-scale deployment and evaluation of our lightweight grid computing approach, under real-life parameters. Under this context, we experiment with CIGRI and a fully transparent system-level checkpointing feature for scheduling and turnaround-time optimisation. We discuss the value of experimentation on computer science, we propose reproducible experiments based on real workload traces and we explain how our experimental methodology contributes on the development and evaluation of our grid platform. Yiannis Georgiou 0002, Olivier Richard, Nicolas Capit |
eScience | 2 |
| 2006 | Resources availability for Peer to Peer systemsabstractNowadays, peer to peer systems are largely studied. But in order to evaluate them in a realistic way, a better knowledge of their environments is needed. In this article we focus on the computers availability in these systems. We characterize this availability behind ADSL lines and we link it with the availability of peer to peer systems participants. We emphasise on the methodology as generalized in other systems such as grids or ad-hoc systems. We finally show how users of ADSL lines are related to peer to peer users and we give some examples of the possible practical use of theses results. The results are based on trace datasets obtained over the first five month of 2003 with around 5000 hosts. Georges Da Costa, Corine Marchand, Olivier Richard, Jean-Marc Vincent |
AINA (1) | 3 |
| 2006 | A tool for environment deployment in clusters and light gridsabstractFocused around the field of the exploitation and the administration of high performance large-scale parallel systems, this article describes the work carried out on the deployment of environment on high computing clusters and grids. We initially present the problems involved in the installation of an environment (OS, middleware, libraries, applications...) on a cluster or grid and how an effective deployment tool, Kadeploy2, can become a new form of exploitation of this type of infrastructures. We present the tool's design choices, its architecture and we describe the various stages of the deployment method, introduced by Kadeploy2. Moreover, we propose methods on the one hand, for the improvement of the deployment time of a new environment; and in addition, for the support of various operating systems. Finally, to validate our approach we present tests and evaluations realized on various clusters of the experimental grid Grid5000. Yiannis Georgiou 0002, Julien Leduc, Brice Videau, Johann Peyrard, Olivier Richard |
IPDPS | 5 |
| 2006 | Lightweight emulation to study peer-to-peer systemsabstractThe current methods used to test and study peer-to-peer systems (namely modeling, simulation, or execution on real testbeds) often show limits regarding scalability, realism and accuracy. This paper describes and evaluates P2PLab, our framework to study peer-to-peer systems by combining emulation (use of the real studied application within a configured synthetic environment) and visualization. P2PLab is scalable (it uses a distributed network model) and has good visualization characteristics (many virtual nodes can be executed on the same physical node by using process-level visualization). Experiments with the BitTorrent file-sharing system complete this paper and demonstrate the usefulness of this platform. Lucas Nussbaum, Olivier Richard |
IPDPS | 2 |
| 2005 | A batch scheduler with high level componentsabstractIn this article we present the design choices and the evaluation of a batch scheduler for large clusters, named OAR. This batch scheduler is based upon an original design that emphasizes on low software complexity by using high level tools. The global architecture is built upon the scripting language Perl and the relational database engine Mysql. The goal of the project OAR is to prove that it is possible today to build a complex system for resource management using such tools without sacrificing efficiency and scalability. Currently, our system offers most of the important features implemented by other batch schedulers such as priority scheduling (by queues), reservations, backfilling and some global computing support. Despite the use of high level tools, our experiments show that our system has performances close to other systems. Furthermore, OAR is currently exploited for the management of 700 nodes (a metropolitan grid) and has shown good efficiency and robustness. Nicolas Capit, Georges Da Costa, Yiannis Georgiou 0002, Guillaume Huard, Cyrille Martin 0002, Grégory Mounié, Pierre Neyron, Olivier Richard |
CCGRID | 8 |
| 2004 | DRAC: Adaptive Control System with Hardware Performance Counters
Maurício Aronne Pillon, Olivier Richard, Georges Da Costa |
Euro-Par | 2 |
| 2002 | The Virtual Cluster: A Dynamic Environment for Exploitation of Idle Network ResourcesabstractStandard environments for exploiting idle time of workstations are based on some kind of spying process that detects low CPU usage and informs to a scheduler so that work can be dispatched. This approach generates local interference and, since the same local environment is used, could lead to security problems. We are investigating the exploitation of idle times in network resources based on a complete mode change in a candidate node. After the detection that some node is idle, a mode-switcher boots a new operating system that will work over a separate disk partition. After the boot phase the node is linked to a logical network topology and is available to receive jobs. Users can allocate nodes from this virtual cluster through a standard frontend as they would do in a "conventional" cluster. Because nodes may leave and join this virtual machine we use a distributed processor management to allow user applications to cope with this dynamic resource behavior. In this paper we describe the architecture of the virtual cluster and present the results obtained with a mode-switcher and a prototype application under real use conditions. César A. F. De Rose, Franco Blanco, Nicolas Maillard, Katia Barbosa Saikoski, Reynaldo Novaes, Olivier Richard, Bruno Richard 0002 |
SBAC-PAD | 6 |
| 2001 | Understanding performance of SMP clusters running MPI programs
Franck Cappello, Olivier Richard, Daniel Etiemble |
Future Gener. Comput. Syst. | 2 |
| 2000 | Investigating the Performance of Two Programming Models for Clusters of SMP PCsabstractMultiprocessors and high performance networks allow building CLUsters of MultiProcessors (CLUMPs). One distinctive feature over traditional parallel computers is their hybrid memory model (message passing between the nodes and shared memory inside the nodes). We evaluate the performance of a cluster of dual processor PCs connected by a Myrinet network for NAS benchmarks using two programming models: a Single Memory Model based on the MPICH-PM/CLUMP library of the RWCP and a Hybrid Memory Model using MPICH-PM and OpenMP. MPI programs are used as the reference in all experiments involving programming models. We compare dual processor node configurations speedup versus uniprocessor node configurations for each model. We demonstrate that the superiority of one model over the other depends on the features of the applications. In particular, we detail the speedup results from breakdowns of the benchmark execution times and from measurements of hardware counters. Franck Cappello, Olivier Richard, Daniel Etiemble |
HPCA | 2 |
| 1999 | A Client/Broker/Server Substrate with µs Round-Trip Overhead
Olivier Richard, Franck Cappello |
Euro-Par | 1 |
| 1998 | On the Self-Similar Nature of Workstations and WWW Servers Workload
Olivier Richard, Franck Cappello |
Euro-Par | 1 |