EDBT 2026 Demo / reviewers in the wild / expert
Olivier Aumage
dblp:86/706
· DBLP profile ↗
24ranked-venue papers
9as first author
6since 2021 · last 2025
0000-0002-5406-8743ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 9 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Task-Based HPC in the Cloud: Price-Performance Analysis of N-Body Simulations with StarPUabstractPublic cloud environments present significant challenges for traditional High Performance Computing (HPC) applications due to infrastructure limitations that differ substantially from dedicated HPC systems. Unlike traditional HPC clusters optimized for tightly coupled parallel workloads, cloud platforms were designed primarily for web services and data processing applications. Key obstacles include high-latency networks, hardware virtualization overhead, and limited availability of specialized accelerators, all of which can severely impact the performance of compute-intensive applications such as physics simulations. This study investigates the feasibility of running HPC workloads on public cloud infrastructure using standard and cost-effective instance configurations rather than expensive specialized "HPC" offerings. We deploy heterogeneous clusters on Amazon Web Services using the HPC@Cloud Toolkit, incorporating various instance types, including GPU-accelerated nodes with different computational capabilities. Our evaluation focuses on N-body simulations implemented using a task-based parallel programming model, leveraging the StarPU runtime system to dynamically schedule computational tasks across various processing units. Our experimental results demonstrate three key findings: (1) smaller GPU-equipped instances (g6.2xlarge) achieve performance comparable to larger instances while costing approximately one-sixth the price, challenging conventional scaling assumptions for cloud-based HPC; (2) strategic GPU utilization yields up to 8.2× performance improvements over CPU-only configurations while reducing total execution costs by 24.4×; and (3) while task-based programming models effectively address network limitations through dynamic scheduling, complex tree-based algorithms like TBFMM face significant optimization challenges in cloud environments due to load balancing issues and expensive parameter tuning requirements. These findings provide practical guidance for researchers and practitioners seeking cost-effective cloud HPC deployments, demonstrating that commodity cloud infrastructures can be viable for regular computational workloads but require careful algorithmic-resource matching for optimal efficiency. Nicolas Vanz, Vanderlei Munhoz, Márcio Castro 0001, Laércio Lima Pilla, Olivier Aumage |
IC2E | 5 |
| 2025 | Compiler, Runtime, and Hardware Parameters Design Space ExplorationabstractHPC systems are increasingly complex with many tunable parameters impacting applications' metrics-e.g., performance, energy consumption. The main challenges of these systems are finding the appropriate configuration per application on any given system and understanding how the configurations affect applications' metrics on a system. Both can be addressed with design space exploration (DSE). However, exploring all the configurations available is costly due to the long execution and setup times of these executions. Indeed, it requires instrumenting the applications to collect data, compiling them with different options and setting the parameters for each execution. DSE algorithms can greatly reduce the exploration time by guiding which configuration to execute next to reach the objective without evaluating all the configurations. A DSE study thus requires implementing an exploration algorithm and automating parameters setting, application instrumentation and compilation, and metrics collection. This represents a huge overhead to the actual study, yet most DSE studies still do it from scratch. To alleviate the setup cost, we propose a unified methodology to perform the exploration and implement it in the CORHPEX framework to setup configurations with compiler, runtime, and hardware parameters, efficiently and flexibly. The framework enables choosing the exploration strategy, the design space to study, the applications to execute and the metrics to collect independently while involving little coding overhead. It is extensible with custom exploration algorithms and data readers. We demonstrate the versatility and robustness of our framework on parallel codes, including NAS, Rodinia, LULESH benchmarks, and real-world applications, on two systems exposing different parameters with various DSE techniques and goals. We show that working with CORHPEX enables getting insights on code optimization strategies by using exploration algorithms that can speedup the execution by a factor of 10X while preserving 95% the possible gains. Finally, we demonstrate the framework's potential for more advanced studies by training surrogate models of complex HPC applications achieving over 93% accuracy. Lana Scravaglieri, Ani Anciaux-Sedrakian, Olivier Aumage, Thomas Guignon, Mihail Popov |
IPDPS | 3 |
| 2025 | Optimal scheduling algorithms for software-defined radio pipelined and replicated task chains on multicore architecturesabstractSoftware-Defined Radio (SDR) represents a move from dedicated hardware to software implementations of digital communication standards. This approach offers flexibility, shorter time to market, maintainability , and lower costs, but it requires an optimized distribution tasks in order to meet performance requirements. Thus, we study the problem of scheduling SDR linear task chains of stateless and stateful tasks for streaming processing. We model this problem as a pipelined workflow scheduling problem based on pipelined and replicated parallelism on homogeneous resources. We propose an optimal dynamic programming solution and an optimal greedy algorithm named OTAC for maximizing throughput while also minimizing resource utilization . Moreover, the optimality of the proposed scheduling algorithm is proved. We evaluate our solutions and compare their execution times and schedules to other algorithms using synthetic task chains and an implementation of the DVB-S2 communication standard on the AFF3CT SDR Domain Specific Language . Our results demonstrate how OTAC quickly finds optimal schedules, leading consistently to better results than other algorithms, or equivalent results with fewer resources. Diane Orhan, Laércio Lima Pilla, Denis Barthou, Adrien Cassagne, Olivier Aumage, Romain Tajan, Christophe Jégo, Camille Leroux |
J. Parallel Distributed Comput. | 5 |
| 2025 | Performance portability of generated cardiac simulation kernels through automatic dimensioning and load balancing on heterogeneous nodes
Vincent Alba, Olivier Aumage, Denis Barthou, Marie Christine Counilh, Amina Guermouche |
J. Supercomput. | 2 |
| 2023 | A DSEL for high throughput and low latency software-defined radio on multicore CPUsabstractSummary This article presents a new Domain Specific Embedded Language (DSEL) dedicated to Software‐Defined Radio (SDR). From a set of carefully designed components, it enables to build efficient software digital communication systems, able to take advantage of the parallelism of modern processor architectures, in a straightforward and safe manner for the programmer. In particular, proposed DSEL enables the combination of pipelining and sequence duplication techniques to extract both temporal and spatial parallelism from digital communication systems. We leverage the DSEL capabilities on a real use case: a fully digital transceiver for the widely used DVB‐S2 standard designed entirely in software. Through evaluation, we show how proposed software DVB‐S2 transceiver is able to get the most from modern, high‐end multicore CPU targets. Adrien Cassagne, Romain Tajan, Olivier Aumage, Camille Leroux, Denis Barthou, Christophe Jégo |
Concurr. Comput. Pract. Exp. | 3 |
| 2023 | Optimizing performance and energy across problem sizes through a search space exploration and machine learning
Lana Scravaglieri, Mihail Popov, Laércio Lima Pilla, Amina Guermouche, Olivier Aumage, Emmanuelle Saillard |
J. Parallel Distributed Comput. | 5 |
| 2020 | InKS: a programming model to decouple algorithm from optimization in HPC codes
Ksander Ejjaaouani, Olivier Aumage, Julien Bigot, Michel Mehrenberger, Hitoshi Murai, Masahiro Nakao, Mitsuhisa Sato |
J. Supercomput. | 2 |
| 2017 | Combining Both a Component Model and a Task-based Model for HPC Applications: a Feasibility Study on GyselaabstractThis paper studies the feasibility of efficiently combining both a software component model and a task-based model. Task based models are known to enable efficient executions on recent HPC computing nodes while component models ease the separation of concerns of application and thus improve their modularity and adaptability. This paper describes a prototype version of the COMET programming model combining concepts of task-based and component models, and a preliminary version of the COMET runtime built on top of StarPU and L2C. Evaluations of the approach have been conducted on a real-world use-case analysis of a subpart of the production application GYSELA. Results show that the approach is feasible and that it enables easy composition of independent software codes without introducing overheads. Performance results are equivalent to those obtained with a plain OpenMP based implementation. Olivier Aumage, Julien Bigot, Hélène Coullon, Christian Pérez, Jérôme Richard |
CCGrid | 1 |
| 2017 | Rewriting System for Profile-Guided Data Layout Transformations on Binaries
Christopher Haine, Olivier Aumage, Denis Barthou |
Euro-Par | 2 |
| 2017 | Bridging the Gap Between OpenMP and Task-Based Runtime Systems for the Fast Multipole MethodabstractWith the advent of complex modern architectures, the low-level paradigms long considered sufficient to build High Performance Computing (HPC) numerical codes have met their limits. Achieving efficiency, ensuring portability, while preserving programming tractability on such hardware prompted the HPC community to design new, higher level paradigms while relying on runtime systems to maintain performance. However, the common weakness of these projects is to deeply tie applications to specific expert-only runtime system APIs. The OpenMP specification, which aims at providing common parallel programming means for shared-memory platforms, appears as a good candidate to address this issue thanks to the latest task-based constructs introduced in its revision 4.0. The goal of this paper is to assess the effectiveness and limits of this support for designing a high-performance numerical library, ScalFMM, implementing the fast multipole method (FMM) that we have deeply re-designed with respect to the most advanced features provided by OpenMP 4. We show that OpenMP 4 allows for significant performance improvements over previous OpenMP revisions on recent multicore processors and that extensions to the 4.0 standard allow for strongly improving the performance, bridging the gap with the very high performance that was so far reserved to expert-only runtime system APIs. Emmanuel Agullo, Olivier Aumage, Bérenger Bramas, Olivier Coulaud, Samuel Pitoiset |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2015 | Towards Seismic Wave Modeling on Heterogeneous Many-Core Architectures Using Task-Based Runtime SystemabstractUnderstanding three-dimensional seismic wave propagation in complex media is still one of the main challenges of quantitative seismology. Because of its simplicity and numerical efficiency, the finite-differences method is one of the standard techniques implemented to consider the elastodynamics equation. Additionally, this class of modeling heavily relies on parallel architectures in order to tackle large scale geometries including a detailed description of the physics. Last decade, significant efforts have been devoted towards efficient implementation of the finite-differences methods on emerging architectures. These contributions have demonstrated their efficiency leading to robust industrial applications. The growing representation of heterogeneous architectures combining general purpose multicore platforms and accelerators leads to re-design current parallel application. In this paper, we consider Star PU task-based runtime system in order to harness the power of heterogeneous CPU+GPU computing nodes. We detail our implementation and compare the performance obtained with the classical CPU or GPU only versions. Preliminary results demonstrate significant speedups in comparison with the best implementation suitable for homogeneous cores. Víctor Martínez, David Michéa, Fabrice Dupros, Olivier Aumage, Samuel Thibault, Hideo Aochi, Philippe Olivier Alexandre Navaux |
SBAC-PAD | 4 |
| 2013 | Adaptive Task Size Control on High Level Programming for GPU/CPU Work Sharing
Tetsuya Odajima, Taisuke Boku, Mitsuhisa Sato, Toshihiro Hanawa, Yuetsu Kodama, Raymond Namyst, Samuel Thibault, Olivier Aumage |
ICA3PP (2) | 8 |
| 2012 | StarPU-MPI: Task Programming over Clusters of Machines Enhanced with Accelerators
Cédric Augonnet, Olivier Aumage, Nathalie Furmento, Raymond Namyst, Samuel Thibault |
EuroMPI | 2 |
| 2010 | Structuring the execution of OpenMP applications for multicore architecturesabstractThe now commonplace multi-core chips have introduced, by design, a deep hierarchy of memory and cache banks within parallel computers as a tradeoff between the user friendliness of shared memory on the one side, and memory access scalability and efficiency on the other side. However, to get high performance out of such machines requires a dynamic mapping of application tasks and data onto the underlying architecture. Moreover, depending on the application behavior, this mapping should favor cache affinity, memory bandwidth, computation synchrony, or a combination of these. The great challenge is then to perform this hardware-dependent mapping in a portable, abstract way. To meet this need, we propose a new, hierarchical approach to the execution of OpenMP threads onto multicore machines. Our ForestGOMP runtime system dynamically generates structured trees out of OpenMP programs. It collects relationship information about threads and data as well. This information is used together with scheduling hints and hardware counter feedback by the scheduler to select the most appropriate threads and data distribution. ForestGOMP features a highlevel platform for developing and tuning portable threads schedulers. We present several applications for which we developed specific scheduling policies that achieve excellent speedups on 16-core machines. François Broquedis, Olivier Aumage, Brice Goglin, Samuel Thibault, Pierre-André Wacrenier, Raymond Namyst |
IPDPS | 2 |
| 2007 | NEW MADELEINE: a Fast Communication Scheduling Engine for High Performance NetworksabstractCommunication libraries have dramatically made progress over the fifteen years, pushed by the success of cluster architectures as the preferred platform for high performance distributed computing. However, many potential optimizations are left unexplored in the process of mapping application communication requests onto low level network commands. The fundamental cause of this situation is that the design of communication subsystems is mostly focused on reducing the latency by shortening the critical path. In this paper, we present a new communication scheduling engine which dynamically optimizes application requests in accordance with the NICs capabilities and activity. The optimizing code is generic and portable. The database of optimizing strategies may be dynamically extended. Olivier Aumage, Elisabeth Brunet, Nathalie Furmento, Raymond Namyst |
IPDPS | 1 |
| 2007 | High-Performance Multi-Rail Support with the NEWMADELEINE Communication LibraryabstractThis paper focuses on message transfers across multiple heterogeneous high-performance networks in the NEWMADELEINE communication library. NEWMADELEINE features a modular design that allows the user to easily implement load-balancing strategies efficiently exploiting the underlying network but without being aware of the low-level interface. Several strategies are studied and preliminary results are given. They show that performance of network transfers can be improved by using carefully designed strategies that take into account NIC activity. Olivier Aumage, Elisabeth Brunet, Guillaume Mercier, Raymond Namyst |
IPDPS | 1 |
| 2006 | Short Paper : Dynamic Optimization of Communications over High Speed NetworksabstractWe present a new communication subsystem for high speed networks featuring an extendable packet optimization engine mixing several communication flows. Optimizations are parameterized by the capabilities of the underlying network drivers, and are triggered by the network cards when they become idle. The database of predefined strategies can be easily extended Elisabeth Brunet, Olivier Aumage, Raymond Namyst |
HPDC | 2 |
| 2005 | NETIBIS: an efficient and dynamic communication system for heterogeneous gridsabstractGrids are more heterogeneous and dynamic than traditional parallel or distributed systems, both in terms of processors and of interconnects. A grid communication system must handle many issues: first, it must run on networks that are not yet determined when the application is launched, including user-space interconnects; second, it must transparently run on different networks at the same time; third, it should yield performance close to that of specialized communication systems. In this paper, we present NETIBIS, a new Java communication system that provides a uniform interface for any underlying inter-cluster or intracluster network. NETIBIS solves the heterogeneity issues posed by grid computing by dynamically constructing network protocol stacks out of drivers, self-contained building blocks for flexible configuration, with limited functionality per driver. We describe the design and implementation of the major NETIBIS drivers for serialization, multicast, reliability, and various underlying networks. We also describe various optimizations for performance, like layer collapsing for the GMdriver. We evaluate the performance of NETIBIS on several platforms, including a European grid. Olivier Aumage, Rutger F. H. Hofman, Henri E. Bal |
CCGRID | 1 |
| 2004 | Wide-Area Communication for Grids: An Integrated Solution to Connectivity, Performance and Security Problems
Alexandre Denis 0001, Olivier Aumage, Rutger F. H. Hofman, Kees Verstoep, Thilo Kielmann, Henri E. Bal |
HPDC | 2 |
| 2003 | MPICH/MADIII: a Cluster of Clusters Enabled MPI ImplementationabstractThis paper presents an MPI implementation that allows an easy and efficient use of the interconnection of several clusters, of potentially heterogeneous nature (as far as the network is concerned). We describe the underlying communication subsystem used. The mechanisms within MPI that inform the user of the underlying topological structure are detailed The performance figures obtained with this MPI implementation are discussed and advocate for the use of such a solution on this particular type of architecture. Olivier Aumage, Guillaume Mercier |
CCGRID | 1 |
| 2002 | Madeleine II: a portable and efficient communication library for high-performance cluster computing
Olivier Aumage, Luc Bougé, Jean-François Méhaut, Raymond Namyst |
Parallel Comput. | 1 |
| 2001 | Efficient Inter-Device Data-Forwarding in the Madeleine Communication LibraryabstractInterconnecting multiple clusters with a high speed network to form a single heterogeneous architecture (i.e. a cluster of clusters) is currently a hot issue. Consequently, new runtime systems that are able to simultaneously deal with multiple high speed networks within the same application have to be designed. This paper presents how we did extend an existing multi-device communication library with fast internal data-forwarding capabilities on gateway nodes. On top of that, efficient high-level routing mechanisms can be implemented. Our approach is easily applicable to many network protocols and is completely independant from the application code. Efficiency is achieved by avoiding extra data copies when possible and by using pipelining techniques. The preliminary experiments show that the observed inter-cluster bandwidth is close to the one that can be delivered by the hardware. Olivier Aumage, Lionel Eyraud-Dubois, Raymond Namyst |
IPDPS | 1 |
| 2001 | MPICH/Madeleine: a True Multi-Protocol MPI for High Performance NetworksabstractThis paper introduces a version of MPICH handling efficiently different networks simultaneously. The core of the implementation relies on a device called ch-mad which is based on a generic multiprotocol communication library called Madeleine. The performance achieved with tested networks such as Fast-Ethernet, Scalable Coherent Interface or Myrinet is very good. Indeed, this multi-protocol version of MPICH generally outperforms other free or commercial implementations of MPI. Olivier Aumage, Guillaume Mercier, Raymond Namyst |
IPDPS | 1 |
| 2000 | Madeleine II: a Portable and Efficient Communication Library for High-Performance Cluster ComputingabstractThis paper introduces Madeleine II, a new adaptive and portable multi-protocol implementation of the Madeleine communication library. Madeleine II has the ability to control multiple network interfaces (BIP, SISCI, VIA) and multiple network adapters (Ethernet, Myrinet, SCI) within the same application session. Moreover it includes advanced mechanisms to dynamically select the most appropriate transfer method for a given network protocol according to various parameters such as data size or responsiveness user requirements. We report on performance measurements obtained using BIP/Myrinet and SISCI/SCI and we present preliminary results about our Nexus/Madeleine II and MPICH/Madeleine II ports. Olivier Aumage, Luc Bougé, Alexandre Denis 0001, Jean-François Méhaut, Guillaume Mercier, Raymond Namyst, Loïc Prylli |
CLUSTER | 1 |