EDBT 2026 Demo / reviewers in the wild / expert
Amina Guermouche
dblp:72/9776
· DBLP profile ↗
15ranked-venue papers
3as first author
6since 2021 · last 2025
0000-0003-3653-3784ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 3 first-author · 5 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Fine-Grain Energy Consumption Modeling of HPC Task-Based ProgramsabstractThe power consumption of supercomputers is and will be a major concern in the future. Therefore, reducing the power consumption of high performance computing (HPC) applications is mandatory. Monitoring the energy consumption of HPC programs is a good first step: using external or software power meters, one can measure the energy consumption of an entire compute node or some of its hardware components. Unfortunately, the differences in scope and time scale between power meters and code level functions prevent the identification of power hungry code blocks. For this work, we propose leveraging the tracing mechanism of the StarPU runtime system in order to estimate task level power consumption. We trace the execution of the application while regularly measuring coarse-grain energy consumption of central processing units (CPUs) and graphics processing units (GPUs) using vendor software interfaces. After execution, we identify the executed tasks on each processing unit for every coarsegrain energy measurement interval. We then use this information to generate an overdetermined linear system linking tasks and energy measurements. Subsequently, solving the system allows us to estimate the fine-grain power consumption of each task independently of its actual duration. We achieve mean average percentage errors (MAPE) ranging from 0.5 % to 5 % on various CPUs, and from 10 % to 28 % on GPUs. We show that a solution generated from a run can be used to predict the energy consumption of other runs with different scheduling policies. Jules Risse, Amina Guermouche, François Trahay |
CLUSTER | 2 |
| 2025 | Performance portability of generated cardiac simulation kernels through automatic dimensioning and load balancing on heterogeneous nodes
Vincent Alba, Olivier Aumage, Denis Barthou, Marie Christine Counilh, Amina Guermouche |
J. Supercomput. | 5 |
| 2023 | Optimizing performance and energy across problem sizes through a search space exploration and machine learning
Lana Scravaglieri, Mihail Popov, Laércio Lima Pilla, Amina Guermouche, Olivier Aumage, Emmanuelle Saillard |
J. Parallel Distributed Comput. | 4 |
| 2022 | Study of the Processor and Memory Power and Energy Consumption of Coupled Sparse/Dense SolversabstractIn the aeronautical industry, aeroacoustics is used to model the propagation of acoustic waves in air flows enveloping an aircraft in flight. This for instance allows one to simulate the noise produced at ground level by an aircraft during the takeoff and landing phases, in order to validate that the regulatory environmental standards are met. Unlike most other complex physics simulations, the method resorts to solving coupled sparse/dense systems. In a previous work, the authors proposed two classes of algorithms for solving such large systems on a relatively small multi-core workstation based on compression techniques and showed their positive impact on time to solution and memory usage. Within a multi-metric experimental study presented in this paper, we establish that this translates to the power and energy consumption as well. Moreover, because of the nature of the problem, coupling dense and sparse matrices, and the underlying solution methods, including dense, sparse direct and compression steps, the experiments yield an interesting processor and memory power profile to share with the community and which we analyze in details. Emmanuel Agullo, Marek Felsöci, Amina Guermouche, Hervé Mathieu, Guillaume Sylvand, Bastien Tagliaro |
SBAC-PAD | 3 |
| 2022 | duf: Dynamic uncore frequency scaling to reduce power consumptionabstractSummary Reducing the power consumption of applications has become one of the key challenges in high‐performance computing. Recent processor architectures differentiate processor core frequency from its uncore frequency. As a consequence, in addition to tuning processor core frequency with dynamic voltage and frequency scaling, power consumption can also be controlled through uncore frequency scaling. This article studies how the uncore frequency can be used as a leverage to improve power consumption. We propose duf, a daemon process that dynamically adapts the uncore frequency to reduce an application power consumption with a user‐defined limit on performance degradation. The evaluation of duf on three different architectures shows that with no performance degradation (<0.6%), duf can reduce socket power consumption by 7.94%. We also show that duf is able to reduce the total energy consumption by up to 18.20%. Étienne André 0002, Rémi Dulong, Amina Guermouche, François Trahay |
Concurr. Comput. Pract. Exp. | 3 |
| 2022 | Thermal design power and vectorized instructions behaviorabstractAbstract Vectorized instructions were introduced to improve the performance of applications. However, they come at the cost of an increase in the power consumption. As a consequence, processors are designed to limit their frequency when such instructions are used in order to respect the thermal design power limit. In this paper, we study and compare the impact of thermal design power and Simple Instruction Multiple Data (SIMD) instructions on performance, power and energy consumption of processors and memory. The study is performed on three different architectures providing different characteristics and four applications with different profiles (including one application with different phases, each phase having a different profile). The study shows that, because of processor frequency, performance and power consumption are strongly related to thermal design power. It also shows that AVX512 has unexpected behavior regarding processor power consumption, while DRAM power consumption is impacted by SIMD instructions because of the generated memory throughput. Finally, this paper tackles the impact of turboboost which shows equivalent to better performance for all the studied cases while not always decreasing energy consumption. Amina Guermouche, Anne-Cécile Orgerie |
Concurr. Comput. Pract. Exp. | 1 |
| 2019 | Using Differential Execution Analysis to Identify Thread InterferenceabstractUnderstanding the performance of a multi-threaded application is difficult. The threads interfere when they access the same shared resource, which slows down their execution. Unfortunately, current profiling tools report the hardware components or the synchronization primitives that saturate, but they cannot tell if the saturation is the cause of a performance bottleneck. In this paper, we propose a holistic metric able to pinpoint the blocks of code that suffer interference the most, regardless of the interference cause. Our metric uses performance variation as a universal indicator of interference problems. With an evaluation of 27 applications we show that our metric can identify interference problems caused by six different kinds of interference in nine applications. We are able to easily remove seven of the bottlenecks, which leads to a performance improvement of up to nine times. Mohamed Said Mosli Bouksiaa, François Trahay, Alexis Lescouet, Gauthier Voron, Rémi Dulong, Amina Guermouche, Elisabeth Brunet, Gaël Thomas 0001 |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2016 | Failure detection and propagation in HPC systemsabstractBuilding an infrastructure for Exascale applications requires, in addition to many other key components, a stable and efficient failure detector. This paper describes the design and evaluation of a robust failure detector, able to maintain and distribute the correct list of alive resources within proven and scalable bounds. The detection and distribution of the fault information follow different overlay topologies that together guarantee minimal disturbance to the applications. A virtual observation ring minimizes the overhead by allowing each node to be observed by another single node, providing an unobtrusive behavior. The propagation stage is using a non-uniform variant of a reliable broadcast over a circulant graph overlay network, and guarantees a logarithmic fault propagation. Extensive simulations, together with experiments on the Titan ORNL supercomputer, show that the algorithm performs extremely well, and exhibits all the desired properties of an Exascale-ready algorithm. George Bosilca, Aurelien Bouteiller, Amina Guermouche, Thomas Hérault, Yves Robert, Pierre Sens 0001, Jack J. Dongarra |
SC | 3 |
| 2014 | Unified model for assessing checkpointing protocols at extreme-scaleabstractSUMMARY In this paper, we present a unified model for several well‐known checkpoint/restart protocols. The proposed model is generic enough to encompass both extremes of the checkpoint/restart space, from coordinated approaches to a variety of uncoordinated checkpoint strategies (with message logging). We identify a set of crucial parameters, instantiate them, and compare the expected efficiency of the fault tolerant protocols, for a given application/platform pair. We then propose a detailed analysis of several scenarios, including some of the most powerful currently available high performance computing platforms, as well as anticipated Exascale designs. The results of this analytical comparison are corroborated by a comprehensive set of simulations. Altogether, they outline comparative behaviors of checkpoint strategies at very large scale, thereby providing insight that is hardly accessible to direct experimentation. Copyright © 2013 John Wiley & Sons, Ltd. George Bosilca, Aurelien Bouteiller, Elisabeth Brunet, Franck Cappello, Jack J. Dongarra, Amina Guermouche, Thomas Hérault, Yves Robert, Frédéric Vivien, Dounia Zaidouni |
Concurr. Comput. Pract. Exp. | 6 |
| 2013 | Multi-criteria Checkpointing Strategies: Response-Time versus Resource Utilization
Aurelien Bouteiller, Franck Cappello, Jack J. Dongarra, Amina Guermouche, Thomas Hérault, Yves Robert |
Euro-Par | 4 |
| 2013 | SPBC: leveraging the characteristics of MPI HPC applications for scalable checkpointingabstractThe high failure rate expected for future supercomputers requires the design of new fault tolerant solutions. Most checkpointing protocols are designed to work with any message-passing application but suffer from scalability issues at extreme scale. We take a different approach: We identify a property common to many HPC applications, namely channel-determinism, and introduce a new partial order relation, called always-happens-before relation, between events of such applications. Leveraging these two concepts, we design a protocol that combines an unprecedented set of features. Our protocol called SPBC combines in a hierarchical way coordinated checkpointing and message logging. It is the first protocol that provides failure containment without logging any information reliably apart from process checkpoints, and this, without penalizing recovery performance. Experiments run with a representative set of HPC workloads demonstrate a good performance of our protocol during both, failure-free execution and recovery. Thomas Ropars, Tatiana V. Martsinkevich, Amina Guermouche, André Schiper, Franck Cappello |
SC | 3 |
| 2012 | HydEE: Failure Containment without Event Logging for Large Scale Send-Deterministic MPI ApplicationsabstractHigh performance computing will probably reach exascale in this decade. At this scale, mean time between failures is expected to be a few hours. Existing fault tolerant protocols for message passing applications will not be efficient anymore since they either require a global restart after a failure (check pointing protocols) or result in huge memory occupation (message logging). Hybrid fault tolerant protocols overcome these limits by dividing applications processes into clusters and applying a different protocol within and between clusters. Combining coordinated check pointing inside the clusters and message logging for the inter-cluster messages allows confining the consequences of a failure to a single cluster, while logging only a subset of the messages. However, in existing hybrid protocols, event logging is required for all application messages to ensure a correct execution after a failure. This can significantly impair failure free performance. In this paper, we propose HydEE, a hybrid rollback-recovery protocol for send-deterministic message passing applications, that provides failure containment without logging any event, and only a subset of the application messages. We prove that HydEE can handle multiple concurrent failures by relying on the send-deterministic execution model. Experimental evaluations of our implementation of HydEE in the MPICH2 library show that it introduces almost no overhead on failure free execution. Amina Guermouche, Thomas Ropars, Marc Snir, Franck Cappello |
IPDPS | 1 |
| 2011 | On the Use of Cluster-Based Partial Message Logging to Improve Fault Tolerance for MPI HPC Applications
Thomas Ropars, Amina Guermouche, Bora Uçar, Esteban Meneses, Laxmikant V. Kalé, Franck Cappello |
Euro-Par (1) | 2 |
| 2011 | Uncoordinated Checkpointing Without Domino Effect for Send-Deterministic MPI ApplicationsabstractAs reported by many recent studies, the mean time between failures of future post-petascale supercomputers is likely to reduce, compared to the current situation. The most popular fault tolerance approach for MPI applications on HPC Platforms relies on coordinated check pointing which raises two major issues: a) global restart wastes energy since all processes are forced to rollback even in the case of a single failure, b) checkpoint coordination may slow down the application execution because of congestions on I/O resources. Alternative approaches based on uncoordinated check pointing and message logging require logging all messages, imposing a high memory/storage occupation and a significant overhead on communications. It has recently been observed that many MPI HPC applications are send-deterministic, allowing to design new fault tolerance protocols. In this paper, we propose an uncoordinated check pointing protocol for send-deterministic MPI HPC applications that (i) logs only a subset of the application messages and (ii) does not require to restart systematically all processes when a failure occurs. We first describe our protocol and prove its correctness. Through experimental evaluations, we show that its implementation in MPICH2 has a negligible overhead on application performance. Then we perform a quantitative evaluation of the properties of our protocol using the NAS Benchmarks. Using a clustering approach, we demonstrate that this protocol actually succeeds to combine the two expected properties: a) it logs only a small fraction of the messages and b) it reduces by a factor approaching 2 the average number of processes to rollback compared to coordinated check pointing. Amina Guermouche, Thomas Ropars, Elisabeth Brunet, Marc Snir, Franck Cappello |
IPDPS | 1 |
| 2010 | On Communication Determinism in Parallel HPC ApplicationsabstractCurrent fault tolerant protocols for high performance computing parallel applications have two major drawbacks: either they require to restart all processes even in the case of only a single process failure or they have a high performance overhead in fault free situation. As a consequence none of existing generic fault tolerant protocols matches needs of HPC applications and surprisingly, there is no fault tolerant protocol dedicated to them. One way to design better fault tolerant protocols for HPC applications is to explore and take advantage of their specific characteristics. In particular we suspect that most of them present some form of determinism in communication patterns. Communication determinism can play an important role in the design of new fault tolerant protocols by reducing their complexity. In this paper, we explore the communication determinism in 27 HPC parallel applications that are representative of production workloads in large scale centers. We show that most of these applications have deterministic or send-deterministic communication patterns. Franck Cappello, Amina Guermouche, Marc Snir |
ICCCN | 2 |