EDBT 2026 Demo / reviewers in the wild / expert
Thomas Ropars
dblp:54/6292
· DBLP profile ↗
23ranked-venue papers
8as first author
4since 2021 · last 2026
0000-0002-9461-1165ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 7 first-author · 4 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fast Checkpointing in Disaggregated Persistent Memory
Ivane Adam, Thomas Ropars, Noel De Palma |
Euro-Par (2) | 2 |
| 2025 | Carbon Footprint of Storage in Data Centers: the Impact of Using Ssds for Key-Value StoresabstractThe greenhouse gas emissions of data centers and the question of how to reduce them are a broad problem that has come into focus with the ongoing climate crisis. Yet, the emissions of storage infrastructures specifically are not well understood. Recent work shows that the manufacturing emissions of SSDs are significantly higher than those of HDDs. Also, the energy consumption of high-end SSDs is in general higher than the one of HDDs. This raises the question of whether the performance improvements provided by SSDs over HDDs are enough to also provide an advantage in terms of carbon footprint. In this paper, we analyze the lifecycle carbon footprint of SSDs and HDDs when used as storage devices for Key-Value stores. Considering state-of-the-art Key-Value stores specifically designed to make best use of HDDs or SSDs, we use specialized hardware to measure the power consumption of NVMe SSDs, as well as of the processor and memory to analyze the impact of different Key-Value stores and workloads on power. We conduct an analysis to determine if SSD-based systems outperform HDD-based systems with respect to their carbon footprint. Our results show that in most cases, the high operational energy efficiency of SSDs allow systems based on SSDs to have lower carbon footprint. However, HDD-based solutions should be considered in cases where the full potential of SSDs cannot be exploited and when the carbon intensity of the electricity powering the data center is low. Jakob Nibler, Thomas Ropars |
CCGrid | 2 |
| 2022 | ResPCT: fast checkpointing in non-volatile memory for multi-threaded applicationsabstractNon-volatile memory (NVMM) technologies are a great opportunity to build fast fault-tolerant programs, as they provide persistent storage in main memory. However, since the processor caches remain volatile, solutions are needed to recover a consistent state from NVMM after a crash. This paper presents ResPCT, a checkpointing approach to make multi-threaded programs fault tolerant, by flushing persistent data structures to NVMM periodically. ResPCT uses In-Cache-Line logging to efficiently track modifications during failure-free execution, and to restore a consistent state after a crash. The ResPCT API enables programmers to position restart points in their program, which simplifies the identification of the persistent program state and can also help improving performance. Experiments with representative benchmarks and applications, show that ResPCT can outperform state-of-the-art solutions by up to 2.7×, and that its overhead can be as low as 4% at large core count. Ana Khorguani, Thomas Ropars, Noel De Palma |
EuroSys | 2 |
| 2021 | CPU overheating prediction in HPC systemsabstractSummary With the increase in size of supercomputers, also increases the number of abnormal events. CPU overheating is one such event that decreases the system efficiency: when a CPU overheats, it reduces its frequency. This paper presents a machine learning solution to predict such events. The proposed algorithm is based on dynamic time warping for feature extraction and on a machine learning algorithm for classification. It predicts overheating events solely by analyzing the trends of the temperature of the CPUs and can deal with very low temperature sampling rates while having a negligible computational cost in practice. Our evaluation, using data coming from a production supercomputer, shows that the proposed solution can make predictions a few minutes in advance with a good accuracy. Furthermore, considering two simple preventive actions to avoid CPU overheating events, we present an analytical study that shows that our predictive solution is good enough to allow a significant reduction of the cost of overheating events. Marc Platini, Thomas Ropars, Benoit Pelletier, Noel De Palma |
Concurr. Comput. Pract. Exp. | 2 |
| 2019 | Data and Thread Placement in NUMA Architectures: A Statistical Learning ApproachabstractNowadays, NUMA architectures are common in compute-intensive systems. Achieving high performance for multi-threaded application requires both a careful placement of threads on computing units and a thorough allocation of data in memory. Finding such a placement is a hard problem to solve, because performance depends on complex interactions in several layers of the memory hierarchy. In this paper we propose a black-box approach to decide if an application execution time can be impacted by the placement of its threads and data, and in such a case, to choose the best placement strategy to adopt. We show that it is possible to reach near-optimal placement policy selection. Furthermore, solutions work across several recent processor architectures and decisions can be taken with a single run of low overhead profiling. Nicolas Denoyelle, Brice Goglin, Emmanuel Jeannot, Thomas Ropars |
ICPP | 4 |
| 2018 | An Efficient Wait-free Resizable Hash TableabstractThis paper presents an efficient wait-free resizable hash table. To achieve high throughput at large core counts, our algorithm is specifically designed to retain the natural parallelism of concurrent hashing, while providing wait-free resizing. An extensive evaluation of our hash table shows that in the common case where resizing actions are rare, our implementation outperforms all existing lock-free hash table implementations while providing a stronger progress guarantee. Panagiota Fatourou, Nikolaos D. Kallimanis, Thomas Ropars |
SPAA | 3 |
| 2015 | Efficient Process Replication for MPI Applications: Sharing Work between ReplicasabstractWith the increased failure rate expected in future extreme scale supercomputers, process replication might become a viable alternative to check pointing. By default, the workload efficiency of replication is limited to 50% because of the additional resources that have to be used to execute the replicas of the application's processes. In this paper, we introduce intra-parallelization, a solution that avoids replicating all computation by introducing work-sharing between replicas. We show on a representative set of benchmarks that intra-parallelization allows achieving more than 50% efficiency without compromising fault tolerance. Thomas Ropars, Arnaud Lefray, André Schiper |
IPDPS | 1 |
| 2014 | High-Throughput Maps on Message-Passing Manycore Architectures: Partitioning versus Replication
Omid Shahmirzadi, Thomas Ropars, André Schiper |
Euro-Par | 2 |
| 2014 | Leveraging hardware message passing for efficient thread synchronizationabstractAs the level of parallelism in manycore processors keeps increasing, providing efficient mechanisms for thread synchronization in concurrent programs is becoming a major concern. On cache-coherent shared-memory processors, synchronization efficiency is ultimately limited by the performance of the underlying cache coherence protocol. This paper studies how hardware support for message passing can improve synchronization performance. Considering the ubiquitous problem of mutual exclusion, we adapt two state-of-the-art solutions used on shared-memory processors, namely the server approach and the combining approach, to leverage the potential of hardware message passing. We propose HybComb, a novel combining algorithm that uses both message passing and shared memory features of emerging hybrid processors. We also introduce MP-Server, a straightforward adaptation of the server approach to hardware message passing. Evaluation on Tilera's TILE-Gx processor shows that MP-Server can execute contended critical sections with unprecedented throughput, as stalls related to cache coherence are removed from the critical path. HybComb can achieve comparable performance, while avoiding the need to dedicate server cores. Consequently, our queue and stack implementations, based on MP-Server and HybComb, largely outperform their most efficient pure-shared-memory counterparts. Darko Petrovic, Thomas Ropars, André Schiper |
PPoPP | 2 |
| 2013 | SPBC: leveraging the characteristics of MPI HPC applications for scalable checkpointingabstractThe high failure rate expected for future supercomputers requires the design of new fault tolerant solutions. Most checkpointing protocols are designed to work with any message-passing application but suffer from scalability issues at extreme scale. We take a different approach: We identify a property common to many HPC applications, namely channel-determinism, and introduce a new partial order relation, called always-happens-before relation, between events of such applications. Leveraging these two concepts, we design a protocol that combines an unprecedented set of features. Our protocol called SPBC combines in a hierarchical way coordinated checkpointing and message logging. It is the first protocol that provides failure containment without logging any information reliably apart from process checkpoints, and this, without penalizing recovery performance. Experiments run with a representative set of HPC workloads demonstrate a good performance of our protocol during both, failure-free execution and recovery. Thomas Ropars, Tatiana V. Martsinkevich, Amina Guermouche, André Schiper, Franck Cappello |
SC | 1 |
| 2012 | Hierarchical Clustering Strategies for Fault Tolerance in Large Scale HPC SystemsabstractFuture high performance computing systems will need to use novel techniques to allow scientific applications to progress despite frequent failures. Checkpoint-Restart is currently the most popular way to mitigate the impact of failures during long-running executions. Different techniques try to reduce the cost of Checkpoint-Restart, some of them such as local check pointing and erasure codes aim to reduce the time to checkpoint while others such as uncoordinated checkpoint and message-logging aim to decrease the cost of recovery. In this paper, we study how to combine all these techniques together in order to optimize both: check pointing and recovery. We present several clustering and topology challenges that lead us to an optimization problem in a four-dimensional space: reliability level, recovery cost, encoding time and message logging overhead. We propose a novel clustering method inspired from brain topology studies in neuroscience and evaluate it with a Tsunami simulation application in TSUBAME2. Our evaluation with 1024 processes shows that our novel clustering method can guarantee good performance for all of the four mentioned dimensions of our optimization problem. Leonardo Arturo Bautista-Gomez, Thomas Ropars, Naoya Maruyama, Franck Cappello, Satoshi Matsuoka |
CLUSTER | 2 |
| 2012 | HydEE: Failure Containment without Event Logging for Large Scale Send-Deterministic MPI ApplicationsabstractHigh performance computing will probably reach exascale in this decade. At this scale, mean time between failures is expected to be a few hours. Existing fault tolerant protocols for message passing applications will not be efficient anymore since they either require a global restart after a failure (check pointing protocols) or result in huge memory occupation (message logging). Hybrid fault tolerant protocols overcome these limits by dividing applications processes into clusters and applying a different protocol within and between clusters. Combining coordinated check pointing inside the clusters and message logging for the inter-cluster messages allows confining the consequences of a failure to a single cluster, while logging only a subset of the messages. However, in existing hybrid protocols, event logging is required for all application messages to ensure a correct execution after a failure. This can significantly impair failure free performance. In this paper, we propose HydEE, a hybrid rollback-recovery protocol for send-deterministic message passing applications, that provides failure containment without logging any event, and only a subset of the application messages. We prove that HydEE can handle multiple concurrent failures by relying on the send-deterministic execution model. Experimental evaluations of our implementation of HydEE in the MPICH2 library show that it introduces almost no overhead on failure free execution. Amina Guermouche, Thomas Ropars, Marc Snir, Franck Cappello |
IPDPS | 2 |
| 2012 | High-performance RMA-based broadcast on the intel SCCabstractMany-core chips with more than 1000 cores are expected by the end of the decade. To overcome scalability issues related to cache coherence at such a scale, one of the main research directions is to leverage the message-passing programming model. The Intel Single-Chip Cloud Computer (SCC) is a prototype of a message-passing many-core chip. It offers the ability to move data between on-chip Message Passing Buffers (MPB) using Remote Memory Access (RMA). Performance of message-passing applications is directly affected by efficiency of collective operations, such as broadcast. In this paper, we study how to make use of the MPBs to implement an efficient broadcast algorithm for the SCC. We propose OC-Bcast (On-Chip Broadcast), a pipelined k-ary tree algorithm tailored to exploit the parallelism provided by on-chip RMA. Using a LogP-based model, we present an analytical evaluation that compares our algorithm to the state-of-the-art broadcast algorithms implemented for the SCC. As predicted by the model, experimental results show that OC-Bcast attains almost three times better throughput, and improves latency by at least 27\%. Furthermore, the analytical evaluation highlights the benefits of our approach: OC-Bcast takes direct advantage of RMA, unlike the other considered broadcast algorithms, which are based on a higher-level send/receive interface. This leads us to the conclusion that RMA-based collective operations are needed to take full advantage of hardware features of future message-passing many-core architectures. Darko Petrovic, Omid Shahmirzadi, Thomas Ropars, André Schiper |
SPAA | 3 |
| 2011 | On the Use of Cluster-Based Partial Message Logging to Improve Fault Tolerance for MPI HPC Applications
Thomas Ropars, Amina Guermouche, Bora Uçar, Esteban Meneses, Laxmikant V. Kalé, Franck Cappello |
Euro-Par (1) | 1 |
| 2011 | Uncoordinated Checkpointing Without Domino Effect for Send-Deterministic MPI ApplicationsabstractAs reported by many recent studies, the mean time between failures of future post-petascale supercomputers is likely to reduce, compared to the current situation. The most popular fault tolerance approach for MPI applications on HPC Platforms relies on coordinated check pointing which raises two major issues: a) global restart wastes energy since all processes are forced to rollback even in the case of a single failure, b) checkpoint coordination may slow down the application execution because of congestions on I/O resources. Alternative approaches based on uncoordinated check pointing and message logging require logging all messages, imposing a high memory/storage occupation and a significant overhead on communications. It has recently been observed that many MPI HPC applications are send-deterministic, allowing to design new fault tolerance protocols. In this paper, we propose an uncoordinated check pointing protocol for send-deterministic MPI HPC applications that (i) logs only a subset of the application messages and (ii) does not require to restart systematically all processes when a failure occurs. We first describe our protocol and prove its correctness. Through experimental evaluations, we show that its implementation in MPICH2 has a negligible overhead on application performance. Then we perform a quantitative evaluation of the properties of our protocol using the NAS Benchmarks. Using a clustering approach, we demonstrate that this protocol actually succeeds to combine the two expected properties: a) it logs only a small fraction of the messages and b) it reduces by a factor approaching 2 the average number of processes to rollback compared to coordinated check pointing. Amina Guermouche, Thomas Ropars, Elisabeth Brunet, Marc Snir, Franck Cappello |
IPDPS | 2 |
| 2011 | Active optimistic and distributed message logging for message-passing applicationsabstractSUMMARY Message logging is an attractive solution to provide fault tolerance for message‐passing applications because it is more scalable than coordinated checkpointing. Sender‐based message logging is a well‐known optimization that allows the saving of message payload in the sender memory. Thus, only message reception events have to be logged reliably by using an event logger. This paper proposes solutions to further improve message logging protocol scalability. In existing works on message logging, the event logger has always been considered as a centralized process. We propose a distributed event logger that takes advantage of multi‐core processors that are to be executed in parallel with application processes, leveraging the volatile memory of the nodes to save events reliably. We also propose the combination of our distributed event logger and O2P, an active optimistic message logging protocol using a gossip‐based protocol to disseminate information on new stable events. Our distributed event logger and O2P are implemented in the Open MPI library. Our results show the following: (i) distributed event logging improves message logging protocol scalability and (ii) using O2P with a distributed event logger provides an efficient and scalable fault‐tolerant solution for message‐passing applications. Copyright © 2011 John Wiley & Sons, Ltd. Thomas Ropars, Christine Morin |
Concurr. Comput. Pract. Exp. | 1 |
| 2010 | Improving Message Logging Protocols Scalability through Distributed Event Logging
Thomas Ropars, Christine Morin |
Euro-Par (1) | 1 |
| 2010 | Semias: Self-Healing Active Replication on Top of a Structured Peer-to-Peer OverlayabstractActive replication on top of a structured peer-to-peer overlay is an attractive solution for transparently providing high availability to distributed applications. However, self-healing is necessary to ensure the availability of the replicated application despite node arrivals, failures or departures in the overlay. Self-healing means to automatically reconfigure the replica groups when changes in the overlay occur. In the case of active replication, reconfigurations must be done carefully, to keep the replicas consistency. Moreover, as every reconfiguration could imply a state transfer between replicas, their number should be limited. In this paper we propose a self-healing solution that limits the number of group reconfigurations and ensures the availability of the replicated application in a dynamic environment. To evaluate the performance of our solution, we implemented it in a framework, called Semias, in the context of Vigne grid middleware. Experiments run on Grid'5000 and Planet Lab show the performance of the framework and the efficiency of our self-healing mechanisms in a dynamic environment. Stefania Costache 0002, Thomas Ropars, Christine Morin |
SRDS | 2 |
| 2009 | Reasons for a pessimistic or optimistic message logging protocol in MPI uncoordinated failure, recoveryabstractWith the growing scale of high performance computing platforms, fault tolerance has become a major issue. Among the various approaches for providing fault tolerance to MPI applications, message logging has been proved to tolerate higher failure rate. However, this advantage comes at the expense of a higher overhead on communications, due to latency intrusive logging of events to a stable storage. Previous work proposed and evaluated several protocols relaxing the synchronicity of event logging to moderate this overhead. Recently, the model of message logging has been refined to better match the reality of high performance network cards, where message receptions are decomposed in multiple interdependent events. According to this new model, deterministic and non-deterministic events are clearly discriminated, reducing the overhead induced by message logging. In this paper we compare, experimentally, a pessimistic and an optimistic message logging protocol, using this new model and implemented in the Open MPI library. Although pessimistic and optimistic message logging are, respectively, the most and less synchronous message logging paradigms, experiments show that most of the time their performance is comparable. Aurelien Bouteiller, Thomas Ropars, George Bosilca, Christine Morin, Jack J. Dongarra |
CLUSTER | 2 |
| 2009 | The Architecture of the XtreemOS Grid Checkpointing Service
John Mehnert-Spahn, Thomas Ropars, Michael Schöttner, Christine Morin |
Euro-Par | 2 |
| 2009 | Active Optimistic Message Logging for Reliable Execution of MPI Applications
Thomas Ropars, Christine Morin |
Euro-Par | 1 |
| 2008 | Fault Tolerance in Cluster Federations with O2P-CFabstractFault tolerance is one of the key issues for large scale applications executed on high performance computing systems. In a cluster federation, clusters are gathered to provide huge computing power. To work efficiently on such systems, networks characteristics have to be taken into account: the latency between two nodes of different clusters is much higher than the latency between two nodes of the same cluster. In this paper, we present O2P-CF a message logging protocol well-suited to provide fault tolerance for message passing applications executed on cluster federations. O2P-CF is based on the combination of O2P, an extremely optimistic message logging protocol, with a pessimistic message logging protocol. Thomas Ropars, Christine Morin |
CCGRID | 1 |
| 2007 | GAMoSe: An Accurate Monitoring Service For Grid ApplicationsabstractMonitoring distributed applications executed on a computational Grid is challenging since they are executed on several heterogeneous nodes belonging to different administrative domains. An application monitoring service should provide users and administrators with useful and dependable data on the applications executed on the Grid. We present in this paper a grid application monitoring service designed to handle high availability and scalability issues. The service supplies information on application state, on failures and on re- source consumption. A set of transparent monitoring mechanisms are used according to grid node nature to effectively monitor the applications. Experiments on the Grid'5000 testbed show that the service provides dependable information with a minimal cost on Grid performances. Thomas Ropars, Emmanuel Jeanvoine, Christine Morin |
ISPDC | 1 |