Rémi Dulong

dblp:256/0797 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
3since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 P4ce: Consensus over RDMA at Line Speed
abstract
P4ce is the first replication protocol that exhibits the same latency and requires the same network capacity as sending data to a single server. P4ce builds upon previous RDMA-based consensus protocols. They achieve consensus with a single network round-trip, but with a reduced network throughput. P4ce also achieves consensus with a single round-trip, but without degrading throughput by decoupling the consensus decisions from the RDMA communications. The decision part of the consensus protocol runs on a commodity server, but the communication part of P4ce is fully implemented on a programmable switch, which replicates data and aggregates the acknowledgements in the network, avoiding the throughput bottleneck at the leader. Although simple in its principle, the implementation of P4ce raises many challenging issues, notably caused by the complexity of RDMA and the underlying network protocols, the intricacies of packet rewriting during replication and aggregation, and the restricted set of operations that can be implemented at wire speed in the programmable switch. We implemented P4ce and deployed it on a commercially-available Intel Tofino switch, achieving up to 4x better through-put and better latency than state-of-the-art consensus protocols.
Rémi Dulong, Nathan Felber, Pascal Felber, Gilles Hopin, Baptiste Lepers, Valerio Schiavoni, Gaël Thomas 0001, Sébastien Vaucher
ICDCS1
2022 duf: Dynamic uncore frequency scaling to reduce power consumption
abstract
Summary Reducing the power consumption of applications has become one of the key challenges in high‐performance computing. Recent processor architectures differentiate processor core frequency from its uncore frequency. As a consequence, in addition to tuning processor core frequency with dynamic voltage and frequency scaling, power consumption can also be controlled through uncore frequency scaling. This article studies how the uncore frequency can be used as a leverage to improve power consumption. We propose duf, a daemon process that dynamically adapts the uncore frequency to reduce an application power consumption with a user‐defined limit on performance degradation. The evaluation of duf on three different architectures shows that with no performance degradation (<0.6%), duf can reduce socket power consumption by 7.94%. We also show that duf is able to reduce the total energy consumption by up to 18.20%.
Étienne André 0002, Rémi Dulong, Amina Guermouche, François Trahay
Concurr. Comput. Pract. Exp.2
2021 NVCache: A Plug-and-Play NVMM-based I/O Booster for Legacy Systems
abstract
This paper introduces NVCACHE, an approach that uses a non-volatile main memory (NVMM) as a write cache to improve the write performance of legacy applications. We compare NVCACHE against file systems tailored for NVMM (Ext4-DAX and NOVA) and with I/O-heavy applications (SQLite, RocksDB). Our evaluation shows that NVCACHE reaches the performance level of the existing state-of-the-art systems for NVMM, but without their limitations: NVCACHE does not limit the size of the stored data to the size of the NVMM, and works transparently with unmodified legacy applications, providing additional persistence guarantees even when their source code is not available.
Rémi Dulong, Rafael Pires 0001, Andreia Correia, Valerio Schiavoni, Pedro Ramalhete, Pascal Felber, Gaël Thomas 0001
DSN1
2019 Using Differential Execution Analysis to Identify Thread Interference
abstract
Understanding the performance of a multi-threaded application is difficult. The threads interfere when they access the same shared resource, which slows down their execution. Unfortunately, current profiling tools report the hardware components or the synchronization primitives that saturate, but they cannot tell if the saturation is the cause of a performance bottleneck. In this paper, we propose a holistic metric able to pinpoint the blocks of code that suffer interference the most, regardless of the interference cause. Our metric uses performance variation as a universal indicator of interference problems. With an evaluation of 27 applications we show that our metric can identify interference problems caused by six different kinds of interference in nine applications. We are able to easily remove seven of the bottlenecks, which leads to a performance improvement of up to nine times.
Mohamed Said Mosli Bouksiaa, François Trahay, Alexis Lescouet, Gauthier Voron, Rémi Dulong, Amina Guermouche, Elisabeth Brunet, Gaël Thomas 0001
IEEE Trans. Parallel Distributed Syst.5