EDBT 2026 Demo / reviewers in the wild / expert
Alexandre Denis 0001
dblp:58/5391-1
· DBLP profile ↗
27ranked-venue papers
19as first author
6since 2021 · last 2025
0000-0001-8606-4344ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 17 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Communication Notification Through User-Level Interrupts for the BXI NetworkabstractTo reduce the cost of communications in highperformance computing, it is possible to overlap communications with computations. Some communication protocols, such as rendez-vous, multi-chunk messages, and collectives, may require a completion notification to be processed before they can further progress. With active polling or passive waiting, completions are not processed while the application is busy with computation, and thus communication does not progress. However, with an eventbased method like interrupts, it is expected to be much more reactive. Nevertheless, using interrupts usually involves system calls, which are avoided with high-performance networks. The Intel Sapphire Rapids processors introduced user-level interrupts (UINTR), hardware interrupts designed to be used directly in user space, without going through the kernel. However, their current implementation is limited to inter-process communication. They cannot be triggered from a device. In this paper, we propose new mechanisms to extend the scope of user-level interrupts, so as to be able to trigger them from a device and not only from a CPU. We have implemented these mechanisms in the BXI network from Eviden. We have evaluated their performance: we obtain a latency only 2.4 times higher than active polling (v.s. 6 times higher for interrupts with system calls). We have assessed their ability to make communication progress when overlapped with computation; we observe a near-perfect computation/communication overlap. Charles Goedefroit, Alexandre Denis 0001, Mathieu Barbe, Brice Goglin, Gregoire Pichon |
CLUSTER | 2 |
| 2025 | NBLFQ: A Lock-Free MPMC Queue Optimized for Low ContentionabstractThe NewMADELEINE communication library relies extensively on lockless queues for its internal data structures in order to operate well in a multi-threaded application. There is ongoing work to make it use interrupt-based network drivers, and thus its queues have to become truly lock-free and not only lockless. In this paper we present NBLFQ, a lock-free MPMC bounded queue optimized for low contention. We show that 97 % of the queue operations in NewMADELEINE are uncontended, hence the idea to design an algorithm optimized primarily for low contention, but that does not collapse under heavy load. We present the algorithm that relies on a single CAS for enqueue and for dequeue, and proofs of its properties. We have implemented the algorithm, measured its performance on four different architectures, and compared it against a wide range of other lock-free queue algorithms. We have observed that the best performance at the network communication library level is obtained by our NBLFQ algorithm in almost all cases, with both single thread and massively multi-threaded communications. Alexandre Denis 0001, Charles Goedefroit |
IPDPS | 1 |
| 2024 | Tracing task-based runtime systems: Feedbacks from the StarPU caseabstractSummary Given the complexity of current supercomputers and applications, being able to trace application executions to understand their behavior is not a luxury. As constraints, tracing systems have to be as little intrusive as possible in the application code and performances, and be precise enough in the collected data. In this article, we present how works the tracing system used by the task‐based runtime systemStarPU. We study the different sources of performance overhead coming from the tracing system and how to reduce these overheads. Then, we evaluate the accuracy of distributed traces with different clock synchronization techniques. Finally, we summarize our experiments and conclusions with the lessons we learned to efficiently trace applications, and the list of characteristics each tracing system should feature to be competitive. The reported experiments and implementation details comprise a feedback of integrating into a task‐based runtime system state‐of‐the‐art techniques to efficiently and precisely trace application executions. We highlight the points every application developer or end‐user should be aware of to seamlessly integrate a tracing system or just trace application executions. Alexandre Denis 0001, Emmanuel Jeannot, Philippe Swartvagher, Samuel Thibault |
Concurr. Comput. Pract. Exp. | 1 |
| 2022 | One core dedicated to MPI nonblocking communication progression? A model to assess whether it is worth itabstractOverlapping communications with computation is an efficient way to amortize the cost of communications of an HPC application. To do so, it is possible to utilize MPI nonblocking primitives so that communications run in back-ground alongside computation. However, these mechanisms rely on communications actually making progress in the background, which may not be true for all MPI libraries. Some MPI libraries leverage a core dedicated to communications to ensure communication progression. However, taking a core away from the application for such purpose may have a negative impact on the overall execution time. It may be difficult to know when such dedicated core is actually helpful. In this paper, we propose a model for the performance of applications using MPI nonblocking primitives running on top of an MPI library with a dedicated core for communications. This model is used to understand the compromise between computation slowdown due to the communication core not being available for computation, and the communication speed-up thanks to the dedicated core; evaluate whether nonblocking communication is actually obtaining the expected performance in the context of the given application; predict the performance of a given application if ran with a dedicated core. We describe the performance model and evaluate it on different applications. We compare the predictions of the model with actual executions. Alexandre Denis 0001, Julien Jaeger, Emmanuel Jeannot, Florian Reynier |
CCGRID | 1 |
| 2022 | A methodology for assessing computation/communication overlap of MPI nonblocking collectivesabstractSummary By allowing computation/communication overlap, MPI nonblocking collectives (NBC) are supposed to improve application scalability and performance. However, it is known that to actually get overlap, the MPI library has to implement progression mechanisms in software or rely on the network hardware. These mechanisms may be present or not, adequate or perfectible, they may have an impact on communication performance or may interfere with computation by stealing CPU cycles. From a user point of view, assessing and understanding the behavior of an MPI library concerning computation/communication overlap is difficult. In this article, we propose a methodology to assess the computation/communication overlap of NBC. We propose new metrics to measure how much communication and computation do overlap, and to evaluate how they interfere with each other. We integrate these metrics into a complete methodology. We compare our methodology with state of the art metrics and benchmarks, and show that ours provides more meaningful informations. We perform experiments on a large panel of MPI implementations and network hardware and show when and why overlap is efficient, nonexistent or even degrades performance. Alexandre Denis 0001, Julien Jaeger, Emmanuel Jeannot, Florian Reynier |
Concurr. Comput. Pract. Exp. | 1 |
| 2021 | Interferences between Communications and Computations in Distributed HPC SystemsabstractInternational audience Alexandre Denis 0001, Emmanuel Jeannot, Philippe Swartvagher |
ICPP | 1 |
| 2020 | Using Dynamic Broadcasts to Improve Task-Based Runtime Performances
Alexandre Denis 0001, Emmanuel Jeannot, Philippe Swartvagher, Samuel Thibault |
Euro-Par | 1 |
| 2019 | Scalability of the NewMadeleine Communication Library for Large Numbers of MPI Point-to-Point RequestsabstractNew kinds of applications with lots of threads or irregular communication patterns which rely a lot on point-to-point MPI communications have emerged. It stresses the MPI library with potentially a lot of simultaneous MPI requests for sending and receiving at the same time. To deal with large numbers of simultaneous requests, the bottleneck lies in two main mechanisms: the tag-matching (the algorithm that matches an incoming packet with a posted receive request), and the progression engine. In this paper, we propose algorithms and implementations that overcome these issues so as to scale up to thousands of requests if needed. In particular our algorithms are able to perform constant-time tag-matching even with any-source and any-tag support. We have implemented these mechanisms in our NewMadeleine communication library. Through micro-benchmarks and computation kernel benchmarks, we demonstrate that our MPI library exhibits better performance than state-of-the-art MPI implementations in cases with many simultaneous requests. Alexandre Denis 0001 |
CCGRID | 1 |
| 2018 | Dynamic Placement of Progress Thread for Overlapping MPI Non-blocking Collectives on Manycore Processor
Alexandre Denis 0001, Julien Jaeger, Emmanuel Jeannot, Marc Pérache, Hugo Taboada |
Euro-Par | 1 |
| 2016 | MPI Overlap: Benchmark and AnalysisabstractIn HPC applications, one of the major overhead compared to sequential code, is communication cost. Application programmers often amortize this cost by overlapping communications with computation. To do so, they post a non-blocking MPI request, perform computation, and wait for communication completion, assuming MPI communication will progress in background. In this paper, we propose to measure what really happens when trying to overlap non-blocking point-to-point communications with computation. We explain how background progression works, we describe relevant test cases, we identify challenges for a benchmark, then we propose a benchmark suite to measure how much overlap happen in various cases. We exhibit overlap benchmark results on a wide panel of MPI libraries and hardware platforms. Finally, we classify, analyze, and explain the results using low-level traces to reveal the internal behavior of the MPI library. Alexandre Denis 0001, François Trahay |
ICPP | 1 |
| 2015 | pioman: A Pthread-Based Multithreaded Communication EngineabstractRecent cluster architectures include dozens of cores per node, with all cores sharing the network resources. To program such architectures, hybrid models mixing MPI+threads, and in particular MPI+OpenMP are gaining popularity. This imposes new requirements on communication libraries, such as the need for MPI_THREAD_MULTIPLE level of multi-threading support. Moreover, the high number of cores brings new opportunities to parallelize communication libraries, so as to have proper background progression of communication and communication/computation overlap. In this paper, we present pioman, a generic framework to be used by MPI implementations, that brings seamless asynchronous progression of communication by opportunistically using available cores. It uses system threads and thus is composable with any runtime system used for multithreading. Through various benchmarks, we demonstrate that our pioman-based MPI implementation exhibits very good properties regarding overlap, progression, and multithreading, and outperforms state-of-art MPI implementations. Alexandre Denis 0001 |
PDP | 1 |
| 2014 | a generic framework for asynchronous progression and multithreaded communicationsabstractRecent cluster architectures include dozens of cores per node, with all cores sharing the network resources. To program such architectures, hybrid models mixing MPI+threads, and in particular MPI+OpenMP are gaining popularity. This imposes new requirements on communication libraries, such as the need for MPI_THREAD_MULTIPLE level of multi-threading support. Moreover, the high number of cores brings new opportunities to parallelize communication libraries, so as to have proper background progression of communication and communication/computation overlap. In this paper, we present pioman, a generic framework to be used by MPI implementations, that brings seamless asynchronous progression of communication by opportunistically using available cores. It uses system threads and thus is composable with any runtime system used for multithreading. Through various benchmarks, we demonstrate that our pioman-based MPI implementation exhibits very good properties regarding overlap, progression, and multithreading, and outperforms state-of-art MPI implementations. Alexandre Denis 0001 |
CLUSTER | 1 |
| 2014 | Toward OpenCL Automatic Multi-Device Support
Sylvain Henry, Alexandre Denis 0001, Denis Barthou, Marie Christine Counilh, Raymond Namyst |
Euro-Par | 2 |
| 2012 | High Performance Checksum Computation for Fault-Tolerant MPI over Infiniband
Alexandre Denis 0001, François Trahay, Yutaka Ishikawa |
EuroMPI | 1 |
| 2011 | A Sampling-Based Approach for Communication Libraries Auto-TuningabstractCommunication performance is a critical issue in HPC applications, and many solutions have been proposed on the literature (algorithmic, protocols, etc.) In the meantime, computing nodes become massively multicore, leading to a real imbalance between the number of communication sources and the number of physical communication resources. Thus it is now mandatory to share network boards between computation flows, and to take this sharing into account while performing communication optimizations. In previous papers, we have proposed a model and a frame work for on-the-fly optimizations of multiplexed concurrent communication flows, and implemented this model in the NEWMADELEINE communication library. This library features optimization strategies able for example to aggregate several messages to reduce the number of packets emitted on the network, or to split messages to use several NICs at the same time. In this paper, we study the tuning of these dynamic optimization strategies. We show that some parameters and thresholds (rendezvous threshold, aggregation packet size) depend on the actual hardware, both host and NICs. We propose and implement a method based on sampling of the actual hardware to auto-tune our strategies. Moreover, we show that multi-rail can greatly benefit from performance predictions. We propose an approach for multi-rail that dynamically balance the data between NICs using predictions based on sampling. Elisabeth Brunet, François Trahay, Alexandre Denis 0001, Raymond Namyst |
CLUSTER | 3 |
| 2011 | A High Performance Superpipeline Protocol for InfiniBand
Alexandre Denis 0001 |
Euro-Par (2) | 1 |
| 2009 | A scalable and generic task scheduling system for communication librariesabstractSince the advent of multi-core processors, the physionomy of typical clusters has dramatically evolved. This new massively multi-core era is a major change in architecture, causing the evolution of programming models towards hybrid MPI+threads, therefore requiring new features at low-level. Modern communication subsystems now have to deal with multi-threading: the impact of thread-safety, the contention on network interfaces or the consequence of data locality on performance have to be studied carefully. In this paper, we present PIOMan, a scalable and generic lightweight task scheduling system for communication libraries. It is designed to ensure concurrent progression of multiple tasks of a communication library (polling, offload, multi-rail) through the use of multiple cores, while preserving locality to avoid contention and allow a scalability to a large number of cores and threads. We have implemented the model, evaluated its performance, and compared it to state of the art solutions regarding overhead, scalability, and communication and computation overlap. François Trahay, Alexandre Denis 0001 |
CLUSTER | 2 |
| 2009 | An analysis of the impact of multi-threading on communication performanceabstractAlthough processors become massively multicore and therefore new programming models mix message passing and multi-threading, the effects of threads on communication libraries remain neglected. Designing an efficient modern communication library requires precautions in order to limit the impact of thread-safety mechanisms on performance. In this paper, we present various approaches to building a thread-safe communication library and we study their benefit and impact on performance. We also describe and evaluate techniques used to exploit idle cores to balance the communication library load across multicore machines. François Trahay, Elisabeth Brunet, Alexandre Denis 0001 |
IPDPS | 3 |
| 2008 | A multicore-enabled multirail communication engineabstractThe current trend in clusters architecture leads toward a massive use of multicore chips. This hardware evolution raises bottleneck issues at the network interface level. The use of multiple parallel networks allows to overcome this problem as it provides an higher aggregate bandwidth. But this bandwidth remains theoretical as only a few communication libraries are able to exploit multiple networks. In this paper, we present an optimization strategy for the NEWMADELEINE communication library. This strategy is able to efficiently exploit parallel interconnect links. By sampling each networkpsilas capabilities, it is possible to estimate a transfer duration a priori. Splitting messages and sending chunks of messages over parallel links can thus be performed efficiently to reach the theoretical aggregate bandwidth. NEWMADELEINE is multithreaded and exploits multicore chips to send small packets, that involve CPU-consuming copies, in parallel. Elisabeth Brunet, François Trahay, Alexandre Denis 0001 |
CLUSTER | 3 |
| 2008 | A multithreaded communication engine for multicore architecturesabstractThe current trend in clusters leads towards an increase of the number of cores per node. As a result, an increasing number of parallel applications is mixing message passing and multithreading as an attempt to better match the underlying architecture's structure. This naturally raises the problem of designing efficient, multithreaded implementations of MPL In this paper, we present the design of a multithreaded communication engine able to exploit idle cores to speed up communications in two ways: it can move CPU- intensive operations out of the critical path (e.g. PIO transfers off load), and is able to let rendezvous transfers progress asynchronously. We have implemented these methods in the PM2 software suite, evaluated their behavior in typical cases, and we have observed good performance results in overlapping communication and computation. François Trahay, Elisabeth Brunet, Alexandre Denis 0001, Raymond Namyst |
IPDPS | 3 |
| 2004 | Wide-Area Communication for Grids: An Integrated Solution to Connectivity, Performance and Security Problems
Alexandre Denis 0001, Olivier Aumage, Rutger F. H. Hofman, Kees Verstoep, Thilo Kielmann, Henri E. Bal |
HPDC | 1 |
| 2004 | Network Communications in Grid Computing: At a Crossroads between Parallel and Distributed WorldsabstractSummary form only given. This article studies a communication model that aims at extending the scope of computational grids by allowing the execution of parallel and/or distributed applications without imposing any programming constraints or the use of a particular communication layer. Such model leads to the design of a communication framework for grids which allows the use of the appropriate middleware for the application rather than the one dictated by the available resources. Such a framework is able to handle any communication middleware - even several at the same time - on any kind of networking technologies. Our proposed dual-abstraction (parallel and distributed) model is organized into three layers: arbitration, abstraction and personalities which are highlighted. The performance obtained with PadicoTM, our available open source implementation of the proposed framework, show that such functionality can be obtained with still providing very high performance. Alexandre Denis 0001, Christian Pérez, Thierry Priol |
IPDPS | 1 |
| 2003 | Achieving portable and efficient parallel CORBA objectsabstractAbstract With the availability of Computational Grids, new kinds of applications are emerging. They raise the problem of how to program them on such computing systems. In this paper, we advocate a programming model based on a combination of parallel and distributed programming models. Compared to previous approaches, this work aims at bringing single program multiple data (SPMD) programming into CORBA in a portable way. For example, we want to interconnect two parallel codes by CORBA without modifying either CORBA or the parallel communication API. We show that such an approach does not entail any loss of performance compared to previous approaches that required modification to the CORBA standard. Moreover, using an ORB that is able to exploit high‐performance networks, we show that portable parallel CORBA objects can efficiently make use of such networks. Copyright © 2003 John Wiley & Sons, Ltd. Alexandre Denis 0001, Christian Pérez, Thierry Priol |
Concurr. Comput. Pract. Exp. | 1 |
| 2003 | PadicoTM: an open integration framework for communication middleware and runtimes
Alexandre Denis 0001, Christian Pérez, Thierry Priol |
Future Gener. Comput. Syst. | 1 |
| 2002 | PadicoTM: An Open Integration Framework for Communication Middleware and RuntimesabstractComputational grids are seen as the future emergent computing infrastructures. Their programming requires the use of several paradigms that are implemented through communication middleware and runtimes. However some of these middleware systems and runtimes are unable to take benefit of the presence of specific networking technologies available in grid infrastructures. In this paper, we describe an open integration framework that allows several communication middleware and runtimes to efficiently share the networking resources available in a computational grid. Such framework encourages grid programmers to use the most suited communication paradigms for their applications independently from the underlying networks. Therefore, there is no obstacle to deploy the applications on a specific grid configuration. Alexandre Denis 0001, Christian Pérez, Thierry Priol |
CCGRID | 1 |
| 2001 | Portable Parallel CORBA Objects: An Approach to Combine Parallel and Distributed Programming for Grid Computing
Alexandre Denis 0001, Christian Pérez, Thierry Priol |
Euro-Par | 1 |
| 2000 | Madeleine II: a Portable and Efficient Communication Library for High-Performance Cluster ComputingabstractThis paper introduces Madeleine II, a new adaptive and portable multi-protocol implementation of the Madeleine communication library. Madeleine II has the ability to control multiple network interfaces (BIP, SISCI, VIA) and multiple network adapters (Ethernet, Myrinet, SCI) within the same application session. Moreover it includes advanced mechanisms to dynamically select the most appropriate transfer method for a given network protocol according to various parameters such as data size or responsiveness user requirements. We report on performance measurements obtained using BIP/Myrinet and SISCI/SCI and we present preliminary results about our Nexus/Madeleine II and MPICH/Madeleine II ports. Olivier Aumage, Luc Bougé, Alexandre Denis 0001, Jean-François Méhaut, Guillaume Mercier, Raymond Namyst, Loïc Prylli |
CLUSTER | 3 |