EDBT 2026 Demo / reviewers in the wild / expert
Patrick Carribault
dblp:08/1403
· DBLP profile ↗
26ranked-venue papers
1as first author
6since 2021 · last 2026
0009-0003-7210-0449ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 1 first-author · 5 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Measuring and interpreting performances of HPC applications with dependent tasksabstractBreaking down the parallel time into work, idleness, and overheads is crucial for assessing the performance of HPC applications but is challenging to measure in runtime systems with dependent tasks. No existing tools allow its measurement accurately. This paper introduces POT: a tool-suite for parallel applications performance analysis with support for dependent tasks. We focus on its low-disturbance methodology consisting of parallel object modeling, discrete-event tracing, and post-mortem simulation-based analysis. The POT tool-suite allows the tracing and analysis of OMPT (OpenMP), PMPI (MPI) and pthreads events. The paper evaluates the accuracy of POT’s analysis on LLVM and MPC-OMP implementations. It shows that measurement bias may be neglected above 16 μ s workload per task, portably across two architectures and OpenMP runtime systems. We also illustrate the benefits unveiled by POT post-mortem simulation approach for analyzing mixed programming models with MPI+OpenMP. Romain Pereira, Adrien Roussel, Patrick Carribault |
Future Gener. Comput. Syst. | 4 |
| 2024 | To Share or Not to Share: A Case for MPI in Shared-Memory
Julien Adam, Jean-Baptiste Besnard, Adrien Roussel, Julien Jaeger, Patrick Carribault, Marc Pérache |
EuroMPI | 5 |
| 2023 | Investigating Dependency Graph Discovery Impact on Task-based MPI+OpenMP Applications PerformancesabstractThe architecture of supercomputers is evolving to expose massive parallelism. MPI and OpenMP are widely used in application codes on the largest supercomputers in the world. The community primarily focused on composing MPI with OpenMP before its version 3.0 introduced task-based programming. Recent advances in OpenMP task model and its interoperability with MPI enabled fine model composition and seamless support for asynchrony. Yet, OpenMP tasking overheads limit the gain of task-based applications over their historical loop parallelization (parallel for construct). Romain Pereira, Adrien Roussel, Patrick Carribault |
ICPP | 3 |
| 2022 | Relative Performance Projection on Arm Architectures
Clément Gavoille, Hugo Taboada, Patrick Carribault, Fabrice Dupros, Brice Goglin, Emmanuel Jeannot |
Euro-Par | 3 |
| 2022 | MPI detach - Towards automatic asynchronous local completion
Joachim Jenke, Marc-André Hermanns, Matthias S. Müller, Van Man Nguyen, Julien Jaeger, Emmanuelle Saillard, Patrick Carribault, Denis Barthou |
Parallel Comput. | 7 |
| 2021 | Enhancing Load-Balancing of MPI Applications with Workshare
Thomas Dionisi, Stéphane Bouhrour, Julien Jaeger, Patrick Carribault, Marc Pérache |
Euro-Par | 4 |
| 2019 | Multi-valued Expression Analysis for Collective Checking
Pierre Huchant, Emmanuelle Saillard, Denis Barthou, Patrick Carribault |
Euro-Par | 4 |
| 2019 | Mixing ranks, tasks, progress and nonblocking collectivesabstractSince the beginning, MPI has defined the rank as an implicit attribute associated with the MPI process' environment. In particular, each MPI process generally runs inside a given UNIX process and is associated with a fixed identifier in its WORLD communicator. However, this state of things is about to change with the rise of new abstractions such as MPI Sessions. In this paper, we propose to outline how such evolution could enable optimizations which were previously linked to specific MPI runtimes executing MPI processes in shared memory (e.g. thread-based MPI). By implementing runtime-level work-sharing through what we define as MPI tasks, enabling the ability to progress indifferently from stream context we show that there is potential for improved asynchronous progress. In the absence of a Session implementation, this assumption is validated in the context of a thread-based MPI where nonblocking Collective (NBC) were implemented on top of Extended Generic Requests progressed by any rank on the node thanks to an MPI extension enabling threads to dynamically share their MPI context. Jean-Baptiste Besnard, Julien Jaeger, Allen D. Malony, Sameer Shende, Hugo Taboada, Marc Pérache, Patrick Carribault |
EuroMPI | 7 |
| 2019 | Checkpoint/restart approaches for a thread-based MPI runtime
Julien Adam, Maxime Kermarquer, Jean-Baptiste Besnard, Leonardo Arturo Bautista-Gomez, Marc Pérache, Patrick Carribault, Julien Jaeger, Allen D. Malony, Sameer Shende |
Parallel Comput. | 6 |
| 2018 | Efficient Communication/Computation Overlap with MPI+OpenMP Runtimes Collaboration
Marc Sergent, Mario Dagrada, Patrick Carribault, Julien Jaeger, Marc Pérache, Guillaume Papauré |
Euro-Par | 3 |
| 2018 | Transparent High-Speed Network Checkpoint/Restart in MPIabstractFault-tolerance has always been an important topic when it comes to running massively parallel programs at scale. Statistically, hardware and software failures are expected to occur more often on systems gathering millions of computing units. Moreover, the larger jobs are, the more computing hours would be wasted by a crash. In this paper, we describe the work done in our MPI runtime to enable transparent checkpointing mechanism. Unlike the MPI 4.0 User-Level Failure Mitigation (ULFM) interface, our work targets solely Checkpoint/Restart (C/R) and ignores wider features such as resiliency. We show how existing transparent checkpointing methods can be practically applied to MPI implementations given a sufficient collaboration from the MPI runtime. Our C/R technique is then measured on MPI benchmarks such as IMB and Lulesh relying on Infiniband high-speed network, demonstrating that the chosen approach is sufficiently general and that performance is mostly preserved. We argue that enabling fault-tolerance without any modification inside target MPI applications is possible, and show how it could be the first step for more integrated resiliency combined with failure mitigation like ULFM. Julien Adam, Jean-Baptiste Besnard, Allen D. Malony, Sameer Shende, Marc Pérache, Patrick Carribault, Julien Jaeger |
EuroMPI | 6 |
| 2017 | Resource-Management Study in HPC Runtime-Stacking ContextabstractWith the advent of multicore and manycore processors as building blocks of HPC supercomputers, many applications shift from relying solely on a distributed programming model (e.g., MPI) to mixing distributed and shared-memory models (e.g., MPI+OpenMP), to better exploit shared-memory communications and reduce the overall memory footprint. One side effect of this programming approach is runtime stacking: mixing multiple models involve various runtime libraries to be alive at the same time and to share the underlying computing resources. This paper explores different configurations where this stacking may appear and introduces algorithms to detect the misuse of compute resources when running a hybrid parallel application. We have implemented our algorithms inside a dynamic tool that monitors applications and outputs resource usage to the user. We validated this tool on applications from CORAL benchmarks. This leads to relevant information which can be used to improve runtime placement, and to an average overhead lower than 1% of total execution time. Arthur Loussert, Benoit Welterlen, Patrick Carribault, Julien Jaeger, Marc Pérache, Raymond Namyst |
SBAC-PAD | 3 |
| 2016 | Introducing Task-Containers as an Alternative to Runtime-StackingabstractThe advent of many-core architectures poses new challenges to the MPI programming model which has been designed for distributed memory message passing. It is now clear that MPI will have to evolve in order to exploit shared-memory parallelism, either by collaborating with other programming models (MPI+X) or by introducing new shared-memory approaches. This paper considers extensions to C and C++ to make it possible for MPI Processes to run into threads. More generally, a thread-local storage (TLS) library is developed to simplify the collocation of arbitrary tasks and services in a shared-memory context called a task-container. The paper discusses how such containers simplify model and service mixing at the OS process level, eventually easing the collocation of arbitrary tasks with MPI processes in a runtime agnostic fashion, opening alternatives to runtime stacking. Jean-Baptiste Besnard, Julien Adam, Sameer Shende, Marc Pérache, Patrick Carribault, Julien Jaeger |
EuroMPI | 5 |
| 2015 | MPI Thread-Level Checking for MPI+OpenMP Applications
Emmanuelle Saillard, Patrick Carribault, Denis Barthou |
Euro-Par | 2 |
| 2015 | Static/Dynamic validation of MPI collective communications in multi-threaded contextabstractScientific applications mainly rely on the MPI parallel programming model to reach high performance on supercomputers. The advent of manycore architectures (larger number of cores and lower amount of memory per core) leads to mix MPI with a thread-based model like OpenMP. But integrating two different programming models inside the same application can be tricky and generate complex bugs. Thus, the correctness of hybrid programs requires a special care regarding MPI calls location. For example, identical MPI collective operations cannot be performed by multiple non-synchronized threads. To tackle this issue, this paper proposes a static analysis and a reduced dynamic instrumentation to detect bugs related to misuse of MPI collective operations inside or outside threaded regions. This work extends PARCOACH designed for MPI-only applications and keeps the compatibility with these algorithms. We validated our method on multiple hybrid benchmarks and applications with a low overhead. Emmanuelle Saillard, Patrick Carribault, Denis Barthou |
PPoPP | 2 |
| 2015 | An MPI Halo-Cell Implementation for Zero-Copy AbstractionabstractIn the race for Exascale, the advent of many-core processors will bring a shift in parallel computing architectures to systems of much higher concurrency, but with a relatively smaller memory per thread. This shift raises concerns for the adaptability of HPC software, for the current generation to the brave new world. In this paper, we study domain splitting on an increasing number of memory areas as an example problem where negative performance impact on computation could arise. We identify the specific parameters that drive scalability for this problem, and then model the halo-cell ratio on common mesh topologies to study the memory and communication implications. Such analysis argues for the use of shared-memory parallelism, such as with OpenMP, to address the performance problems that could occur. In contrast, we propose an original solution based entirely on MPI programming semantics, while providing the performance advantages of hybrid parallel programming. Our solution transparently replaces halo-cells transfers with pointer exchanges when MPI tasks are running on the same node, effectively removing memory copies. The results we present demonstrate gains in terms of memory and computation time on Xeon Phi (compared to OpenMP-only and MPI-only) using a representative domain decomposition benchmark. Jean-Baptiste Besnard, Allen D. Malony, Sameer Shende, Marc Pérache, Patrick Carribault, Julien Jaeger |
EuroMPI | 5 |
| 2015 | Correctness Analysis of MPI-3 Non-Blocking Communications in PARCOACHabstractMPI-3 provide functions for non-blocking collectives. To help programmers introduce non-blocking collectives to existing MPI programs, we improve the PARCOACH tool for checking correctness of MPI call sequences. These enhancements focus on correct call sequences of all flavor of collective calls, and on the presence of completion calls for all non-blocking communications. The evaluation shows an overhead under 10% of original compilation time. Julien Jaeger, Emmanuelle Saillard, Patrick Carribault, Denis Barthou |
EuroMPI | 3 |
| 2015 | Fine-grain data management directory for OpenMP 4.0 and OpenACCabstractSummary Today's trend to use accelerators in heterogeneous systems forces a paradigm shift in programming models. The use of low‐level APIs for accelerator programming is tedious and not intuitive for casual programmers. To tackle this problem, recent approaches focused on high‐level directive‐based models, with a standardization effort made with OpenACC and the directives for accelerator in the latest OpenMP 4.0 release. The pragmas for data management automatically handle data exchange between the host and the device. To keep the runtime simple and efficient, severe restrictions hinder the use of these pragmas. To address this issue, we propose the design for a directory, along with a reduced runtime application binary interface, to handle correctly data management in these standards. A few improvements to our directory allow a more flexible use of data management pragmas, with negligible overhead. Our design fits a multi‐accelerator system. Copyright © 2014 John Wiley & Sons, Ltd. Julien Jaeger, Patrick Carribault, Marc Pérache |
Concurr. Comput. Pract. Exp. | 2 |
| 2014 | Evaluation of OpenMP Task Scheduling Algorithms for Large NUMA Architectures
Jérôme Clet-Ortega, Patrick Carribault, Marc Pérache |
Euro-Par | 2 |
| 2013 | Combining static and dynamic validation of MPI collective communicationsabstractCollective MPI communications have to be executed in the same order by all processes in their communicator and the same number of times, otherwise a deadlock occurs. As soon as the control-flow involving these collective operations becomes more complex, in particular including conditionals on process ranks, ensuring the correction of such code is error-prone. We propose in this paper a static analysis to detect when such situation occurs, combined with a code transformation that prevents from deadlocking. We show on several benchmarks the small impact on performance and the ease of integration of our techniques in the development process. Emmanuelle Saillard, Patrick Carribault, Denis Barthou |
EuroMPI | 2 |
| 2012 | Hierarchical Local Storage: Exploiting Flexible User-Data Sharing Between MPI TasksabstractWith the advent of the multicore era, the number of cores per computational node is increasing faster than the amount of memory. This diminishing memory to core ratio sometimes even prevents pure MPI applications to benefit from all cores available on each node. A possible solution is to add a shared memory programming model like Open MP inside the application to share variables between Open MP threads that would otherwise be duplicated for each MPI task. Going to hybrid can thus improve the overall memory consumption, but may be a tedious task on large applications. To allow this data sharing without the overhead of mixing multiple programming models, we propose an MPI extension called Hierarchical Local Storage (HLS) that allows application developers to share common variables between MPI tasks on the same node. HLS is designed as a set of directives that preserve the original parallel semantics of the code and are compatible with C, C++ and Fortran languages and the Open MP programming model. This new mechanism is implemented inside a state-of-the-art MPI 1.3 compliant runtime called MPC. Experiments show that the HLS mechanism can effectively reduce memory consumption of HPC applications. Moreover, by reducing data duplication in the shared cache of modern multicores, the HLS mechanism can also improve performances of memory intensive applications. Marc Tchiboukdjian, Patrick Carribault, Marc Pérache |
IPDPS | 2 |
| 2012 | Improving MPI Communication Overlap with Collaborative Polling
Sylvain Didelot, Patrick Carribault, Marc Pérache, William Jalby |
EuroMPI | 2 |
| 2008 | Scheduling strategies for optimistic parallel execution of irregular programsabstractRecent application studies have shown that many irregular applications have a generalized data parallelism that manifests itself as iterative computations over worklists of different kinds. In general, there are complex dependencies between iterations. These dependencies cannot be elucidated statically because they depend on the inputs to the program; thus, optimistic parallel execution is the only tractable approach to parallelizing these applications. Milind Kulkarni 0001, Patrick Carribault, Keshav Pingali, Ganesh Ramanarayanan, Bruce Walter, Kavita Bala, L. Paul Chew |
SPAA | 2 |
| 2007 | Loop Optimization using Hierarchical Compilation and Kernel DecompositionabstractThe increasing complexity of hardware features for recent processors makes high performance code generation very challenging. In particular, several optimization targets have to be pursued simultaneously (minimizing L1/L2/L3/TLB misses and maximizing instruction level parallelism). Very often, these optimization goals impose different and contradictory constraints on the transformations to be applied. We propose a new hierarchical compilation approach for the generation of high performance code relying on the use of state-of-the-art compilers. This approach is not application-dependent and do not require any assembly hand-coding. It relies on the decomposition of the original loop nest into simpler kernels, typically 1D to 2D loops, much simpler to optimize. We successfully applied this approach to optimize dense matrix muliply primitives (not only for the square case but to the more general rectangular cases) and convolution. The performance of the optimized codes on Itanium 2 and Pentium 4 architectures outperforms ATLAS and in most cases, matches hand-tuned vendor libraries (e.g. MKL) Denis Barthou, Sébastien Donadio, Patrick Carribault, Alexandre Duchateau, William Jalby |
CGO | 3 |
| 2005 | Collisions of SHA-0 and Reduced SHA-1
Eli Biham, Rafi Chen, Antoine Joux, Patrick Carribault, Christophe Lemuet, William Jalby |
EUROCRYPT | 4 |
| 2004 | Applications of storage mapping optimization to register promotionabstractStorage mapping optimization is a flexible approach to folding array dimensions in numerical codes. It is designed to reduce the memory footprint after a wide spectrum of loop transformations, whether based on uniform dependence vectors or more expressive polyhedral abstractions. Conversely, few loop transformations have been proposed to facilitate register promotion, namely loop fusion, unroll-and-jam or tiling. Building on array data-flow analysis and expansion, we extend storage mapping optimization to improve opportunities for register promotion.Our work is motivated by the empirical study of a computational biology benchmark, the approximate string matching algorithm BPR from NR-grep, on a wide issue micro-architecture. Our experiments confirm the major benefit of register tiling (even on non-numerical benchmarks) but also shed the light on two novel issues: prior array expansion may be necessary to enable loop transformations that finally authorize profitable register promotion, and more advanced scheduling techniques (beyond tiling and unroll-and-jam) may significantly improve performance in fine-tuning register usage and instruction-level parallelism. Patrick Carribault, Albert Cohen 0001 |
ICS | 1 |