EDBT 2026 Demo / reviewers in the wild / expert
Pierre-André Wacrenier
dblp:33/4908
· DBLP profile ↗
15ranked-venue papers
0as first author
3since 2021 · last 2025
0000-0003-3351-844XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 3 since 2021Theory of computation · 4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Optimizing parallel heterogeneous system efficiency: Dynamic task graph adaptation with recursive tasks
Nathalie Furmento, Abdou Guermouche, Gwenolé Lucas, Thomas Morin, Samuel Thibault, Pierre-André Wacrenier |
J. Parallel Distributed Comput. | 6 |
| 2023 | Programming heterogeneous architectures using hierarchical tasksabstractSummary Task‐based systems have become popular due to their ability to utilize the computational power of complex heterogeneous systems. A typical programming model used is the Sequential Task Flow (STF) model, which unfortunately only supports static task graphs. This can result in submission overhead and a static task graph that is not well‐suited for execution on heterogeneous systems. A common approach is to find a balance between the granularity needed for accelerator devices and the granularity required by CPU cores to achieve optimal performance. To address these issues, we have extended the STF model in the StarPU runtime system by introducing the concept of hierarchical tasks . This allows for a more dynamic task graph and, when combined with an automatic data manager, it is possible to adjust granularity at runtime to best match the targeted computing resource. That data manager makes it possible to switch between various data layout without programmer input and allows us to enforce the correctness of the DAG as hierarchical tasks alter it during runtime. Additionally, submission overhead is reduced by using large‐grain hierarchical tasks, as the submission process can now be done in parallel. We have shown that the hierarchical task model is correct and have conducted an early evaluation on shared memory heterogeneous systems using the Chameleon dense linear algebra library. Mathieu Faverge, Nathalie Furmento, Abdou Guermouche, Gwenolé Lucas, Raymond Namyst, Samuel Thibault, Pierre-André Wacrenier |
Concurr. Comput. Pract. Exp. | 7 |
| 2021 | EasyPAP: A framework for learning parallel programming
Alice Lasserre, Raymond Namyst, Pierre-André Wacrenier |
J. Parallel Distributed Comput. | 3 |
| 2019 | Resource aggregation for task-based Cholesky Factorization on top of modern architectures
Terry Cojean, Abdou Guermouche, Andra Hugo, Raymond Namyst, Pierre-André Wacrenier |
Parallel Comput. | 5 |
| 2014 | A Runtime Approach to Dynamic Resource Allocation for Sparse Direct SolversabstractTo face the advent of multicore processors and the ever increasing complexity of hardware architectures, programming models based on DAG-of-tasks parallelism regained popularity in the high performance, scientific computing community. In this context, enabling HPC applications to perform efficiently when dealing with graphs of parallel tasks that could potentially run simultaneously is a great challenge. Even if a uniform runtime system is used underneath, scheduling multiple parallel tasks over the same set of hardware resources introduces many issues, such as undesirable cache flushes or memory bus contention. In this paper, we show how runtime system-based scheduling contexts can be used to dynamically enforce locality of parallel tasks on multicore machines. We extend an existing generic sparse direct solver to use our mechanism and introduce a new decomposition method based on proportional mapping that is used to build the scheduling contexts. We propose a runtime-level dynamic context management policy to cope with the very irregular behaviour of the application. A detailed performance analysis shows significant performance improvements of the solver over various multicore hardware. Andra Hugo, Abdou Guermouche, Pierre-André Wacrenier, Raymond Namyst |
ICPP | 3 |
| 2011 | StarPU: a unified platform for task scheduling on heterogeneous multicore architecturesabstractAbstract In the field of HPC, the current hardware trend is to design multiprocessor architectures featuring heterogeneous technologies such as specialized coprocessors (e.g. Cell/BE) or data‐parallel accelerators (e.g. GPUs). Approaching the theoretical performance of these architectures is a complex issue. Indeed, substantial efforts have already been devoted to efficiently offload parts of the computations. However, designing an execution model that unifies all computing units and associated embedded memory remains a main challenge. We therefore designed StarPU, an original runtime system providing a high‐level, unified execution model tightly coupled with an expressive data management library. The main goal of StarPU is to provide numerical kernel designers with a convenient way to generate parallel tasks over heterogeneous hardware on the one hand, and easily develop and tune powerful scheduling algorithms on the other hand. We have developed several strategies that can be selected seamlessly at run‐time, and we have analyzed their efficiency on several algorithms running simultaneously over multiple cores and a GPU. In addition to substantial improvements regarding execution times, we have obtained consistent superlinear parallelism by actually exploiting the heterogeneous nature of the machine. We eventually show that our dynamic approach competes with the highly optimized MAGMA library and overcomes the limitations of the corresponding static scheduling in a portable way. Copyright © 2010 John Wiley & Sons, Ltd. Cédric Augonnet, Samuel Thibault, Raymond Namyst, Pierre-André Wacrenier |
Concurr. Comput. Pract. Exp. | 4 |
| 2010 | Structuring the execution of OpenMP applications for multicore architecturesabstractThe now commonplace multi-core chips have introduced, by design, a deep hierarchy of memory and cache banks within parallel computers as a tradeoff between the user friendliness of shared memory on the one side, and memory access scalability and efficiency on the other side. However, to get high performance out of such machines requires a dynamic mapping of application tasks and data onto the underlying architecture. Moreover, depending on the application behavior, this mapping should favor cache affinity, memory bandwidth, computation synchrony, or a combination of these. The great challenge is then to perform this hardware-dependent mapping in a portable, abstract way. To meet this need, we propose a new, hierarchical approach to the execution of OpenMP threads onto multicore machines. Our ForestGOMP runtime system dynamically generates structured trees out of OpenMP programs. It collects relationship information about threads and data as well. This information is used together with scheduling hints and hardware counter feedback by the scheduler to select the most appropriate threads and data distribution. ForestGOMP features a highlevel platform for developing and tuning portable threads schedulers. We present several applications for which we developed specific scheduling policies that achieve excellent speedups on 16-core machines. François Broquedis, Olivier Aumage, Brice Goglin, Samuel Thibault, Pierre-André Wacrenier, Raymond Namyst |
IPDPS | 5 |
| 2009 | StarPU: A Unified Platform for Task Scheduling on Heterogeneous Multicore Architectures
Cédric Augonnet, Samuel Thibault, Raymond Namyst, Pierre-André Wacrenier |
Euro-Par | 4 |
| 2007 | Building Portable Thread Schedulers for Hierarchical Multiprocessors: The BubbleSched Framework
Samuel Thibault, Raymond Namyst, Pierre-André Wacrenier |
Euro-Par | 3 |
| 2005 | An Efficient Multi-level Trace Toolkit for Multi-threaded Applications
Vincent Danjean, Raymond Namyst, Pierre-André Wacrenier |
Euro-Par | 3 |
| 2001 | A Distributed Algorithm for Computing a Spanning Tree in Anonymous Tprime Graph
Yves Métivier, Mohamed Mosbah 0001, Pierre-André Wacrenier, Stefan Gruner |
OPODIS | 3 |
| 1997 | About the local detection of termination of local computations in graphs
Yves Métivier, Anca Muscholl, Pierre-André Wacrenier |
SIROCCO | 3 |
| 1995 | Computing the Closure of Sets of Words Under Partial Commutations
Yves Métivier, Gwénaël Richomme, Pierre-André Wacrenier |
ICALP | 3 |
| 1993 | On Regular Compatibility of Semi-Commutations
Edward Ochmanski, Pierre-André Wacrenier |
ICALP | 2 |
| 1991 | Composition of Two Semi Commutations
Yves Roos, Pierre-André Wacrenier |
MFCS | 2 |