Raymond Namyst

dblp:75/4443 · DBLP profile ↗
← Back
38ranked-venue papers
0as first author
4since 2021 · last 2023
0000-0001-7734-1258ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 34 · 4 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2023 Programming heterogeneous architectures using hierarchical tasks
abstract
Summary Task‐based systems have become popular due to their ability to utilize the computational power of complex heterogeneous systems. A typical programming model used is the Sequential Task Flow (STF) model, which unfortunately only supports static task graphs. This can result in submission overhead and a static task graph that is not well‐suited for execution on heterogeneous systems. A common approach is to find a balance between the granularity needed for accelerator devices and the granularity required by CPU cores to achieve optimal performance. To address these issues, we have extended the STF model in the StarPU runtime system by introducing the concept of hierarchical tasks . This allows for a more dynamic task graph and, when combined with an automatic data manager, it is possible to adjust granularity at runtime to best match the targeted computing resource. That data manager makes it possible to switch between various data layout without programmer input and allows us to enforce the correctness of the DAG as hierarchical tasks alter it during runtime. Additionally, submission overhead is reduced by using large‐grain hierarchical tasks, as the submission process can now be done in parallel. We have shown that the hierarchical task model is correct and have conducted an early evaluation on shared memory heterogeneous systems using the Chameleon dense linear algebra library.
Mathieu Faverge, Nathalie Furmento, Abdou Guermouche, Gwenolé Lucas, Raymond Namyst, Samuel Thibault, Pierre-André Wacrenier
Concurr. Comput. Pract. Exp.5
2022 Exploring Scheduling Algorithms for Parallel Task Graphs: A Modern Game Engine Case Study
Mustapha Regragui, Baptiste Coye, Laércio Lima Pilla, Raymond Namyst, Denis Barthou
Euro-Par4
2021 TEXTAROSSA: Towards EXtreme scale Technologies and Accelerators for euROhpc hw/Sw Supercomputing Applications for exascale
abstract
To achieve high performance and high energy efficiency on near-future exascale computing systems, three key technology gaps needs to be bridged. These gaps include: energy efficiency and thermal control; extreme computation efficiency via HW acceleration and new arithmetics; methods and tools for seamless integration of reconfigurable accelerators in heterogeneous HPC multi-node platforms. TEXTAROSSA aims at tackling this gap through a co-design approach to heterogeneous HPC solutions, supported by the integration and extension of HW and SW IPs, programming models and tools derived from European research.
Giovanni Agosta, Daniele Cattaneo 0002, William Fornaciari, Andrea Galimberti, Giuseppe Massari, Federico Reghenzani, Federico Terraneo, Davide Zoni, Carlo Brandolese, Massimo Celino, Francesco Iannone, Paolo Palazzari, Giuseppe Zummo, Massimo Bernaschi, Pasqua D'Ambra, Sergio Saponara, Marco Danelutto, Massimo Torquati, Marco Aldinucci, Yasir Arfat, Barbara Cantalupo, Iacopo Colonnelli, Roberto Esposito, Alberto Riccardo Martinelli, Gianluca Mittone, Olivier Beaumont, Bérenger Bramas, Lionel Eyraud-Dubois, Brice Goglin, Abdou Guermouche, Raymond Namyst, Samuel Thibault, Antonio Filgueras, Miquel Vidal, Carlos Álvarez 0001, Xavier Martorell, Ariel Oleksiak, Michal Kulczewski, Alessandro Lonardo, Piero Vicini, Francesca Lo Cicero, Francesco Simula, Andrea Biagioni, Paolo Cretaro, Ottorino Frezza, Pier Stanislao Paolucci, Matteo Turisini, Francesco Giacomini, Tommaso Boccali, Simone Montangero, Roberto Ammendola
DSD31
2021 EasyPAP: A framework for learning parallel programming
Alice Lasserre, Raymond Namyst, Pierre-André Wacrenier
J. Parallel Distributed Comput.2
2019 Resource aggregation for task-based Cholesky Factorization on top of modern architectures
Terry Cojean, Abdou Guermouche, Andra Hugo, Raymond Namyst, Pierre-André Wacrenier
Parallel Comput.4
2018 Combining Task-based Parallelism and Adaptive Mesh Refinement Techniques in Molecular Dynamics Simulations
abstract
Modern parallel architectures require applications to generate massive parallelism so as to feed their large number of cores and their wide vector units. We revisit the extensively studied classical Molecular Dynamics N-body problem in the light of these hardware constraints. We use Adaptive Mesh Refinement techniques to store particles in memory, and to optimize the force computation loop using multi-threading and vectorization-friendly data structures. Our design is guided by the need for load balancing and adaptivity raised by highly dynamic particle sets, as typically observed in simulations of strong shocks resulting in material micro-jetting. We analyze performance results on several simulation scenarios, over nodes equipped by Intel Xeon Phi Knights Landing (KNL) or Intel Xeon Skylake (SKL) processors. Performance obtained with our OpenMP implementation outperforms state-of-the-art implementations (LAMMPS) on both steady and micro-jetting particles simulations. In the latter case, our implementation is 4.7 times faster on KNL, and 2 times faster on SKL.
Raphaël Prat, Laurent Colombet, Raymond Namyst
ICPP3
2017 Resource-Management Study in HPC Runtime-Stacking Context
abstract
With the advent of multicore and manycore processors as building blocks of HPC supercomputers, many applications shift from relying solely on a distributed programming model (e.g., MPI) to mixing distributed and shared-memory models (e.g., MPI+OpenMP), to better exploit shared-memory communications and reduce the overall memory footprint. One side effect of this programming approach is runtime stacking: mixing multiple models involve various runtime libraries to be alive at the same time and to share the underlying computing resources. This paper explores different configurations where this stacking may appear and introduces algorithms to detect the misuse of compute resources when running a hybrid parallel application. We have implemented our algorithms inside a dynamic tool that monitors applications and outputs resource usage to the user. We validated this tool on applications from CORAL benchmarks. This leads to relevant information which can be used to improve runtime placement, and to an average overhead lower than 1% of total execution time.
Arthur Loussert, Benoit Welterlen, Patrick Carribault, Julien Jaeger, Marc Pérache, Raymond Namyst
SBAC-PAD6
2015 Automatic OpenCL Code Generation for Multi-device Heterogeneous Architectures
abstract
Using multiple accelerators, such as GPUs or Xeon Phis, is attractive to improve the performance of large data parallel applications and to increase the size of their workloads. However, writing an application for multiple accelerators remains today challenging because going from a single accelerator to multiple ones indeed requires to deal with potentially non-uniform domain decomposition, inter-accelerator data movements, and dynamic load balancing. Writing such code manually is time consuming and error-prone. In this paper, we propose a new programming tool called STEPOCL along with a new domain specific language designed to simplify the development of an application for multiple accelerators. We evaluate both the performance and the usefulness of STEPOCL with three applications and show that: (i) the performance of an application written with STEPOCL scales linearly with the number of accelerators, (ii) the performance of an application written using STEPOCL competes with a handwritten version, (iii) larger workloads run on multiple devices that do not fit in the memory of a single device, (iv) thanks to STEPOCL, the number of lines of code required to write an application for multiple accelerators is roughly divided by ten.
Pei Li 0003, Elisabeth Brunet, François Trahay, Christian Parrot, Gaël Thomas 0001, Raymond Namyst
ICPP6
2014 Toward OpenCL Automatic Multi-Device Support
Sylvain Henry, Alexandre Denis 0001, Denis Barthou, Marie Christine Counilh, Raymond Namyst
Euro-Par5
2014 A Runtime Approach to Dynamic Resource Allocation for Sparse Direct Solvers
abstract
To face the advent of multicore processors and the ever increasing complexity of hardware architectures, programming models based on DAG-of-tasks parallelism regained popularity in the high performance, scientific computing community. In this context, enabling HPC applications to perform efficiently when dealing with graphs of parallel tasks that could potentially run simultaneously is a great challenge. Even if a uniform runtime system is used underneath, scheduling multiple parallel tasks over the same set of hardware resources introduces many issues, such as undesirable cache flushes or memory bus contention. In this paper, we show how runtime system-based scheduling contexts can be used to dynamically enforce locality of parallel tasks on multicore machines. We extend an existing generic sparse direct solver to use our mechanism and introduce a new decomposition method based on proportional mapping that is used to build the scheduling contexts. We propose a runtime-level dynamic context management policy to cope with the very irregular behaviour of the application. A detailed performance analysis shows significant performance improvements of the solver over various multicore hardware.
Andra Hugo, Abdou Guermouche, Pierre-André Wacrenier, Raymond Namyst
ICPP4
2013 Adaptive Task Size Control on High Level Programming for GPU/CPU Work Sharing
Tetsuya Odajima, Taisuke Boku, Mitsuhisa Sato, Toshihiro Hanawa, Yuetsu Kodama, Raymond Namyst, Samuel Thibault, Olivier Aumage
ICA3PP (2)6
2012 Programmability and performance portability aspects of heterogeneous multi-/manycore systems
abstract
We discuss three complementary approaches that can provide both portability and an increased level of abstraction for the programming of heterogeneous multicore systems. Together, these approaches also support performance portability, as currently investigated in the EU FP7 project PEPPHER. In particular, we consider (1) a library-based approach, here represented by the integration of the SkePU C++ skeleton programming library with the StarPU runtime system for dynamic scheduling and dynamic selection of suitable execution units for parallel tasks; (2) a language-based approach, here represented by the Offload-C++ high-level language extensions and Offload compiler to generate platform-specific code; and (3) a component-based approach, specifically the PEPPHER component system for annotating user-level application components with performance metadata, thereby preparing them for performance-aware composition. We discuss the strengths and weaknesses of these approaches and show how they could complement each other in an integrational programming framework for heterogeneous multicore systems.
Christoph W. Kessler, Usman Dastgeer, Samuel Thibault, Raymond Namyst, Andrew Richards, Uwe Dolinsky, Siegfried Benkner, Jesper Larsson Träff, Sabri Pllana
DATE4
2012 High-Level Support for Pipeline Parallelism on Many-Core Architectures
Siegfried Benkner, Enes Bajrovic, Erich Marth, Martin Sandrieser, Raymond Namyst, Samuel Thibault
Euro-Par5
2012 StarPU-MPI: Task Programming over Clusters of Machines Enhanced with Accelerators
Cédric Augonnet, Olivier Aumage, Nathalie Furmento, Raymond Namyst, Samuel Thibault
EuroMPI4
2011 EZTrace: A Generic Framework for Performance Analysis
abstract
Modern supercomputers with multi-core nodes enhanced by accelerators, as well as hybrid programming models introduce more complexity in modern applications. Exploiting efficiently all the resources requires a complex analysis of the performance of applications in order to detect time-consuming sections. We present eztrace, a generic trace generation framework that aims at providing a simple way to analyze applications. eztrace is based on plugins that allow it to trace different programming models such as MPI, pthread or OpenMP as well as user-defined libraries or applications. eztrace uses two steps: one to collect the basic information during execution and one post-mortem analysis. This permits tracing the execution of applications with low overhead while allowing to refine the analysis after the execution. We also present a script language for eztrace that gives the user the opportunity to easily define the functions to instrument without modifying the source code of the application.
François Trahay, François Rué, Mathieu Faverge, Yutaka Ishikawa, Raymond Namyst, Jack J. Dongarra
CCGRID5
2011 A Sampling-Based Approach for Communication Libraries Auto-Tuning
abstract
Communication performance is a critical issue in HPC applications, and many solutions have been proposed on the literature (algorithmic, protocols, etc.) In the meantime, computing nodes become massively multicore, leading to a real imbalance between the number of communication sources and the number of physical communication resources. Thus it is now mandatory to share network boards between computation flows, and to take this sharing into account while performing communication optimizations. In previous papers, we have proposed a model and a frame work for on-the-fly optimizations of multiplexed concurrent communication flows, and implemented this model in the NEWMADELEINE communication library. This library features optimization strategies able for example to aggregate several messages to reduce the number of packets emitted on the network, or to split messages to use several NICs at the same time. In this paper, we study the tuning of these dynamic optimization strategies. We show that some parameters and thresholds (rendezvous threshold, aggregation packet size) depend on the actual hardware, both host and NICs. We propose and implement a method based on sampling of the actual hardware to auto-tune our strategies. Moreover, we show that multi-rail can greatly benefit from performance predictions. We propose an approach for multi-rail that dynamically balance the data between NICs using predictions based on sampling.
Elisabeth Brunet, François Trahay, Alexandre Denis 0001, Raymond Namyst
CLUSTER4
2011 StarPU: a unified platform for task scheduling on heterogeneous multicore architectures
abstract
Abstract In the field of HPC, the current hardware trend is to design multiprocessor architectures featuring heterogeneous technologies such as specialized coprocessors (e.g. Cell/BE) or data‐parallel accelerators (e.g. GPUs). Approaching the theoretical performance of these architectures is a complex issue. Indeed, substantial efforts have already been devoted to efficiently offload parts of the computations. However, designing an execution model that unifies all computing units and associated embedded memory remains a main challenge. We therefore designed StarPU, an original runtime system providing a high‐level, unified execution model tightly coupled with an expressive data management library. The main goal of StarPU is to provide numerical kernel designers with a convenient way to generate parallel tasks over heterogeneous hardware on the one hand, and easily develop and tune powerful scheduling algorithms on the other hand. We have developed several strategies that can be selected seamlessly at run‐time, and we have analyzed their efficiency on several algorithms running simultaneously over multiple cores and a GPU. In addition to substantial improvements regarding execution times, we have obtained consistent superlinear parallelism by actually exploiting the heterogeneous nature of the machine. We eventually show that our dynamic approach competes with the highly optimized MAGMA library and overcomes the limitations of the corresponding static scheduling in a portable way. Copyright © 2010 John Wiley & Sons, Ltd.
Cédric Augonnet, Samuel Thibault, Raymond Namyst, Pierre-André Wacrenier
Concurr. Comput. Pract. Exp.3
2010 Data-Aware Task Scheduling on Multi-accelerator Based Platforms
abstract
To fully tap into the potential of heterogeneous machines composed of multicore processors and multiple accelerators, simple offloading approaches in which the main trunk of the application runs on regular cores while only specific parts are offloaded on accelerators are not sufficient. The real challenge is to build systems where the application would permanently spread across the entire machine, that is, where parallel tasks would be dynamically scheduled over the full set of available processing units. To face this challenge, we previously proposed StarPU, a runtime system capable of scheduling tasks over multicore machines equipped with GPU accelerators. StarPU uses a software virtual shared memory (VSM) that provides a highlevel programming interface and automates data transfers between processing units so as to enable a dynamic scheduling of tasks. We now present how we have extended StarPU to minimize the cost of transfers between processing units in order to efficiently cope with multi-GPU hardware configurations. To this end, our runtime system implements data prefetching based on asynchronous data transfers, and uses data transfer cost prediction to influence the decisions taken by the task scheduler. We demonstrate the relevance of our approach by benchmarking two parallel numerical algorithms using our runtime system. We obtain significant speedups and high efficiency over multicore machines equipped with multiple accelerators. We also evaluate the behaviour of these applications over clusters featuring multiple GPUs per node, showing how our runtime system can combine with MPI.
Cédric Augonnet, Jérôme Clet-Ortega, Samuel Thibault, Raymond Namyst
ICPADS4
2010 Structuring the execution of OpenMP applications for multicore architectures
abstract
The now commonplace multi-core chips have introduced, by design, a deep hierarchy of memory and cache banks within parallel computers as a tradeoff between the user friendliness of shared memory on the one side, and memory access scalability and efficiency on the other side. However, to get high performance out of such machines requires a dynamic mapping of application tasks and data onto the underlying architecture. Moreover, depending on the application behavior, this mapping should favor cache affinity, memory bandwidth, computation synchrony, or a combination of these. The great challenge is then to perform this hardware-dependent mapping in a portable, abstract way. To meet this need, we propose a new, hierarchical approach to the execution of OpenMP threads onto multicore machines. Our ForestGOMP runtime system dynamically generates structured trees out of OpenMP programs. It collects relationship information about threads and data as well. This information is used together with scheduling hints and hardware counter feedback by the scheduler to select the most appropriate threads and data distribution. ForestGOMP features a highlevel platform for developing and tuning portable threads schedulers. We present several applications for which we developed specific scheduling policies that achieve excellent speedups on 16-core machines.
François Broquedis, Olivier Aumage, Brice Goglin, Samuel Thibault, Pierre-André Wacrenier, Raymond Namyst
IPDPS6
2010 hwloc: A Generic Framework for Managing Hardware Affinities in HPC Applications
abstract
The increasing numbers of cores, shared caches and memory nodes within machines introduces a complex hardware topology. High-performance computing applications now have to carefully adapt their placement and behavior according to the underlying hierarchy of hardware resources and their software affinities. We introduce the Hardware Locality (hwloc) software which gathers hardware information about processors, caches, memory nodes and more, and exposes it to applications and runtime systems in a abstracted and portable hierarchical manner. hwloc may significantly help performance by having runtime systems place their tasks or adapt their communication strategies depending on hardware affinities. We show that hwloc can already be used by popular high-performance OpenMP or MPI software. Indeed, scheduling OpenMP threads according to their affinities or placing MPI processes according to their communication patterns shows interesting performance improvement thanks to hwloc. An optimized MPI communication strategy may also be dynamically chosen according to the location of the communicating processes in the machine and its hardware characteristics.
François Broquedis, Jérôme Clet-Ortega, Stéphanie Moreaud, Nathalie Furmento, Brice Goglin, Guillaume Mercier, Samuel Thibault, Raymond Namyst
PDP8
2010 Adaptive MPI Multirail Tuning for Non-uniform Input/Output Access
Stéphanie Moreaud, Brice Goglin, Raymond Namyst
EuroMPI3
2009 StarPU: A Unified Platform for Task Scheduling on Heterogeneous Multicore Architectures
Cédric Augonnet, Samuel Thibault, Raymond Namyst, Pierre-André Wacrenier
Euro-Par3
2008 MPC: A Unified Parallel Runtime for Clusters of NUMA Machines
Marc Pérache, Hervé Jourdren, Raymond Namyst
Euro-Par3
2008 A multithreaded communication engine for multicore architectures
abstract
The current trend in clusters leads towards an increase of the number of cores per node. As a result, an increasing number of parallel applications is mixing message passing and multithreading as an attempt to better match the underlying architecture's structure. This naturally raises the problem of designing efficient, multithreaded implementations of MPL In this paper, we present the design of a multithreaded communication engine able to exploit idle cores to speed up communications in two ways: it can move CPU- intensive operations out of the critical path (e.g. PIO transfers off load), and is able to let rendezvous transfers progress asynchronously. We have implemented these methods in the PM2 software suite, evaluated their behavior in typical cases, and we have observed good performance results in overlapping communication and computation.
François Trahay, Elisabeth Brunet, Alexandre Denis 0001, Raymond Namyst
IPDPS4
2007 Building Portable Thread Schedulers for Hierarchical Multiprocessors: The BubbleSched Framework
Samuel Thibault, Raymond Namyst, Pierre-André Wacrenier
Euro-Par2
2007 NEW MADELEINE: a Fast Communication Scheduling Engine for High Performance Networks
abstract
Communication libraries have dramatically made progress over the fifteen years, pushed by the success of cluster architectures as the preferred platform for high performance distributed computing. However, many potential optimizations are left unexplored in the process of mapping application communication requests onto low level network commands. The fundamental cause of this situation is that the design of communication subsystems is mostly focused on reducing the latency by shortening the critical path. In this paper, we present a new communication scheduling engine which dynamically optimizes application requests in accordance with the NICs capabilities and activity. The optimizing code is generic and portable. The database of optimizing strategies may be dynamically extended.
Olivier Aumage, Elisabeth Brunet, Nathalie Furmento, Raymond Namyst
IPDPS4
2007 High-Performance Multi-Rail Support with the NEWMADELEINE Communication Library
abstract
This paper focuses on message transfers across multiple heterogeneous high-performance networks in the NEWMADELEINE communication library. NEWMADELEINE features a modular design that allows the user to easily implement load-balancing strategies efficiently exploiting the underlying network but without being aware of the low-level interface. Several strategies are studied and preliminary results are given. They show that performance of network transfers can be improved by using carefully designed strategies that take into account NIC activity.
Olivier Aumage, Elisabeth Brunet, Guillaume Mercier, Raymond Namyst
IPDPS4
2006 Short Paper : Dynamic Optimization of Communications over High Speed Networks
abstract
We present a new communication subsystem for high speed networks featuring an extendable packet optimization engine mixing several communication flows. Optimizations are parameterized by the capabilities of the underlying network drivers, and are triggered by the network cards when they become idle. The database of predefined strategies can be easily extended
Elisabeth Brunet, Olivier Aumage, Raymond Namyst
HPDC3
2005 An Efficient Multi-level Trace Toolkit for Multi-threaded Applications
Vincent Danjean, Raymond Namyst, Pierre-André Wacrenier
Euro-Par2
2003 Controlling Kernel Scheduling from User Space: An Approach to Enhancing Applications' Reactivity to I/O Events
Vincent Danjean, Raymond Namyst
HiPC2
2002 Improving Reactivity to I/O Events in Multithreaded Environments Using a Uniform, Scheduler-Centric API
Luc Bougé, Vincent Danjean, Raymond Namyst
Euro-Par3
2002 Madeleine II: a portable and efficient communication library for high-performance cluster computing
Olivier Aumage, Luc Bougé, Jean-François Méhaut, Raymond Namyst
Parallel Comput.4
2001 Efficient Inter-Device Data-Forwarding in the Madeleine Communication Library
abstract
Interconnecting multiple clusters with a high speed network to form a single heterogeneous architecture (i.e. a cluster of clusters) is currently a hot issue. Consequently, new runtime systems that are able to simultaneously deal with multiple high speed networks within the same application have to be designed. This paper presents how we did extend an existing multi-device communication library with fast internal data-forwarding capabilities on gateway nodes. On top of that, efficient high-level routing mechanisms can be implemented. Our approach is easily applicable to many network protocols and is completely independant from the application code. Efficiency is achieved by avoiding extra data copies when possible and by using pipelining techniques. The preliminary experiments show that the observed inter-cluster bandwidth is close to the one that can be delivered by the hardware.
Olivier Aumage, Lionel Eyraud-Dubois, Raymond Namyst
IPDPS3
2001 MPICH/Madeleine: a True Multi-Protocol MPI for High Performance Networks
abstract
This paper introduces a version of MPICH handling efficiently different networks simultaneously. The core of the implementation relies on a device called ch-mad which is based on a generic multiprotocol communication library called Madeleine. The performance achieved with tested networks such as Fast-Ethernet, Scalable Coherent Interface or Myrinet is very good. Indeed, this multi-protocol version of MPICH generally outperforms other free or commercial implementations of MPI.
Olivier Aumage, Guillaume Mercier, Raymond Namyst
IPDPS3
2001 The Hyperion system: Compiling multithreaded Java bytecode for distributed execution
Gabriel Antoniu, Luc Bougé, Philip J. Hatcher, Mark MacBeth, Keith McGuigan, Raymond Namyst
Parallel Comput.6
2000 Madeleine II: a Portable and Efficient Communication Library for High-Performance Cluster Computing
abstract
This paper introduces Madeleine II, a new adaptive and portable multi-protocol implementation of the Madeleine communication library. Madeleine II has the ability to control multiple network interfaces (BIP, SISCI, VIA) and multiple network adapters (Ethernet, Myrinet, SCI) within the same application session. Moreover it includes advanced mechanisms to dynamically select the most appropriate transfer method for a given network protocol according to various parameters such as data size or responsiveness user requirements. We report on performance measurements obtained using BIP/Myrinet and SISCI/SCI and we present preliminary results about our Nexus/Madeleine II and MPICH/Madeleine II ports.
Olivier Aumage, Luc Bougé, Alexandre Denis 0001, Jean-François Méhaut, Guillaume Mercier, Raymond Namyst, Loïc Prylli
CLUSTER6
2000 Compiling Multithreaded Java Bytecode for Distributed Execution (Distinguished Paper)
Gabriel Antoniu, Luc Bougé, Philip J. Hatcher, Mark MacBeth, Keith McGuigan, Raymond Namyst
Euro-Par6
1999 Technology transfer within the ProHPC TTN at ENS Lyon
Christophe Barberet, Lionel Brunie, Frédéric Desprez, Gilles Lebourgeois, Raymond Namyst, Yves Robert, Stéphane Ubéda, Karine Van Heumen
Future Gener. Comput. Syst.5