Luc Giraud

dblp:10/2046 · DBLP profile ↗
← Back
20ranked-venue papers
3as first author
3since 2021 · last 2025
0000-0002-7062-7672ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 17 · 3 first-author · 3 since 2021Theory of computation · 2
YearPublicationVenuePosition
2025 Multifacets of lossy compression for scientific data in the Joint-Laboratory of Extreme Scale Computing
Franck Cappello, Mario C. Acosta, Emmanuel Agullo, Hartwig Anzt, Jon Calhoun 0001, Sheng Di, Luc Giraud, Thomas Grützmacher, Sian Jin, Kentaro Sano, Kento Sato, Amarjit Singh, Dingwen Tao, Jiannan Tian, Tomohiro Ueno, Robert Underwood, Frédéric Vivien, Xavier Yepes, Kazutomo Yoshii, Boyuan Zhang 0002
Future Gener. Comput. Syst.7
2023 Combining reduction with synchronization barrier on multi-core processors
abstract
Summary With the rise of multi‐core processors with a large number of cores, the need for shared memory reduction that performs efficiently on a large number of cores is more pressing. Efficient shared memory reduction on these multi‐core processors will help share memory programs be more efficient. In this article, we propose a reduction method combined with a barrier method that uses SIMD read/write instructions to combine barrier signaling and reduction value to minimize memory/cache traffic between cores, thereby reducing barrier latency. We compare different barriers and reduction methods on three multi‐core processors and show that the proposed combining barrier/reduction methods are 4 and 3.5 times faster than respectively GCC 11.1 and Intel 21.2 OpenMP 4.5 reduction.
Aboul-Karim Mohamed El Maarouf, Luc Giraud, Abdou Guermouche, Thomas Guignon
Concurr. Comput. Pract. Exp.2
2021 Guest editorial: Virtual special issue on parallel matrix algorithms and applications (PMAA'18)
Olaf Schenk, Peter Arbenz, Luc Giraud, Wim Vanroose
Parallel Comput.3
2019 Energy Analysis of a Solver Stack for Frequency-Domain Electromagnetics
abstract
High-performance computing (HPC) aims at developing models and simulations for applications in numerous scientific fields. Yet, the energy consumption of these HPC facilities currently limits their size and performance, and consequently the size of the tackled problems. The complexity of the HPC software stacks and their various optimizations makes it difficult to finely understand the energy consumption of scientific applications. To highlight this difficulty on a concrete use-case, we perform an energy and power analysis of a software stack for the simulation of frequency-domain electromagnetic wave propagation. This solver stack combines a high order finite element discretization framework of the system of three-dimensional frequency -domain Maxwell equations with an algebraic hybrid iterative-direct sparse linear solver. This analysis is conducted on the KNL-based PRACE-PCP system. Our results illustrate the difficulty in predicting how to trade energy and runtime.
Emmanuel Agullo, Luc Giraud, Stéphane Lanteri, Gilles Marait, Anne-Cécile Orgerie, Louis Poirel
PDP2
2018 Special issue on parallel matrix algorithms and applications (PMAA'16)
Emmanuel Agullo, Peter Arbenz, Luc Giraud, Olaf Schenk
Parallel Comput.3
2015 On the Resilience of Parallel Sparse Hybrid Solvers
abstract
As the computational power of high performance computing (HPC) systems continues to increase by using a huge number of CPU cores or specialized processing units, extreme-scale applications are increasingly prone to faults. Consequently, the HPC community has proposed many contributions to design resilient HPC applications. These contributions may be system-oriented, theoretical or numerical. In this study we consider an actual fully-featured parallel sparse hybrid (direct/iterative) linear solver, MaPHyS, and we propose numerical remedies to design a resilient version of the solver. The solver being hybrid, we focus in this study on the iterative solution step, which is often the dominant step in practice. We furthermore assume that a separate mechanism ensures fault detection and that a system layer provides support for setting back the environment (processes, ...) in a running state. The present manuscript therefore focuses on (and only on) strategies for recovering lost data after the faulthas been detected (a separate concern beyond the scope of this study), once the system is restored (another separate concern not studied here). The numerical remedies we propose are twofold. Whenever possible, we exploit the natural data redundancy between processes from the solver to perform exact recovery through clever copies over processes. Otherwise, data that has been lost and no longer available on any process is recovered through a so-called interpolation-restart mechanism. This mechanism is derived from a previous work by carefully taking into account the properties of the target hybrid solver. These numerical remedies have been implemented in the MaPhys parallel solver so that we can assess their efficiency on a large number of processing units (up to 12,288 CPU cores) for solving large-scale real-life problems.
Emmanuel Agullo, Luc Giraud, Mawussi Zounon
HiPC2
2011 Introduction
Martin Berzins, Daniela di Serafino, Martin J. Gander, Luc Giraud
Euro-Par (2)4
2010 Using multiple levels of parallelism to enhance the performance of domain decomposition solvers
Luc Giraud, Azzam Haidar, S. Pralet
Parallel Comput.1
2008 Parallel scalability study of hybrid preconditioners in three dimensions
Luc Giraud, Azzam Haidar, Layne T. Watson
Parallel Comput.1
2008 Algorithm 881: A Set of Flexible GMRES Routines for Real and Complex Arithmetics on High-Performance Computers
abstract
In this article we describe our implementations of the FGMRES algorithm for both real and complex, single and double precision arithmetics suitable for serial, shared-memory, and distributed-memory computers. For the sake of portability, simplicity, flexibility, and efficiency, the FGMRES solvers have been implemented in Fortran 77 using the reverse communication mechanism for the matrix-vector product, the preconditioning, and the dot-product computations. For distributed-memory computation, several orthogonalization procedures have been implemented to reduce the cost of the dot-product calculation, which is a well-known bottleneck of efficiency for Krylov methods. Furthermore, either implicit or explicit calculation of the residual at restart is possible depending on the actual cost of the matrix-vector product. Finally, the implemented stopping criterion is based on a normwise backward error.
Valérie Frayssé, Luc Giraud, Serge Gratton
ACM Trans. Math. Softw.2
2007 A distributed packed storage for large dense parallel in-core calculations
abstract
Abstract In this paper we propose a distributed packed storage format that exploits the symmetry or the triangular structure of a dense matrix. This format stores only half of the matrix while maintaining most of the efficiency compared with a full storage for a wide range of operations. This work has been motivated by the fact that, in contrast to sequential linear algebra libraries (e.g. LAPACK), there is no routine or format that handles packed matrices in the currently available parallel distributed libraries. The proposed algorithms exclusively use the existing ScaLAPACK computational kernels, which proves the generality of the approach, provides easy portability of the code and provides efficient re‐use of existing software. The performance results obtained for the Cholesky factorization show that our packed format performs as good as or better than the ScaLAPACK full storage algorithm for a small number of processors. For a larger number of processors, the ScaLAPACK full storage routine performs slightly better until each processor runs out of memory. Copyright © 2006 John Wiley & Sons, Ltd.
Marc Baboulin, Luc Giraud, Serge Gratton, Julien Langou
Concurr. Comput. Pract. Exp.2
2005 Algorithm 842: A set of GMRES routines for real and complex arithmetics on high performance computers
abstract
In this article we describe our implementations of the GMRES algorithm for both real and complex, single and double precision arithmetics suitable for serial, shared memory and distributed memory computers. For the sake of portability, simplicity, flexibility and efficiency the GMRES solvers have been implemented in Fortran 77 using the reverse communication mechanism for the matrix-vector product, the preconditioning and the dot product computations. For distributed memory computation, several orthogonalization procedures have been implemented to reduce the cost of the dot product calculation, which is a well-known bottleneck of efficiency for the Krylov methods. Either implicit or explicit calculation of the residual at restart are possible depending on the actual cost of the matrix-vector product. Finally the implemented stopping criterion is based on a normwise backward error.
Valérie Frayssé, Luc Giraud, Serge Gratton, Julien Langou
ACM Trans. Math. Softw.2
2003 Topic Introduction
Iain S. Duff, Luc Giraud, Henk A. van der Vorst, Peter Zinterhof
Euro-Par2
2002 Numerical Algorithms
Iain S. Duff, Wolfgang Borchers, Luc Giraud, Henk A. van der Vorst
Euro-Par3
1999 Emerging Topics in Advanced Computing in Europe - Introduction
Renato Campo, Luc Giraud
Euro-Par2
1999 Parallel Subdomain-Based Preconditioner for the Schur Complement
Luiz Mariano Carvalho, Luc Giraud
Euro-Par2
1999 Some Investigations of Domain Decomposition Techniques in Parallel CFD
F. Chalot, G. Chevalier, Luc Giraud
Euro-Par4
1999 A Parallel Distributed Fast 3D Poisson Solver for Méso-NH
Luc Giraud, Ronan Guivarch, Joël Stein
Euro-Par1
1999 Parallelization of the French Meteorological Mesoscale Model MésoNH
Patrick Jabouille, Ronan Guivarch, Philippe Kloos, Didier Gazen, Nicolas Gicquel, Luc Giraud, Nicole Asencio, Veronique Ducrocq, Juan Escobar, Jean-Luc Redelsperger, Joël Stein, Jean-Pierre Pinty
Euro-Par6
1998 On the Influence of the Orthogonalization Scheme on the Parallel Performance of GMRES
Valérie Frayssé, Luc Giraud, Hatim Kharraz Aroussi
Euro-Par2