Patrick Amestoy

dblp:a/PatrickAmestoy · also Patrick R. Amestoy · DBLP profile ↗
← Back
12ranked-venue papers
10as first author
2since 2021 · last 2026
0000-0002-8559-9600ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 5 first-authorTheory of computation · 5 · 5 first-author · 2 since 2021
YearPublicationVenuePosition
2026 BLAS-based Mixed Precision Block Memory Accessor with Applications to Sparse Direct Solvers
abstract
International audience
Patrick Amestoy, Antoine Jego, Jean-Yves L'Excellent, Théo Mary, Gregoire Pichon
ACM Trans. Math. Softw.1
2023 Combining Sparse Approximate Factorizations with Mixed-precision Iterative Refinement
abstract
The standard LU factorization-based solution process for linear systems can be enhanced in speed or accuracy by employing mixed-precision iterative refinement. Most recent work has focused on dense systems. We investigate the potential of mixed-precision iterative refinement to enhance methods for sparse systems based on approximate sparse factorizations. In doing so, we first develop a new error analysis for LU- and GMRES-based iterative refinement under a general model of LU factorization that accounts for the approximation methods typically used by modern sparse solvers, such as low-rank approximations or relaxed pivoting strategies. We then provide a detailed performance analysis of both the execution time and memory consumption of different algorithms, based on a selected set of iterative refinement variants and approximate sparse factorizations. Our performance study uses the multifrontal solver MUMPS, which can exploit block low-rank factorization and static pivoting. We evaluate the performance of the algorithms on large, sparse problems coming from a variety of real-life and industrial applications showing that mixed-precision iterative refinement combined with approximate sparse factorization can lead to considerable reductions of both the time and memory consumption.
Patrick Amestoy, Alfredo Buttari, Nicholas J. Higham, Jean-Yves L'Excellent, Théo Mary, Bastien Vieublé
ACM Trans. Math. Softw.1
2019 Performance and Scalability of the Block Low-Rank Multifrontal Factorization on Multicore Architectures
abstract
Matrices coming from elliptic Partial Differential Equations have been shown to have a low-rank property that can be efficiently exploited in multifrontal solvers to provide a substantial reduction of their complexity. Among the possible low-rank formats, the Block Low-Rank format (BLR) is easy to use in a general purpose multifrontal solver and its potential compared to standard (full-rank) solvers has been demonstrated. Recently, new variants have been introduced and it was proved that they can further reduce the complexity but their performance has never been analyzed. In this article, we present a multithreaded BLR factorization and analyze its efficiency and scalability in shared-memory multicore environments. We identify the challenges posed by the use of BLR approximations in multifrontal solvers and put forward several algorithmic variants of the BLR factorization that overcome these challenges by improving its efficiency and scalability. We illustrate the performance analysis of the BLR multifrontal factorization with numerical experiments on a large set of problems coming from a variety of real-life applications.
Patrick Amestoy, Alfredo Buttari, Jean-Yves L'Excellent, Théo Mary
ACM Trans. Math. Softw.1
2010 Parallel Numerical Algorithms
Patrick Amestoy, Daniela di Serafino, Rob H. Bisseling, Enrique S. Quintana-Ortí, Marián Vajtersic
Euro-Par (2)1
2010 Analysis of the solution phase of a parallel multifrontal approach
Patrick Amestoy, Iain S. Duff, Abdou Guermouche, Tzvetomila Slavova
Parallel Comput.1
2009 Introduction
Peter Arbenz, Martin B. van Gijzen, Patrick Amestoy, Pasqua D'Ambra
Euro-Par3
2006 Hybrid scheduling for the parallel solution of linear systems
Patrick Amestoy, Abdou Guermouche, Jean-Yves L'Excellent, Stéphane Pralet
Parallel Comput.1
2004 Algorithm 837: AMD, an approximate minimum degree ordering algorithm
abstract
AMD is a set of routines that implements the approximate minimum degree ordering algorithm to permute sparse matrices prior to numerical factorization. There are versions written in both C and Fortran 77. A MATLAB interface is included.
Patrick Amestoy, Enseeiht-Irit, Timothy A. Davis 0001, Iain S. Duff
ACM Trans. Math. Softw.1
2003 Impact of the implementation of MPI point-to-point communications on the performance of two general sparse solvers
Patrick Amestoy, Iain S. Duff, Jean-Yves L'Excellent, Xiaoye S. Li
Parallel Comput.1
2003 Adapting a parallel sparse direct solver to architectures with clusters of SMPs
Patrick Amestoy, Iain S. Duff, Stéphane Pralet, Christof Vömel
Parallel Comput.1
2001 Analysis and comparison of two general sparse solvers for distributed memory computers
abstract
This paper provides a comprehensive study and comparison of two state-of-the-art direct solvers for large sparse sets of linear equations on large-scale distributed-memory computers. One is a multifrontal solver called MUMPS, the other is a supernodal solver called superLU. We describe the main algorithmic features of the two solvers and compare their performance characteristics with respect to uniprocessor speed, interprocessor communication, and memory requirements. For both solvers, preorderings for numerical stability and sparsity play an important role in achieving high parallel efficiency. We analyse the results with various ordering algorithms. Our performance analysis is based on data obtained from runs on a 512-processor Cray T3E using a set of matrices from real applications. We also use regular 3D grid problems to study the scalability of the two solvers.
Patrick Amestoy, Iain S. Duff, Jean-Yves L'Excellent, Xiaoye S. Li
ACM Trans. Math. Softw.1
2000 Hybridizing Nested Dissection and Halo Approximate Minimum Degree for Efficient Sparse Matrix Ordering
abstract
Minimum degree and nested dissection are the two most popular reordering schemes used to reduce fill-in and operation count when factoring and solving sparse matrices. Most of the state-of-the-art ordering packages hybridize these methods by performing incomplete nested dissection and ordering by minimum degree the subgraphs associated with the leaves of the separation tree, but most often only loose couplings have been achieved, resulting in poorer performance than could have been expected. This paper presents a tight coupling of the nested dissection and halo approximate minimum degree algorithms, which allows the minimum degree algorithm to use exact degrees on the boundaries of the subgraphs passed to it and to yield back not only the ordering of the nodes of the subgraph, but also the amalgamated assembly subtrees, for efficient block computations. Experimental results show the performance improvement of this hybridization, both in terms of fill-in reduction and increase of concurrency on a parallel sparse block symmetric solver. Copyright © 2000 John Wiley & Sons, Ltd.
François Pellegrini, Jean Roman, Patrick Amestoy
Concurr. Pract. Exp.3