Lisandro Dalcín

dblp:06/4626 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
3since 2021 · last 2023
0000-0001-8086-0155ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2023 MPI Application Binary Interface Standardization
abstract
MPI is the most widely used interface for high-performance computing (HPC) workloads. Its success lies in its embrace of libraries and ability to evolve while maintaining backward compatibility for older codes, enabling them to run on new architectures for many years. In this paper, we propose a new level of MPI compatibility: a standard Application Binary Interface (ABI). We review the history of MPI implementation ABIs, identify the constraints from the MPI standard and ISO C, and summarize recent efforts to develop a standard ABI for MPI. We provide the current proposal from the MPI Forum’s ABI working group, which has been prototyped both within MPICH and as an independent abstraction layer called Mukautuva. We also list several use cases that would benefit from the definition of an ABI while outlining the remaining constraints.
Jeff R. Hammond, Lisandro Dalcín, Erik Schnetter, Marc Pérache, Jean-Baptiste Besnard, Jed Brown, Gonzalo Brito Gadeschi, Simon Byrne, Joseph Schuchart, Hui Zhou 0012
EuroMPI2
2023 Improvements to SLEPc in Releases 3.14-3.18
abstract
This short article describes the main new features added to SLEPc, the Scalable Library for Eigenvalue Problem Computations, in the past two and a half years, corresponding to five release versions. The main novelty is the extension of the SVD module with new problem types, such as the generalized SVD or the hyperbolic SVD. Additionally, many improvements have been incorporated in different parts of the library, including contour integral eigensolvers, preconditioning, and GPU support.
José E. Román, Fernando Alvarruiz, Carmen Campos, Lisandro Dalcín, Pierre Jolivet, Alejandro Lamas Daviña
ACM Trans. Math. Softw.4
2023 mpi4py.futures: MPI-Based Asynchronous Task Execution for Python
abstract
We present mpi4py.futures, a lightweight, asynchronous task execution framework targeting the Python programming language and using the Message Passing Interface (MPI) for interprocess communication. mpi4py.futures follows the interface of the concurrent.futures package from the Python standard library and can be used as its drop-in replacement, while allowing applications to scale over multiple compute nodes. We discuss the design, implementation, and feature set of mpi4py.futures and compare its performance to other solutions on both shared and distributed memory architectures. On a shared-memory system, we show mpi4py.futures to consistently outperform Python's concurrent.futures with speedup ratios between 1.4X and 3.7X in throughput (tasks per second) and between 1.9X and 2.9X in bandwidth. On a Cray XC40 system, we compare mpi4py.futures to Dask – a well-known Python parallel computing package. Although we note more varied results, we show mpi4py.futures to outperform Dask in most scenarios.
Marcin Rogowski, Samar Aseeri, David E. Keyes, Lisandro Dalcín
IEEE Trans. Parallel Distributed Syst.4
2020 Performance study of sustained petascale direct numerical simulation on Cray XC40 systems
abstract
Summary We present in this paper a comprehensive performance study of highly efficient extreme scale direct numerical simulations of secondary flows, using an optimized version of Nek5000. Our investigations are conducted on various Cray XC40 systems, using a very high‐order spectral element method. Single‐node efficiency is achieved by auto‐generated assembly implementations of small matrix multiplies and key vector‐vector operations, streaming lossless I/O compression, aggressive loop merging, and selective single precision evaluations. Comparative studies across different Cray XC40 systems at scale, Trinity (LANL), Cori (NERSC), and ShaheenII (KAUST) show that a Cray programming environment, network configuration, parallel file system, and burst buffer all have a major impact on the performance. All three systems possess a similar hardware with similar CPU nodes and parallel file system, but they have different theoretical peak network bandwidths, different OSs, and different versions of the programming environment. Our study reveals how these slight configuration differences can be critical in terms of performance of the application. We also find that with 9216 nodes (294 912 cores) on Trinity XC40 the applications sustain petascale performance, as well as 50% of peak memory bandwidth over the entire solver (500 TB/s in aggregate). On 3072 Xeon Phi nodes of Cori, we reach 378 TFLOP/s with an aggregated bandwidth of 310 TB/s, corresponding to time‐to‐solution 2.11× faster than obtained with the same number of (dual‐socket) Xeon nodes.
Bilel Hadri, Matteo Parsani, Maxwell Hutchinson, Alexander Heinecke, Lisandro Dalcín, David E. Keyes
Concurr. Comput. Pract. Exp.5
2019 Fast parallel multidimensional FFT using advanced MPI
Lisandro Dalcín, Mikael Mortensen, David E. Keyes
J. Parallel Distributed Comput.1
2008 MPI for Python: Performance improvements and MPI-2 extensions
abstract
MPI for Python provides bindings of the message passing interface (MPI) standard for the Python programming language and allows any Python program to exploit multiple processors. In its first release, MPI for Python was constructed on top of the MPI-1 specification defining an object-oriented interface that closely followed the MPI-2 C ++ bindings, and provided support for communications of general Python objects. In the latest release, this package is improved to enable direct blocking/non-blocking communication of numeric arrays, and to support almost all MPI-2 features. Improvements in communication performance have been tested in a Beowulf class cluster. Results showed a negligible overhead in comparison to compiled C code. MPI for Python is open source and available for download on the web ( http://mpi4py.scipy.org/ ).
Lisandro Dalcín, Rodrigo R. Paz, Mario A. Storti, Jorge D'Elía
J. Parallel Distributed Comput.1
2005 MPI for Python
Lisandro Dalcín, Rodrigo R. Paz, Mario A. Storti
J. Parallel Distributed Comput.1