Vincent Maillou

dblp:392/8249 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2026
0000-0003-4861-3298ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Parallel Quadratic Selected Inversion in Quantum Transport Simulation
abstract
Driven by Moore’s law, the dimensions of transistors have been pushed down to the nanometer scale so that advanced quantum transport (QT) solvers are nowadays required to reliably design such nano-devices. The non-equilibrium Green’s function (NEGF) formalism is suited to this task but is computationally intensive, involving the selected inversion (SI) and the selected solution of quadratic matrix (SQ) equations. Existing algorithms to tackle these numerical problems are ideally suited to GPU acceleration, e.g., the recursive Green’s function (RGF) technique. However, they are typically sequential, limited to block-tridiagonal (BT) matrices, and their implementation has been restricted so far to shared-memory parallelism, limiting the achievable device sizes. To address these shortcomings, we introduce distributed methods that build on RGF and enable parallel SI and SQ. We further extend them to handle BT matrices with arrowhead, allowing for the inclusion of gate leakage currents, a major limiting factor at ultra-scaled device dimensions. We evaluate the performance of our approach on a real dataset from the QT simulation of a nano-ribbon field-effect transistor and perform a comparison with the sparse direct solvers PARDISO and cuDSS. Our SI solver is at least one order of magnitude faster than PARDISO (cuDSS) on CPUs (GPUs), regardless of the system size. When fused, our SI+SQ implementation outperforms the SI-only module of PARDISO by a factor of 1.56 × for the same device dimensions. Performing weak scaling up to 8 CPUs (GPU), our SI+SQ solver achieves a parallel efficiency of \(\eta = 17.1\%\) (\(\eta = 18.5\%\)), thus enabling distributed memory nano-device simulations.
Vincent Maillou, Matthias Bollhöfer, Olaf Schenk, Alexandros Nikolaos Ziogas, Mathieu Luisier
ICS1
2025 Parallel Selected Inversion of Block-Tridiagonal with Arrowhead Matrices
abstract
The inversion of structured sparse matrices is a fundamental yet computationally and memory-intensive task in many scientific applications, such as Bayesian statistical modeling and material science. In certain cases, only particular entries of the full inverse are required. This has motivated the development of so-called selected inversion algorithms (SIA), capable of computing only specific elements of the full inverse. Currently, most SIA implementations are restricted to shared-/distributed-memory CPU architectures or to single GPUs. Here, we introduce novel numerical methods to perform the parallel selected inversion and Cholesky decomposition of positive-definite, block-tridiagonal with arrowhead matrices. A distributed memory, GPU-accelerated implementation of our approach is presented and integrated into the structured solver library Serinv. We demonstrate its performance on synthetic and real datasets from statistical air temperature prediction models and achieve CPU (GPU) speedups of up to$2.6 \times(71.4 \times)$over the SIA of the PARDISO library and up to$14 \times(380.9 \times)$over the MUMPS library, when scaling to 16 processes.
Vincent Maillou, Lisa Gaedke-Merzhäuser, Alexandros Nikolaos Ziogas, Olaf Schenk, Mathieu Luisier
CLUSTER1
2025 Accelerated Spatio-Temporal Bayesian Modeling for Multivariate Gaussian Processes
abstract
Multivariate Gaussian processes (GPs) offer a powerful probabilistic framework to represent complex interdependent phenomena. They pose, however, significant computational challenges in high-dimensional settings, which frequently arise in spatio-temporal applications. We present DALIA, a highly scalable framework for performing Bayesian inference tasks on spatio-temporal multivariate GPs, based on the methodology of integrated nested Laplace approximations. Our approach relies on a sparse inverse covariance matrix formulation of the GP, puts forward a GPU-accelerated block-dense approach, and introduces a hierarchical, triple-layer, distributed-memory parallel scheme. We showcase weak-scaling performance surpassing the state of the art by two orders of magnitude on a model whose parameter space is 8 × larger and measure strong-scaling speedups of three orders of magnitude when running on 496 GH200 superchips on the Alps supercomputer. Applying DALIA to an air pollution study over northern Italy spanning 48 days, we showcase refined spatial resolutions over the aggregated pollutant measurements.
Lisa Gaedke-Merzhäuser, Vincent Maillou, Fernando Rodriguez Avellaneda, Olaf Schenk, Paula Moraga, Mathieu Luisier, Alexandros Nikolaos Ziogas, Håvard Rue
SC2
2025 Ab-initio Quantum Transport with the GW Approximation, 42, 240 Atoms, and Sustained Exascale Performance
abstract
Designing nanoscale electronic devices such as the currently manufactured nanoribbon field-effect transistors (NRFETs) requires advanced modeling tools capturing all relevant quantum mechanical effects. State-of-the-art approaches combine the non-equilibrium Green’s function (NEGF) formalism and density functional theory (DFT). However, as device dimensions do not exceed a few nanometers anymore, electrons are confined in ultra-small volumes, giving rise to strong electron-electron interactions. To account for these critical effects, DFT+NEGF solvers should be extended with the GW approximation, which massively increases their computational intensity. Here, we present the first implementation of the NEGF+GW scheme capable of handling NRFET geometries with dimensions comparable to experiments. This package, called QuaTrEx, makes use of a novel spatial domain decomposition scheme, can treat devices made of up to 84,480 atoms, scales very well on the Alps and Frontier supercomputers (> 80% weak scaling efficiency), and sustains an exascale FP64 performance on 42,240 atoms (1.15 Eflop/s).
Nicolas Vetsch, Alexander Maeder, Vincent Maillou, Anders Winka, Jiang Cao, Grzegorz Kwasniewski, Leonard Deuschle, Torsten Hoefler, Alexandros Nikolaos Ziogas, Mathieu Luisier
SC3
2024 Towards Exascale Simulations of Nanoelectronic Devices in the GW Approximation
abstract
Experimental development of gate-all-around silicon nanowire field-effect transistors (NWFETs), a viable replacement for FinFETs, can be complemented by technology computer-aided design. This requires the availability of advanced device simulators relying on a quantum transport (QT) approach without any empirical parameters as inputs. Concretely, all material properties should be described from first-principles, and the whole physics at play should be accurately modeled, particularly the strong electron-electron interactions occurring in highly confined structures such as NWFETs. To shed light on these many-body effects, we implement them within the self-consistent GW approximation into an ab initio QT solver called QuaTrEx, based on density functional theory and the Non-equilibrium Green’s Function formalism. We then simulate transistors made of up to 10,560 atoms on the LUMI supercomputer’s GPU partition, reaching a parallel efficiency of $\mathbf{7 4 \%}(\mathbf{6 0 \%}$) in weak (strong) scaling and an overall computational performance of 69.3 Pflop/s in double precision on 1,800 nodes.
Leonard Deuschle, Alexander Maeder, Vincent Maillou, Nicolas Vetsch, Anders Winka, Jiang Cao, Alexandros Nikolaos Ziogas, Mathieu Luisier
SC3