EDBT 2026 Demo / reviewers in the wild / expert
Mathieu Luisier
dblp:49/2897
· DBLP profile ↗
20ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0002-2212-7972ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Parallel Quadratic Selected Inversion in Quantum Transport SimulationabstractDriven by Moore’s law, the dimensions of transistors have been pushed down to the nanometer scale so that advanced quantum transport (QT) solvers are nowadays required to reliably design such nano-devices. The non-equilibrium Green’s function (NEGF) formalism is suited to this task but is computationally intensive, involving the selected inversion (SI) and the selected solution of quadratic matrix (SQ) equations. Existing algorithms to tackle these numerical problems are ideally suited to GPU acceleration, e.g., the recursive Green’s function (RGF) technique. However, they are typically sequential, limited to block-tridiagonal (BT) matrices, and their implementation has been restricted so far to shared-memory parallelism, limiting the achievable device sizes. To address these shortcomings, we introduce distributed methods that build on RGF and enable parallel SI and SQ. We further extend them to handle BT matrices with arrowhead, allowing for the inclusion of gate leakage currents, a major limiting factor at ultra-scaled device dimensions. We evaluate the performance of our approach on a real dataset from the QT simulation of a nano-ribbon field-effect transistor and perform a comparison with the sparse direct solvers PARDISO and cuDSS. Our SI solver is at least one order of magnitude faster than PARDISO (cuDSS) on CPUs (GPUs), regardless of the system size. When fused, our SI+SQ implementation outperforms the SI-only module of PARDISO by a factor of 1.56 × for the same device dimensions. Performing weak scaling up to 8 CPUs (GPU), our SI+SQ solver achieves a parallel efficiency of \(\eta = 17.1\%\) (\(\eta = 18.5\%\)), thus enabling distributed memory nano-device simulations. Vincent Maillou, Matthias Bollhöfer, Olaf Schenk, Alexandros Nikolaos Ziogas, Mathieu Luisier |
ICS | 5 |
| 2025 | Parallel Selected Inversion of Block-Tridiagonal with Arrowhead MatricesabstractThe inversion of structured sparse matrices is a fundamental yet computationally and memory-intensive task in many scientific applications, such as Bayesian statistical modeling and material science. In certain cases, only particular entries of the full inverse are required. This has motivated the development of so-called selected inversion algorithms (SIA), capable of computing only specific elements of the full inverse. Currently, most SIA implementations are restricted to shared-/distributed-memory CPU architectures or to single GPUs. Here, we introduce novel numerical methods to perform the parallel selected inversion and Cholesky decomposition of positive-definite, block-tridiagonal with arrowhead matrices. A distributed memory, GPU-accelerated implementation of our approach is presented and integrated into the structured solver library Serinv. We demonstrate its performance on synthetic and real datasets from statistical air temperature prediction models and achieve CPU (GPU) speedups of up to$2.6 \times(71.4 \times)$over the SIA of the PARDISO library and up to$14 \times(380.9 \times)$over the MUMPS library, when scaling to 16 processes. Vincent Maillou, Lisa Gaedke-Merzhäuser, Alexandros Nikolaos Ziogas, Olaf Schenk, Mathieu Luisier |
CLUSTER | 5 |
| 2025 | Learning the Electronic Hamiltonian of Large Atomic StructuresabstractGraph neural networks (GNNs) have shown promise in learning the ground-state electronic properties of materials, subverting ab initio density functional theory (DFT) calculations when the underlying lattices can be represented as small and/or repeatable unit cells (i.e., molecules and periodic crystals). Realistic systems are, however, non-ideal and generally characterized by higher structural complexity. As such, they require large (10+ {Å}) unit cells and thousands of atoms to be accurately described. At these scales, DFT becomes computationally prohibitive, making GNNs especially attractive. In this work, we present a strictly local equivariant GNN capable of learning the electronic Hamiltonian (H) of realistically extended materials. It incorporates an augmented partitioning approach that enables training on arbitrarily large structures while preserving local atomic environments beyond boundaries. We demonstrate its capabilities by predicting the electronic Hamiltonian of various systems with up to 3,000 nodes (atoms), 500,000+ edges, 28 million orbital interactions (nonzero entries of H), and $\leq$0.53% error in the eigenvalue spectra. Our work expands the applicability of current electronic property prediction methods to some of the most challenging cases encountered in computational materials science, namely systems with disorder, interfaces, and defects. Chen Hao Xia, Manasa Kaniselvan, Alexandros Nikolaos Ziogas, Marko Mladenovic, Rayen Mahjoub, Alexander Maeder, Mathieu Luisier |
ICML | 7 |
| 2025 | Accelerated Spatio-Temporal Bayesian Modeling for Multivariate Gaussian ProcessesabstractMultivariate Gaussian processes (GPs) offer a powerful probabilistic framework to represent complex interdependent phenomena. They pose, however, significant computational challenges in high-dimensional settings, which frequently arise in spatio-temporal applications. We present DALIA, a highly scalable framework for performing Bayesian inference tasks on spatio-temporal multivariate GPs, based on the methodology of integrated nested Laplace approximations. Our approach relies on a sparse inverse covariance matrix formulation of the GP, puts forward a GPU-accelerated block-dense approach, and introduces a hierarchical, triple-layer, distributed-memory parallel scheme. We showcase weak-scaling performance surpassing the state of the art by two orders of magnitude on a model whose parameter space is 8 × larger and measure strong-scaling speedups of three orders of magnitude when running on 496 GH200 superchips on the Alps supercomputer. Applying DALIA to an air pollution study over northern Italy spanning 48 days, we showcase refined spatial resolutions over the aggregated pollutant measurements. Lisa Gaedke-Merzhäuser, Vincent Maillou, Fernando Rodriguez Avellaneda, Olaf Schenk, Paula Moraga, Mathieu Luisier, Alexandros Nikolaos Ziogas, Håvard Rue |
SC | 6 |
| 2025 | Ab-initio Quantum Transport with the GW Approximation, 42, 240 Atoms, and Sustained Exascale PerformanceabstractDesigning nanoscale electronic devices such as the currently manufactured nanoribbon field-effect transistors (NRFETs) requires advanced modeling tools capturing all relevant quantum mechanical effects. State-of-the-art approaches combine the non-equilibrium Green’s function (NEGF) formalism and density functional theory (DFT). However, as device dimensions do not exceed a few nanometers anymore, electrons are confined in ultra-small volumes, giving rise to strong electron-electron interactions. To account for these critical effects, DFT+NEGF solvers should be extended with the GW approximation, which massively increases their computational intensity. Here, we present the first implementation of the NEGF+GW scheme capable of handling NRFET geometries with dimensions comparable to experiments. This package, called QuaTrEx, makes use of a novel spatial domain decomposition scheme, can treat devices made of up to 84,480 atoms, scales very well on the Alps and Frontier supercomputers (> 80% weak scaling efficiency), and sustains an exascale FP64 performance on 42,240 atoms (1.15 Eflop/s). Nicolas Vetsch, Alexander Maeder, Vincent Maillou, Anders Winka, Jiang Cao, Grzegorz Kwasniewski, Leonard Deuschle, Torsten Hoefler, Alexandros Nikolaos Ziogas, Mathieu Luisier |
SC | 10 |
| 2024 | Invariant subspaces and PCA in nearly matrix multiplication timeabstractApproximating invariant subspaces of generalized eigenvalue problems (GEPs) is a fundamental computational problem at the core of machine learning and scientific computing. It is, for example, the root of Principal Component Analysis (PCA) for dimensionality reduction, data visualization, and noise filtering, and of Density Functional Theory (DFT), arguably the most popular method to calculate the electronic structure of materials.
Given Hermitian $H,S\in\mathbb{C}^{n\times n}$, where $S$ is positive-definite, let $\Pi_k$ be the true spectral projector on the invariant subspace that is associated with the $k$ smallest (or largest) eigenvalues of the GEP $HC=SC\Lambda$, for some $k\in[n]$.
We show that we can compute a matrix $\widetilde\Pi_k$ such that $\lVert\Pi_k-\widetilde\Pi_k\rVert_2\leq \epsilon$, in $O\left( n^{\omega+\eta}\mathrm{polylog}(n,\epsilon^{-1},\kappa(S),\mathrm{gap}_k^{-1}) \right)$ bit operations in the floating point model, for some $\epsilon\in(0,1)$, with probability $1-1/n$. Here, $\eta>0$ is arbitrarily small, $\omega\lesssim 2.372$ is the matrix multiplication exponent, $\kappa(S)=\lVert S\rVert_2\lVert S^{-1}\rVert_2$, and $\mathrm{gap}_k$ is the gap between eigenvalues $k$ and $k+1$.
To achieve such provable "forward-error" guarantees, our methods rely on a new $O(n^{\omega+\eta})$ stability analysis for the Cholesky factorization, and a smoothed analysis for computing spectral gaps, which can be of independent interest.
Ultimately, we obtain new matrix multiplication-type bit complexity upper bounds for PCA problems, including classical PCA and (randomized) low-rank approximation. Alexandros Sobczyk, Marko Mladenovic, Mathieu Luisier |
NeurIPS | 3 |
| 2024 | Towards Exascale Simulations of Nanoelectronic Devices in the GW ApproximationabstractExperimental development of gate-all-around silicon nanowire field-effect transistors (NWFETs), a viable replacement for FinFETs, can be complemented by technology computer-aided design. This requires the availability of advanced device simulators relying on a quantum transport (QT) approach without any empirical parameters as inputs. Concretely, all material properties should be described from first-principles, and the whole physics at play should be accurately modeled, particularly the strong electron-electron interactions occurring in highly confined structures such as NWFETs. To shed light on these many-body effects, we implement them within the self-consistent GW approximation into an ab initio QT solver called QuaTrEx, based on density functional theory and the Non-equilibrium Green’s Function formalism. We then simulate transistors made of up to 10,560 atoms on the LUMI supercomputer’s GPU partition, reaching a parallel efficiency of $\mathbf{7 4 \%}(\mathbf{6 0 \%}$) in weak (strong) scaling and an overall computational performance of 69.3 Pflop/s in double precision on 1,800 nodes. Leonard Deuschle, Alexander Maeder, Vincent Maillou, Nicolas Vetsch, Anders Winka, Jiang Cao, Alexandros Nikolaos Ziogas, Mathieu Luisier |
SC | 8 |
| 2024 | Accelerated Atomistic Kinetic Monte Carlo Simulations of Resistive Memory ArraysabstractSimulating emerging resistive switching memory devices, such as memristors, requires modeling frameworks that can treat the motion of point defects across nanoscale domains. Field-driven Kinetic Monte Carlo (d-KMC) methods that simulate the discrete structural evolution of atomic coordinates in the presence of external potential and heat fields can be used for this purpose. While physically similar to conventional KMC methods, field-driven approaches present different computational motifs and introduce global communication. Here, we develop the first scalable d-KMC code for resistive memory arrays at atomistic resolution. We accelerate this latency-sensitive simulation on the GPU partition of the LUMI Supercomputer, exploiting the high-speed interconnects between GPUs on the same node. Applied to the technologically relevant HfOx material stack, our code enables the first atomistic simulation of $3 \times 3$ arrays of resistive switching memory cells with more than 1 million atoms, matching the dimensions of fabricated structures. Manasa Kaniselvan, Alexander Maeder, Marko Mladenovic, Mathieu Luisier, Alexandros Nikolaos Ziogas |
SC | 4 |
| 2022 | Approximate Euclidean lengths and distances beyond Johnson-LindenstraussabstractA classical result of Johnson and Lindenstrauss states that a set of $n$ high dimensional data points can be projected down to $O(\log n/\epsilon^2)$ dimensions such that the square of their pairwise distances is preserved up to a small distortion $\epsilon\in(0,1)$. It has been proved that the JL lemma is optimal for the general case, therefore, improvements can only be explored for special cases. This work aims to improve the $\epsilon^{-2}$ dependency based on techniques inspired by the Hutch++ Algorithm, which reduces $\epsilon^{-2}$ to $\epsilon^{-1}$ for the related problem of implicit matrix trace estimation. We first present an algorithm to estimate the Euclidean lengths of the rows of a matrix. We prove for it element-wise probabilistic bounds that are at least as good as standard JL approximations in the worst-case, but are asymptotically better for matrices with decaying spectrum. Moreover, for any matrix, regardless of its spectrum, the algorithm achieves $\epsilon$-accuracy for the total, Frobenius norm-wise relative error using only $O(\epsilon^{-1})$ queries. This is a quadratic improvement over the norm-wise error of standard JL approximations. We also show how these results can be extended to estimate (i) the Euclidean distances between data points and (ii) the statistical leverage scores of tall-and-skinny data matrices, which are ubiquitous for many applications, with analogous theoretical improvements. Proof-of-concept numerical experiments are presented to validate the theoretical analysis. Alexandros Sobczyk, Mathieu Luisier |
NeurIPS | 2 |
| 2020 | Countdown Slack: A Run-Time Library to Reduce Energy Footprint in Large-Scale MPI ApplicationsabstractThe power consumption of supercomputers is a major challenge for system owners, users, and society. It limits the capacity of system installations, it requires large cooling infrastructures, and it is the cause of a large carbon footprint. Reducing power during application execution without changing the application source code or increasing time-to-completion is highly desirable in real-life high-performance computing scenarios. The power management run-time frameworks proposed in the last decade are based on the assumption that the duration of communication and application phases in an MPI application can be predicted and used at run-time to trade-off communication slack with power consumption. In this article, we first show that this assumption is too general and leads to mispredictions, slowing down applications, thereby jeopardizing the claimed benefits. We then propose a new approach based on (i) the separation of communication phases and slack during MPI calls and (ii) a timeout algorithm to cope with the hardware power management latency, which jointly makes it possible to achieve performance-neutral power saving in MPI applications without requiring labor-intensive and risky application source code modifications. We validate our approach in a tier-1 production environment with widely adopted scientific applications. Our approach has a time-to-completion overhead lower than 1 percent, while it successfully exploits slack in communication phases to achieve an average energy saving of 10 percent. If we focus on a large-scale application runs, the proposed approach achieves 22 percent energy saving with an overhead of only 0.4 percent. With respect to state-of-the-art approaches, COUNTDOWN Slack is the only that always leads to an energy saving with negligible overhead (<; 3 percent). Daniele Cesarini, Andrea Bartolini, Andrea Borghesi, Carlo Cavazzoni, Mathieu Luisier, Luca Benini |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2019 | A data-centric approach to extreme-scale ab initio dissipative quantum transport simulationsabstractThe Predictive Science Academic Alliance Program (PSAAP) II at Stanford University is developing an exascale-ready multi-physics solver to investigate particle-laden turbulent flows in a radiation environment for solar energy receiver applications. In order to simulate the proposed concentrated particle-based receiver design three distinct but coupled physical phenomena must be modeled: fluid flows, Lagrangian particle dynamics, and the transport of thermal radiation. Therefore, three different physics solvers (fluid, particles, and radiation) must run concurrently with significant cross-communication in an integrated multi-physics simulation. However, each solver uses substantially different algorithms and data access patterns. Coordinating the overall data communication, computational load balancing, and scaling these different physics solvers together on modern massively parallel, heterogeneous high performance computing systems presents several major challenges. We have adopted the Legion programming system, via the Regent programming language, and its task parallel programming model to address these challenges. Our multi-physics solver Soleil-X is written entirely in the high level Regent programming language and is one of the largest and most complex applications written in Regent to date. At this workshop we will give an overview of the software architecture of Soleil-X as well as discuss how our multi-physics solver was designed to use the task parallel programming model provided by Legion. We will also discuss the development experience, scaling, performance, portability, and multi-physics simulation results. Alexandros Nikolaos Ziogas, Tal Ben-Nun, Guillermo Indalecio Fernández, Timo Schneider, Mathieu Luisier, Torsten Hoefler |
SC | 5 |
| 2019 | Optimizing the data movement in quantum transport simulations via data-centric parallel programmingabstractDesigning efficient cooling systems for integrated circuits (ICs) relies on a deep understanding of the electro-thermal properties of transistors. To shed light on this issue in currently fabricated Fin-FETs, a quantum mechanical solver capable of revealing atomically-resolved electron and phonon transport phenomena from first-principles is required. In this paper, we consider a global, data-centric view of a state-of-the-art quantum transport simulator to optimize its execution on supercomputers. The approach yields coarse-and fine-grained data-movement characteristics, which are used for performance and communication modeling, communication-avoidance, and data-layout transformations. The transformations are tuned for the Piz Daint and Summit supercomputers, where each platform requires different caching and fusion strategies to perform optimally. The presented results make ab initio device simulation enter a new era, where nanostructures composed of over 10,000 atoms can be investigated at an unprecedented level of accuracy, paving the way for better heat management in next-generation ICs. Alexandros Nikolaos Ziogas, Tal Ben-Nun, Guillermo Indalecio Fernández, Timo Schneider, Mathieu Luisier, Torsten Hoefler |
SC | 5 |
| 2015 | Pushing back the limit of ab-initio quantum transport simulations on hybrid supercomputersabstractThe capabilities of CP2K, a density-functional theory package and OMEN, a nano-device simulator, are combined to study transport phenomena from first-principles in unprecedentedly large nanostructures. Based on the Hamiltonian and overlap matrices generated by CP2K for a given system, OMEN solves the Schrödinger equation with open boundary conditions (OBCs) for all possible electron momenta and energies. To accelerate this core operation a robust algorithm called SplitSolve has been developed. It allows to simultaneously treat the OBCs on CPUs and the Schrödinger equation on GPUs, taking advantage of hybrid nodes. Our key achievements on the Cray-XK7 Titan are (i) a reduction in time-to-solution by more than one order of magnitude as compared to standard methods, enabling the simulation of structures with more than 50000 atoms, (ii) a parallel efficiency of 97% when scaling from 756 up to 18564 nodes, and (iii) a sustained performance of 15 DP-PFlop/s. Mauro Calderara, Sascha Brück, Andreas Pedersen, Mohammad H. Bani-Hashemian, Joost VandeVondele, Mathieu Luisier |
SC | 6 |
| 2013 | Fast Methods for Computing Selected Elements of the Green's Function in Massively Parallel Nanoelectronic Device Simulations
Andrey Kuzmin, Mathieu Luisier, Olaf Schenk |
Euro-Par | 2 |
| 2011 | Atomistic nanoelectronic device engineering with sustained performances up to 1.44 PFlop/sabstractWe present a multi-dimensional, atomistic, quantum transport simulation approach to investigate the performances of realistic nanoscale transistors for various geometries and material systems. The central computation consists in solving the Schrödinger equation with open boundary conditions several thousand times. To do that, a Wave Function approach is used since it can be relatively easily parallelized. To further improve the computational efficiency, three additional levels of parallelization are identified, the work load is optimally balanced between the CPUs, computational interleaving is applied where possible, and a mixed precision scheme is introduced. Using two different device types, a high electron mobility and a band-to-band tunneling transistor, sustained performances up to 1.28 PFlop/s in double precision (55% of the peak performance) and 1.44 PFlop/s in mixed precision are reached on 221,400 cores on the CRAY-XT5 Jaguar at Oak Ridge National Lab. Mathieu Luisier, Timothy B. Boykin, Gerhard Klimeck, Wolfgang Fichtner |
SC | 1 |
| 2010 | A Parallel Implementation of Electron-Phonon Scattering in Nanoelectronic Devices up to 95k CoresabstractA quantum transport approach based on the Non-equilibrium Green's Function formalism and the tight-binding method has been developed to investigate the performances of atomistically resolved nanoelectronic devices in the presence of electron-phonon scattering. The model is integrated into a quad-level parallel environment (bias, momentum, energy, and spatial domain decomposition) that scales almost perfectly up to 220k cores in the ballistic limit of electron transport. In this case, the momentum and energy points form a quasi-embarrassingly parallel problem. The novelty in this paper is the inclusion of scattering self-energies that couple all the momenta and several energies together, requiring substantial inter-processor communication. An efficient parallel implementation of electron-phonon scattering is therefore proposed and applied to a realistically extended transistor structure. A good scaling of the simulation walltime up to 95,256 cores and a sustained performance of 142 TFlop/s are reported on the Cray-XT5 Jaguar. Mathieu Luisier |
SC | 1 |
| 2010 | Numerical strategies towards peta-scale simulations of nanoelectronics devices
Mathieu Luisier, Gerhard Klimeck |
Parallel Comput. | 1 |
| 2008 | A Parallel Sparse Linear Solver for Nearest-Neighbor Tight-Binding Problems
Mathieu Luisier, Gerhard Klimeck, Andreas Schenk, Wolfgang Fichtner, Timothy B. Boykin |
Euro-Par | 1 |
| 2008 | Rapid Parallel Systems Deployment: Techniques for Overnight Clustering
Donna Cumberland, Randy Herban, Rick Irvine, Michael Shuey, Mathieu Luisier |
LISA | 5 |
| 2008 | A multi-level parallel simulation approach to electron transport in nano-scale transistorsabstractPhysics-based simulation of electron transport in nanoelectronic devices requires the solution of thousands of highly complex equations to obtain the output characteristics of one single input voltage. The only way to obtain a complete set of bias points within a reasonable amount of time is the recourse to supercomputers offering several hundreds to thousands of cores. To profit from the rapidly increasing availability of such machines we have developed a state-of-the-art quantum mechanical transport simulator dedicated to nanodevices and working with four levels of parallelism. Using these four levels we demonstrate that an almost ideal scaling of the walltime up to 32768 processors with a parallel efficiency of 86% is reached in the simulation of realistically extended and gated field-effect transistors. Obtaining the current characteristics of these devices is reduced to some hundreds of seconds instead of days on a small cluster or months on a single CPU. Mathieu Luisier, Gerhard Klimeck |
SC | 1 |