Matthias Bollhöfer

dblp:51/2266 · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
3since 2021 · last 2026
0000-0002-8093-5812ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
High-performance computing · 67% Parallel and multicore computing · 33%
Databases, data mining, and information retrieval
1 paper
Data mining · 100%
Theoretical computer science
1 paper
Mathematical optimization · 100%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Data mining › structured data mining › graph mining
graph learning
0.712023
Sparse Quadratic Approximation for Graph Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Parallel and multicore computing › parallel algorithms
distributed memory algorithm
0.312018
Distributed memory sparse inverse covariance matrix estimation on high-performance computing architectures · SC 2018
High-performance computing
parallel numerical algorithms
0.312018
Distributed memory sparse inverse covariance matrix estimation on high-performance computing architectures · SC 2018
High-performance computing › sparse linear algebra
sparse matrix computation
0.312018
Distributed memory sparse inverse covariance matrix estimation on high-performance computing architectures · SC 2018
Mathematical optimization
constrained optimization
0.212023
Sparse Quadratic Approximation for Graph Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2023
Mathematical optimization › continuous optimization › nonlinear optimization
sequential quadratic programming
0.212023
Sparse Quadratic Approximation for Graph Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2023

Methods — techniques the papers use, named apart from their topics

sequential quadratic programming · 1.3m-matrix learning · 1.3gaussian maximum likelihood · 1.3sparse matrix estimation · 0.3
YearPublicationVenuePosition
2026 Parallel Quadratic Selected Inversion in Quantum Transport Simulation
abstract
Driven by Moore’s law, the dimensions of transistors have been pushed down to the nanometer scale so that advanced quantum transport (QT) solvers are nowadays required to reliably design such nano-devices. The non-equilibrium Green’s function (NEGF) formalism is suited to this task but is computationally intensive, involving the selected inversion (SI) and the selected solution of quadratic matrix (SQ) equations. Existing algorithms to tackle these numerical problems are ideally suited to GPU acceleration, e.g., the recursive Green’s function (RGF) technique. However, they are typically sequential, limited to block-tridiagonal (BT) matrices, and their implementation has been restricted so far to shared-memory parallelism, limiting the achievable device sizes. To address these shortcomings, we introduce distributed methods that build on RGF and enable parallel SI and SQ. We further extend them to handle BT matrices with arrowhead, allowing for the inclusion of gate leakage currents, a major limiting factor at ultra-scaled device dimensions. We evaluate the performance of our approach on a real dataset from the QT simulation of a nano-ribbon field-effect transistor and perform a comparison with the sparse direct solvers PARDISO and cuDSS. Our SI solver is at least one order of magnitude faster than PARDISO (cuDSS) on CPUs (GPUs), regardless of the system size. When fused, our SI+SQ implementation outperforms the SI-only module of PARDISO by a factor of 1.56 × for the same device dimensions. Performing weak scaling up to 8 CPUs (GPU), our SI+SQ solver achieves a parallel efficiency of \(\eta = 17.1\%\) (\(\eta = 18.5\%\)), thus enabling distributed memory nano-device simulations.
Vincent Maillou, Matthias Bollhöfer, Olaf Schenk, Alexandros Nikolaos Ziogas, Mathieu Luisier
ICS2
2024 Algorithm 1042: Sparse Precision Matrix Estimation with SQUIC
abstract
We present SQUIC , a fast and scalable package for sparse precision matrix estimation. The algorithm employs a second-order method to solve the \(\ell_{1}\) -regularized maximum likelihood problem, utilizing highly optimized linear algebra subroutines. In comparative tests using synthetic datasets, we demonstrate that SQUIC not only scales to datasets of up to a million random variables but also consistently delivers runtimes that are significantly faster than other well-established sparse precision matrix estimation packages. Furthermore, we showcase the application of the introduced package in classifying microarray gene expressions. We demonstrate that by utilizing a matrix form of the tuning parameter (also known as the regularization parameter), SQUIC can effectively incorporate prior information into the estimation procedure, resulting in improved application results with minimal computational overhead.
Aryan Eftekhari, Lisa Gaedke-Merzhäuser, Dimosthenis Pasadakis, Matthias Bollhöfer, Simon Scheidegger, Olaf Schenk
ACM Trans. Math. Softw.4
2023 Sparse Quadratic Approximation for Graph Learning
abstract
-regularized Gaussian maximum-likelihood method is a popular approach, but also one that poses computational challenges for large scale datasets. Recently proposed methods cast this problem as a constrained optimization variant of precision matrix estimation. In this paper, we build on a state-of-the-art sparse precision matrix estimation method and introduce two algorithms that learn M-matrices, that can be subsequently used for the estimation of graph Laplacian matrices. In the first one, we propose an unconstrained method that follows a post processing approach in order to learn an M-matrix, and in the second one, we implement a constrained approach based on sequential quadratic programming. We also demonstrate the effectiveness, accuracy, and performance of both algorithms. Our numerical examples and comparative results with modern open-source packages reveal that the proposed methods can accelerate the learning of graphs by up to 3 orders of magnitude, while accurately retrieving the latent graphical structure of the data. Furthermore, we conduct large scale case studies for the clustering of COVID-19 daily cases and the classification of image datasets to highlight the applicability in real-world scenarios.
Dimosthenis Pasadakis, Matthias Bollhöfer, Olaf Schenk
IEEE Trans. Pattern Anal. Mach. Intell.2
2018 Distributed memory sparse inverse covariance matrix estimation on high-performance computing architectures
Aryan Eftekhari, Matthias Bollhöfer, Olaf Schenk
SC2
2017 Communication in task-parallel ILU-preconditioned CG solvers using MPI + OmpSs
abstract
Summary We target the parallel solution of sparse linear systems via iterative Krylov subspace–based methods enhanced with incomplete LU (ILU)‐type preconditioners on clusters of multicore processors. In order to tackle large‐scale problems, we develop task‐parallel implementations of the classical iteration for the CG method, accelerated via ILUPACK and ILU(0) preconditioners, using MPI + OmpSs. In addition, we integrate several communication‐avoiding strategies into the codes, including the butterfly communication scheme and Eijkhout's formulation of the CG method. For all these implementations, we analyze the communication patterns and perform a comparative analysis of their performance and scalability on a cluster consisting of 16 nodes, with 16 cores each.
José Ignacio Aliaga, Maria Barreda, Goran Flegar, Matthias Bollhöfer, Enrique S. Quintana-Ortí
Concurr. Comput. Pract. Exp.4
2016 Exploiting Task-Parallelism in Message-Passing Sparse Linear System Solvers Using OmpSs
José Ignacio Aliaga, Maria Barreda, Matthias Bollhöfer, Enrique S. Quintana-Ortí
Euro-Par3
2016 Exploiting task and data parallelism in ILUPACK's preconditioned CG solver on NUMA architectures and many-core accelerators
José Ignacio Aliaga, Rosa M. Badia, Maria Barreda, Matthias Bollhöfer, Ernesto Dufrechu, Pablo Ezzatti, Enrique S. Quintana-Ortí
Parallel Comput.4
2014 Leveraging Data-Parallelism in ILUPACK using Graphics Processors
abstract
In this paper, we address the exploitation of data parallelism for the solution of sparse symmetric positive definite linear systems via iterative methods on Graphics Processing Units (GPUs). In particular, we accelerate the preconditioned CG-based iterative solver underlying the incomplete LU decomposition package (ILUPACK) by off-loading the most expensive computations i.e., The solution of sparse triangular systems and sparse matrix-vector products-to the hardware accelerator. The results collected using GPUs from the two most recent generations from NVIDIA ("Fermi" and "Kepler") and a benchmark test bed of sparse linear systems show that the GPU-enabled implementations deliver a notable reduction of the execution time, while maintaining the convergence rate and numerical properties of the original ILUPACK solver.
José Ignacio Aliaga, Matthias Bollhöfer, Ernesto Dufrechu, Pablo Ezzatti, Enrique S. Quintana-Ortí
ISPDC2
2014 Leveraging Task-Parallelism with OmpSs in ILUPACK's Preconditioned CG Method
abstract
In this paper we describe how to efficiently exploit task parallelism for the solution of sparse linear systems on multithreaded processors via ILUPACK's multi-level preconditioned CG method. Using a pair of data structures, we capture the task dependencies that appear in the two most challenging operations in the method (calculation of the preconditioned and its application), passing this information to the OmpSs runtime which can then implement a correct and efficient schedule of the entire solver. Our results with high-end multicore platforms equipped with Intel and AMD processors report significant performance gains, demonstrating that OmpSs provides an efficient and close-to seamless means to leverage the concurrency in a complex scientific code like ILUPACK.
José Ignacio Aliaga, Rosa M. Badia, Maria Barreda, Matthias Bollhöfer, Enrique S. Quintana-Ortí
SBAC-PAD4
2011 Exploiting thread-level parallelism in the iterative solution of sparse linear systems
José Ignacio Aliaga, Matthias Bollhöfer, Alberto F. Martín, Enrique S. Quintana-Ortí
Parallel Comput.2
2007 Topic 10 Parallel Numerical Algorithms
Iain S. Duff, Michel J. Daydé, Matthias Bollhöfer, Anne E. Trefethen
Euro-Par3