EDBT 2026 Demo / reviewers in the wild / expert
Pieter Ghysels
dblp:118/0456
· DBLP profile ↗
10ranked-venue papers
3as first author
5since 2021 · last 2023
0000-0002-5981-5234ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 3 first-author · 3 since 2021Theory of computation · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | VoxImp: Impedance Extraction Simulator for Voxelized StructuresabstractAn impedance extractor, called VoxImp, is proposed to compute the impedances of the structures discretized by voxels. VoxImp iteratively solves the volume-surface integral equations discretized by a carefully selected set of basis functions. During iterative solution, matrix-vector multiplications are accelerated by the fast Fourier transform, while the rapid convergence of the iterative solution at resonant frequencies is ensured by a novel sparse preconditioner. The memory requirement of the sparse preconditioner is reduced by a sparse LU decomposition obtained by a multifrontal algorithm with compressed frontal matrices. The overall memory requirement of the VoxImp is further reduced by the Tucker decomposition. VoxImp’s accuracy, efficiency, and capability are demonstrated through impedance extraction of various voxelized structures, including an RF coil array discretized by more than four million voxels and 15 million panels and analyzed on a commodity desktop computer. Yang Liu 0179, Pieter Ghysels, Abdulkadir C. Yucel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | Sparse Approximate Multifrontal Factorization with Composite Compression MethodsabstractThis article presents a fast and approximate multifrontal solver for large sparse linear systems. In a recent work by Liu et al., we showed the efficiency of a multifrontal solver leveraging the butterfly algorithm and its hierarchical matrix extension, HODBF (hierarchical off-diagonal butterfly) compression to compress large frontal matrices. The resulting multifrontal solver can attain quasi-linear computation and memory complexity when applied to sparse linear systems arising from spatial discretization of high-frequency wave equations. To further reduce the overall number of operations and especially the factorization memory usage to scale to larger problem sizes, in this article we develop a composite multifrontal solver that employs the HODBF format for large-sized fronts, a reduced-memory version of the nonhierarchical block low-rank format for medium-sized fronts, and a lossy compression format for small-sized fronts. This allows us to solve sparse linear systems of dimension up to 2.7 × larger than before and leads to a memory consumption that is reduced by 70% while ensuring the same execution time. The code is made publicly available in GitHub. Lisa Claus, Pieter Ghysels, Yang Liu 0179, Thái Anh Nhan, Ramakrishnan Thirumalaisamy, Amneet Pal Singh Bhalla, Xiaoye S. Li |
ACM Trans. Math. Softw. | 2 |
| 2022 | Addressing Irregular Patterns of Matrix Computations on GPUs and Their Impact on Applications Powered by Sparse Direct SolversabstractMany scientific applications rely on sparse direct solvers for their numerical robustness. However, performance optimization for these solvers remains a challenging task, especially on GPUs. This is due to workloads of small dense matrices that are different in size. Matrix decompositions on such irregular workloads are rarely addressed on GPUs. This paper addresses irregular workloads of matrix computations on GPUs, and their application to accelerate sparse direct solvers. We design an interface for the basic matrix operations supporting problems of different sizes. The interface enables us to develop irrLU-GPU, an LU decomposition on matrices of different sizes. We demonstrate the impact of irrLU-GPU on sparse direct LU solvers using NVIDIA and AMD GPUs. Experimental results are shown for a sparse direct solver based on a multifrontal sparse LU decomposition applied to linear systems arising from the simulation, using finite element discretization on unstructured meshes, of a high-frequency indefinite Maxwell problem. Ahmad Abdelfattah, Pieter Ghysels, Wajih Halim Boukaram, Stanimire Tomov, Xiaoye S. Li, Jack J. Dongarra |
SC | 2 |
| 2022 | Graph Partitioning and Sparse Matrix Ordering using Reinforcement Learning and Graph Neural NetworksabstractWe present a novel method for graph partitioning, based on reinforcement learning and graph convolutional neural networks. Our approach is to recursively partition coarser representations of a given graph. The neural network is implemented using SAGE graph convolution layers, and trained using an advantage actor critic (A2C) agent. We present two variants, one for finding ean edge separator that minimizes the normalized cut or quotient cut, and one that finds a small vertex separator. The vertex separators are then used to construct a nested dissection ordering to permute a sparse matrix so that its triangular factorization will incur less fill-in. The partitioning quality is compared with partitions obtained using METIS and SCOTCH, and the nested dissection ordering is evaluated in the sparse solver SuperLU. Our results show that the proposed method achieves similar partitioning quality as METIS, SCOTCH and spectral partitioning. Furthermore, the method generalizes across different classes of graphs, and works well on a variety of graphs from the SuiteSparse sparse matrix collection. Alice Gatti, Zhixiong Hu, Tess E. Smidt, Esmond G. Ng, Pieter Ghysels |
J. Mach. Learn. Res. | 5 |
| 2022 | High performance sparse multifrontal solvers on modern GPUs
Pieter Ghysels, Ryan Synk |
Parallel Comput. | 1 |
| 2020 | Scalable and Memory-Efficient Kernel Ridge RegressionabstractWe present a scalable and memory-efficient framework for kernel ridge regression. We exploit the inherent rank deficiency of the kernel ridge regression matrix by constructing an approximation that relies on a hierarchy of low-rank factorizations of tunable accuracy, rather than leverage scores or other subsampling techniques. Without ever decompressing the kernel matrix approximation, we propose factorization and solve methods to compute the weight(s) for a given set of training and test data. We show that our method performs an optimal number of operations $\mathcal{O}\left( {{r^2}n} \right)$ with respect to the number of training samples (n) due to the underlying numerical low-rank (r) structure of the kernel matrix. Furthermore, each algorithm is also presented in the context of a massively parallel computer system, exploiting two levels of concurrency that take into account both shared-memory and distributed-memory inter-node parallelism. In addition, we present a variety of experiments using popular datasets – small, and large – to show that our approach provides sufficient accuracy in comparison with state-of-the-art methods and with the exact (i.e. non-approximated) kernel ridge regression method. For datasets, in the order of 106data points, we show that our framework strong-scales to 103cores. Finally, we provide a Python interface to the scikit-learn library so that scikit-learn can leverage our high-performance solver library to achieve much-improved performance and memory footprint. Gustavo Chavez, Yang Liu 0179, Pieter Ghysels, Xiaoye S. Li, Elizaveta Rebrova |
IPDPS | 3 |
| 2017 | A Robust Parallel Preconditioner for Indefinite Systems Using Hierarchical Matrices and Randomized SamplingabstractWe present the design and implementation of a parallel and fully algebraic preconditioner based on an approximate sparse factorization using low-rank matrix compression. The sparse factorization uses a multifrontal algorithm with fill-in occurring in dense frontal matrices. These frontal matrices are approximated as hierarchically semi-separable matrices, which are constructed using a randomized sampling technique. The resulting preconditioner has (close to) optimal complexity in terms of flops and memory usage for many discretized partial differential equations. We illustrate the robustness and performance of this new preconditioner for a number of unstructured grid problems. Initial results show that the rank-structured preconditioner could be a viable alternative to algebraic multigrid and incomplete LU, for instance. Our implementation uses MPI and OpenMP and supports real and complex arithmetic and 32 and 64 bit integers. We present a detailed performance analysis. The code is released as the STRUMPACK library with a BSD license, and a PETSc interface is available to allow for easy integration in existing applications. Pieter Ghysels, Xiaoye S. Li, Christopher Gorman, François-Henry Rouet |
IPDPS | 1 |
| 2016 | A Distributed-Memory Package for Dense Hierarchically Semi-Separable Matrix Computations Using RandomizationabstractWe present a distributed-memory library for computations with dense structured matrices. A matrix is considered structured if its off-diagonal blocks can be approximated by a rank-deficient matrix with low numerical rank. Here, we use Hierarchically Semi-Separable (HSS) representations. Such matrices appear in many applications, for example, finite-element methods, boundary element methods, and so on. Exploiting this structure allows for fast solution of linear systems and/or fast computation of matrix-vector products, which are the two main building blocks of matrix computations. The compression algorithm that we use, that computes the HSS form of an input dense matrix, relies on randomized sampling with a novel adaptive sampling mechanism. We discuss the parallelization of this algorithm and also present the parallelization of structured matrix-vector product, structured factorization, and solution routines. The efficiency of the approach is demonstrated on large problems from different academic and industrial applications, on up to 8,000 cores. This work is part of a more global effort, the STRUctured Matrices PACKage (STRUMPACK) software package for computations with sparse and dense structured matrices. Hence, although useful on their own right, the routines also represent a step in the direction of a distributed-memory sparse solver. François-Henry Rouet, Xiaoye S. Li, Pieter Ghysels, Artem Napov |
ACM Trans. Math. Softw. | 3 |
| 2014 | Hiding global synchronization latency in the preconditioned Conjugate Gradient algorithm
Pieter Ghysels, Wim Vanroose |
Parallel Comput. | 1 |
| 2012 | The Impact of Global Communication Latency at Extreme Scales on Krylov Methods
Thomas J. Ashby, Pieter Ghysels, Wim Heirman, Wim Vanroose |
ICA3PP (1) | 2 |