Chao Chen 0008

dblp:66/3019-8 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
3since 2021 · last 2024
0000-0002-5385-3651ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2024 An O(N) distributed-memory parallel direct solver for planar integral equations
abstract
Boundary value problems involving elliptic PDEs such as the Laplace and the Helmholtz equations are ubiquitous in mathematical physics and engineering. Many such problems can be alternatively formulated as integral equations that are mathematically more tractable. However, an integral-equation formulation poses a significant computational challenge: solving large dense linear systems that arise upon discretization. In cases where iterative methods converge rapidly, existing methods that draw on fast summation schemes such as the Fast Multipole Method are highly efficient and well-established. More recently, linear complexity direct solvers that sidestep convergence issues by directly computing an invertible factorization have been developed. However, storage and computation costs are high, which limits their ability to solve large-scale problems in practice. In this work, we introduce a distributed-memory parallel algorithm based on an existing direct solver named "strong recursive skeletonization factorization [1]." Specifically, we apply low-rank compression to certain off-diagonal matrix blocks in a way that minimizes computation and data movement. Compared to iterative algorithms, our method is particularly suitable for problems involving ill-conditioned matrices or multiple righthand sides. Large-scale numerical experiments are presented to show the performance of our Julia implementation.
Tianyu Liang, Chao Chen 0008, Per-Gunnar Martinsson, George Biros
IPDPS2
2022 Solving Linear Systems on a GPU with Hierarchically Off-Diagonal Low-Rank Approximations
abstract
We are interested in solving linear systems arising from three applications: (1) kernel methods in machine learning, (2) discretization of boundary integral equations from mathematical physics, and (3) Schur complements formed in the factorization of many large sparse matrices. The coefficient matrices are often data-sparse in the sense that their off-diagonal blocks have low numerical ranks; specifically, we focus on “hierarchically off-diagonal low-rank (HODLR)” matrices. We introduce algorithms for factorizing HODLR matrices and for applying the factorizations on a GPU. The algorithms leverage the efficiency of batched dense linear algebra, and they scale nearly linearly with the matrix size when the numerical ranks are fixed. The accuracy of the HODLR-matrix approximation is a tunable parameter, so we can construct high-accuracy fast direct solvers or low-accuracy robust preconditioners. Numerical results show that we can solve problems with several millions of unknowns in a couple of seconds on a single GPU.
Chao Chen 0008, Per-Gunnar Martinsson
SC1
2021 PBBFMM3D: A parallel black-box algorithm for kernel matrix-vector multiplication
Chao Chen 0008, Jonghyun Lee 0005, Eric Darve
J. Parallel Distributed Comput.2
2018 A distributed-memory hierarchical solver for general sparse linear systems
Chao Chen 0008, Hadi Pouransari, Sivasankaran Rajamanickam, Erik G. Boman, Eric Darve
Parallel Comput.1