Jianqi Zhao 0001

dblp:238/2725-1 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2023
0009-0001-6731-4317ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
High-performance computing · 50% Parallel and multicore computing · 28% Electronic design automation · 11%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing › sparse linear solver
sparse direct solver
1.222023
PanguLU: A Scalable Regular Two-Dimensional Block-Cyclic Sparse Direct Solver on Distributed Heterogeneous Systems · SC 2023
SFLU: Synchronization-Free Sparse LU Factorization for Fast Circuit Simulation on GPUs · DAC 2021
Parallel and multicore computing › data distribution
block-cyclic distribution
0.712023
PanguLU: A Scalable Regular Two-Dimensional Block-Cyclic Sparse Direct Solver on Distributed Heterogeneous Systems · SC 2023
Parallel and multicore computing › parallelization strategies
distributed-memory parallelization
0.712023
PanguLU: A Scalable Regular Two-Dimensional Block-Cyclic Sparse Direct Solver on Distributed Heterogeneous Systems · SC 2023
High-performance computing
scientific computing systems
0.712023
PanguLU: A Scalable Regular Two-Dimensional Block-Cyclic Sparse Direct Solver on Distributed Heterogeneous Systems · SC 2023
Electronic design automation
circuit simulation
0.512021
SFLU: Synchronization-Free Sparse LU Factorization for Fast Circuit Simulation on GPUs · DAC 2021
GPUs and heterogeneous computing
GPU computing
0.512021
SFLU: Synchronization-Free Sparse LU Factorization for Fast Circuit Simulation on GPUs · DAC 2021
High-performance computing › sparse linear solver
sparse LU factorization
0.512021
SFLU: Synchronization-Free Sparse LU Factorization for Fast Circuit Simulation on GPUs · DAC 2021

Methods — techniques the papers use, named apart from their topics

supernodal method · 0.7multifrontal method · 0.7BLAS · 0.7synchronization-free elimination · 0.5level-set parallelization · 0.5
YearPublicationVenuePosition
2023 PanguLU: A Scalable Regular Two-Dimensional Block-Cyclic Sparse Direct Solver on Distributed Heterogeneous Systems
abstract
Sparse direct solvers play a vital role in large-scale high performance computing in science and engineering. Existing distributed sparse direct methods employ multifrontal/supernodal patterns to aggregate columns of nearly identical forms and to exploit dense basic linear algebra subprograms (BLAS) for computation. However, such a data layout may bring more unevenness when the structure of the input matrix is not ideal, and using dense BLAS may waste many floating-point operations on zero fill-ins.
Xu Fu, Bingbin Zhang, Tengcheng Wang, Wenhao Li 0020, Yuechen Lu, Enxin Yi, Jianqi Zhao 0001, Xiaohan Geng, Fangying Li, Zhou Jin 0001, Weifeng Liu 0002
SC7
2021 SFLU: Synchronization-Free Sparse LU Factorization for Fast Circuit Simulation on GPUs
abstract
Sparse LU factorization is one of the key building blocks of sparse direct solvers and often dominates the computing time of circuit simulation programs. Existing GPU-accelerated sparse LU factorization methods either offload relatively small dense matrix-matrix multiplications to GPU cores, or extract level-set information to parallelize elimination operations in each level. However, because of the insufficient parallelism, neither of the methods can saturate a large amount of compute units on modern GPUs.We in this paper propose a synchronization-free sparse LU factorization algorithm called SFLU. To saturate GPU cores, our method lets each thread block eliminate a column and runs all the thread blocks at the same time. Through communicating dependency information stored on global memory, all the thread blocks either busy wait to run or get updated by their previous columns. Because elimination of all the columns work concurrently, our method avoids any barrier synchronization and saturates GPU resources. By benchmarking over 1000 sparse matrices on an NVIDIA Titan RTX GPU, our SFLU outperforms SuperLU and GLU by a factor of on average 155.71 and 8.21 (up to 3585.62 and 252.66), respectively.
Jianqi Zhao 0001, Zhou Jin 0001, Weifeng Liu 0002, Zhenya Zhou
DAC1