Qi Liu 0065

dblp:95/2446-65 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2024
0009-0004-7083-6644ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021
YearPublicationVenuePosition
2024 swPTS: an efficient parallel Thomas split algorithm for tridiagonal systems on Sunway manycore processors
Min Tian 0005, Qi Liu 0065, Jingshan Pan, Ying Gou, Zanjun Zhang
J. Supercomput.2
2023 AH-TDMA: An Adaptive Heterogeneous Tridiagonal Matrix Algorithm on the New Sunway Supercomputer
abstract
The solution of tridiagonal linear systems is used in in various fields and plays a crucial role in numerical simulations. However, there is few efficient solver for tridiagonal linear systems on the new Sunway supercomputer. Based on a three-dimensional heat conduction problem, we propose an adaptive heterogeneous tridiagonal matrix algorithm (AH-TDMA). Our major innovations include: (1) To address computational hotspots within AH-TDMA, a multi-level parallel approach involving MPI+Athread has been adopted. (2) An adaptive data partitioning scheme has been set up to achieve load balance. (3) Employing direct memory access and establishing shared space between registers and main memory, instead of employing discrete memory access, is done to enhance memory access efficiency. (4) The optimization of the loop structure has been made in adjusting the sequencing of the dual-layered loops to reduce communication overhead. The experimental results show that, with one core group, the AH-TDMA achieves a speedup of 99.6 times for hotspot and total time speedup up to 59.4 times compared to the Parallel and Scalable Library for Tridiagonal Matrix Algorithm (PaScal TDMA). The AH-TDMA is scalable up to 2048 core groups, with a parallel efficiency of 69.2%.
Min Tian 0005, Qi Liu 0065, Zenghui Ren, Yue Liu 0027
ICPADS2
2023 hcaPCG: A Heterogeneous and communication-avoid PCG with Jacobi preconditioner on SW26010-Pro architecture
abstract
Due to its efficiency and versatility, the preconditioned conjugate gradient algorithm has long been a staple in the realm of iterative linear system solvers. In this paper, we proposed an optimized preconditioned conjugate gradient algorithm tailored for the SW26010-Pro manycore processor, The main work includes: Optimizing data block sizes based on the processor’s storage structure; combining thread-level and data-level parallelism, utilizing manual SIMD to improve efficiency; optimizing the memory access pattern by employing direct memory access and overlapping of computation and communication; employing a shared-memory approach to store long vectors across all cores. Furthermore, we design an accelerated algorithm for reduction operations, to avoid data communication. Experimental results show that the hcaPCG yields up to 28.1× speedups on average compared to the original implementation.
Min Tian 0005, Yue Liu 0027, Qi Liu 0065, Jingshan Pan
ICPADS5
2023 SPM-GCN: An adaptive reordering algorithm for sparse LU factorization via GCN
abstract
Sparse LU factorization is a critical kernel in scientific computing and engineering applications. A better nonzero pattern of sparse matrixes can accelerate LU factorization by reordering. Traditionally it’s difficult to predict which non-zero pattern is optimal for a sparse matrix. In this paper, we proposed a graph convolutional neural network (GCN) for adaptively selecting the optimal reordering algorithm from five candidate reordering algorithms in the Matrix preprocessing step, referred to as SPM-GCN. After using our SPM-GCN, the average numerical factorization time outperformed the other five algorithms, and the average numerical factorization time was reduced by 17.3% compared with the default method of swSuperLU.
Min Tian 0005, Huazeng Liu, Qi Liu 0065, Zanjun Zhang
ICPADS4