EDBT 2026 Demo / reviewers in the wild / expert
Guixia He
dblp:73/2702
· DBLP profile ↗
11ranked-venue papers
4as first author
4since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Parallel Dynamic Sparse Approximate Inverse Preconditioning Algorithm on GPUabstractThe dynamic sparse approximate inverse (SPAI) preconditioner has proven to be effective in accelerating the convergence of iterative methods for large linear systems. Recently, accelerating it on graphics processing unit (GPU) has attracted considerable attention due to the fact that the cost of constructing the preconditioner is high. However, the existing parallel dynamic SPAI preconditioning algorithms on GPU are usually ineffective because of the out-of-memory error for large matrices. This motivates us to investigate how to accelerate the construction of dynamic SPAI preconditioners on GPU. In this article, we propose an efficient dynamic SPAI preconditioning algorithm on GPU, called GDSPAI. For our proposed GDSPAI, there are the following novelties: (1) a well-known dynamic SPAI preconditioning algorithm is substantially modified to address the main challenges of parallelization on GPU, (2) a parallel framework of constructing the dynamic SPAI preconditioner on GPU is presented on the basis of the modified dynamic SPAI preconditioning algorithm; and (3) each component of the preconditioner is computed in parallel inside a group of threads. Experimental results show that the proposed GDSPAI is effective for large matrices, and outperforms the popular preconditioning algorithms in three public libraries, as well as a recent parallel static SPAI preconditioning algorithm. Jiaquan Gao, Xinyue Chu, Xiaotong Wu, Jun Wang 0077, Guixia He |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2021 | A new diagonal storage for efficient implementation of sparse matrix-vector multiplication on graphics processing unitabstractSummary The sparse matrix–vector multiplication (SpMV) is of great importance in computational science. For multidiagonal sparse matrices that have many long zero sections or scatter points, a great number of zeros are filled to maintain the diagonal structure when using the popular DIA format to store them. This leads to the performance degradation of the DIA kernel. To alleviate the drawback of DIA, we present a novel diagonal storage format, called RBDCS (diagonal compressed storage based on row‐blocks), for multidiagonal sparse matrices, and thus propose an efficient SpMV kernel that corresponds to RBDCS. Given that the RBDCS kernel codes must be manually rewritten for different multidiagonal sparse matrices, a code generator is presented to automatically generate RBDCS kernel codes. Experimental results show that the proposed RBDCS kernel is effective, and outperforms HYBMV in the CUSPARSE library, and three popular diagonal SpMV kernels: DIA, HDI, and CRSD. Guixia He, Jiaquan Gao |
Concurr. Comput. Pract. Exp. | 1 |
| 2021 | Adaptive diagonal sparse matrix-vector multiplication on GPU
Jiaquan Gao, Renjie Yin, Guixia He |
J. Parallel Distributed Comput. | 4 |
| 2021 | A thread-adaptive sparse approximate inverse preconditioning algorithm on multi-GPUs
Jiaquan Gao, Guixia He |
Parallel Comput. | 3 |
| 2020 | An efficient sparse approximate inverse preconditioning algorithm on GPUabstractSummary The sparse approximate inverse (SPAI) preconditioner has proven to be effective in accelerating the convergence of iterative methods. Recently, accelerating it on the graphics processing unit (GPU) has attracted considerable attention due to the fact that the cost of constructing it is high. This motivates us to investigate how to accelerate the construction of SPAI preconditioners on GPU in this paper. We propose an efficient sparse approximate inverse algorithm on GPU, called SPAI‐Adaptive. For our proposed SPAI‐Adaptive, there are the following novelties: (1) an adaptive thread allocation strategy for SPAI‐Adaptive is proposed to assign the optimal thread number for each column of the preconditioner, and (2) Each component of the preconditioner, which includes finding indices I and J , constructing local submatrix, decomposing the local matrix into QR, and solving the upper triangular linear system, is computed in parallel inside a thread group of GPU. Experimental results show that the proposed SPAI‐Adaptive is effective, and has good performance and high parallelism. Guixia He, Renjie Yin, Jiaquan Gao |
Concurr. Comput. Pract. Exp. | 1 |
| 2018 | Efficient dense matrix-vector multiplication on GPUabstractSummary Given that the dense matrix‐vector multiplication (Ax or ATx) is of great importance in scientific computations, how to accelerate it is investigated on the graphics processing unit (GPU) in this paper. We present a warp‐based implementation of Ax on the GPU, called GEMV‐Adaptive, and a thread‐based implementation of ATx on the GPU, called GEMV‐T‐Adaptive. For our proposed GEMV‐Adaptive and GEMV‐T‐Adaptive, there are the following novelties: (1) an adaptive warp allocation strategy for GEMV‐Adaptive is proposed to assign the optimal warp number for each matrix row, (2) an adaptive thread allocation strategy for GEMV‐T‐Adaptive is designed to assign the optimal thread number to each matrix row, and (3) several optimization schemes are formulated. Experimental results show that the proposed GEMV‐Adaptive and GEMV‐T‐Adaptive mitigate the performance fluctuations of the implementations in the CUBLAS library, always have high performance, and outperform the most recently proposed GEMV and GEMV‐T kernels by Gao et al, respectively, for all test matrices. Guixia He, Jiaquan Gao, Jun Wang 0077 |
Concurr. Comput. Pract. Exp. | 1 |
| 2017 | A multi-GPU parallel optimization model for the preconditioned conjugate gradient algorithm
Jiaquan Gao, Yuanshen Zhou, Guixia He |
Parallel Comput. | 3 |
| 2013 | Modified Incomplete Cholesky Preconditioned Conjugate Gradient Algorithm on GPU for the 3D Parabolic Equation
Jiaquan Gao, Guixia He |
NPC | 3 |
| 2010 | A Quantum-Inspired Artificial Immune System for Multiobjective 0-1 Knapsack Problems
Jiaquan Gao, Guixia He |
ISNN (1) | 3 |
| 2009 | A Novel Weight-Based Immune Genetic Algorithm for Multiobjective Optimization Problems
Guixia He, Jiaquan Gao |
ISNN (2) | 1 |
| 2008 | Multi-objective scheduling problems subjected to special process constraintabstractThe problem of parallel machine multi-objective scheduling subjected to special process constraint in the textile industries, as one of the most important combinational optimization problems, is different from other parallel machine scheduling problems in the following characteristics. On one hand, processing machines are non-identical; on the other hand, the sort of job processed on every machine can be restricted Considering one of the multi-objective problems, either minimizing the maximum completion time among all the machines(makespan) or minimizing the total earliness/tardiness penalty of all the jobs has been cornerstone of most studies done so far. However, under special process constraint, taking them into account as a multi-objective problem has not been well studied Therefore, in this paper, a multi-objective model based on them is presented and a new parallel genetic algorithm based on a vector group coding method is also proposed in order to effectively solve this model. The algorithm shows the following advantages: the coding method is simple and can effectively reflect the virtual scheduling policy, which can vividly reflect the numbers and sequences of these processed jobs on every machine, and then enables the individuals generated by crossover and mutation to satisfy process constraint. Numerical experiments show that it is efficient, and is better than the common genetic algorithm, and has the better parallel efficiency. A much better prospect of application can be optimistically expected. Jiaquan Gao, Guixia He, Yushun Wang |
IEEE Congress on Evolutionary Computation | 2 |