VLDB 2026 Research / reviewers in the wild / expert
Jianhua Gao 0001
dblp:92/1608-1
· DBLP profile ↗
16ranked-venue papers
7as first author
14since 2021 · last 2026
0000-0002-3828-0015ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 6 first-author · 13 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ARROW: Adaptive Row Reorganization for Warp-Balanced SpMV
Jianhua Gao 0001, Weixing Ji |
APPT | 2 |
| 2026 | ApproxRAG: A Systematic Framework for Mitigating I/O Overheads in Retrieval-Augmented Generation
Danying Ge, Qizhi Jiang, Jianhua Gao 0001, Weixing Ji |
Euro-Par (2) | 3 |
| 2026 | HRPF: A parallel programming framework for recursive algorithms on heterogeneous CPU-GPU systems
Yizhuo Wang 0001, Senhao Shao, Jianhua Gao 0001, Weixing Ji, Hongbo Xing |
Parallel Comput. | 4 |
| 2025 | Adaptive point cloud compression based on precision-aware floating-point encoding
Yanpeng Han, Yizhuo Wang 0001, Fawang Liu, Jianhua Gao 0001, Weixing Ji |
CCF Trans. High Perform. Comput. | 4 |
| 2025 | RaNAS: Resource-Aware Neural Architecture Search for Edge ComputingabstractNeural architecture search (NAS) for edge devices is often time-consuming because of long-latency deploying and testing on edge devices. The ability to accurately predict the computation cost and memory requirement for convolutional neural networks (CNNs) in advance holds substantial value. Existing work primarily relies on analytical models, which can result in high prediction errors. This article proposes a resource-aware NAS (RaNAS) model based on various features. Additionally, a new graph neural network is introduced to predict inference latency and maximum memory requirements for CNNs on edge devices. Experimental results show that, within the error bound of ±1%, RaNAS achieves an accuracy improvement of approximately 8% for inference latency prediction and about 25% for maximum memory occupancy prediction over the state-of-the-art approaches. Jianhua Gao 0001, Zeming Liu, Yizhuo Wang 0001, Weixing Ji |
ACM Trans. Archit. Code Optim. | 1 |
| 2025 | PTPS: Precision-Aware Task Partitioning and Scheduling for SpMV on CPU-FPGA Heterogeneous PlatformsabstractThe CPU-FPGA heterogeneous computing architecture is extensively employed in the embedded domain due to its low cost and power efficiency, with numerous sparse matrix-vector multiplication (SpMV) acceleration efforts already targeting this architecture. However, existing work rarely includes collaborative SpMV computations between CPU and FPGA, which limits the exploration of hybrid architectures that could potentially offer enhanced performance and flexibility. This article introduces an FPGA architecture design that supports multiprecision SpMV computations, including FP16, FP32, and FP64. Building on this, PTPS, a precision-aware SpMV task partitioning and dynamic scheduling algorithm tailored for the CPU-FPGA heterogeneous architecture, is proposed. The core idea of PTPS is lossless partitioning of sparse matrices across multiple precisions, prioritizing low-precision SpMV computations on the FPGA and high-precision computations on the CPU. PTPS not only leverages the strengths of CPU and FPGA for collaborative SpMV computations but also reduces data transmission overhead between them, thereby improving the overall computational efficiency. Experimental evaluation demonstrates that the proposed approach offers an average speedup of$1.57\times $over the CPU-only approach and$2.58\times $over the FPGA-only approach. Jianhua Gao 0001, Xingze Huang, Yizhuo Wang 0001, Weixing Ji |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | Load Balancing Optimizations for Distributed GMRES Algorithm
Shuaizhe Guo, Jianhua Gao 0001, Weixing Ji, Yizhuo Wang 0001 |
ICA3PP (6) | 3 |
| 2024 | pSpMv: precision-based sparse matrix partition and SpMV optimizationabstractAbstract The new generation of computing devices tends to support multiple floating-point formats and different computing precision. Besides single and double precision, half precision is embraced and widely supported by new computing devices. Low-precision representations have compact memory size and lightweight computing strength, and they also bring opportunities to the optimization of BLAS routines. This paper proposes a new sparse matrix partition approach based on IEEE 754 standard floating-point format. An input sparse matrix in double precision is partitioned and transformed into several sub-matrices in different precision without loss of accuracy. Most non-zero elements can be stored in half or single precision, if the most significant bits of exponent and the least significant bits of mantissa are zeros in double-precision representation. Based on this mixed-precision representation of sparse matrix, we also present a new SpMV algorithm pSpMV for GPU devices. pSpMV not only reduces the memory access overhead, but also reduces the computing strength of floating-point numbers. Experimental results on two GPU devices show that pSpMV achieves a geometric mean speedup of 1.39x on Tesla V100 and 1.45x on Tesla P100 over double-precision SpMV for 2,554 sparse matrices. Yizhuo Wang 0001, Jianhua Gao 0001, Weixing Ji |
CCF Trans. High Perform. Comput. | 3 |
| 2024 | Revisiting thread configuration of SpMV kernels on GPU: A machine learning based approach
Jianhua Gao 0001, Weixing Ji, Yizhuo Wang 0001, Feng Shi 0009 |
J. Parallel Distributed Comput. | 1 |
| 2024 | Optimization of Large-Scale Sparse Matrix-Vector Multiplication on Multi-GPU SystemsabstractSparse matrix-vector multiplication (SpMV) is one of the important kernels of many iterative algorithms for solving sparse linear systems. The limited storage and computational resources of individual GPUs restrict both the scale and speed of SpMV computing in problem-solving. As real-world engineering problems continue to increase in complexity, the imperative for collaborative execution of iterative solving algorithms across multiple GPUs is increasingly apparent. Although the multi-GPU-based SpMV takes less kernel execution time, it also introduces additional data transmission overhead, which diminishes the performance gains derived from parallelization across multi-GPUs. Based on the non-zero elements distribution characteristics of sparse matrices and the tradeoff between redundant computations and data transfer overhead, this article introduces a series of SpMV optimization techniques tailored for multi-GPU environments and effectively enhances the execution efficiency of iterative algorithms on multiple GPUs. First, we propose a two-level non-zero elements-based matrix partitioning method to increase the overlap of kernel execution and data transmission. Then, considering the irregular non-zero elements distribution in sparse matrices, a long-row-aware matrix partitioning method is proposed to hide more data transmissions. Finally, an optimization using redundant and inexpensive short-row execution to exchange costly data transmission is proposed. Our experimental evaluation demonstrates that, compared with the SpMV on a single GPU, the proposed method achieves an average speedup of 2.00× and 1.85× on platforms equipped with two RTX 3090 and two Tesla V100-SXM2, respectively. The average speedup of 2.65× is achieved on a platform equipped with four Tesla V100-SXM2. Jianhua Gao 0001, Weixing Ji, Yizhuo Wang 0001 |
ACM Trans. Archit. Code Optim. | 1 |
| 2024 | Optimization of Sparse Matrix Computation for Algebraic Multigrid on GPUsabstractAMG is one of the most efficient and widely used methods for solving sparse linear systems. The computational process of AMG mainly consists of a series of iterative calculations of generalized sparse matrix-matrix multiplication (SpGEMM) and sparse matrix-vector multiplication (SpMV). Optimizing these sparse matrix calculations is crucial for accelerating solving linear systems. In this paper, we first focus on optimizing the SpGEMM algorithm in AmgX, a popular AMG library for GPUs. We propose a new algorithm called SpGEMM-upper, which achieves an average speedup of 2.02× on Tesla V100 and 1.96× on RTX 3090 against the original algorithm. Next, through experimental investigation, we conclude that no single SpGEMM library or algorithm performs optimally for most sparse matrices, and the same holds true for SpMV. Therefore, we build machine learning-based models to predict the optimal SpGEMM and SpMV used in the AMG calculation process. Finally, we integrate the prediction models, SpGEMM-upper, and other selected algorithms into a framework for adaptive sparse matrix computation in AMG. Our experimental results prove that the framework achieves promising performance improvements on the test set. Yizhuo Wang 0001, Fangli Chang, Bingxin Wei, Jianhua Gao 0001, Weixing Ji |
ACM Trans. Archit. Code Optim. | 4 |
| 2022 | TaiChi: A Hybrid Compression Format for Binary Sparse Matrix-Vector Multiplication on GPUabstractBinary Sparse Matrix-Vector Multiplication (SpMV) is a heavy computational kernel in weblink analysis, integer factorization, compressed sensing, spectral graph theory, and other domains. Testing several popular GPU-based SpMV implementations on 400 sparse matrices, we observed that data transfer to GPU memory accounts for a large part of the total computation time. The transfer of constant value “1”s can be easily eliminated for binary sparse matrices. However, compressing index arrays has always been a great challenge. This article proposes a new compression format TaiChi to further reduce index data copies and improve the performance of SpMV, especially for diagonally dominant binary sparse matrices. Input matrices are first partitioned into relatively dense and ultra-sparse areas. Then the dense areas are encoded inversely by marking “0”s, while the ultra-sparse area is encoded by marking “1”s. We also designed a new SpMV algorithm only using addition and subtraction for binary matrices based on our partition and encoding format. Evaluation results on real-world binary sparse matrices show that our hybrid encoding for binary matrix significantly reduces the data transfer and speeds up the kernel execution. It achieves the highest transfer and kernel execution speedups of 5.63x and 3.84x on GTX 1080 Ti, 3.39x and 3.91x on Tesla V100. Jianhua Gao 0001, Weixing Ji, Zhaonian Tan, Yizhuo Wang 0001, Feng Shi 0009 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2021 | AMF-CSR: Adaptive Multi-Row Folding of CSR for SpMV on GPUabstractSpMV is a cost-dominant operation used in many iterative methods for solving large-scale sparse linear systems. However, irregular memory access of SpMV to the multiplied vector leads to low data locality and then harms the performance. This paper presents an adaptive multi-row folding of CSR (AMF-CSR) format for SpMV calculation on GPU. This new storage format supports the folding of the variable number of rows in order to achieve better load balancing in computation. AMF-CSR not only increases the density of non-zero elements in a folded row, thereby improving the access locality of the multiplied vector, but also merges an approximately equal number of nonzero elements in a folded row, hence achieving load balancing. The performance evaluation using 28 sparse matrices shows that the proposed SpMV algorithm based on AMF-CSR achieves the highest speedup of 4.11x and 3.62x on GTX 1080 Ti and Tesla V100 respectively against a fixed multi-row folding-based SpMV algorithm. Evaluation results using 450 regular sparse matrices and 450 irregular sparse matrices also show that AMF-CSR is superior to other SpMV implementations. Jianhua Gao 0001, Weixing Ji, Senhao Shao, Yizhuo Wang 0001, Feng Shi 0009 |
ICPADS | 1 |
| 2021 | Towards Optimal Fast Matrix Multiplication on CPU-GPU Platforms
Senhao Shao, Yizhuo Wang 0001, Weixing Ji, Jianhua Gao 0001 |
PDCAT | 4 |
| 2020 | MMSparse: 2D partitioning of sparse matrix based on mathematical morphology
Zhaonian Tan, Weixing Ji, Jianhua Gao 0001, Yueyan Zhao, Akrem Benatia, Yizhuo Wang 0001, Feng Shi 0009 |
Future Gener. Comput. Syst. | 3 |
| 2020 | Cube-based incremental outlier detection for streaming computing
Jianhua Gao 0001, Weixing Ji, Anmin Li, Yizhuo Wang 0001, Zongyu Zhang |
Inf. Sci. | 1 |