Jinliang Shi

dblp:89/7606 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0001-5420-906XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Omnia: Efficient RAG Serving through Speculative Scheduling
abstract
Retrieval-Augmented Generation (RAG) has emerged for enhancing Large Language Models (LLMs) by improving factual accuracy and mitigating hallucinations. A typical RAG pipeline executes in three cascaded stages: retrieval, reranking, and generation. The existing serving systems suffer from two critical system-level bottlenecks when applying to RAG serving: the first is the cumulative latency caused by rigid sequential dependencies between reranking and generation, and the second is the system saturation triggered by bursty, high fan-in reranking workloads.
Rongtian Fu, Shigang Li 0002, Youxuan Xu, Tong Wu 0024, Jinliang Shi
HPDC6
2025 FlashSparse: Minimizing Computation Redundancy for Fast Sparse Matrix Multiplications on Tensor Cores
abstract
Sparse Matrix-matrix Multiplication (SpMM) and Sampled Dense-dense Matrix Multiplication (SDDMM) are important sparse operators in scientific computing and deep learning. Tensor Core Units (TCUs) enhance modern accelerators with superior computing power, which is promising to boost the performance of matrix operators to a higher level. However, due to the irregularity of unstructured sparse data, it is difficult to deliver practical speedups on TCUs. To this end, we propose FlashSparse, a novel approach to bridge the gap between sparse workloads and the TCU architecture. Specifically, FlashSparse minimizes the sparse granularity for SpMM and SDDMM on TCUs through a novel swap-and-transpose matrix multiplication strategy. Benefiting from the minimum sparse granularity, the computation redundancy is remarkably reduced while the computing power of TCUs is fully utilized. Besides, FlashSparse is equipped with a memory-efficient thread mapping strategy for coalesced data access and a sparse matrix storage format to save memory footprint. Extensive experimental results on H100 and RTX 4090 GPUs show that FlashSparse sets a new state-of-the-art for sparse matrix multiplications (geometric mean 5.5x speedup over DTC-SpMM and 3.22x speedup over RoDe).
Jinliang Shi, Shigang Li 0002, Youxuan Xu, Rongtian Fu, Xueying Wang 0003, Tong Wu 0024
PPoPP1
2024 Parallel optimization and application of unstructured sparse triangular solver on new generation of Sunway architecture
Jianjiang Li, Wei Xue 0003, Jiabi Liang, Jinliang Shi
Parallel Comput.6
2024 Toward efficient structured-grid triangular solver on sunway many-core processors
Jianjiang Li, Jiabi Liang, Wei Xue 0003, Zhengding Hu, Jinliang Shi
J. Supercomput.6
2023 A parallel and balanced SVM algorithm on spark for data-intensive computing
abstract
Support Vector Machine (SVM) is a machine learning with excellent classification performance, which has been widely used in various fields such as data mining, text classification, face recognition and etc. However, when data volume scales to a certain level, the computational time becomes too long and the efficiency becomes low. To address this issue, we propose a parallel balanced SVM algorithm based on Spark, named PB-SVM, which is optimized on the basis of the traditional Cascade SVM algorithm. PB-SVM contains three parts, i.e., Clustering Equal Division, Balancing Shuffle and Iteration Termination, which solves the problems of data skew of Cascade SVM and the large difference between local support vector and global support vector. We implement PB-SVM in AliCloud Spark distributed cluster with five kinds of public datasets. Our experimental results show that in the two-classification test on the dataset covtype, compared with MLlib-SVM and Cascade SVM on Spark, PB-SVM improves efficiency by 38.9% and 75.4%, and the accuracy is improved by 7.16% and 8.38%. Moreover, in the multi-classification test, compared with Cascade SVM on Spark on the dataset covtype, PB-SVM improves efficiency and accuracy by 94.8% and 18.26% respectively.
Jianjiang Li, Jinliang Shi, Can Feng
Intell. Data Anal.2
2023 New YARN sharing GPU based on graphics memory granularity scheduling
Jinliang Shi, Dewu Chen, Jiabi Liang, Jianjiang Li
Parallel Comput.1