Deshun Bi

dblp:370/1918 · DBLP profile ↗
← Back
4ranked-venue papers
4as first author
4since 2021 · last 2026
0009-0009-4488-5473ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 first-author · 4 since 2021
YearPublicationVenuePosition
2026 A Memory-Aware Sparse Matrix-Matrix Multiplication on Multicore Architectures
abstract
Sparse matrix–matrix multiplication (SpMM) is a fundamental operation in scientific computing with broad applications across numerous domains. Tiling is a key optimization technique for improving data locality and is widely adopted in high-performance computing. However, the irregular data access patterns inherent to SpMM make it challenging to exploit tiling effectively for data reuse. In this article, we propose MaSpMM , a memory-aware SpMM framework that integrates cache-aware tiling with a segment-oriented data layout. MaSpMM stores matrices as continuous segments to enhance data locality within each tile. Moreover, since many sparse matrices in real-world applications exhibit symmetry, we further develop MaSpMM-Sym, an extension that recursively partitions symmetric matrices to eliminate write conflicts and further improve locality. To adapt to diverse scenarios, we finally introduce MaSpMM-Adap, which adaptively selects the most suitable approach for each input matrix. Comprehensive evaluations on both x86 and ARM CPUs demonstrate that MaSpMM-Adap achieves average speedups of up to 1.86× over Intel oneMKL, 1.84× over ASpT, and 1.75× over J-Stream.
Deshun Bi, Shengguo Li, Haozhong Qiu, Chuanfu Xu, Xiaojian Yang, Dezun Dong, Tiaojie Xiao, Jie Liu 0002
ACM Trans. Archit. Code Optim.1
2024 Optimizing SpMV on Heterogeneous Multi-Core DSPs through Improved Locality and Vectorization
abstract
The sparse matrix-vector multiplication (SpMV) is widely used in large-scale scientific computing and engineering. However, optimizing SpMV for high-performance digital signal processors (DSPs) has received limited attention. We present HaLAV, a method to accelerate SpMV on CPU-DSP heterogeneous platforms, using the FT-M7032 DSP platform as a case study. HaLAV partitions the input matrix into ‘dense’ and ‘sparse’ parts through column reordering. For the dense part, HaLAV automatically selects storage formats optimized for vectorization to run on the DSP. At the same time, it offloads the sparse component to be processed by the CPU using the standard CSR algorithm. We evaluate our approach on the FT-M7032 platform and an Intel Xeon CPU. Experimental results show that our techniques achieve average speedups of 2.09 × and 1.66 × over the competing baselines on the FT-M7032 and the Xeon platform, respectively.
Deshun Bi, Shengguo Li, Dezun Dong, Peng Zhang 0061, Jianbin Fang
ICPP1
2023 Efficiently Running SpMV on Multi-core DSPs for Banded Matrix
Deshun Bi, Shengguo Li, Xiaojian Yang, Dezun Dong
ICA3PP (5)1
2023 Efficiently Running SpMV on Multi-Core DSPs for Block Sparse Matrix
abstract
Sparse Matrix-Vector Multiplication (SpMV) is a fundamental operation in sparse computations. Although many techniques have been developed to speed up SpMV, optimizing this process on low-power multicore digital signal processors (DSPs) has often been neglected. This paper presents the FT-M7032, a cutting-edge CPU-DSP hybrid multi-core processor. We assess the data transfer efficiency among various units to identify performance constraints of SpMV on multicore DSPs. Based on our evaluation, we develop a method for block sparse matrices, namely SpMV_BLOCK, which can break the bandwidth bottleneck of SpMV and achieve significant performance gains. We then propose a load-balancing strategy for each thread by using binary search and devise a pipeline that overlaps data transfers and computations to improve SpMV performance. To measure our method’s effectiveness, we compared its performance against a baseline on the FT-M7032’s general-purpose CPU cores. Our experiments show that our approach delivers a notable 5.80 × speedup over the baseline.
Deshun Bi, Xiaowen Tian, Shengguo Li, Dezun Dong
ICPADS1