Chuhe Hong

dblp:383/3074 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2026
0009-0000-5746-164XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Block-Aware Adaptive State Management for Optimistic Parallel Discrete Event Simulation
Gencheng Liu, Chuhe Hong, Xinhai Chen 0001, Qingyang Zhang 0009, Jie Liu 0002
ICS4
2025 VES: Vectorized Sparse General Matrix-Matrix Multiplication on Multi-Core DSPs
abstract
The Sparse General Matrix-Matrix Multiplication (SpGEMM) is widely used in a variety of applications. However, research on optimizing SpGEMM for high-performance digital signal processors (DSPs) has been limited. We present VES, a method to accelerate SpGEMM on multi-core DSPs, using the FT-M7032 platform as a case study. Based on the ESC algorithm, VES enhances computational efficiency through vectorized expansion operations, a double-buffering strategy, and an optimized vectorized sorting method. We provide an in-depth analysis of the bottlenecks in vectorized sorting and introduce an efficient vectorized reduction method that significantly improves instruction-level pipeline throughput. Experimental results show that VES outperforms existing methods HASH, ESC, and SPA by an average of 1.12×, 1.90×, and 22.34× on 1,931 sparse matrices, with maximum speedups of 8.69×, 65.9×, and 821.6×, respectively.
Chuhe Hong, Gencheng Liu, Qingyang Zhang 0009, Xinhai Chen 0001, Jie Liu 0002
ICPP1
2025 An Efficient Adaptive Dual-Threshold Svm Based on Heterogeneous Collaboration
abstract
Support Vector Machine (SVM) is highly effective at processing high-dimensional, nonlinear data. However, more than 90 % of the training time is spent on kernel matrix calculations, and existing approaches encounter challenges in adapting to heterogeneous architectures. This paper presents an adaptive dual-threshold method leveraging heterogeneous collaboration to accelerate kernel matrix computations. We also analyze the impact of the working set size on training time and accuracy in ThunderSVM to minimize training time. Tasks with distinct characteristics are allocated to appropriate computing cores through heterogeneous collaboration, with dynamic load balancing via adaptive dual thresholds. On a CPU-DSP heterogeneous platform, our method delivers an average speedup of$5.52 \times$compared to the optimal CPU-only implementation.
Chuhe Hong, Gencheng Liu, Xinhai Chen 0001, Jie Liu 0002
IPDPS3
2024 SaSpGEMM: Sorting-Avoiding Sparse General Matrix-Matrix Multiplication on Multi-Core Processors
abstract
We propose the SaSpGEMM: a parallel sparse general matrix-matrix multiplication (SpGEMM) to avoid the overhead of sorting. The typical workflow of SpGEMM contains: size prediction, memory allocation, numeric calculation and sorting. However, sorting has always been overlooked as a bottleneck in the performance of SpGEMM. It constitutes an average of 30% in HASH and 55% in ESC during the calculation stage. The key idea behind SaSpGEMM is to leverage the compressed sparse row (CSR) storage format’s feature of increasing index order of elements in the same row which is sorted during preprocessing, preserving intermediate products ordered consistently. To achieve this, we introduce a linked list-based accumulator (LLA) designed for batch insertion while maintaining order with a low time complexity. We provide a comprehensive empirical evidence showing that SaSpGEMM outperforms other methods based on time complexity analysis. Compared to three state-of-the-art methods ESC, SPA, and HASH on both x86 (Intel Xeon Gold 6348) and ARM (Phytium2000+) architectures, our method achieves an average speedup of 2.82x, 5.24x, 1.16x (with a maximum speedup of 34.4x, 195x, 23.3x) on the Intel Xeon Gold 6348. On the Phytium2000+, it achieves an average speedup of 2.21x, 4,65x, 1.05x (with a maximum speedup of 40.93x, 146.7x, 9.03x).
Chuhe Hong, Runzhang Mao, Yuechao Liang, Jie Liu 0002
ICPP1