Shengle Lin

dblp:304/5542 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0003-3329-0924ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021
YearPublicationVenuePosition
2025 SSpMV: A Sparsity-aware SpMV Framework Empowered by Multimodal Machine Learning
abstract
Sparse Matrix-Vector Multiplication (SpMV) is an essential sparse operation in scientific computing and artificial intelligence. Efficiently adapting SpMV algorithms to diverse matrices and architectures requires a framework capable of accurately recognizing sparse patterns and selecting the optimal implementation. In this work, we introduce Sparsity-aware SpMV (SSpMV), a framework that integrates expert-designed features with multimodal representations to adaptively predict the best-performing algorithm and parameters. For this purpose, we design a multimodal neural network called MM-Adapter, to capture diverse modalities to represent the computational features of SpMV. Experimental results demonstrate that MMAdapter achieves the highest accuracy of $81.05 \%$, outperforming existing SpMV prediction models. Furthermore, SSpMV consistently delivers substantial performance improvements over state-of-the-art sparse libraries across various multi-core platforms.
Shengle Lin, Chubo Liu, Yan Ding 0004, Joey Tianyi Zhou, Kenli Li 0001, Wangdong Yang
DAC1
2025 MM-AutoSolver: A multimodal machine learning method for the auto-selection of iterative solvers and preconditioners
Hantao Xiong, Wangdong Yang, Weiqing He, Shengle Lin, Keqin Li 0001, Kenli Li 0001
J. Parallel Distributed Comput.4
2025 High Performance OpenCL-Based GEMM Kernel Auto-Tuned by Bayesian Optimization
Shengle Lin, Guoqing Xiao 0001, Haotian Wang 0006, Wangdong Yang, Kenli Li 0001, Keqin Li 0001
IEEE Trans. Parallel Distributed Syst.1
2024 Parallel algorithm design and optimization of geodynamic numerical simulation application on the Tianhe new-generation high-performance computer
Wangdong Yang, Ruixuan Qi, Qinyun Tsai, Shengle Lin, Fengkun Dong, Kenli Li 0001, Keqin Li 0001
J. Supercomput.5
2021 STM-multifrontal QR: streaming task mapping multifrontal QR factorization empowered by GCN
abstract
Multifrontal QR algorithm, which consists of symbolic analysis and numerical factorization, is a high-performance algorithm for orthogonal factorizing sparse matrix. In this work, a graph convolutional network (GCN) for adaptively selecting the optimal reordering algorithm is proposed in symbolic analysis. Using our GCN adaptive classifier, the average numerical factorization time is reduced by 20.78% compared with the default approach, and the additional memory overhead is approximately 4% higher than that of prior work. Moreover, for numerical factorization, an optimized tasks stream parallel processing strategy is proposed and a more efficient computing task mapping framework for NUMA architecture is adopted in this paper, which called STM-Multifrontal QR factorization. Numerical experiments on the TaiShan Server show average 1.22x performance gains over the original SuiteSparseQR. Nearly 80% of datasets have achieved better performance compared with the MKL sparse QR on Intel Xeon 6248.
Shengle Lin, Wangdong Yang, Haotian Wang 0006, Qinyun Tsai, Kenli Li 0001
SC1