Huiyuan Li 0002

dblp:85/2098-2 · DBLP profile ↗
← Back
10ranked-venue papers
0as first author
10since 2021 · last 2026
0000-0002-6326-9926ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Factorization-in-Loop: Proximal Fill-in Minimization for Sparse Matrix Reordering
abstract
Fill-ins are new nonzero elements in the summation of the upper and lower triangular factors generated during LU factorization. For large sparse matrices, they will increase the memory usage and computational time, and be reduced through proper row or column arrangement, namely matrix reordering. Finding a row or column permutation with the minimal fill-ins is NP-hard, and surrogate objectives are designed to derive fill-in reduction permutations or learn a reordering function. However, there is no theoretical guarantee between the golden criterion and these surrogate objectives. Here we propose to learn a reordering network by minimizing l1 norm of triangular factors of the reordered matrix to approximate the exact number of fill-ins. The reordering network utilizes a graph encoder to predict row or column node scores. For inference, it is easy and fast to derive the permutation from sorting algorithms for matrices. For gradient based optimization, there is a large gap between the predicted node scores and resultant triangular factors in the optimization objective. To bridge the gap, we first design two reparameterization techniques to obtain the permutation matrix from node scores. The matrix is reordered by multiplying the permutation matrix. Then we introduce the factorization process into the objective function to arrive at target triangular factors. The overall objective function is optimized with the alternating direction method of multipliers and proximal gradient descent. Experimental results on benchmark sparse matrix collection SuiteSparse show the fill-in and LU factorization time reduction of our proposed method is 0.2% and 17.8% compared with state-of-the-art baselines.
Shuzi Niu, Huiyuan Li 0002, Wenjia Wu
AAAI4
2026 Self-supervised Learning for Sparse Matrix Reordering
Fangfang Liu 0004, Shuzi Niu, Huiyuan Li 0002, Wenjia Wu
DASFAA (6)5
2026 AMGNet: Dual-Domain Scale-Specific Graph Learning for Multivariate Time Series Forecasting
Shuzi Niu, Huiyuan Li 0002, Changyou Zhang, Changmao Wu
ICIC (3)3
2025 Bridging the Gap Between Sparse Matrix Reordering and Factorization: A Deep Learning Framework for Fill-in Reduction
Shuzi Niu, Huiyuan Li 0002
DASFAA (1)4
2025 Multigrid-Inspired Graph Neural Networks for Selection of Solvers and Preconditioners in Sparse Linear Systems
abstract
Solving sparse linear systems is a fundamental task in science and engineering computing. Multiple iterative solvers have been developed to iteratively refining an initial guess to con-verge to the solution, with preconditioning techniques for better convergence. However, it’s challenging to select a quasi-optimal solver and preconditioner without background knowledge and domain expertise to reduce time and space cost. Meanwhile, a suboptimal selection may incur exponential computational overhead and even solution divergence. We designed a graph neural network model to help select optimal solver and preconditioner for sparse linear systems, where a multigrid-inspired GNN structure was introduced to synergistically aggregate local matrix patterns and global spectral features through hierarchical structure. Experimental results on benchmark sparse matrix collection show that our model outperforms state-of-the-art baseline selectors.
Shuzi Niu, Huiyuan Li 0002
IJCNN3
2025 DH_Aligner: A fast short-read aligner on multicore platforms with AVX vectorization
Qiao Sun 0005, Leisheng Li, Huiyuan Li 0002
J. Parallel Distributed Comput.4
2024 A novel HPL-AI approach for FP16-only accelerator and its instantiation on Kunpeng+Ascend AI-specific platform
Zijian Cao 0006, Qiao Sun 0005, Changcheng Song, Huiyuan Li 0002
J. Parallel Distributed Comput.6
2023 Evolving the HPL benchmark towards multi-GPGPU clusters
Qiao Sun 0005, Wenjing Ma, Jiachang Sun, Huiyuan Li 0002
CCF Trans. High Perform. Comput.4
2023 MFFT: A GPU Accelerated Highly Efficient Mixed-Precision Large-Scale FFT Framework
abstract
Fast Fourier transform (FFT) is widely used in computing applications in large-scale parallel programs, and data communication is the main performance bottleneck of FFT and seriously affects its parallel efficiency. To tackle this problem, we propose a new large-scale FFT framework, MFFT, which optimizes parallel FFT with a new mixed-precision optimization technique, adopting the “high precision computation, low precision communication” strategy. To enable “low precision communication”, we propose a shared-exponent floating-point number compression technique, which reduces the volume of data communication, while maintaining higher accuracy. In addition, we apply a two-phase normalization technique to further reduce the round-off error. Based on the mixed-precision MFFT framework, we apply several optimization techniques to improve the performance, such as streaming of GPU kernels, MPI message combination, kernel optimization, and memory optimization. We evaluate MFFT on a system with 4,096 GPUs. The results show that shared-exponent MFFT is 1.23 × faster than that of double-precision MFFT on average, and double-precision MFFT achieves performance 3.53× and 9.48× on average higher than open source library 2Decomp&FFT (CPU-based version) and heFFTe (AMD GPU-based version), respectively. The parallel efficiency of double-precision MFFT increased from 53.2% to 78.1% compared with 2Decomp&FFT, and shared-exponent MFFT further increases the parallel efficiency to 83.8%.
Fangfang Liu 0004, Wenjing Ma, Huiyuan Li 0002, Yuanchi Peng
ACM Trans. Archit. Code Optim.4
2021 Topological Interpretable Multi-scale Sequential Recommendation
Shuzi Niu, Huiyuan Li 0002
DASFAA (3)3