EDBT 2026 Demo / reviewers in the wild / expert
Yi Zong
dblp:36/11467
· DBLP profile ↗
6ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0001-7179-6593ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Entropy Model of FIRO-Based TRNGs With Differential Structure
Huiying Han, Lihua Dong, Yi Zong |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2025 | A Fast Sparse Triangular Solve for Structured-grid Problems on Heterogeneous ProcessorsabstractStructured-grid problems are common in scientific computing, particularly in applications like fluid dynamics and electromagnetic simulation. One of the key kernels in solving these problems is Sparse Triangular Solve (SpTRSV), which often becomes a performance bottleneck due to its low computing intensity and inherent internal data dependencies. In structured-grid SpTRSV, the regularity of non-zero distributions and the high parallelism of sparse matrices present opportunities to harness the architectural strengths of modern heterogeneous processors. However, existing SpTRSV algorithms fail to fully exploit these advantages, due to their mismatches in data dependencies, computational order, and memory layouts. In this paper, we introduce a novel SpTRSV algorithm tailored for structured-grids on modern heterogeneous processors. Our approach introduces a two-level blocking strategy to enhance data locality and reduce communication overhead, while a vertical tiling-based pipeline balances parallelism with computational granularity. Additionally, we design hardware-specific adaptive scheduling strategies to accommodate varying degrees of parallelism across distinct architectures. The algorithm has been implemented on two types of heterogeneous processors, NVIDIA GPUs and SW26010-Pro, with hardware-specific optimizations to further improve the performance. Experimental results show that our implementations achieve speedups of more than 1.87 × over state-of-the-art baselines and provide efficient end-to-end solutions with lightweight preprocessing. Zhengding Hu, Yi Zong, Jingwei Sun 0001, Wei Xue 0003, Guangzhong Sun |
ICPP | 2 |
| 2025 | An Efficient 2D Fusion Method for High-Performance Two-Stage Eigensolvers on Modern Heterogeneous ArchitecturesabstractSolving a significant portion of the eigensystem is a critical problem in numerical linear algebra and is widely applied in real-world applications.As problem sizes increase, the twostage tridiagonalization method has emerged as the state-ofthe-art approach and has been implemented in well-known libraries such as LAPACK, PLASMA, and MAGMA.Its major performance bottleneck is the tridiagonal-to-band back transformation of eigenvectors (st2sb) due to the dilemma between limited operational intensity and excessive computational cost.This challenge is further exacerbated by the growing imbalance between computational speed and memory bandwidth in modern heterogeneous architectures.To address this challenge, this paper introduces a 2D Fusion method to decouple the operational intensity from the computational cost of st2sb.To reduce the intrinsic overhead of 2D Fusion for large fusion factors, we further propose an effective skipping strategy.Our 2D Fusion enhances the performance of all existing two-stage eigensolvers without loss of accuracy.We evaluated the effectiveness of 2D Fusion in MAGMA and LAPACK across various problem sizes: on the Nvidia A100 GPU, 2D Fusion improves the performance of eigenvalue decomposition in MAGMA by an average speedup of 1.06× for matrices larger than 24k×24k Yongxiao Zhou, Yi Zong, Yuyang Jin 0001, Wei Xue 0003 |
ICS | 2 |
| 2025 | Semi-StructMG: A Fast and Scalable Semi-Structured Algebraic MultigridabstractParallel multigrid methods are widely used as preconditioned for solving large sparse linear systems. Most multigrids rely on general sparse matrix formats, which prevent them from achieving optimal performance. There is an emerging trend towards semi-structured multigrids that balance flexibility with performance. However, existing libraries often fall short in terms of speed and scalability for semi-structured problems. To address these limitations, we have designed and implemented Semi-StructMG. It employs multi-dimensional coarsening to reduce complexity and simplify communication patterns. It also considers the special role of inter-block connections in smoothers and triple-matrix products to improve convergence under large-scale parallelism. We evaluated Semi-StructMG using two benchmark problems and four real-world applications from petroleum reservoir simulation, ship manufacturing, numerical weather prediction, and ocean modeling. Compared to hypre's multigrids, Semi-StructMG achieves the fastest time-to-solution across all cases, with average speedups of 5.97x, 15.2x, and 3.85x over SSAMG, Split, and BoomerAMG, respectively. Additionally, Semi-StructMG significantly improves both strong and weak scaling efficiencies in all tests. These results suggest that it can serve as an effective alternative to SSAMG and Split. Yi Zong, Longjiang Mu, Jianchun Wang, Peinan Yu, Wei Xue 0003 |
PPoPP | 1 |
| 2024 | FP16 Acceleration in Structured Multigrid Preconditioner for Real-World ApplicationsabstractHalf-precision hardware support is now almost ubiquitous. In contrast to its active use in AI, half-precision is less commonly employed in scientific and engineering computing. The valuable proposition of accelerating scientific computing applications using half-precision prompted this study. Focusing on solving sparse linear systems in scientific computing, we explore the technique of utilizing FP16 in multigrid preconditioners. Based on observations of sparse matrix formats, numerical features of scientific applications, and the performance characteristics of multigrid, this study formulates four guidelines for FP16 utilization in multigrid. The proposed algorithm demonstrates how to avoid FP16 overflow through scaling. A setup-then-scale strategy prevents FP16’s limited accuracy and narrow range from interfering with the multigrid’s numerical properties. Another strategy, recover-and-rescale on the fly, reduces the memory footprint of hotspot kernels. The extra precision-conversion overhead in mix-precision kernels is addressed by the transformation of storage formats and SIMD implementation. Two ablation experiments validate the effectiveness of our algorithm and parallel kernel implementation on ARM and X86 architectures. We further evaluate three idealized and five real-world problems to demonstrate the advantage of utilizing FP16 in a multigrid preconditioner. The average speedups are approximately 2.75x and 1.95x in preconditioner and end-to-end workflow, respectively. Yi Zong, Peinan Yu, Haopeng Huang, Wei Xue 0003 |
ICPP | 1 |
| 2024 | POSTER: StructMG: A Fast and Scalable Structured MultigridabstractParallel multigrid is widely used as preconditioners in solving large-scale sparse linear systems. However, the current multigrid library still needs more satisfactory performance for structured grid problems regarding speed and scalability. To this end, we design and implement StructMG, a fast and scalable multigrid that constructs hierarchical grids automatically based on the original matrix. As a preconditioner, StructMG can achieve both low cost per iteration and good convergence. Two idealized and five real-world problems from four application fields, including radiation hydrodynamics, petroleum reservoir simulation, numerical weather prediction, and solid mechanics, are evaluated on ARM and X86 platforms. In comparison to hypre's multigrid preconditioners, StructMG achieves the fastest time-to-solutions in all cases with average speedups of 17.6x, 5.7x, 4.6x, 8.5x over SMG, PFMG, SysPFMG, and BoomerAMG, respectively. Additionally, StructMG significantly improves strong and weak scaling efficiencies in most tests. Yi Zong, Haopeng Huang, Sicong Li 0005, Zhaohui Ding, Wei Xue 0003 |
PPoPP | 1 |