EDBT 2026 Demo / reviewers in the wild / expert
Zenghui Ren
dblp:352/1973
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2025
0009-0007-7873-3078ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Towards Efficient SpMV: A Multi-Aware Optimization Framework for Heterogeneous Architecture
Xinyin Zhang, Zenghui Ren, Yonghua Zhao |
ICA3PP (3) | 4 |
| 2024 | swDarknet: A Heterogeneous Parallel Deep Learning Framework Suitable for SW26010 Pro Processor
Huazeng Liu, Meihong Yang, Zenghui Ren |
NPC (1) | 5 |
| 2023 | hsSpMV: A Heterogeneous and SPM-aggregated SpMV for SW26010-Pro many-core processorabstractSparse matrix vector multiplication (SpMV) is a critical performance bottleneck for numerical simulation and artificial intelligence training. The new generation of Sunway supercomputer is the advanced exascale supercomputer in China. The SW26010-Pro many-core processor renders itself as a competitive candidate for its attractive computational power in both numerical simulation and artificial intelligence training. In this paper, we propose a heterogeneous and SPM-aggregated SpMV kernel, specifically designed for the SW26010-Pro many-core processor. To fully exploit the computational power of the SW26010-Pro and balance the load of each core group(CG) during computation, we employ asynchronous computation workflow and propose the SPM-aggregated strategy and vector adaptive mapping algorithm. In addition, we propose the two-level data partition scheme to implement computational load balance. In order to improve memory access efficiency, we directly access memory via DMA controller to replace the discrete memory access. Using several optimizations, we achieve a 77.16x speedup compared to the original implementation. Our experimental results show that the hsSpMV yields up to 3.82× speedups on average compared to the SpMV kernel of the state-of-the-art Sunway math library xMath2.0. Jingshan Pan, Chaochao Yang, Renjiang Chen, Zenghui Ren, Anjun Liu |
CCGrid | 7 |
| 2023 | SW-TRRM: Parallel Optimization Research of the Random Ray Method Based on Sunway Bluelight II Supercomputer
Zenghui Ren, Tao Liu 0029, Zhaoyuan Liu, Ying Guo 0028, Jingshan Pan, Meihong Yang |
ICA3PP (5) | 1 |
| 2023 | SW-LeNet: Implementation and Optimization of LeNet-1 Algorithm on Sunway Bluelight II Supercomputer
Zenghui Ren, Tao Liu 0029, Zhaoyuan Liu, Min Tian 0005, Ying Guo 0028, Jingshan Pan |
ICA3PP (5) | 1 |
| 2023 | AH-TDMA: An Adaptive Heterogeneous Tridiagonal Matrix Algorithm on the New Sunway SupercomputerabstractThe solution of tridiagonal linear systems is used in in various fields and plays a crucial role in numerical simulations. However, there is few efficient solver for tridiagonal linear systems on the new Sunway supercomputer. Based on a three-dimensional heat conduction problem, we propose an adaptive heterogeneous tridiagonal matrix algorithm (AH-TDMA). Our major innovations include: (1) To address computational hotspots within AH-TDMA, a multi-level parallel approach involving MPI+Athread has been adopted. (2) An adaptive data partitioning scheme has been set up to achieve load balance. (3) Employing direct memory access and establishing shared space between registers and main memory, instead of employing discrete memory access, is done to enhance memory access efficiency. (4) The optimization of the loop structure has been made in adjusting the sequencing of the dual-layered loops to reduce communication overhead. The experimental results show that, with one core group, the AH-TDMA achieves a speedup of 99.6 times for hotspot and total time speedup up to 59.4 times compared to the Parallel and Scalable Library for Tridiagonal Matrix Algorithm (PaScal TDMA). The AH-TDMA is scalable up to 2048 core groups, with a parallel efficiency of 69.2%. Min Tian 0005, Qi Liu 0065, Zenghui Ren, Yue Liu 0027 |
ICPADS | 5 |