EDBT 2026 Demo / reviewers in the wild / expert
Zheng Wang 0079
dblp:181/2834-79
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2024
0009-0001-7858-6238ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Mentor: A Memory-Efficient Sparse-dense Matrix Multiplication Accelerator Based on Column-Wise ProductabstractSparse-dense matrix multiplication (SpMM) is the performance bottleneck of many high-performance and deep-learning applications, making it attractive to design specialized SpMM hardware accelerators. Unfortunately, existing hardware solutions do not take full advantage of data reuse opportunities of the input and output matrices or suffer from irregular memory access patterns. Their strategies increase the off-chip memory traffic and bandwidth pressure, leaving much room for improvement. We present Mentor , a new approach to designing SpMM accelerators. Our key insight is that column-wise dataflow, while rarely exploited in prior works, can address these issues in SpMM computations. Mentor is a software-hardware co-design approach for leveraging column-wise dataflow to improve data reuse and regular memory accesses of SpMM. On the software level, Mentor incorporates a novel streaming construction scheme to preprocess the input matrix for enabling a streaming access pattern. On the hardware level, it employs a fully pipelined design to unlock the potential of column-wise dataflow further. The design of Mentor is underpinned by a carefully designed analytical model to find the tradeoff between performance and hardware resources. We have implemented an FPGA prototype of Mentor . Experimental results show that Mentor achieves speedup by geomean 2.05× (up to 3.98×), reduces the memory traffic by geomean 2.92× (up to 4.93×), and improves bandwidth utilization by geomean 1.38× (up to 2.89×), compared with the state-of-the-art hardware solutions. Xiaobo Lu, Jianbin Fang, Lin Peng 0001, Chun Huang 0006, Zidong Du, Yongwei Zhao 0001, Zheng Wang 0079 |
ACM Trans. Archit. Code Optim. | 7 |
| 2024 | Optimizing Full-Spectrum Matrix Multiplications on ARMv8 Multi-Core CPUsabstractGeneral Matrix Multiplication (GEMM) is a key subroutine in high-performance computing. While the mainstream Basic Linear Algebra Subprograms (BLAS) libraries can deliver good performance on large and regular-shaped GEMMs, they are inadequate for optimizing small and irregular-shaped GEMMs, which are commonly seen in emerging HPC applications. Recent research has focused on improving GEMM performance on GPUs, but there is still significant room for improvement on emerging HPC hardware based on multi-core CPUs. We presentLibShalom2, an open-source library to optimize full-spectrum GEMMs, taking small, irregular-shaped, and large-scale regular-shaped matrices.LibShalom2explicitly targets the ARMv8 architecture, which is becoming common in HPC systems.LibShalom2is designed to minimize the expensive memory accessing overhead for data packing and processing small matrices. It uses analytic methods to determine GEMM kernel optimization parameters, enhancing the computation and parallelization efficiency of the GEMM kernels. We evaluateLibShalom2by applying it to three ARMv8 multi-core architectures and comparing it against five mainstream linear algebra libraries. Experimental results show thatLibShalom2consistently outperforms existing solutions across full-spectrum GEMM workloads and hardware architectures. We also show thatLibShalom2delivers an average speedup of 2.2x for real-life neural network workloads. Weiling Yang, Jianbin Fang, Dezun Dong, Xing Su 0004, Zheng Wang 0079 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2024 | Optimizing Multi-Grid Preconditioned Conjugate Gradient Method on Multi-CoresabstractMultigrid preconditioned conjugate gradient (MGPCG) is commonly used in high-performance computing (HPC) workloads. However, MGPCG is notoriously challenging to optimize since most of its computation kernels are memory-bounded with low arithmetic intensity and non-trivial communication patterns among parallel processes. This article presents new techniques to improve the data locality and reduce the communication overhead of MGPCG by first merging the kernels of multigrid (MG). We then develop an asynchronous neighboring communication algorithm to reduce the data communications across parallel processes. We demonstrated the benefits of our approach by applying it to the high-performance conjugate gradient (HPCG) benchmark and integrating it with a real-life algebraic multigrid package. We test the resulting software implementations on three ARMv8 and one Intel Xeon system. Experimental results show that our approach leads to a 1.62x-2.54x speedup over the engineer- and vendor-tuned HPCG implementations across various workloads and platforms. Xiaojian Yang, Shengguo Li, Dezun Dong, Chun Huang 0006, Zheng Wang 0079 |
IEEE Trans. Parallel Distributed Syst. | 6 |