Peng Zhang 0061

dblp:21/1048-61 · DBLP profile ↗
← Back
21ranked-venue papers
3as first author
17since 2021 · last 2026
0000-0001-8364-9793ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 19 · 3 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Optimizing small matrix multiplications via batch grouping on multi-core DSPs
Xiaotian Chen, Jianbin Fang, Peng Zhang 0061, Chun Huang 0006
CCF Trans. High Perform. Comput.4
2026 High-performance matrix multiplication micro-kernel generation for GPDSPs via MLIR progressive lowering
Kainan Yu, Peng Zhang 0061, Chun Huang 0006, Jianbin Fang
CCF Trans. High Perform. Comput.3
2026 Optimizing Attention for Large Language Model Inference on the MT-3000 Many-Core Processor
abstract
Transformer-based large language models (LLM) are increasingly deployed in high-performance computing environments, where the attention mechanism often becomes a key bottleneck during inference. Although state-of-the-art attention algorithms (e.g., FlashAttention) achieve high efficiency on GPUs, they are ill-suited to emerging heterogeneous many-core processors. In this work, we focus on MT-3000, a representative architecture deployed in the new-generation Tianhe supercomputer, and identify three principal challenges in realizing high-performance attention: complex multi-tier memory requiring manual data movement, excessive reduction overhead caused by sub-tile softmax operations, and static execution pipelines that fail to adapt to inference phases and sequence lengths. To overcome these challenges, we propose DeferAttention , a high-performance attention implementation designed for the MT-3000 many-core processor. DeferAttention introduces a novel deferred-reduction attention strategy to decouple reduction from the fused compute pipeline, enabling more efficient aggregation over large tiles. Moreover, DeferAttention adopts a memory-centric operator design, including data tiling, multi-level software pipelining, and modular micro-kernels, to maximize data reuse and execution throughput. Finally, to support runtime-adaptive execution, DeferAttention integrates a lightweight kernel selection strategy guided by an analytical cost model. Experimental results show that DeferAttention achieves up to 98% of the theoretical peak at the micro-kernel level and 85% at the operator level, outperforming baseline implementations and significantly accelerating end-to-end inference.
Xinxin Qi, Jianbin Fang, Peng Zhang 0061, Yonggang Che
ACM Trans. Archit. Code Optim.3
2026 mtGEMM: An Efficient GEMM Library for Modern Multi-Core DSPs
abstract
The General Matrix Multiplication (GEMM) is a crucial subprogram in high-performance computing (HPC). With the increasing importance of power and energy consumption, modern Digital Signal Processors (DSPs) are being integrated into general-purpose HPC systems. However, due to architecture disparities, traditional optimizations for CPUs and GPUs are not easily applicable to modern DSPs. This paper shares our experience of optimizing the GEMM operation using a CPU-DSP platform as a case study. Our work employs a set of strategies to improve the performance and scalability of GEMM. These strategies focus on developing micro-kernels based on heterogeneous on-chip memory, addressing the memory access bottleneck in multi-core parallelism, and facilitating efficient transpose-GEMM. These approaches, collectively referred to as an efficient and practical library (a.k.a.mtGEMM), maximize computational capabilities and bandwidth utilization of multi-core DSPs, while achieving high performance for variously-shaped GEMMs. Our experimental results demonstrate thatmtGEMMcan attain between 92% and 96% of the hardware peak, with the multi-core scalability being almost linear.
Jianbin Fang, Kainan Yu, Peng Zhang 0061, Dezun Dong, Xinxin Qi, Xingyu Hou, Ruibo Wang, Kai Lu 0001
IEEE Trans. Parallel Distributed Syst.3
2025 Selection of Supervised Learning-Based Sparse Matrix Reordering Algorithms
Tao Tang 0001, Youfu Jiang, Yingbo Cui 0001, Jianbin Fang, Peng Zhang 0061, Lin Peng 0001, Chun Huang 0006
HiPC5
2025 Optimizing Direct Convolutions on High-Performance Multi-Core DSPs
abstract
Convolution operations form the computational backbone of deep learning inference but often become performance bottlenecks on conventional architectures. While multi-core Digital Signal Processors (DSPs) offer energy-efficient alternatives through long vector units and software-managed memory hierarchies, existing convolution optimizations designed for CPUs/GPUs underperform due to architectural mismatches in memory systems and execution pipelines. We present mtConv, an optimizing convolution method for multi-core DSPs. mtConv achieves high performance by exploiting data reuse, managing on-chip memory, and designing efficient micro-kernels. It maximizes the overlap between computation and communication to hide data transfer latency, leveraging DSPs’ long vector units and hierarchical scratchpad memories. We evaluate mtConv against state-of-the-art convolution optimizations on DSPs. Experimental results show that mtConv delivers the best overall performance across various convolution layers, achieving up to 93.25% of the hardware’s peak performance on a single DSP core and 92.31% when using all 8 cores of a single DSP cluster.
Xiaotian Chen, Jianbin Fang, Peng Zhang 0061, Yonggang Che, Chun Huang 0006, Jie Ren 0007
ICPP4
2025 Constraint-Driven Auto-Tuning of GEMM-like Operators for MT-3000 Many-core Processor
abstract
Optimizing deep learning (DL) operators, particularly GEMM-like operations, for emerging heterogeneous many-core processors like MT-3000 is challenging due to the large search space and hardware-specific constraints. Existing approaches - such as hand-crafted libraries or general-purpose auto-tuners - are either expensive to develop or deliver sub-optimal performance due to expensive search overheads. We present DynaChain, an operator-level optimization framework for MT-3000. DynaChain decouples the computation and data movement of operators, allowing each to be optimized independently and maximizing global data reuse across the operator schedule. To reduce the search space, DynaChain introduces constraint dependency chains that dynamically eliminate invalid scheduling options during exploration. It then applies an integer linear programming (ILP) based decomposition to handle irregular matrix dimensions, avoiding padding and improving hardware utilization. For low-level code generation, DynaChain offers a hardware-aware micro-kernel design optimized for the MT-3000’s VLIW+SIMD architecture, supporting irregular operations through improved register allocation and instruction pipelining. Experimental results on a range of representative DL operators demonstrate that DynaChain simplifies kernel development for heterogeneous many-core architectures while delivering performance on par with expert-optimized libraries.
Xinxin Qi, Jianbin Fang, Peng Zhang 0061, Yonggang Che, Jie Ren 0007
SC3
2025 An empirical performance evaluation of SYCL on ARM multi-core processors
Hanzheng Liang, Chencheng Deng, Peng Zhang 0061, Jianbin Fang, Tao Tang 0001, Chun Huang 0006
CCF Trans. High Perform. Comput.3
2025 nDirect2: A High-Performance Library for Direct Convolutions on Multicore CPUs
abstract
Convolution kernels are widely seen in high-performance computing (HPC) and deep learning (DL) workloads and are often responsible for performance bottlenecks. Prior works have demonstrated that the direct convolution approach can outperform the conventional convolution implementation. Although well-studied, the existing approaches for direct convolution are either incompatible with the mainstream DL data layouts or lead to suboptimal performance. We designnDirect2, a novel direct convolution approach that targets multi-core CPUs commonly found in smartphones and HPC systems.nDirect2is compatible with the data layout formats used by mainstream DL frameworks and offers new optimizations for the computational kernel, data packing, advanced operator fusion, and parallelization. We evaluatenDirect2by applying it to representative convolution kernels and demonstrating how well it performs on four distinct ARM-based CPUs and an X86-based CPU. Experimental results show thatnDirect2outperforms four state-of-the-art convolution approaches across most evaluation cases and hardware architectures.
Weiling Yang, Jianbin Fang, Dezun Dong, Zhengbin Pang, Runxi He, Peng Zhang 0061, Tao Tang 0001, Chun Huang 0006, Yonggang Che, Jie Ren 0007
IEEE Trans. Computers7
2024 Optimizing SpMV on Heterogeneous Multi-Core DSPs through Improved Locality and Vectorization
abstract
The sparse matrix-vector multiplication (SpMV) is widely used in large-scale scientific computing and engineering. However, optimizing SpMV for high-performance digital signal processors (DSPs) has received limited attention. We present HaLAV, a method to accelerate SpMV on CPU-DSP heterogeneous platforms, using the FT-M7032 DSP platform as a case study. HaLAV partitions the input matrix into ‘dense’ and ‘sparse’ parts through column reordering. For the dense part, HaLAV automatically selects storage formats optimized for vectorization to run on the DSP. At the same time, it offloads the sparse component to be processed by the CPU using the standard CSR algorithm. We evaluate our approach on the FT-M7032 platform and an Intel Xeon CPU. Experimental results show that our techniques achieve average speedups of 2.09 × and 1.66 × over the competing baselines on the FT-M7032 and the Xeon platform, respectively.
Deshun Bi, Shengguo Li, Dezun Dong, Peng Zhang 0061, Jianbin Fang
ICPP4
2024 Optimizing Stencil Computation on Multi-core DSPs
abstract
Stencil is a common computation pattern in high-performance computing (HPC) applications. While extensive work has been proposed to optimize stencil kernels on CPUs and GPUs, there is no consensus on how to best optimize stencils on multi-core Digital Signal Processors (DSPs) used in emerging HPC systems. This paper shares our experience in optimizing stencil kernels on multi-core DSPs. Our approach combines coarse and fine-grained parallel optimization techniques to enhance the performance of stencil computations. Our optimizations include a vectorization-enabled micro-kernel to utilize instruction parallelism, a memory-aware data reuse strategy to maximize data locality across multiple memory levels and a triple-buffering mechanism to overlap computation and memory communications. Experimental results show that our approach can effectively utilize the memory bandwidth and the computation capability of the underlying hardware. Our integrated optimizations can yield a 3.72x speedup over the 16-core CPU counterpart.
Fugeng Zhu, Xinxin Qi, Peng Zhang 0061, Jianbin Fang, Tao Tang 0001, Yonggang Che, Kainan Yu, Jing Xie 0023, Chun Huang 0006, Jie Ren 0007
ICPP3
2024 Optimizing General Matrix Multiplications on Modern Multi-core DSPs
abstract
General Matrix Multiplication (GEMM) is a key subprogram in high-performance computing (HPC) and deep learning workloads. With the rising significance of power and energy consumption in HPC systems, accelerators based on Digital Signal Processors (DSPs) have been integrated into general-purpose HPC systems. Due to the architecture disparities, the GEMM optimization techniques used on conventional multi-core CPUs and GPGPUs are not always applicable to DSPs. This paper shares our experience in optimizing GEMM on multi-core GPDSPs, using a CPU-DSP processor as a case study. Our approach employs a range of techniques to optimize performance for DSP architectures. These include data partitioning, three-level pipelining, dedicated micro-kernel design, and improved vector reduction. These optimizations maximize the overlap between computation and communication while fully exploiting the capabilities of floating-point arithmetic units to achieve high performance. Our experimental results demonstrate that the performance attained by our optimization is up to 96% of the theoretical peak performance of the hardware.
Kainan Yu, Xinxin Qi, Peng Zhang 0061, Jianbin Fang, Dezun Dong, Ruibo Wang, Tao Tang 0001, Chun Huang 0006, Yonggang Che, Zheng Wang 0001
IPDPS3
2024 thSORT: an efficient parallel sorting algorithm on multi-core DSPs
Mouzhi Yang, Peng Zhang 0061, Jianbin Fang, Weifeng Liu 0002, Chun Huang 0006
CCF Trans. High Perform. Comput.2
2023 MTMap: A Long-Read Alignment Tool based on Multi-Core DSPs
abstract
Read alignment is a basic and important task in genomic data analysis. The popularity of the third—generation sequencing technology has brought the need of sequence alignment algorithms to analyze long-read sequences with longer read length and high error rate. Moreover, the rapid growth of sequence data has also presented challenges for read alignment. To improve the ability to process large volume of sequencing reads, we developed a long-read sequence alignment algorithm MTMap on the heterogeneous processor FT-m7032. MTMap utilizes multi-level parallel technologies: firstly, we tailored the data structure for the wide vector processing units of DSP to speedup the score matrix computation. Secondly, we developed multithread parallelization for base-level alignment on each DSP cluster. Finally, we implemented multi-process parallelization between DSP clusters to fully exploit the computing power of FT-m7032. Experiments show that, MTMap achieves up to 16 times of parallel acceleration performance compared with the original algorithm under the condition of ensuring accuracy.
Xinjie An, Shijie Li 0002, Yingbo Cui 0001, Peng Zhang 0061, Biao Long
BIBM5
2023 Optimizing Direct Convolutions on ARM Multi-Cores
abstract
Convolution kernels are widely seen in deep learning workloads and are often responsible for performance bottlenecks. Recent research has demonstrated that a direct convolution approach can outperform the traditional convolution implementation based on tensor-to-matrix conversions. However, existing approaches for direct convolution still have room for performance improvement. We present nDirect, a new direct convolution approach that targets ARM-based multi-core CPUs commonly found in smartphones and HPC systems. nDirect is designed to be compatible with the data layout formats used by mainstream deep learning frameworks but offers new optimizations for the computational kernel, data packing, and parallelization. We evaluate nDirect by applying it to representative convolution kernels and demonstrating its performance on four distinct ARM multi-core CPU platforms. We compare nDirect against state-of-the-art convolution optimization techniques. Experimental results show that nDirect gives the best overall performance across evaluation scenarios and platforms.
Weiling Yang, Jianbin Fang, Dezun Dong, Chun Huang 0006, Peng Zhang 0061, Tao Tang 0001, Zheng Wang 0001
SC6
2023 Programming bare-metal accelerators with heterogeneous threading models: a case study of Matrix-3000
abstract
As the hardware industry moves toward using specialized heterogeneous many-core processors to avoid the effects of the power wall, software developers are finding it hard to deal with the complexity of these systems. In this paper, we share our experience of developing a programming model and its supporting compiler and libraries for Matrix-3000, which is designed for next-generation exascale supercomputers but has a complex memory hierarchy and processor organization. To assist its software development, we have developed a software stack from scratch that includes a low-level programming interface and a high-level OpenCL compiler. Our low-level programming model offers native programming support for using the bare-metal accelerators of Matrix-3000, while the high-level model allows programmers to use the OpenCL programming standard. We detail our design choices and highlight the lessons learned from developing system software to enable the programming of bare-metal accelerators. Our programming models have been deployed in the production environment of an exascale prototype system.
Jianbin Fang, Peng Zhang 0061, Chun Huang 0006, Tao Tang 0001, Kai Lu 0001, Ruibo Wang, Zheng Wang 0001
Frontiers Inf. Technol. Electron. Eng.2
2021 Large-Scale Parallel Alignment Algorithm for SMRT Reads
Yingbo Cui 0001, Peng Zhang 0061, Tao Tang 0001, Lin Peng 0001, Chun Huang 0006, Canqun Yang, Xiangke Liao
ICA3PP (2)4
2020 Optimizing Streaming Parallelism on Heterogeneous Many-Core Architectures
abstract
As many-core accelerators keep integrating more processing units, it becomes increasingly more difficult for a parallel application to make effective use of all available resources. An effective way of improving hardware utilization is to exploit spatial and temporal sharing of the heterogeneous processing units by multiplexing computation and communication tasks - a strategy known as heterogeneous streaming. Achieving effective heterogeneous streaming requires carefully partitioning hardware among tasks, and matching the granularity of task parallelism to the resource partition. However, finding the right resource partitioning and task granularity is extremely challenging, because there is a large number of possible solutions and the optimal solution varies across programs and datasets. This article presents an automatic approach to quickly derive a good solution for hardware resource partition and task granularity for task-based parallel applications on heterogeneous many-core architectures. Our approach employs a performance model to estimate the resulting performance of the target application under a given resource partition and task granularity configuration. The model is used as a utility to quickly search for a good configuration at runtime. Instead of hand-crafting an analytical model that requires expert insights into low-level hardware details, we employ machine learning techniques to automatically learn it. We achieve this by first learning a predictive model offline using training programs. The learned model can then be used to predict the performance of any unseen program at runtime. We apply our approach to 39 representative parallel applications and evaluate it on two representative heterogeneous many-core platforms: a CPU-XeonPhi platform and a CPU-GPU platform. Compared to the single-stream version, our approach achieves, on average, a 1.6x and 1.1x speedup on the XeonPhi and the GPU platform, respectively. These results translate to over 93 percent of the performance delivered by a theoretically perfect predictor.
Peng Zhang 0061, Jianbin Fang, Canqun Yang, Chun Huang 0006, Tao Tang 0001, Zheng Wang 0001
IEEE Trans. Parallel Distributed Syst.1
2019 The Communication-Overlapped Hybrid Decomposition Parallel Algorithm for Multi-Scale Fluid Simulations
abstract
The MCDPar (Parallel algorithm for multi-scale simulations based on Mesh and BCF Decomposition) algorithm significantly reduced the execution time and improved the parallel scalability for the multi-scale fluid simulations. However, the performance bottleneck still exists for extremely large-scale parallel simulations. In this paper, we designed a communication-overlapped hybrid decomposition parallel algorithm to improve the performance of the original MCDPar on large-scale clusters. Through non-blocking communication and code scheduling, the communication overhead between the master and slave groups have been overlapped with the computation of more microscopic configuration fields for the master process. Thus the parallel efficiency and scalability of the multi-scale solver could be improved on large-scale parallel simulations. In the test case with the number of configuration fields NBCF = 1000 and mesh cells Ncell = 64000, the communication percentage between the corresponding master and slave processes is reduced by 39.71%. In the test case with NBCF = 3000 and Ncell = 64000, the time cost of the fastest execution is reduced by 31.13% using the communication-overlapped algorithm, which offers a better parallel scaling on 256 cores compared to original 128 cores.
Yi Liu 0083, Chao Li 0070, Canqun Yang, Xinbiao Gan, Peng Zhang 0061, Sijiang Fan
ICPP6
2018 MOCL: an efficient openCL implementation for the matrix-2000 architecture
abstract
This paper presents the design and implementation of an Open Computing Language (OpenCL) framework for the Matrix-2000 many-core architecture. This architecture is designed to replace the Intel XeonPhi accelerators of the TianHe-2 supercomputer. We share our experience and insights on how to design an effective OpenCL system for this new hardware accelerator. We propose a set of new analysis and optimizations to unlock the potential of the hardware. We extensively evaluate our approach using a wide range of OpenCL benchmarks on a single and multiple computing nodes. We present our design choices and provide guidance how to optimize code on the new Matrix-2000 architecture.
Peng Zhang 0061, Tao Tang 0001, Jianbin Fang, Chun Huang 0006, Canqun Yang, Zheng Wang 0001
CF1
2018 Auto-tuning Streamed Applications on Intel Xeon Phi
abstract
Many-core accelerators, as represented by the XeonPhi coprocessors and GPGPUs, allow software to exploit spatial and temporal sharing of computing resources to improve the overall system performance. To unlock this performance potential requires software to effectively partition the hardware resource to maximize the overlap between host-device communication and accelerator computation, and to match the granularity of task parallelism to the resource partition. However, determining the right resource partition and task parallelism on a per program, per dataset basis is challenging. This is because the number of possible solutions is huge, and the benefit of choosing the right solution may be large, but mistakes can seriously hurt the performance. In this paper, we present an automatic approach to determine the hardware resource partition and the task granularity for any given streamed application, targeting the Intel XeonPhi architecture. Instead of hand-crafting the heuristic for which the process will have to repeat for each hardware generation, we employ machine learning techniques to automatically learn it. We achieve this by first learning a predictive model offline using training programs; we then use the learned model to predict the resource partition and task granularity for any unseen programs at runtime. We apply our approach to 23 representative parallel applications and evaluate it on a CPU-XeonPhi mixed heterogenous many-core platform. Our approach achieves, on average, a 1.6x (upto 5.6x) speedup, which translates to 94.5% of the performance delivered by a theoretically perfect predictor.
Peng Zhang 0061, Jianbin Fang, Tao Tang 0001, Canqun Yang, Zheng Wang 0001
IPDPS1