Pengyu Liu 0004

dblp:73/7783-4 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2026
0009-0002-1366-5055ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 3 first-author · 6 since 2021
YearPublicationVenuePosition
2026 Leveraging Tensor Dataflow for Improved Thermal Performance on 3D-Stacked SRAM Architecture
abstract
While 3D-stacked SRAM architectures have demonstrated prominent performance speedup for tensor computing by exploiting higher bandwidth, larger buffer and reduced latency, they suffer from thermal challenges owing to vertical stacking nature of chips. In this paper, we identify that tensor dataflow may further exacerbate the thermal issues, so we propose T3D, the first thermal-aware tensor framework for 3D-stacked SRAM architectures, leveraging tensor dataflow characteristics to significantly enhance thermal performance. Specifically, we first perform a quantitative formulation to identify the most energy-efficient tensor dataflow with given 3D constraints, effectively reducing heat generation without performance loss. Then, we develop a thermal-aware 3D architectural floorplan to improve heat spreading by optimizing the spatial arrangement of multiple SRAM macros with varying power overheads, which is caused by mismatched data access rates of tensor computing. Experimental results show that, our proposed T3D can reduce the peak chip temperature by 12.9°C on certain LLM and DNN workloads over the state-of-the-art 3D solutions.
Pengyu Liu 0004, Zelong Yuan, Yingkun Liu, Jianfei Jiang 0001, Qin Wang 0009, Zhigang Mao, Naifeng Jing
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2025 Bridge-NDP: Efficient Communication-Computation Overlap in Near Data Processing System
abstract
Near data processing (NDP), enabled by near data accelerators (NDAs) within DIMM-based main memory, enhances performance by providing more aggregated bandwidth and reducing long-distance data transfers. While the performance of NDAs has received widespread attention, the overhead of host-NDA communication has been overlooked, becoming a bottleneck in NDP systems. To alleviate performance degradation from communication, we propose Bridge-NDP, the first NDP architecture that implements a workflow with efficient communication-computation overlap. Bridge-NDP is built upon the conventional NDP architecture and can be easily applied to existing NDP designs, regardless of the memory level where NDAs are attached. Specifically, we introduce a novel direct host-NDA communication method that utilizes existing memory buses as bridge buses, avoiding the need for new interconnections. It enables seamless integration with other memory accesses while achieving high bandwidth utilization with minimal hardware overhead. For the system-level workflow design, we optimize and extend existing dataflow to achieve richer computing paradigms with fewer redundant memory accesses. Additionally, we provide programming support with efficient API designs and data management to hide low-level resource details and ensure correctness guarantees. Comprehensive experiments demonstrate that Bridge-NDP achieves significant performance improvements, with speedups of$1.8\times $–$3.1\times $and bandwidth utilization improvement of$2.0\times $–$2.9\times $over the state-of-the-art NDP solutions.
Pengyu Liu 0004, Dongxu Lyu, Jianfei Jiang 0001, Qin Wang 0009, Zhigang Mao, Naifeng Jing
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2024 A Flexible and High-Precision Activation Function Unit Based on Equi-Error Partitioning Algorithm
abstract
The diversity of activation functions has gradually increased to accommodate different tasks in modern deep neural networks (DNNs). However, these novel activation functions involve more nonlinear operations relative to traditional activation functions, which increases the computational complexity. To address these issues, a piecewise linear (PWL) approximation algorithm called Equi-Error Partitioning Algorithm is proposed in this paper. The algorithm aims at balancing the errors between segments and solves the problem of excessive precision that exists in other PWL approximation methods and achieves on average 30.07× better mean squared error compared to the previous works. Based on this algorithm, we propose an activation function unit (AFU) which enables the addressing scheme of non-uniform segments and provides reconfigurability for all common activation functions by reloading parameters. End-to-end evaluation with several DNNs shows the accuracy loss is all less than 0.06% with 64 segments.
Zelong Yuan, Siwei Yuan, Pengyu Liu 0004, Weiguang Sheng, Naifeng Jing
ISCAS3
2024 A Comprehensive Dataflow-Mapping Optimization for Fully Pipelined Execution in Spatial Programmable Architecture
abstract
Although spatial programmable architectures have demonstrated high-performance and programmability for a variety of applications, they suffer from the pipeline unbalancing issue which restricts resource utilization and degrades the performance. In this paper, we identify that spatial initiation interval (SpII) can quantitatively describe the impact of pipeline unbalancing on performance, so we formulate SpII for the first time in spatial architectures. To achieve an optimal SpII, we propose dataflow decomposing and integrated mapping to enable high performance dataflow-mapping on spatial architectures. Dataflow decomposing decomposes the application graph into subgraphs and runs them serially, so that it adapts the regular spatial architecture to various application dataflows, particularly for extremely unbalanced datapaths without incurring large buffering overhead. Based on the quantitative SpII, we propose integrated mapping to consider operator placing, operand routing and pipeline balancing at the same time that can find a better SpII for fully-pipelined execution on spatial architectures. The experiment results show that our proposal can gain an average of 2.1× performance speedup on a variety of application kernels over the state-of-the-art approaches.
Pengyu Liu 0004, Ang Li 0045, Jianfei Jiang 0001, Qin Wang 0009, Zhigang Mao, Naifeng Jing
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2023 Pipeline Balancing for Integrated Mapping in High Performance Spatial Programmable Architecture
abstract
Recently, spatial programmable architectures have gained increasing popularity owing to their performance and programmability, while the achievable performance is highly related to how the operators are mapped onto a number of processing elements (PEs) in the spatial architectures. In this paper, we first identify that in the spatial mapping process, the pipeline balancing problem is essential by affecting the spatial initial interval (SpII). Hence, we formulate the SpII for the first time in spatial architecture. The quantitative formulation enables an integrated mapping algorithm which combines operator placement, operand routing and pipeline balancing at the same time. In addition, to reduce the balancing hardware cost, we propose a bridge-buffer structure to facilitate operand routing and buffering on demand. To reduce the mapping searching space, we propose three optimization techniques to trade off solution quality, mapping time and hardware overhead. The experiment results show that the proposed integrated mapping algorithm can reduce the SpII by 42.3%, which in turn improves the throughput and algorithm runtime up to 1.74× and 3.08× over the state-of-the-art heuristic spatial mapping.
Pengyu Liu 0004, Jianfei Jiang 0001, Qin Wang 0009, Zhigang Mao, Naifeng Jing
FPL1
2023 Exploiting bit sparsity in both activation and weight in neural networks accelerators
Naifeng Jing, Yongshuai Sun, Pengyu Liu 0004, Qin Wang 0009, Jianfei Jiang 0001
Integr.4