Fengzhe Zhang

dblp:76/4754 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0002-7584-1045ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FHPSAC: FPGA-based High-Parallelism SAC Accelerator
abstract
Reinforcement learning (RL) enables autonomous decision-making in applications such as robotics and control, and Soft Actor-Critic (SAC) is a leading model-free algorithm for continuous tasks. However, SAC’s small-batch training updates and fine-grained computation lead to heavy scheduling and kernel-launch overheads on GPU, limiting efficiency. In this work, we present FHPSAC (FPGA-based High-Parallelism SAC Accelerator), the first FPGA-accelerated architecture dedicated to SAC training. First, we propose a hardware–software co-designed on-chip memory hierarchy to statically partition and allocate SAC’s training data for conflict-free parallel access. Second, we build a high-parallelism accelerator with a tensor core for GEMM (General Matrix Multiply) and a lightweight unit for irregular elementwise/reduction kernels. Finally, we implement the full system on a Xilinx XCVU9P FPGA and demonstrate significant speedup with low power. Experimental results show that FHPSAC obtains 5.33–14.85× speedup compared with the Intel Xeon Gold 6130 CPU, while outperforming an NVIDIA A100-SXM4 GPU by 2.90–10.43× in training latency with an average power of 40.17 W. FHPSAC substantially reduces SAC training latency, providing a computational foundation for large-scale SAC deployments.
Jiabin Xu, Wang Fan, Xuegong Zhou, Wei Cao 0002, Fengzhe Zhang, Fan Zhang 0044, Xinsheng Yu 0001
FCCM6
2025 FbsPipe: Forward-Backward Separation Pipeline Parallelism Method for Deep Learning
Fengzhe Zhang, Yanzhao Gao, Xiaofeng Qi, Shuaikang Hou
ICA3PP (8)3
2025 OLMP: Operator-Level Computational Graph Partition Mapping for Deep Learning
abstract
Expanding the scale of deep neural networks (DNNs) is a fundamental approach to improving the performance and accuracy of the model. However, as DNN models continue to grow in complexity, a single computing accelerator is often inadequate to accommodate the entire model, leading to computational and memory bottlenecks. Distributing the computation graph of a DNN across multiple accelerators offers an effective solution to this challenge, enabling better resource utilization and improved performance scalability. Existing methods for coarse-grained partitioning and mapping of computational graph, often resulting in suboptimal solutions. This separation can lead to imbalanced workloads across accelerators, reducing pipeline efficiency and overall system performance. In this work, we propose OLMP to address these limitations, which jointly optimizes the partitioning and mapping processes. We formulate this problem as a combinatorial optimization task and employ Mixed Integer Programming to find an optimal solution that balances execution time, workload distribution, and memory constraints. To validate our approach, we conducted extensive experiments on a diverse set of DNN models with varying scales and architectures. Compared to existing approaches, OLMP achieves up to 2.64× speedup in training time. Experimental results demonstrate that OLMP outperforms existing methods, significantly reducing training time and enhancing load balancing across accelerators.
Fengzhe Zhang, Yijing Song, Xiaofeng Qi, Yanzhao Gao
SMC3
2025 DVHetero: A Framework for Designing and Validating Heterogeneous SoC with RISC-V Processor and CGRA
abstract
CGRA, as a coprocessor in SoCs, has been widely studied. However, there is limited research on how to efficiently debug and verify SoCs composed of CGRAs and processors during the design process. To address this gap, we introduce DVHetero. DVHetero incorporates a simulation and validation framework, SoCDiff, which enables comprehensive SoC simulation, debugging, and rapid error localization. Using this verification framework, we successfully implemented and validated the entire SoC. The SoC includes a Chisel-based CGRA generator and provides a pipelined CGRA architecture template. The CGRA is tightly integrated with the RISC-V processor, allowing for efficient DMA-based data transfer and MMIO support within the SoC. The pipelined CGRA architecture generated by DVHetero shows a 1.27× improvement in area efficiency and a 10.54× increase in mapping speed compared to the state-of-the-art CGRA framework, HierCGRA. Additionally, compared to state-of-the-art CGRA-SoC systems FDRA, DVHetero demonstrates a 1.67× increase in execution speed and a 4.34× improvement in area efficiency.
Guowei Zhu, Liming Deng, Kaisen Zhang, Wang Fan, Boyin Jin, Wei Cao 0002, Fengzhe Zhang, Xuegong Zhou, Fan Zhang 0044, Xinsheng Yu 0001
ACM Trans. Reconfigurable Technol. Syst.7
2023 Boosting Performance and QoS for Concurrent GPU B+trees by Combining-Based Synchronization
abstract
Concurrent B+trees have been widely used in many systems. With the scale of data requests increasing exponentially, the systems are facing tremendous performance pressure. GPU has shown its potential to accelerate concurrent B+trees performance. When many concurrent requests are processed, the conflicts should be detected and resolved. Prior methods guarantee the correctness of concurrent GPU B+trees through lock-based or software transactional memory (STM)-based approaches. However, these methods complicate the request processing logic, increase the number of memory accesses and bring execution path divergence. They lead to performance degradation and variance in response time increasing. Moreover, previous methods do not guarantee linearizability among concurrent requests.
Chuanlei Zhao, Lu Peng 0001, Yuzhe Lin, Fengzhe Zhang, Yunping Lu
PPoPP5
2022 High performance GPU concurrent B+tree
abstract
Concurrent B+trees have been widely used in many systems from file systems to databases. With the volume of data requests expanding exponentially, the systems are facing tremendous performance pressure. GPUs have shown their potential to accelerate the concurrent B+trees operations with their high volume of parallel computing resources and large memory bandwidth. In concurrent B+tree, the conflicts should be detected and resolved when multiple concurrent requests are traversing and operating on the tree. However, conflict detection and handling in concurrent B+tree complicates the request processing logic, increases the number of memory accesses and leads to execution path divergence. That leads to performance degradation and increased response time variance.
Chuanlei Zhao, Lu Peng 0001, Yuzhe Lin, Fengzhe Zhang, Jinhu Jiang
PPoPP5
2012 Mercury: Combining Performance with Dependability Using Self-Virtualization
Haibo Chen 0001, Fengzhe Zhang, Rong Chen 0001, Binyu Zang, Pen-Chung Yew
J. Comput. Sci. Technol.2
2011 CloudVisor: retrofitting protection of virtual machines in multi-tenant cloud with nested virtualization
abstract
Multi-tenant cloud, which usually leases resources in the form of virtual machines, has been commercially available for years. Unfortunately, with the adoption of commodity virtualized infrastructures, software stacks in typical multi-tenant clouds are non-trivially large and complex, and thus are prone to compromise or abuse from adversaries including the cloud operators, which may lead to leakage of security-sensitive data.
Fengzhe Zhang, Haibo Chen 0001, Binyu Zang
SOSP1
2007 Mercury: Combining Performance with Dependability Using Self-virtualization
abstract
There has recently been increasing interests in using system virtualization to improve the dependability of HPC cluster systems. However, it is not cost-free and may come with some performance degradation, uncertain QoS and loss of functionalities. Meanwhile, many virtualization-enabled features such as online maintenance and fault tolerance do not require virtualization being always on. This paper proposes a technique, called self-virtualization, that supports dynamically attaching and detaching a full-fledged virtual machine monitor (VMM) beneath an operating system, without disturbing applications thereon, and rid the system of potential overhead when the virtualization is not needed. This technique enables HPC clusters to reap most benefits from virtualization without sacrificing performance. This paper presents the design and implementation of Mercury, a working prototype based on Linux and Xen VMM. Our performance measurement shows that Mercury incurs very little overhead: about 0.2 ms to complete a mode switch, and negligible performance degradation compared to Linux.
Haibo Chen 0001, Rong Chen 0001, Fengzhe Zhang, Binyu Zang, Pen-Chung Yew
ICPP3
2006 Live updating operating systems using virtualization
abstract
Many critical IT infrastructures require non-disruptive operations. However, the operating systems thereon are far from perfect that patches and upgrades are frequently applied, in order to close vulnerabilities, add new features and enhance performance. To mitigate the loss of availability, such operating systems need to provide features such as live update through which patches and upgrades can be applied without having to stop and reboot the operating system. Unfortunately, most current live updating approaches cannot be easily applied to existing operating systems: some are tightly bound to specific design approaches (e.g. object-oriented); others can only be used under particular circumstances (e.g. quiescence states).In this paper, we propose using virtualization to provide the live update capability. The proposed approach allows a broad range of patches and upgrades to be applied at any time without the requirement of a quiescence state. Moreover, such approach shares good portability for its OS-transparency and is suitable for inclusion in general virtualization systems. We present a working prototype, LUCOS, which supports live update capability on Linux running on Xen virtual machine monitor. To demonstrate the applicability of our approach, we use real-life kernel patches from Linux kernel 2.6.10 to Linux kernel 2.6.11, and apply some of those kernel patches on the fly. Performance measurements show that our implementation incurs negligible performance overhead: a less than 1% performance degradation compared to a Xen-Linux. The time to apply a patch is also very minimal.
Haibo Chen 0001, Rong Chen 0001, Fengzhe Zhang, Binyu Zang, Pen-Chung Yew
VEE3