Jianzhuang Lu

dblp:78/3499 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
8since 2021 · last 2026
0009-0008-3049-5573ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 A hardware-efficient FPGA-based YOLOv5 accelerator with operator fusion and unified dataflow scheduling
Libo Huang 0002, Run Yan, Lei Wang 0011, Jianzhuang Lu
Future Gener. Comput. Syst.5
2026 WSGraph: A Framework for Tackling Redundant and Irregular Data Access in Streaming Graph Processing
abstract
The demand for real-time streaming graph analysis has grown significantly, as hundreds of thousands of updates come every second. Monotonic graph algorithms such as Shortest Path are widely used in real-time analytics, but there are two bottlenecks that limit their performance, specially on planar graphs. One is massive redundant data accesses due to irregular state propagations and the other is high memory latency caused by irregular data accesses. We observe that existing systems mainly focus on general-purpose graph algorithms. If the properties of specific graph algorithms are exploited, the analysis performance can be further improved. Moreover, these systems typically tackle these two bottlenecks separately through either software or hardware mechanisms, but not both. However, both bottlenecks need to be addressed simultaneously in real scenarios such as road navigation. This article proposes WSGraph, a software-hardware co-design framework for high-performance streaming graph processing. WSGraph tackles these two challenges by enforcing regularized processing orders and enabling precise data prefetching. Specifically, at the software level, WSGraph integrates a priority-based work scheduler with sliding-window bucket mapping scheme to regulate state propagations, thereby drastically reducing redundant data accesses. At the hardware level, WSGraph incorporates a lightweight in-core Proactive Data Engine (PDE). By exploiting intra-vertex access regularity, the PDE accurately prefetches relevant graph data to effectively hide the high latency of irregular memory accesses. Experimental results demonstrate that WSGraph achieves significant performance improvements over existing systems. Compared with the state-of-the-art software system KickStarter, WSGraph gains a 2.13× speedup primarily by reducing graph data accesses by an average of 78.6%.
Xuanyi Li, Chen Li 0015, Zhengyi Dai, Jianzhuang Lu, Yang Guo 0003
ACM Trans. Archit. Code Optim.5
2025 LAMP: A Locality-Aware TLB Sharing Framework for Multi-Tenant GPUs via TLB Comprehensive Profiling
abstract
With the growing prevalence of cloud services, GPUs are increasingly shared across multiple tenants, making address translation a performance-critical path. Our study makes three key observations: 1) Shared L2 TLB becomes the bottleneck in multi-tenant systems. Intense competition for entries can lead to high miss rates and thrashing under heavy pressure. 2) Benchmarks vary significantly in TLB sensitivity due to differing access patterns and memory reuse potential. 3) The prevailing management strategies, including free-for-all sharing and static partitioning, fail to account for tenant access behavior. And the state-of-the-art token-based strategy relies solely on miss rate, an inadequate metric that fails to capture per-tenant access patterns. Based on these observations, we propose LAMP, a localityaware TLB sharing framework. LAMP comprehensively profiles the access pattern considering three key aspects: spatial locality, temporal locality, and access density. These aspects allow LAMP to assess each tenant's sensitivity to TLB capacity and their reuse potential. Based on these insights, LAMP optimizes the sharing by dynamically partitioning L2 TLB, granting more entries to reuse-efficient tenants while limiting wasteful occupancy. This reduces thrashing and cross-tenant interference. Experimental results show that LAMP improves system performance by$\mathbf{1 2. 3 \%}$on average across a range of multi-tenant workloads, with negligible hardware overhead.
Chen Li 0015, Xuanyi Li, Jianzhuang Lu, Yang Guo 0003
HPCC5
2025 A Multi-objective Mapping Optimization Strategy for Large-Scale NoC Design
Chen Li 0015, Jianzhuang Lu, Yang Guo 0003
ICA3PP (5)4
2024 AFTP: An Adaptive High Cost-Effectiveness Fault-Tolerant NoC Design Based on Prediction
abstract
Network-on-Chip (NoC), known for its high bandwidth and scalability, is extensively utilized in chip multiprocessors. However, as technology advances to the nanometer scale, NoC is becoming increasingly vulnerable to errors caused by crosstalk, radiation, electromagnetic interference, and other factors. In addition to ensuring network reliability, designers must consider the overhead instead of blindly pursuing fault-tolerance capabilities in NoC design. In this paper, we analyze the characteristics of conventional End-to-End (E2E) and Switch-to-Switch (S2S) designs and propose AFTP, an adaptive high cost-effectiveness fault-tolerant NoC design based on prediction. Our design prioritizes stronger protection for head flits and relaxed protection for body and tail flits, and it is capable of dynamically adjusting the decoding times for flits in response to changes in error rate, aiming to achieve a balance between the overhead and the reliability of NoC. Our design demonstrates a significant improvement in cost-effectiveness compared to conventional E2E and S2S designs, achieving a 5.9x and 8.9x improvement, respectively, under common synthetic traffic and PARSEC benchmarks.
Ta Tan, Chen Li 0015, Jianzhuang Lu, Qijin Zhu
ISPA4
2024 Load balancing-oriented fault-tolerant NoC design
abstract
Network-on-Chip (NoC) has been widely applied in modern chip multiprocessors due to its high bandwidth and scalability. However, as technology advances to the nanometer scale, NoC is increasingly vulnerable to errors caused by crosstalk, radiation, electromagnetic interference, etc. Conventional Switch-to-Switch (S2S) fault-tolerant designs based on ECC have overlooked the characteristic of the distribution of traffic load. This oversight not only increases area overhead significantly but also leads to low average utilization of ECC decoder modules. In this paper, we analyze the distribution of traffic load in mesh network and propose a load balancing-oriented fault-tolerant NoC design. The core idea is to allocate different numbers of ECC decoder modules to each router based on the distribution of traffic load, aiming to improve the average utilization of ECC decoder modules and reduce the area overhead without compromising fault-tolerant capability of NoC. The experiment under 6 common synthetic traffic patterns shows that compared to the baseline, our design exhibits an average delay performance loss of less than 0.88%. Additionally, the maximum reduction in the number of ECC decoder modules is 160, the maximum reduction in the area overhead of NoC is 15.06%, and the maximum improvement in the average utilization of ECC decoder modules is 1.21x. Furthermore, the experiment under PARSEC benchmarks shows that compared to the baseline, our design exhibits an average delay performance loss of less than 0.08%. Additionally, the maximum reduction in the number of ECC decoder modules is 156, the maximum reduction in total NoC area overhead is 14.69%, and the maximum improvement in the average utilization of ECC decoder modules is 1.13x.
Ta Tan, Jianzhuang Lu
ITC-Asia4
2022 Convolutional Neural Network Accelerator for Compression Based on Simon k-means
abstract
Convolutional Neural Networks (CNN) are popular models widely used in image classification, target recognition, and other fields. FPGA-based accelerators for CNN are a standard method in recent years to reduce CNN's inference time and energy efficiency. However, the limitations of on-chip storage space and computing resources introduce deep compression. Contrary to most compression algorithms that pay no attention to the underlying hardware acceleration strategy and hardware-only accelerators, this paper introduces a novel model compression scheme with software and hardware collaboration for accelerating inference. First, we propose a pre-processing algorithm named Simon k-means based on clustering to quantify trained weight to speed up inference. Next, we propose a new encoding method for the quantized weight, significantly reducing the model's storage size. Finally, we give the architecture design of the accelerator using the quantized weight to accelerate the convolution. We have evaluated many popular CNNs in image classification tasks on various data sets. Experiments show that the number of multiply-accumulate operations on the convolutional layer can be reduced 66.6% with a slight loss of precision.
Yunping Zhao, Jianzhuang Lu, Zerun Li
IJCNN3
2021 Accelerating Depthwise Separable Convolutions with Vector Processor
Yuekai Zhao, Jianzhuang Lu
ICANN (2)2
2020 Dynamic GMMU Bypass for Address Translation in Multi-GPU Systems
Jinhui Wei, Jianzhuang Lu, Yunping Zhao
NPC2
2011 Design and chip implementation of a heterogeneous multi-core DSP
abstract
This paper presents a novel heterogeneous multi-core Digital Signal Processor, named YHFT-QDSP, hosting one RISC CPU core and four VLIW DSP cores. The CPU core is responsible for task scheduling and management, while the DSP cores take charge of speeding up data processing. The YHFT-QDSP provides three kinds of interconnection communication. One is for inner-chip communication between the CPU core and the four DSP cores, the other two for both inner-chip and inter-chip communication amongst DSP cores. The YHFT-QDSP is implemented under SMIC®130nm LVT CMOS technology and can run [email protected] with 114.49 mm2die area.
Shuming Chen, Jianghua Wan, Jianzhuang Lu, Xiangyuan Liu, Shenggang Chen
ASP-DAC5
2010 YHFT-QDSP: High-Performance Heterogeneous Multi-Core DSP
Shuming Chen, Jianghua Wan, Jianzhuang Lu, Hai-Yan Sun, Yong-Jie Sun, Hengzhu Liu, Xiang-Yuan Liu, Zhentao Li
J. Comput. Sci. Technol.3
2004 A Case of SCMP with TLS
Jianzhuang Lu, Chunyuan Zhang
ISPA1