EDBT 2026 Demo / reviewers in the wild / expert
Lu Zhang 0074
dblp:82/10609-74
· DBLP profile ↗
7ranked-venue papers
3as first author
4since 2021 · last 2024
0000-0002-5136-1119ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 3 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | ULSeq-TA: Ultra-Long Sequence Attention Fusion Transformer Accelerator Supporting Grouped Sparse Softmax and Dual-Path Sparse LayerNormabstractTransformer networks have been increasingly successful in various fields. The input sequence lengths have become much larger as the algorithm and task complexity develops, which is challenging due to high computational and storage cost. Softmax and LayerNorm are bottleneck nonlinear operators in ultra-long sequence Transformer networks. To improve the efficiency of Softmax, assumption-based and quantization-based Softmax approaches are introduced. However, the sparsity potential to accelerate Softmax itself is not fully discovered. To improve the efficiency of LayerNorm, some works reduce the input size, and some works explore the pipeline. However, the sparsity potential is also not yet explored. To address these challenges, this article presents the ULSeq-TA software–hardware co-design framework. The software includes 1) the grouped sparse Softmax method to leverage the data magnifying characteristic to explore the middle and post-Softmax sparse processing and 2) the dual-path sparse LayerNorm method which explores the dimensional significance for sparse calculation. The hardware includes 1) an attention fusion architecture which reduces the on-chip memory with fused operators; 2) the grouped sparse Softmax core; and 3) the dual-path sparse LayerNorm core. Experiments show that the software achieves$4.45\times $and$7.59\times $computation reduction with little output difference for Softmax and LayerNorm, respectively. The hardware architecture supports at most 32768 sequence length with only 186-kB on-chip memory and achieves$1.75\times -1.98\times $and$3.22\times -4.32\times $speedups for sparse Softmax core and sparse LayerNorm core with little accuracy loss, respectively. Jingyu Wang 0004, Lu Zhang 0074, Xueqing Li 0002, Huazhong Yang, Yongpan Liu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2024 | RE-Specter: Examining the Architectural Features of Configurable CNN With Power Side-ChannelabstractAs domain-specific training data is recognized as valuable intellectual property, acquiring well-trained weights in Convolutional Neural Networks (CNN) has emerged as a new threat to the neural network design community. To design a CNN accelerator that is resilient to side-channel threats, it is crucial to have an accurate and efficient security-driven framework at the early design stage. However, there is no standard way to perform root-cause analysis on the power side channel that exists in FPGA-based CNN accelerators. Therefore, we build RE-Specter, a framework that facilitates security-driven design space exploration (DSE) across various building components, combination patterns, and parallelism configurations in CNNs. The goal is to fully understand the power side-channel effects resulting from architectural modifications or optimization decisions. We further compare the benchmarks considering precision, resource utilization, and power side-channel leakage. Finally, we experimentally explore the design space of various architectural features. The experimental results show that low-bit precision delivers more secure architectures (68.9× among DSPs, 2439× among LUTs) in Measurement-To-Disclosure (MTD), but mixed-precision strategies are necessary to maintain the model accuracy. For loop optimization, in 16-parallel scenario, accumulator-based architecture outperforms the architecture featuring an adder tree with the improvements of 8.28× in MTD and 1.38× in PST. Lu Zhang 0074, Jingyu Wang 0004, Ruoyang Liu, Yifan He 0003, Yaolei Li, Yu Tai, Shengbing Zhang, Xiaoya Fan, Huazhong Yang, Yongpan Liu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2022 | C-RRAM: A Fully Input Parallel Charge-Domain RRAM-based Computing-in-Memory Design with High Tolerance for RRAM VariationsabstractPrevious RRAM-based computing-in-memory works mainly focus on the current-domain approach. However, the performance and accuracy of current computation are limited by large read currents of RRAM cells and their variations. This work presents a novel RRAM-based charge-domain design, C-RRAM, to resolve these limitations. A 3TlRlC cell is proposed to execute MAC operations by capacitor discharging. The resistance variations can be tolerated with reasonable discharging time, which is accelerated by a positive feedback loop. Also, the output from each cell is accumulated by charge sharing instead of summing currents to eliminate the static current path in readout circuits. In this way, robust and efficient RRAM-based CIM operation is enabled with fully input parallelism. A $512\times 514$ RRAM array is implemented to evaluate the benefits of the proposed charge-domain approach. The experiment results show that C-RRAM can suppress the output variations by $41\times$ and incur negligible accuracy loss for ResNet-18 on Cifar10 dataset. Compared to previous ITIR current-domain RRAM designs, it achieves $1.2\times$ energy efficiency and $127\times$ area efficiency due to improved parallelism. Yifan He 0003, Jinshan Yue, Wenyu Sun, Lu Zhang 0074, Yongpan Liu |
ISCAS | 5 |
| 2021 | PETRI: Reducing Bandwidth Requirement in Smart Surveillance by Edge-Cloud Collaborative Adaptive Frame Clustering and Pipelined Bidirectional TrackingabstractNeural networks running on cloud servers have been widely used in smart surveillance, but they require high bandwidth to upload videos. Edge-cloud collaborative encoding based on ROI (Region-Of-Interest) can reduce bandwidth requirement, but it suffers from inaccurate ROI detection due to feedback latency and undetected new targets. To address the above challenges, we propose an object detection system named PETRI. It adopts a latency-hiding pipeline workflow with adaptive keyframe interval selection for different input videos, and utilizes a retro-tracking method to find undetected targets. While achieving negligible impact on model accuracy, the proposed PETRI can save up to 66.44% and 30.25% bandwidth compared with the cloud only method and the previous state-of-art work respectively. Ruoyang Liu, Lu Zhang 0074, Jingyu Wang 0004, Huazhong Yang, Yongpan Liu |
DAC | 2 |
| 2020 | Memory-Based High-Level Synthesis Optimizations Security Exploration on the Power Side-ChannelabstractHigh-level synthesis (HLS) allows hardware designers to think algorithmically and not worry about low-level, cycle-by-cycle details. This provides the ability to quickly explore the architectural design space and tradeoffs between resource utilization and performance. Unfortunately, security evaluation is not a standard part of the HLS design flow. In this article, we aim to understand the effects of memory-based HLS optimizations on power side-channel leakage. We use Xilinx Vivado HLS to develop different cryptographic cores, implement them on a Spartan-6 FPGA, and collect power traces. We evaluate the designs with respect to resource utilization, performance, and information leakage through power consumption. We have two important observations and contributions. First, the choice of resource optimization directive results in different levels of side-channel vulnerabilities. Second, the partitioning optimization directive can greatly compromise the hardware cryptographic system through power side-channel leakage due to the deployment of memory control logic. We describe an evaluation procedure for power side-channel leakage and use it to make best-effort recommendations about how to design more secure architectures in the cryptographic domain. Lu Zhang 0074, Wei Hu 0008, Yu Tai, Jeremy Blackstone, Ryan Kastner |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2018 | Examining the consequences of high-level synthesis optimizations on power side-channelabstractHigh-level synthesis (HLS) allows hardware designers to think algorithmically and not have to worry about low-level, cycle-by-cycle details. This provides the ability to quickly explore the architectural design space and tradeoff between resource utilization and performance. Unfortunately, evaluating the security is not a standard part of the HLS design flow. In this work, we aim to understand the effects of HLS optimizations with respect to power side-channel leakage. We use Vivado HLS to develop different cryptographic cores, implement them on a Xilinx Spartan 6 FPGA, and collect power traces. We evaluate the designs with respect to resource utilization, performance, and side-channel leakage through power consumption. Furthermore, we analyze the first-order leakage of the HLS-based designs alongside well-known register transfer level (RTL) cryptographic cores. We describe an evaluation procedure for hardware designers and use it to make insightful recommendations on how to design the best architecture in cryptographic domain. Lu Zhang 0074, Wei Hu 0008, Armita Ardeshiricham, Yu Tai, Jeremy Blackstone, Ryan Kastner |
DATE | 1 |
| 2017 | Why you should care about don't cares: Exploiting internal don't care conditions for hardware TrojansabstractHardware Trojans are a significant security threat due to the globalization of hardware design and supply chain. We demonstrate a new type of hardware Trojan hidden behind internal don't care conditions. The proposed Trojans can pass through formal equivalence checking; they may reside after logic synthesis optimizations; and they are resilient to switching probability and side channel analysis. The new Trojans can create a surface for fault attack to retrieve secret information or downgrade performance by increasing power consumption. Experimental results show that these Trojans may stay after logic synthesis and that secret information can be retrieved using fault attack. We present detectability analysis and suggest synthesis optimizations as well as countermeasures that can help mitigate this new Trojan. Wei Hu 0008, Lu Zhang 0074, Armita Ardeshiricham, Jeremy Blackstone, Bochuan Hou, Yu Tai, Ryan Kastner |
ICCAD | 2 |