EDBT 2026 Demo / reviewers in the wild / expert
Ruoyang Liu
dblp:233/8119
· DBLP profile ↗
7ranked-venue papers
2as first author
5since 2021 · last 2024
0000-0001-9873-6574ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | RE-Specter: Examining the Architectural Features of Configurable CNN With Power Side-ChannelabstractAs domain-specific training data is recognized as valuable intellectual property, acquiring well-trained weights in Convolutional Neural Networks (CNN) has emerged as a new threat to the neural network design community. To design a CNN accelerator that is resilient to side-channel threats, it is crucial to have an accurate and efficient security-driven framework at the early design stage. However, there is no standard way to perform root-cause analysis on the power side channel that exists in FPGA-based CNN accelerators. Therefore, we build RE-Specter, a framework that facilitates security-driven design space exploration (DSE) across various building components, combination patterns, and parallelism configurations in CNNs. The goal is to fully understand the power side-channel effects resulting from architectural modifications or optimization decisions. We further compare the benchmarks considering precision, resource utilization, and power side-channel leakage. Finally, we experimentally explore the design space of various architectural features. The experimental results show that low-bit precision delivers more secure architectures (68.9× among DSPs, 2439× among LUTs) in Measurement-To-Disclosure (MTD), but mixed-precision strategies are necessary to maintain the model accuracy. For loop optimization, in 16-parallel scenario, accumulator-based architecture outperforms the architecture featuring an adder tree with the improvements of 8.28× in MTD and 1.38× in PST. Lu Zhang 0074, Jingyu Wang 0004, Ruoyang Liu, Yifan He 0003, Yaolei Li, Yu Tai, Shengbing Zhang, Xiaoya Fan, Huazhong Yang, Yongpan Liu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2023 | Block-Wise Dynamic-Precision Neural Network Training Acceleration via Online Quantization Sensitivity AnalyticsabstractData quantization is an effective method to accelerate neural network training and reduce power consumption. However, it is challenging to perform low-bit quantized training: the conventional equal-precision quantization will lead to either high accuracy loss or limited bit-width reduction, while existing mixed-precision methods offer high compression potential but failed to perform accurate and efficient bit-width assignment. In this work, we propose DYNASTY, a block-wise dynamic-precision neural network training framework. DYNASTY provides accurate data sensitivity information through fast online analytics, and maintains stable training convergence with an adaptive bit-width map generator. Network training experiments on CIFAR-100 and ImageNet dataset are carried out, and compared to 8-bit quantization baseline, DYNASTY brings up to 5.1× speedup and 4.7× energy consumption reduction with no accuracy drop and negligible hardware overhead. Ruoyang Liu, Chenhan Wei, Yixiong Yang, Wenxun Wang, Huazhong Yang, Yongpan Liu |
ASP-DAC | 1 |
| 2022 | Toward Low-Bit Neural Network Training Accelerator by Dynamic Group AccumulationabstractLow-bit quantization is a big challenge for neural network training. Conventional training hardware adopts FP32 to accumulate the partial-sum result, which seriously degrades energy efficiency. In this paper, a technology called dynamic group accumulation (DGA) is proposed to reduce the accumulation error. First, we model the proposed group accumulation method and give the optimal DGA algorithm. Second, we design a training architecture and implement a hardware-efficient DGA unit. Third, we make a comprehensive analysis of the DGA algorithm and training architecture. The proposed method is evaluated on CIFAR and ImageNet datasets, and results show that DGA can reduce accumulation bit-width by 6 bits while achieving the same precision as the static group method. With the FP12 DGA, the CNN algorithm only loses 0.11% accuracy in ImageNet training, and our architecture saves 32% of power consumption compared to the FP32 baseline. Yixiong Yang, Ruoyang Liu, Wenyu Sun, Jinshan Yue, Huazhong Yang, Yongpan Liu |
ASP-DAC | 2 |
| 2022 | PACA: A Pattern Pruning Algorithm and Channel-Fused High PE Utilization Accelerator for CNNsabstractIn recent years, convolutional neural networks (CNNs) have achieved significant advancements in various fields. However, the computation and storage overheads of CNNs are overwhelming for Internet-of-Things devices. Both network pruning algorithms and hardware accelerators have been introduced to empower CNN inference at the edge. Network pruning algorithms reduce the size and computational cost of CNNs by regularizing unimportant weights to zeros. However, existing works lack intrakernel structured types to tradeoff between sparsity and hardware efficiency, and the index storage for irregularly pruned networks is significant. Hardware accelerators leverage the sparsity of pruned CNNs to improve energy efficiency. However, their process element (PE) utilization rate is low because of uneven sparsity among input convolutional kernels. To overcome these problems, we propose PACA: a Pattern pruning Algorithm and Channel-fused high PE utilization Accelerator for CNNs. It includes three parts: a pattern pruning algorithm to explore the intrakernel sparsity type and reduce the index storage, a channel-fused hardware architecture to reduce the PEs’ idle rate and improve the performance, and a heuristic and taboo search-based smart fusion scheduler to analyze the idle PE problem and schedule the channel fusion in hardware. To demonstrate the effectiveness of PACA, we have implemented the software parts by Python and the hardware architecture by RTL codes. Experimental results on various datasets show that compared with an existing work, PACA can reduce the index storage overhead by$3.47\times $–$5.63\times $with 3.85–9.12 average patterns, and it can improve the hardware performance by$2.01\times $–$5.53\times $because of PEs’ idle rate reduction. Jingyu Wang 0004, Songming Yu, Zhuqing Yuan, Jinshan Yue, Ruoyang Liu, Yanzhi Wang 0001, Huazhong Yang, Xueqing Li 0002, Yongpan Liu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2021 | PETRI: Reducing Bandwidth Requirement in Smart Surveillance by Edge-Cloud Collaborative Adaptive Frame Clustering and Pipelined Bidirectional TrackingabstractNeural networks running on cloud servers have been widely used in smart surveillance, but they require high bandwidth to upload videos. Edge-cloud collaborative encoding based on ROI (Region-Of-Interest) can reduce bandwidth requirement, but it suffers from inaccurate ROI detection due to feedback latency and undetected new targets. To address the above challenges, we propose an object detection system named PETRI. It adopts a latency-hiding pipeline workflow with adaptive keyframe interval selection for different input videos, and utilizes a retro-tracking method to find undetected targets. While achieving negligible impact on model accuracy, the proposed PETRI can save up to 66.44% and 30.25% bandwidth compared with the cloud only method and the previous state-of-art work respectively. Ruoyang Liu, Lu Zhang 0074, Jingyu Wang 0004, Huazhong Yang, Yongpan Liu |
DAC | 1 |
| 2020 | GAAS: An Efficient Group Associated Architecture and Scheduler Module for Sparse CNN AcceleratorsabstractConvolutional neural networks (CNNs) have become powerful algorithms in various tasks. Application-specific integrated circuit (ASIC) has been widely used to accelerate CNN on mobile platforms because of its tremendous energy efficiency and performance. Meanwhile, CNNs have become much sparser with the development of network pruning algorithms. Recent works have employed different methods to improve the energy efficiency and performance of ASIC accelerators by utilizing the sparsity character of CNN. However, some of these methods suffer from tremendous output memory overhead and performance degradation induced by hash collisions. To overcome the aforementioned problem, we propose GAAS: an efficient group associated architecture and scheduler module for sparse CNN accelerators. It achieves smaller output memory overhead and higher performance compared with the state-of-the-art accelerator. Our proposed method GAAS mainly consists of two parts: 1) an$n$-way group associated architecture to reduce the output memory overhead and 2) a scheduler module to improve the performance. Besides, a load-balancing algorithm is proposed and implemented in the scheduler module to improve the performance by reducing the hash collision rate. To demonstrate the efficiency of GAAS, we implement a 4-way image-principal associated architecture with a$16\times 16$PE array and the scheduler module on our proposed method. The experimental results on AlexNet, VGG16, ResNet18, and MobileNet show that GAAS can reduce the output memory overhead by 50%, and it can surely improve the performance of them by$1.53\times $,$1.62\times $,$1.46\times $, and$1.55\times $, respectively. Jingyu Wang 0004, Ruoyang Liu, Xiaoyu Feng, Huazhong Yang, Yongpan Liu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2019 | An N-way group association architecture and sparse data group association load balancing algorithm for sparse CNN acceleratorsabstractIn recent years, ASIC CNN Accelerators have attracted great attention among researchers for the high performance and energy efficiency. Some former works utilize the sparsity of CNN networks to improve the performance and the energy efficiency. However, these methods bring tremendous overhead to the output memory, and the performance suffers from the hash collision. This paper presents: 1) an N-Way Group Association Architecture to reduce the memory overhead for Sparse CNN Accelerators; 2) a Sparse Data Group Association Load Balancing Algorithm which is implemented by the Scheduler module in the architecture to reduce the collision rate and improve the performance. Compared with the state-of-art accelerator, this work achieves either 1) 1.74x performance with 50% memory overhead reduction in the 4-way associated design or 2) 1.91x performance without memory overhead reduction the 2-way associated design, which is close to the theoretical performance limit (without collision). Jingyu Wang 0004, Ruoyang Liu, Huazhong Yang, Yongpan Liu |
ASP-DAC | 3 |