Beilei Jiang

dblp:251/3222 · DBLP profile ↗
← Back
5ranked-venue papers
3as first author
3since 2021 · last 2023
0009-0001-3985-333XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 3 first-author · 3 since 2021Computer networks · 1
YearPublicationVenuePosition
2023 Output-Directed Dynamic Quantization for DNN Acceleration
abstract
Quantization is an effective technique for reducing the number of computations and improving the performance of deep neural networks (DNNs). Weight quantization is popular because weights can be trained beforehand. However, weight quantization only targets the kernel weights and ignores the sensitivity of input features, which can lead to reduced accuracy. Fine-grained input quantization has gained attention as a way to speed up DNNs while maintaining accuracy. Existing approaches determine computation precision based on input sensitivity but do not effectively reduce computations for insensitive outputs or retain the precision of sensitive outputs. These limitations motivate us to develop an output-directed dynamic quantization method named ODQ in this paper. ODQ is a two-stage DNN quantization scheme designed to improve performance, reduce energy consumption, and maintain and often improve accuracy, compared with existing quantization methods. Specifically, inputs and weights go through sensitivity prediction and result generation. The high-order 2 bits of input and weight are used to predict output sensitivity. Result generation is performed only for predicted sensitive outputs. We designed an FPGA accelerator to optimize ODQ quantization performance for DNNs. We implement a prototype of ODQ and evaluate its performance using several state-of-the-art DNNs. Compared with a state-of-the-art input-directed quantization approach, ODQ achieves a 67.6% performance speedup and a 66.9% energy saving, with minimal accuracy degradation (≤ 0.6%).
Beilei Jiang, Xianwei Cheng, Yuan Li 0054, Jocelyn Zhang, Song Fu, Qing Yang 0003, Mingxiong Liu, Alejandro Olvera
ICPP1
2022 MLCNN: Cross-Layer Cooperative Optimization and Accelerator Architecture for Speeding Up Deep Learning Applications
abstract
The ever-increasing number of layers, millions of parameters, and large data volume make deep learning workloads resource-intensive and power-hungry. In this paper, we develop a convolutional neural network (CNN) acceleration framework, named MLCNN, which explores algorithm-hardware co-design to achieve cross-layer cooperative optimization and acceleration. MLCNN dramatically reduces computation and on-off chip communication, improving CNN's performance. To achieve this, MLCNN reorders the position of nonlinear activation layers and pooling layers, which we prove results in a negligible accuracy loss; then the convolutional layer and pooling layer are co-optimized by means of redundant multiplication elimination, local addition reuse, and global addition reuse. To the best of our knowledge, MLCNN is the first of its kind that incorporates cooperative optimization across convolutional, activation, and pooling layers. We further customize the MLCNN accelerator to take full advantage of cross-layer CNN optimization to reduce both computation and on-off chip communication. Our analysis shows that MLCNN can significantly reduce (up to 98%) multiplications and additions. We have implemented a prototype of MLCNN and evaluated its performance on several widely used CNN models using both an accelerator-level cycle and energy model and RTL implementation. Experimental results show that MLCNN achieves 3.2x speedup and 2.9x energy efficiency compared with dense CNNs. MLCNN's optimization methods are orthogonal to other CNN acceleration techniques, such as quantization and pruning. Combined with quantization, our quantized MLCNN gains a 12.8x speedup and 11.3x energy efficiency compared with DCNN.
Beilei Jiang, Xianwei Cheng, Sihai Tang, Xu Ma 0005, Zhaochen Gu, Song Fu, Qing Yang 0003, Mingxiong Liu
IPDPS1
2021 APCNN: Explore Multi-Layer Cooperation for CNN Optimization and Acceleration on FPGA
abstract
In this paper, we introduce APCNN, which explores algorithm-hardware co-design and provides a CNN acceleration framework with multi-layer cooperative optimization and customized design on FPGA. In terms of the algorithm design, the pooling layer is moved before the non-linear activation function and normalization in APCNN, which we prove causes negligible accuracy loss; the pooling layer is then co-optimized with the convolutional layer by means of redundant multiplication elimination, local addition reuse, and global addition reuse. We further design a dedicated accelerator to take full advantage of convolutional-pooling cross-layer optimization to not only accelerate computation but also reduce on-off chip data communication on FPGA. We demonstrate that our novel APCNN can achieve 75% multiplication and 75% addition reduction in the best case. For on-off chip data communication, a max{Row,Col} /(Row x Col) percent of memory footprint can be eliminated, where Row and Col are the number of rows and columns in the activation feature map respectively. We have implemented a prototype of APCNN and evaluated its performance on LeNet-5 and VGG16 using both an accelerator-level cycle and energy model and an RTL implementation. Our experimental results show that APCNN achieves a 2.5× speedup and 4.7× energy efficiency compared with the dense CNN. (This research was supported in part by NSF grants CCF-1563750, OAC-2017564, and CNS-2037982.)
Beilei Jiang, Xianwei Cheng, Sihai Tang, Xu Ma 0005, Zhaochen Gu, Hui Zhao 0013, Song Fu
FPGA1
2020 AMOEBA: a coarse grained reconfigurable architecture for dynamic GPU scaling
abstract
Different GPU applications exhibit varying scalability patterns with network-on-chip (NoC), coalescing, memory and control divergence, and L1 cache behavior. A GPU consists of several Streaming Multi-processors (SMs) that collectively determine how shared resources are partitioned and accessed. Recent years have seen divergent paths in SM scaling towards scale-up (fewer, larger SMs) vs. scale-out (more, smaller SMs). However, neither scaling up nor scaling out can meet the scalability requirement of all applications running on a given GPU system, which inevitably results in performance degradation and resource under-utilization for some applications. In this work, we investigate major design parameters that influence GPU scaling. We then propose AMOEBA, a solution to GPU scaling through reconfigurable SM cores. AMOEBA monitors and predicts application scalability at run-time and adjusts the SM configuration to meet program requirements. AMOEBA also enables dynamic creation of heterogeneous SMs through independent fusing or splitting. AMOEBA is a microarchitecture-based solution and requires no additional programming effort or custom compiler support. Our experimental evaluations with application programs from various benchmark suites indicate that AMOEBA is able to achieve a maximum performance gain of 4.3x, and generates an average performance improvement of 47% when considering all benchmarks tested.
Xianwei Cheng, Hui Zhao 0013, Mahmut T. Kandemir, Beilei Jiang, Gayatri Mehta
ICS4
2019 A Low-Cost and Energy-Efficient NoC Architecture for GPGPUs
abstract
GPGPU accelerated systems demand high throughput in data communication in order to fully exploit thread-level parallelism. Most of current GPGPU Network-on-Chips (NoCs) employ topology adapted from CPUs, such as mesh and crossbar. However, the trade-off between performance and cost for such networks is sub-optimal, due to the unique traffic pattern of GPUs. In this work, we propose a novel NoC architecture called fused fat tree which modifies the fat tree to match GPU traffic pattern. By separately connecting memory controllers and computing cores to tree roots and leaves, protocol deadlocks can be avoided using just one physical network. However, this modification removes the advantage of path diversity in the original fat tree topology and makes the network vulnerable to hotspot-caused congestion. To solve this problem, we propose to fuse routers with side links to create multiple paths. A load-balancing routing algorithm is also proposed in order to increase network throughput. We also propose a novel preemptive bandwidth allocation scheme to improve resource utilization by taking advantage of request message slacks. Our evaluation results show that our design can improve performance by 46% while achieving 27 % and 25 % area and energy savings on the average.
Xianwei Cheng, Yang Zhao 0013, Mohammadreza Robaei, Beilei Jiang, Hui Zhao 0013, Juan Fang 0004
ANCS4