Zhe Li 0001

dblp:11/751-1 · DBLP profile ↗
← Back
19ranked-venue papers
5as first author
0since 2021 · last 2020
0000-0001-7056-4133ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 4 first-authorArtificial intelligence and machine learning · 5 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
6 papers
Hardware accelerators and domain-specific architectures · 72% Reconfigurable computing and FPGAs · 14% Emerging computing paradigms · 12%
Artificial intelligence
7 papers
Efficient and distributed learning · 72% Learning theory · 19% Deep learning architectures and training · 6%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 100%

Topics — the 22 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
machine learning accelerator
1.752019
HEIF: Highly Efficient Stochastic Computing-Based Inference Framework for Deep Neural Networks · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019
E-RNN: Design Optimization for Efficient Recurrent Neural Networks in FPGAs · HPCA 2019
Towards Ultra-High Performance and Energy Efficiency of Deep Learning Systems: An Algorithm-Hardware Co-Optimization Framework · AAAI 2018
Machine learning › Efficient and distributed learning
model compression
1.142019
C-LSTM: Enabling Efficient LSTM using Structured Compression Techniques on FPGAs · FPGA 2018
Towards Ultra-High Performance and Energy Efficiency of Deep Learning Systems: An Algorithm-Hardware Co-Optimization Framework · AAAI 2018
CirCNN: accelerating and compressing deep neural networks using block-circulant weight matrices · MICRO 2017
Emerging computing paradigms › approximate and stochastic computing
stochastic computing
0.722019
HEIF: Highly Efficient Stochastic Computing-Based Inference Framework for Deep Neural Networks · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019
SC-DCNN: Highly-Scalable Deep Convolutional Neural Network using Stochastic Computing · ASPLOS 2017
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
stochastic computing accelerator
0.722019
HEIF: Highly Efficient Stochastic Computing-Based Inference Framework for Deep Neural Networks · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019
SC-DCNN: Highly-Scalable Deep Convolutional Neural Network using Stochastic Computing · ASPLOS 2017
Machine learning › Efficient and distributed learning › model compression › weight compression
block-circulant matrix compression
0.422019
Towards Ultra-High Performance and Energy Efficiency of Deep Learning Systems: An Algorithm-Hardware Co-Optimization Framework · AAAI 2018
E-RNN: Design Optimization for Efficient Recurrent Neural Networks in FPGAs · HPCA 2019
Reconfigurable computing and FPGAs
FPGA accelerator
0.412019
E-RNN: Design Optimization for Efficient Recurrent Neural Networks in FPGAs · HPCA 2019
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
RNN accelerator
0.412019
E-RNN: Design Optimization for Efficient Recurrent Neural Networks in FPGAs · HPCA 2019
Hardware accelerators and domain-specific architectures › machine learning accelerator › DNN inference accelerator
RNN inference accelerator
0.412019
E-RNN: Design Optimization for Efficient Recurrent Neural Networks in FPGAs · HPCA 2019
Machine learning › Efficient and distributed learning › model compression › pruning
structured pruning
0.312018
C-LSTM: Enabling Efficient LSTM using Structured Compression Techniques on FPGAs · FPGA 2018
Medical and health informatics › computer-aided diagnosis
thorax disease classification
0.312018
Thoracic Disease Identification and Localization With Limited Supervision · CVPR 2018
Medical and health informatics › computational pathology
tissue abnormality localization
0.312018
Thoracic Disease Identification and Localization With Limited Supervision · CVPR 2018
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
FPGA-based neural network accelerator
0.312018
Towards Ultra-High Performance and Energy Efficiency of Deep Learning Systems: An Algorithm-Hardware Co-Optimization Framework · AAAI 2018
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
LSTM accelerator
0.312018
C-LSTM: Enabling Efficient LSTM using Structured Compression Techniques on FPGAs · FPGA 2018
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator
0.312018
C-LSTM: Enabling Efficient LSTM using Structured Compression Techniques on FPGAs · FPGA 2018
Machine learning › Learning theory › approximation theory
approximation error bound
0.312017
Theoretical Properties for Neural Networks with Weight Matrices of Low Displacement Rank · ICML 2017
Machine learning › Efficient and distributed learning › model compression
neural network compression
0.312017
Theoretical Properties for Neural Networks with Weight Matrices of Low Displacement Rank · ICML 2017
Machine learning › Learning theory › approximation theory › neural network approximation
universal approximation
0.312017
Theoretical Properties for Neural Networks with Weight Matrices of Low Displacement Rank · ICML 2017
Energy-efficient computing › energy-efficient machine learning
energy-efficient inference
0.112019
HEIF: Highly Efficient Stochastic Computing-Based Inference Framework for Deep Neural Networks · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019
Computer vision › Image recognition and object detection
medical image analysis
0.112018
Thoracic Disease Identification and Localization With Limited Supervision · CVPR 2018
Reconfigurable computing and FPGAs
FPGA-based neural network inference
0.112018
C-LSTM: Enabling Efficient LSTM using Structured Compression Techniques on FPGAs · FPGA 2018
Machine learning › Deep learning architectures and training
backpropagation
0.112017
Theoretical Properties for Neural Networks with Weight Matrices of Low Displacement Rank · ICML 2017
Machine learning › Deep learning architectures and training
convolutional neural network
0.112017
SC-DCNN: Highly-Scalable Deep Convolutional Neural Network using Stochastic Computing · ASPLOS 2017

Methods — techniques the papers use, named apart from their topics

block-circulant matrix · 1.4pipelining · 1.0quantization · 0.8ADMM · 0.8stochastic computing · 0.7weakly supervised learning · 0.7structured compression · 0.7reconfiguration · 0.7pruning · 0.7multi-task learning · 0.7weight clustering · 0.4approximate parallel counter · 0.4backpropagation · 0.3
YearPublicationVenuePosition
2020 Database and Benchmark for Early-stage Malicious Activity Detection in 3D Printing
abstract
Increasing malicious users have sought practices to leverage 3D printing technology to produce unlawful tools in criminal activities. It is of vital importance to enable 3D printers to identify the objects to be printed and terminate at early stage if illegal objects are identified. Deep learning yields significant rises in performance in the object recognition tasks. However, the lack of large-scale databases in 3D printing domain stalls the advancement of automatic illegal weapon recognition. This paper presents a new 3D printing image database, namely C3PO, which compromises two subsets for the different system working scenarios. We extract images from the numerical control programming code files of 22 3D models, and then categorize the images into 10 distinct labels. These two sets are designed for identifying: (i). printing knowledge source (G-code) at beginning of manufacturing, (ii). printing procedure during manufacturing. Importantly, we demonstrate that the weapons can be recognized in either scenario using deep learning based approaches using our proposed database. The quantitative results are promising, and the future exploration of the database and the crime prevention in 3D printing are demanding tasks.
Zhe Li 0001, Hongjia Li 0003, Qiyuan An, Qinru Qiu, Wenyao Xu, Yanzhi Wang 0001
ASP-DAC2
2019 E-RNN: Design Optimization for Efficient Recurrent Neural Networks in FPGAs
abstract
Recurrent Neural Networks (RNNs) are becoming increasingly important for time series-related applications which require efficient and real-time implementations. The two major types are Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU) networks. It is a challenging task to have real-time, efficient, and accurate hardware RNN implementations because of the high sensitivity to imprecision accumulation and the requirement of special activation function implementations. Recently two works have focused on FPGA implementation of inference phase of LSTM RNNs with model compression. First, ESE uses a weight pruning based compressed RNN model but suffers from irregular network structure after pruning. The second work C-LSTM mitigates the irregular network limitation by incorporating block-circulant matrices for weight matrix representation in RNNs, thereby achieving simultaneous model compression and acceleration. A key limitation of the prior works is the lack of a systematic design optimization framework of RNN model and hardware implementations, especially when the block size (or compression ratio) should be jointly optimized with RNN type, layer size, etc. In this paper, we adopt the block-circulant matrixbased framework, and present the Efficient RNN (E-RNN) framework for FPGA implementations of the Automatic Speech Recognition (ASR) application. The overall goal is to improve performance/energy efficiency under accuracy requirement. We use the alternating direction method of multipliers (ADMM) technique for more accurate block-circulant training, and present two design explorations providing guidance on block size and reducing RNN training trials. Based on the two observations, we decompose E-RNN in two phases: Phase I on determining RNN model to reduce computation and storage subject to accuracy requirement, and Phase II on hardware implementations given RNN model, including processing element design/optimization, quantization, activation implementation, etc. 1 Experimental results on actual FPGA deployments show that E-RNN achieves a maximum energy efficiency improvement of 37.4× compared with ESE, and more than 2× compared with C-LSTM, under the same accuracy.
Zhe Li 0001, Caiwen Ding, Siyue Wang, Wujie Wen, Youwei Zhuo, Qinru Qiu, Wenyao Xu, Xue Lin 0001, Xuehai Qian, Yanzhi Wang 0001
HPCA1
2019 Fast and Accurate Trajectory Tracking for Unmanned Aerial Vehicles based on Deep Reinforcement Learning
abstract
Continuous trajectory control of fixed-wing unmanned aerial vehicles (UAVs) is complicated when considering hidden dynamics. Due to UAV multi degrees of freedom, tracking methodologies based on conventional control theory, such as Proportional-Integral-Derivative (PID) has limitations in response time and adjustment robustness, while a model based approach that calculates the force and torques based on UAV's current status is complicated and rigid. We present an actor-critic reinforcement learning framework that controls UAV trajectory through a set of desired waypoints. A deep neural network is constructed to learn the optimal tracking policy and reinforcement learning is developed to optimize the resulting tracking scheme. The experimental results show that our proposed approach can achieve 58.14% less position error, 21.77% less system power consumption and 9.23% faster attainment than the baseline. The actor network consists of only linear operations, hence Field Programmable Gate Arrays (FPGA) based hardware acceleration can easily be designed for energy efficient real-time control.
Yilan Li, Hongjia Li 0003, Zhe Li 0001, Haowen Fang, Amit K. Sanyal, Yanzhi Wang 0001, Qinru Qiu
RTCSA3
2019 Normalization and dropout for stochastic computing-based deep convolutional neural networks
Ji Li 0006, Zhe Li 0001, Ao Ren, Caiwen Ding, Jeffrey T. Draper, Shahin Nazarian, Qinru Qiu, Bo Yuan 0001, Yanzhi Wang 0001
Integr.3
2019 HEIF: Highly Efficient Stochastic Computing-Based Inference Framework for Deep Neural Networks
abstract
Deep convolutional neural networks (DCNNs) are one of the most promising deep learning techniques and have been recognized as the dominant approach for almost all recognition and detection tasks. The computation of DCNNs is memory intensive due to large feature maps and neuron connections, and the performance highly depends on the capability of hardware resources. With the recent trend of wearable devices and Internet of Things, it becomes desirable to integrate the DCNNs onto embedded and portable devices that require low power and energy consumptions and small hardware footprints. Recently stochastic computing (SC)-DCNN demonstrated that SC as a low-cost substitute to binary-based computing radically simplifies the hardware implementation of arithmetic units and has the potential to satisfy the stringent power requirements in embedded devices. In SC, many arithmetic operations that are resource-consuming in binary designs can be implemented with very simple hardware logic, alleviating the extensive computational complexity. It offers a colossal design space for integration and optimization due to its reduced area and soft error resiliency. In this paper, we present HEIF, a highly efficient SC-based inference framework of the large-scale DCNNs, with broad applications including (but not limited to) LeNet-5 and AlexNet, that achieves high energy efficiency and low area/hardware cost. Compared to SC-DCNN, HEIF features: 1) the first (to the best of our knowledge) SC-based rectified linear unit activation function to catch up with the recent advances in software models and mitigate degradation in application-level accuracy; 2) the redesigned approximate parallel counter and optimized stochastic multiplication using transmission gates and inverse mirror adders; and 3) the new optimization of weight storage using clustering. Most importantly, to achieve maximum energy efficiency while maintaining acceptable accuracy, HEIF considers holistic optimizations on cascade connection of function blocks in DCNN, pipelining technique, and bit-stream length reduction. Experimental results show that in large-scale applications HEIF outperforms previous SC-DCNN by the throughput of 4.1×, by area efficiency of up to 6.5×, and achieves up to 5.6× energy improvement.
Zhe Li 0001, Ji Li 0006, Ao Ren, Ruizhe Cai, Caiwen Ding, Xuehai Qian, Jeffrey T. Draper, Bo Yuan 0001, Jian Tang 0008, Qinru Qiu, Yanzhi Wang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2018 Towards Ultra-High Performance and Energy Efficiency of Deep Learning Systems: An Algorithm-Hardware Co-Optimization Framework
abstract
Hardware accelerations of deep learning systems have been extensively investigated in industry and academia. The aim of this paper is to achieve ultra-high energy efficiency and performance for hardware implementations of deep neural networks (DNNs). An algorithm-hardware co-optimization framework is developed, which is applicable to different DNN types, sizes, and application scenarios. The algorithm part adopts the general block-circulant matrices to achieve a fine-grained tradeoff of accuracy and compression ratio. It applies to both fully-connected and convolutional layers and contains a mathematically rigorous proof of the effectiveness of the method. The proposed algorithm reduces computational complexity per layer from O(n2) to O(n log n) and storage complexity from O(n2) to O(n), both for training and inference. The hardware part consists of highly efficient Field Programmable Gate Array (FPGA)-based implementations using effective reconfiguration, batch processing, deep pipelining, resource re-using, and hierarchical control. Experimental results demonstrate that the proposed framework achieves at least 152X speedup and 71X energy efficiency gain compared with IBM TrueNorth processor under the same test accuracy. It achieves at least 31X energy efficiency gain compared with the reference FPGA-based work.
Yanzhi Wang 0001, Caiwen Ding, Zhe Li 0001, Geng Yuan, Siyu Liao, Bo Yuan 0001, Xuehai Qian, Jian Tang 0008, Qinru Qiu, Xue Lin 0001
AAAI3
2018 Thoracic Disease Identification and Localization With Limited Supervision
abstract
Accurate identification and localization of abnormalities from radiology images play an integral part in clinical diagnosis and treatment planning. Building a highly accurate prediction model for these tasks usually requires a large number of images manually annotated with labels and finding sites of abnormalities. In reality, however, such annotated data are expensive to acquire, especially the ones with location annotations. We need methods that can work well with only a small amount of location annotations. To address this challenge, we present a unified approach that simultaneously performs disease identification and localization through the same underlying model for all images. We demonstrate that our approach can effectively leverage both class information as well as limited location annotation, and significantly outperforms the comparative reference baseline in both classification and localization tasks.
Zhe Li 0001, Chong Wang 0002, Wei Wei 0019, Li-Jia Li 0001, Li Fei-Fei 0001
CVPR1
2018 C-LSTM: Enabling Efficient LSTM using Structured Compression Techniques on FPGAs
abstract
Recently, significant accuracy improvement has been achieved for acoustic recognition systems by increasing the model size of Long Short-Term Memory (LSTM) networks. Unfortunately, the ever-increasing size of LSTM model leads to inefficient designs on FPGAs due to the limited on-chip resources. The previous work proposes to use a pruning based compression technique to reduce the model size and thus speedups the inference on FPGAs. However, the random nature of the pruning technique transforms the dense matrices of the model to highly unstructured sparse ones, which leads to unbalanced computation and irregular memory accesses and thus hurts the overall performance and energy efficiency.
Shuo Wang 0009, Zhe Li 0001, Caiwen Ding, Bo Yuan 0001, Qinru Qiu, Yanzhi Wang 0001, Yun Liang 0001
FPGA2
2018 Learning Topics Using Semantic Locality
abstract
The topic modeling discovers the latent topic probability of the given text documents. To generate the more meaningful topic that better represents the given document, we proposed a new feature extraction technique which can be used in the data preprocessing stage. The method consists of three steps. First, it generates the word/word-pair from every single document. Second, it applies a two-way TF-IDF algorithm to word/word-pair for semantic filtering. Third, it uses the K-means algorithm to merge the word pairs that have the similar semantic meaning. Experiments are carried out on the Open Movie Database (OMDb), Reuters Dataset and 20NewsGroup Dataset. The mean Average Precision score is used as the evaluation metric. Comparing our results with other state-of-the-art topic models, such as Latent Dirichlet allocation and traditional Restricted Boltzmann Machines. Our proposed data preprocessing can improve the generated topic accuracy by up to 12.99 %.
Krittaphat Pugdeethosapol, Sheng Lin 0001, Zhe Li 0001, Caiwen Ding, Yanzhi Wang 0001, Qinru Qiu
ICPR4
2017 Towards acceleration of deep convolutional neural networks using stochastic computing
abstract
In recent years, Deep Convolutional Neural Network (DCNN) has become the dominant approach for almost all recognition and detection tasks and outperformed humans on certain tasks. Nevertheless, the high power consumptions and complex topologies have hindered the widespread deployment of DCNNs, particularly in wearable devices and embedded systems with limited area and power budget. This paper presents a fully parallel and scalable hardware-based DCNN design using Stochastic Computing (SC), which leverages the energy-accuracy trade-off through optimizing SC components in different layers. We first conduct a detailed investigation of the Approximate Parallel Counter (APC) based neuron and multiplexer-based neuron using SC, and analyze the impacts of various design parameters, such as bit stream length and input number, on the energy/power/area/accuracy of the neuron cell. Then, from an architecture perspective, the influence of inaccuracy of neurons in different layers on the overall DCNN accuracy (i.e., software accuracy of the entire DCNN) is studied. Accordingly, a structure optimization method is proposed for a general DCNN architecture, in which neurons in different layers are implemented with optimized SC components, so as to reduce the area, power, and energy of the DCNN while maintaining the overall network performance in terms of accuracy. Experimental results show that the proposed approach can find a satisfactory DCNN configuration, which achieves 55X, 151X, and 2X improvement in terms of area, power and energy, respectively, while the error is increased by 2.86%, compared with the conventional binary ASIC implementation.
Ji Li 0006, Ao Ren, Zhe Li 0001, Caiwen Ding, Bo Yuan 0001, Qinru Qiu, Yanzhi Wang 0001
ASP-DAC3
2017 SC-DCNN: Highly-Scalable Deep Convolutional Neural Network using Stochastic Computing
abstract
With the recent advance of wearable devices and Internet of Things (IoTs), it becomes attractive to implement the Deep Convolutional Neural Networks (DCNNs) in embedded and portable systems. Currently, executing the software-based DCNNs requires high-performance servers, restricting the widespread deployment on embedded and mobile IoT devices. To overcome this obstacle, considerable research efforts have been made to develop highly-parallel and specialized DCNN accelerators using GPGPUs, FPGAs or ASICs.
Ao Ren, Zhe Li 0001, Caiwen Ding, Qinru Qiu, Yanzhi Wang 0001, Ji Li 0006, Xuehai Qian, Bo Yuan 0001
ASPLOS2
2017 Structural design optimization for deep convolutional neural networks using stochastic computing
abstract
Deep Convolutional Neural Networks (DCNNs) have been demonstrated as effective models for understanding image content. The computation behind DCNNs highly relies on the capability of hardware resources due to the deep structure. DCNNs have been implemented on different large-scale computing platforms. However, there is a trend that DCNNs have been embedded into light-weight local systems, which requires low power/energy consumptions and small hardware footprints. Stochastic Computing (SC) radically simplifies the hardware implementation of arithmetic units and has the potential to satisfy the small low-power needs of DCNNs. Local connectivities and down-sampling operations have made DCNNs more complex to be implemented using SC. In this paper, eight feature extraction designs for DCNNs using SC in two groups are explored and optimized in detail from the perspective of calculation precision, where we permute two SC implementations for inner-product calculation, two down-sampling schemes, and two structures of DCNN neurons. We evaluate the network in aspects of network accuracy and hardware performance for each DCNN using one feature extraction design out of eight. Through exploration and optimization, the accuracies of SC-based DCNNs are guaranteed compared with software implementations on CPU/GPU/binary-based ASIC synthesis, while area, power, and energy are significantly reduced by up to 776x, 190x, and 32835x.
Zhe Li 0001, Ao Ren, Ji Li 0006, Qinru Qiu, Bo Yuan 0001, Jeffrey T. Draper, Yanzhi Wang 0001
DATE1
2017 Softmax Regression Design for Stochastic Computing Based Deep Convolutional Neural Networks
abstract
Recently, Deep Convolutional Neural Networks (DCNNs) have made tremendous advances, achieving close to or even better accuracy than human-level perception in various tasks. Stochastic Computing (SC), as an alternate to the conventional binary computing paradigm, has the potential to enable massively parallel and highly scalable hardware implementations of DCNNs. In this paper, we design and optimize the SC based Softmax Regression function. Experiment results show that compared with a binary SR, the proposed SC-SR under longer bit stream can reach the same level of accuracy with the improvement of 295X, 62X, 2617X in terms of power, area and energy, respectively. Binary SR is suggested for future DCNNs with short bit stream length input whereas SC-SR is recommended for longer bit stream.
Ji Li 0006, Zhe Li 0001, Caiwen Ding, Ao Ren, Bo Yuan 0001, Qinru Qiu, Jeffrey T. Draper, Yanzhi Wang 0001
ACM Great Lakes Symposium on VLSI3
2017 Energy-efficient, high-performance, highly-compressed deep neural network design using block-circulant matrices
abstract
Deep neural networks (DNNs) have emerged as the most powerful machine learning technique in numerous artificial intelligent applications. However, the large sizes of DNNs make themselves both computation and memory intensive, thereby limiting the hardware performance of dedicated DNN accelerators. In this paper, we propose a holistic framework for energy-efficient high-performance highly-compressed DNN hardware design. First, we propose block-circulant matrix-based DNN training and inference schemes, which theoretically guarantee Big-O complexity reduction in both computational cost (from O(n2) to O(n log n)) and storage requirement (from O(n2) to O(n)) of DNNs. Second, we dedicatedly optimize the hardware architecture, especially on the key fast Fourier transform (FFT) module, to improve the overall performance in terms of energy efficiency, computation performance and resource cost. Third, we propose a design flow to perform hardware-software co-optimization with the purpose of achieving good balance between test accuracy and hardware performance of DNNs. Based on the proposed design flow, two block-circulant matrix-based DNNs on two different datasets are implemented and evaluated on FPGA. The fixed-point quantization and the proposed block-circulant matrix-based inference scheme enables the network to achieve as high as 3.5 TOPS computation performance and 3.69 TOPS/W energy efficiency while the memory is saved by 108X ~ 116X with negligible accuracy degradation.
Siyu Liao, Zhe Li 0001, Xue Lin 0001, Qinru Qiu, Yanzhi Wang 0001, Bo Yuan 0001
ICCAD2
2017 A Hierarchical Framework of Cloud Resource Allocation and Power Management Using Deep Reinforcement Learning
abstract
Automatic decision-making approaches, such as reinforcement learning (RL), have been applied to (partially) solve the resource allocation problem adaptively in the cloud computing system. However, a complete cloud resource allocation framework exhibits high dimensions in state and action spaces, which prohibit the usefulness of traditional RL techniques. In addition, high power consumption has become one of the critical concerns in design and control of cloud computing systems, which degrades system reliability and increases cooling cost. An effective dynamic power management (DPM) policy should minimize power consumption while maintaining performance degradation within an acceptable level. Thus, a joint virtual machine (VM) resource allocation and power management framework is critical to the overall cloud computing system. Moreover, novel solution framework is necessary to address the even higher dimensions in state and action spaces. In this paper, we propose a novel hierarchical framework for solving the overall resource allocation and power management problem in cloud computing systems. The proposed hierarchical framework comprises a global tier for VM resource allocation to the servers and a local tier for distributed power management of local servers. The emerging deep reinforcement learning (DRL) technique, which can deal with complicated control problems with large state space, is adopted to solve the global tier problem. Furthermore, an autoencoder and a novel weight sharing structure are adopted to handle the high-dimensional state space and accelerate the convergence speed. On the other hand, the local tier of distributed server power managements comprises an LSTM based workload predictor and a model-free RL based power manager, operating in a distributed manner. Experiment results using actual Google cluster traces show that our proposed hierarchical framework significantly saves the power consumption and energy usage than the baseline while achieving no severe latency degradation. Meanwhile, the proposed framework can achieve the best trade-off between latency and power/energy consumption in a server cluster.
Ning Liu 0007, Zhe Li 0001, Jielong Xu, Sheng Lin 0001, Qinru Qiu, Jian Tang 0008, Yanzhi Wang 0001
ICDCS2
2017 Theoretical Properties for Neural Networks with Weight Matrices of Low Displacement Rank
abstract
Recently low displacement rank (LDR) matrices, or so-called structured matrices, have been proposed to compress large-scale neural networks. Empirical results have shown that neural networks with weight matrices of LDR matrices, referred as LDR neural networks, can achieve significant reduction in space and computational complexity while retaining high accuracy. This paper gives theoretical study on LDR neural networks. First, we prove the universal approximation property of LDR neural networks with a mild condition on the displacement operators. We then show that the error bounds of LDR neural networks are as efficient as general neural networks with both single-layer and multiple-layer structure. Finally, we propose back-propagation based training algorithm for general LDR neural networks.
Liang Zhao 0002, Siyu Liao, Yanzhi Wang 0001, Zhe Li 0001, Jian Tang 0008, Bo Yuan 0001
ICML4
2017 Hardware-driven nonlinear activation for stochastic computing based deep convolutional neural networks
abstract
Recently, Deep Convolutional Neural Networks (DCNNs) have made unprecedented progress, achieving the accuracy close to, or even better than human-level perception in various tasks. There is a timely need to map the latest software DCNNs to application-specific hardware, in order to achieve orders of magnitude improvement in performance, energy efficiency and compactness. Stochastic Computing (SC), as a low-cost alternative to the conventional binary computing paradigm, has the potential to enable massively parallel and highly scalable hardware implementation of DCNNs. One major challenge in SC based DCNNs is designing accurate nonlinear activation functions, which have a significant impact on the network-level accuracy but cannot be implemented accurately by existing SC computing blocks. In this paper, we design and optimize SC based neurons, and we propose highly accurate activation designs for the three most frequently used activation functions in software DCNNs, i.e, hyperbolic tangent, logistic, and rectified linear units. Experimental results on LeNet-5 using MNIST dataset demonstrate that compared with a binary ASIC hardware DCNN, the DCNN with the proposed SC neurons can achieve up to 61X, 151X, and 2X improvement in terms of area, power, and energy, respectively, at the cost of small precision degradation. In addition, the SC approach achieves up to 21X and 41X of the area, 41X and 72X of the power, and 198200X and 96443X of the energy, compared with CPU and GPU approaches, respectively, while the error is increased by less than 3.07%. ReLU activation is suggested for future SC based DCNNs considering its superior performance under a small bit stream length.
Ji Li 0006, Zhe Li 0001, Caiwen Ding, Ao Ren, Qinru Qiu, Jeffrey T. Draper, Yanzhi Wang 0001
IJCNN3
2017 CirCNN: accelerating and compressing deep neural networks using block-circulant weight matrices
abstract
Large-scale deep neural networks (DNNs) are both compute and memory intensive. As the size of DNNs continues to grow, it is critical to improve the energy efficiency and performance while maintaining accuracy. For DNNs, the model size is an important factor affecting performance, scalability and energy efficiency. Weight pruning achieves good compression ratios but suffers from three drawbacks: 1) the irregular network structure after pruning, which affects performance and throughput; 2) the increased training complexity; and 3) the lack of rigirous guarantee of compression ratio and inference accuracy.
Caiwen Ding, Siyu Liao, Yanzhi Wang 0001, Zhe Li 0001, Ning Liu 0007, Youwei Zhuo, Chao Wang 0051, Xuehai Qian, Yu Bai 0004, Geng Yuan, Jian Tang 0008, Qinru Qiu, Xue Lin 0001, Bo Yuan 0001
MICRO4
2016 DSCNN: Hardware-oriented optimization for Stochastic Computing based Deep Convolutional Neural Networks
abstract
Deep Convolutional Neural Networks (DCNN), a branch of Deep Neural Networks which use the deep graph with multiple processing layers, enables the convolutional model to finely abstract the high-level features behind an image. Large-scale applications using DCNN mainly operate in high-performance server clusters, GPUs or FPGA clusters; it is restricted to extend the applications onto mobile/wearable devices and Internet-of-Things (IoT) entities due to high power/energy consumption. Stochastic Computing is a promising method to overcome this shortcoming used in specific hardware-based systems. Many complex arithmetic operations can be implemented with very simple hardware logic in the SC framework, which alleviates the extensive computation complexity. The exploration of network-wise optimization and the revision of network structure with respect to stochastic computing based hardware design have not been discussed in previous work. In this paper, we investigate Deep Stochastic Convolutional Neural Network (DSCNN) for DCNN using stochastic computing. The essential calculation components using SC are designed and evaluated. We propose a joint optimization method to collaborate components guaranteeing a high calculation accuracy in each stage of the network. The structure of original DSCNN is revised to accommodate SC hardware design's simplicity. Experimental Results show that as opposed to software inspired feature extraction block in DSCNN, an optimized hardware oriented feature extraction block achieves as higher as 59.27% calculation precision. And the optimized DSCNN can achieve only 3.48% network test error rate compared to 27.83% for baseline DSCNN using software inspired feature extraction block.
Zhe Li 0001, Ao Ren, Ji Li 0006, Qinru Qiu, Yanzhi Wang 0001, Bo Yuan 0001
ICCD1