Lixue Xia

dblp:160/1534 · DBLP profile ↗
← Back
35ranked-venue papers
8as first author
2since 2021 · last 2023
0000-0002-7731-7028ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 34 · 7 first-author · 2 since 2021Software engineering, systems software and programming languages · 5 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
14 papers
Memory systems · 31% Hardware accelerators and domain-specific architectures · 31% Parallel and multicore computing · 9%
Artificial intelligence
5 papers
Efficient and distributed learning · 95% Deep learning architectures and training · 5%

Topics — the 30 heaviest of 36, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
machine learning accelerator
2.572023
MNSIM 2.0: A Behavior-Level Modeling Tool for Processing-In-Memory Architectures · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2023
Low Bit-Width Convolutional Neural Network on RRAM · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
TIME: A Training-in-Memory Architecture for RRAM-Based Deep Neural Networks · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019
Memory systems
non-volatile memory
1.132020
Low Bit-Width Convolutional Neural Network on RRAM · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
TIME: A Training-in-Memory Architecture for RRAM-Based Deep Neural Networks · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019
TIME: A Training-in-memory Architecture for Memristor-based Deep Neural Networks · DAC 2017
Memory systems
processing-in-memory
1.022023
MNSIM 2.0: A Behavior-Level Modeling Tool for Processing-In-Memory Architectures · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2023
A Configurable Multi-Precision CNN Computing Framework Based on Single Bit RRAM · DAC 2019
Emerging computing paradigms
neuromorphic computing
1.042019
MNSIM: Simulation Platform for Memristor-Based Neuromorphic Computing System · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
Fault-Tolerant Training with On-Line Fault Detection for RRAM-Based Neural Computing Systems · DAC 2017
Switched by input: power efficient structure for RRAM-based convolutional neural network · DAC 2016
Memory systems › non-volatile memory
resistive memory
0.822020
Low Bit-Width Convolutional Neural Network on RRAM · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
TIME: A Training-in-Memory Architecture for RRAM-Based Deep Neural Networks · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019
Memory systems
in-memory computing
0.822020
Long Live TIME: Improving Lifetime and Security for NVM-Based Training-in-Memory Systems · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Long live TIME: improving lifetime for training-in-memory engines by structured gradient sparsification · DAC 2018
Hardware accelerators and domain-specific architectures › machine learning accelerator › CNN accelerator
RRAM-based CNN accelerator
0.722020
Low Bit-Width Convolutional Neural Network on RRAM · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Switched by input: power efficient structure for RRAM-based convolutional neural network · DAC 2016
Performance modeling and evaluation › system modeling
architecture modeling
0.712023
MNSIM 2.0: A Behavior-Level Modeling Tool for Processing-In-Memory Architectures · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2023
Memory systems › emerging memory technologies
RRAM crossbar
0.622019
TIME: A Training-in-Memory Architecture for RRAM-Based Deep Neural Networks · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019
Switched by input: power efficient structure for RRAM-based convolutional neural network · DAC 2016
Machine learning › Efficient and distributed learning
model compression
0.522020
Low Bit-Width Convolutional Neural Network on RRAM · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Long live TIME: improving lifetime for training-in-memory engines by structured gradient sparsification · DAC 2018
Machine learning › Efficient and distributed learning
distributed training
0.512021
DAPPLE: a pipelined data parallel approach for training large models · PPoPP 2021
Machine learning › Efficient and distributed learning › distributed training
hybrid parallel training
0.512021
DAPPLE: a pipelined data parallel approach for training large models · PPoPP 2021
Parallel and multicore computing
data parallelism
0.512021
DAPPLE: a pipelined data parallel approach for training large models · PPoPP 2021
Parallel and multicore computing › parallel computing › parallel machine learning
parallel training
0.512021
DAPPLE: a pipelined data parallel approach for training large models · PPoPP 2021
Parallel and multicore computing
pipeline parallelism
0.512021
DAPPLE: a pipelined data parallel approach for training large models · PPoPP 2021
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator
0.422019
A Configurable Multi-Precision CNN Computing Framework Based on Single Bit RRAM · DAC 2019
Merging the interface: power, area and accuracy co-optimization for RRAM crossbar-based mixed-signal computing system · DAC 2015
Machine learning › Efficient and distributed learning › model compression › quantization
low-bit quantization
0.412020
Low Bit-Width Convolutional Neural Network on RRAM · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Hardware security and side channels › hardware attacks
side-channel and fault attacks
0.412020
Long Live TIME: Improving Lifetime and Security for NVM-Based Training-in-Memory Systems · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Hardware accelerators and domain-specific architectures › machine learning accelerator › DNN inference
CNN inference
0.412019
PAI-FCNN: FPGA Based CNN Inference System · FPGA 2019
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN training
0.412019
TIME: A Training-in-Memory Architecture for RRAM-Based Deep Neural Networks · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2019
Reconfigurable computing and FPGAs › FPGA accelerator
FPGA-based CNN inference
0.412019
PAI-FCNN: FPGA Based CNN Inference System · FPGA 2019
Memory systems › processing-in-memory
ReRAM-based processing-in-memory
0.412019
A Configurable Multi-Precision CNN Computing Framework Based on Single Bit RRAM · DAC 2019
Performance modeling and evaluation › simulation › architectural simulation
accelerator simulation
0.312018
MNSIM: Simulation Platform for Memristor-Based Neuromorphic Computing System · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
Distributed systems › fault tolerance › fault-tolerant distributed systems
fault-tolerant training
0.312017
Fault-Tolerant Training with On-Line Fault Detection for RRAM-Based Neural Computing Systems · DAC 2017
Memory systems › in-memory computing
in-memory computing for neural networks
0.312017
TIME: A Training-in-memory Architecture for Memristor-based Deep Neural Networks · DAC 2017
Emerging computing paradigms
memristive computing
0.312017
TIME: A Training-in-memory Architecture for Memristor-based Deep Neural Networks · DAC 2017
Integrated circuit design › analog and mixed-signal circuits
mixed-signal circuit design
0.212015
Merging the interface: power, area and accuracy co-optimization for RRAM crossbar-based mixed-signal computing system · DAC 2015
Machine learning › Deep learning architectures and training
convolutional neural network
0.112019
A Configurable Multi-Precision CNN Computing Framework Based on Single Bit RRAM · DAC 2019
Machine learning › Efficient and distributed learning › distributed training › gradient compression
gradient sparsification
0.112018
Long live TIME: improving lifetime for training-in-memory engines by structured gradient sparsification · DAC 2018
Electronic design automation
design space exploration
0.112018
MNSIM: Simulation Platform for Memristor-Based Neuromorphic Computing System · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018

Methods — techniques the papers use, named apart from their topics

pipeline parallelism · 1.0parallelization strategy planning · 1.0data parallelism · 1.0threshold training · 0.7remapping · 0.7quiescent-voltage comparison · 0.7deep reinforcement learning · 0.7backpropagation · 0.7neural network quantization · 0.7behavior-level modeling · 0.7row swapping · 0.4refresh · 0.4random swapping · 0.4quantization-aware training · 0.4pipelined line buffer · 0.4matrix splitting · 0.4gradient sparsification · 0.4network quantization · 0.4
YearPublicationVenuePosition
2023 MNSIM 2.0: A Behavior-Level Modeling Tool for Processing-In-Memory Architectures
abstract
In the age of Artificial Intelligence (AI), the huge data movements between memory and computing units become the bottleneck of von Neumann architectures, i.e., the “memory wall” problem. In order to tackle this challenge, Processing-In-Memory (PIM) architectures are proposed, which perform in-situ computations in memory and give alternative solutions to boost the computing energy efficiency and performance. Because of the large-scale Neural Network (NN) algorithm models and the huge hardware design space, various factors affect computing accuracy and performance, bringing the need for efficient PIM modeling and evaluation tools. In this work, we propose a behavior-level modeling tool, MNSIM 2.0, to model the performance of PIM architectures efficiently. At the hardware level, MNSIM 2.0 provides a hierarchical PIM modeling structure with flexible architecture configurability and components extensibility. Moreover, the first unified PIM memory array model is proposed for describing both digital and analog PIM. At the algorithm level, MNSIM 2.0 supports the PIM-based NN computing accuracy simulation considering various architecture and device parameters. A PIM-oriented NN model training and quantization flow is also integrated to improve the performance gain brought by PIM. At the scheduling level, MNSIM 2.0 adopts a universal scheduling description compatible with different scheduling strategies. Validation using fabricated PIM macros shows the relative modeling error rate of MNSIM 2.0 is 3:8 5:5%. Case studies show that MNSIM 2.0 enables PIM design space explorations, influences analysis of device parameters, and architecture design insight discoveries.
Zhenhua Zhu 0002, Hanbo Sun, Tongxin Xie, Guohao Dai 0001, Lixue Xia, Dimin Niu, Xiaoming Chen 0003, Xiaobo Sharon Hu, Yu Cao 0001, Yuan Xie 0001, Huazhong Yang, Yu Wang 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2021 DAPPLE: a pipelined data parallel approach for training large models
abstract
It is a challenging task to train large DNN models on sophisticated GPU platforms with diversified interconnect capabilities. Recently, pipelined training has been proposed as an effective approach for improving device utilization. However, there are still several tricky issues to address: improving computing efficiency while ensuring convergence, and reducing memory usage without incurring additional computing costs. We propose DAPPLE, a synchronous training framework which combines data parallelism and pipeline parallelism for large DNN models. It features a novel parallelization strategy planner to solve the partition and placement problems, and explores the optimal hybrid strategies of data and pipeline parallelism. We also propose a new runtime scheduling algorithm to reduce device memory usage, which is orthogonal to re-computation approach and does not come at the expense of training throughput. Experiments show that DAPPLE planner consistently outperforms strategies generated by PipeDream's planner by up to 3.23× speedup under synchronous training scenarios, and DAPPLE runtime outperforms GPipe by 1.6× speedup of training throughput and saves 12% of memory consumption at the same time.
Shiqing Fan, Zongyan Cao, Siyu Wang 0006, Zhen Zheng, Chuan Wu 0001, Guoping Long, Jun Yang 0052, Lixue Xia, Lansong Diao, Wei Lin 0016
PPoPP10
2020 MNSIM 2.0: A Behavior-Level Modeling Tool for Memristor-based Neuromorphic Computing Systems
abstract
Memristor based neuromorphic computing systems give alternative solutions to boost the computing energy efficiency of Neural Network (NN) algorithms. Because of the large-scale applications and the large architecture design space, many factors will affect the computing accuracy and system's performance. In this work, we propose a behavior-level modeling tool for memristor-based neuromorphic computing systems, MNSIM 2.0, to model the performance and help researchers to realize an early-stage design space exploration. Compared with the former version and other benchmarks, MNSIM 2.0 has the following new features: 1. In the algorithm level, MNSIM 2.0 supports the inference accuracy simulation for mixed-precision NNs considering non-ideal factors. 2. In the architecture level, a hierarchical modeling structure for PIM systems is proposed. Users can customize their designs from the aspects of devices, interfaces, processing units, buffer designs, and interconnections. 3. Two hardware-aware algorithm optimization methods are integrated in MNSIM 2.0 to realize software-hardware co-optimization.
Zhenhua Zhu 0002, Hanbo Sun, Kaizhong Qiu, Lixue Xia, Guohao Dai 0001, Dimin Niu, Xiaoming Chen 0003, Xiaobo Sharon Hu, Yu Cao 0001, Yuan Xie 0001, Yu Wang 0002, Huazhong Yang
ACM Great Lakes Symposium on VLSI4
2020 Long Live TIME: Improving Lifetime and Security for NVM-Based Training-in-Memory Systems
abstract
Nonvolatile memory (NVM)-based training-in-memory (TIME) systems have emerged that can process the neural network (NN) training in an energy-efficient manner. However, the endurance of NVM cells is disappointing, rendering concerns about the lifetime of TIME systems, because the weights of NN models always need to be updated for thousands to millions of times during training. Gradient sparsification (GS) can alleviate this problem by preserving only a small portion of the gradients to update the weights. However, conventional GS will introduce nonuniform writes on different cells across the whole NVM crossbars, which significantly reduces the excepted available lifetime. Moreover, an adversary can easily launch malicious training tasks to exactly wear-out the target cells and fast break down the system. In this article, we propose an efficient and effective framework, referred as SGS-ARS, to improve the lifetime and security of TIME systems. The framework mainly contains a structured GS (SGS) scheme for reducing the write frequency, and an aging-aware row swapping (ARS) scheme to make the writes uniform. Meanwhile, we show that the back-propagation mechanism allows the attacker to localize and update fixed memory locations and wear them out. Therefore, we introduce Random-ARS and Refresh techniques to thwart adversarial training attacks, preventing the systems from being fast broken in an extremely short time. Our experiments show that when TIME is programmed to train ResNet-50 on ImageNet dataset, $356\times $ lifetime extension can be achieved without sacrificing the accuracy much or incurring much hardware overhead. Under the adversarial environment, the available lifetime of TIME systems can still be improved by $84\times $ .
Yi Cai 0003, Yujun Lin 0001, Lixue Xia, Xiaoming Chen 0003, Song Han 0003, Yu Wang 0002, Huazhong Yang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2020 Low Bit-Width Convolutional Neural Network on RRAM
abstract
The emerging resistive random-access memory (RRAM) has been widely applied in accelerating the computing of deep neural networks. However, it is challenging to achieve highprecision computations based on RRAM due to the limits of the resistance level and the interfaces. Low bit-width convolutional neural networks (CNNs) provide promising solutions to introduce low bit-width RRAM devices and low bit-width interfaces in RRAM-based computing system (RCS). While open questions still remain regarding: 1) how to make matrix splitting when a single crossbar is not large enough to hold all parameters of one weight matrix; 2) how to design a pipeline to accelerate the inference based on line buffer structure; and 3) how to reduce the accuracy drop due to the parameter splitting and data quantization. In this paper, we propose an RRAM crossbar-based low bit-width CNN (LB-CNN) accelerator. We make detailed discussion on the system design, including the matrix splitting strategies to enhance the scalability, and the pipelined implementation based on line buffers to accelerate the inference. In addition, we propose a splitting and quantizing while training method to incorporate the actual hardware constraints with the training. In our experiments, low bit-width LeNet-5 on RRAM show much better robustness than multibit models with device variation. The pipeline strategy achieves approximately 6.0× speedup to process each image on ResNet-18. For low-bit VGG-8 on CIFAR-10, the proposed accelerator saves 54.9% of the energy consumption and 48.3% of the area compared with the multibit VGG-8 structure.
Yi Cai 0003, Tianqi Tang 0001, Lixue Xia, Boxun Li, Yu Wang 0002, Huazhong Yang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2020 Algorithmic Fault Detection for RRAM-based Matrix Operations
abstract
An RRAM-based computing system (RCS) provides an energy-efficient hardware implementation of vector-matrix multiplication for machine-learning hardware. However, it is vulnerable to faults due to the immature RRAM fabrication process. We propose an efficient fault tolerance method for RCS; the proposed method, referred to as extended-ABFT (X-ABFT), is inspired by algorithm-based fault tolerance (ABFT). We utilize row checksums and test-input vectors to extract signatures for fault detection and error correction. We present a solution to alleviate the overflow problem caused by the limited number of voltage levels for the test-input signals. Simulation results show that for a Hopfield classifier with faults in 5% of its RRAM cells, X-ABFT allows us to achieve nearly the same classification accuracy as in the fault-free case.
Lixue Xia, Yu Wang 0002, Krishnendu Chakrabarty
ACM Trans. Design Autom. Electr. Syst.2
2019 PAI-FCNN: FPGA Based Inference System for Complex CNN Models
abstract
Convolutional Neural Network (CNN) models are becoming complex with advanced OPs and structures, which introduces design challenges for FPGA-based system. In this paper, we present the design of an FPGA-based CNN inference system, PAI-FCNN, to support modern complex CNN models. PAI-FCNN consists of scalable hardware design and a model reconstruction flow in software compiler. In this way, advanced OPs like Deconv, Conv with upsampling, Dilated Conv, Concatenation can be processed by PAI-FCNN with high performance and hardware efficiency. PAI-FCNN also incorporates reduced precision to boost computing capacity, and the emerging CNN-RNN (Recurrent Neural Network) hybrid models are supported. Our experiments on both PC and embedded FPGA platforms show that the system consistently performs in an efficient manner. PAI-FCNN achieves better throughput and power efficiency than GPU solutions.
Lixue Xia, Lansong Diao, Zhao Jiang, Hao Liang 0003, Kai Chen 0008, Shunli Dou, Zibin Su, Jiansong Zhang 0001, Wei Lin 0016
ASAP1
2019 Fault tolerance in neuromorphic computing systems
abstract
Resistive Random Access Memory (RRAM) and RRAM-based computing systems (RCS) provide energy-efficient technology options for neuromorphic computing. However, the applicability of RCS is limited by reliability problems that arise from the immature fabrication process. In order to take advantage of RCS in practical applications, fault-tolerant design is a key challenge. We present a survey of fault-tolerant designs for RRAM-based neuromorphic computing systems. We first describe RRAM-based crossbars and training architectures in RCS. Following this, we classify fault models into different categories, and review post-fabrication testing methods. Subsequently, online testing methods are presented. Finally, we present various fault-tolerant techniques that were designed to tolerate different types of RRAM faults. The methods reviewed in this survey represent recent trends in fault-tolerant designs of RCS, and are expected motivate further research in this field.
Lixue Xia, Yu Wang 0002, Krishnendu Chakrabarty
ASP-DAC2
2019 A Configurable Multi-Precision CNN Computing Framework Based on Single Bit RRAM
abstract
Convolutional Neural Networks (CNNs) play a vital role in machine learning. Emerging resistive random-access memories (RRAMs) and RRAM-based Processing-In-Memory architectures have demonstrated great potentials in boosting both the performance and energy efficiency of CNNs. However, restricted by the immature process technology, it is hard to implement and fabricate a CNN accelerator chip based on multi-bit RRAM devices. In addition, existing single bit RRAM based CNN accelerators only focus on binary or ternary CNNs which have more than 10% accuracy loss compared with full precision CNNs. This paper proposes a configurable multi-precision CNN computing framework based on single bit RRAM, which consists of an RRAM computing overhead aware network quantization algorithm and a configurable multi-precision CNN computing architecture based on single bit RRAM. The proposed method can achieve equivalent accuracy as full precision CNN but also with lower storage consumption and latency via multiple precision quantization. The designed architecture supports for accelerating the multi-precision CNNs even with various precision among different layers. Experiment results show that the proposed framework can reduce 70% computing area and 75% computing energy on average, with nearly no accuracy loss. And the equivalent energy efficiency is 1.6 ~ 8.6× compared with existing RRAM based architectures with only 1.07% area overhead.
Zhenhua Zhu 0002, Hanbo Sun, Yujun Lin 0001, Guohao Dai 0001, Lixue Xia, Song Han 0003, Yu Wang 0002, Huazhong Yang
DAC5
2019 PAI-FCNN: FPGA Based CNN Inference System
abstract
We describe the FPGA subsystem of the Platform of Artificial Intelligence (PAI) in Alibaba Group, called PAI-FCNN. PAI-FCNN plays the role of a heterogeneous back-end for CNN inference, together with other CPU, GPU and ASIC subsystems in PAI. Driven by various business needs, we built PAI-FCNN from scratch since two years ago. We present our experience from FPGA/compiler design and implementation, to system evaluation and deployment. In particular, in order to address three practical challenges: (1) Efficient processing for diverse operators and model structure such as Deconv, Dilated Conv, Up-sampling, PReLu and Concatenation. (2) Serving multiple highly-different models on single FPGA hardware. (3) Competitive performance with alternative GPU or ASIC solutions, we extensively perform joint software & hardware design to optimize system efficiency across multiple CNN models, which includes model reconstruction in compiler software and flexible data access in data-flow CNN processor. We also incorporate reduced precision and model retraining to boost system capacity. Using U-net as an example, on Xilinx KU115 chip, with the help of 74.9% efficiency on Int16-precision hardware (with 3.226TOPS capacity) and 72.9% efficiency on mixed-int8/int3-precision hardware (with 14.746TOPS capacity), we achieve slightly better throughput and 2X higher power efficiency than P4.
Lansong Diao, Zhao Jiang, Hao Liang 0003, Chang'an Ye, Kai Chen 0008, Shunli Dou, Lixue Xia, Jiansong Zhang 0001, Wei Lin 0016
FPGA9
2019 Ouroboros: An Inference Engine for Deep Learning Based TTS on Embedded Devices
abstract
This article consists of a collection of slides from the author's conference presentation.
Jiansong Zhang 0001, Lixue Xia, Zhao Jiang, Hao Liang 0003, Shouda Liu, Wei Lin 0016, Yuan Xie 0001
Hot Chips Symposium2
2019 TIME: A Training-in-Memory Architecture for RRAM-Based Deep Neural Networks
abstract
The training of neural networks (NN) is usually time-consuming and resource intensive. The emerging metaloxide resistive random-access memory (RRAM) device has shown potential for the computation of NN. RRAM crossbar structure and multibit characteristics can perform the matrix-vector product in high energy efficiency, which is the most common operation of NN. Two challenges exist for realizing training NN based on RRAM. First, the current architectures based on RRAM only support the inference in training NN and cannot perform the backpropagation (BP) and the weight update of training NN. Second, training NN requires enormous iterations to constantly update the weights for reaching the convergence. However, this weight update leads to large energy consumption because of the nonideal factors of RRAM. In this paper, we propose a training-in-memory based on RRAM (TIME) architecture and the peripheral circuit design to enable training NN on RRAM. TIME supports the BP and the weight update while maximizing the re-usage of peripheral circuits of the inference operation on RRAM. Meanwhile, a set of optimization strategies focusing on the nonideal factors are designed to reduce the cost of tuning RRAM. We explore the performance of both supervised learning (SL) and deep reinforcement learning (DRL) on TIME. A specific mapping method of DRL is also introduced to further improve energy efficiency. Simulation results show that in SL, TIME can achieve 5.3× higher energy efficiency on average compared with DaDianNao, an application-specific integrated circuits (ASIC) in CMOS technology. In DRL, TIME can perform an average 126× higher than GPU in energy efficiency. If the cost of tuning RRAM can be further reduced, TIME has the potential to boost the energy efficiency by two orders of magnitudes compared with ASIC.
Lixue Xia, Zhenhua Zhu 0002, Yi Cai 0003, Yuan Xie 0001, Yu Wang 0002, Huazhong Yang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2019 Fault-Tolerant Training Enabled by On-Line Fault Detection for RRAM-Based Neural Computing Systems
abstract
An resistive random-access memory (RRAM)-based computing system (RCS) is an attractive hardware platform for implementing neural computing algorithms. On-line training for RCS enables hardware-based learning for a given application and reduces the additional error caused by device parameter variations. However, a high occurrence rate of hard faults due to immature fabrication processes and limited write endurance restrict the applicability of on-line training for RCS. We propose a fault-tolerant on-line training method that alternates between a fault-detection phase and a fault-tolerant training phase. In the fault-detection phase, a quiescent-voltage comparison method is utilized. In the training phase, a threshold-training method and a remapping scheme is proposed. Our results show that, compared to neural computing without fault tolerance, the recognition accuracy for the Cifar-10 dataset improves from 37% to 83% when using low-endurance RRAM cells, and from 63% to 76% when using RRAM cells with high endurance but a high percentage of initial faults.
Lixue Xia, Xuefei Ning, Krishnendu Chakrabarty, Yu Wang 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2018 Training low bitwidth convolutional neural network on RRAM
abstract
Convolutional Neural Networks (CNNs) have achieved excellent performance on various artificial intelligence (AI) applications, while a higher demand on energy efficiency is required for future AI. Resistive Random-Access Memory (RRAM)-based computing system provides a promising solution to energy-efficient neural network training. However, it's difficult to support high-precision CNN in RRAM-based hardware systems. Firstly, multi-bit digital-analog interfaces will take up most energy overhead of the whole system. Secondly, it's difficult to write the RRAM to expected resistance states accurately; only low-precision numbers can be represented. To enable CNN training based on RRAM, we propose a low-bitwidth CNN training method, using low-bitwidth convolution outputs (CO), activations (A), weights (W) and gradients (G) to train CNN models based on RRAM. Furthermore, we design a system to implement the training algorithms. We explore the accuracy under different bitwidth combinations of (A, CO, W, G), and propose a practical tradeoff between accuracy and energy overhead. Our experiments demonstrate that the proposed system perform well on low-bitwidth CNN training tasks. For example, training LeNet-5 with 4-bit convolution outputs, 4-bit weights, 4-bit activations and 4-bit gradients on MNIST can still achieve 97.67% accuracy. Moreover, the proposed system can achieve 23.0X higher energy efficiency than GPU when processing the training task of LeNet-5, and 4.4X higher energy efficiency when processing the training task of ResNet-20.
Yi Cai 0003, Tianqi Tang 0001, Lixue Xia, Zhenhua Zhu 0002, Yu Wang 0002, Huazhong Yang
ASP-DAC3
2018 Long live TIME: improving lifetime for training-in-memory engines by structured gradient sparsification
abstract
Deeper and larger Neural Networks (NNs) have made breakthroughs in many fields. While conventional CMOS-based computing platforms are hard to achieve higher energy efficiency. RRAM-based systems provide a promising solution to build efficient Training-In-Memory Engines (TIME). While the endurance of RRAM cells is limited, it's a severe issue as the weights of NN always need to be updated for thousands to millions of times during training. Gradient sparsification can address this problem by dropping off most of the smaller gradients but introduce unacceptable computation cost. We proposed an effective framework, SGS-ARS, including Structured Gradient Sparsification (SGS) and Aging-aware Row Swapping (ARS) scheme, to guarantee write balance across whole RRAM crossbars and prolong the lifetime of TIME. Our experiments demonstrate that 356× lifetime extension is achieved when TIME is programmed to train ResNet-50 on Imagenet dataset with our SGS-ARS framework.
Yi Cai 0003, Yujun Lin 0001, Lixue Xia, Xiaoming Chen 0003, Song Han 0003, Yu Wang 0002, Huazhong Yang
DAC3
2018 Rescuing memristor-based computing with non-linear resistance levels
abstract
Emerging memristor devices like metal oxide resistive switching random access memory (RRAM) and memristor crossbar have shown great potential in computing matrix-vector multiplication. However, due to the nonlinear distribution of resistance levels in memristor devices, the state-of-the-art multi-bit cell cannot accomplish the multi-bit computing task accurately. In this paper, we propose fault-tolerant schemes to rescue memristor-based computation with nonlinear resistance levels. We classify the resistance level distributions in memristor devices into three types, and the corresponding models are proposed to analyze the computation characteristics. We propose two theoretical conditions to determine if a memristor device can support multi-bit matrix computation. For the deviated linear model, the least squares method is used to reduce the computing error. When the resistance distribution obeys the proposed power model, a logarithmic operation circuit is used to decode the multiplication results and then accomplish the computing accurately. For the exponential model, since the device cannot complete typical matrix-vector multiplication from hardware level, we propose online and offline quantization methods to make the neural computing algorithms friendly to memristor device. Simulation results show that the root-mean-square error improves around 4% with the linear model and more than 99% with the power model. After quantization, the accuracy of ResNet-18 using memristor with exponential conductance levels can be improved to the same accuracy with ideal linear devices.
Jilan Lin, Lixue Xia, Zhenhua Zhu 0002, Hanbo Sun, Yi Cai 0003, Xiaoming Chen 0003, Yu Wang 0002, Huazhong Yang
DATE2
2018 A peripheral circuit reuse structure integrated with a retimed data flow for low power RRAM crossbar-based CNN
abstract
Convolutional computations implemented in RRAM crossbar-based Computing System (RCS) demonstrate the outstanding advantages of high performance and low power. However, current designs are energy-unbalanced among the three parts of RRAM crossbar computation, peripheral circuits and memory accesses, and the latter two factors can significantly limit the potential gains of RCS. Addressing the problem of high power overhead of peripheral circuits in RCS, this paper proposes a Peripheral Circuit Unit (PeriCU)-Reuse scheme to meet power budgets in energy constrained embedded systems. The underlying idea is to put the expensive ADCs/DACs onto spotlight and arrange multiple convolution layers to be sequentially served by the same PeriCU. In the solution, the first step is to determine the number of PeriCUs which are organized by cycle frames. Inside a cycle frame, the layers are computed in parallel inter-PeriCUs while sequentially intra-PeriCU. Furthermore, a layer retiming technique is exploited to further improve the energy of RCS by assigning two adjacent layers within the same PeriCU so as to bypass the energy consuming memory accesses. The experiments of five convolutional applications validate that the PeriCU-Reuse scheme integrated with the retiming technique can efficiently meet variable power budgets, and further reduce energy consumption efficiently.
Keni Qiu, Weiwen Chen, Yuanchao Xu 0002, Lixue Xia, Yu Wang 0002, Zili Shao
DATE4
2018 Design of fault-tolerant neuromorphic computing systems
abstract
Neuromorphic computing is rapidly becoming mainstream, and Resistive Random Access Memory (RRAM) and RRAM-based computing systems (RCS) provide a promising hardware implementation of neuromorphic computing. This emerging computing system helps us to realize vector-matrix multiplications in a time complexity of 0(1), and it improves energy efficiency dramatically. However, due to the immature fabrication process, RCS is susceptible to defects; the resulting errors lead to a significant accuracy drop in neuromorphic computing applications. In order to take advantage of RCS in practical applications, fault-tolerant design is necessary. We present a survey of fault-tolerant designs for RRAM-based neuromorphic computing systems. We first describe RRAM-based crossbars and their role in neuromorphic computing systems. Following this, we classify fault models into different categories, and review the test solutions. Subsequently, the framework of fault-tolerant design for RCS is presented, which contains an online testing phase and a fault-tolerant training phase. The techniques proposed for these two phases are classified and explained to highlight their similarities and differences. The methods reviewed in this survey represent recent trends in fault-tolerant designs of RCS, and are expected motivate further research in this field.
Lixue Xia, Yu Wang 0002, Krishnendu Chakrabarty
ETS2
2018 Mixed size crossbar based RRAM CNN accelerator with overlapped mapping method
abstract
Convolutional Neural Networks (CNNs) play a vital role in machine learning. CNNs are typically both computing and memory intensive. Emerging resistive random-access memories (RRAMs) and RRAM crossbars have demonstrated great potentials in boosting the performance and energy efficiency of CNNs. Compared with small crossbars, large crossbars show better energy efficiency with less interface overhead. However, conventional workload mapping methods for small crossbars cannot make full use of the computation ability of large crossbars. In this paper, we propose an Overlapped Mapping Method (OMM) and MIxed Size Crossbar based RRAM CNN Accelerator (MISCA) to solve this problem. MISCA with OMM can reduce the energy consumption caused by the interface circuits, and improve the parallelism of computation by leveraging the idle RRAM cells in crossbars. The simulation results show that MISCA with OMM can achieve 2.7× speedup, 30% utilization rate improvement, and 1.2× energy efficiency improvement on average compared with fixed size crossbars based accelerator using the conventional mapping method. In comparison with GPU platform, MISCA with OMM can perform 490.4× higher on average in energy efficiency and 20× higher on average in speedup. Compared with PRIME, an existing RRAM based accelerator, MISCA has 26.4× speedup and 1.65× energy efficiency improvement.
Zhenhua Zhu 0002, Jilan Lin, Lixue Xia, Hanbo Sun, Xiaoming Chen 0003, Yu Wang 0002, Huazhong Yang
ICCAD4
2018 Fault Tolerance for RRAM-Based Matrix Operations
abstract
An RRAM-based computing system (RCS) provides an energy efficient hardware implementation of vector-matrix multiplication for machine-learning hardware. However, it is vulnerable to faults due to the immature RRAM fabrication process. We propose an efficient fault tolerance method for RCS; the proposed method, referred to as extended-ABFT (X-ABFT), is inspired by algorithm-based fault tolerance (ABFT). We utilize row checksums and test-input vectors to extract signatures for fault detection and error correction. We present a solution to alleviate the overflow problem caused by the limited number of voltage levels for the test-input signals. Simulation results show that for a Hopfield classifier with faults in 5% of its RRAM cells, X-ABFT allows us to achieve nearly the same classification accuracy as in the fault-free case.
Lixue Xia, Yu Wang 0002, Krishnendu Chakrabarty
ITC2
2018 MNSIM: Simulation Platform for Memristor-Based Neuromorphic Computing System
abstract
Memristor-based computation provides a promising solution to boost the power efficiency of the neuromorphic computing system. However, a behavior-level memristor-based neuromorphic computing simulator, which can model the performance and realize an early stage design space exploration, is still missing. In this paper, we propose a simulation platform for the memristor-based neuromorphic system, called MNSIM. A hierarchical structure for memristor-based neuromorphic computing accelerator is proposed to provides flexible interfaces for customization. A detailed reference design is provided for large-scale applications. A behavior-level computing accuracy model is incorporated to evaluate the computing error rate affected by interconnect lines and nonideal device factors. Experimental results show that MNSIM achieves over 7000 times speed-up than SPICE simulation. MNSIM can optimize the design and estimate the tradeoff relationships among different performance metrics for users.
Lixue Xia, Boxun Li, Tianqi Tang 0001, Peng Gu 0008, Pai-Yu Chen, Shimeng Yu, Yu Cao 0001, Yu Wang 0002, Yuan Xie 0001, Huazhong Yang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2017 Computation-oriented fault-tolerance schemes for RRAM computing systems
abstract
The emerging metal-oxide resistive switching random-access memory (RRAM) devices and RRAM crossbar arrays have demonstrated their potential in enormously boosting the speed and energy-efficiency of analog matrix-vector multiplication. Unfortunately, due to the immature fabrication technology, commonly occurring Stuck-At-Faults (SAFs) seriously degrade the computational accuracy of RRAM crossbar based Computing System (RCS). In this paper, we propose a Mapping Algorithm with inner fault-tolerant ability (MAO) to convert matrix parameters into RRAM conductances in RCS by providing larger mapping space and fully exploring the available mapping space. Furthermore, we present two computation-oriented redundancy schemes — ‘Redundant Crossbars’ (RX) and ‘Independent Redundant Columns’ (IRC) to alleviate the loss of computational accuracy due to SAFs. RX adds redundant RRAM crossbar arrays and IRC introduces independent redundant RRAM columns to compensate the computational errors brought by SAFs.
Wenqin Huangfu, Lixue Xia, Xiling Yin, Tianqi Tang 0001, Boxun Li, Krishnendu Chakrabarty, Yuan Xie 0001, Yu Wang 0002, Huazhong Yang
ASP-DAC2
2017 Binary convolutional neural network on RRAM
abstract
Recent progress in the machine learning field makes low bit-level Convolutional Neural Networks (CNNs), even CNNs with binary weights and binary neurons, achieve satisfying recognition accuracy on ImageNet dataset. Binary CNNs (BCNNs) make it possible for introducing low bit-level RRAM devices and low bit-level ADC/DAC interfaces in RRAM-based Computing System (RCS) design, which leads to faster read-and-write operations and better energy efficiency than before. However, some design challenges still exist: (1) how to make matrix splitting when one crossbar is not large enough to hold all parameters of one layer; (2) how to design the pipeline to accelerate the whole CNN forward process. In this paper, an RRAM crossbar-based accelerator is proposed for BCNN forward process. Moreover, the special design for BCNN is well discussed, especially the matrix splitting problem and the pipeline implementation. In our experiment, BCNNs on RRAM show much smaller accuracy loss than multi-bit CNNs for LeNet on MNIST when considering device variation. For AlexNet on ImageNet, the RRAM-based BCNN accelerator saves 58.2% energy consumption and 56.8% area compared with multi-bit CNN structure.
Tianqi Tang 0001, Lixue Xia, Boxun Li, Yu Wang 0002, Huazhong Yang
ASP-DAC2
2017 TIME: A Training-in-memory Architecture for Memristor-based Deep Neural Networks
abstract
The training of neural network (NN) is usually time-consuming and resource intensive. Memristor has shown its potential in computation of NN. Especially for the metal-oxide resistive random access memory (RRAM), its crossbar structure and multi-bit characteristic can perform the matrix-vector product in high precision, which is the most common operation of NN. However, there exist two challenges on realizing the training of NN. Firstly, the current architecture can only support the inference phase of training and cannot perform the backpropagation (BP), the weights update of NN. Secondly, the training of NN requires enormous iterations and constantly updates the weights to reach the convergence, which leads to large energy consumption because of lots of write and read operations. In this work, we propose a novel architecture, TIME, and peripheral circuit designs to enable the training of NN in RRAM. TIME supports the BP and the weights update while maximizing the reuse of peripheral circuits for the inference operation on RRAM. Meanwhile, a variability-free tuning scheme and gradually-write circuits are designed to reduce the cost of tuning RRAM. We explore the performance of both SL (supervised learning) and DRL (deep reinforcement learning) in TIME, and a specific mapping method of DRL is also introduced to further improve the energy efficiency. Experimental results show that, in SL, TIME can achieve 5.3x higher energy efficiency on average compared with the most powerful application-specific integrated circuits (ASIC) in the literature. In DRL, TIME can perform averagely 126x higher than GPU in energy efficiency. If the cost of tuning RRAM can be further reduced, TIME have the potential of boosting the energy efficiency by 2 orders of magnitude compared with ASIC.
Lixue Xia, Zhenhua Zhu 0002, Yi Cai 0003, Yuan Xie 0001, Yu Wang 0002, Huazhong Yang
DAC2
2017 Fault-Tolerant Training with On-Line Fault Detection for RRAM-Based Neural Computing Systems
abstract
An RRAM-based computing system (RCS) is an attractive hardware platform for implementing neural computing algorithms. Online training for RCS enables hardware-based learning for a given application and reduces the additional error caused by device parameter variations. However, a high occurrence rate of hard faults due to immature fabrication processes and limited write endurance restrict the applicability of on-line training for RCS. We propose a fault-tolerant on-line training method that alternates between a fault-detection phase and a fault-tolerant training phase. In the fault-detection phase, a quiescent-voltage comparison method is utilized. In the training phase, a threshold-training method and a re-mapping scheme is proposed. Our results show that, compared to neural computing without fault tolerance, the recognition accuracy for the Cifar-10 dataset improves from 37% to 83% when using low-endurance RRAM cells, and from 63% to 76% when using RRAM cells with high endurance but a high percentage of initial faults.
Lixue Xia, Xuefei Ning, Krishnendu Chakrabarty, Yu Wang 0002
DAC1
2016 RRAM based learning acceleration
abstract
Deep Learning (DL) is becoming popular in a wide range of domains. Many emerging applications, ranging from image and speech recognition to natural language processing and information retrieval, rely heavily on deep learning techniques, especially the Neural Networks (NNs). NNs have led to great advances in recognition accuracy compared with other traditional methods in recent years. NN-based methods demand much more computation and memory resource, and therefore a number of NN accelerators have been proposed on CMOS-based platforms, such as FPGA and GPU [1]. However, it becomes more and more difficult to obtain substantial power efficiency and gains directly through the scaling down of traditional CMOS technique. Meanwhile, the large data amount in DL applications also meets an ever-increasing "memory wall" challenge because of the efficiency of von Neumann architecture. Consequently, there is a growing research interest of exploring emerging nano-devices and new computing architectures to further improve power efficiency [2].
Yu Wang 0002, Lixue Xia, Tianqi Tang 0001, Boxun Li, Huazhong Yang
CASES2
2016 Switched by input: power efficient structure for RRAM-based convolutional neural network
abstract
Convolutional Neural Network (CNN) is a powerful technique widely used in computer vision area, which also demands much more computations and memory resources than traditional solutions. The emerging metal-oxide resistive random-access memory (RRAM) and RRAM crossbar have shown great potential on neuromorphic applications with high energy efficiency. However, the interfaces between analog RRAM crossbars and digital peripheral functions, namely Analog-to-Digital Converters (ADCs) and Digital-to-Analog Converters (DACs), consume most of the area and energy of RRAM-based CNN design due to the large amount of intermediate data in CNN. In this paper, we propose an energy efficient structure for RRAM-based CNN. Based on the analysis of data distribution, a quantization method is proposed to transfer the intermediate data into 1 bit and eliminate DACs. An energy efficient structure using input data as selection signals is proposed to reduce the ADC cost for merging results of multiple crossbars. The experimental results show that the proposed method and structure can save 80% area and more than 95% energy while maintaining the same or comparable classification accuracy of CNN on MNIST.
Lixue Xia, Tianqi Tang 0001, Wenqin Huangfu, Xiling Yin, Boxun Li, Yu Wang 0002, Huazhong Yang
DAC1
2016 Sparsity-oriented sparse solver design for circuit simulation
Xiaoming Chen 0003, Lixue Xia, Yu Wang 0002, Huazhong Yang
DATE2
2016 MNSIM: Simulation platform for memristor-based neuromorphic computing system
Lixue Xia, Boxun Li, Tianqi Tang 0001, Peng Gu 0008, Xiling Yin, Wenqin Huangfu, Pai-Yu Chen, Shimeng Yu, Yu Cao 0001, Yu Wang 0002, Yuan Xie 0001, Huazhong Yang
DATE1
2016 Low power Convolutional Neural Networks on a chip
abstract
Deep learning, and especially Convolutional Neural Network (CNN, is among the most powerful and widely used techniques in computer vision. Applications range from image classification to object detection, segmentation, Optical Character Recognition (OCR), etc. At the same time, CNNs are both computationally intensive and memory intensive, making them difficult to be deployed on low power lightweight embedded systems. In this work, we introduce an on-chip convoltional neural network implementation for low-power embedded system. We point out that the high precision of weights limits the low-power CNN implementation on both FPGA and RRAM platform. A dynamic quantization method is introduced to reduce the precision while maintaining the same or comparable accuracy at the same time. Finally, the de ailed designs of low-power FPGA-based CNN and RRAM-based CNN are provided and compared. The results show that FPGA-based design gets 2× energy efficiency compared with GPU implementation, and toe RRAM-based design can further obtain more than 40× energy efficiency gains.
Yu Wang 0002, Lixue Xia, Tianqi Tang 0001, Boxun Li, Huazhong Yang
ISCAS2
2016 Technological Exploration of RRAM Crossbar Array for Matrix-Vector Multiplication
Lixue Xia, Peng Gu 0008, Boxun Li, Tianqi Tang 0001, Xiling Yin, Wenqin Huangfu, Shimeng Yu, Yu Cao 0001, Yu Wang 0002, Huazhong Yang
J. Comput. Sci. Technol.1
2015 An accurate and low-cost PM2.5 estimation method based on Artificial Neural Network
abstract
PM2.5has already been a major pollutant in many cities in China. It is a kind of harmful pollutant which may cause several kinds of lung diseases. However, the existing methods to monitor PM2.5with high accuracy are too expensive to popularize. The high cost also limits the further researches about PM2.5. This paper implements a method to estimate PM2.5with low cost and high accuracy by Artificial Neural Network (ANN) technique using other pollutants and meteorological factors that are easy to be monitored. An Entropy Maximization step is proposed to avoid the over-fitting related to the data distribution of pollutant data. Also, how to choose the input attributes is abstracted to an optimization problem. An iterative greedy algorithm is proposed to solve it, which reduces the cost and increases the estimation accuracy at the same time. The experiment shows that the linear correlation coefficient between the estimated value and real value is 0.9488. Our model can also classify PM2.5levels with a high accuracy. Additionally, the trade-off between accuracy and cost is investigated according to the price and error rate of each sensor.
Lixue Xia, Yu Wang 0002, Huazhong Yang
ASP-DAC1
2015 Merging the interface: power, area and accuracy co-optimization for RRAM crossbar-based mixed-signal computing system
abstract
The invention of resistive-switching random access memory (RRAM) devices and RRAM crossbar-based computing system (RCS) demonstrate a promising solution for better performance and power efficiency. The interfaces between analog and digital units, especially AD/DAs, take up most of the area and power consumption of RCS and are always the bottleneck of mixed-signal computing systems. In this work, we propose a novel architecture, MEI, to minimize the overhead of AD/DA by MErging the Interface into the RRAM crossbar. An optional ensemble method, the Serial Array Adaptive Boosting (SAAB), is also introduced to take advantage of the area and power saved by MEI and boost the accuracy and robustness of RCS. On top of these two methods, a design space exploration is proposed to achieve trade-offs among accuracy, area, and power consumption. Experimental results on 6 diverse benchmarks demonstrate that, compared with the traditional architecture with AD/DAs, MEI is able to save 54.63%~86.14% area and reduce 61.82%~86.80% power consumption under quality guarantees; and SAAB can further improve the accuracy by 5.76% on average and ensure the system performance under noisy conditions.
Boxun Li, Lixue Xia, Peng Gu 0008, Yu Wang 0002, Huazhong Yang
DAC2
2015 Spiking neural network with RRAM: can we use it for real-world application?
Tianqi Tang 0001, Lixue Xia, Boxun Li, Yiran Chen 0001, Yu Wang 0002, Huazhong Yang
DATE2
2015 Energy Efficient RRAM Spiking Neural Network for Real Time Classification
abstract
Inspired by the human brain's function and efficiency, neuromorphic computing offers a promising solution for a wide set of tasks, ranging from brain machine interfaces to real-time classification. The spiking neural network (SNN), which encodes and processes information with bionic spikes, is an emerging neuromorphic model with great potential to drastically promote the performance and efficiency of computing systems. However, an energy efficient hardware implementation and the difficulty of training the model significantly limit the application of the spiking neural network. In this work, we address these issues by building an SNN-based energy efficient system for real time classification with metal-oxide resistive switching random-access memory (RRAM) devices. We implement different training algorithms of SNN, including Spiking Time Dependent Plasticity (STDP) and Neural Sampling method. Our RRAM SNN systems for these two training algorithms show good power efficiency and recognition performance on realtime classification tasks, such as the MNIST digit recognition. Finally, we propose a possible direction to further improve the classification accuracy by boosting multiple SNNs.
Yu Wang 0002, Tianqi Tang 0001, Lixue Xia, Boxun Li, Peng Gu 0008, Huazhong Yang, Hai Li 0001, Yuan Xie 0001
ACM Great Lakes Symposium on VLSI3