Peng Gu 0008

dblp:33/5231-8 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
2since 2021 · last 2021
0000-0002-2663-4568ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 19 · 5 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
10 papers
Memory systems · 42% Hardware accelerators and domain-specific architectures · 31% Emerging computing paradigms · 10%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 100%

Topics — the 30 heaviest of 39, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
processing-in-memory
1.952021
DLUX: A LUT-Based Near-Bank Accelerator for Data Center Deep Learning Training Workloads · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2021
SpaceA: Sparse Matrix Vector Multiplication on Processing-in-Memory Accelerator · HPCA 2021
iPIM: Programmable In-Memory Image Processing Accelerator Using Near-Bank Architecture · ISCA 2020
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.822020
NNBench-X: A Benchmarking Methodology for Neural Network Accelerator Designs · ACM Trans. Archit. Code Optim. 2020
SCOPE: A Stochastic Computing Engine for DRAM-Based In-Situ Accelerator · MICRO 2018
Memory systems
3d-stacked memory
0.512021
SpaceA: Sparse Matrix Vector Multiplication on Processing-in-Memory Accelerator · HPCA 2021
Memory systems › processing-in-memory
3D-stacked PIM
0.512021
DLUX: A LUT-Based Near-Bank Accelerator for Data Center Deep Learning Training Workloads · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2021
Memory systems › DRAM › DRAM microarchitecture
bank-level parallelism
0.512021
SpaceA: Sparse Matrix Vector Multiplication on Processing-in-Memory Accelerator · HPCA 2021
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN training accelerator
0.512021
DLUX: A LUT-Based Near-Bank Accelerator for Data Center Deep Learning Training Workloads · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2021
Memory systems › in-memory computing
near-bank computing
0.512021
SpaceA: Sparse Matrix Vector Multiplication on Processing-in-Memory Accelerator · HPCA 2021
High-performance computing › sparse linear algebra › sparse matrix computation
sparse matrix-vector multiplication
0.512021
SpaceA: Sparse Matrix Vector Multiplication on Processing-in-Memory Accelerator · HPCA 2021
Hardware accelerators and domain-specific architectures › sparse matrix multiplication accelerator
sparse matrix-vector multiplication accelerator
0.512021
SpaceA: Sparse Matrix Vector Multiplication on Processing-in-Memory Accelerator · HPCA 2021
Compilers and program optimization
accelerator compilation
0.412020
iPIM: Programmable In-Memory Image Processing Accelerator Using Near-Bank Architecture · ISCA 2020
Compilers and program optimization › domain-specific compilation
image processing pipeline compilation
0.412020
iPIM: Programmable In-Memory Image Processing Accelerator Using Near-Bank Architecture · ISCA 2020
Performance modeling and evaluation
benchmarking
0.412020
NNBench-X: A Benchmarking Methodology for Neural Network Accelerator Designs · ACM Trans. Archit. Code Optim. 2020
Hardware accelerators and domain-specific architectures
image processing accelerator
0.412020
iPIM: Programmable In-Memory Image Processing Accelerator Using Near-Bank Architecture · ISCA 2020
Hardware accelerators and domain-specific architectures › machine learning accelerator › neural network accelerator
neural network accelerator design
0.412020
NNBench-X: A Benchmarking Methodology for Neural Network Accelerator Designs · ACM Trans. Archit. Code Optim. 2020
Hardware accelerators and domain-specific architectures
bioinformatics accelerator
0.412019
MEDAL: Scalable DIMM based Near Data Processing Accelerator for DNA Seeding Algorithm · MICRO 2019
Memory systems › processing-in-memory › near-memory processing
DIMM-based near-memory processing
0.412019
MEDAL: Scalable DIMM based Near Data Processing Accelerator for DNA Seeding Algorithm · MICRO 2019
Hardware accelerators and domain-specific architectures
graph processing accelerator
0.412019
Alleviating Irregularity in Graph Analytics Acceleration: a Hardware/Software Co-Design Approach · MICRO 2019
Electronic design automation
hardware/software co-design
0.412019
Alleviating Irregularity in Graph Analytics Acceleration: a Hardware/Software Co-Design Approach · MICRO 2019
Performance modeling and evaluation › simulation › architectural simulation
accelerator simulation
0.312018
MNSIM: Simulation Platform for Memristor-Based Neuromorphic Computing System · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN accelerator
0.312018
SCOPE: A Stochastic Computing Engine for DRAM-Based In-Situ Accelerator · MICRO 2018
Emerging computing paradigms
neuromorphic computing
0.312018
MNSIM: Simulation Platform for Memristor-Based Neuromorphic Computing System · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018
Memory systems › processing-in-memory
processing-using-DRAM
0.312018
SCOPE: A Stochastic Computing Engine for DRAM-Based In-Situ Accelerator · MICRO 2018
Emerging computing paradigms › approximate and stochastic computing › stochastic computing
stochastic arithmetic
0.312018
SCOPE: A Stochastic Computing Engine for DRAM-Based In-Situ Accelerator · MICRO 2018
Emerging computing paradigms › approximate and stochastic computing
stochastic computing
0.312018
SCOPE: A Stochastic Computing Engine for DRAM-Based In-Situ Accelerator · MICRO 2018
Memory systems
DRAM
0.322021
DLUX: A LUT-Based Near-Bank Accelerator for Data Center Deep Learning Training Workloads · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2021
iPIM: Programmable In-Memory Image Processing Accelerator Using Near-Bank Architecture · ISCA 2020
Emerging computing paradigms
approximate computing
0.212015
RRAM-Based Analog Approximate Computing · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015
Integrated circuit design › analog and mixed-signal circuits
mixed-signal circuit design
0.212015
Merging the interface: power, area and accuracy co-optimization for RRAM crossbar-based mixed-signal computing system · DAC 2015
Memory systems
non-volatile memory
0.212015
RRAM-Based Analog Approximate Computing · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015
Memory systems › non-volatile memory
resistive memory
0.212015
RRAM-Based Analog Approximate Computing · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2015
Machine learning › Efficient and distributed learning
model compression
0.112020
NNBench-X: A Benchmarking Methodology for Neural Network Accelerator Designs · ACM Trans. Archit. Code Optim. 2020

Methods — techniques the papers use, named apart from their topics

software-hardware co-design · 0.9quantitative metric analysis · 0.9workload balancing · 0.5scratchpad buffer · 0.5near-bank architecture · 0.5loop tiling · 0.5data mapping · 0.5content-addressable memory · 0.5LUT-based FP multiplier · 0.5DIMM-based NDP · 0.4
YearPublicationVenuePosition
2021 SpaceA: Sparse Matrix Vector Multiplication on Processing-in-Memory Accelerator
abstract
Sparse matrix-vector multiplication (SpMV) is an important primitive across a wide range of application domains such as scientific computing and graph analytics. Due to its intrinsic memory-bound characteristics, the performance of SpMV on throughput-oriented architectures such as GPU is bounded by the limited bandwidth between processors and memory. Processing-in-memory (PIM) architectures, made feasible by advances in 3D stacking, provide new opportunities to utilize ultra-high bandwidth by integrating compute-logic into memory.In this paper, we develop an SpMV accelerator, named as SpaceA, based on PIM architectures. SpaceA integrates compute logic near memory banks to exploit bank-level bandwidth. SpaceA contains both hardware and data-mapping design features to alleviate irregular memory access patterns which hinder full utilization of high memory bandwidth. In terms of hardware design features, SpaceA consists of two unique features: (1) it utilizes the capability of outstanding memory requests to hide the memory access latency to data located in non-local memory banks; (2) it integrates Content Addressable Memory (CAM) at the bank level to exploit data reuse of the input vectors. In addition, we develop a mapping scheme that partitions the sparse matrix into different memory banks, to maximize the data locality of the input vector and to achieve workload balance among processing elements (PEs) near each bank. Overall, SpaceA together with the proposed mapping method achieves 13.54x speedup and 87.49% energy saving on average over the GPU baseline on SpMV computation. In addition to SpMV primitives, we conduct a case study on graph analytics to demonstrate the benefits of SpaceA for applications built on SpMV. Compared to Tesseract and GraphP, state-of-the-art graph accelerators, SpaceA obtains better performance due to its higher effective bandwidth provided by near-bank integration.
Xinfeng Xie, Zheng Liang 0003, Peng Gu 0008, Abanti Basak, Lei Deng 0003, Ling Liang 0003, Xing Hu 0001, Yuan Xie 0001
HPCA3
2021 DLUX: A LUT-Based Near-Bank Accelerator for Data Center Deep Learning Training Workloads
abstract
The frequent data movement between the processor and the memory has become a severe performance bottleneck for deep neural network (DNN) training workloads in data centers. To solve this off-chip memory access challenge, the 3-D stacking processing-in-memory (3D-PIM) architecture provides a viable solution. However, existing 3D-PIM designs for DNN training suffer from the limited memory bandwidth in the base logic die. To overcome this obstacle, integrating the DNN related logic near each memory bank becomes a promising yet challenging solution, since naively implementing the floating-point (FP) unit and the cache in the memory die incurs a large area overhead. To address these problems, we propose DLUX, a high performance and energy-efficient 3D-PIM accelerator for DNN training using the near-bank architecture. From the hardware perspective, to support the FP multiplier with low area overhead, an in-DRAM lookup table (LUT) mechanism is invented. Then, we propose to use a small scratchpad buffer together with a lightweight transformation engine to exploit the locality and enable flexible data layout without the expensive cache. From the software aspect, we split the mapping/scheduling tasks during DNN training into intralayer and interlayer phases. During the intralayer phase, to maximize data reuse in the LUT buffer and the scratchpad buffer, achieve high concurrency, and reduce data movement among banks, a 3D-PIM customized loop tiling technique is adopted. During the interlayer phase, efficient techniques are invented to ensure the input-output data layout consistency and realize the forward-backward layout transposition. Experiment results show that DLUX can reduce FP32 multiplier area overhead by 60% against the direct implementation. Compared with a Tesla V100 GPU, end-to-end evaluations show that DLUX can provide on average 6.3× speedup and 42× energy efficiency improvement.
Peng Gu 0008, Xinfeng Xie, Shuangchen Li, Dimin Niu, Hongzhong Zheng, Krishna T. Malladi, Yuan Xie 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2020 NEST: DIMM based Near-Data-Processing Accelerator for K-mer Counting
abstract
With the ability to help wildlife conservation, precise medical care, and disease understanding, genomics analysis is becoming more and moe important. Recently, with the development and wide adoption of the Next-Generation Sequencing (NGS) technology, bio-data grows exponentially, putting forward great challenges for k-mer counting - a widely used application in genomics analysis.
Wenqin Huangfu, Krishna T. Malladi, Shuangchen Li, Peng Gu 0008, Yuan Xie 0001
ICCAD4
2020 iPIM: Programmable In-Memory Image Processing Accelerator Using Near-Bank Architecture
abstract
Image processing is becoming an increasingly important domain for many applications on workstations and the datacenter that require accelerators for high performance and energy efficiency. GPU, which is the state-of-the-art accelerator for image processing, suffers from the memory bandwidth bottleneck. To tackle this bottleneck, near-bank architecture provides a promising solution due to its enormous bank-internal bandwidth and low-energy memory access. However, previous work lacks hardware programmability, while image processing workloads contain numerous heterogeneous pipeline stages with diverse computation and memory access patterns. Enabling programmable near-bank architecture with low hardware overhead remains challenging.This work proposes iPIM, the first programmable in-memory image processing accelerator using near-bank architecture. We first design a decoupled control-execution architecture to provide lightweight programmability support. Second, we propose the SIMB (Single-Instruction-Multiple-Bank) ISA to enable flexible control flow and data access. Third, we present an end-to-end compilation flow based on Halide that supports a wide range of image processing applications and maps them to our SIMB ISA. We further develop iPIM-aware compiler optimizations, including register allocation, instruction reordering, and memory order enforcement to improve performance. We evaluate a set of representative image processing applications on iPIM and demonstrate that on average iPIM obtains 11.02× acceleration and 79.49% energy saving over an NVIDIA Tesla V100 GPU. Further analysis shows that our compiler optimizations contribute 3.19× speedup over the unoptimized baseline.
Peng Gu 0008, Xinfeng Xie, Yufei Ding 0001, Guoyang Chen, Weifeng Zhang 0003, Dimin Niu, Yuan Xie 0001
ISCA1
2020 NNBench-X: A Benchmarking Methodology for Neural Network Accelerator Designs
abstract
The tremendous impact of deep learning algorithms over a wide range of application domains has encouraged a surge of neural network (NN) accelerator research. Facilitating the NN accelerator design calls for guidance from an evolving benchmark suite that incorporates emerging NN models. Nevertheless, existing NN benchmarks are not suitable for guiding NN accelerator designs. These benchmarks are either selected for general-purpose processors without considering unique characteristics of NN accelerators or lack quantitative analysis to guarantee their completeness during the benchmark construction, update, and customization. In light of the shortcomings of prior benchmarks, we propose a novel benchmarking methodology for NN accelerators with a quantitative analysis of application performance features and a comprehensive awareness of software-hardware co-design. Specifically, we decouple the benchmarking process into three stages: First, we characterize the NN workloads with quantitative metrics and select the representative applications for the benchmark suite to ensure diversity and completeness. Second, we refine the selected applications according to the customized model compression techniques provided by specific software-hardware co-design. Finally, we evaluate a variety of accelerator designs on the generated benchmark suite. To demonstrate the effectiveness of our benchmarking methodology, we conduct a case study of composing an NN benchmark from the TensorFlow Model Zoo and compress these selected models with various model compression techniques. Finally, we evaluate compressed models on various architectures, including GPU, Neurocube, DianNao, and Cambricon-X.
Xinfeng Xie, Xing Hu 0001, Peng Gu 0008, Shuangchen Li, Yu Ji 0002, Yuan Xie 0001
ACM Trans. Archit. Code Optim.3
2019 MEDAL: Scalable DIMM based Near Data Processing Accelerator for DNA Seeding Algorithm
abstract
Computational genomics has proven its great potential to support precise and customized health care. However, with the wide adoption of the Next Generation Sequencing (NGS) technology, 'DNA Alignment', as the crucial step in computational genomics, is becoming more and more challenging due to the booming bio-data. Consequently, various hardware approaches have been explored to accelerate DNA seeding - the core and most time consuming step in DNA alignment.
Wenqin Huangfu, Xueqi Li 0001, Shuangchen Li, Xing Hu 0001, Peng Gu 0008, Yuan Xie 0001
MICRO5
2019 Alleviating Irregularity in Graph Analytics Acceleration: a Hardware/Software Co-Design Approach
abstract
Graph analytics is an emerging application which extracts insights by processing large volumes of highly connected data, namely graphs. The parallel processing of graphs has been exploited at the algorithm level, which in turn incurs three irregularities onto computing and memory patterns that significantly hinder an efficient architecture design. Certain irregularities can be partially tackled by the prior domain-specific accelerator designs with well-designed scheduling of data access, while others remain unsolved.
Mingyu Yan, Xing Hu 0001, Shuangchen Li, Abanti Basak, Han Li 0011, Itir Akgun, Yujing Feng, Peng Gu 0008, Lei Deng 0003, Xiaochun Ye, Zhimin Zhang 0004, Dongrui Fan, Yuan Xie 0001
MICRO9
2018 SCOPE: A Stochastic Computing Engine for DRAM-Based In-Situ Accelerator
abstract
Memory-centric architecture, which bridges the gap between compute and memory, is considered as a promising solution to tackle the memory wall and the power wall. Such architecture integrates the computing logic and the memory resources close to each other, in order to embrace large internal memory bandwidth and reduce the data movement overhead. The closer the compute and memory resources are located, the greater these benefits become. DRAM-based in-situ accelerators [1] tightly couple processing units to every memory bitline, achieving the maximum benefits among various memory-centric architectures. However, the processing units in such architectures are typically limited to simple functions like AND/OR due to strict area and power overhead constraints in DRAMs, making it difficult to accomplish complex tasks while providing high performance. In this paper, we address the challenge by applying stochastic computing arithmetic to the DRAM-based in-situ accelerator, targeting at the acceleration of error-tolerant applications such as deep learning. In stochastic computing, binary numbers are converted into stochastic bitstreams, which turns integer multiplications into simple bitwise AND operations, but at the expense of larger memory capacity/bandwidth demands. Stochastic computing is a perfect match for the DRAM-based in-situ accelerators because it addresses the in-situ accelerator's low performance problem by simplifying the operations, while leveraging the in-situ accelerator's advantage of large memory capacity/bandwidth. To further boost the performance and compensate for the numerical precision loss, we propose a novel Hierarchical and Hybrid Deterministic (H2D) stochastic computing arithmetic. Finally, we consider quantized deep neural network inference and training applications as a case study. The proposed architecture provides 2.3× improvement in performance per unit area compared with the binary arithmetic baseline, and 3.8× improvement over GPU. The proposed H2D arithmetic contributes 11× performance boost and 60% numerical precision improvement.
Shuangchen Li, Alvin Oliver Glova, Xing Hu 0001, Peng Gu 0008, Dimin Niu, Krishna T. Malladi, Hongzhong Zheng, Bob Brennan, Yuan Xie 0001
MICRO4
2018 MNSIM: Simulation Platform for Memristor-Based Neuromorphic Computing System
abstract
Memristor-based computation provides a promising solution to boost the power efficiency of the neuromorphic computing system. However, a behavior-level memristor-based neuromorphic computing simulator, which can model the performance and realize an early stage design space exploration, is still missing. In this paper, we propose a simulation platform for the memristor-based neuromorphic system, called MNSIM. A hierarchical structure for memristor-based neuromorphic computing accelerator is proposed to provides flexible interfaces for customization. A detailed reference design is provided for large-scale applications. A behavior-level computing accuracy model is incorporated to evaluate the computing error rate affected by interconnect lines and nonideal device factors. Experimental results show that MNSIM achieves over 7000 times speed-up than SPICE simulation. MNSIM can optimize the design and estimate the tradeoff relationships among different performance metrics for users.
Lixue Xia, Boxun Li, Tianqi Tang 0001, Peng Gu 0008, Pai-Yu Chen, Shimeng Yu, Yu Cao 0001, Yu Wang 0002, Yuan Xie 0001, Huazhong Yang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2017 Security Threats and Countermeasures in Three-Dimensional Integrated Circuits
abstract
Existing works on Three-dimensional (3D) hardware security focus on leveraging the unique 3D characteristics to address the supply chain attacks that exist in 2D design. However, 3D ICs introduce specific and unexplored challenges as well as new opportunities for managing hardware security. In this paper, we analyze new security threats unique to 3D ICs. The corresponding attack models are summarized for future research. Furthermore, existing representative countermeasures, including split manufacturing, camouflaging, transistor locking, techniques against thermal signal based side-channel attacks, and network-on-chip based shielding plane (NoCSIP) for different hardware threats are reviewed and categorized. Moreover, preliminary countermeasures are proposed to thwart TSV-based hardware Trojan insertion attacks.
Jaya Dofe, Peng Gu 0008, Dylan C. Stow, Qiaoyan Yu, Eren Kursun, Yuan Xie 0001
ACM Great Lakes Symposium on VLSI2
2016 MNSIM: Simulation platform for memristor-based neuromorphic computing system
Lixue Xia, Boxun Li, Tianqi Tang 0001, Peng Gu 0008, Xiling Yin, Wenqin Huangfu, Pai-Yu Chen, Shimeng Yu, Yu Cao 0001, Yu Wang 0002, Yuan Xie 0001, Huazhong Yang
DATE4
2016 Leveraging 3D Technologies for Hardware Security: Opportunities and Challenges
abstract
3D die stacking and 2.5D interposer design are promising technologies to improve integration density, performance and cost. Current approaches face serious issues in dealing with emerging security challenges such as side channel attacks, hardware trojans, secure IC manufacturing and IP piracy. By utilizing intrinsic characteristics of 2.5D and 3D technologies, we propose novel opportunities in designing secure systems. We present: (i) a 3D architecture for shielding side-channel information; (ii) split fabrication using active interposers; (iii) circuit camouflage on monolithic 3D IC, and (iv) 3D IC-based security processing-in-memory (PIM). Advantages and challenges of these designs are discussed, showing that the new designs can improve existing countermeasures against security threats and further provide new security features.
Peng Gu 0008, Shuangchen Li, Dylan C. Stow, Russell Barnes, Liu Liu 0017, Yuan Xie 0001, Eren Kursun
ACM Great Lakes Symposium on VLSI1
2016 NVSim-CAM: a circuit-level simulator for emerging nonvolatile memory based content-addressable memory
abstract
Ternary Content-Addressable Memory (TCAM) is widely used in networking routers, fully associative caches, search engines, etc. While the conventional SRAM-based TCAM suffers from the poor scalability, the emerging nonvolatile memories (NVM, i.e., MRAM, PCM, and ReRAM) bring evolution for the TCAM design. It effectively reduces the cell size, and makes significant energy reduction and scalability improvement. New applications such as associative processors/accelerators are facilitated by the emergence of the nonvolatile TCAM (nvTCAM). However, nvTCAM design is challenging. In addition to the emerging device's uncertainty, the nvTCAM cell structure is so diverse that it results in a design space too large to explore manually. To tackle these challenges, we propose a circuit-level model and develop a simulation tool, NVSim-CAM, which helps researchers to make early design decisions, and to evaluate device/circuit innovations. The tool is validated by HSPICE simulations and data from fabricated chips. We also present a case study to illustrate how NVSim-CAM benefits the nvTCAM design. In the case study, we propose a novel 3D vertical ReRAM based TCAM cell, the 3DvTCAM. We project the advantages/disadvantages and explore the design space for the proposed cell with NVSim-CAM.
Shuangchen Li, Liu Liu 0017, Peng Gu 0008, Cong Xu 0002, Yuan Xie 0001
ICCAD3
2016 Cost analysis and cost-driven IP reuse methodology for SoC design based on 2.5D/3D integration
abstract
Due to the increasing fabrication and design complexity with new process nodes, the cost per transistor trend originally identified in Moore's Law is slowing when using traditional integration methods. However, emerging die-level integration technologies may be viable alternatives that can scale the number of transistors per integrated device while reducing the cost per transistor through yield improvements across multiple smaller dies. Additionally, the escalating overheads of non-recurring engineering costs like masks and verification can be curtailed through die integration-enabled reuse of intellectual property across heterogeneous process technologies. In this paper, we present an analytical cost model for 3D and interposer-based 2.5D die integration and employ it to demonstrate the potential cost reductions across semiconductor markets. We also propose a methodology and platform for IP reuse based on these integration technologies and explore the available reductions in overall product cost through reduction in non-recurring engineering effort.
Dylan C. Stow, Itir Akgun, Russell Barnes, Peng Gu 0008, Yuan Xie 0001
ICCAD4
2016 Thermal-aware 3D design for side-channel information leakage
abstract
Side-channel attacks are important security challenges as they reveal sensitive information about on-chip activities. Among such attacks, the thermal side-channel has been shown to disclose the activities of key functional blocks and even encryption keys. This paper proposes a novel approach to proactively conceal critical activities in the functional layers while minimizing the power dissipation by (i) leveraging inherent characteristics of 3D integration to protect from side-channel attacks and (ii) dynamically generating custom activity patterns to match the activity to be concealed in the functional layers. Experimental analysis shows that 3D technology combined with the proposed run-time algorithm effectively reduces the Side-channel Vulnerability Factor (SVF) below 0.05 and the Spatial Thermal Side-channel Factor (STSF) below 0.59.
Peng Gu 0008, Dylan C. Stow, Russell Barnes, Eren Kursun, Yuan Xie 0001
ICCD1
2016 Technological Exploration of RRAM Crossbar Array for Matrix-Vector Multiplication
Lixue Xia, Peng Gu 0008, Boxun Li, Tianqi Tang 0001, Xiling Yin, Wenqin Huangfu, Shimeng Yu, Yu Cao 0001, Yu Wang 0002, Huazhong Yang
J. Comput. Sci. Technol.2
2015 Technological exploration of RRAM crossbar array for matrix-vector multiplication
abstract
The matrix-vector multiplication is the key operation for many computationally intensive algorithms. In recent years, the emerging metal oxide resistive switching random access memory (RRAM) device and RRAM crossbar array have demonstrated a promising hardware realization of the analog matrix-vector multiplication with ultra-high energy efficiency. In this paper, we analyze the impact of nonlinear voltage-current relationship of RRAM devices and the interconnect resistance as well as other crossbar array parameters on the circuit performance and present a design guide. On top of that, we propose a technological exploration flow for device parameter configuration to overcome the impact of nonideal factors and achieve a better trade-off among performance, energy and reliability for each specific application. The simulation results of a support vector machine (SVM) and MNIST pattern recognition dataset show that the RRAM crossbar array-based SVM is robust to the input signal fluctuation but sensitive to the tunneling gap deviation. A further resistance resolution test presents that a 4-bit RRAM device is able to realize a recognition accuracy of ∼ 90%, indicating the physical feasibility of RRAM crossbar array-based SVM. In addition, the proposed technological exploration flow is able to achieve 10.98% improvement of recognition accuracy on the MNIST dataset and 26.4% energy savings compared with previous work.
Peng Gu 0008, Boxun Li, Tianqi Tang 0001, Shimeng Yu, Yu Cao 0001, Yu Wang 0002, Huazhong Yang
ASP-DAC1
2015 Merging the interface: power, area and accuracy co-optimization for RRAM crossbar-based mixed-signal computing system
abstract
The invention of resistive-switching random access memory (RRAM) devices and RRAM crossbar-based computing system (RCS) demonstrate a promising solution for better performance and power efficiency. The interfaces between analog and digital units, especially AD/DAs, take up most of the area and power consumption of RCS and are always the bottleneck of mixed-signal computing systems. In this work, we propose a novel architecture, MEI, to minimize the overhead of AD/DA by MErging the Interface into the RRAM crossbar. An optional ensemble method, the Serial Array Adaptive Boosting (SAAB), is also introduced to take advantage of the area and power saved by MEI and boost the accuracy and robustness of RCS. On top of these two methods, a design space exploration is proposed to achieve trade-offs among accuracy, area, and power consumption. Experimental results on 6 diverse benchmarks demonstrate that, compared with the traditional architecture with AD/DAs, MEI is able to save 54.63%~86.14% area and reduce 61.82%~86.80% power consumption under quality guarantees; and SAAB can further improve the accuracy by 5.76% on average and ensure the system performance under noisy conditions.
Boxun Li, Lixue Xia, Peng Gu 0008, Yu Wang 0002, Huazhong Yang
DAC3
2015 Energy Efficient RRAM Spiking Neural Network for Real Time Classification
abstract
Inspired by the human brain's function and efficiency, neuromorphic computing offers a promising solution for a wide set of tasks, ranging from brain machine interfaces to real-time classification. The spiking neural network (SNN), which encodes and processes information with bionic spikes, is an emerging neuromorphic model with great potential to drastically promote the performance and efficiency of computing systems. However, an energy efficient hardware implementation and the difficulty of training the model significantly limit the application of the spiking neural network. In this work, we address these issues by building an SNN-based energy efficient system for real time classification with metal-oxide resistive switching random-access memory (RRAM) devices. We implement different training algorithms of SNN, including Spiking Time Dependent Plasticity (STDP) and Neural Sampling method. Our RRAM SNN systems for these two training algorithms show good power efficiency and recognition performance on realtime classification tasks, such as the MNIST digit recognition. Finally, we propose a possible direction to further improve the classification accuracy by boosting multiple SNNs.
Yu Wang 0002, Tianqi Tang 0001, Lixue Xia, Boxun Li, Peng Gu 0008, Huazhong Yang, Hai Li 0001, Yuan Xie 0001
ACM Great Lakes Symposium on VLSI5
2015 RRAM-Based Analog Approximate Computing
abstract
Approximate computing is a promising design paradigm for better performance and power efficiency. In this paper, we propose a power efficient framework for analog approximate computing with the emerging metal-oxide resistive switching random-access memory (RRAM) devices. A programmable RRAM-based approximate computing unit (RRAM-ACU) is introduced first to accelerate approximated computation, and an approximate computing framework with scalability is then proposed on top of the RRAM-ACU. In order to program the RRAM-ACU efficiently, we also present a detailed configuration flow, which includes a customized approximator training scheme, an approximator-parameter-to-RRAM-state mapping algorithm, and an RRAM state tuning scheme. Finally, the proposed RRAM-based computing framework is modeled at system level. A predictive compact model is developed to estimate the configuration overhead of RRAM-ACU and help explore the application scenarios of RRAM-based analog approximate computing. The simulation results on a set of diverse benchmarks demonstrate that, compared with a x86-64 CPU at 2 GHz, the RRAM-ACU is able to achieve 4.06-196.41× speedup and power efficiency of 24.59-567.98 GFLOPS/W with quality loss of 8.72% on average. And the implementation of hierarchical model and X application demonstrates that the proposed RRAM-based approximate computing framework can achieve 12.8× power efficiency than its pure digital implementation counterparts (CPU, graphics processing unit, and field- programmable gate arrays).
Boxun Li, Peng Gu 0008, Yu Wang 0002, Yiran Chen 0001, Huazhong Yang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2