Bonan Yan

dblp:161/0839 · DBLP profile ↗
← Back
39ranked-venue papers
7as first author
19since 2021 · last 2026
0000-0002-3052-9330ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 36 · 5 first-author · 18 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2026 M3DKV: Monolithic 3D Gain Cell Memory Enabled Efficient KV Cache & Processing
abstract
Transformer-based generative large language models (LLMs) have revolutionized natural language processing, yet their quadratic growth in computational complexity in context length creates severe inference bottlenecks. While LLM keyvalue cache (KV cache) enhances decoding efficiency, prolonged contexts infer frequent KV cache reloads that exacerbate memory bandwidth constraints. To address this hardware challenge, we propose M3DKV-a monolithic three-dimensional (3D) gain cell near-memory computing accelerator featuring back-end-of-line (BEOL) cache layers for in-situ KV matrix buffering and computation and a front-end-of-line (FEOL) base layer for full selfattention operations. Through optimized 3D data organization, inter-layer dataflow management, and intelligent computation scheduling, our design achieves $0.29 \mathrm{~TB} / \mathrm{s} /$ core on-die bandwidth while demonstrating $97.03 \times / 268.01 \times$ speedup over GPU/CPU in the decoding stage and $1.72 \times-262.16 \times$ better area efficiency per parameter versus state-of-the-art accelerators.
Jiaqi Yang 0009, Yanbo Su, Yihan Fu, Jianshi Tang, Bonan Yan
ASP-DAC5
2026 Reconfigurable Computing Challenge: FPGA-Gym-v2: FPGA-Based RL Environment Acceleration with LLM-Assisted Onboarding
abstract
Reinforcement learning (RL) repeatedly executes environment rollouts computation, where environment stepping may dominate training time when many environments run in parallel. Building on our prior FPGA-Gym/PEARL acceleration backbone, this work introduces FPGA-Gym-v2, a upgraded framework that keeps RL inference, training, and replay-buffer management on the host CPU/GPU, but offloads massively parallel environment stepping to an FPGA. It reduces host–FPGA traffic with compact encoding, on-chip state residency, and compute–communication overlap. Across representative Gymnasium workloads, FPGA-Gym-v2 achieves 4.36×–972.6× throughput speedup over EnvPool and up to 4699.3× over VectorEnv, while reducing DQN/PPO training time by 7%–15% over EnvPool and 18%–61% over VectorEnv. Further, FPGA-Gym-v2 adds an large-language-model-assisted (LLM-assisted) onboarding flow that generates environment-specific specifications, wrappers, Verilog HDL skeletons, and golden tests under explicit validation gates. The code is available at https://github.com/Selinaee/FPGA_Gym.
Hongxiao Zhao, Wenshuo Yue, Yihan Fu, Daijing Shi, Anjunyi Fan, Bonan Yan
FCCM7
2026 Reconfigurable Computing Challenge: FPGA-Based WebAssembly Stack Co-Processor
abstract
Large language models suffer from hallucinations when performing scientific computing, motivating the use of AI agents such as IronClaw that offload computation to specialized tools. IronClaw invokes tools implemented as WebAssembly (Wasm) plugins for security and extensibility, but the stack-based Wasm bytecode is mismatched with register-based processors (x86, ARM), causing runtime overhead. We propose PAWS, a native Wasm coprocessor that directly executes Wasm bytecode in hardware. PAWS features: (1) full support for all five Wasm instruction types; (2) dual digital stack circuits (operand stack and control stack) replacing register files to minimize memory access latency; (3) dedicated control logic for block-based branching; and (4) a sliding-window instruction fetch unit that decodes variable-length Wasm instructions. Evaluated on the PolyBench suite, PAWS achieves average execution latencies 28.6× lower than an Intel Xeon processor and 40.6× lower than an Nvidia Jetson TX2, making it highly suitable for IronClaw’s compute-intensive scientific applications. The design is available at https://github.com/Iris-WQP/PAWS_FPGA_softcore.
Qiuping Wu, Mugeng Liu 0001, Hongxiao Zhao, Yihan Fu, Gang Huang 0001, Yun Ma 0002, Bonan Yan
FCCM8
2026 ESTroM: Element-Flow Architecture for Processing Sparse Tractable Probabilistic Models
abstract
Probabilistic Circuits (PCs) models are emerging popular tractable probabilistic models. Their internal connections are represented in the form of directed acyclic graphs (DAGs) with sum nodes and product nodes, ensuring their internal parameter efficiency and model expressiveness in terms of probabilistic inference. Despite these algorithmic advantages, executing PC still faces graph structure deployment issues. PyJuice on GPU with the block-sparse parallel computation methods causes a parallelism-sparsity gap, while DAG-style processing does not take advantage of the repetitive characteristics of PC internal nodes, resulting in low throughput. To address this challenge, this work proposes the ESTroM, an efficient architecture that provides novel graph-element (nodes/edges) parallelism with sparsity-aware compilation. Through analysis of the sum/product node computational requirements, ESTroM core uses compressed matrices for sum/product nodes DAG representations, edge-based dataflow for product node processing, and node-based dataflow for sum node processing. With intra-core rewind and intercore multicast optimizations, we develop a prototype ESTrom chip and a demonstrative system for a PC-based neural lossless compression application. Our ablation experiments show ESTrom offers a speed improvement of$2.11 \sim 3.79 \times$compared to the state-of-the-art DAG processing unit (DPU)-v2 with the same computing resources. Under various typical PC structures, ESTrom achieves a speedup of$18.7 \times$compared to DPU-v2 and$3.9 \times$compared to NVIDIA RTX 4090 GPU with PyJuice framework. In terms of neural lossless compression, ESTroM demonstrates a$1.39 \times$improvement in compression ratio compared to the industrial-standard Z-standard (Zstd) algorithms with the highest compression level, while offering$16.3 \sim 65.2 \times$improvement in compression speed compared to Zstd on Intel Xeon Gold 6230. In a nutshell, this work develops novel graph element parallelism and element-flow architecture theory with practical prototype chips and systems, revealing a new hardware-perspective path for the “scaling law” of emerging tractable probabilistic models.
Anjunyi Fan, Xuejie Liu, Anji Liu, Qiuping Wu, Jiaqi Yang 0009, Yuchao Qin, Guy Van den Broeck, Yitao Liang, Bonan Yan
HPCA9
2026 Phase-Change Memory Based Passive Timing Circuits for DRAM Refresh Management
Anjunyi Fan, Yingming Lu, Bonan Yan, Yuchao Yang 0001
ISCAS5
2026 Efficient SRAM-PIM Co-Design by Joint Exploration of Value-Level and Bit-Level Sparsity
abstract
Processing-in-memory (PIM) architectures mitigate the Von Neumann bottleneck by integrating computation units into memory arrays. Among PIM architectures, digital SRAMPIM has become a prominent approach, directly integrating digital logic within the SRAM array. However, the rigid crossbar architecture and full array activation pose challenges in efficiently utilizing value-level sparsity. Moreover, neural network models exhibit a high proportion of zero bits within non-zero values, which remain underutilized due to architectural constraints. To overcome these limitations, we present Dyadic Block PIM (DB-PIM), a groundbreaking algorithm-architecture co-design framework to harness both value-level and bit-level sparsity. At the algorithm level, our hybrid-grained pruning technique, combined with a novel sparsity pattern, enables effective sparsity management. Architecturally, DB-PIM incorporates a sparse network and customized digital SRAM-PIM macros, including input pre-processing unit (IPU), dyadic block multiply units (DBMUs), and Canonical Signed Digit (CSD)-based adder trees. It circumvents structured zero values in weights and bypasses unstructured zero bits within non-zero weights and block-wise all-zero bit columns in input features. As a result, the DBPIM framework skips a majority of unnecessary computations, thereby driving significant gains in computational efficiency. Experimental results demonstrate that our DB-PIM framework achieves up to 8.01× speedup and 85.28% energy savings, significantly boosting computational efficiency in digital SRAMPIM systems.
Cenlin Duan, Jianlei Yang 0001, Yiou Wang, Yingjie Qi, Xiaolin He, Bonan Yan, Xiaotao Jia, Weisheng Zhao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2025 PEARL: FPGA-Based Reinforcement Learning Acceleration with Pipelined Parallel Environments
abstract
Reinforcement learning (RL) is an effective machine learning approach that enables artificial intelligence agents to perform complex tasks and make decisions in dynamic situations. Training an RL agent demands its repetitive interaction with the environment to learn optimal policies. To efficiently collect training data, parallelizing environments is a widely used technique by enabling simultaneous interactions between multiple agents and environments. However, existing CPU-based RL software frameworks face a key challenge of slow multi-environmental update computation. To solve this problem, we present a novel FPGA-based RL accelerating framework-PEARL. PEARL instantiates multiple parallel environments and accelerates them with a carefully designed pipeline scheme to hide data transfer latency within the computation time. We evaluate PEARL on respective RL environments and achieve 4.36 x to 972.6 x speedup over the existing fastest software-based framework for parallel environment execution. When scaling the number of environments from 1024 to 43008 (42x) in CliffWalking benchmark, the power consumption increases marginally by 3%, while LUT and flip-flops utilization rise by 2.24 x and 3.08 x, respectively. This demonstrates efficient resource usage and power management in PEARL. Further, PEARL allows users to define and add their environments within the framework flexibly. We have established an open-source repository for users to utilize and expand. We also implement PEARL with the existing RL algorithm and save 7% -15% training time. All the source code is available online https://github.com/Selinaee/FPGA_Gym.
Hongxiao Zhao, Wenshuo Yue, Yihan Fu, Daijing Shi, Anjunyi Fan, Yuchao Yang 0001, Bonan Yan
DATE8
2025 PROCA: Programmable Probabilistic Processing Unit Architecture with Accept/Reject Prediction & Multicore Pipelining for Causal Inference
abstract
Causal inference is an important field in data science and cognitive artificial intelligence. It requires the construction of complex probabilistic models to describe the causal relationships between random variables. Probabilistic models rely on probabilistic programming as a flexible framework. However, the computing speed of probabilistic programming is often hindered by the extensive use of Markov chain Monte Carlo (MCMC) algorithms, even though they are powerful in Bayesian inference. To accelerate MCMC, this work presents PROCA, a programmable MCMC-based probabilistic processing unit architecture. PROCA exploits processing-in-memory function units to generate new samples of Markov chains. PROCA is programmable to execute the computation for arbitrary forms of posterior distribution formulas that software probabilistic programming frameworks support. We develop a novel accept/reject prediction methodology to accelerate the sequential MCMC computation, thereby introducing efficient multi-core pipelining methods. We implement and validate the PROCA architecture with commercial process development kits. The implementation is evaluated based on 9 representative benchmarks, covering PyMC official tutorial probabilistic problems, single-variable probabilistic problems, and real-world causal inference problems. Our comprehensive experiments demonstrate that PROCA achieves a speedup of 172~4871 $\times$ compared to Intel Xeon Gold CPU, $42 \sim 1058 \times$ compared to NVIDIA A100 GPU, and $1.765 \times$ over state-of-the-art MCMC accelerators, respectively. PROCA achieves comparable statistical robustness to the software probabilistic programming frameworks. Compared with state-of-the-art MCMC domain-specific accelerators, our design boosts the energy efficiency by $9.47 \times$.
Yihan Fu, Anjunyi Fan, Wenshuo Yue, Hongxiao Zhao, Daijing Shi, Qiuping Wu, Yaoyu Tao, Yuchao Yang 0001, Bonan Yan
HPCA11
2024 Towards Efficient SRAM-PIM Architecture Design by Exploiting Unstructured Bit-Level Sparsity
abstract
Bit-level sparsity in neural network models harbors immense untapped potential. Eliminating redundant calculations of randomly distributed zero-bits significantly boosts computational efficiency. Yet, traditional digital SRAM-PIM architecture, limited by rigid crossbar architecture, struggles to effectively exploit this unstructured sparsity. To address this challenge, we propose Dyadic Block PIM (DB-PIM), a groundbreaking algorithm-architecture co-design framework. First, we propose an algorithm coupled with a distinctive sparsity pattern, termed a dyadic block (DB), that preserves the random distribution of non-zero bits to maintain accuracy while restricting the number of these bits in each weight to improve regularity. Architecturally, we develop a custom PIM macro that includes dyadic block multiplication units (DBMUs) and Canonical Signed Digit (CSD)-based adder trees, specifically tailored for Multiply-Accumulate (MAC) operations. An input pre-processing unit (IPU) further refines performance and efficiency by capitalizing on block-wise input sparsity. Results show that our proposed co-design framework achieves a remarkable speedup of up to 7.69× and energy savings of 83.43%.
Cenlin Duan, Jianlei Yang 0001, Yiou Wang, Yingjie Qi, Xiaolin He, Bonan Yan, Xiaotao Jia, Weisheng Zhao 0001
DAC7
2024 MeMCISA: Memristor-Enabled Memory-Centric Instruction-Set Architecture for Database Workloads
abstract
The exponential growth of data exerts great pressure on hardware design for database systems. Memory-centric computing (MCC) architecture, which enable compute capabilities near or inside memory storage, demonstrate great potential in enhancing the efficiency of database operations with higher compute parallelism and reduced data movements. However, existing MCC architecture mainly focus on artificial intelligence (AI) computations and those designed for database applications can only run a limited number of standalone queries such as SORT or JOIN, lacking efficient support for increasingly diverse and complex database workloads. For example, realizing a commercial recommendation engine on database requires supporting workloads including but not limited to vector aggregation, convolution or$N$-hop neighborhoods computing, etc. In this work, we develop a memristor-enabled memory-centric instruction-set architecture (MeMCISA) aiming to efficiently accelerate versatile workloads in modern database systems. MeMCISA features scalable multi-bank memristor-based storage organization with near-memory circuitries and caches in banks. An out-of-order (O0O) scheduling scheme is designed for MeMCISA based on a vector instruction set with four types of instructions (bit-level, element-level, vector-level, and control-level), combining memristor-enabled in-memory computing and near-memory computing to efficiently run workloads with varying computational kernels and data sizes. MeMCISA can support parallel instruction executions across different memristor banks as well as different hardware modules within a memristor bank. Furthermore, we develop data dependency handling mechanisms to support vector dependency scenarios in MeMCISA that do not exist in conventional scalar-based instruction sets. A prototype MeMCISA is implemented based on a 40nm CMOS technology with necessary peripheral hardware including instruction buffer and instruction scheduler. To accurately study MeMCISA performance in real-world database systems, a software-hardware co-designed framework integrating reconfigurable MeMCISA prototype is created that can support end-to-end simulations for database workloads starting from raw software codes. Based on this framework, we evaluate MeMCISA performance with standalone database queries as well as complex database workloads from representative benchmarks including UniBench, neural collaborative filtering (NCF), and ResNet-18. Simulation results demonstrate that MeMCISA achieves up to 41.84 × ~ 1767.70 × in speed compared to general-purpose processors (CPUs/GPUs).
Yihang Zhu, Lianfeng Yu, Anjunyi Fan, Longhao Yan, Zhaokun Jing, Bonan Yan, Pek Jun Tiw, Yaoyu Tao, Yuchao Yang 0001
MICRO7
2024 DDC-PIM: Efficient Algorithm/Architecture Co-Design for Doubling Data Capacity of SRAM-Based Processing-in-Memory
abstract
Processing-in-memory (PIM), as a novel computing paradigm, provides significant performance benefits from the aspect of effective data movement reduction. SRAM-based PIM has been demonstrated as one of the most promising candidates due to its endurance and compatibility. However, the integration density of SRAM-based PIM is much lower than other nonvolatile memory-based ones, due to its inherent 6T structure for storing a single bit. Within comparable area constraints, SRAM-based PIM exhibits notably lower capacity. Thus, aiming to unleash its capacity potential, we propose DDC-PIM, an efficient algorithm/architecture co-design methodology that effectively doubles the equivalent data capacity. At the algorithmic level, we propose a filter-wise complementary correlation (FCC) algorithm to obtain a bitwise complementary pair. At the architecture level, we exploit the intrinsic cross-coupled structure of 6T SRAM to store the bitwise complementary pair in their complementary states$(Q/\overline {Q})$, thereby maximizing the data capacity of each SRAM cell. The dual-broadcast input structure and reconfigurable unit support both depthwise and pointwise convolution, adhering to the requirements of various neural networks. Evaluation results show that DDC-PIM yields about$2.84\times $speedup on MobileNetV2 and$2.69\times $on EfficientNet-B0 with negligible accuracy loss compared with PIM baseline implementation. Compared with state-of-the-art SRAM-based PIM macros, DDC-PIM achieves up to$8.41\times $and$2.75\times $improvement in weight density and area efficiency, respectively.
Cenlin Duan, Jianlei Yang 0001, Xiaolin He, Yingjie Qi, Yiou Wang, Ziyan He, Bonan Yan, Xiaotao Jia, Weitao Pan, Weisheng Zhao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.8
2024 Probabilistic Compute-in-Memory Design for Efficient Markov Chain Monte Carlo Sampling
abstract
Markov chain Monte Carlo (MCMC) is a widely used sampling method in modern artificial intelligence and probabilistic computing systems. It involves repetitive random number generations and thus often dominates the latency of probabilistic model computing. Hence, we propose a compute-in-memory (CIM) based MCMC design as a hardware acceleration solution. This work investigates SRAM bitcell stochasticity and proposes a novel “pseudo-read” operation, based on which we offer a block-wise random number generation circuit scheme for fast random number generation. Moreover, this work proposes a novel multi-stage exclusive-OR gate (MSXOR) design method to generate strictly uniformly distributed random numbers. The probability error deviating from a uniform distribution is suppressed under$10^{-6}$. Also, this work presents a novel in-memory copy circuit scheme to realize data copy inside a CIM sub-array, significantly reducing the use of R/W circuits for power saving. Evaluated in a commercial 28-nm process development kit, this CIM-based MCMC design generates 4-bit$\sim$32-bit samples with an energy efficiency of 0.53 pJ/sample and high throughput of up to 1066.7M samples/s. Compared to conventional processors, the overall energy efficiency improves$2.12\times10^{9}$to$9.58\times10^{9}$times.
Yihan Fu, Daijing Shi, Anjunyi Fan, Wenshuo Yue, Yuchao Yang 0001, Ru Huang 0001, Bonan Yan
IEEE Trans. Circuits Syst. I Regul. Pap.7
2023 STAR: An Efficient Softmax Engine for Attention Model with RRAM Crossbar
abstract
RRAM crossbars have been studied to construct in-memory accelerators for neural network applications due to their in-situ computing capability. However, prior RRAM-based accelerators show efficiency degradation when executing the popular attention models. We observed that the frequent softmax operations arise as the efficiency bottleneck and also are insensitive to computing precision. Thus, we propose STAR, which boosts the computing efficiency with an efficient RRAM-based softmax engine and a fine-grained global pipeline for the attention models. Specifically, STAR exploits the versatility and flexibility of RRAM crossbars to trade off the model accuracy and hardware efficiency. The experimental results evaluated on several datasets show STAR achieves up to 30.63×and 1.31× computing efficiency improvements over the GPU and the state-of-the-art RRAM-based attention accelerators, respectively.
Yifeng Zhai, Bonan Yan, Jing Wang 0055
DATE3
2023 Accelerating Neural-ODE Inference on FPGAs with Two-Stage Structured Pruning and History-based Stepsize Search
abstract
Neural ordinary differential equation (Neural-ODE) outperforms conventional deep neural networks (DNNs) in modeling continuous-time or dynamical systems by adopting numerical ODE integration onto a shallow embedded NN. However, Neural-ODE suffers from slow inference due to the costly iterative stepsize search in numerical integration, especially when using higher-order Runge-Kutta (RK) methods and smaller error tolerance for improved integration accuracy. In this work, we first present algorithmic techniques to speedup RK-based Neural-ODE inference: a two-stage coarse-grained/fine-grained structured pruning method based on top-K sparsification that reduces the overall computations by more than 60% in the embedded NN and a history-based stepsize search method based on past integration steps that reduces the latency for reaching accepted stepsize by up to 77% in RK methods. A reconfigurable hardware architecture is co-designed based on proposed speedup techniques, featuring three processing loops to support programmable embedded NN and a variety of higher-order RK methods. Sparse activation processor with multi-dimensional sorters is designed to exploit structured sparsity in activations. Implemented on a Xilinx Virtex-7 XC7VX690T FPGA and experimented on a variety of datasets, the prototype accelerator using a more complex 3rd-order RK method achieves more than 2.6x speedup compared to the latest Neural-ODE FPGA accelerator using the simplest Euler method. Compared to a software execution on Nvidia A100 GPU, the inference speedup can be up to 18x.
Jing Wang 0172, Lianfeng Yu, Bonan Yan, Yaoyu Tao, Yuchao Yang 0001
FPGA4
2023 SRAM-Based Processing-In-Memory Design with Kullback-Leibler Divergence-Based Dynamic Precision Quantization
abstract
Deep convolutional neural networks (CNNs) are widely used in Artificial Intelligence of Things (AIoT) systems. Limited by power and area, conventional edge devices are insufficient to handle the cost of CNN computation. The idea of SRAM based Processing-In-Memory (SRAM-PIM) has been advocated to implement CNN on edge devices because of its high area and power efficiency. To further excavate the potential of SRAM-PIM on edge inferences, this paper proposes an SRAM-PIM design with Kullback-Leibler (KL) divergence-based dynamic precision quantization. The proposed quantization method decouples the effect of different CNN layers on accuracy and introduces the SRAM-PIM hardware performance in quantization, realizing SRAM-PIM-aware layer-wise precision adjustment. The proposed SRAM-PIM design has been applied in image classification tasks on edge devices. Our evaluation shows that the implemented design achieves up to 2.03x energy efficiency improvement and 2.54% accuracy improvement compared with existing dynamic precision PIM design. Compared with existing reinforcement-learning-based dynamic quantization method that requires several hours quantization time, the proposed dynamic precision quantization method takes only 26.28us to get the optimal quantization results.
Chunshan Zu, Bingqian Wang, Zhenhua Zhu 0002, Yaojun Zhang, Ran Duan 0004, Bing Li 0017, Bonan Yan
ACM Great Lakes Symposium on VLSI8
2023 Memristive dynamics enabled neuromorphic computing systems
Bonan Yan, Yuchao Yang 0001, Ru Huang 0001
Sci. China Inf. Sci.1
2022 Heterogeneous Memory Architecture Accommodating Processing-in-Memory on SoC for AIoT Applications
abstract
Processing-In-Memory (PIM) technologies is one of most promising candidates for AIoT applications due to its attractive characteristics, such as low computation latency, large throughput and high power efficiency. However, how to efficiently utilize PIM with System-on-Chip (SoC) architecture has been scarcely discussed. In this paper, we demonstrate a series of solution from hardware architecture to algorithm to maximize the benefits of PIM design. First, we propose a Heterogeneous Memory Architecture (HMA) that facilitates the existing SoC with PIM via high-throughput on-chip buses. Then, based on given HMA structure, we also propose an HMA tensor mapping approach to partition tensors and deploy general matrix multiplication operations on PIM structures. Both HMA hardware and HMA tensor mapping approach harnesses the programmability of the mature embedded CPU solution stack and maximize the high efficiency of PIM technology. The whole HMA system can save 416 x power as well as 44.6% design area compare with the latest accelerator solutions. The evaluation also shows that our design can reduce the operation latency by 430 × and 11 × for TinyML applications, compare with state-of-art baseline and PIM without optimization, respectively.
Kangyi Qiu, Yaojun Zhang, Bonan Yan, Ru Huang 0001
ASP-DAC3
2022 ASTERS: adaptable threshold spike-timing neuromorphic design with twin-column ReRAM synapses
abstract
Complex event-driven neuron dynamics was an obstacle to implementing efficient brain-inspired computing architectures with VLSI circuits. To solve this problem and harness the event-driven advantage, we propose ASTERS, a resistive random-access memory (ReRAM) based neuromorphic design to conduct the time-to-first-spike SNN inference. In addition to the fundamental novel axon and neuron circuits, we also propose two techniques through hardware-software co-design: "Multi-Level Firing Threshold Adjustment" to mitigate the impact of ReRAM device process variations, and "Timing Threshold Adjustment" to further speed up the computation. Experimental results show that our cross-layer solution ASTERS achieves more than 34.7% energy savings compared to the existing spiking neuromorphic designs, meanwhile maintaining 90.1% accuracy under the process variations with a 20% standard deviation.
Ziru Li, Qilin Zheng, Bonan Yan, Ru Huang 0001, Bing Li 0005, Yiran Chen 0001
DAC3
2022 VSDCA: A Voltage Sensing Differential Column Architecture Based on 1T2R RRAM Array for Computing-in-Memory Accelerators
abstract
Non-volatile memory (NVM) such as RRAM and PCM has become the key component in high energy efficiency computing-in-memory (CIM) architectures. However, the computing accuracy and energy efficiency improvement of conventional 1T1R RRAM array based current sensing CIM scheme is hindered by device variation and large output current. In this work, we propose a voltage sensing differential column architecture (VSDCA) based on 1T2R RRAM array for binary memory and CIM applications. The memory mode of VSDCA macro can improve$1.12\times $to$5.29\times $relative read margin compared to conventional 1T1R current sensing memory. The computing mode supports 8-bit input, 9-bit weight and 18-bit output high precision and rows fully parallel computing. The VSDCA macro design is evaluated under SMIC 40 nm technology node, the energy efficiency for the high precision CIM reaches 39.52 TOPS/W. The CIFAR10 inference accuracy of the simulated VGG16 and ResNet18 model is 85.91% and 89.32% respectively.
Zhaokun Jing, Bonan Yan, Yuchao Yang 0001, Ru Huang 0001
IEEE Trans. Circuits Syst. I Regul. Pap.2
2020 ReSiPE: ReRAM-based Single-Spiking Processing-In-Memory Engine
abstract
Processing-in-memory (PIM) designs that leverage emerging nanotechnologies like resistive random access memory (ReRAM) have demonstrated enormous potential in accelerating deep learning applications due to high energy efficiency and integration density. The common approach of existing ReRAM-based PIM designs is to encode data into either voltage levels with the assist of power-thirsty analog/digital conversion circuits or spike series by sacrificing computing latency. In this paper, we introduce ReSiPE, a ReRAM-based Single-spiking PIM Engine, which uses the arrival time of a single spike to represent the data. We analyze how to encode data into a set of single spikes and develop the circuit to realize the matrix-based computation. The proposed design can minimize the spike numbers, shorten the computation period, and thus improve energy efficiency dramatically. Our simulation results show that ReSiPE achieves 67.1% power reduction and 1.97× power efficiency improvement compared to rate-coding based ReRAM PIM designs under the comparable area, throughput, and accuracy.
Ziru Li, Bonan Yan, Hai Li 0001
DAC2
2020 Lattice: An ADC/DAC-less ReRAM-based Processing-In-Memory Architecture for Accelerating Deep Convolution Neural Networks
abstract
Nonvolatile Processing-In-Memory (NVPIM) has demonstrated its great potential in accelerating Deep Convolution Neural Networks (DCNN). However, most of existing NVPIM designs require costly analog-digital conversions and often rely on excessive data copies or writes to achieve performance speedup. In this paper, we propose a new NVPIM architecture, namely, Lattice, which calculates the partial sum of the dot products between the feature map and weights of network layers in a CMOS peripheral circuit to eliminate the analog-digital conversions. Lattice also naturally offers an efficient data mapping scheme to align the data of the feature maps and the weights and hence, avoiding the excessive data copies or writes in the previous NVPIM designs. Finally, we develop a zero-flag encoding scheme to save the energy of processing zero-values in sparse DCNNs. Our experimental results show that Lattice improves the system energy efficiency by 4× ~ 13.22× compared to three state-of-the-art NVPIM designs: ISAAC, PipeLayer, and FloatPIM.
Qilin Zheng, Zongwei Wang 0001, Zishun Feng, Bonan Yan, Yimao Cai, Ru Huang 0001, Yiran Chen 0001, Chia-Lin Yang, Hai Li 0001
DAC4
2020 ReTransformer: ReRAM-based Processing-in-Memory Architecture for Transformer Acceleration
abstract
Transformer has emerged as a popular deep neural network (DNN) model for Neural Language Processing (NLP) applications and demonstrated excellent performance in neural machine translation, entity recognition, etc. However, its scaled dot-product attention mechanism in auto-regressive decoder brings a performance bottleneck during inference. Transformer is also computationally and memory intensive and demands for a hardware acceleration solution. Although researchers have successfully applied ReRAM-based Processing-in-Memory (PIM) to accelerate convolutional neural networks (CNNs) and recurrent neural networks (RNNs), the unique computation process of the scaled dot-product attention in Transformer makes it difficult to directly apply these designs. Besides, how to handle intermediate results in Matrix-matrix Multiplication (MatMul) and how to design a pipeline at a finer granularity of Transformer remain unsolved. In this work, we propose ReTransformer - a ReRAM-based PIM architecture for Transformer acceleration. ReTransformer can not only accelerate the scaled dot-product attention of Transformer using ReRAM-based PIM but also eliminate some data dependency by avoiding writing the intermediate results using the proposed matrix decomposition technique. Moreover, we propose a new sub-matrix pipeline design for multi-head self-attention. Experimental results show that compared to GPU and Pipelayer, ReTransformer improves computing efficiency by 23.21× and 3.25×, respectively. The corresponding overall power is reduced by 1086× and 2.82×, respectively.
Xiaoxuan Yang 0001, Bonan Yan, Hai Li 0001, Yiran Chen 0001
ICCAD2
2019 Build reliable and efficient neuromorphic design with memristor technology
abstract
Neuromorphic computing is a revolutionary approach of computation, which attempts to mimic the human brain's mechanism for extremely high implementation efficiency and intelligence. Latest research studies showed that the memristor technology has a great potential for realizing power- and area-efficient neuromorphic computing systems (NCS). On the other hand, the memristor device processing is still under development. Unreliable devices can severely degrade system performance, which arises as one of the major challenges in developing memristor-based NCS. In this paper, we first review the impacts of the limited reliability of memristor devices and summarize the recent research progress in building reliable and efficient memristor-based NCS. In the end, we discuss the main difficulties and the trend in memristor-based NCS development.
Bing Li 0017, Bonan Yan, Hai Li 0001
ASP-DAC2
2019 An Overview of In-memory Processing with Emerging Non-volatile Memory for Data-intensive Applications
abstract
The conventional von Neumann architecture has been revealed as a major performance and energy bottleneck for rising data-intensive applications. The decade-old idea of leveraging in-memory processing to eliminate substantial data movements has returned and led extensive research activities. The effectiveness of in-memory processing heavily relies on memory scalability, which cannot be satisfied by traditional memory technologies. Emerging non-volatile memories (eNVMs) that pose appealing qualities such as excellent scaling and low energy consumption, on the other hand, have been heavily investigated and explored for realizing in-memory processing architecture. In this paper, we summarize the recent research progress in eNVM-based in-memory processing from various aspects, including the adopted memory technologies, locations of the in-memory processing in the system, supported arithmetics, as well as applied applications.
Bing Li 0017, Bonan Yan, Hai Li 0001
ACM Great Lakes Symposium on VLSI2
2019 Hardware Fault Tolerance for Binary RRAM Crossbars
abstract
Resistive random-access memory (RRAM)-based computing systems (RCS) are being advocated for neural network acceleration. The memristor is the unit cell of an RCS and it is susceptible to process variations and manufacturing defects. Therefore, it is essential to tolerate faulty memristors to ensure intended system operation. We present the architecture of a novel processing element to tolerate both stuck-at and undefined-state faults in binary RRAM cells. We also describe a 4T1R reconfigurable cell-based crossbar design with an ancillary 3T mesh to provide 100% hardware fault tolerance for random and clustered fault distributions for up to 50% fault density. The proposed 4T1R cell is 2.04× smaller than the state-of-the-art neuromorphic SRAM cell. Evaluation results for binary pattern-matching and digit recognition applications demonstrate the effectiveness of our fault tolerance methodology.
Arjun Chaudhuri, Bonan Yan, Yiran Chen 0001, Krishnendu Chakrabarty
ITC2
2018 A neuromorphic design using chaotic mott memristor with relaxation oscillation
abstract
The recent proposed nanoscale Mott memristor features negative differential resistance and chaotic dynamics. This work proposes a novel neuromorphic computing system that utilizes Mott memristors to simplify peripheral circuitry. According to the analytic description of chaotic dynamics and relaxation oscillation, we carefully tune the working point of Mott memristors to balance the chaotic behavior weighing testing accuracy and training efficiency. Compared with conventional designs, the proposed design accelerates the training by 1.893× averagely and saves 27.68% and 43.32% power consumption with 36.67% and 26.75% less area for single-layer and two-layer perceptrons, respectively.
Bonan Yan, Xiong Cao, Hai Li 0001
DAC1
2018 Exploring the opportunity of implementing neuromorphic computing systems with spintronic devices
abstract
Many cognitive algorithms such as neural networks cannot be efficiently executed by von Neumann architectures, the performance of which is constrained by the memory wall between microprocessor and memory hierarchy. Hence, researchers started to investigate new computing paradigms such as neuromorphic computing that can adapt their structure to the topology of the algorithms and accelerate their executions. New computing units have been also invented to support this effort by leveraging emerging nano-devices. In this work, we will discuss the opportunity of implementing neuromorphic computing systems with spintronic devices. We will also provide insights on how spintronic devices fit into different part of neuromorphic computing systems. Approaches to optimize the circuits are also discussed.
Bonan Yan, Fan Chen 0001, Yaojun Zhang, Chang Song 0001, Hai Li 0001, Yiran Chen 0001
DATE1
2018 Challenges of memristor based neuromorphic computing system
Bonan Yan, Yiran Chen 0001, Hai Li 0001
Sci. China Inf. Sci.1
2017 Low-power neuromorphic speech recognition engine with coarse-grain sparsity
abstract
In recent years, we have seen a surge of interest in neuromorphic computing and its hardware design for cognitive applications. In this work, we present new neuromorphic architecture, circuit, and device co-designs that enable spike-based classification for speech recognition task. The proposed neuromorphic speech recognition engine supports a sparsely connected deep spiking network with coarse granularity, leading to large memory reduction with minimal index information. Simulation results show that the proposed deep spiking neural network accelerator achieves phoneme error rate (PER) of 20.5% for TIMIT database, and consume 2.57mW in 40nm CMOS for real-time performance. To alleviate the memory bottleneck, the usage of non-volatile memory is also evaluated and discussed.
Shunti Yin, Deepak Kadetotad, Bonan Yan, Chang Song 0001, Yiran Chen 0001, Chaitali Chakrabarti, Jae-sun Seo
ASP-DAC3
2017 A closed-loop design to enhance weight stability of memristor based neural network chips
abstract
Compared with the algorithm optimizations, brain-inspired neural network chips aim to fundamentally change the computer architecture and therefore enhance the computation capability and performance in advanced data processing. In recent years, memristor technology has been investigated in developing high-speed and large-capacity neural network chips. However, it has been observed that memristance values that represent the well-trained network weights can be disturbed by electrical or thermal perturbations. It severely degrades overall system reliability and emerges as a major design challenge. In this work, we systematically analyze the impacts of low-voltage induced memristance drift upon weight disturbance after times of recall operations. A closed-loop design by introducing a real-time feedback controller is proposed to enhance the weight stability of memristor based neural network chips. By mimicking the training process, the controller adaptively compensates the memristance deviation, according to the relation of the input data and recall output. In view of tiny disturbance per access, we integrate the memristance compensation into regular recall operation to avoid the execution speed degradation. Our simulations based on the implementation of representative single-layer (two-layer) network show that the proposed closed-loop design can prolong the service time of memristor-based neural network chip by 14.85x (14.94x), without reducing computational speed. Extra circuitry of the feedback controller induces a negligible overhead about 1.16% on overall power consumption.
Bonan Yan, J. Joshua Yang, Qing Wu 0002, Yiran Chen 0001, Hai Li 0001
ICCAD1
2017 Giant Spin-Hall assisted STT-RAM and logic design
Enes Eken, Ismail Bayram, Yaojun Zhang, Bonan Yan, Hai Li 0001, Yiran Chen 0001
Integr.4
2017 Persistent and Nonpersistent Error Optimization for STT-RAM Cell Design
abstract
Rapidly increasing demands for memory capacity and severe technical scaling challenges of conventional memory technologies motivated recent investments on next-generation nonvolatile memory technologies. As a promising candidate, spin-transfer torque random access memory (STT-RAM) has demonstrated many attractive properties, such as nanosecond access time, high integration density, nonvolatility, and excellent CMOS integration compatibility. However, similar to all other nano-devices, the performance and reliability of STT-RAM cells are greatly affected by process variations, device operating uncertainties, and environmental fluctuations. As a result, the read and write operations of STT-RAM demonstrate some variabilities and errors. In this paper, we systematically analyze the impacts of CMOS and magnetic tunneling junction (MTJ) process variations, MTJ resistance switching randomness that are induced by intrinsic thermal fluctuations, and working temperature changes on STT-RAM cell designs. The STT-RAM cell reliability issues in both read and write operations are first investigated. A combined circuit and magnetic simulation platform is then established to quantitatively study the persistent and nonpersistent errors in STT-RAM cell operations. Our analysis proved the importance of a full statistical design method in STT-RAM designs for design pessimism minimization.
Yaojun Zhang, Bonan Yan, Xiaobin Wang, Yiran Chen 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2016 A neuromorphic ASIC design using one-selector-one-memristor crossbar
abstract
The applications of memristors in neuromorphic computing have been extensively studied for its analogy to synapse. To overcome sneak path issue, nonlinear resistive selectors have been introduced to the design of memristor crossbar, enabling a high integration density and robust computing capability. However, the nonlinearity of such selectors also influences the computation accuracy of the vector-matrix multiplication performed on the memristor crossbar. In this work, we evaluate the impact of nonlinear resistive selectors on the computation robustness of a Hopfield spike-based pattern recognition system based on memristor crossbar technology. The methods that can suppress the adverse impact of the nonlinear selector on the system performance are also studied.
Bonan Yan, Amr Mahmoud Mahmoud, J. Joshua Yang, Qing Wu 0002, Yiran Chen 0001, Hai Li 0001
ISCAS1
2015 Spiking-based matrix computation by leveraging memristor crossbar array
abstract
As process technology continues scaling down, the memory barrier becomes more severe. Thus, spiking neuromorphic computing that can significantly enhance computing and communication efficiencies has been widely studied. Both conventional CMOS technology and emerging devices have been used in hardware implementation of spiking neuromorphic computing. Particularly, the memristor technology that can naturally emulate plasticity and energy efficiency of biological synapses have gained a lot of attention. However, the use of memristors in high density computation, such as matrix-vector operation, is still missing. In this work, a spiking (pulse-based) computing component that leverages memristor crossbar array is proposed for matrix-vector operation. We adopt the rate coding model and count the produced spike number during a given time period of T as the computation output. We carefully design the crossbar array structure and the integrate-and-fire circuit. The linear relationship between output spike numbers and the sum-of-production of input vector and matrix entries is observed in our simulation results. The proposed spiking computing design realizes matrix computation successfully and demonstrates good adaptability in neural network.
Hai Li 0001, Bonan Yan, Chaofei Yang, Linghao Song, Yiran Chen 0001, Qing Wu 0002, Hao Jiang 0014
CISDA3
2015 A spiking neuromorphic design with resistive crossbar
abstract
Neuromorphic systems recently gained increasing attention for their high computation efficiency. Many designs have been proposed and realized with traditional CMOS technology or emerging devices. In this work, we proposed a spiking neuromorphic design built on resistive crossbar structures and implemented with IBM 130nm technology. Our design adopts a rate coding scheme where pre- and post-neuron signals are represented by digitalized pulses. The weighting function of pre-neuron signals is executed on the resistive crossbar in analog format. The computing result is transferred into digitalized output spikes via an integrate-and-fire circuit (IFC) as the post-neuron. We calibrated the computation accuracy of the entire system through circuit simulations. The results demonstrated a good match to our analytic modeling. Furthermore, we implemented both feedforward and Hopfield networks by utilizing the proposed neuromorphic design. The system performance and robustness were studied through massive Monte-Carlo simulations based on the application of digital image recognition. Comparing to the previous crossbar-based computing engine that represents data with voltage amplitude, our design can achieve >50% energy savings, while the average probability of failed recognition increase only 1.46% and 5.99% in the feedforward and Hopfield implementations, respectively.
Bonan Yan, Chaofei Yang, Linghao Song, Beiye Liu, Yiran Chen 0001, Hai Li 0001, Qing Wu 0002, Hao Jiang 0014
DAC2
2015 Giant spin hall effect (GSHE) logic design for low power application
Yaojun Zhang, Bonan Yan, Hai Li 0001, Yiran Chen 0001
DATE2
2015 A High-Speed Robust NVM-TCAM Design Using Body Bias Feedback
abstract
As manufacture process scales down rapidly, the design of ternary content-addressable memory (TCAM) requiring high storage density, fast access speed and low power consumption becomes very challenging. In recent years, many novel TCAM designs have been inspired by the research on emerging nonvolatile memory technologies, such as magnetic tunneling junction (MTJ), phase change memory (PCM), and memristor. These designs store a data as the resistive variable of a nonvolatile device, which usually results in limited sensing margin and therefore constrains the searching speed of TCAM architecture severely. To further enhance the performance and robustness of TCAMs, we proposed two novel cell designs that utilize MTJs as data storage units - the symmetrical dual-N structure and the asymmetrical P-N scheme. In both designs, a body bias feedback circuit is integrated to enlarge the sensing margins. Compared with an existing MTJ-based TCAM structure, the tolerance in gate voltage variation of the symmetrical dua-N (asymmetrical P-N) scheme can significantly improve 59.5% (21.2%). The latency and the dynamic energy consumption in one searching operation at the word length of 256 bits are merely 590.35ps (97.89ps) and 65.05fJ/bit (36.85fJ/bit), not even mentioning that the use of nonvolatile MTJ devices avoids unnecessary leakage power consumption.
Bonan Yan, Yaojun Zhang, Jianlei Yang 0001, Hai Li 0001, Weisheng Zhao 0001, Pierre Chor-Fung Chia
ACM Great Lakes Symposium on VLSI1
2015 A new self-reference sensing scheme for TLC MRAM
abstract
Density is one of the major design factors of magnetic random access memory (MRAM). Very recently, a tri-level cell (TLC) structure was proposed to enhance the storage density of MRAM. In this work, we propose a new self-reference sensing scheme for the TLC MRAM cell based on its unique property called state ordering. Simulation results show that compared to conventional design, our proposed self-reference scheme achieves on average 61% saving on sensing delay while also demonstrating significantly enhanced resilience to device parametric variations.
Bonan Yan, Lun Yang, Weisheng Zhao 0001, Yiran Chen 0001, Hai Li 0001
ISCAS2
2015 An overview on memristor crossabr based neuromorphic circuit and architecture
abstract
As technology advances, artificial intelligence becomes pervasive in society and ubiquitous in our lives, which stimulates the desire for embedded-everywhere and human-centric intelligent computation paradigm. However, conventional instruction-based computer architecture was designed for algorithmic and exact calculations. It is not suitable for handling the applications of machine learning and neural networks that usually involve a large sets of noisy and incomplete natural data. Instead, neuromorphic systems inspired by the working mechanism of human brains create promising potential. Neuromorphic systems possess a massively parallel architecture with closely coupled memory and computing. Moreover, through the sparse utilizations of hardware resources in time and space, extremely high power efficiency can be achieved. In recent years, the use of memristor technology in neuromorphic systems has attracted growing attention for its distinctive properties, such as nonvolatility, reconfigurability, and analog processing capability. In this paper, we summarize the research efforts in the development of memristor crossbar based neuromorphic design from the perspectives of device modeling, circuit, architecture, and design automation.
Yandan Wang, Bonan Yan, Chaofei Yang, Jianlei Yang 0001, Hai Li 0001
VLSI-SoC4