EDBT 2026 Demo / reviewers in the wild / expert
Bin Gao 0006
dblp:181/2330-6
· DBLP profile ↗
20ranked-venue papers
0as first author
11since 2021 · last 2026
0000-0002-2417-983XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MiniBuf: An On-Chip Buffer Allocation Framework Toward Minimizing Buffer Size and Latency for Memristor-Based CNN AcceleratorabstractMemristor-based Convolutional Neural Network (CNN) accelerators have gained considerable attention due to their low latency and high energy efficiency, making them promising candidates for edge acceleration. Alongside statically stored model weights, dynamically generated intermediate feature maps during inference occupy a significant portion of the on-chip buffer capacity and directly affect the efficiency of hardware pipeline execution. However, there is a lack of theoretical analysis and methods for efficiently allocating on-chip buffer for feature map data. To address this gap, this paper implements three innovative aspects. Firstly, a mathematical model is developed to estimate the minimal buffer size required for pipelined inference of CNNs on memristor-based accelerators, offering accurate and swift evaluations of the buffer size requirement. Secondly, based on this model, the paper establishes mathematical conditions for buffer requirements to maintain a blocking-free pipeline during CNN inference, providing theoretical guidance for on-chip buffer allocation strategies. Thirdly, a simulation-in-loop optimization method is proposed to further reduce latency by efficiently increasing the buffer size of critical layers. To validate our proposed model and method, evaluations were conducted on five representative models: ResNet-18, ResNet-50, YOLO-v5, U-Net, and Faster-RCNN-FPN. The results reveal a remarkably low average estimation error of only 2.6% between the mathematical model and the experimentally measured results, with the maximum error still below 10%. Moreover, our simulation-in-loop optimization strategy achieved significant latency reductions ranging from 5.3% to 57.5% across the five models. Ruihua Yu, Chenhuan Zou, Jiaming Li 0006, Bin Gao 0006, Jianshi Tang, He Qian, Huaqiang Wu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2024 | Compensation Architecture to Alleviate Noise Effects in RRAM-based Computing-in-memory Chips with Residual ResourceabstractResistive random access memory (RRAM) is a promising technology for energy-efficient in-memory computing. However, due to technology limits, RRAM device faces a series of reliability issues. Deep neural network (DNN) computing based on RRAM suffers from accuracy degradation. On the one hand, offline DNN training solutions are difficult to fully consider and simulate all nonidealities. Worse still, new error or nonideality may come up with the usage of RRAM, which further deteriorates the effectiveness of offline training. On the other hand, online training poses great challenges on programming overhead and device lifetime. The iterative write-verify technique to program multi-bit RRAM cells prolongs write latency more than 10× longer than read latency. To overcome these issues, we propose a compensation architecture and a software and hardware co-training design to mitigate the realistic network accuracy loss in RRAM-based computing-in-memory chips. Firstly, we add trainable compensation channels in crossbars utilizing the residual resource after original weight mapping. Secondly, an offline training procedure with computing output from hardware is triggered to settle down appropriate weight value in compensation channels. Experimental results demonstrate that the proposed design can guarantee ≤ 0.8% loss of accuracy in DNN on MNIST and CIFAR10 dataset even when nonidealities reduce the original accuracy down to ≤73%. Longjun Liu, Yuyi Liu, Bin Gao 0006, Hongbin Sun 0001 |
ISCAS | 4 |
| 2023 | Architecture-circuit-technology co-optimization for resistive random access memory-based computation-in-memory chips
Yuyi Liu, Bin Gao 0006, Jianshi Tang, Huaqiang Wu, He Qian |
Sci. China Inf. Sci. | 2 |
| 2023 | CLEAR: a full-stack chip-in-loop emulator for analog RRAM based computing-in-memory system
Ruihua Yu, Bin Gao 0006, Yiwen Geng, Yuyi Liu, Qingtian Zhang, Jianshi Tang, Hu He 0001, Ning Deng 0008, He Qian, Huaqiang Wu |
Sci. China Inf. Sci. | 3 |
| 2023 | An Error-Free 64KB ReRAM-Based nvSRAM Integrated to a Microcontroller Unit Supporting Real-Time Program Storage and RestorationabstractNonvolatile SRAM (nvSRAM), which integrates the nonvolatile elements with SRAM using a direct bit-to-bit connection has raised much attention in the past few years, owing to its fast parallel data transfer and fast power-on/off speed. However, few nvSRAM macros have been silicon verified to be enacted through the power-failure event. On the other hand, the capacity of fabricated nvSRAM macro is small (~ Kbit) to date, inhibiting its practical application. This study presents a novel ReRAM-based nvSRAM bitcell with improved reliability and scalability. A 64KB nvSRAM macro was designed and integrated into a 32-bit microcontroller unit (MCU). The chip was fabricated using HfOx-based BEOL ReRAM and a 130nm CMOS technology. To pursue fast storage, a write-without-verify scheme is adopted to program ReRAM, measurement results show that the raw bit error rate between the power outages is < 0.1% for the full macro under such constraint. Cryptography and machine learning applications are successfully performed on the MCU system. For the first time, with the help of correction techniques, we achieved an error-free nvSRAM macro that is reliable enough to store/restore programs and demonstrated a real-time robotic control system empowered by the nvSRAM. The proposed nvSRAM macro has the largest capacity to date. Hanwen Gong, Hu He 0001, Liyang Pan, Bin Gao 0006, Jianshi Tang, Sining Pan, Dabin Wu, He Qian, Huaqiang Wu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2022 | Complementary Memtransistor-Based Multilayer Neural Networks for Online Supervised Learning Through (Anti-)Spike-Timing-Dependent PlasticityabstractWe propose a complete hardware-based architecture of multilayer neural networks (MNNs), including electronic synapses, neurons, and periphery circuitry to implement supervised learning (SL) algorithm of extended remote supervised method (ReSuMe). In this system, complementary (a pair of n- and p-type) memtransistors (C-MTs) are used as an electrical synapse. By applying the learning rule of spike-timing-dependent plasticity (STDP) to the memtransistor connecting presynaptic neuron to the output one whereas the contrary anti-STDP rule to the other memtransistor connecting presynaptic neuron to the teacher one, extended ReSuMe with multiple layers is realized without the usage of those complicated supervising modules in previous approaches. In this way, both the C-MT-based chip area and power consumption of the learning circuit for weight updating operation are drastically decreased comparing with the conventional single memtransistor (S-MT)-based designs. Two typical benchmarks, the linearly nonseparable benchmark XOR problem and Mixed National Institute of Standards and Technology database (MNIST) recognition have been successfully tackled using the proposed MNN system while impact of the nonideal factors of realistic devices has been evaluated. Nuo Xu 0002, Bin Gao 0006, Fuwei Zhuge, Zijian Tang, Xinchen Deng, Yi Li 0049, Yuhui He, Xiangshui Miao |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | An On-chip Layer-wise Training Method for RRAM based Computing-in-memory ChipsabstractRRAM-based computing-in-memory (CIM) chips have shown great potentials to accelerate deep neural networks on edge devices by reducing data transfer between the memory and the computing unit. However, due to the non-ideal characteristics of RRAM, the accuracy of the neural network on the RRAM chip is usually lower than the software. Here we propose an on-chip layer-wise training (LWT) method to alleviate the adverse effect of RRAM imperfections and improve the accuracy of the chip. Using a locally validated dataset, LWT can reduce the communication between the edge and the cloud, which benefits personalized data privacy. The simulation results on the CIFAR-10 dataset show that the LWT method can improve the accuracy of VGG-16 and ResNet-18 by more than 5% and 10%, respectively, with only 25% operations and 35% buffer compared with the back-propagation method. Moreover, the pipe-LWT method is presented to improve the throughput by three times further. Yiwen Geng, Bin Gao 0006, Qingtian Zhang, Yudeng Lin, Jianshi Tang, Huaqiang Wu, He Qian |
DATE | 2 |
| 2021 | Recent progress of integrated circuits and optoelectronic chips
Yue Hao 0001, Genquan Han, Jincheng Zhang 0001, Xiaohua Ma 0001, Zhangming Zhu, Yanan Han, Ling Yang 0003, Jiangyi Shi, Wei Zhang 0343, Biao Pan, Yangqi Huang, Qi Liu 0010, Yimao Cai, Xin Ou, Tiangui You, Huaqiang Wu, Bin Gao 0006, Guoping Guo, Yonghua Chen, Xiangfei Chen, Chunlai Xue, Lixia Zhao, Xihua Zou, Lianshan Yan |
Sci. China Inf. Sci. | 26 |
| 2021 | Array-level boosting method with spatial extended allocation to improve the accuracy of memristor based computing-in-memory chips
Bin Gao 0006, Jianshi Tang, He Qian, Huaqiang Wu |
Sci. China Inf. Sci. | 2 |
| 2021 | In-memory Learning with Analog Resistive Switching Memory: A Review and PerspectiveabstractIn this article, we review the existing analog resistive switching memory (RSM) devices and their hardware technologies for in-memory learning, as well as their challenges and prospects. Since the characteristics of the devices are different for in-memory learning and digital memory applications, it is important to have an in-depth understanding across different layers from devices and circuits to architectures and algorithms. First, based on a top-down view from architecture to devices for analog computing, we define the main figures of merit (FoMs) and perform a comprehensive analysis of analog RSM hardware including the basic device characteristics, hardware algorithms, and the corresponding mapping methods for device arrays, as well as the architecture and circuit design considerations for neural networks. Second, we classify the FoMs of analog RSM devices into two levels. Level 1 FoMs are essential for achieving the functionality of a system (e.g., linearity, symmetry, dynamic range, level numbers, fluctuation, variability, and yield). Level 2 FoMs are those that make a functional system more efficient and reliable (e.g., area, operational voltage, energy consumption, speed, endurance, retention, and compatibility with back-end-of-line processing). By constructing a device-to-application simulation framework, we perform an in-depth analysis of how these FoMs influence in-memory learning and give a target list of the device requirements. Lastly, we evaluate the main FoMs of most existing devices with analog characteristics and review optimization methods from programming schemes to materials and device structures. The key challenges and prospects from the device to system level for analog RSM devices are discussed. Bin Gao 0006, Jianshi Tang, Meng-Fan Chang, Xiaobo Sharon Hu, Jan Van der Spiegel, He Qian, Huaqiang Wu |
Proc. IEEE | 2 |
| 2021 | Diagonal Matrix Regression Layer: Training Neural Networks on Resistive Crossbars With Interconnect Resistance EffectabstractResistive crossbars implement parallel vector-matrix multiplication (VMM) in analog fashion, and thus enable fast and energy-efficient neuromorphic systems. However, interconnect resistance and resistive switching devices form a complex resistance network with sneak paths. It could result in severe distortions on the output currents. When implementing neural networks, current distortions also cause significant accuracy loss. This article proposes an accurate and computationally efficient model of VMM in resistive crossbars, called diagonal matrix regression (DMR), and incorporates the model into the topology of neural networks as DMR layer (DMRL). Given an m×n crossbar, two diagonal matrices are calculated directly according to the resistance network in a time complexity of only O(m2+n2). No hyper-parameter needs to be determined manually. Modeling of VMM is implemented in a time complexity of only O(mn). DMRL is developed to replace the weight matrix of neural networks so that the effect of interconnect resistance and the sneak path problem are well handled during ex-situ training. Using this technique, for the task of MNIST and fashion-MNIST classification, the accuracy is dramatically restored. Yan Liao, Bin Gao 0006, Jianshi Tang, Huaqiang Wu, He Qian |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2019 | Design Guidelines of RRAM based Neural-Processing-Unit: A Joint Device-Circuit-Algorithm AnalysisabstractRRAM based neural-processing-unit (NPU) is emerging for processing general purpose machine intelligence algorithms with ultra-high energy efficiency, while the imperfections of the analog devices and cross-point arrays make the practical application more complicated. In order to improve accuracy and robustness of the NPU, device-circuit-algorithm codesign with consideration of underlying device and array characteristics should outperform the optimization of individual device or algorithm. In this work, we provide a joint device-circuit-algorithm analysis and propose the corresponding design guidelines. Key innovations include: 1) An end-to-end simulator for RRAM NPU is developed with an integrated framework from device to algorithm. 2) The complete design of circuit and architecture for RRAM NPU is provided to make the analysis much close to the real prototype. 3) A large-scale neural network as well as other general-purpose networks are processed for the study of device-circuit interaction. 4) Accuracy loss from non-idealities of RRAM, such as I-V nonlinearity, noises of analog resistance levels, voltage-drop for interconnect, ADC/DAC precision, are evaluated for the NPU design. Xiaochen Peng, Huaqiang Wu, Bin Gao 0006, Hu He 0001, Youhui Zhang, Shimeng Yu, He Qian |
DAC | 4 |
| 2019 | Three-Dimensional nand Flash for Vector-Matrix MultiplicationabstractThree-Dimensional NAND flash technology is one of the most competitive integrated solutions for the high-volume massive data storage. So far, there are few investigations on how to use 3-D NAND flash for in-memory computing in the neural network accelerator. In this brief, we propose using the 3-D vertical channel NAND array architecture to implement the vector-matrix multiplication (VMM) with for the first time. Based on the array-level SPICE simulation, the bias condition including the selector layer and the unselected layers is optimized to achieve high computation accuracy of VMM. Since the VMM can be performed layer by layer in a 3-D NAND array, the read-out latency is largely improved compared to the conventional single-cell read-out operation. The impact of device-to-device variation on the computation accuracy is also analyzed. Panni Wang, Bo Wang 0067, Bin Gao 0006, Huaqiang Wu, He Qian, Shimeng Yu |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2018 | Sign backpropagation: An on-chip learning algorithm for analog RRAM neuromorphic computing systemsabstractCurrently, powerful deep learning models usually require significant resources in the form of processors and memory, which leads to very high energy consumption. The emerging resistive random access memory (RRAM) has shown great potential for constructing a scalable and energy-efficient neural network. However, it is hard to port a high-precision neural network from conventional digital CMOS hardware systems to analog RRAM systems owing to the variability of RRAM devices. A suitable on-chip learning algorithm should be developed to retrain or improve the performance of the neural network. In addition, determining how to integrate the periphery digital computations and analog RRAM crossbar is still a challenge. Here, we propose an on-chip learning algorithm, named sign backpropagation (SBP), for RRAM-based multilayer perceptron (MLP) with binary interfaces (0, 1) in forward process and 2-bit (±1, 0) in backward process. The simulation results show that the proposed method and architecture can achieve a comparable classification accuracy with MLP on MNIST dataset, meanwhile it can save area and energy cost by the calculation and storing of the intermediate results and take advantages of the RRAM crossbar potential in neuromorphic computing. Qingtian Zhang, Huaqiang Wu, Bin Gao 0006, Ning Deng 0008, He Qian |
Neural Networks | 5 |
| 2017 | Neuromorphic Computing based on Resistive RAMabstractResistive random access memory (RRAM) has gained significant attentions because of its excellent characteristics which are suitable for next-generation non-volatile memory applications. It is also very attractive to build neuromorphic computing chip based on RRAM cells due to non-volatile and analog properties. Neuromorphic computing hardware technologies using analog weight storage allow the scaling-up of the system size to complete cognitive tasks such as face classification much faster while consuming much lower energy. In this paper, RRAM technology development from material selection to device structure, from small array to full chip will be discussed in detail. Neuromorphic computing using RRAM devices is demonstrated, and speed & energy consumption are compared with Xeon Phi processor. Huaqiang Wu, Bin Gao 0006, He Qian |
ACM Great Lakes Symposium on VLSI | 3 |
| 2017 | Resistive Random Access Memory for Future Information Processing SystemabstractResistive random access memory (RRAM) is regarded as one of the most promising emerging memory technologies for next-generation embedded, standalone nonvolatile memory (NVM), and storage class memory (SCM) due to its speed, density, cost, and scalability. Considerable progress has been made in recent years on the manufacturability of RRAM, with low-density RRAM products now in production and the path to higher density parts becoming clearer. This review updates the learning on the fundamental materials and process integration needed for high-volume manufacturing and summarizes very recent progress on array level performance improvement methodology using novel techniques, and circuit level contributions for different applications. The device performance, array integration, and device/circuit codesign for memory systems are discussed. Novel applications besides embedded memory and standalone memory are addressed, including hardware security, neuromorphic computing, and nonvolatile logic systems. Huaqiang Wu, Xiao Hu Wang, Bin Gao 0006, Ning Deng 0008, Zhichao Lu, Brent Haukness, Gary Bronner, He Qian |
Proc. IEEE | 3 |
| 2015 | Modeling and design optimization of ReRAMabstractResistive switching memories (ReRAM) have been widely studied for applications in next-generation data storage and neurormorphic computing systems. To enable device-circuit-system co-design and optimization, a SPICE model of ReRAM that can reproduce the device characteristics in circuit simulations is needed. In this paper, we present a novel tool for ReRAM design including a physics-based SPICE model, the model parameters extraction strategy, as well as the system assessment method. This physics-based SPICE model can capture all the essential features of HfOx-based ReRAM including the DC/AC and multi-level switching behaviors, switching reliability, and intrinsic device variations. A strategy is developed to extract the critical model parameters from the fabricated ReRAM devices. A variety of electrical measurements on various ReRAMs are performed to verify and calibrate the model. The assessment method based on the experimentally verified SPICE model can be applied to explore a wide range of applications including: 1) variation-aware and reliability-emphasized system design; 2) system performance evaluation; 3) array architecture optimization. This verified design tool not only enables system design but also enables system optimization that capitalizes on device/circuit interaction for both data storage and neuromorphic computing applications. Jinfeng Kang, Haitong Li, Peng Huang 0004, Bin Gao 0006, Zizhen Jiang, H.-S. Philip Wong |
ASP-DAC | 5 |
| 2015 | Variation-aware, reliability-emphasized design and optimization of RRAM using SPICE model
Haitong Li, Zizhen Jiang, Peng Huang 0004, Hong-Yu Chen, Bin Gao 0006, Jinfeng Kang, H.-S. Philip Wong |
DATE | 6 |
| 2014 | Scaling and operation characteristics of HfOx based vertical RRAM for 3D cross-point architectureabstractStacked HfOxbased vertical RRAM with interface engineering for 3D cross-point architecture is fabricated using a cost-effective fabrication process. The excellent performances such as low reset current, fast switching speed, high switching endurance and disturbance immunity, good retention and self-selectivity are demonstrated in the fabricated HfOxbased vertical RRAM devices. The scaling limit and the functionality along with a viable write/read scheme of the presented vertical RRAM are investigated. The experiments show that the pillar electrode thickness and the plane electrode thickness of the vertical RRAM can be scaled down to 3nm and 5nm without significant performance degradation, respectively. Jinfeng Kang, Bin Gao 0006, Peng Huang 0004, Hong-Yu Chen, Zizhen Jiang, H.-S. Philip Wong, Shimeng Yu |
ISCAS | 2 |
| 2014 | Design guidelines for 3D RRAM cross-point architectureabstractDesign guidelines were proposed to evaluate and optimize the 3D RRAM cross-point architecture by a full-size 3D circuit simulation in SPICE. The performance metrics that were evaluated include the write/read margin, access latency, energy consumption per programming, and the density per bit. Different 3D cross-point architecture including the horizontally stacked or the vertically stacked structure were compared in terms of these metrics, revealing the advantages of the vertical RRAM structure. Then the scaling trend of the vertical RRAM based 3D array with respect to the scaling of lateral feature size, vertical electrode thickness and vertical isolation layer thickness were evaluated. The design parameters that affect the scaling trend include the metal interconnect resistance, RRAM on-state cell resistance (or the nonlinearity of the I-V). The design trade-offs are discussed considering those parameters constraints. Shimeng Yu, Yexin Deng, Bin Gao 0006, Peng Huang 0004, Jinfeng Kang, Hong-Yu Chen, Zizhen Jiang, H.-S. Philip Wong |
ISCAS | 3 |