EDBT 2026 Demo / reviewers in the wild / expert
Huaqiang Wu
dblp:149/5060
· DBLP profile ↗
19ranked-venue papers
2as first author
12since 2021 · last 2026
0000-0001-8359-7997ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MiniBuf: An On-Chip Buffer Allocation Framework Toward Minimizing Buffer Size and Latency for Memristor-Based CNN AcceleratorabstractMemristor-based Convolutional Neural Network (CNN) accelerators have gained considerable attention due to their low latency and high energy efficiency, making them promising candidates for edge acceleration. Alongside statically stored model weights, dynamically generated intermediate feature maps during inference occupy a significant portion of the on-chip buffer capacity and directly affect the efficiency of hardware pipeline execution. However, there is a lack of theoretical analysis and methods for efficiently allocating on-chip buffer for feature map data. To address this gap, this paper implements three innovative aspects. Firstly, a mathematical model is developed to estimate the minimal buffer size required for pipelined inference of CNNs on memristor-based accelerators, offering accurate and swift evaluations of the buffer size requirement. Secondly, based on this model, the paper establishes mathematical conditions for buffer requirements to maintain a blocking-free pipeline during CNN inference, providing theoretical guidance for on-chip buffer allocation strategies. Thirdly, a simulation-in-loop optimization method is proposed to further reduce latency by efficiently increasing the buffer size of critical layers. To validate our proposed model and method, evaluations were conducted on five representative models: ResNet-18, ResNet-50, YOLO-v5, U-Net, and Faster-RCNN-FPN. The results reveal a remarkably low average estimation error of only 2.6% between the mathematical model and the experimentally measured results, with the maximum error still below 10%. Moreover, our simulation-in-loop optimization strategy achieved significant latency reductions ranging from 5.3% to 57.5% across the five models. Ruihua Yu, Chenhuan Zou, Jiaming Li 0006, Bin Gao 0006, Jianshi Tang, He Qian, Huaqiang Wu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2025 | MTVSM-CIM: A Magnified-TMR VC-SOT-MRAM Computing-in-Memory Macro for Edge AIabstractPursuit of energy efficiency and accuracy in intelligent edge devices drives emerging non-volatile memories (NVM) developing more advanced computing-in-memory (CIM) architectures. The application of high performance magnetoresistive random-access memory (MRAM) in CIM has been limited due to its inherent non-ideal characteristics such as low resistance, low switching ratio, high area overhead and low accuracy. In this work, we introduce an innovative weighted two-transistor-one-MTJ (W-2T1R) VC-SOT-MRAM CIM macro that addresses these challenges with: 1) a novel 2T1R cell with high TMR; 2) an excellent approach to assign multi-bit array weight and the redundancy sub-track for linear promotion; 3) a bitwise input sparsity dataflow and architecture to improve energy and latency. Our proposal, excels at high linear computing, with energy efficiencies of 105.5 TOPS/W, area efficiency of 0.73 TOPS/mm2, throughput of 7.2 TOPS and classification accuracy of 91.46% on CIFAR-10 dataset for configurations comprising 6-bit input, 3-bit weight and 6-bit output on VGG-8 model. Bingqian Song, Cancheng Xiao, Fantao Gao, Mengzhu Li, Ziwei Han, Jianshi Tang, Huaqiang Wu, Tianxiang Nan |
ISCAS | 8 |
| 2024 | HXNOR-PBNN: A Scalable and Parallel Spintronics Synaptic Architecture for Probabilistic Binary Neural NetworksabstractThe combination of computing-in-memory (CIM) architecture and emerging non-volatile memory (NVM) is considered as a promising alternative for low-power electronics. Among different emerging NVMs, the insufficient conductance level of magnetoresistive random-access memory (MRAM) is a serious limitation to introduce its high performance into analogue CIM architecture. In this paper, we propose a probabilistic binary neural networks(PBNN) hardware based on voltage-controlled spin-orbit torque MRAM (VC-SOT-MRAM), which employs 2T2MTJ sub-tracks integrated of the probabilistic switching MTJs to achieve the bit-wise weights sampling and activation. Selective writing is discussed to co-optimized. A pipeline readout circuit with hybrid dynamic reference was suggested to improve the precision and throughput. In addition, a weight mapping method is applied for multi-tracks computing in parallel and kernel duplicating to improve the efficiency. Our proposal, simulated in 28 nm technology, achievean energy efficiency of 312.5 bTOPS/W, throughput of 527.3 bGOPS, accuracies of >90% on MNIST with more acceleration (58×), less resource requirement (32×) and high robust row-parallel computing. Cancheng Xiao, Dingsong Jiang, Jianle Liu, Bingqian Song, Jianshi Tang, Huaqiang Wu, Tianxiang Nan |
ISCAS | 7 |
| 2023 | Architecture-circuit-technology co-optimization for resistive random access memory-based computation-in-memory chips
Yuyi Liu, Bin Gao 0006, Jianshi Tang, Huaqiang Wu, He Qian |
Sci. China Inf. Sci. | 4 |
| 2023 | CLEAR: a full-stack chip-in-loop emulator for analog RRAM based computing-in-memory system
Ruihua Yu, Bin Gao 0006, Yiwen Geng, Yuyi Liu, Qingtian Zhang, Jianshi Tang, Hu He 0001, Ning Deng 0008, He Qian, Huaqiang Wu |
Sci. China Inf. Sci. | 13 |
| 2023 | An Error-Free 64KB ReRAM-Based nvSRAM Integrated to a Microcontroller Unit Supporting Real-Time Program Storage and RestorationabstractNonvolatile SRAM (nvSRAM), which integrates the nonvolatile elements with SRAM using a direct bit-to-bit connection has raised much attention in the past few years, owing to its fast parallel data transfer and fast power-on/off speed. However, few nvSRAM macros have been silicon verified to be enacted through the power-failure event. On the other hand, the capacity of fabricated nvSRAM macro is small (~ Kbit) to date, inhibiting its practical application. This study presents a novel ReRAM-based nvSRAM bitcell with improved reliability and scalability. A 64KB nvSRAM macro was designed and integrated into a 32-bit microcontroller unit (MCU). The chip was fabricated using HfOx-based BEOL ReRAM and a 130nm CMOS technology. To pursue fast storage, a write-without-verify scheme is adopted to program ReRAM, measurement results show that the raw bit error rate between the power outages is < 0.1% for the full macro under such constraint. Cryptography and machine learning applications are successfully performed on the MCU system. For the first time, with the help of correction techniques, we achieved an error-free nvSRAM macro that is reliable enough to store/restore programs and demonstrated a real-time robotic control system empowered by the nvSRAM. The proposed nvSRAM macro has the largest capacity to date. Hanwen Gong, Hu He 0001, Liyang Pan, Bin Gao 0006, Jianshi Tang, Sining Pan, Dabin Wu, He Qian, Huaqiang Wu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 11 |
| 2022 | Improving the accuracy of neural networks in analog computing-in-memory systems by analog weightabstractCrossbar-enabled analog computing-in-memory (CACIM) systems can significantly improve the computation speed and energy efficiency of deep neural networks (DNNs). However, an important issue is that the performance of DNNs degrades severely when deploying the DNNs onto the CACIM systems. Because the devices in the CACIM systems have low precision to present the weights, which is caused by the intrinsic variation and high programming overhead. The computational paradigms of the CACIM systems and the digital systems are essentially different. One of the main differences is that the weights are expressed in analog terms, and it has no encoding and decoding process during the computation. We can take advantage of the characteristic of data presentation to get better performance in limited data precision. A generalized quantization method that does not constrain the range of quanta and can obtain less quantization error will be effective in the CACIM systems. For the first time, we introduced a generalized quantization method into CACIM systems and showed superior performance on a series of computer vision tasks, such as image classification, object detection, and semantic segmentation. Using the generalized quantization method, the DNN with 8-level analog weights can outperform the 32-bit networks. With fewer levels, the generalized quantization method can obtain less accuracy loss than other uniform quantization methods. Lingjun Dai, Qingtian Zhang, Huaqiang Wu |
ICPR | 3 |
| 2021 | An On-chip Layer-wise Training Method for RRAM based Computing-in-memory ChipsabstractRRAM-based computing-in-memory (CIM) chips have shown great potentials to accelerate deep neural networks on edge devices by reducing data transfer between the memory and the computing unit. However, due to the non-ideal characteristics of RRAM, the accuracy of the neural network on the RRAM chip is usually lower than the software. Here we propose an on-chip layer-wise training (LWT) method to alleviate the adverse effect of RRAM imperfections and improve the accuracy of the chip. Using a locally validated dataset, LWT can reduce the communication between the edge and the cloud, which benefits personalized data privacy. The simulation results on the CIFAR-10 dataset show that the LWT method can improve the accuracy of VGG-16 and ResNet-18 by more than 5% and 10%, respectively, with only 25% operations and 35% buffer compared with the back-propagation method. Moreover, the pipe-LWT method is presented to improve the throughput by three times further. Yiwen Geng, Bin Gao 0006, Qingtian Zhang, Yudeng Lin, Jianshi Tang, Huaqiang Wu, He Qian |
DATE | 10 |
| 2021 | Recent progress of integrated circuits and optoelectronic chips
Yue Hao 0001, Genquan Han, Jincheng Zhang 0001, Xiaohua Ma 0001, Zhangming Zhu, Yanan Han, Ling Yang 0003, Jiangyi Shi, Wei Zhang 0343, Biao Pan, Yangqi Huang, Qi Liu 0010, Yimao Cai, Xin Ou, Tiangui You, Huaqiang Wu, Bin Gao 0006, Guoping Guo, Yonghua Chen, Xiangfei Chen, Chunlai Xue, Lixia Zhao, Xihua Zou, Lianshan Yan |
Sci. China Inf. Sci. | 25 |
| 2021 | Array-level boosting method with spatial extended allocation to improve the accuracy of memristor based computing-in-memory chips
Bin Gao 0006, Jianshi Tang, He Qian, Huaqiang Wu |
Sci. China Inf. Sci. | 6 |
| 2021 | In-memory Learning with Analog Resistive Switching Memory: A Review and PerspectiveabstractIn this article, we review the existing analog resistive switching memory (RSM) devices and their hardware technologies for in-memory learning, as well as their challenges and prospects. Since the characteristics of the devices are different for in-memory learning and digital memory applications, it is important to have an in-depth understanding across different layers from devices and circuits to architectures and algorithms. First, based on a top-down view from architecture to devices for analog computing, we define the main figures of merit (FoMs) and perform a comprehensive analysis of analog RSM hardware including the basic device characteristics, hardware algorithms, and the corresponding mapping methods for device arrays, as well as the architecture and circuit design considerations for neural networks. Second, we classify the FoMs of analog RSM devices into two levels. Level 1 FoMs are essential for achieving the functionality of a system (e.g., linearity, symmetry, dynamic range, level numbers, fluctuation, variability, and yield). Level 2 FoMs are those that make a functional system more efficient and reliable (e.g., area, operational voltage, energy consumption, speed, endurance, retention, and compatibility with back-end-of-line processing). By constructing a device-to-application simulation framework, we perform an in-depth analysis of how these FoMs influence in-memory learning and give a target list of the device requirements. Lastly, we evaluate the main FoMs of most existing devices with analog characteristics and review optimization methods from programming schemes to materials and device structures. The key challenges and prospects from the device to system level for analog RSM devices are discussed. Bin Gao 0006, Jianshi Tang, Meng-Fan Chang, Xiaobo Sharon Hu, Jan Van der Spiegel, He Qian, Huaqiang Wu |
Proc. IEEE | 9 |
| 2021 | Diagonal Matrix Regression Layer: Training Neural Networks on Resistive Crossbars With Interconnect Resistance EffectabstractResistive crossbars implement parallel vector-matrix multiplication (VMM) in analog fashion, and thus enable fast and energy-efficient neuromorphic systems. However, interconnect resistance and resistive switching devices form a complex resistance network with sneak paths. It could result in severe distortions on the output currents. When implementing neural networks, current distortions also cause significant accuracy loss. This article proposes an accurate and computationally efficient model of VMM in resistive crossbars, called diagonal matrix regression (DMR), and incorporates the model into the topology of neural networks as DMR layer (DMRL). Given an m×n crossbar, two diagonal matrices are calculated directly according to the resistance network in a time complexity of only O(m2+n2). No hyper-parameter needs to be determined manually. Modeling of VMM is implemented in a time complexity of only O(mn). DMRL is developed to replace the weight matrix of neural networks so that the effect of interconnect resistance and the sneak path problem are well handled during ex-situ training. Using this technique, for the task of MNIST and fashion-MNIST classification, the accuracy is dramatically restored. Yan Liao, Bin Gao 0006, Jianshi Tang, Huaqiang Wu, He Qian |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2019 | Design Guidelines of RRAM based Neural-Processing-Unit: A Joint Device-Circuit-Algorithm AnalysisabstractRRAM based neural-processing-unit (NPU) is emerging for processing general purpose machine intelligence algorithms with ultra-high energy efficiency, while the imperfections of the analog devices and cross-point arrays make the practical application more complicated. In order to improve accuracy and robustness of the NPU, device-circuit-algorithm codesign with consideration of underlying device and array characteristics should outperform the optimization of individual device or algorithm. In this work, we provide a joint device-circuit-algorithm analysis and propose the corresponding design guidelines. Key innovations include: 1) An end-to-end simulator for RRAM NPU is developed with an integrated framework from device to algorithm. 2) The complete design of circuit and architecture for RRAM NPU is provided to make the analysis much close to the real prototype. 3) A large-scale neural network as well as other general-purpose networks are processed for the study of device-circuit interaction. 4) Accuracy loss from non-idealities of RRAM, such as I-V nonlinearity, noises of analog resistance levels, voltage-drop for interconnect, ADC/DAC precision, are evaluated for the NPU design. Xiaochen Peng, Huaqiang Wu, Bin Gao 0006, Hu He 0001, Youhui Zhang, Shimeng Yu, He Qian |
DAC | 3 |
| 2019 | On-Chip Analog Trojan Detection Framework for Microprocessor TrustworthinessabstractWith the globalization of semiconductor industry, hardware security issues have been gaining increasing attention. Among all hardware security threats, the insertion of hardware Trojans is one of the main concerns. Meanwhile, many current Trojan detection solutions follow the assumption that the hardware Trojan itself should be composed of digital logic. This assumption is invalidated by recently proposed analog Trojans which are extremely small and can detect rare events. This paper proposes a runtime hardware Trojan detection method which is geared toward detecting such advanced Trojans. The principle of this method is to guard a set of concerned signals, and initiate a hardware interrupt request when abnormal toggling events occur in these guarded signals. To prove the effectiveness of this method, we design a processor based on ARMv7-A&R ISA, and insert an analog Trojan into the processor. We fabricated the design in an SMIC 130-nm process and demonstrate the effectiveness of the proposed methodology. Yumin Hou, Hu He 0001, Kaveh Shamsi, Yier Jin, Huaqiang Wu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2019 | Three-Dimensional nand Flash for Vector-Matrix MultiplicationabstractThree-Dimensional NAND flash technology is one of the most competitive integrated solutions for the high-volume massive data storage. So far, there are few investigations on how to use 3-D NAND flash for in-memory computing in the neural network accelerator. In this brief, we propose using the 3-D vertical channel NAND array architecture to implement the vector-matrix multiplication (VMM) with for the first time. Based on the array-level SPICE simulation, the bias condition including the selector layer and the unselected layers is optimized to achieve high computation accuracy of VMM. Since the VMM can be performed layer by layer in a 3-D NAND array, the read-out latency is largely improved compared to the conventional single-cell read-out operation. The impact of device-to-device variation on the computation accuracy is also analyzed. Panni Wang, Bo Wang 0067, Bin Gao 0006, Huaqiang Wu, He Qian, Shimeng Yu |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2018 | Sign backpropagation: An on-chip learning algorithm for analog RRAM neuromorphic computing systemsabstractCurrently, powerful deep learning models usually require significant resources in the form of processors and memory, which leads to very high energy consumption. The emerging resistive random access memory (RRAM) has shown great potential for constructing a scalable and energy-efficient neural network. However, it is hard to port a high-precision neural network from conventional digital CMOS hardware systems to analog RRAM systems owing to the variability of RRAM devices. A suitable on-chip learning algorithm should be developed to retrain or improve the performance of the neural network. In addition, determining how to integrate the periphery digital computations and analog RRAM crossbar is still a challenge. Here, we propose an on-chip learning algorithm, named sign backpropagation (SBP), for RRAM-based multilayer perceptron (MLP) with binary interfaces (0, 1) in forward process and 2-bit (±1, 0) in backward process. The simulation results show that the proposed method and architecture can achieve a comparable classification accuracy with MLP on MNIST dataset, meanwhile it can save area and energy cost by the calculation and storing of the intermediate results and take advantages of the RRAM crossbar potential in neuromorphic computing. Qingtian Zhang, Huaqiang Wu, Bin Gao 0006, Ning Deng 0008, He Qian |
Neural Networks | 2 |
| 2017 | Neuromorphic Computing based on Resistive RAMabstractResistive random access memory (RRAM) has gained significant attentions because of its excellent characteristics which are suitable for next-generation non-volatile memory applications. It is also very attractive to build neuromorphic computing chip based on RRAM cells due to non-volatile and analog properties. Neuromorphic computing hardware technologies using analog weight storage allow the scaling-up of the system size to complete cognitive tasks such as face classification much faster while consuming much lower energy. In this paper, RRAM technology development from material selection to device structure, from small array to full chip will be discussed in detail. Neuromorphic computing using RRAM devices is demonstrated, and speed & energy consumption are compared with Xeon Phi processor. Huaqiang Wu, Bin Gao 0006, He Qian |
ACM Great Lakes Symposium on VLSI | 2 |
| 2017 | Resistive Random Access Memory for Future Information Processing SystemabstractResistive random access memory (RRAM) is regarded as one of the most promising emerging memory technologies for next-generation embedded, standalone nonvolatile memory (NVM), and storage class memory (SCM) due to its speed, density, cost, and scalability. Considerable progress has been made in recent years on the manufacturability of RRAM, with low-density RRAM products now in production and the path to higher density parts becoming clearer. This review updates the learning on the fundamental materials and process integration needed for high-volume manufacturing and summarizes very recent progress on array level performance improvement methodology using novel techniques, and circuit level contributions for different applications. The device performance, array integration, and device/circuit codesign for memory systems are discussed. Novel applications besides embedded memory and standalone memory are addressed, including hardware security, neuromorphic computing, and nonvolatile logic systems. Huaqiang Wu, Xiao Hu Wang, Bin Gao 0006, Ning Deng 0008, Zhichao Lu, Brent Haukness, Gary Bronner, He Qian |
Proc. IEEE | 1 |
| 2014 | Stack engineering for ReRAM devices performance improvementabstractAl/W:AlOx/WOy/W and Pt/AlOδ/Ta2O5-x/TaOy/Pt multiple layers ReRAM devices have been fabricated and carefully studied. Experimental results exhibit significant performance improvement through the insertion of AlOxlayer between the switching layer and the top electrode. Operation current is remarkably reduced, ON/OFF ratio is greatly increased, and stable multi-level operations have been successfully achieved. Multiple layers stack engineering has been proved as an efficient method to improve the performances of ReRAM devices. Huaqiang Wu, Minghao Wu, Zhiping Yu, He Qian |
ISCAS | 1 |