EDBT 2026 Demo / reviewers in the wild / expert
He Qian
dblp:04/4991
· DBLP profile ↗
19ranked-venue papers
0as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 4 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MiniBuf: An On-Chip Buffer Allocation Framework Toward Minimizing Buffer Size and Latency for Memristor-Based CNN AcceleratorabstractMemristor-based Convolutional Neural Network (CNN) accelerators have gained considerable attention due to their low latency and high energy efficiency, making them promising candidates for edge acceleration. Alongside statically stored model weights, dynamically generated intermediate feature maps during inference occupy a significant portion of the on-chip buffer capacity and directly affect the efficiency of hardware pipeline execution. However, there is a lack of theoretical analysis and methods for efficiently allocating on-chip buffer for feature map data. To address this gap, this paper implements three innovative aspects. Firstly, a mathematical model is developed to estimate the minimal buffer size required for pipelined inference of CNNs on memristor-based accelerators, offering accurate and swift evaluations of the buffer size requirement. Secondly, based on this model, the paper establishes mathematical conditions for buffer requirements to maintain a blocking-free pipeline during CNN inference, providing theoretical guidance for on-chip buffer allocation strategies. Thirdly, a simulation-in-loop optimization method is proposed to further reduce latency by efficiently increasing the buffer size of critical layers. To validate our proposed model and method, evaluations were conducted on five representative models: ResNet-18, ResNet-50, YOLO-v5, U-Net, and Faster-RCNN-FPN. The results reveal a remarkably low average estimation error of only 2.6% between the mathematical model and the experimentally measured results, with the maximum error still below 10%. Moreover, our simulation-in-loop optimization strategy achieved significant latency reductions ranging from 5.3% to 57.5% across the five models. Ruihua Yu, Chenhuan Zou, Jiaming Li 0006, Bin Gao 0006, Jianshi Tang, He Qian, Huaqiang Wu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2026 | Phase Shift Design for Multiple Incident Beams in Low Complexity RIS Under 5G Commercial Networks: Simulation and MeasurementabstractWith the capability to configure wireless propagation environment, reconfigurable intelligent surface (RIS) has attracted wide attention from both academia and industry, in which measurement campaign of RIS in commercial networks is of great importance for performance evaluation. However, existing RIS configuration schemes in practical environments are generally angle-based which require accurate angle information and mainly consider single incident beam, or statistical methods with high sampling overhead. In this paper, we propose a phase shift design scheme for RIS referred to as multiple incident beam superposition (MIBS) scheme, which can be applied in scenarios with incident signals on the RIS from multiple directions. Instead of utilizing random sampling, the proposed scheme firstly employs an incident beam scanning process to extract the incident beam information by pre-calculated codebook to reduce sampling cost, and requires no complex channel estimation or prior channel information. Then the proposed scheme concentrates the incident beams toward the desired reflection direction through a phase shift superposition algorithm. Numerical simulations verify the excellent performance and a 90% sample reduction of the proposed scheme compared with existing methods under the condition of 1-bit RIS hardware for practical applications. Furthermore, a measurement campaign in 5G commercial networks is conducted to validate the advantages of the proposed MIBS scheme in optimizing crucial signal metrics, yielding a 6.76 dB RSRP gain, a 4.9 dB SINR gain and a 37% throughput improvement, showing great potential of RIS for coverage enhancement. Wankai Tang, He Qian, Weicong Chen 0001, Xin Su 0010, Yifei Yuan 0003, Xiao Li 0001, Shi Jin 0002, Tiejun Cui |
IEEE Trans. Wirel. Commun. | 3 |
| 2025 | A survey on machine learning methods for food safety risk assessment: Approaches, challenges, and future outlook
Zhiyao Zhao, Bojian Qi, Nuo Duan, He Qian |
Eng. Appl. Artif. Intell. | 5 |
| 2023 | Architecture-circuit-technology co-optimization for resistive random access memory-based computation-in-memory chips
Yuyi Liu, Bin Gao 0006, Jianshi Tang, Huaqiang Wu, He Qian |
Sci. China Inf. Sci. | 5 |
| 2023 | CLEAR: a full-stack chip-in-loop emulator for analog RRAM based computing-in-memory system
Ruihua Yu, Bin Gao 0006, Yiwen Geng, Yuyi Liu, Qingtian Zhang, Jianshi Tang, Hu He 0001, Ning Deng 0008, He Qian, Huaqiang Wu |
Sci. China Inf. Sci. | 12 |
| 2023 | An Error-Free 64KB ReRAM-Based nvSRAM Integrated to a Microcontroller Unit Supporting Real-Time Program Storage and RestorationabstractNonvolatile SRAM (nvSRAM), which integrates the nonvolatile elements with SRAM using a direct bit-to-bit connection has raised much attention in the past few years, owing to its fast parallel data transfer and fast power-on/off speed. However, few nvSRAM macros have been silicon verified to be enacted through the power-failure event. On the other hand, the capacity of fabricated nvSRAM macro is small (~ Kbit) to date, inhibiting its practical application. This study presents a novel ReRAM-based nvSRAM bitcell with improved reliability and scalability. A 64KB nvSRAM macro was designed and integrated into a 32-bit microcontroller unit (MCU). The chip was fabricated using HfOx-based BEOL ReRAM and a 130nm CMOS technology. To pursue fast storage, a write-without-verify scheme is adopted to program ReRAM, measurement results show that the raw bit error rate between the power outages is < 0.1% for the full macro under such constraint. Cryptography and machine learning applications are successfully performed on the MCU system. For the first time, with the help of correction techniques, we achieved an error-free nvSRAM macro that is reliable enough to store/restore programs and demonstrated a real-time robotic control system empowered by the nvSRAM. The proposed nvSRAM macro has the largest capacity to date. Hanwen Gong, Hu He 0001, Liyang Pan, Bin Gao 0006, Jianshi Tang, Sining Pan, Dabin Wu, He Qian, Huaqiang Wu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 10 |
| 2021 | An On-chip Layer-wise Training Method for RRAM based Computing-in-memory ChipsabstractRRAM-based computing-in-memory (CIM) chips have shown great potentials to accelerate deep neural networks on edge devices by reducing data transfer between the memory and the computing unit. However, due to the non-ideal characteristics of RRAM, the accuracy of the neural network on the RRAM chip is usually lower than the software. Here we propose an on-chip layer-wise training (LWT) method to alleviate the adverse effect of RRAM imperfections and improve the accuracy of the chip. Using a locally validated dataset, LWT can reduce the communication between the edge and the cloud, which benefits personalized data privacy. The simulation results on the CIFAR-10 dataset show that the LWT method can improve the accuracy of VGG-16 and ResNet-18 by more than 5% and 10%, respectively, with only 25% operations and 35% buffer compared with the back-propagation method. Moreover, the pipe-LWT method is presented to improve the throughput by three times further. Yiwen Geng, Bin Gao 0006, Qingtian Zhang, Yudeng Lin, Jianshi Tang, Huaqiang Wu, He Qian |
DATE | 11 |
| 2021 | Array-level boosting method with spatial extended allocation to improve the accuracy of memristor based computing-in-memory chips
Bin Gao 0006, Jianshi Tang, He Qian, Huaqiang Wu |
Sci. China Inf. Sci. | 5 |
| 2021 | In-memory Learning with Analog Resistive Switching Memory: A Review and PerspectiveabstractIn this article, we review the existing analog resistive switching memory (RSM) devices and their hardware technologies for in-memory learning, as well as their challenges and prospects. Since the characteristics of the devices are different for in-memory learning and digital memory applications, it is important to have an in-depth understanding across different layers from devices and circuits to architectures and algorithms. First, based on a top-down view from architecture to devices for analog computing, we define the main figures of merit (FoMs) and perform a comprehensive analysis of analog RSM hardware including the basic device characteristics, hardware algorithms, and the corresponding mapping methods for device arrays, as well as the architecture and circuit design considerations for neural networks. Second, we classify the FoMs of analog RSM devices into two levels. Level 1 FoMs are essential for achieving the functionality of a system (e.g., linearity, symmetry, dynamic range, level numbers, fluctuation, variability, and yield). Level 2 FoMs are those that make a functional system more efficient and reliable (e.g., area, operational voltage, energy consumption, speed, endurance, retention, and compatibility with back-end-of-line processing). By constructing a device-to-application simulation framework, we perform an in-depth analysis of how these FoMs influence in-memory learning and give a target list of the device requirements. Lastly, we evaluate the main FoMs of most existing devices with analog characteristics and review optimization methods from programming schemes to materials and device structures. The key challenges and prospects from the device to system level for analog RSM devices are discussed. Bin Gao 0006, Jianshi Tang, Meng-Fan Chang, Xiaobo Sharon Hu, Jan Van der Spiegel, He Qian, Huaqiang Wu |
Proc. IEEE | 8 |
| 2021 | Diagonal Matrix Regression Layer: Training Neural Networks on Resistive Crossbars With Interconnect Resistance EffectabstractResistive crossbars implement parallel vector-matrix multiplication (VMM) in analog fashion, and thus enable fast and energy-efficient neuromorphic systems. However, interconnect resistance and resistive switching devices form a complex resistance network with sneak paths. It could result in severe distortions on the output currents. When implementing neural networks, current distortions also cause significant accuracy loss. This article proposes an accurate and computationally efficient model of VMM in resistive crossbars, called diagonal matrix regression (DMR), and incorporates the model into the topology of neural networks as DMR layer (DMRL). Given an m×n crossbar, two diagonal matrices are calculated directly according to the resistance network in a time complexity of only O(m2+n2). No hyper-parameter needs to be determined manually. Modeling of VMM is implemented in a time complexity of only O(mn). DMRL is developed to replace the weight matrix of neural networks so that the effect of interconnect resistance and the sneak path problem are well handled during ex-situ training. Using this technique, for the task of MNIST and fashion-MNIST classification, the accuracy is dramatically restored. Yan Liao, Bin Gao 0006, Jianshi Tang, Huaqiang Wu, He Qian |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2019 | Design Guidelines of RRAM based Neural-Processing-Unit: A Joint Device-Circuit-Algorithm AnalysisabstractRRAM based neural-processing-unit (NPU) is emerging for processing general purpose machine intelligence algorithms with ultra-high energy efficiency, while the imperfections of the analog devices and cross-point arrays make the practical application more complicated. In order to improve accuracy and robustness of the NPU, device-circuit-algorithm codesign with consideration of underlying device and array characteristics should outperform the optimization of individual device or algorithm. In this work, we provide a joint device-circuit-algorithm analysis and propose the corresponding design guidelines. Key innovations include: 1) An end-to-end simulator for RRAM NPU is developed with an integrated framework from device to algorithm. 2) The complete design of circuit and architecture for RRAM NPU is provided to make the analysis much close to the real prototype. 3) A large-scale neural network as well as other general-purpose networks are processed for the study of device-circuit interaction. 4) Accuracy loss from non-idealities of RRAM, such as I-V nonlinearity, noises of analog resistance levels, voltage-drop for interconnect, ADC/DAC precision, are evaluated for the NPU design. Xiaochen Peng, Huaqiang Wu, Bin Gao 0006, Hu He 0001, Youhui Zhang, Shimeng Yu, He Qian |
DAC | 8 |
| 2019 | Three-Dimensional nand Flash for Vector-Matrix MultiplicationabstractThree-Dimensional NAND flash technology is one of the most competitive integrated solutions for the high-volume massive data storage. So far, there are few investigations on how to use 3-D NAND flash for in-memory computing in the neural network accelerator. In this brief, we propose using the 3-D vertical channel NAND array architecture to implement the vector-matrix multiplication (VMM) with for the first time. Based on the array-level SPICE simulation, the bias condition including the selector layer and the unselected layers is optimized to achieve high computation accuracy of VMM. Since the VMM can be performed layer by layer in a 3-D NAND array, the read-out latency is largely improved compared to the conventional single-cell read-out operation. The impact of device-to-device variation on the computation accuracy is also analyzed. Panni Wang, Bo Wang 0067, Bin Gao 0006, Huaqiang Wu, He Qian, Shimeng Yu |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2018 | Sign backpropagation: An on-chip learning algorithm for analog RRAM neuromorphic computing systemsabstractCurrently, powerful deep learning models usually require significant resources in the form of processors and memory, which leads to very high energy consumption. The emerging resistive random access memory (RRAM) has shown great potential for constructing a scalable and energy-efficient neural network. However, it is hard to port a high-precision neural network from conventional digital CMOS hardware systems to analog RRAM systems owing to the variability of RRAM devices. A suitable on-chip learning algorithm should be developed to retrain or improve the performance of the neural network. In addition, determining how to integrate the periphery digital computations and analog RRAM crossbar is still a challenge. Here, we propose an on-chip learning algorithm, named sign backpropagation (SBP), for RRAM-based multilayer perceptron (MLP) with binary interfaces (0, 1) in forward process and 2-bit (±1, 0) in backward process. The simulation results show that the proposed method and architecture can achieve a comparable classification accuracy with MLP on MNIST dataset, meanwhile it can save area and energy cost by the calculation and storing of the intermediate results and take advantages of the RRAM crossbar potential in neuromorphic computing. Qingtian Zhang, Huaqiang Wu, Bin Gao 0006, Ning Deng 0008, He Qian |
Neural Networks | 7 |
| 2017 | Neuromorphic Computing based on Resistive RAMabstractResistive random access memory (RRAM) has gained significant attentions because of its excellent characteristics which are suitable for next-generation non-volatile memory applications. It is also very attractive to build neuromorphic computing chip based on RRAM cells due to non-volatile and analog properties. Neuromorphic computing hardware technologies using analog weight storage allow the scaling-up of the system size to complete cognitive tasks such as face classification much faster while consuming much lower energy. In this paper, RRAM technology development from material selection to device structure, from small array to full chip will be discussed in detail. Neuromorphic computing using RRAM devices is demonstrated, and speed & energy consumption are compared with Xeon Phi processor. Huaqiang Wu, Bin Gao 0006, He Qian |
ACM Great Lakes Symposium on VLSI | 6 |
| 2017 | A micro-array bio detection system based on a GMR sensor with 50-ppm sensitivity
Lei Zhang 0033, Jinwen Geng, Xizeng Shi, He Qian |
Sci. China Inf. Sci. | 5 |
| 2017 | Resistive Random Access Memory for Future Information Processing SystemabstractResistive random access memory (RRAM) is regarded as one of the most promising emerging memory technologies for next-generation embedded, standalone nonvolatile memory (NVM), and storage class memory (SCM) due to its speed, density, cost, and scalability. Considerable progress has been made in recent years on the manufacturability of RRAM, with low-density RRAM products now in production and the path to higher density parts becoming clearer. This review updates the learning on the fundamental materials and process integration needed for high-volume manufacturing and summarizes very recent progress on array level performance improvement methodology using novel techniques, and circuit level contributions for different applications. The device performance, array integration, and device/circuit codesign for memory systems are discussed. Novel applications besides embedded memory and standalone memory are addressed, including hardware security, neuromorphic computing, and nonvolatile logic systems. Huaqiang Wu, Xiao Hu Wang, Bin Gao 0006, Ning Deng 0008, Zhichao Lu, Brent Haukness, Gary Bronner, He Qian |
Proc. IEEE | 8 |
| 2014 | Stack engineering for ReRAM devices performance improvementabstractAl/W:AlOx/WOy/W and Pt/AlOδ/Ta2O5-x/TaOy/Pt multiple layers ReRAM devices have been fabricated and carefully studied. Experimental results exhibit significant performance improvement through the insertion of AlOxlayer between the switching layer and the top electrode. Operation current is remarkably reduced, ON/OFF ratio is greatly increased, and stable multi-level operations have been successfully achieved. Multiple layers stack engineering has been proved as an efficient method to improve the performances of ReRAM devices. Huaqiang Wu, Minghao Wu, Zhiping Yu, He Qian |
ISCAS | 7 |
| 2011 | Understanding dynamic behavior of mm-wave CML divider with injection-locking conceptabstractAn analytical framework has been developed to describe the locking behavior of millimeter (mm) wave current mode logic (CML) frequency divider. Unlike traditional analysis based on RC delay, the proposed model is established with injection-locking concept from analog perspective. Both analytical formulas and graphic interpretation are provided for design insights. The model has been validated by exhaustive simulations and important guidelines have been concluded for circuit design. Dajie Zeng, Li Zhang 0046, Lei Zhang 0033, Yan Wang 0023, He Qian, Zhiping Yu |
ISCAS | 7 |
| 2005 | Growth and characterization of 0.8-µm gate length AlGaN/GaN HEMTs on sapphire substrates
Cuimei Wang, Guoxin Hu, Junxue Ran, Cebao Fang, Yiping Zeng, Jinmin Li, He Qian |
Sci. China Ser. F Inf. Sci. | 11 |